NanoTLM: A Fast Laboratory for AI-Driven Architecture Research
The goal of NanoTLM is to make architecture research fast: turn a model idea into a controlled comparison with a traditional baseline, then use the result to drive the next AI-assisted optimization.
Its primary output is not a single language model. It is a research loop: propose an architecture, validate its real graph and budget, train it under a frozen contract, compare it with the current champion and a traditional model, then preserve the result for the next experiment.
The current laboratory focuses on thermodynamic language models whose central computation is expressed as sparse energy graphs over probabilistic bits, or pbits. The traditional reference is a parameter-matched Transformer built from the core of nanochat. Both see the same contexts, targets, training exposure, seeds, and evaluation examples.
This makes each result immediately useful. A new thermodynamic architecture is not evaluated in isolation. It is compared with the previous thermodynamic champion and with nanochat under one common contract. If it improves, it earns more expensive replication. If it does not, the laboratory rejects it early and records why.
Autoresearch closes the loop. AI can read the evidence, propose a falsifiable optimization, create a new architecture profile, and launch the next experiment. The laboratory—not the AI—owns the dataset, comparator, baselines, metrics, and promotion rules.
Autoresearch loop
AI changes the architecture. The laboratory keeps the comparison fixed.
Each result becomes evidence for the next hypothesis, while the dataset, budget, baselines, metrics, and promotion rules stay outside AI control.
- 01 Hypothesis
AI proposes one falsifiable architectural change.
- 02 Architecture
A TOML profile describes the candidate and its limits.
- 03 Cheap gates
Static validation, smoke run, then one full seed.
- 04 Compare
The same contract scores the champion and nanochat.
- 05 Decide
Reject, or replicate and pass physical gates to promote.
Project status — August 16, 2026. “Fast” refers to research iteration: cheap validation, staged training, and early rejection of weak candidates. It is not a claim about inference speed or physical hardware. The benchmark numbers are a frozen snapshot of the phase 13 champion and the rejected phase 14 candidate.
Fast Research Needs A Small Contract
The fastest experiment is useless if its result cannot be compared with the previous one.
“Can thermodynamic computers run language models?” hides several different problems. Can an energy-based model learn a next-character distribution? Can that distribution be normalized? Can a sampler reproduce it? Is an improvement architectural, or did the experiment simply spend more parameters or see more data? If AI searches the design space, how do we stop it from optimizing the measurement instead of the model?
NanoTLM makes the task deliberately small and the comparison deliberately strict.
The task is character-level prediction on Tiny Shakespeare. There are 65 characters, a context of 16 characters, and one next-character target. Each model trains for 5,000 updates with 64 targets per update, for 320,000 target exposures per seed. Training is repeated with seeds 0, 1, and 2. Evaluation uses 512 examples selected by a separate frozen seed.
The traditional control is not the complete nanochat system. There is no BPE tokenizer, FineWeb corpus, or post-training pipeline. It is the GPT core adapted to the same character task and constrained to nearly the same parameter budget.
| Contract field | Value |
|---|---|
| Vocabulary | 65 characters |
| Context | 16 characters |
| Training seeds | 0, 1, 2 |
| Updates per seed | 5,000 |
| Target exposures per seed | 320,000 |
| Frozen evaluation examples | 512 |
| NanoTLM parameters | 9,790 |
| nanochat parameters | 9,824 |
| Primary metric | Bits per character |
This benchmark is tiny by modern language-model standards. That is a feature. At this scale I can afford exact probabilities, multiple seeds, independent samplers, and enough failed experiments to learn something. A larger model would produce a more impressive bill before it produced a more trustworthy answer.
Speed comes from staging the cost. Every candidate first faces static graph validation and a smoke run. Only a candidate that survives those gates receives a complete seed-0 run. Seeds 1 and 2, physical sampling gates, and reserved confirmation are paid for only when the cheaper evidence says the architecture is still promising.
The Unit Of Work Is An Architecture
Every experiment begins as a declarative TOML profile under research/dtm/architectures. A profile names one hypothesis, selects a model family, changes its configuration, and declares the budget it expects to satisfy.
schema_version = 1
name = "dtm-my-hypothesis-v1"
family = "dtm"
status = "draft"
description = "Tests one concrete architectural change."
mutable = true
[config]
n_latent = 104
[expected]
max_params = 9790
Before training begins, the laboratory instantiates the real graph and counts its pbits, edges, stages, and trainable parameters. A profile that exceeds its declared budget fails without consuming a competitive run. A valid run writes a manifest next to its artifact, so the architecture that was proposed and the architecture that was measured cannot silently diverge.
Architectures do not need to imitate the current DTM implementation. Any thermodynamic or traditional backend can enter the laboratory by emitting the backend-neutral LabRecord schema. That common record is what lets the same comparator rank different architectures by seed, exposure, budget, and predictive quality.
The Model Is an Energy Landscape
The current champion inside this laboratory is a two-stage denoising thermodynamic model, or DTM.
The 16-character context becomes 161 clamped pbits. A noisy seven-bit character code enters the first energy-based model. Two stages progressively transform that code into a distribution over the next character. The stages share sparse interaction weights, while their output biases may differ. A pool of 104 latent pbits gives the system capacity to represent dependencies that are not directly visible in the output bits.
16-character context ──> 161 clamped pbits
│
7-bit noisy code ─────> noisy EBM ─────> clean EBM ─────> 128 codes
↕ ↕ │
104 latent pbits 65 characters
For context , output , and latent state , each stage defines a conditional distribution through an energy function:
Low-energy configurations receive more probability. Sampling is part of the computation, not noise added after a conventional network has already produced its logits.
Seven output bits create 128 physical codes, but the vocabulary contains only 65 characters. That mismatch could be hidden by clipping the output to the vocabulary and renormalizing. NanoTLM reports it instead. The current champion assigns a median of roughly 0.021% probability to invalid codes.
The small output is also what makes the model unusually inspectable. All 128 codes can be enumerated. The latent pbits can be marginalized analytically for each code. The exact next-character distribution becomes an oracle against which the samplers can be tested.
This is the core design tradeoff of NanoTLM. The model is too small to make a useful product, but small enough that its claims do not have to depend on faith.
Comparing Architecture Generations
The first architecture in NanoTLM was a sparse conditional restricted Boltzmann machine. It learned above chance, normalized correctly, and could be sampled through Gibbs updates. It also scored 4.9386 BPC after adding a learning-rate decay schedule, far behind nanochat at 3.1905.
That model was useful because it established the plumbing. A checkpoint could contain everything required for evaluation. Gibbs samples could be compared with the exact oracle. Training and evaluation contexts were reproducible. The model and the control could be given the same data exposure.
The architecture then moved from one monolithic RBM to a chain of denoising energy models. Early DTM work reached 4.4999 BPC. Later experiments changed how context was represented rather than simply making the model larger.
The current champion spends dense bits on the three most recent characters and adds 49 binary products describing interactions between the final two. Older context remains sparse. This change brought the median to 3.4026 BPC with 34 fewer parameters than the control.
87.9% of the original gap closed
- Constant RBM 5.342
- RBM + decay 4.939
- Initial DTM 4.703
- DTM phase 10 4.500
- Recent cross 3.537
- DTM phase 13 3.403
nanochat baseline 3.1905
From the 4.9386 BPC RBM to the 3.4026 BPC DTM, NanoTLM improved by 1.5360 bits per character. That closes approximately 87.9% of the original distance between the RBM and nanochat.
Closing most of a gap is not the same as closing it. The remaining 0.2 BPC is the entire research problem now.
The important artifact is not only the last point on the chart. It is a comparable sequence of architectures: constant-rate RBM, decayed RBM, initial DTM, phase 10 DTM, recent-interaction representation, and phase 13 champion. Each transition can be evaluated against the same reference rather than remembered as an isolated anecdote.
Autoresearch Closes The Loop
NanoTLM’s autoresearch controller turns the architecture workflow into a machine-readable loop. AI reads the current champion, rejected hypotheses, and active protocol. It forms one falsifiable hypothesis, creates a candidate profile, runs the laboratory, and consumes a structured promote, parity, or reject result before proposing the next optimization.
read champion + rejected hypotheses + active protocol
form one falsifiable architectural hypothesis
write one candidate profile
run the frozen laboratory
read promote · parity · reject
repeat
The AI can choose the hypothesis and configuration inside a bounded sandbox. It cannot edit the dataset, evaluator, nanochat baseline, seed set, promotion margin, or previously recorded evidence.
The distinction is simple: the AI is the researcher, not the judge.
Every candidate moves through the same funnel:
- Static validation builds the real graph and counts pbits, edges, stages, and parameters.
- A smoke run catches compilation, memory, and numerical failures without entering the leaderboard.
- Seed 0 runs under the full training contract and must clear a predeclared improvement margin.
- Seeds 1 and 2 run only if seed 0 passes. Promotion normally requires improvement in all three.
- Exactness, THRML, and Torx gates test normalization, sampler error, valid outputs, and graph budgets.
- A fresh confirmation batch is opened once, after every earlier gate passes.
The controller fingerprints the protocol, source, model profile, baselines, and evidence with SHA-256. It refuses changed profiles, retries of completed candidates, and overwritten artifacts. A metric failure is recorded as a result, not treated as an invitation to quietly reroll the seed.
This sounds strict for a small model. It became necessary the moment AI entered the optimization loop.
An optimizer will exploit whatever it controls. If the same system can change the architecture and reinterpret the benchmark, “autonomous research” can become an elaborate way to confirm its own ideas. NanoTLM keeps the creative surface flexible and the measurement surface boring.
Fast Failure Is Part Of The Throughput
The laboratory is fastest when a wrong idea fails at the cheapest reliable stage and stays failed. The most recent experiment shows what that means.
Residual analysis suggested that the remaining NanoTLM error had structure around line breaks. A new candidate, line-cross-v1, explicitly crossed the distance from the last newline with the bits of the most recent character. Its budget remained below nanochat at 9,810 parameters.
On seed 0, it improved by 0.0173 BPC. The gate required 0.012, so the result was good enough to continue.
Then replication failed.
| Seed | Champion BPC | Candidate BPC | Candidate improvement |
|---|---|---|---|
| 0 | 3.4038 | 3.3865 | +0.0173 |
| 1 | 3.4026 | 3.4121 | −0.0095 |
| 2 | 3.3279 | 3.3695 | −0.0416 |
The apparent win survived in one of three seeds. Physical gates were not run, and the reserved confirmation set remained unopened. The candidate was rejected. The phase 13 model stayed champion.
There is a tempting alternative story where seed 0 becomes a promising preliminary result and the other two seeds become implementation details. NanoTLM does not permit that story. The replication rule was written before the run, so the result has only one valid interpretation: the hypothesis did not produce a reliable improvement.
Nothing was wasted. The experiment retired three versions of the line-structure idea and narrowed the next search. A laboratory should reduce uncertainty even when it does not increase a score.
Sampling Must Agree With Something
Predictive quality is only half of a thermodynamic model.
NanoTLM can calculate its output distribution exactly, but the intended computation is probabilistic. The project therefore samples the same model through two independent paths: Gibbs sampling with THRML and a local probabilistic dataflow graph in Torx.
Both are compared with the exact oracle using total variation distance. For the current champion, mean THRML error ranges from 0.0369 to 0.0575 across seeds. Torx ranges from 0.0316 to 0.0339. Normalization, explicit two-stage composition, valid generation, and sampler gates pass for all three seeds.
That validates the software representations against a small exact model. It does not validate physical hardware.
No NanoTLM experiment has run on a Z1 system or any other thermodynamic accelerator. The repository contains no energy measurement and no comparable latency result. Torx parity means the graph was expressed and sampled consistently in a local simulator. Turning that into a hardware claim would require an entirely different experiment.
This boundary is easy to state and important to preserve. “Compatible with a probabilistic graph” and “advantageous on a physical machine” are not synonyms.
The Current Scoreboard
The scoreboard is not the product. It is the common reference that tells every new architecture where it stands.
+0.2056 BPC paired gap nanochat wins 3/3 seeds
| Model | Parameters | BPC by seed | Median BPC | Median accuracy |
|---|---|---|---|---|
| NanoTLM DTM phase 13 | 9,790 | 3.4037 / 3.4026 / 3.3279 | 3.4026 | 0.3730 |
| nanochat-core matched | 9,824 | 3.1609 / 3.1970 / 3.1905 | 3.1905 | 0.3809 |
The median paired gap is 0.2056 BPC, and nanochat wins all three seeds. NanoTLM has not demonstrated predictive parity, sustained linguistic coherence, or a compute advantage.
It has demonstrated something narrower:
- a language model whose central operation is a sparse energy graph;
- exact normalized next-character probabilities over its physical output codes;
- local agreement among exact evaluation, THRML, and Torx;
- large reproducible gains across several architectural generations;
- a matched control that remains better;
- promotion rules capable of rejecting an attractive result.
I am more interested in this list than I would be in a fragile win produced by a weaker comparison.
Running The Laboratory
NanoTLM requires Python 3.12 and uv. The complete test suite runs on CPU:
git clone https://github.com/Luisgarcav/nanotlm
cd nanotlm
uv sync --group dev
JAX_PLATFORMS=cpu uv run pytest -q
The champion architecture and its budget can be inspected without training:
JAX_PLATFORMS=cpu uv run python -m research.dtm.lab show \
dtm-phase13-champion
After copying the sandbox profile and changing one hypothesis, the AI-driven loop can register and run it:
uv run python -m research.autoresearch_lab init
uv run python -m research.autoresearch_lab register \
--profile research/dtm/architectures/dtm-my-hypothesis-v1.toml \
--hypothesis "The proposed mechanism should lower paired BPC because ..."
uv run python -m research.autoresearch_lab run dtm-my-hypothesis-v1
uv run python -m research.autoresearch_lab status
If replication passes, the same controller advances the candidate to independently owned physical gates:
uv run python -m research.autoresearch_lab physical dtm-my-hypothesis-v1
New experiments are TOML profiles. A candidate can change one hypothesis, validate the resulting graph, and run under the same frozen contract. The repository README contains the complete quick start, while the phase 13 analysis and phase 14 negative result preserve the evidence behind the current state.
The values rendered in this article mirror NanoTLM’s source charts, which are generated from versioned JSON records. The repository’s chart script also has a check mode, so its source-of-truth visuals cannot silently drift away from the data they claim to show.
What NanoTLM Is For
NanoTLM is not built to protect one thermodynamic architecture. It is built to replace architectures quickly when the evidence points somewhere better.
Its job is to make a complete research cycle fast, comparable, and automatable:
- express a new architecture as a small, reviewable profile;
- reject invalid or over-budget graphs before expensive training;
- compare every surviving candidate with its champion and with nanochat;
- replicate improvements across the same frozen seeds;
- validate thermodynamic sampling against an exact oracle;
- return structured evidence that AI can use for the next optimization.
The current DTM champion reaches 3.4026 BPC while nanochat reaches 3.1905. That gap is not the thesis of the project. It is the present state of the laboratory—the reference from which the next AI-proposed architecture must improve.
I do not know whether this model family will catch the Transformer under the current budget, or whether another energy-based architecture will replace it first. NanoTLM exists so that either answer can be reached through short iterations and honest comparisons rather than isolated demonstrations.
The model should keep changing. The contract should survive. The faster the laboratory can turn an AI-generated hypothesis into comparable evidence, the more useful NanoTLM becomes.