NanoTLM Update: Comparing Extropic's Z1T with nanochat
The latest experiment in NanoTLM compares the core of nanochat with a small model built from Extropic’s official Z1T implementation. Two practical questions drive the comparison:
- Which model learns better with the same resources? Give both a shared parameter cap and the same training examples, then compare what they learn.
- Which model delivers better quality for the same cost? Give both the same compute budget, measured here as active training time on the same GPU, then compare the quality they reach.
The first question asks how effectively each model learns from a fixed amount of data within a shared parameter limit. The second asks how much predictive quality each implementation can achieve within a fixed training budget. Equal data exposure does not imply equal training cost, so each question needs its own measurement.
Z1T had a small mean advantage in both comparisons, but the uncertainty prevents establishing consistent superiority. The result is no clear winner on either question, and no demonstrated equivalence.
A Brief Introduction to Z1T
Z1T is Extropic’s family of sparse, transformer-like models designed for its Z1 probabilistic hardware. Its building blocks include sparse tanh-linear operations and gated convolutional attention. The design studies how a neural architecture can fit the connectivity constraints of a different computing substrate.
Extropic’s article also describes inference split between Z1 and digital coprocessors, with energy and latency estimates for that arrangement. This experiment evaluates a different setting: digital training on a conventional GPU. It does not test those hardware estimates.
NanoTLM imports the official Z1T implementation, fixes its source revision, and trains a small character-level configuration from scratch. These are our own trained weights, not Extropic’s large pretrained checkpoint.
The other model uses the nanochat GPT core, also trained from scratch on the character task. This campaign evaluates its small predictive model; it does not reproduce nanochat’s complete tokenizer, pretraining corpus, or chat-training pipeline.
Two Questions, Two Budgets
The task is next-character prediction on Tiny Shakespeare. Both models receive a context of 16 characters, use the same vocabulary of 65 characters, and train with one supervised target per context. Parameters and computation use float32.
The common parameter cap is 9,824. Nanochat uses 9,824 trainable parameters; Z1T uses 9,684, or 1.43% fewer. This is a shared limit, not an exact parameter match.
| Comparison | What it answers | Fixed budget |
|---|---|---|
| Same parameter cap and training data (equal exposure) | Which learns better with those resources? | 5,000 updates × 64 targets: 320,000 target exposures |
| Same compute budget (equal active time) | Which delivers better quality for the same cost? | 15 seconds of active training on one RTX 4060 Laptop GPU |
In the first comparison, the time required to train can differ. In the second, the number of updates and target exposures completed can differ. The shared parameter cap applies to both.
The active timer includes generating and transferring batches, training, and synchronization. Compilation and warm-up happen first, after which model and optimizer state are reset. Checkpoint copying, saving, and evaluation stay outside that timer. The cutoff is the first synchronized boundary after 15 seconds, with a predeclared maximum overshoot of 0.25 seconds; every run records its actual time.
For the second question, cost means active GPU training time. This is a wall-clock implementation comparison: JAX and PyTorch are part of what is measured. It is not an equal-FLOP experiment, a measurement of electrical energy, or a monetary cost comparison.
Give Both Models a Chance to Improve
A comparison against one arbitrary training recipe would leave an obvious question: did the architecture lose, or did its optimizer settings need work?
I kept the architectures fixed and gave each family the same 12-recipe AdamW search:
- learning rates of 0.0003, 0.001, and 0.003;
- beta pairs of (0.9, 0.999) and (0.8, 0.95);
- weight decay of 0 or 0.01.
Each recipe ran with seeds 0, 1, and 2. That produced 72 tuning runs. The best mean development BPC was selected separately for each family and each budget. Equal numbers of trials do not guarantee equal total search compute; those costs are recorded separately.
The selected recipes were frozen before confirmation. Five new seeds—10 through 14—then trained from scratch. Nanochat selected the same recipe for both budgets. Z1T selected different weight decay for each budget, so confirmation required 15 additional training runs. Four earlier technical pilots checked the pipeline. All 91 runs completed successfully.
Both families selected a constant learning rate of 0.003 and betas of (0.9, 0.999). Nanochat selected no weight decay for either budget. Z1T selected weight decay 0.01 for equal exposure and zero for equal active time.
This search is deliberately bounded. It does not include architecture search or nanochat’s native Muon-based training recipe, and it does not establish the best possible implementation of either family.
Keep Confirmation Separate
The new training split excludes the final 100,000 characters of the original training file. That region is reserved for confirmation; the original validation file supplies development examples. No training or evaluation window crosses a split boundary.
There is a limitation to this reserve: it belongs to a known corpus and was used to train historical models in the project. It was excluded from this campaign’s training and recipe selection, but it is not external data that the entire research project has never encountered.
Development and confirmation each use 64 disjoint text blocks. Each block supplies 64 consecutive targets with 16-character contexts, for 4,096 evaluated targets. Their indices and hashes were fixed before scoring. An independent check also verified that both Python environments generate identical complete training streams for all eight training seeds.
The earlier NanoTLM article describes a different frozen benchmark, including the DTM architecture. Its scores should not be merged with this campaign’s results: the training region, evaluation sample, and selection procedure have changed.
What the Experiment Says About Each Question
The metric is bits per character, or BPC: the average negative log-probability assigned to the correct next character, measured in bits. Lower is better. Each mean below comes from five fresh training seeds evaluated on the same confirmation targets.
| Comparison | nanochat mean BPC | Z1T mean BPC |
|---|---|---|
| Same 320,000 target exposures | 3.1175 | 3.0954 |
| Same 15 seconds of active training | 3.1050 | 3.0882 |
Which Learns Better with the Same Resources?
With the same 320,000 target exposures and a shared parameter cap, Z1T has the lower mean BPC. It scores better in three of five seeds. That is a small observed advantage in learning from the allotted data, but it does not establish that Z1T reliably learns better under these constraints.
Which Delivers Better Quality for the Same Cost?
With the same 15 seconds of active GPU training, Z1T again has the lower mean BPC, scoring better in four of five seeds. This is the comparison that addresses quality for a fixed compute budget: each implementation can complete as many training updates as the time allows. The observed advantage still falls short of establishing a reliable winner on quality for the same active training cost.
How Certain Are These Answers?
The difference changes sign across seeds in both comparisons, so the averages do not describe a uniform advantage.
The uncertainty calculation resamples both training seeds and shared text blocks, keeping the comparison paired. It uses 20,000 bootstrap replicates and a 97.5% interval for each question, giving nominal 95% joint coverage through a Bonferroni adjustment. With five seeds and one corpus, these remain approximate intervals.
Define the difference as Z1T BPC minus nanochat BPC. Negative values favor Z1T; positive values favor nanochat.
| Comparison | Mean difference | 97.5% bootstrap interval |
|---|---|---|
| Equal exposure | −0.0221 BPC | −0.0652 to +0.0207 BPC |
| Equal active time | −0.0168 BPC | −0.0583 to +0.0290 BPC |
For equal exposure, the interval spans a Z1T advantage of about 0.065 BPC through a nanochat advantage of about 0.021 BPC. Both intervals include zero.
Before the runs, the protocol set a practical margin of ±0.01 BPC. A winner required an interval entirely beyond that margin in its favor and a consistent direction across all five seeds. Practical equivalence required the entire interval to fit inside the margin. Neither condition was met.
Calling the result a demonstrated tie would therefore go further than the evidence. The supported statement is narrower: close observed averages, with no conclusive winner and no demonstrated equivalence.
Training Time Has a Boundary
The 15-second comparison begins after preparation. During confirmation, preparation averaged about 3.9 seconds for nanochat and 8.0–8.2 seconds for Z1T. Those costs matter for short jobs, even though they are outside the active training budget.
A fixed total budget beginning at process launch would be another experiment. So would inference latency, throughput, or energy per generated character. The answer to which model delivers better quality for the same cost therefore depends on where the cost measurement starts and ends. Here it covers active training on this local machine; it does not establish an energy advantage for Z1T or physical Z1 hardware.
The complete processes for all 91 runs, including pilots, took about 39.82 minutes in aggregate. That includes preparation, training, checkpoint work, and evaluation, but excludes code editing, review, and the project test suite. The full report separates these costs by phase and family.
What Remains Open
The best learning rate for both families was the highest one in the tested grid. That does not prove that a higher rate will help, but it leaves a clear question for another campaign. Learning-rate schedules, native optimizers, and architecture choices also remain open.
A useful next study would declare a new search budget, extend both families’ recipes, and use renewed validation. Larger parameter budgets, longer contexts, and additional corpora would test whether the behavior persists beyond this small character task. Any extension should fix its stopping rule before collecting more results.
For this campaign, both questions remain without a conclusive winner: which model learns better with the same resources, and which delivers better quality for the same training cost? Z1T’s observed averages are encouraging enough to justify further investigation, but they do not establish superiority over nanochat on either question. The purpose of NanoTLM is to make both comparisons reproducible, with the budgets and uncertainty preserved alongside the scores.
Resources and Evidence
- Extropic: Z1T — Sparse Transformer-Like Models for Probabilistic Hardware.
- Official Z1T implementation, pinned revision.
- Nanochat repository.
- NanoTLM repository and README.
- Complete research notebook, including decisions and the run ledger.
- Frozen protocol, recipe selection, and machine-readable results.
- Run-level CSV and reproduction instructions.
The chart uses a local snapshot of the versioned NanoTLM results. The campaign’s verification records preserve matching training streams, checkpoint and source hashes, and an independent recalculation of the intervals. The project suite passed 371 tests on its documented CPU backend; the training comparisons ran on GPU.