Study 001 — running

9 done · 1 running · 1 queued · 4 failed · conclusions not final

Open ledger →

Independent research replication

We rerun the experiments.

NULSPEC independently replicates published AI research on hardware we control—and publishes the run ledger in public before we know how it ends.

No pay-to-confirm. No hidden reruns. No success-only drawer.

What this is

Independent replication, done by people curious enough to check.

NULSPEC is a small team of enthusiasts and accelerationrationalists—a fused word, on purpose. We want the field to move fast, and we think the fastest route runs through checking the work.

We reproduce recent papers on our own machines, freeze the protocol before the first run, and publish the ledger whether or not the result cooperates.

Our operating protocol

The artifact is the argument.

A paper is not a vibe. A replication should leave enough evidence for a stranger to disagree productively.

  1. 01

    Freeze the specification

    The protocol, comparison rules, and exclusions enter Git before the first full-matrix run.

  2. 02

    Reproduce before extending

    Released code and manuscript-faithful interpretations stay separate. New ideas cannot rewrite the primary result.

  3. 03

    Number every deviation

    Hardware, stack, and implementation substitutions get an ID, a reason, and an impact control.

  4. 04

    Publish the miss

    Failures, null results, and irreproducible recipes receive the same artifact trail as a match.

  5. 05

    Make rerunning cheaper

    Commands, digests, checkpoints, and analysis code are preserved so the next person starts ahead of us.

Now running · Study 001

Towards Robust Reinforcement Learning for Small-Scale Language Model Agents

We are testing the complete 15-configuration released-code matrix before opening claim-level analysis. Operational status is public; partial conclusions are deliberately withheld.

Live run ledger

15 frozen configurations

9 done · 1 running · 1 queued · 4 failed
as of Jul 31, 2026, 1:29 AM UTC

Arm 1: Pythia 70M, TinyStories, DoneArm 2: Pythia 70M, CNN / DailyMail, DoneArm 3: Pythia 70M, WikiText, DoneArm 4: Pythia 160M, TinyStories, DoneArm 5: Pythia 160M, CNN / DailyMail, DoneArm 6: Pythia 160M, WikiText, DoneArm 7: Pythia 410M, TinyStories, DoneArm 8: Pythia 410M, CNN / DailyMail, DoneArm 9: Pythia 410M, WikiText, DoneArm 10: SmolLM2 135M, TinyStories, FailedArm 11: SmolLM2 135M, CNN / DailyMail, RunningArm 12: SmolLM2 135M, WikiText, QueuedArm 13: SmolLM2 360M, TinyStories, FailedArm 14: SmolLM2 360M, CNN / DailyMail, FailedArm 15: SmolLM2 360M, WikiText, Failed
Track R execution ledger. DONE means only that a run finished; it is not a replication verdict.
ArmStateConfigurationGPUProvenanceVerdict
001Done. Run finished; no study verdict implied.Pythia 70MTinyStoriesRTX 4090MonkeyPCEXACT
002Done. Run finished; no study verdict implied.Pythia 70MCNN / DailyMailRTX 3090wtatum84EXACT
003Done. Run finished; no study verdict implied.Pythia 70MWikiTextRTX 4090MonkeyPCEXACT
004Done. Run finished; no study verdict implied.Pythia 160MTinyStoriesRTX 4090MonkeyPCEXACT
005Done. Run finished; no study verdict implied.Pythia 160MCNN / DailyMailRTX 4090MonkeyPCEXACT
006Done. Run finished; no study verdict implied.Pythia 160MWikiTextRTX 3090wtatum84EXACT
007Done. Run finished; no study verdict implied.Pythia 410MTinyStoriesRTX 3090wtatum84EXACT
008Done. Run finished; no study verdict implied.Pythia 410MCNN / DailyMailRTX 4090MonkeyPCEXACT
009Done. Run finished; no study verdict implied.Pythia 410MWikiTextRTX 3090wtatum84EXACT
010Failed. Run ended without a valid completion.SmolLM2 135MTinyStoriesRTX 4090MonkeyPCEXACT
011Running. Actively executing; no result implied.SmolLM2 135MCNN / DailyMailRTX 3090wtatum84EXACT
012Queued. Assigned, not yet started.SmolLM2 135MWikiTextRTX 4090MonkeyPCEXACT
013Failed. Run ended without a valid completion.SmolLM2 360MTinyStoriesRTX PRO 6000wtatum84COMPAT1
014Failed. Run ended without a valid completion.SmolLM2 360MCNN / DailyMailRTX PRO 6000wtatum84COMPAT1
015Failed. Run ended without a valid completion.SmolLM2 360MWikiTextRTX PRO 6000wtatum84COMPAT1

1 COMPAT marks RTX PRO 6000 Blackwell arms. The paper-pinned PyTorch build cannot target sm_120; the substitution and required exact-stack re-evaluation are recorded as D-001.

State is operational. Verdict remains blank until the frozen 15-configuration family is complete and the analysis gate opens.

A null result is a result.

A deviation hidden is a claim faked.

If you cannot rerun it, you are reading marketing.

Put a claim on the bench

Seen a result you want tested?

Nominate it. We choose papers we can honestly attempt on local compute and a fixed budget. If we take yours on, the protocol goes public before the first arm launches—and so does every deviation we are forced to make.

Nominate a paper

Keep the apparatus alive

Support buys compute, not conclusions.

Donations go to GPU-hours, storage, and time. They cannot touch a verdict: protocols and decision rules are frozen before analysis. If you want more papers checked, faster, buy the lab monkey a little more runway.

Fund GPU-hours on Ko-fi