Season 6 · Ch. 5

Shipping Something Boring

For a long time the app only existed on my laptop. Getting it off my laptop meant two boring problems and one thing I did not expect. Problem one: too heavy The database of FDA labels I had been building against was 169 megabytes, most of it raw label text I did not need at lookup time. A web demo does not want to drag that around. So I built a slimmed copy that keeps the drug names and the label text a user actually reads, and drops everything else. It came out to 77 megabytes. Same answers, less than half the weight. ...

June 1, 2026 · 7 min · Jun Park
Season 6 · Ch. 4

The Weekend I Tried to Break My Own App

Once the stack was built, I did the only responsible thing. I tried to destroy it. I spent a weekend typing the worst inputs I could imagine into my own search box. When I ran out of bad ideas, I handed a list of nearly three hundred nasty queries to two other AIs, Gemini and Claude, and asked them to be as adversarial as they could. Their whole assignment was to find the input that made my app hand someone the wrong drug. It is a strange way to spend a Saturday, paying for compute so one robot can try to make your other robot commit malpractice. ...

May 28, 2026 · 6 min · Jun Park
Season 6 · Ch. 3

Bigger Is Not Safer

The deterministic matcher was supposed to be the safe part. It mostly is. But safe does not mean finished, and the matcher had failure modes of its own. Most of the work in this project was not building it. It was tuning it. The dial The matcher has one main dial, but the dial only makes sense once you have watched the thing work. It is worth seeing once, because the whole safety argument rests on how dumb and how legible this machine is. ...

May 25, 2026 · 8 min · Jun Park
Season 6 · Ch. 2

Where the Deterministic Line Goes

The obvious fix is to let the model guess and check its work. The model reads zoloff and suggests a drug. Before you show that drug to anyone, you look it up in the database. If it is real, you trust it. If it is not, you throw it out. A guardrail on the output. Everyone reaches for this first. I did too. It did not work. Here is why. ...

May 18, 2026 · 5 min · Jun Park
Season 6 · Ch. 1

What I Was Trying to Build

What I actually wanted to build was bigger than this. In fact, much bigger. Here was the daydream. A doctor on a night shift, no signal in the hospital basement, pulls out her phone and asks it whether two drugs are safe together. It answers instantly and offline, with the FDA citation sitting right there. She squints. “Huh,” she says. “That is better than the thing we pay six hundred dollars a year for.” A pharmacist tells another pharmacist. Someone on a medical forum posts that it is the first one of these that does not just make things up. In my head I had already gotten the email from hospital procurement, and turned it down for being too generous. ...

May 15, 2026 · 7 min · Jun Park
Season 5 · Ch. 4

It Took Me Two Weeks to Read My Own Code

I drafted this chapter on May 14, the morning after the 3B benchmarks landed. The opening was a banger about “the inflection point at 3B parameters.” The conclusion explained how alignment tax shrinks with scale and flips positive between 2B and 3B. There was a chart. The chart had a smooth curve. I was very pleased with the chart. For two weeks I was going to publish that chapter. I told three people. I rehearsed a tweet thread. I picked which checkpoint to put in the OG image. ...

May 6, 2026 · 12 min · Jun Park
Season 5 · Ch. 3

Nothing Happened for 75,000 Steps and It Was Glorious

After Chapter 1 (three days of cloud chaos) and Chapter 2 (twelve hours of blaming the wrong thing), you have earned the right to expect another disaster chapter. I am sorry. There is no disaster here. The training worked. Here is the diary. Day What happened 1 Loss went down 2 Loss went down 3 Loss went down 4 Loss went down 5 Loss reached 2.2475. Run complete. That is the whole season, basically. We can stop now if you want. ...

April 19, 2026 · 6 min · Jun Park
Season 5 · Ch. 2

My Code Agent Said It Was a Moose. I Said No. It Was a Moose.

The H200 was working. The 3B was training. After three days of fighting the cloud, the model was finally putting tokens through the GPU at 23,200 per second. I had a checkpoint at step 1,000. I had a checkpoint at step 1,200. I went to bed feeling, briefly, like a person. Six hours later the run was dead. The checkpoint at step 1,200 was corrupted. The next run got to step 25 and froze. The one after that got to step 17 and silently disappeared. ...

April 12, 2026 · 10 min · Jun Park
Season 5 · Ch. 1

I Have an A100. I Have 528 Shards of Data. I Cannot Combine Them.

I had a 3B model expanded from the 2B-75K base. Code tested. Smoke test passed. 528 shards on my NAS, ~70 GB, ~38 billion tokens of FineWeb, FineMath, PubMed, and cleaned Python. Three days later I had spent zero training tokens and was 1,200 words deep into a Notion page about VRAM accounting. This is that story. Why a 3B Two reasons. One: I wanted the next model to know what a kinase is. The 2B was clean, polite, and had read a lot of FineWeb. It had also never seen a single PubMed abstract. I have plans for this model that involve answering biomedical questions, and you cannot retrieve your way out of a model that does not know what “phosphorylation” means. The 3B’s data plan added 256 shards of PubMed, ~5.5B tokens, all fresh. The 2B is a polite generalist. The 3B is a polite generalist who also took two semesters of biochemistry. ...

April 7, 2026 · 8 min · Jun Park
Season 4 · Ch. 4

Verbatim: The Proof Is in the Output

Benchmarks say the 1B and 2B are basically the same model. The outputs say otherwise. Here are the receipts - same 8 prompts, same temperature (0.7), same top-p (0.9), same max tokens (200). 1B-160K-Chat vs 2B-75K-Chat-DPO, head to head. Why the 1B’s Chat model and not its DPO version? Because DPO made the 1B worse - the best DPO run scored 4/8 garbage, worse than the Chat baseline. The Chat model is the 1B at its best. This is as fair as it gets. ...

March 30, 2026 · 11 min · Jun Park
GPUburnout
GPUburnout
Will Code for Tokens
S1 GPT-2 134M
S2 Llama 1B
S3 1B SFT
S4 Llama 2B
S5 Llama 3B