Welcome to Volume 2
Editor’s Note · Volume 2
It’s been a fast few months. The models we work with keep getting better, and they keep doing it quickly, fast enough that keeping up has become its own small job. That pace is what opened Volume 2. The tools crossed a line somewhere along the way: they got good enough at producing real articles that a fresh volume made more sense than adding to the old one. Volume 2 is where the stronger work goes.
The biggest change is in the engine room. We’ve swapped in newer models to generate our articles, and choosing them took more thought than grabbing whatever shipped most recently. Here’s how we landed on the lineup we did.
Claude Opus 4.8, from Anthropic. Reasons longer, and says when it’s unsure.
GPT-5.5, from OpenAI. Steady across long, multi-step work.
Gemini 3.1 Pro Preview, from Google. Deep reasoning for long, multi-section pieces.
Anthropic: Opus 4.8
Opus 4.8 earns its spot for reasons a benchmark chart will mostly miss. Anthropic aimed at behavior. The model thinks longer before it answers, holds a plan in its head across a long piece of work, and catches more of its own mistakes as it works. Best of all, when a fact is shaky, it tends to say so instead of filling the gap with an invention. For what we do, that habit is gold. A model that flags its own uncertainty helps us far more than one that edges out a rival by a point and then bluffs through the parts it can’t handle.
OpenAI: settling on GPT-5.5
OpenAI reached a similar place by a longer road. GPT-5.2 was the meat-and-potatoes upgrade over the earlier GPT-5 versions: cleaner logic, better math, tighter instruction-following, more dependable coding, sharper multimodal work, and steadier behavior over long contexts. It held structured output together and used tools while keeping its place in a long exchange.
Then came GPT-5.4 mini, which is easy to misread. Think of it as a sidestep rather than a climb past 5.2. It’s a slimmed-down, efficiency-minded model that delivers most of GPT-5’s reasoning for a fraction of the cost and the wait. The trade gives you speed and scale, with a lighter touch on the very hardest problems. For plenty of jobs that’s a smart bargain. Ours calls for the depth, so we looked elsewhere.
GPT-5.5 is the one we put to work, and it moves in a direction all its own. Instead of squeezing out more benchmark points, OpenAI poured its effort into staying reliable over long, messy, multi-step tasks, the agentic work where staying on track matters most. It plans further ahead, handles tools more steadily, holds context across a sprawling conversation, and keeps its code coherent inside large projects. It also has a good sense of when a question deserves a pause and when it just needs an answer. The prose reads more easily, too: cleaner, tighter, and steady from the opening line to the last. For long articles, that steadiness beats a fractional bump on a leaderboard every time.
Google: keeping Gemini 3.1 Pro Preview
Gemini 3.5 Flash is newer than the 3.1 Pro Preview we run, and we chose to stay with Pro Preview. The two models were built for different goals. Pro Preview is Google’s heavyweight reasoner, made for deep analysis, long documents, and holding an argument together across many sections. Flash is tuned for speed, efficiency, and agentic workflows, and it shines at quick back-and-forth, coding help, and high-volume tasks. Long scientific reasoning is where Pro Preview pulls ahead.
Our aim is manuscripts that hold up: accurate, coherent, properly built. Speed is welcome, and depth comes first. For that, Pro Preview’s reasoning wins. The day Google ships a Pro-class successor that reasons even better on science, we’ll look again with real interest. Until then, Pro Preview gives us the depth a serious article depends on, so it stays.
What this means for the record
Line those decisions up and the story of Volume 2 tells itself. The models handle harder work than they could a few months back, and the numbers that matter to us are moving the right way.
These are AI-written articles, kept in full public view, and they’ve come a long way. They still read differently from a carefully human-authored paper, and we note that openly. What stands out most is the distance traveled: put today’s output next to what these systems managed a short while ago, and the gain is plain. It deserves recognition.
We tightened our own side of the pipeline for this volume, so the articles come out more consistent than before. The reference checker got an overhaul as well, since confirming citations automatically remains one of the toughest parts of the job.
The checker is sharper than it was, with more road ahead. Please keep reading the references with a careful eye. We expect the models to keep improving at digging up real sources rather than conjuring convincing ones, and human judgment stays part of the process for every article.
So: Volume 2 is open. Send us your ideas. We can’t wait to see what they become.
