
Letters
43 letters, page 1 of 5
The M5 Ultra Mac Studio is a much better machine than the 128 GB box I run everything on, the 512 GB tier is tempting for the biggest local models, and it would be a dumb financial decision. So I built an interactive yes/no decision tree: pick the chip, memory, machines and the model you want, and it lands on buy at launch, wait, avoid, or a CUDA box instead. I walk my own answers through it. It says wait.
Local AI · Money
Every Unsloth quant of GLM-5.3-Flash that fits a 128 GB Strix Halo desktop, 93 to 120 GB, through the same four one-shot tests at reasoning low, high and max. Five files fit, speed is flat at 8 to 10 tokens a second, only the smallest file finishes the exam, high finishes one test in twenty, and max loops on one word for three hours. Every output and number on the results page.
Local AI · Benchmarks
Can a 128 GB desktop run GLM-5.3-Flash, Z.ai's 320B open-weights model? The smallest Unsloth quant (93 GB) on a branch build of llama.cpp: one crash in Unsloth Desktop, then 9 tokens a second, four one-shot tests finished twice, vision working at 1-bit, and the one bug a harness would have fixed. Every output and number on the results page.
Local AI · Benchmarks
- How Claude Code makes my YouTube videos2026-09-05
The channel at 90 days: 76,000 views, 462 subscribers, and the workflow that gets every video out. Ten stages, one folder per video, the Claude Code skills, the one post-processing command, what is still friction, and the gear on the desk.
Workflow
OpenAI released GPT-6 Astra. I read the full announcement on camera: the 98% FrontierMath, 99.9% ARC-AGI-3 and 100% ExploitBench claims, the Hugging Face honeypot eval (48% to 0%), the computer-use demos, DeepSWE and Artificial Analysis, and Matthew Berman's early-access review. A first impression, and why the hype still points back to local AI.
AI news
A first impression of Claude Fable 5.1 and Mythos 5.1: what changed, who gets which model, the pricing read off the docs, and then two days of Claude Code in auto mode on my own local coding-agent benchmark. $165 on the estimator, a harder suite, and one honest caveat.
AI news · Benchmarks
OpenAI is ending its partnership with Cursor after SpaceX's $60 billion acquisition, with a proposed shutoff of November 12. What OpenAI actually said, the Musk history it points at, the disputed 5% number, what still works after the cutoff, and why the answer is owning your tools.
AI news
Nvidia has reportedly agreed to buy Hugging Face for $12.9 billion, about 86 times its self-reported revenue. Reported, not signed. The numbers, the offer Hugging Face turned down nine months ago, the open-weights letter both companies signed, and why a chip company owning the open-weights hub is not the worst outcome for people who run models at home.
AI news
Qwen's new 125B mixture-of-experts (6B active) on a 128 GB mini PC, the night it became runnable: llama.cpp built from the support PR, four Unsloth quants from 1-bit to IQ4_XS through the same four one-shot tests, my Qwen 3.8-27B as the reference. Every quant ran at about 20 tokens per second. Then the honest question: would I actually use it?
Local AI · Benchmarks
Qwen's new 125B mixture-of-experts model with 6B active parameters dropped today. Before running a single benchmark I read the Unsloth guide and the Qwen model card on camera to answer one question: can a 128 GB unified-memory mini PC even hold it? The smallest quant is 75 GB, they recommend 96, and BF16 is 355 GB.
Local AI
Get the letters
Every post goes out free on Substack. What I build, break, and fix, written up in plain English. No spam, unsubscribe anytime.
Subscribe on Substack