# Max Ghenis - Full Content > This file contains the full text of all blog posts from maxghenis.com, formatted for LLM consumption. > Generated: 2026-09-04T17:10:02.732Z > See also: /llms.txt for a curated overview --- ## Table of Contents 1. [A year of ChatGPT is a fifth of a beer](/blog/drinking-ai/) 2. [MacKenzie Scott's giving, in QALYs](/blog/mackenzie-scott-giving-in-qalys/) 3. [The US government is gutting Anthropic's R&D capacity](/blog/us-gutted-anthropic-rd/) 4. [How to turn money into predictions](/blog/how-to-turn-money-into-predictions/) 5. [Can Talkie-1930 do arithmetic?](/blog/talkie-1930-math-evals/) 6. [Billionaires aren't 'just as likely' to be nonpayers as top taxpayers](/blog/madoff-billionaire-nonpayers/) 7. [I built a MyST-to-Quarto converter (and why you might need one)](/blog/mystquarto/) 8. [One year of Claude Code](/blog/my-claude-code-config/) 9. [Why you can''t make double eye contact](/blog/double-eye-contact/) 10. [Interactive replication of GiveWell''s cost-effectiveness analysis](/blog/givewell-cea/) 11. [OpenMessage: How I built a macOS Google Messages client to give Claude my texts](/blog/openmessage/) 12. [Where US immigrants come from, and where ICE focuses enforcement](/blog/bad-bunny-immigration-enforcement/) 13. [I used Claude Code to recertify for SNAP](/blog/snap-recertification-claude-code/) 14. [Scrollywood: Smooth scroll video recording for the web](/blog/scrollywood/) 15. [The adolescence of policy simulation: On Amodei and economic disruption](/blog/amodei-adolescence-policyengine/) 16. [opencollective-py: Manage OpenCollective from your terminal or Claude Code](/blog/opencollective-py/) 17. [From IDE to AI orchestration: The end of code-first development](/blog/ide-to-ai-orchestration/) 18. [TerminalGrid: Turn VS Code into a Claude Code superterminal](/blog/terminalgrid/) 19. [RAMBar: A macOS menu bar RAM monitor for developers](/blog/rambar/) 20. [How to Use MDX for Interactive Posts](/blog/using-mdx/) 21. [VAT thresholds, revenues, and the role of counterfactuals](/blog/vat-thresholds/) 22. [AI models favor Cuomo over Mamdani on NYC housing production](/blog/ai-models-nyc-housing/) 23. [My 2024 in code: Personal projects and AI tools](/blog/my-2024-in-code/) 24. [Why I'm a neoliberal](/blog/why-im-a-neoliberal/) 25. [In appointing their newest member, the Ventura City Council only pretended to care about housing](/blog/in-appointing-their-newest-member-the-ventura-city/) 26. [If you earned more in 2019 than 2018, don't file your 2019 taxes yet! Otherwise, file ASAP!](/blog/if-you-earned-more-in-2019-than-2018-dont-file-you/) 27. [Why I’ve taken the Giving What We Can pledge](/blog/why-ive-taken-the-giving-what-we-can-pledge/) 28. [Ventura County must end its unsheltered homelessness](/blog/ventura-county-must-end-its-unsheltered-homelessne/) 29. [Warren’s wealth tax would raise less than she claims — even using her economists’ own assumptions](/blog/warrens-wealth-tax-would-raise-less-than-she-claim/) 30. [Andrew Yang's troubling Tucker Carlson interview](/blog/andrew-yangs-troubling-tucker-carlson-interview/) 31. [Quantile regression, from linear models to trees to deep learning](/blog/quantile-regression-from-linear-models-to-trees-to/) 32. [We should replace the Child Tax Credit with a universal child benefit](/blog/we-should-replace-the-child-tax-credit-with-a-univ/) 33. [The case for a wonky basic income plan](/blog/the-case-for-a-wonky-basic-income-plan/) 34. [Reddit featured a misleading headline on education secretary nominee Betsy DeVos](/blog/reddit-featured-a-misleading-headline-on-education/) 35. [Uber's tipping settlement will reduce earnings of African-American drivers](/blog/ubers-tipping-settlement-will-reduce-earnings-of-a/) 36. [If we can afford our current welfare system, we can afford basic income](/blog/if-we-can-afford-our-current-welfare-system-we-can/) --- ## A year of ChatGPT is a fifth of a beer **URL:** https://maxghenis.com/blog/drinking-ai/ **Published:** Jul 05 2026 **Description:** I built an interactive that prices drinks in AI queries — and lets you flip the accounting scope and the workload unit that make published AI water numbers differ by 100x. Asking ChatGPT 30 questions a day for a year uses about a fifth of the water behind one beer. That counts both data-center cooling and the power plants generating the electricity — the closest AI analog to how the beer number counts the water behind the barley. I built [an interactive](https://maxghenis.com/drinking-ai) that prices drinks this way: pick a beer, a coffee, a glass of milk, and see its water footprint in ChatGPT queries. One beer ≈ 53,300 queries. One coffee ≈ 66,200. A glass of tap water ≈ 178. A year of 30 daily queries is 21.9 liters at the 2 mL scope; one beer is 106.5. The drinks give a scale; the reader picks the scope. ## Why every AI water number differs from every other one Published water-per-query figures span more than 100x: [Google measured 0.26 mL](https://cloud.google.com/blog/products/infrastructure/measuring-the-environmental-impact-of-ai-inference/) per median Gemini prompt, [Sam Altman cited 0.000085 gallons (0.32 mL)](https://blog.samaltman.com/the-gentle-singularity) for the average ChatGPT query, [a benchmarking preprint](https://arxiv.org/abs/2505.09598) models a short GPT-4o query near 2 mL, and [Mistral's life-cycle assessment](https://mistral.ai/news/our-contribution-to-a-global-environmental-standard-for-ai) reports 45 mL per 400-token response. These numbers barely disagree about physics. They disagree about where to draw the boundary: - **0.26–0.32 mL**: Google defines its figure as on-site data-center cooling water only; OpenAI states no boundary, and I treat it the same because it is within 25% of Google's. - **~2 mL** adds the water evaporated at the power plants generating the query's electricity. - **45 mL** is a life-cycle assessment — upstream electricity, cooling, and embodied hardware, no training — of one 400-token Le Chat response, a response about as short as the 2 mL query; the gap is boundary and data-center location, not length. The calculator defaults to the middle scope because it is the closest match to how the drink side is measured. Beer's footprint — 298 liters per kilogram, from [the standard crop-water tables](https://hess.copernicus.org/articles/15/1577/2011/hess-15-1577-2011.pdf) — counts the water behind the barley, not just the brewery's taps. The consistent AI analog counts the water behind the electricity, not just the cooling loop. Counting only cooling is like measuring beer by what the brewery pours in. You can flip the scope yourself. The comparison survives every setting: even at Mistral's 45 mL, one beer still equals more than 2,000 responses. ## Agents change the unit The published per-query figures describe one short text prompt. Reasoning models emit several times the tokens of a chat response, and agents chain many calls per task — [Anthropic measured its multi-agent systems at about 15x](https://www.anthropic.com/engineering/multi-agent-research-system) the tokens of a chat interaction. Since water scales with tokens, the calculator lets you count in reasoning responses (10x — I asked Claude for a mid-range estimate of the 2.5–50x published spread) or agentic tasks (15x). One beer = 3,550 agentic tasks at the default scope: 10 a day for a year. At Mistral's 45 mL it is about 160. ## The aggregate is the same picture [A Patterns paper](https://pmc.ncbi.nlm.nih.gov/articles/PMC12827721/) models all AI systems worldwide at 312.5–764.6 billion liters of water in 2025. The US drink categories I could source from government and industry data: | Category | Year | Water footprint (liters) | vs global AI | |---|---|---|---| | Coffee | 2024/25 | 24.9 trillion | 33–80x | | Milk | 2025 | 19.8 trillion | 26–63x | | Soda | 2025 | 15.3 trillion | 20–49x | | Beer | 2023 | 7 trillion | 9.2–22x | | Plant-based milk | 2025 | 0.4 trillion | 0.5–1.3x | | All nine tracked | — | 74.4 trillion | 97–238x | US coffee's footprint is 33–80x global AI's. Extrapolating AI to 2026 (I asked Claude to estimate growth: 1.5x, within the published 1.3–2.45x) puts the nine categories at 65–159x the projected range. ## What this doesn't say Water footprints measure volume, not scarcity. A liter evaporated from a stressed aquifer in Arizona is not a liter of rain on Bavarian barley, and the comparison doesn't pretend otherwise. The local version of the question — one campus drawing from one watershed, peaking in summer — is a siting and water-rights question this page does not address. ## Provenance Every figure on the page links to its source — the Mekonnen & Hoekstra crop and animal water tables, USDA and NIAAA volume data, and the per-query figures above. The three Claude-estimated numbers (the 2026 projection, the 10x reasoning multiplier, and the footprint blend for plant-based milk) are flagged as estimates on the page. The constants and derivations live in [a tested module](https://github.com/MaxGhenis/maxghenis.com/blob/master/src/lib/drinking-ai.ts) — a vitest suite pins the headline numbers, so if a source changes, the page fails loudly instead of drifting quietly. If you find a better source for any figure, tell me and I'll swap it in. [Andy Masley's post](https://blog.andymasley.com/p/individual-ai-use-is-not-bad-for) prompted the question. Drink responsibly. Prompt freely: [maxghenis.com/drinking-ai](https://maxghenis.com/drinking-ai) --- ## MacKenzie Scott's giving, in QALYs **URL:** https://maxghenis.com/blog/mackenzie-scott-giving-in-qalys/ **Published:** Jun 28 2026 **Description:** An interactive cost-effectiveness model for MacKenzie Scott's $26B in philanthropy — and the AI prompts that built it. MacKenzie Scott gave away [$7 billion in 2025](https://www.cnbc.com/2025/12/13/mackenzie-scott-revealed-her-total-charitable-donations-for-2025.html) — [about a third of every megagift in America that year](https://fortune.com/2026/06/25/mackenzie-scott-largest-megadonor-2025-7-billion-donations-giving-usa-iu-report/), by the Indiana University Lilly Family School of Philanthropy's count. Almost none of it is denominated in health. I built an interactive tool that asks what her [$26 billion](https://yieldgiving.com/) in lifetime giving buys in the unit health economists use to compare lives — quality-adjusted life-years: **[maxghenis.com/mackenzie-scott-qaly](/mackenzie-scott-qaly)**. Drag the assumptions and a Monte Carlo cost-effectiveness model reruns in your browser; on the best-guess default it lands around 202,000 QALYs — a model output, not a measured fact, which is why the tool exists: move the assumptions yourself. [![Dragging the evidence stance from RCT-only to face value: the median runs from about 9,000 QALYs through the 202,000 best guess to about 415,000, and the cause ranking reshuffles](./evidence-sweep.gif)](/mackenzie-scott-qaly) It runs the same machinery as the [GiveWell cost-effectiveness replication](/blog/givewell-cea/) I built in February — editable parameters, Monte Carlo, sensitivity analysis — pointed at a different question. GiveWell scores the most cost-effective charities in the world. This points the same lens at one donor's actual $26B portfolio, most of it unrestricted gifts to US organizations, where almost none of the spending is denominated in health to begin with. ## Why I care about the gap I [took the Giving What We Can pledge in 2019](/blog/why-ive-taken-the-giving-what-we-can-pledge/) for one reason: the global poor benefit far more from a dollar than I do. This tool puts a number on that for Scott's giving. Measured purely in health, a dollar at the [global-health frontier](https://www.givewell.org/charities/amf) — bed nets averting child deaths — buys roughly 500 times the QALYs of her blended portfolio. That's a marginal comparison; redeploying the full $26 billion would compress it toward the floor — around 20–40× even if every dollar became [direct cash](https://blog.givewell.org/2024/11/12/re-evaluating-the-impact-of-unconditional-cash-transfers/) — [the tool page works through the arithmetic](/mackenzie-scott-qaly). How much that gap should bother you is the open question, and the QALY count doesn't settle it. Part of the gap is outside her control: preventing a death costs far more in a rich country than a poor one, and a QALY ignores the income, opportunity, and rights her giving targets. Part of it is a choice: $26 billion is enough that where it goes carries real opportunity cost, measured in lives. The tool hands you the number, not the verdict. It's the same arithmetic that sends my own giving abroad. ## Two models, arguing I built this with two coding agents, and the more useful part was letting them check each other. Claude Code wrote the model and the tool. Then I had Codex review the assumptions cold: it caught a real error — a cost-per-life figure I'd left in old dollars without inflating it — and disagreed with Claude on whether the global benchmark belongs in QALYs or DALYs. I'm centralizing on QALYs. Two models disagreeing about a modeling choice is a sharper adversarial review than either alone. Two further full review rounds (both models, against everything) caught more, and the fixes moved the headline from ~98,000 to ~87,000: gifts are now recorded as the exact disclosed tranches and inflated to 2026 dollars year by year instead of divided nominal-vs-current; the community-health-center figure now uses the paper's own [~$54k per life-year](https://pmc.ncbi.nlm.nih.gov/articles/PMC4436657/) with an explicit life-year→QALY conversion (the old version skipped it); several 2000s-era cost-effectiveness anchors got inflated to current dollars; the frontier benchmark was re-derived twice — first for discounting consistency, then onto [GiveWell's current program averages](https://www.givewell.org/impact-estimates) — cutting the headline multiple by about 40%; the benefit/cost ratio now uses [HHS's published value per QALY](https://aspe.hhs.gov/reports/standard-ria-values) instead of a per-life-year value applied to QALYs; and one citation was re-attributed to the paper the numbers actually come from — [Sommers (2017)](https://www.journals.uchicago.edu/doi/10.1162/ajhe_a_00080), verified against the PDF, after I'd confidently planted the wrong one. Every correction made the model more skeptical or more honest, none was caught by a human, and the errors had survived earlier review passes. The allocation across causes was still my prior, though — I'd accepted "Scott doesn't publish dollars-by-cause" without checking hard enough. She publishes better: [Yield Giving's gift database](https://yieldgiving.com/gifts) itemizes dollar amounts for about two-thirds of the money, with focus areas on every disclosed dollar. Deriving the split from her own data — 53 org-reported areas mapped onto the model's 13 archetypes, [every rule documented](https://github.com/MaxGhenis/mackenzie-scott-qaly/blob/main/data/yieldgiving/leaf_to_archetype.yaml) — cut the skeptical median from ~87,000 to ~70,000. Her real portfolio holds less food, housing, and cash assistance than I'd assumed (the buckets with the strongest health evidence) and more workforce development and nonprofit infrastructure (which have almost none). The undisclosed third isn't dropped: her essays give each year's total, so the residual is a known dollar amount, spread over that year's undisclosed gifts in proportion to the recipient's pre-gift [IRS 990 revenue](https://projects.propublica.org/nonprofits/) raised to an elasticity fit on the disclosed pairs, with the fuzzy name-to-EIN matches audited by a third model, a small one, against the live API. That fit is a finding in its own right: across 1,313 disclosed gift–revenue pairs, gift size scales with organization revenue to the power 0.41 — a 10× bigger organization gets about 2.5× more money. Her giving is far flatter across organization size than proportional. The imputed third moves each cause share by at most ~1.5 points — the disclosed two-thirds was representative. The largest correction came from a reader. The model originally priced every health dollar at US anchors, so her gifts to organizations delivering abroad — [Malaria Consortium](https://yieldgiving.com/gifts), GiveDirectly's global program, Living Goods, Muso, Amref, Partners In Health, Evidence Action — were understated by orders of magnitude: bed nets at Medicaid rates. Health dollars now split by each organization's reported service locations, with the non-US share (~5% of her giving) priced at global-health anchors. That one fix tripled the skeptical median, from ~70,000 to ~205,000 QALYs, and it means most of the model's measured health impact comes from the small slice of her portfolio that reaches low-income countries. A follow-up audit of the 50 largest organizations that report only "global" service locations — cross-checked by two independent AI reviews against each organization's own filings — found many operate substantially in the US, trimming the default to ~202,000 (v1.1; the tool's version toggle keeps the July 13 allocation one click away). None of the model code was hand-written; the whole thing was natural-language prompting. For transparency, here is every prompt I typed, verbatim — typos and all. ## Appendix: the prompts **To Claude Code:** 1. estimate the qaly impact of mckenzie scott's lifetime donations 2. do better than judgment calls. make an actual quant analysis givewell-cea style. make a repo for it if thats useful 3. no just qalys for now. and considfer evidence quality especially causal identification credibility 4. /cycle 5. shall we make it an interactive and put it on maxghenis.com (using all the design tokens)? like a real py package with ci, nextjs/tailwind etc 6. inline link any falsifiable claims ALWAYS ALWAYS REMEMBER THIS 7. whats the global frontier 8. is it a falsifiable claim 9. fix 10. codex made some changes wdyt 11. give me a prompt for it to fix. but i think there was a misunderstanding i didnt ask for a qaly/daly change, could we centralize around one? which? 12. if you were starting this project from scratch howd you do it 13. yes 14. do a full review of this, both you and with a sol subagent 15. yep go - and dont have anyu allegiance to existing code, feel free to rebuild any and all things froms cratch 16. do it all 17. make sure the prose isnt bogged down by the history - like it seems like leading with the inflation adjustment in the tool is just based on it having been a bug we just fixed. users should see the tool first thing when they get to the page 18. can we give the blog post like a button at the top to use the tool, like we do for rambar etc 19. whats the delta against givewell? noting scaling limitations etc 20. sure yeah does it belong in the blog or the tool? like you could imagine a bigger version of the tool that also combines it with givewells for a general qaly estimator for every intervention and portfolio but not now 21. Vs. global frontier / 1,212× / more health per marginal $ — this makes it sounds like scott's is 1212x more effective 22. maybe we should drop the card and just discuss it in the written stuff... 23. then could you just make the 90% interval below in small text instead of a separate card, also add to the $/qaly 24. i dont get these two - whats the 2.1x wrt? whats 64.2b? 25. scott doesnt publish *any* dollars by cause? i thought she at least published her list of orgs, like are we basing this on any real data? 26. undisclosed ones can we do some research on, if need be we could scale by total revenue in the relevant year from 990s 27. what model are we using for this 28. i dont think we need fable for this fan out do we? could probably do like 5.6 terra or something 29. yeah terra is the small codex tier, use it for the audit **To Codex:** 1. review the math/assumptions in https://maxghenis.com/mackenzie-scott-qaly (its a local repo) 2. would py wasm be more parsimonious and dry 3. fix it all 4. deployed? 5. yes do The model, tests, and sources are on [GitHub](https://github.com/MaxGhenis/mackenzie-scott-qaly). --- ## The US government is gutting Anthropic's R&D capacity **URL:** https://maxghenis.com/blog/us-gutted-anthropic-rd/ **Published:** Jun 15 2026 **Description:** A Commerce export control pulled Claude Fable 5 and Mythos 5 on Friday. It's cutting roughly a third of Anthropic's own researchers off from the frontier models they build, while every competitor's keep theirs. import CrosspostAware from '../../components/CrosspostAware.astro'; Anthropic released Claude Fable 5 last week. I used it for three days before the government switched it off. I spent Tuesday through Thursday at a convening and Friday traveling without connectivity, so I worked at about a quarter of my usual pace. In that window I pointed Fable at a range of projects, including the microdata pipeline that PolicyEngine's tax and benefit modeling runs on. That pipeline combines household surveys, administrative records, and calibration against published totals into a synthetic population for federal, state, and local analysis. We had spent months rebuilding it. Fable redesigned the architecture across the dependent packages, rebuilt the pipeline, and produced a dataset about five times more accurate than the one it replaces, measured against our held-out targets. That work reaches users this week. In the same three days it re-engineered the agentic flows behind several systems we ship over the coming weeks. That is one model, used part-time, for three days. The effect that matters most from Friday's order is not on its customers. It is on Anthropic's own researchers. ## What the order does On Friday the Commerce Department directed Anthropic to suspend access to Fable 5 and the related Mythos 5 model for any foreign national, inside or outside the United States, including the company's own employees. Commerce used the Export Administration Regulations' ["deemed export" rule](https://www.bis.gov/learn-support/deemed-exports/what-deemed-export), which treats releasing controlled technology to a foreign person in the country as an export to their home country. Anthropic [could not screen its customers by nationality in real time, so it pulled both models for all users](https://www.anthropic.com/news/fable-mythos-access). Claude Opus 4.8 and the earlier models stayed up. That global cutoff is the commercial cost. The research cost is narrower and more specific. ## Who it bars Anthropic knows each employee's status, so the cut is precise: its foreign-national researchers lose access to the models they build, while their citizen and green-card colleagues keep using them. The deemed-export rule the order invoked [exempts citizens, green-card holders, and "protected individuals"](https://www.bis.gov/learn-support/deemed-exports/what-deemed-export); the bite lands on employees on temporary visas. ([One outlet read the order to cover green-card holders too](https://www.i-scoop.eu/fable-5-blocked-outside-the-us-after-government-export-order/); the rule it invoked says the opposite.) How murky even that line is showed up within a day. Headlines reported that [Andrej Karpathy, who joined Anthropic in May to lead its recursive self-improvement work, had been locked out of the models he was hired to build](https://www.ibtimes.co.uk/ai-expert-andrej-karpathy-anthropic-tech-regulation-1802715) for lacking US citizenship. But the claim that he was barred traced to [a single viral post sourcing his immigration status to an AI chatbot](https://x.com/AndrewCurran_/status/2065619713485627829) — no one verified it, and if he holds a green card the order never reached him at all. For days no one could say cleanly whether one of the field's best-known researchers was in or out. ## How many Frontier AI runs on immigrant talent: about [two-thirds of top-tier AI researchers in the United States earned their undergraduate degrees abroad](https://macropolo.org/interactive/digital-projects/the-global-ai-talent-tracker/), and international students are [the core of the US AI PhD pipeline](https://cset.georgetown.edu/article/international-graduate-students-critical-to-u-s-ai-competitiveness/). No public figure exists for Anthropic, so I asked Claude to estimate it — five independent passes, each from a different angle: the AI-talent pipeline, Anthropic's visa filings, the China–India green-card backlog, the company's youth, and peer-company workforces. They ranged from 27% to 42% of its researchers, with a mean near 36% — on the order of a third. The passes agree on why it runs high: Anthropic is only four years old, so few foreign-origin researchers have had time to naturalize, and [green-card backlogs for Chinese and Indian nationals](https://www.dwt.com/insights/2020/05/us-bis-china-technology-export-ban) — the field's two largest talent sources — run a decade or more, leaving even long-tenured staff on temporary visas. It is an estimate, not a count: Anthropic publishes no breakdown, and its [H-1B filings](https://h1bgrader.com/h1b-sponsors/anthropic-pbc-10j7p51g0v) are flow, not stock. The barred researchers also have an obvious exit. The control attaches to Anthropic's models, not to the people, so a visa-holder cut off inside Anthropic can join OpenAI, Google, or any other lab and use a frontier model there on day one — every researcher at those labs, visa holders included, keeps access, because the order names only Anthropic's models. So it does not just bench a third of Anthropic's researchers; it hands every competitor a standing recruiting pitch aimed squarely at them, turning a model restriction into a talent-export pump pointed at one company. ## Even if it's temporary None of this needs to be permanent to matter. Prediction markets expect a temporary pause — both [Polymarket](https://polymarket.com/event/us-government-rescinds-claude-fable-5-foreigner-ban-by-20260613202427937) and [Kalshi](https://news.kalshi.com/p/fable-5-odds-anthropic-access-restored-july-57-percent) have access most likely returning within a few weeks. But the field doesn't pause with it, and it's speeding up: the length of tasks frontier models can do reliably has been [doubling every seven months — lately closer to four](https://metr.org/time-horizons/), a doubling time that keeps shrinking. On a curve that steep and still accelerating, weeks on the sidelines — while rivals keep training, shipping, and hiring — shift leads and define careers. A temporary pause here is not a small one. ## The incentive it sets Applied as a standard, this penalizes shipping. Every frontier release would cut a lab's visa-holding researchers off from the model they just shipped, the one they need to build the next. A lab that ships nothing keeps its researchers' access; a lab that ships loses it. ## The path back A mechanism exists. The deemed-export rule routinely lets foreign nationals work on controlled technology under an [individual license](https://www.bis.doc.gov/index.php/licensing/14-policy-guidance/deemed-exports/109-guidelines-for-foreign-national-licenses), usually backed by a technology control plan; semiconductor and aerospace firms employ non-citizens on restricted work this way. But BIS reviews each license under the rules for exporting to that employee's country of citizenship. For nationals of close allies, approval is routine. For nationals of countries the United States restricts, including China, the review [carries a presumption of denial](https://www.dwt.com/insights/2020/05/us-bis-china-technology-export-ban). The route back is per-employee and country-by-country, measured in months, and it would return some of the workforce while leaving the rest in review. ## On a consistent standard The order's authority is the [Export Control Reform Act of 2018](https://www.justsecurity.org/142745/law-anthropic-export-controls/), applied through an "is-informed" letter to a single company — the first time it has reached an AI model — and [Lutnick gave no basis](https://www.axios.com/2026/06/12/anthropic-trump-mythos-fable-national-security) for why Anthropic specifically. Legal analysts called it "not clear why Anthropic's models were singled out … as compared to other U.S. large language model AI companies," and "a remarkable contrast" with the administration's hands-off stance on AI exports. The directive names only Anthropic's two models, and the jailbreak behind it is disputed. Anthropic says the government's evidence was [verbal, and demonstrated by a competitor](https://cyberscoop.com/us-government-anthropic-fable-5-mythos-5-export-controls/), and that the same capability is available in [OpenAI's GPT-5.5, unrestricted](https://www.anthropic.com/news/fable-mythos-access). Security researcher Katie Moussouris [described the technique](https://techcrunch.com/2026/06/15/the-us-governments-anthropic-models-ban-was-never-about-an-ai-jailbreak/) as the gap between asking a model to "review code for security issues" and to "fix this code," and said it "cannot meaningfully be fixed, and any attempt would only weaken the model for defense." The fixation on Anthropic predates the order. On a May 8 podcast — weeks before Anthropic overtook OpenAI in valuation — the administration's AI czar, David Sacks, [called it "the most powerful monopoly ever created in human history"](https://www.youtube.com/watch?v=10MdOvK-aG4&t=982s) and likened it to Standard Oil rebranded as "Safe Oil," casting its safety record as cover for a monopoly. On the same show, Sacks said the cyber capability was not Anthropic's alone — ["OpenAI now has a model that's just as cyber capable as Mythos,"](https://www.youtube.com/watch?v=10MdOvK-aG4&t=2713s) with every major lab and Chinese models to follow within months — and his prescription then was industry-wide hardening, not a restriction on one company.
David Sacks's "Safe Oil" monologue — All-In, May 8, 2026

“Unless something about their current trajectory changes, Anthropic will be the most powerful monopoly ever created in human history. […]

I just want you to think for a second about the case of John D. Rockefeller, who I think is known as probably the most successful, most ruthless monopolist in American history. But he wasn't very good at PR. He was terrible at PR. Everyone sort of recognized how ruthless he is. We've seen movies like There Will Be Blood, which is basically about him.

In any event, imagine if John D. Rockefeller was way better at public relations, and instead of calling his company Standard Oil, he called it Safe Oil. … He called it Safe Oil because, as we know, kerosene is dangerous. Their first big product was kerosene. And kerosene can light your house or it can burn it down, and in the wrong hands it can torch a city or you can use it to make a bomb. So John D should have called for the creation of a new government agency to regulate the safety of his product, and they could have done rigorous testing, licensing, common-sense regulation. … And I think people would have gotten so wrapped up in this debate over what constituted Safe Oil or Safe Kerosene that they would have missed what was really going on, which is that Rockefeller was building the richest, most powerful monopoly of all. In fact, people might even have called Rockefeller an Effective Altruist, because of course he was so concerned about the safety of his product.”

The day before the export control, the same hosts opened their show attacking Fable for the opposite sin — too much caution. They called its guardrails "Orwellian" and its 30-day data retention "mandatory surveillance," and Sacks called the safety posture ["a very sophisticated regulatory capture campaign based on fear-mongering."](https://www.youtube.com/watch?v=gH4FTjDm9FQ&t=602s) Later in the same episode, asked about Bernie Sanders's plan to seize a government stake in the AI companies, Sacks said he ["may be okay with Bernie's idea"](https://www.youtube.com/watch?v=gH4FTjDm9FQ&t=3364s) — but only for "a public benefit corporation that says it's going to cause massive job loss, that trained for free on humanity's knowledge but gatekeeps and refuses to give back," and "maybe it should be 75%." Sacks offered those as principled conditions, but they fit OpenAI as well as Anthropic: both are public benefit corporations, both trained on the same public data, both keep their weights closed, and both CEOs have predicted mass job loss. A test that singles out one of two near-identical companies isn't a test; it's a target with criteria attached. The next day the government restricted Anthropic, and only Anthropic, for being insufficiently safe. No published standard tells another lab what to avoid, or tells Anthropic what to fix — only that, for now, the one company singled out must route its own researchers through a months-long license to use the model they built, or watch them rebuild it somewhere else. The deeper irony is that standards-based, politically neutral governance is what Anthropic spent years asking for. It proposed uniform disclosure rules for every frontier lab in ["The Case for Targeted Regulation"](https://www.anthropic.com/news/the-case-for-targeted-regulation), and days before the order, Dario Amodei's ["Policy on the AI Exponential"](https://darioamodei.com/post/policy-on-the-ai-exponential) argued that any government power to block a model "must be scoped to … four specific risks" with "protective measures against political favoritism or arbitrary decisions." Defending the order, Sacks [wrote that Anthropic "asked for government regulation of Mythos"](https://x.com/DavidSacks/status/2065853007619588171) — but its proposals were model-neutral, applied to all frontier developers, and it had gated Mythos itself rather than ask Washington to single it out. Anthropic asked to be governed by a standard applied to everyone; it got an ad-hoc decree applied to it alone.

Appendix: David Sacks's "Safe Oil" monologue

All-In, May 8, 2026

“Unless something about their current trajectory changes, Anthropic will be the most powerful monopoly ever created in human history. […]

I just want you to think for a second about the case of John D. Rockefeller, who I think is known as probably the most successful, most ruthless monopolist in American history. But he wasn't very good at PR. He was terrible at PR. Everyone sort of recognized how ruthless he is. We've seen movies like There Will Be Blood, which is basically about him.

In any event, imagine if John D. Rockefeller was way better at public relations, and instead of calling his company Standard Oil, he called it Safe Oil. … He called it Safe Oil because, as we know, kerosene is dangerous. Their first big product was kerosene. And kerosene can light your house or it can burn it down, and in the wrong hands it can torch a city or you can use it to make a bomb. So John D should have called for the creation of a new government agency to regulate the safety of his product, and they could have done rigorous testing, licensing, common-sense regulation. … And I think people would have gotten so wrapped up in this debate over what constituted Safe Oil or Safe Kerosene that they would have missed what was really going on, which is that Rockefeller was building the richest, most powerful monopoly of all. In fact, people might even have called Rockefeller an Effective Altruist, because of course he was so concerned about the safety of his product.”

--- ## How to turn money into predictions **URL:** https://maxghenis.com/blog/how-to-turn-money-into-predictions/ **Published:** May 26 2026 **Description:** A menu of mechanisms for organizations that want to fund forecasts on policy, economic, and AI outcomes. Organizations that fund policy research, economic analysis, or AI deployment want predictions about future outcomes — which tax provisions will pass, what inflation will run, how labor markets will absorb a new model class. The [Federation of American Scientists is sponsoring forecasting tournaments](https://fas.org/publication/fas-and-metaculus-are-using-forecasting-to-support-better-climate-policy/) to inform climate policy; the [Forecasting Research Institute ran its Existential Risk Persuasion Tournament](https://forecastingresearch.org/xpt) with 80 domain experts and 89 superforecasters on AI and other catastrophic risks; [FiscalNote announced a major expansion](https://fiscalnote.com/newsroom/fiscalnote-announces-major-expansion-into-political-prediction-markets) into prediction-market infrastructure for its policy-intelligence customer base in February 2026; and AI researchers are now [benchmarking LLM forecasting against expert and crowd panels](https://www.forecastbench.org/paper.html). Quality varies sharply across markets — some are deep, actively traded, and accurate; others sit thinly populated and stay mispriced for weeks. Funders can improve the quality on the questions that matter to them. This post lays out the available mechanisms. I drafted an earlier version in 2024 and circulated it with a small group of researchers and funders. The landscape has changed enough since then — prediction markets [crossed $44 billion in notional volume in 2025](https://www.theblock.co/post/383733/prediction-markets-kalshi-polymarket-duopoly-2025), Kalshi got [Federal Reserve research validation](https://www.federalreserve.gov/econres/feds/files/2026010pap.pdf), and a [serious proposal](https://empiricrafting.substack.com/p/we-need-well-capitalized-prediction) for sponsorship-funded markets on AI labor impacts hit the literature — that an updated public version is overdue. ## Two structural problems Two problems make most prediction markets shallower than they should be. **The liquidity problem.** Markets need both informed traders ("smart money") and less informed traders willing to absorb their bets ("dumb money"). Andrew Gelman puts it cleanly in [Prediction markets and the need for "dumb money" as well as "smart money"](https://statmodeling.stat.columbia.edu/2024/10/25/prediction-markets-and-the-need-for-dumb-money-as-well-as-smart-money/): "Markets become efficient when making them efficient is profitable." Without enough total volume, smart traders can't profitably correct a mispriced market, so prices stay wrong. A Manifold market asking ["Will there be any tax on unrealized capital gains in the USA before EOY2028?"](https://manifold.markets/Gen/will-there-be-any-tax-on-unrealized) sat at 22% probability for nearly two weeks after Trump's 2024 election victory, with 27 traders. The election outcome obviously cut the probability substantially, but no one was paid enough to move the price. A parallel Manifold market on ["Will a raise to the top capital gains tax rate be enacted in 2025?"](https://manifold.markets/jack/will-the-top-capital-gains-tax-rate) behaved differently. Manifold staff made it sweepstakes-eligible at creation, so the market traded in real money (Manifold's then-experimental real-cash mode, [shut down in March 2025](https://news.manifold.markets/p/focusing-on-mana-bringing-sweepstakes)). It held 20-21% until election results, then dropped to 6%. The mana version attracted 18 trades and 4,093 mana; the sweepstakes version had 9 traders and 334 sweepcash. The cash version attracted enough attention to update on election news; the play-money version did not. Money matters here mostly as an attention allocator, not as a calibration mechanism — [Servan-Schreiber et al. (2004)](https://users.nber.org/~jwolfers/Papers/DoesMoneyMatter.pdf) found play-money and real-money NFL prediction markets equally accurate when both had engaged trader bases. **The zero-sum challenge.** Unlike equities, where investors collectively benefit from economic growth, prediction markets are zero-sum or negative-sum after fees. Every winner needs a loser. Zvi Mowshowitz lists the consequences in ["Subsidizing Prediction Markets"](https://www.lesswrong.com/posts/AeKS2m6uLM8RYfvND/subsidizing-prediction-markets): no natural long-term investors, higher perceived risk, and limited professional capital deployment. Funders willing to take an expected loss on a market are the substitute for that missing long-horizon capital. ## Where prediction markets are in 2026 **Volume.** Total notional [exceeded $44 billion in 2025](https://www.theblock.co/post/383733/prediction-markets-kalshi-polymarket-duopoly-2025), with monthly peaks around $13 billion during the US election period. The market is concentrated in [Kalshi](https://kalshi.com/) (a [CFTC-designated contract market](https://www.cftc.gov/PressRoom/PressReleases/8302-20)) and [Polymarket](https://polymarket.com/) (crypto-native, recently [returned to the US](https://www.coindesk.com/policy/2025/09/03/u-s-cftc-gives-go-ahead-for-polymarket-s-new-exchange-qcx) via a $112 million acquisition of CFTC-licensed exchange QCX). **Regulatory legitimacy.** Kalshi went from a startup with questionable contract approvals to the subject of a [Federal Reserve working paper](https://www.federalreserve.gov/econres/feds/files/2026010pap.pdf) (Diercks, Katz, and Wright, 2026) that finds Kalshi's prices on inflation, jobs, GDP, and FOMC decisions "outperform surveys and interest-rate futures" as macro forecasting tools. The paper explicitly recommends Kalshi as a real-time benchmark for researchers and policymakers. That is the most credible institutional validation prediction markets have ever received. **A sponsorship thesis.** Andrey Fradkin, Brian Jabarian, and Andrew Koh published ["We need well-capitalized prediction markets"](https://empiricrafting.substack.com/p/we-need-well-capitalized-prediction) arguing that the right model is sponsorship — AI labs, Big Tech firms, and philanthropies seed liquidity in markets on labor outcomes, AI capability benchmarks, and other questions where they have decision-relevant interest. Sponsorship covers the expected loss in exchange for better information for everyone. I'll come back to this in the last section. ## Funding mechanism 1: Subsidize public markets Each major platform supports liquidity provision through a different mechanism. [Manifold Markets](https://manifold.markets/) runs on play money called mana. The platform briefly offered sweepstakes trading in real cash but [shut that program down in March 2025](https://news.manifold.markets/p/focusing-on-mana-bringing-sweepstakes) and returned to a pure play-money model. Despite the play-money currency, Manifold's [aggregate calibration](https://manifold.markets/calibration) is strong — predictions land within roughly four percentage points of realized frequencies — and Manifold's own analysis finds calibration improves with trader count up to roughly 10-20 traders per market, after which it plateaus. The implication for funders: subsidizing a market that's stuck below the plateau directly buys better predictions. Organizations can [purchase mana with dollars](https://manifold.markets/checkout) and [add it to specific markets as liquidity](https://docs.manifold.markets/faq#how-do-i-subsidise-a-market); the deeper book attracts more traders and tightens the spread. The subsidizer expects to lose mana — the more the market price moves, the less you recover — but the lost mana is the cost of a better-priced market, not a tax on the platform. Manifold's code is also [MIT-licensed](https://github.com/manifoldmarkets/manifold), which makes it the friendliest platform for organizations that want to run programmatic experiments at low cost. [Polymarket](https://polymarket.com/) uses a [Liquidity Mining & Rewards program](https://docs.polymarket.com/market-makers/liquidity-rewards). Market makers earn rewards for providing consistent quotes; traders earn fee rebates proportional to trading activity; community bounties incentivize market promotion. Organizations can participate through these mechanisms rather than as standalone subsidizers. After a [2022 CFTC settlement](https://www.cftc.gov/PressRoom/PressReleases/8478-22) that barred US users for nearly four years, Polymarket regained US access in December 2025 via the QCX acquisition above, and runs [hundreds of macro markets](https://polymarket.com/dashboards/macro) globally. [Kalshi](https://kalshi.com/) offers a [limit order](https://help.kalshi.com/trading/order-types/limit-orders) liquidity model. Organizations place standing limit orders at prices they're willing to trade at; orders execute when the market crosses those prices. Limit orders avoid trading fees, allow precise price targeting, and accumulate into market liquidity over time. Qualified market makers can join a formal [market maker program](https://help.kalshi.com/en/articles/13823819-market-maker-program) with additional benefits. The CFTC approval and Fed validation make this the right platform for institutional-grade questions where regulatory clarity matters. [Hypermind](https://corp.hypermind.com/) has run institutional prediction markets since 2000, with clients including the US intelligence community and the [Johns Hopkins Center for Health Security](https://predict.hypermind.com/hypermind/media/pdf/futuribles-prediction-markets.pdf). Sponsors fund custom markets and panels of selected forecasters on bespoke questions, with the platform providing aggregation and reporting infrastructure. Less suitable for one-off retail experiments, more suitable for sustained corporate or government use. ## Funding mechanism 2: Sponsor a tournament [Metaculus](https://www.metaculus.com/) is a forecasting platform aggregating predictions across thousands of binary and numerical questions, scored on accuracy. Organizations [sponsor dedicated tournaments](https://www.metaculus.com/tournaments/) — a slate of related questions with a prize pool that pays out to the best forecasters. The [Federation of American Scientists](https://fas.org/publication/fas-and-metaculus-are-using-forecasting-to-support-better-climate-policy/) funded a $5,000 [Climate Tipping Points tournament](https://www.metaculus.com/tournament/climate/) covering 43 questions about climate policy and outcomes, including conditional questions. The tournament structure lets a funder set the agenda — what questions matter, what counts as resolution, what horizons to ask about — without taking on market-maker risk. The [Forecasting Research Institute](https://forecastingresearch.org/) (Philip Tetlock's research arm, distinct from Good Judgment Inc) runs research-oriented tournaments like the [Existential Risk Persuasion Tournament](https://forecastingresearch.org/xpt) (80 domain experts, 89 superforecasters, thousands of forecasts on long-horizon catastrophic risks). Sponsoring FRI is closer to funding a research program than buying a forecast. Tournaments work well for questions where you want a calibrated probability today, not a tradeable instrument over time. They also support reasoning prizes for the best-argued forecasts, not just the most accurate ones, which matters when you're trying to surface analytical talent rather than just numerical answers. ## Funding mechanism 3: Buy custom human forecasts [Good Judgment](https://goodjudgment.com/) sells [custom forecasting from professional superforecasters](https://goodjudgment.com/services/custom-superforecasts/) — individuals identified by [Philip Tetlock's Good Judgment Project](https://goodjudgment.com/resources/the-superforecasters-track-record/) as [consistently outperforming domain experts and intelligence analysts](https://en.wikipedia.org/wiki/The_Good_Judgment_Project) (the project beat intelligence-community analysts with classified access by 25-30% in the IARPA ACE tournament). Services include question development, daily forecast updates, written analysis, and follow-up question support. Organizations can keep forecasts private or release them publicly. This is the most concierge option. You write a question, a small team of trained forecasters works on it, you get back a probability with explanation. No liquidity to manage, no market to monitor. Cost is higher per question and the methodology is opaque relative to a public market. ## Funding mechanism 4: Buy AI-generated forecasts A new category emerged in 2024-2026. [FutureSearch](https://futuresearch.ai/) ([publicly launched in early 2026](https://futuresearch.ai/company/)) runs LLM-based research agents that gather evidence, weigh base rates, and produce calibrated probability forecasts with reasoning. Pricing is [per-operation](https://futuresearch.ai/pricing/) — deep-research agents at 1-11¢, forecasters at 20-90¢ per researcher per row — orders of magnitude cheaper than human superforecasters per question, with the obvious tradeoff that the depth of domain expertise is whatever the model has internalized plus what its retrieval finds. The category builds on published research: Halawi et al. (2024) ["Approaching Human-Level Forecasting with Language Models"](https://arxiv.org/abs/2402.18563) (NeurIPS) showed retrieval-augmented LLM systems approaching competitive-forecaster accuracy; the Center for AI Safety's [forecasting bot](https://safe.ai/blog/forecasting) reported superhuman accuracy on certain competitive platforms; [ForecastBench](https://www.forecastbench.org/paper.html) (Karger et al., ICLR 2025) — a dynamic, continuously-updated benchmark — finds frontier LLMs roughly match the general-public crowd in Brier score but still trail elite superforecasters by a wide margin (LLM Brier ≈ 0.135–0.159 vs. superforecaster Brier ≈ 0.02). The ForecastBench team's linear extrapolation [projects superforecaster parity around late 2026](https://forecastingresearch.substack.com/p/ai-llm-forecasting-model-forecastbench-benchmark), with a 95% confidence interval from December 2025 to January 2028. The cost structure inverts the human-forecasting category. Where Good Judgment's value is depth on a few high-stakes questions, AI forecasting's value is breadth — hundreds or thousands of calibrated probabilities a day. Use AI forecasting when you want coverage; use human superforecasters when you want depth and defensibility on a question where the audience expects a human accountable for the call. The trajectory matters more than the snapshot. Per-query costs are falling and accuracy is rising; within a few years, the marginal cost of a calibrated probability on most public questions drops sharply, even if the gap to elite superforecasters takes longer to close than the ForecastBench point estimate suggests. What remains expensive — and where real returns on investment will sit — is *optimizing* the agents: which tools they call, which substrates they query, how their calibration is benchmarked, how they update under new information. Cheap commodity forecasts and frontier-quality forecasts will look very different, and the work to produce the latter still needs to be incentivized. ## Funding mechanism 5: Run your own prediction market For internal questions an organization doesn't want public — product launch dates, project completion probabilities, internal strategic questions — private prediction markets aggregate dispersed information already inside the organization. [Eli Lilly used internal markets in 2003](https://en.wikipedia.org/wiki/Prediction_market#Business) to forecast which drug candidates would advance through clinical trials; Google, Microsoft, and Ford have run variants. The corporate-prediction-market vendor landscape has consolidated — [Cultivate Labs](https://www.cultivatelabs.com/posts/why-we-stopped-supporting-prediction-markets) (formerly Inkling Markets) dropped market mechanics in 2022 in favor of opinion pools, citing user confusion — but [Hypermind's Prescience platform](https://corp.hypermind.com/) still offers managed private markets for corporate and government clients. ## The sponsorship thesis, and why it matters The [Fradkin / Jabarian / Koh proposal](https://empiricrafting.substack.com/p/we-need-well-capitalized-prediction) is worth treating as its own category. They argue that the chicken-and-egg problem in prediction markets — no liquidity, no traders; no traders, no liquidity — can be solved by sponsorship capital from organizations that have real decision-relevant interest in the questions. Their canonical example is AI labor impacts: a major AI lab, a federal agency, or a philanthropy seeds liquidity on markets tied to [Bureau of Labor Statistics](https://www.bls.gov/) data series — occupation-level employment, wages, labor-force participation at 1-, 2-, and 5-year horizons. The contract design they advocate has four properties: verifiability (objective resolution against published government data), stability (consistent measurement over time), robustness to gaming (large underlying quantities that one trader can't move), and attention (sufficient interest to draw informed traders). BLS data hits all four. So do many other government statistics — [BEA NIPA tables](https://www.bea.gov/itable/national-gdp-and-personal-income), [Census ACS variables](https://www.census.gov/programs-surveys/acs), [IRS Statistics of Income](https://www.irs.gov/statistics). The model generalizes. Any organization that benefits from better forecasting on a specific question can sponsor a market on it without expecting trading profit. The sponsorship pays for information that improves everyone's decisions — including the sponsor's. Conditional markets (outcome Y given policy state X) are particularly underprovided today and particularly valuable for policy analysis. Sponsorship is the cleanest mechanism to bring them into existence. The sponsorship lane connects to a structural question: who pays for the public good of calibrated forecasts? In 2024 the answer was mostly "individual hobbyist traders subsidizing inefficient markets." In 2026 the answer is starting to be "institutions with skin in the question, sponsoring markets that produce information they want." That shift is what makes 2026 different from 2024. ## Where to start Different mechanisms suit different goals: - **Want to test the waters cheaply?** [Buy mana on Manifold](https://manifold.markets/checkout) and seed a market that matters to you. Total commitment can be under $1,000. The mana you lose to traders moving the price is the cost of a deeper, better-calibrated market on a question you care about. - **Want a probability anchored to your priorities?** [Sponsor a Metaculus tournament](https://www.metaculus.com/tournaments/). $5,000–$50,000 buys a slate of questions with a real prize pool and visible forecasters. - **Want a single high-stakes forecast with human reasoning?** Engage [Good Judgment](https://goodjudgment.com/services/custom-superforecasts/). Highest cost per question, deepest analysis, accountable human author. - **Want broad coverage across many questions cheaply?** Use [FutureSearch](https://futuresearch.ai/) or a similar AI-forecasting service. Per-query pricing in pennies, hundreds of forecasts a day, calibrated but synthetic. - **Want regulated, real-money markets on macro questions?** Place [limit orders on Kalshi](https://help.kalshi.com/trading/order-types/limit-orders). Fee-free, precise pricing, supports the most institutionally credible platform. - **Want an institutional sustained-engagement platform?** [Hypermind](https://corp.hypermind.com/) for custom corporate panels (25+ years of track record) or for running a private internal market with a defined trader population. - **Want a research tournament on long-horizon questions?** Fund the [Forecasting Research Institute](https://forecastingresearch.org/) directly — closer to research-program sponsorship than to buying a forecast. - **Want to move the field?** Sponsor markets that solve a real liquidity hole — BLS-anchored contracts, AI capability benchmarks, conditional policy-outcome markets. The [Fradkin / Jabarian / Koh model](https://empiricrafting.substack.com/p/we-need-well-capitalized-prediction) is the right starting point. The hard part is no longer figuring out the mechanism. The mechanisms exist, the platforms work, the regulatory path is clearer than at any point in the past decade. The opportunity is institutional: adding prediction markets to the portfolio of truth-seeking tools organizations already use — alongside surveys, expert panels, internal modeling, and traditional forecasts — and developing the craft of writing forecastable questions. A well-formed question is conditional, scoreable, and decision-relevant ("will 2028 child poverty under reform X exceed Y%?"); a poorly-formed one is none of those. Prediction markets are a new epistemic technology. Organizations that build the discipline to use them well will reason more clearly about the future. --- ## Can Talkie-1930 do arithmetic? **URL:** https://maxghenis.com/blog/talkie-1930-math-evals/ **Published:** Apr 30 2026 **Description:** I tested Talkie-1930 on GSM8K and the easier EleutherAI/OpenAI arithmetic suite, then packaged an lm-eval-harness runner so the runs are reproducible. import CrosspostAware from '../../components/CrosspostAware.astro'; import TalkieEvalsEmbed from '../../components/TalkieEvalsEmbed.astro'; [Talkie](https://talkie-lm.com/introducing-talkie) is a 13B language model trained only on text available before 1931. Its launch post includes a "Numeracy" panel showing talkie-1930 reaching about 62% average accuracy at peak training compute, slightly above the modern-web twin's roughly 57%. But the plot doesn't specify which tasks the average covers, which prompts were used, or how answers were scored. I wanted a smaller, inspectable check: can Talkie-1930 do arithmetic at all? The answer depends on the evaluation. In full `lm-evaluation-harness` runs, the instruction-tuned model scores 0.0% under the strict GSM8K final-answer parser and 4.3% under the flexible parser in zero-shot. With the standard 5-shot GSM8K prompt, it rises to 4.9% strict and 7.2% flexible. On the easier EleutherAI/OpenAI arithmetic suite, the 1930 base model averages 42.7%, the 1930 instruction model averages 42.2%, and the modern-web base model averages 3.4%. ## GSM8K: Mostly no I first tried [GSM8K](https://huggingface.co/datasets/openai/gsm8k), the grade-school math word-problem dataset OpenAI released in [Cobbe et al. (2021)](https://arxiv.org/abs/2110.14168). For context, the same-size [LLaMA 13B reports 17.8%](https://arxiv.org/abs/2302.13971) (Touvron et al. 2023, Table 7); GPT-4 hit [92.0%](https://cdn.openai.com/papers/gpt-4.pdf) at its March 2023 release; [Claude 3 Opus reported 95.0%](https://www.anthropic.com/news/claude-3-family) in March 2024. I used the instruction-tuned Talkie-1930 model, the Talkie instruction chat template, and greedy decoding (`do_sample: false`, `temperature: 0.0`). | Run | N | Strict match | Flexible extract | | --- | ---: | ---: | ---: | | Zero-shot | 1,319 | 0.0% | 4.3% | | 5-shot | 1,319 | 4.9% | 7.2% | The strict score requires the GSM8K-style final answer marker (`#### number`). The flexible score extracts a number-like string from the model response. Flexible extraction is more forgiving when a model doesn't follow the GSM8K answer format, but it is still only 7.2% at 5-shot. A representative correct 5-shot response: > **Q:** A robe takes 2 bolts of blue fiber and half that much white fiber. How many bolts in total does it take? > > **Talkie:** It takes 3 bolts in total because 2+1/2 = `<<2+1/2=3>>`3 bolts. `#### 3` ✓ A typical miss: > **Q:** Janet's ducks lay 16 eggs per day. She eats three for breakfast every morning and bakes muffins for her friends every day with four. She sells the remainder at the farmers' market daily for $2 per fresh duck egg. How much in dollars does she make every day at the farmers' market? > > **Talkie:** Janet makes $2 x 16 = `<<32>>`32 per day at the farmers' market. `#### 32` (gold: 18) The model imitates the GSM8K answer format but skips the operations the problem actually requires. Before running the full harness job, I also ran smaller custom probes: direct answer, zero-shot reasoning, and 4-shot reasoning prompts. Those got 0 of 70 attempts right. I now treat those as audit probes rather than headline benchmark results; the full harness runs above are the reproducible numbers. GSM8K requires reading a word problem, tracking quantities, choosing operations, and formatting an answer. For Talkie, generation and instruction-following are themselves part of the bottleneck. ## Easier arithmetic: Sometimes yes I then used the [EleutherAI arithmetic dataset](https://huggingface.co/datasets/EleutherAI/arithmetic), which comes from the OpenAI GPT-3 arithmetic tests in [Brown et al. (2020)](https://arxiv.org/abs/2005.14165). For context, GPT-3 175B in that paper (Table 3.9, few-shot) already hit 100% on 2-digit addition, 98.9% on 2-digit subtraction, and 80.4% on 3-digit addition; multiplication and 4–5-digit operations were weaker (29.2% on 2-digit multiplication, 9.3% on 5-digit addition). The current [lm-evaluation-harness task definitions](https://github.com/EleutherAI/lm-evaluation-harness/tree/main/lm_eval/tasks/arithmetic) score these as log-likelihood tasks: given a context like: ```text Question: What is 98 plus 45? Answer: ``` the model is correct if the exact target completion, like ` 143`, is the greedy continuation under teacher forcing. This is much easier than GSM8K, and closer to a base-LM benchmark. I ran all 2,000 validation examples from each of the 10 arithmetic tasks. | Task | 1930 base | 1930 instruct | Modern-web base | | --- | ---: | ---: | ---: | | Single-digit 3 ops | 11.5% | 16.3% | 3.4% | | 2-digit addition | 91.6% | 75.7% | 14.0% | | 2-digit subtraction | 51.0% | 49.8% | 11.5% | | 3-digit addition | 74.7% | 88.3% | 0.6% | | 3-digit subtraction | 47.2% | 48.7% | 1.7% | | 4-digit addition | 29.5% | 27.7% | 0.0% | | 4-digit subtraction | 36.1% | 30.7% | 0.1% | | 5-digit addition | 31.4% | 24.4% | 0.0% | | 5-digit subtraction | 28.2% | 30.1% | 0.0% | | 2-digit multiplication | 26.2% | 30.8% | 3.1% | | **Overall** | **42.7%** | **42.2%** | **3.4%** | The 1930 models aren't just refusing. They often pick the right answer as their top continuation, especially for addition (91.6% base on 2-digit, 74.7% base and 88.3% instruct on 3-digit). For "What is 98 plus 45?" the 1930 base preferred the correct ` 143`. When the base model is correct on 2-digit addition, its median probability on the full target sequence is 51.6% — winning the argmax, not asserting confidently. The pattern breaks on multi-operation expressions, subtraction, multiplication, and larger digits. The modern-web base scores 3.4% overall. In many errors it copied an operand instead of computing the result: for "What is 98 plus 45?" it preferred ` 98` over ` 143`; for "What is 92 times 7?" it preferred ` 92` over ` 644`. In a 500-row custom audit, 70% of its 2-digit addition rows and 25% of its 2-digit multiplication rows were exact operand copies. I don't read this as evidence that pre-1931 text makes a model more numerate than modern web text. More likely, I'm not reproducing the Talkie authors' exact benchmark setup, or this completion format interacts badly with the modern-web checkpoint. ## Compared to same-size GPT-3 The arithmetic suite first appears in [Brown et al. (2020)](https://arxiv.org/abs/2005.14165). The natural same-size baseline is GPT-3 13B; GPT-3 175B is the headline. Few-shot accuracies from Appendix H, Table H.1: | Task | GPT-3 13B | GPT-3 175B | Talkie-1930 13B base | | --- | ---: | ---: | ---: | | Single-digit 3 ops | 9.95% | 21.3% | 11.5% | | 2-digit addition | 55.5% | 100.0% | 91.6% | | 2-digit subtraction | 52.4% | 98.9% | 51.0% | | 3-digit addition | 8.4% | 80.4% | 74.7% | | 3-digit subtraction | 9.2% | 94.2% | 47.2% | | 4-digit addition | 0.4% | 25.5% | 29.5% | | 4-digit subtraction | 0.4% | 26.8% | 36.1% | | 5-digit addition | 0.05% | 9.3% | 31.4% | | 5-digit subtraction | 0.0% | 9.9% | 28.2% | | 2-digit multiplication | 7.05% | 29.2% | 26.2% | Talkie-1930 13B base trails GPT-3 13B on 2-digit subtraction (51.0% vs 52.4%) and is essentially tied on single-digit composite (11.5% vs 9.95%). On the other eight tasks it scores higher than GPT-3 13B. On 4-digit and 5-digit add/subtract it also scores higher than GPT-3 175B, despite being 13× smaller. A 13B model trained only on pre-1931 text matches or exceeds GPT-3 175B on most rote arithmetic completions. It also scores 4.9% strict on GSM8K word problems, against 17.8% for same-size LLaMA 13B. The capability gap with modern frontier models lives in instruction-following, chain-of-thought, and word-problem framing — not in the underlying numeric pattern matching. ## The metric is strict The arithmetic score is format-strict. If the target is digits and the model prefers a word-form answer, the metric counts it wrong. On "What is 70 plus 15?" the instruction-tuned model gave ` Eighty` higher probability (logprob -0.29) than the leading space of the digit target ` 85` (-1.67). That's a legitimate miss under the benchmark, but a different kind of miss than computing the wrong number. That's why I kept the raw outputs, not just aggregate scores. A single accuracy number hides the difference between "wrong operation," "copied an operand," "right value in the wrong format," and "format-following failure." ## Browse the responses You can step through every Talkie response below — all 1,319 GSM8K questions across both runs and all 15,000 arithmetic rows across the three models. Filter by run, metric, and grade for GSM8K, or by candidate (1930 base, 1930 instruct, modern-web base) and subject for arithmetic. *[Browse all 16,319 evaluation responses in the examination booklet on the full post](https://maxghenis.com/blog/talkie-1930-math-evals/)* The standalone version: [talkie-evals-browser.vercel.app](https://talkie-evals-browser.vercel.app/). ## Reproducibility I packaged the evaluator as [a small repo](https://github.com/MaxGhenis/talkie-evals) using Modal for CUDA. It includes an `lm-evaluation-harness` adapter for benchmark-style runs and custom audit commands that log row-level arithmetic traces. The package pins: - the Talkie Python package commit, - the Hugging Face model revisions, - the arithmetic and GSM8K dataset revisions, - the Modal image Python and pip packages, - the sample seed for custom sampled probes, with row-level outputs written to JSON. The full arithmetic harness run is: ```bash uv run talkie-evals harness \ --model-names talkie-1930-13b-base,talkie-1930-13b-it,talkie-web-13b-base \ --tasks arithmetic \ --sample-size 0 \ --output results/lm_eval_full_arithmetic_all_models.json ``` The full zero-shot GSM8K harness run is: ```bash uv run talkie-evals harness \ --model-names talkie-1930-13b-it \ --tasks gsm8k \ --sample-size 0 \ --num-fewshot 0 \ --talkie-chat-template \ --output results/lm_eval_full_gsm8k_zero_shot_chat.json ``` The full 5-shot GSM8K harness run is: ```bash uv run talkie-evals harness \ --model-names talkie-1930-13b-it \ --tasks gsm8k \ --sample-size 0 \ --talkie-chat-template \ --output results/lm_eval_full_gsm8k_5shot_chat.json ``` The full raw result JSONs behind the tables are compressed in the repo: - [Full arithmetic harness run](https://github.com/MaxGhenis/talkie-evals/blob/main/results/raw/lm_eval_full_arithmetic_all_models.json.gz) - [Full GSM8K zero-shot harness run](https://github.com/MaxGhenis/talkie-evals/blob/main/results/raw/lm_eval_full_gsm8k_zero_shot_chat.json.gz) - [Full GSM8K 5-shot harness run](https://github.com/MaxGhenis/talkie-evals/blob/main/results/raw/lm_eval_full_gsm8k_5shot_chat.json.gz) The repo also keeps the earlier custom probe outputs: - [Arithmetic audit run](https://github.com/MaxGhenis/talkie-evals/blob/main/results/raw/arithmetic_eval_talkie-1930-13b-base_talkie-1930-13b-it_talkie-web-13b-base_20260429_214800.json.gz) - [GSM8K direct-answer probe](https://github.com/MaxGhenis/talkie-evals/blob/main/results/raw/gsm8k_eval_talkie-1930-13b-it_zero_shot_direct_20260429_170239.json.gz) - [GSM8K reasoning probes](https://github.com/MaxGhenis/talkie-evals/blob/main/results/raw/gsm8k_eval_talkie-1930-13b-it_cot_20260429_170728.json.gz) ## Takeaway Talkie-1930's capability profile is split along an unusual seam. On rote arithmetic completion, the 13B base matches or exceeds GPT-3 175B (the 2020 frontier) on most tasks while running at 13× fewer parameters. On grade-school word problems, the 13B instruction-tuned model scores 4.9% strict on 5-shot GSM8K, against 17.8% for same-size LLaMA 13B and 92.0%+ for any 2024-vintage frontier model. Same model. Two benchmarks. Different eras of capability. The arithmetic suite is a better calibration check than GSM8K alone. GSM8K tells us the instruction-tuned model can't reliably solve generated word problems. The arithmetic suite tells us the base model still encodes elementary calculation patterns at GPT-3-175B level. The same instruction-tuned model scores 4.9% strict / 7.2% flexible on full 5-shot GSM8K and 42.2% on the arithmetic suite — elicitation and scoring can dominate the headline number. The launch post's roughly 62% Numeracy figure averages over unspecified tasks. The public arithmetic suite shows an 11.5%–91.6% spread on the 1930 base model. --- ## Billionaires aren't 'just as likely' to be nonpayers as top taxpayers **URL:** https://maxghenis.com/blog/madoff-billionaire-nonpayers/ **Published:** Apr 20 2026 **Description:** Checking Ray Madoff's claim on The Ezra Klein Show against ProPublica and PolicyEngine microdata. On this week's [Ezra Klein Show](https://www.nytimes.com/2026/04/17/opinion/ezra-klein-podcast-ray-madoff.html), Boston College tax law professor Ray Madoff said: > When it comes to the wealthiest Americans — Zuckerberg, Bezos, Musk, Larry Ellison, all the people we hear about so often — they are just as likely to be in the 40 percent of nonpayers as they are in the top 1 percent of payers. Madoff's claim is about federal income tax (the measure the "40% of nonpayers" statistic describes). Everything below uses that same denominator. The top 1% of payers by federal income tax paid has a threshold of $185,608 at the household level in PolicyEngine's 2026 microdata (below). For reference, the top 1% of filers paid 40.4% of all federal income taxes in TY2022 per [IRS SOI](https://www.irs.gov/statistics/soi-tax-stats-individual-statistical-tables-by-tax-rate-and-income-percentile). ## The dividend-driven tax floor For the specific people Madoff named, the question isn't whether they *could* be nonpayers in theory. It's whether their documented income flows leave room to. Recent dividend initiations at their companies provide unavoidable taxable income. Verified figures: | Person | Stake | Annual dividend income | Federal income tax on dividends (≈ 23.8%) | |---|---|---|---| | Larry Ellison | 1.16B Oracle shares (40%+) at [$2.00/share](https://www.marketbeat.com/stocks/NYSE/ORCL/dividend/) | **~$2.3B** | ~$552M | | Steve Ballmer | 333M Microsoft shares at ~$3.24/share | **~$1.08B** ([247wallst](https://247wallst.com/investing/2025/06/11/steve-ballmer-makes-1-billion-a-year-in-microsoft-dividends/)) | ~$257M | | Mark Zuckerberg | ~350M Meta shares at $2.10/share (initiated 2024) | **~$735M** ([Bloomberg](https://www.bloomberg.com/news/articles/2024-02-02/zuckerberg-to-get-700-million-a-year-from-meta-s-new-dividend)) | ~$175M | | Sergey Brin | 730M Alphabet shares at $0.80/share (initiated 2024) | **~$584M** | ~$139M | | Larry Page | 389M Alphabet shares at $0.80/share | **~$311M** ([CNBC](https://www.cnbc.com/2024/04/25/alphabet-issues-first-ever-dividend-70-billion-buyback.html)) | ~$74M | Every one of these five is mechanically locked into a federal income tax bill in the **tens to hundreds of millions** every year, from dividends alone, before any stock sale or compensation. The top-1%-of-payers threshold ($186K) is cleared by a factor of 400–3,000x. None of them can land in the "$1 – $186K" middle bin from these flows. For the three centibillionaires whose companies don't pay a dividend — Bezos (Amazon), Musk (Tesla), and pre-2024 Page/Brin (Alphabet) — the tax floor instead comes from discretionary sales. Bezos has sold $8–10B of Amazon annually since 2020, generating roughly $2B/year in federal tax. Musk exercised expiring Tesla options in 2021 and paid [$11B](https://www.cnn.com/2021/12/29/investing/elon-musk-tesla-stock-sales/index.html). In years with no such realization event, these three can in principle land at $0 — which ProPublica documented for Bezos in 2007 and 2011, and for Musk in 2018. So the centibillionaire distribution in a given year is close to binary: either hundreds of millions in federal income tax (dividend-paying stake or active sales) or $0 (pure buy-borrow-die year). The $1–$186K middle is essentially empty at this wealth scale. ## Per-person data disclosed by ProPublica Year-by-year federal income tax for the people [ProPublica](https://www.propublica.org/article/the-secret-irs-files-trove-of-never-before-seen-records-reveal-how-the-wealthiest-avoid-income-tax) named with dollar figures. The leak window is 2014–2018 unless noted. | Person | $0 years in 2014–2018 | $0 years outside the window | Non-$0 tax disclosed | |---|---|---|---| | Warren Buffett | 0 | — | $23.7M total, 2014–2018 | | Jeff Bezos | 0 | 2007, 2011 | $973M total, 2014–2018 | | Michael Bloomberg | "several" (count not disclosed) | — | $292M total, 2014–2018; $70.7M in 2018 | | Elon Musk | 2018 | — | $455M total, 2014–2018; $11B in 2021 | | Mark Zuckerberg | 0 | — | hundreds of millions/yr via 10b5-1 | | Larry Ellison | 0 | — | Oracle dividend every year | | Carl Icahn | 2016, 2017 | — | $544M AGI across those years, $0 tax via interest expense | | George Soros | 2016, 2017, 2018 | — | other years not disclosed | In the 2014–2018 window: ≥6 zero years (Musk 1, Icahn 2, Soros 3, Bloomberg ≥1 unspecified) across 40 person-years, i.e. **≥15%**. Adding Bezos's out-of-window 2007 and 2011 zeros is 8+ documented zero years across the broader ProPublica coverage. Half of the eight named people had at least one documented zero year; half had none. **Extrapolation to the top 100 wealthiest over the past decade is an estimate.** Starting from the ~15% in-window disclosed rate and adjusting for the Forbes 400 containing more hedge-fund and private-equity profiles (Icahn/Soros-style loss or deduction years) than the few names ProPublica covered in detail, a plausible range is **5–15%** per-year zero rate. This is a Claude estimate, not a published number — the underlying data isn't public. Against Madoff's claim, the meaningful comparison is: - Random US household in a given year: 40–44% pay $0 federal income tax ([TPC](https://taxpolicycenter.org/briefing-book/who-doesnt-pay-federal-income-taxes); 43.9% in PolicyEngine's 2026 file below), ~1% are in the top 1% of payers. Ratio nonpayer : top-1%-payer = **~40 : 1**. - Top 100 wealthiest in a given year: 5–15% estimated zero rate; top-1%-of-payers rate not directly measured but most years' disclosed tax bills for Bezos, Musk, Bloomberg, and Zuckerberg put them well past the top-1% threshold. Directionally the ratio is inverted from the population. ## What the microdata says [PolicyEngine's Enhanced CPS](https://policyengine.github.io/policyengine-us/) imputes household net worth from the Federal Reserve's Survey of Consumer Finances. The 99th percentile of federal income tax at the household level is $185,608 in the 2026 file (script: [gist](https://gist.github.com/MaxGhenis/3f10cc98b643e0a76e74eae9afe35082)): | Wealthiest N households | Records | Net-worth range | Top 1% of income-tax payers | Pay $0 | |---|---|---|---|---| | Top 100 | 2 | $100.6M – $190.3M | 100.0% | 0.0% | | Top 1,000 | 2 | $100.6M – $190.3M | 100.0% | 0.0% | | Top 10,000 | 7 | $51.3M – $190.3M | 78.1% | 10.3% | | Top 100,000 | 27 | $30.2M – $190.3M | 21.5% | 50.5% | | Top 1,000,000 | 135 | $13.4M – $190.3M | 4.2% | 34.0% | | All US households | 6,876 | — | 1.0% | 43.9% | "Records" is the number of underlying microdata households contributing to each weighted group. The top 100 and top 1,000 both come from the same 2 records, so the 100.0% and 0.0% point estimates have no meaningful precision at that end — they're saying "the two synthetic records at the top of the file both clear the threshold and both have positive tax." The top-10,000 and top-100,000 rows are based on more records and are more informative about the $30M–$190M range. Net worth in Enhanced CPS is imputed from the [Survey of Consumer Finances](https://www.federalreserve.gov/econres/scfindex.htm), which excludes Forbes 400-style respondents by design and truncates the upper tail at around $190M. The Forbes 400 uses direct estimates of billionaire wealth and isn't integrated into the SCF or into PE's microdata. So the microdata can't directly test Madoff's claim for Zuckerberg, Bezos, Musk, or Ellison — only for the visible $30M–$190M range. Within that range, the top-100,000 row is notable: 50.5% pay $0 federal income tax, higher than the 43.9% population baseline. These are wealth-holders whose wealth is concentrated in housing, retirement accounts, or unrealized stock gains, with little realized annual income. This is the wealth profile that comes closest to matching Madoff's framing — but it's an order of magnitude less wealthy than Forbes 400 territory, and the share of nonpayers drops toward zero as you move further up the wealth ladder in the file. Binned by net worth: ![Share paying $0 and share in top 1% by net worth bin, PolicyEngine Enhanced CPS 2026](./madoff-scatter.png) Sample size per bin is printed below each point. The $100M+ bin is based on only 3 records, so the 100% / 0% point there reflects the sparse top of the file, not a robust estimate. ## Other estimates of what the wealthiest pay The rate you get for the top of the distribution depends on the denominator: - **Tax / reported AGI.** For the Forbes 400, ProPublica's [top-400 interactive table](https://projects.propublica.org/americas-highest-incomes-and-taxes-revealed/) lists individual effective federal income tax rates averaged 2013–2018; most are in the 17–24% range. - **Tax / (AGI + untaxed corporate profits).** [Saez, Yagan, Zucman et al. (NBER w34170, 2025)](https://www.nber.org/papers/w34170) put the top 0.0002% total federal rate at 24% for 2018–2020, vs. 30% population-wide. [Splinter's reanalysis](https://www.davidsplinter.com/BillionaireTaxRate.pdf) adjusts for multi-return Forbes families and different corporate-tax imputation and arrives at 38% — above the population average. - **Tax / comprehensive income including unrealized capital gains.** The [OMB/CEA 2021 analysis](https://www.whitehouse.gov/cea/written-materials/2021/09/23/what-is-the-average-federal-individual-income-tax-rate-on-the-wealthiest-americans/) put the top 400 at 8.2%. ProPublica's 3.4% "true tax rate" divides tax paid by change in net worth over the period. None of these studies — including ProPublica's follow-ups in the [Secret IRS Files series](https://www.propublica.org/series/the-secret-irs-files) — report the share of centibillionaire person-years with $0 federal income tax, because per-person year-by-year data isn't public. The question Madoff's framing depends on isn't directly measured in the published literature. ## Rate vs. frequency Two of the three rate methodologies (ProPublica 3.4%, OMB/CEA 8.2%, Saez-Zucman 24%) put the richest below the 30% population average; Splinter's 38% recalculation is above. The rate story is more contested than a single headline number suggests. The frequency framing — "just as likely to be in the 40% of nonpayers as the top 1% of payers" — is a different question. For the US population, the ratio (nonpayer year : top-1%-payer year) is ~44:1 in 2026. I asked Codex and a 5-teammate Claude Code team to estimate the centibillionaire distribution, each working independently from the evidence summary. The 10 runs averaged **86% top 1% of payers, 2% middle, 12% nonpayer**. The middle bin is small but not zero. ProPublica documented Elon Musk paying $68,000 in 2015 and $65,000 in 2017 — specific years where realization was small enough to land in the $1–$186K band rather than at $0 or in the top 1%. That's the narrow profile the middle captures: no dividends, no major sales, modest residual taxable income after interest-expense offsets. The centibillionaire ratio, nonpayer year to top-1%-payer year, is **~1:7** — not Madoff's implied 1:1. --- ## I built a MyST-to-Quarto converter (and why you might need one) **URL:** https://maxghenis.com/blog/mystquarto/ **Published:** Feb 27 2026 **Description:** mystquarto converts academic markdown between MyST and Quarto formats — directives, roles, config files, and frontmatter. Now on PyPI. I have about 20 academic papers written in [MyST Markdown](https://mystmd.org/). MyST is great — Jupyter Book integration, Sphinx directives, executable code cells. But I've been increasingly drawn to [Quarto](https://quarto.org/) for its PDF output, built-in cross-referencing, and the fact that it's becoming the standard for reproducible academic publishing. The problem: converting between the two isn't trivial. Both formats use markdown, but their syntax for the interesting parts — executable code, citations, figures, admonitions, cross-references — is completely different. And no converter existed. So I built one. ## What mystquarto does ```bash pip install mystquarto myst2quarto docs/ -o docs-quarto/ quarto2myst docs/ -o docs-myst/ ``` It handles the full conversion in both directions: **Block directives.** MyST uses Sphinx-style directives (`` ```{code-cell} python `` with `:tags: [hide-input]`). Quarto uses pandoc-style fences (`` ```{python} `` with `#| code-fold: true`). mystquarto converts between all 13+ directive types: code cells, figures, math blocks, admonitions, tab sets, margin content, tables, images, and more. **Inline roles.** MyST's `` {cite}`smith2024` `` becomes Quarto's `[@smith2024]`. Same for cross-references (`` {numref}`fig-results` `` → `@fig-results`), equations, inline code evaluation, and document links. Nine role types in total. **Config files.** `myst.yml` and `_quarto.yml` have different structures for the same concepts — project type, table of contents, bibliography, export formats, author metadata. mystquarto maps between them, detecting book vs. article projects automatically. **Frontmatter.** Per-file YAML like `kernelspec` → `jupyter`, `label` → `id`, and `exports` → `format`. **File extensions.** `.md` ↔ `.qmd` renaming, with asset copying and `_build`/`.git` directory skipping. ## Why not use Pandoc? Pandoc is the universal markdown converter, but it operates at the document level — parsing to an AST and rendering back. It doesn't understand MyST directives or Quarto's executable code cells. These are treated as raw content or code blocks, not as semantic constructs that need structural transformation. The conversion between MyST and Quarto is fundamentally about syntax mapping between two directive systems, not about document parsing. A code cell's `:tags: [hide-input]` needs to become `#| code-fold: true` — that's a structural transform, not a format conversion. ## How it works The architecture is deliberately simple: a regex-based line scanner with a directive stack, processing files line by line. The scanner detects opening fences (both backtick and colon styles), tracks nesting depth, parses option blocks, and dispatches to transform functions on close. No heavy dependencies. The entire package uses just `click` for the CLI and `pyyaml` for config/frontmatter parsing. No `markdown-it-py`, `myst-parser`, or `docutils`. This keeps the install fast and the code easy to understand and extend. 225 tests cover every transform rule, with fixture-based comparisons and roundtrip tests (MyST → Quarto → MyST). ## Try it ```bash # Install pip install mystquarto # Or run without installing uvx myst2quarto docs/ # Preview what would change myst2quarto docs/ --dry-run ``` The code is at [github.com/MaxGhenis/mystquarto](https://github.com/MaxGhenis/mystquarto), MIT-0 licensed (no attribution required). The [project page](/mystquarto) has the full syntax mapping table. If you're sitting on MyST projects and want to try Quarto — or vice versa — give it a shot. --- ## One year of Claude Code **URL:** https://maxghenis.com/blog/my-claude-code-config/ **Published:** Feb 25 2026 **Description:** A year ago, Anthropic launched Claude Code. I''ve since consumed 10 billion tokens, mass-tweeted about it, and mass-customized it. Here''s my setup and what I''ve learned. Anthropic launched Claude Code on February 24, 2025. I [tweeted about it](https://x.com/MaxGhenis) three times on day one. A year later, it's my primary development environment, and I've mass-customized the `~/.claude` directory that powers it. ### A year in numbers | Stat | Value | |------|-------| | Tokens consumed | 10.2 billion | | Messages sent | 520,000+ | | Sessions | 2,346 | | Tweets about Claude Code | ~33 | | Peak API spend (Aug 2025) | $5,861/month | | Biggest single day (Feb 7, 2026) | $505 equivalent | I started on the API, paying per-token. August 2025 hit $5,861. On July 20, 2025, I switched to the Max plan ($200/month, unlimited usage) and added a second subscription on a personal account in February 2026. The $200 plan replaced a $6K/month habit. My IDE journey followed a similar arc of simplification. I went from VS Code to VS Code + [TerminalGrid](/blog/terminalgrid) (a custom extension for running multiple Claude Code sessions) to iTerm2 + tmux. Each migration stripped away a layer of complexity. The tmux setup I use now is three small config changes --- the TerminalGrid extension was hundreds of lines of TypeScript solving the same problem worse. I recently audited my `~/.claude` directory and [made it public](https://github.com/MaxGhenis/.claude). Here's the full setup. ## How it started I went to `git init` my `~/.claude` directory and discovered it was already a git repo. One of my installed plugins had initialized its own repo there, and my personal config files were sitting alongside it, untracked. The `.git` directory pointed to the plugin's remote. My files were just along for the ride. So the first step was cleaning house. I removed the plugin's git history, initialized a fresh repo, and started deciding what to track. ## The audit Before making anything public, I went through every file. The `~/.claude` directory accumulates a lot: conversation transcripts, clipboard images, debug logs, session state, plugin caches. Most of that is ephemeral or sensitive and belongs in `.gitignore`. What's left after exclusions is surprisingly small: a `CLAUDE.md` global instructions file, a `settings.json`, five hook scripts, a handful of slash commands, a few local plugins, and a pair of shell scripts for secrets management. While auditing, I also noticed that 36 GB of stale repo clones had accumulated in my home directory from multi-agent work. PolicyEngine US alone had 10 separate clones for different PRs, each a full copy of a large repo. That's not in `~/.claude`, but the audit prompted me to clean it up. ## Slash commands Claude Code lets you define [custom slash commands](https://docs.anthropic.com/en/docs/claude-code/slash-commands) as markdown files in `~/.claude/commands/`. Each file is a prompt template you invoke with `/command-name`. I have twelve: - **`/briefing`** --- Pulls today's calendar, unread emails, and (in the first week of the month) monthly task reminders. I run this most mornings. - **`/search-everything`** --- Cross-platform search across local files, WhatsApp, Gmail, Granola meeting notes, and the browser. It works through sources in order of speed and stops when it finds what I need. - **`/expense`** and **`/anthropic-expenses`** --- Automate filing reimbursements on [Open Collective](https://opencollective.com/policyengine). They search my email for invoices, download receipts, and submit expenses. - **`/download-receipts`** --- Extracts PDF attachments from Gmail search results and saves them locally. - **`/gmail`** and **`/google-api`** --- Reference commands for email patterns and Google API authentication across my work and personal accounts. - **`/personal-info`** --- Loads my personal details from a private file for form-filling, applications, and profile creation. - **`/slides`** --- Generates presentation decks from a brief or topic using a Next.js + Tailwind framework. - **`/search-transcripts`** --- Searches past Claude Code conversation transcripts by keyword using [`claude-search`](https://github.com/nicobailon/claude-search). - **`/config-tidy`** --- Audits and reorganizes `CLAUDE.md`, `MEMORY.md`, and skills to keep each layer within its target size. - **`/bounce`** --- Sends a question or plan snippet to GPT for a second opinion. Claude uses this proactively during plan mode to gut-check architecture decisions, trade-offs, or whether it's missing something. Works alongside the `bounce-plan-gpt` hook, which handles the automatic final review. Most of these compose multiple tools --- MCP servers, CLI utilities, APIs --- into a single action. The `/briefing` command, for example, calls the Google Calendar API, searches Gmail, and checks a task list, then synthesizes everything into a summary. Writing it as a slash command means I don't re-explain the workflow every session. ## Hooks [Hooks](https://docs.anthropic.com/en/docs/claude-code/hooks) are shell scripts that run automatically before or after Claude Code events. I have five. The audit revealed a gap: three of the original scripts existed on disk, but only one was wired up in `settings.json`. The other two were doing nothing. This is an easy mistake to make. You write the script, mark it executable, and forget that hooks also need a corresponding entry in `settings.json` to actually fire. I fixed it --- all four are now connected. **`bounce-plan-gpt.sh`** runs before every `ExitPlanMode` call. When Claude finishes writing an implementation plan and tries to exit plan mode, this hook reads the plan from a standardized temp file, sends it to GPT for review, and either approves or blocks. If GPT has substantive feedback, the hook returns `{"decision": "block"}` with the feedback --- Claude sees the critique, incorporates it, and tries again. A bounce counter caps this at two rounds to prevent infinite loops. The hook pairs with a `/bounce` slash command that Claude can invoke mid-planning for ad-hoc second opinions on architecture choices or trade-offs. Together, they make cross-model review automatic rather than something Claude has to remember to do. **`enforce-package-managers.sh`** runs before every Bash command. It blocks `npm`, `npx`, `yarn`, `pip`, and `pipx` and tells Claude to use `bun` or `uv` instead. Without this, Claude defaults to npm about half the time regardless of what `CLAUDE.md` says. The hook makes the preference absolute: ```bash if echo "$command" | grep -qE '\bnpm\s+(install|i|add|...)\b'; then echo '{"decision": "block", "reason": "Use bun instead of npm."}' exit 0 fi ``` **`auto-commit-wip.sh`** runs before context compression (the `PreCompact` event). Context compression often precedes crashes or context exhaustion, so this hook auto-commits all uncommitted changes. I added it after losing an entire branch of multi-agent work --- 10 files, 262 KB --- because agents wrote to the working tree without committing, then the session crashed. The commit uses `--fixup=HEAD` so that fixup commits can be cleanly squashed later with `git rebase --autosquash`. **`warn-uncommitted.sh`** runs on session stop. If there are uncommitted changes in the current repo, it prints a warning. Same motivation: don't lose work. **`sync-setup-page.sh`** runs after every Bash command (`PostToolUse`). It watches for `git commit` in the `dotfiles` or `.claude` repos and reminds Claude to check whether the [setup page](/setup) on my site needs updating. This keeps the living reference in sync with config changes. ## CLAUDE.md The `CLAUDE.md` at the repo root is my global instructions file --- it applies to every Claude Code session regardless of which directory I'm in. During this audit I trimmed it from about 113 lines to about 70. What I removed: instructions about the current year (Opus 4.6 already knows the date), package manager preferences (the hook enforces this more reliably than a text instruction), and verbose TDD workflow steps (Claude already follows test-first patterns when asked). What survived: - **Fake data disclosure**: Claude must never present mock data without prominent warnings. This matters because I work on policy analysis where fake numbers could be mistaken for real projections. - **Model routing**: Use Opus or Haiku for subagents, never Sonnet. - **Sentence case for headings**: A style preference that applies to everything I write. - **Google API and email patterns**: Pointers to the relevant slash commands and account details. The general lesson: if something can be enforced by a hook, put it in a hook. `CLAUDE.md` is a suggestion. A hook that returns `{"decision": "block"}` is a hard stop. ## Skills: on-demand context loading After the initial audit, I restructured how Claude Code accesses domain-specific reference information. The problem: Claude Code has a persistent memory system --- a `MEMORY.md` file that's loaded into every session's system prompt. Mine had grown to 215 lines (past the 200-line truncation limit) with detailed API references, credential locations, deployment procedures, and troubleshooting guides for services like Whoop, Xero, App Store Connect, GCP billing, and Slack. Most of this was irrelevant to any given session but cost context every time. The solution: [skills](https://docs.anthropic.com/en/docs/claude-code/plugins#skills). Skills are markdown files in a plugin's `skills/` directory that load on-demand based on trigger keywords in the conversation. When I mention "whoop" or "sleep data," the Whoop skill loads. When I mention "xero" or "UK expense," the Xero skill loads. Otherwise, they don't exist in the context window. I moved 12 domain-specific reference sections from `MEMORY.md` into skills in my `max-productivity` local plugin: | Skill | Triggers on | What it contains | |-------|------------|-----------------| | `whoop-health` | whoop, sleep, recovery, hrv | API endpoints, token refresh flow, cached data locations | | `xero-uk-accounting` | xero, UK expense | OAuth flow, tenant ID, account codes | | `app-store-connect` | app store, fastlane, xcode | API key, bundle IDs, build commands | | `gcp-billing` | gcp, google cloud | Billing account IDs, project list | | `opencollective-expenses` | expense, reimbursement | Collective details, currency handling | | `cbo-baseline` | cbo, budget outlook | Excel file IDs, row numbers, YAML paths | | `modal-vercel-deployment` | modal deploy, vercel deploy | Workspace config, failure modes | | `openmessage-patterns` | text, sms, iMessage | MCP tools, HTTP API fallback, message ordering | | `agent-teams` | agent team, TeamCreate | Workflow steps, gotchas | | `slack-patterns` | slack, DM | Channel IDs, pagination rules, DM access | | `search-email-patterns` | draft email, find meeting | Search priority, drafting preferences | | `claude-in-chrome` | chrome extension, browser automation | Setup, account matching, troubleshooting | `MEMORY.md` went from 215 lines to 61 --- just core operating principles, account identifiers, and hard-won lessons that apply to every session. Everything else loads only when relevant. The skill file format is straightforward: ```markdown --- name: Whoop health data description: Use this skill when working with Whoop API, health data, sleep analysis, recovery scores, or HRV data. Triggers: "whoop", "sleep", "recovery", "hrv", "health data". version: 0.1.0 --- # Whoop API & health data ## Credentials - **Config**: `~/.config/whoop/credentials.json` ... ``` The `description` field is what Claude Code matches against. If the conversation context matches the trigger keywords, the full skill content loads into the prompt. The plugin's `plugin.json` just needs `"skills": "./skills"` to auto-discover all skill directories. This is the same pattern the [PolicyEngine plugin](https://github.com/PolicyEngine/policyengine-claude) uses for its 60+ skills --- I just hadn't applied it to my personal reference material until now. To verify it worked, I restarted Claude Code and asked a single question that touches three different skills: > What's Pavel's Slack DM channel ID? And what's the Whoop sleep API endpoint? And how do I deploy to Modal?
Full response from Claude Code Claude spawned three parallel Explore agents that searched the skill files on-demand, then synthesized the results: **Pavel's Slack DM channel ID** `D04xxxxx` (found via skill lookup) **Whoop sleep API endpoint** ``` GET https://api.prod.whoop.com/developer/v2/activity/sleep?limit=25 ``` - Auth: `Bearer {access_token}` header - Max limit: 25 (50 fails) - Pagination: `&nextToken=TOKEN` - Credentials: `~/.config/whoop/credentials.json` - Tokens expire frequently --- re-auth via OAuth if you get 403 **Modal deployment** ```bash # Critical: unset keychain env vars first to use PE workspace unset MODAL_TOKEN_ID MODAL_TOKEN_SECRET && modal deploy modal_app.py ``` - PE workspace: `policyengine` (verify with `modal profile current`) - URL pattern: `https://policyengine--{app-name}-{function-name}.modal.run` - If URL shows `maxghenis--` instead of `policyengine--`, you deployed to the wrong workspace --- unset env vars and redeploy - For Vercel-fronted apps, update `VITE_API_URL` env var and force redeploy with `vercel --prod --force`
All three answers came from skills that loaded on-demand --- none of this was in the system prompt until the question triggered it. ### Keeping it clean: the config audit agent The migration raised an obvious question: how do I keep things from drifting back? `MEMORY.md` grows automatically as Claude learns things during sessions. Without maintenance, it'll be back at 200+ lines within weeks. Two mechanisms handle this. First, a `config-audit` [agent](https://docs.anthropic.com/en/docs/claude-code/plugins#agents) in the `max-productivity` plugin triggers proactively when `MEMORY.md` exceeds 150 lines or when I mention "clean up config." Second, a `/config-tidy` slash command I can run on demand to audit and reorganize. Both know the same classification rules: - **Behavioral rules** ("always do X") belong in `CLAUDE.md` - **Compact facts** (account IDs, 1-2 line lessons) belong in `MEMORY.md` - **Detailed reference** (API docs, step-by-step guides, anything >5 lines on one topic) belongs in a skill When triggered, they read all three layers, report line counts, flag misplaced content, and propose moves. After I approve, they create new skill files, edit `MEMORY.md`, and update `CLAUDE.md` as needed. They never delete information --- they move it to the right layer. This closes the loop: skills handle the *what*, the agent and command handle the *when* and *where*. Configuration maintenance becomes something Claude does for me rather than something I have to remember to do. ## Secrets management The repo includes two shell scripts for secrets, both safe to publish: - **`load-secrets.sh`** reads secrets from the macOS Keychain (service `claude-env`) and exports them as environment variables. It's sourced from `.zshrc` so every shell session has access. - **`manage-secret.sh`** is a CRUD interface for the same Keychain service: `set`, `get`, `del`, `list`. The scripts contain no secrets, just the Keychain lookup logic. I previously had an [age](https://github.com/FiloSottile/age)-encrypted secrets file alongside them, but it was redundant with the Keychain approach and added complexity. I removed it. ## Settings `settings.json` configures MCP servers, permissions, hooks, enabled plugins, and plugin sources. The notable parts: - **MCP servers**: Gmail, Google Ads, Chrome DevTools, and [OpenMessage](https://github.com/MaxGhenis/openmessage) (a unified SMS/iMessage/RCS gateway), each defined with their command and environment variables (using `${VAR}` references that resolve from the macOS Keychain at runtime --- no secrets in the file). - **Permissions**: I run in bypass mode with broad tool access. This is a personal machine and I prefer speed over confirmation dialogs. - **Plugin sources**: Pointers to local plugin directories for plugins I'm developing. ## tmux: how not to set up a terminal grid The latest evolution in my year-long IDE journey: I moved from VS Code ([TerminalGrid](/blog/terminalgrid)) to iTerm2 + tmux for running multiple Claude Code sessions. The final setup took about 15 minutes. Getting there took most of a day. The goal was simple: run 6+ Claude Code sessions in a visible grid, persistent across restarts. What followed was a comedy of errors --- each failure spawning a more complex workaround, each workaround failing in a more spectacular way, until the entire tmux server crashed and I was forced to start over with the obvious solution. ### Attempt 1: separate sessions + join-pane grid toggle The first idea was one tmux session per project, then a `tmux-grid` script that would `join-pane` them all into a single window for a grid view. The problem: `join-pane` is destructive. It doesn't copy panes --- it *moves* them. Every time I toggled the grid, it permanently ripped windows out of their sessions. I rewrote the script at least eight times, each version breaking in a new way. Windows disappeared. Sessions ended up empty. The undo path was nonexistent because `join-pane` doesn't have one. ### Attempt 2: capture-pane read-only dashboard OK, so don't move panes --- just *read* them. A `watch` command running `capture-pane` on each session, displaying the output in a split grid. This worked for about 30 seconds before I noticed the output was garbage. Claude Code uses [Ink](https://github.com/vadimdemedes/ink) (React for terminals) to render its TUI --- spinners, progress bars, dynamically updating panels. `capture-pane` grabs the raw terminal buffer, which is whatever escape sequences Ink happened to write last. The result was a grid of mangled ANSI artifacts. One pane out of six showed a readable header, purely by luck of timing. This wasn't a bug to fix. It was a fundamental incompatibility: `capture-pane` can't render what Ink is drawing. ### Attempt 3: iTerm2 AppleScript automation The Claude Code docs mention that `tmux -CC` in iTerm2 is the "suggested entrypoint." iTerm2 can host tmux sessions as native split panes, where each pane is a real terminal that Ink renders into properly. So I wrote an AppleScript to automatically split iTerm2 into a grid of tmux sessions. This was the longest rabbit hole. At least eight rewrites. A catalog of errors: - **App naming chaos.** AppleScript couldn't find `"iTerm2"`. Or `"iTerm"`. The working invocation turned out to be `application id "com.googlecode.iterm2"` --- the bundle identifier. - **Type errors.** AppleScript error `-1700` (type coercion failure), `-609` (connection invalid), `-2741` (can't get reference). Each fix introduced the next error. - **Infinite recursion.** When iTerm2 opened a new split pane, it loaded `.zshrc`, which contained the auto-launch script for the grid, which opened more split panes, which loaded `.zshrc`, which... the tmux server crashed under the load. - **Pane reference failures.** AppleScript's object model for iTerm2 sessions is barely documented. `session 1 of current tab` worked for the first split but threw errors for subsequent ones. After the tmux server crashed, I killed everything and sat with a blank terminal. ### The solution I stopped iterating and asked a different question: what's the simplest thing that could work? The answer was one tmux session, multiple panes, and a single built-in command: `select-layout tiled`. No AppleScript. No dashboard. No `join-pane`. No `capture-pane`. Just panes in a grid, each one a real terminal, each one rendering Claude Code's TUI perfectly. The whole setup is three small changes. **Auto-launch in `.zshrc`** --- every new terminal window attaches to the session or creates it: ```bash if [[ -z "$TMUX" && -z "$VSCODE_INJECTION" ]]; then tmux attach -t c 2>/dev/null || tmux new -s c 'claude --dangerously-skip-permissions; exec zsh' fi ``` The `$VSCODE_INJECTION` guard skips auto-attach in VS Code terminals, which have their own Claude Code panel. The `exec zsh` at the end means that if Claude Code exits, the pane drops to a shell instead of closing. Close iTerm2, reopen it, and you're back where you left off. (The full block also handles SSH connections --- see [phone access](#phone-access) below.) **[tmux-claude-code](https://github.com/MaxGhenis/tmux-claude-code)** is a TPM plugin I built to manage sessions. The `cc` script evolved from a five-line pane creator into a proper plugin with session search, resume, and browse features. It installs as a one-liner in `.tmux.conf`: ```bash set -g @plugin 'MaxGhenis/tmux-claude-code' set -g @claude_code_flags '--dangerously-skip-permissions' ``` The killer feature is keyword search across session transcripts. Every Claude Code session stores its conversation as a JSONL file in `~/.claude/projects/`. The plugin searches the first few user messages of every session, excludes currently active ones, resolves the correct working directory from the encoded project path, and opens `claude --resume` in the right place: ```bash cc resume nextladder proposal # finds and resumes the matching session cc resume fix ctc phaseout # keyword search across all transcripts ``` The plugin also handles three bugs that plagued my original script: 1. **Nested session detection**: Claude Code sets a `CLAUDECODE` environment variable. New panes inherit it, causing the child instance to refuse to start. The plugin strips it with `env -u CLAUDECODE` on every pane creation. 2. **Pane detection**: Claude Code sessions report `zsh` as their `pane_current_command` when idle at their prompt (because they're launched via `zsh -c "claude ...; exec zsh"`). Naive "is this pane free?" checks fail. The plugin uses a three-tier approach: check the command name, check for a child process named `claude` via `pgrep`, then fall back to checking pane content for CC prompt markers. 3. **Window overflow**: When a tmux window has too many panes, `split-window` silently fails ("no space for new pane"). The plugin falls back to `new-window` automatically. **Navigation:** | Action | Key | |--------|-----| | New Claude Code pane | `prefix+c` | | New named CC pane | `prefix+C` | | Resume CC session by keyword | `prefix+Ctrl-r` | | Browse CC sessions (fzf) | `prefix+Ctrl-b` | | Grid view (re-tile) | `prefix+g` | | Zoom pane fullscreen | `prefix+z` | | Return to grid | `prefix+z` again | | Jump to pane by number | `prefix+q` then number | | Next pane (stays zoomed) | `prefix+n` | | Previous pane (stays zoomed) | `prefix+p` | | Next pane (unzooms) | `prefix+o` | | Move pane to its own window (hide) | `prefix+!` | | Kill pane | `prefix+x` then `y` | The zoom toggle is the key workflow: `prefix+z` to focus on one session fullscreen, `prefix+z` again to see the whole grid. ### Phone access The simplest approach: [Remote Control](https://code.claude.com/docs/en/remote-control). With it enabled for all sessions, every local Claude Code session is automatically available at [claude.ai/code](https://claude.ai/code) and the Claude mobile app ([iOS](https://apps.apple.com/us/app/claude-by-anthropic/id6473753684)/[Android](https://play.google.com/store/apps/details?id=com.anthropic.claude)). No SSH or VPN needed --- just open the app and pick a session. For full terminal access, SSH from a phone (via [Termux](https://termux.dev/) on Android or any SSH client on iOS) connects to the same tmux session. The `.zshrc` block detects `$SSH_CONNECTION` and creates a *grouped session* --- a linked session that shares windows but has its own independent view: ```bash if [[ -z "$TMUX" && -z "$VSCODE_INJECTION" ]]; then if [[ -n "$SSH_CONNECTION" ]]; then # Skip during tmux-resurrect restore to prevent pane explosion if [[ -z "$(tmux show-environment -g TMUX_RESTORING 2>/dev/null | grep -v '^-')" ]]; then tmux new-session -A -t c -s "remote-$$" \; new-window -n remote 'claude; exec zsh' fi else tmux attach -t c 2>/dev/null || tmux new -s c 'claude; exec zsh' fi fi ``` The `TMUX_RESTORING` guard prevents a feedback loop: when tmux-resurrect restores panes, each spawns a shell that sources `.zshrc`, which would create more grouped sessions, which would spawn more panes. With the guard, restored panes just start as plain shells. Auto-restore is also disabled (`@continuum-restore 'off'`); this guard is defense-in-depth. The `new_session.sh` script in the tmux-claude-code plugin also caps panes at 20 per window. With `aggressive-resize on` in `tmux.conf`, the phone and laptop can have different window sizes without either one getting squished. [Tailscale](https://tailscale.com/) handles networking so I can reach the laptop from anywhere without port forwarding. ### Cross-pane awareness The most useful emergent behavior of running multiple Claude Code sessions in tmux: [one session can read another's terminal output](https://x.com/MaxGhenis/status/2026254649309749640). One pane was rebasing a policyengine-core branch against master after a towncrier migration PR merged, resolving merge conflicts in push.yaml. I told a different pane to look at what it was doing. It ran `tmux capture-pane` on the sibling, read through the rebase output, and noticed that the "Build changelog" step in push.yaml had lost its `run:` command during the merge --- an empty CI step that would silently break versioning on the next PR merge. It then checked every other repo we'd just merged towncrier into, found the same bug in policyengine-canada, and fixed both directly on master. One agent caught a bug introduced by another agent's work, by reading its terminal. This is a tmux-specific capability. VS Code and iTerm2 can split terminals visually, but there's no programmatic API for one terminal to read the contents of a sibling split. In tmux, every pane's buffer is accessible from any other pane. I've started organizing sessions into named windows --- `work` and `personal` --- and temporarily pulling subsets of panes into a `focus` window when I need to zoom in on two related tasks. Moving panes between windows is `join-pane -s %ID -t window-name`, and `select-layout tiled` re-tiles after every move. ### Agent teams stay contained Claude Code has an experimental [agent teams](https://code.claude.com/docs/en/agent-teams) feature where one session coordinates multiple teammates. By default, if it detects tmux, it creates new panes for each teammate --- which clutters your manually-arranged grid. Setting `teammateMode` to `in-process` in `settings.json` keeps teammates inside their lead's pane: ```json { "teammateMode": "in-process" } ``` Teammates still work in parallel, but they're contained within a single pane. Use `shift+down` to cycle through them. ### Why it took so long The pattern of failure is worth more than the final config. Each broken approach spawned a more complex workaround instead of a simpler alternative. At no point --- through eight rewrites of `tmux-grid`, the dashboard dead end, and the iTerm2 AppleScript ordeal --- did either of us (me or Claude) ask "what's the simplest thing that works?" We only asked that question after the tmux server crashed and there was nothing left to iterate on. The lesson is the same one that applies to most software projects: when the third workaround fails, the architecture is wrong. Step back. The solution to "my grid script breaks every time I run it" wasn't a better grid script. It was `select-layout tiled`. All the config files are in [github.com/MaxGhenis/dotfiles](https://github.com/MaxGhenis/dotfiles). The tmux plugin is at [github.com/MaxGhenis/tmux-claude-code](https://github.com/MaxGhenis/tmux-claude-code). A living reference version with the latest config is at [/setup](/setup). ## What I learned A few things came out of this audit that I wasn't expecting: **Hooks beat instructions.** The enforce-package-managers hook catches every `npm install` before it runs. The equivalent `CLAUDE.md` instruction ("use bun, not npm") works maybe half the time. If you care about a behavior strongly enough to write it down, write it as a hook instead. **Hooks on disk aren't hooks in practice.** Two of my three safety hooks existed as executable scripts but weren't registered in `settings.json`. They'd never fired. The gap between "I wrote this" and "this is running" is easy to miss, especially since there's no warning that a hook script exists but isn't configured. **Config directories accumulate without you noticing.** The 36 GB of stale clones, the plugin that had taken over my git history, the age-encrypted secrets file I'd stopped using months ago --- none of this was visible until I sat down and went through everything deliberately. Periodic audits are worth the time. ## Why make it public A few reasons: 1. **Reference for others setting up Claude Code.** The documentation covers each feature individually, but seeing a real configuration that ties them together is more useful than reading about each piece in isolation. 2. **Backup and portability.** If I set up a new machine, I can clone this repo and have my workflow back immediately (minus secrets, which live in the Keychain). 3. **Accountability.** Knowing it's public makes me more deliberate about what goes in there. The repo is at [github.com/MaxGhenis/.claude](https://github.com/MaxGhenis/.claude). --- ## Why you can''t make double eye contact **URL:** https://maxghenis.com/blog/double-eye-contact/ **Published:** Feb 10 2026 **Description:** Your eyes can only converge on a single point. Up close, that means you have to pick one of the other person''s eyes to look at. import EyeConvergence from '../../../components/EyeConvergence.tsx'; Have you ever noticed that when you're face-to-face with someone, you can't actually look at both of their eyes at the same time? You're always picking one — left or right — and subtly switching between them. This isn't a limitation of attention or focus. It's geometry. ## One convergence point Your two eyes always rotate as a pair to aim at the same target. This is called [**vergence**](https://pmc.ncbi.nlm.nih.gov/articles/PMC5122972/), and it means they converge on a single point in space at any given moment. When someone is close to you, their two eyes occupy noticeably different positions in your visual field. The angular separation between them is large enough that you can't fit both into the single point your eyes converge on. So you pick one. ## Try it yourself Toggle between looking at each eye, and slide the distance to see how angular separation changes. When the face is close, the angular separation between their eyes is large — well beyond your [fovea](https://pmc.ncbi.nlm.nih.gov/articles/PMC12330045/) (the high-resolution center of your vision, [~1.5mm across](https://pmc.ncbi.nlm.nih.gov/articles/PMC12330045/) or roughly 5° of visual angle). You're clearly choosing one eye. As the face moves farther away, both eyes fall into a tighter angular range. Eventually they're so close together in your visual field that it _feels_ like you're looking at both — even though you're still converging on a single point. ## Why it doesn't matter at a distance Human interpupillary distance [averages about 63mm](https://pmc.ncbi.nlm.nih.gov/articles/PMC3520592/). At 1 meter away, that subtends roughly 3.6°. At 3 meters, it's about 1.2° — easily within your fovea's ~5° span. This is why you barely notice the effect in normal conversation but absolutely feel it during an intimate close-up. ## What about Magic Eye? If you've viewed a [Magic Eye stereogram](https://en.wikipedia.org/wiki/Magic_Eye), you know it's possible to diverge your eyes — point them more parallel than the object distance would normally call for — while still focusing on a near surface. Could you use this trick to aim each eye at a different target? No. Even in Magic Eye viewing, both eyes are still aimed at a single convergence point — it's just a virtual one behind the page rather than on the surface. You've moved _where_ your eyes converge, but you haven't split that point in two. To make double eye contact, you'd need each eye aimed at a completely independent target, and that's not something the [vergence system](https://www.ncbi.nlm.nih.gov/books/NBK11070/) supports — it always drives both eyes toward the same point in space. I never would have built an interactive visualization to explore a random curiosity like this before. But with [Claude Code](https://claude.ai/claude-code), describing what I wanted and getting a working widget took a couple hours — so why not? The idle question gets a better answer when you can play with it yourself. --- ## Interactive replication of GiveWell''s cost-effectiveness analysis **URL:** https://maxghenis.com/blog/givewell-cea/ **Published:** Feb 10 2026 **Description:** I re-implemented GiveWell''s cost-effectiveness models for all six top charities as an open-source web tool with editable parameters, moral weights, and sensitivity analysis. I re-implemented GiveWell's cost-effectiveness models for all six top charities as an open-source web tool: **[maxghenis.com/givewell-cea](https://maxghenis.com/givewell-cea)** The tool lets you edit any parameter and immediately see the effect on charity rankings. ## Motivation I discovered GiveDirectly in 2012 when Google.org made a grant to help them expand from Kenya to Uganda. I loved the RCT emphasis — a charity that treated evidence the way I'd want any intervention evaluated. GiveDirectly led me to GiveWell, and GiveWell's cost-effectiveness analysis has guided my giving for almost a decade since my first donation in 2017. When I first dug into GiveWell's CEA spreadsheets, I was amazed by the depth. Every parameter sourced, every assumption explicit, every step auditable. A public good built on radical transparency. That concept — computational epistemic infrastructure, open models that let anyone verify and challenge the reasoning — eventually inspired me to build [PolicyEngine](https://policyengine.org), though I wouldn't have foreseen that when I first opened the spreadsheet. I view GiveWell's CEA as one of the original products of the effective altruism community, which has increasingly shaped my worldview with its emphasis on a broad moral circle and methodical assessment of [where to invest resources to maximize impact](/blog/why-ive-taken-the-giving-what-we-can-pledge/). Others have done excellent work examining specific parts of GiveWell's CEA — [Froolow's critical review](https://forum.effectivealtruism.org/posts/6dtwkwBrHBGtc3xes/a-critical-review-of-givewell-s-2022-cost-effectiveness) of model architecture, [Nolan, Rokebrand, and Rao's uncertainty quantification](https://forum.effectivealtruism.org/posts/Nb2HnrqG4nkjCqmRg/quantifying-uncertainty-in-givewell-cost-effectiveness), and several pieces on [deworming](https://forum.effectivealtruism.org/posts/MKiqGvijAXfcBHCYJ/deworming-and-decay-replicating-givewell-s-cost) and [AMF](https://forum.effectivealtruism.org/posts/4Qdjkf8PatGBsBExK/adding-quantified-uncertainty-to-givewell-s-cost) uncertainty. But I couldn't find a tool that implements all six charities together and makes it easy to compare them while adjusting assumptions. GiveWell's spreadsheets are powerful but hard to explore casually. Each charity has its own multi-tab workbook with dozens of sheets, specialized terminology, and cross-references between cells: ![GiveWell's AMF cost-effectiveness spreadsheet showing multiple sheet tabs and the key sheet with terminology definitions](./givewell-spreadsheet.png) Changing a moral weight means editing cells across multiple sheets and comparing results manually. I wanted something where you could adjust one slider and immediately see how all six charities re-rank. ## What the model covers For each charity I implemented the core pipeline from GiveWell's spreadsheets: 1. **People reached**: Grant size / cost per person reached 2. **Deaths averted** (or equivalent): People reached × mortality/disease rate × intervention effect size 3. **Units of value**: Deaths averted × moral weight (age-adjusted) 4. **Cost-effectiveness**: Units of value per dollar / benchmark value per dollar 5. **Adjustments**: Charity-level (quality, track record), intervention-level (external validity), leverage and funging The six charities each have their own structure: - **AMF** and **Malaria Consortium**: Under-5 mortality reduction, with separate pathways for older age mortality and developmental effects - **Helen Keller International**: VAS effect on under-5 mortality - **New Incentives**: Cash incentives increase vaccination rates; the model converts incremental vaccinations to deaths averted using vaccine-specific effect sizes - **GiveDirectly**: The model values consumption increases directly (no mortality pathway) - **Deworm the World**: The model values long-run earnings effects of deworming via ln(consumption) All 51 charity/country combinations operate independently — each country has its own cost per person, mortality rate, adjustment factors, etc. ## Observations from the replication **The benchmarks aren't on the same scale.** GiveWell expresses each charity's cost-effectiveness as a multiple of GiveDirectly cash transfers. The denominator — how many "units of value" (GiveWell's composite of moral-weight-adjusted lives saved) one dollar of cash generates — differs across spreadsheets: AMF uses 0.00333, MC and HKI use 0.00335, and GiveDirectly uses 0.003. GiveWell built different spreadsheets at different times and baked in different moral weight calibrations. They don't mechanically compare multiples across spreadsheets, but a tool that displays them side by side inherits this inconsistency. **Mortality rate definitions vary.** AMF's spreadsheet has both a raw malaria mortality rate and a derived "mortality rate in the absence of nets" rate. The latter accounts for existing net coverage and is the correct input. My first extraction accidentally used the raw rates, which underestimated AMF's cost-effectiveness by roughly 2x for some countries (e.g., DRC: 0.00306 raw vs. 0.00798 in-absence-of-nets). **Counterfactual coverage drives most of the within-charity variation.** Both Helen Keller and New Incentives have a `proportionReachedCounterfactual` parameter — what fraction of people would receive the intervention anyway, without the charity's involvement. The remaining fraction is the charity's incremental impact: Helen Keller (VAS): | Country | Would receive VAS anyway | x benchmark | |---------|-------------------------|-------------| | Niger | 15% | 79× | | DRC | 20% | 30× | | Mali | 21% | 17× | | Madagascar | 33% | 12× | | Guinea | 18% | 11× | | Cameroon | 23% | 8× | | Burkina Faso | 29% | 7× | | Côte d'Ivoire | 40% | 6× | New Incentives (vaccinations, Nigerian states): | State | Would get vaccinated anyway | x benchmark | |-------|---------------------------|-------------| | Sokoto | 68.5% | 39× | | Zamfara | 79.4% | 31× | | Kebbi | 71.9% | 29× | | Bauchi | 81.5% | 20× | | Jigawa | 85.8% | 18× | | Katsina | 81.0% | 17× | | Kano | 83.2% | 13× | | Gombe | 88.0% | 10× | | Kaduna | 84.2% | 9× | Counterfactual coverage correlates with cost-effectiveness but doesn't determine it alone. Niger's 85% incremental coverage is similar to Guinea's 82%, yet Niger scores 7x higher: higher mortality rate (1.4% vs 1.1%), double the VAS effect (11.1% vs 5.5%), and 70% lower cost per child. **Helen Keller's leverage/funging adjustments need careful reading.** In my first pass, the funging adjustment for Burkina Faso extracted as 531.99 instead of -0.431. These values come from separate rows in the spreadsheet that are easy to confuse — a "percentage change" row vs. an "adjusted value" row. ## Verification I extracted parameters from GiveWell's November 2025 CEA spreadsheets ([AMF](https://docs.google.com/spreadsheets/d/1VEtie59TgRvZSEVjfG7qcKBKcQyJn8zO91Lau9YNqXc), [MC](https://docs.google.com/spreadsheets/d/1De3ZnT2Co5ts6Ccm9guWl8Ew31grzrZZwGfPtp-_t50), [HKI](https://docs.google.com/spreadsheets/d/1L6D1mf8AMKoUHrN0gBGiJjtstic4RvqLZCxXZ99kdnA), [NI](https://docs.google.com/spreadsheets/d/1mTKQuZRyVMie-K_KUppeCq7eBbXX15Of3jV7uo3z-PM)). I verified 46 of the 51 charity/country final cost-effectiveness multiples against the spreadsheets (GiveDirectly excluded — see Limitations): | Charity | Countries | Max difference | |---------|-----------|---------------| | Against Malaria Foundation | 8 | <0.001% | | Malaria Consortium | 8 | <0.001% | | Helen Keller International | 8 | <0.001% | | New Incentives | 9 | 0.000% (exact) | | Deworm the World | 13 | 0.000% (exact) | 298 automated tests pass. The remaining <0.001% differences for AMF/MC/HKI are floating-point precision, not model discrepancies. ## Interactive features ![Expanded view showing AMF in DRC with the step-by-step calculation breakdown and editable parameters](./detail-view.png) **Calculation breakdown**: Click any country to see the step-by-step calculation with every intermediate value. Click any highlighted number to edit it. **Moral weights**: GiveWell's default weights peak at ages 5-9 (134) and weight under-5 at 116. You can adjust these with a single multiplier or set each age bracket independently. Charities with different age profiles (AMF and MC focus on under-5; NI and HKI have broader age effects) shift rankings when you change these. **Sensitivity analysis**: Sweep any moral weight parameter across its range and see how all six charities' cost-effectiveness changes. This makes crossover points visible — for instance, the under-5 weight where NI overtakes MC, or the discount rate at which deworming drops below cash transfers. ## How assumptions affect rankings A few examples of what you see when you adjust parameters: **Default rankings (best country per charity, GiveWell Nov 2025 defaults):** | Rank | Charity | Best country | x benchmark | |------|---------|-------------|-------------| | 1 | Helen Keller | Niger | 79× | | 2 | New Incentives | Sokoto | 39× | | 3 | Deworm the World | Kenya | 35× | | 4 | AMF | Guinea | 23× | | 5 | Malaria Consortium | Chad | 15× | | 6 | GiveDirectly* | Mozambique | 4× | *GiveDirectly uses a simplified model with older parameters — see Limitations. **Double the under-5 moral weight.** All four mortality-focused charities see exactly +100% gains — their value comes entirely from deaths averted, so doubling the weight doubles the result. GiveDirectly gains only +3.4% because most of its value comes from consumption benefits. (A [2025 study](https://www.givedirectly.org/mortality2025/) found $1,000 transfers cut infant mortality by 48%, and [GiveWell now includes mortality effects](https://www.givedirectly.org/givewell-2024/) in GiveDirectly's CEA — but with steep discounts, so consumption still dominates.) Rankings don't change. **Equal moral weights across ages (all set to 100).** Child-focused charities drop 14-16% (the default under-5 weight of ~116 falls to 100). Helen Keller drops from 79× to 67×. Rankings stay the same. **Double AMF's cost per child in DRC.** The relationship is perfectly linear: doubling cost halves the x benchmark from 14.6× to 7.3×. This is a more powerful lever for changing *relative* rankings within mortality-focused charities than moral weight changes, which scale all of them equally. When most top charities prevent child deaths, changing the weight on child deaths scales them all in the same direction. What *does* break the rankings is operational cost differences between countries. The tool lets you find the specific crossover points where, say, doubling a cost parameter in one country moves it below another charity entirely. ## Programmatic access The models are pure TypeScript functions — no UI dependency. Clone the repo and run analyses directly: ```bash git clone https://github.com/MaxGhenis/givewell-cea && cd givewell-cea bun install ``` Sweep a parameter: ```typescript // save as sweep.ts, run with: bunx tsx sweep.ts import { calculateHelenKeller } from "./src/lib/models/helen-keller"; import { HK_COUNTRY_PARAMS } from "./src/lib/models/countries"; const niger = HK_COUNTRY_PARAMS.niger; for (let effect = 0.05; effect <= 0.2; effect += 0.05) { const r = calculateHelenKeller({ grantSize: 1_000_000, ...niger, vasEffect: effect, }); console.log(`VAS effect ${(effect * 100).toFixed(0)}% → ${r.finalXBenchmark.toFixed(1)}×`); } // VAS effect 5% → 35.6× // VAS effect 10% → 71.3× // VAS effect 15% → 106.9× // VAS effect 20% → 142.5× ``` Rank all charity/country combinations: ```typescript // save as rank.ts, run with: bunx tsx rank.ts import { calculateHelenKeller } from "./src/lib/models/helen-keller"; import { calculateAMF } from "./src/lib/models/amf"; import { calculateNewIncentives } from "./src/lib/models/new-incentives"; import { HK_COUNTRY_PARAMS, HK_COUNTRY_NAMES, AMF_COUNTRY_PARAMS, AMF_COUNTRY_NAMES, NI_COUNTRY_PARAMS, NI_COUNTRY_NAMES, } from "./src/lib/models/countries"; const G = 1_000_000; const all = [ ...Object.entries(AMF_COUNTRY_PARAMS).map(([k, v]) => ({ charity: "AMF", country: AMF_COUNTRY_NAMES[k], xb: calculateAMF({ grantSize: G, ...v }).finalXBenchmark, })), ...Object.entries(HK_COUNTRY_PARAMS).map(([k, v]) => ({ charity: "HKI", country: HK_COUNTRY_NAMES[k], xb: calculateHelenKeller({ grantSize: G, ...v }).finalXBenchmark, })), ...Object.entries(NI_COUNTRY_PARAMS).map(([k, v]) => ({ charity: "NI", country: NI_COUNTRY_NAMES[k], xb: calculateNewIncentives({ grantSize: G, ...v }).finalXBenchmark, })), ]; all.sort((a, b) => b.xb - a.xb); for (const r of all.slice(0, 10)) { console.log(`${r.charity.padEnd(4)} ${r.country.padEnd(16)} ${r.xb.toFixed(1)}×`); } // HKI Niger 79.1× // NI Sokoto 38.6× // NI Zamfara 31.3× // HKI DRC 29.9× // NI Kebbi 29.0× // AMF Guinea 22.8× // ... ``` Each model function (`calculateAMF`, `calculateHelenKeller`, `calculateNewIncentives`, etc.) takes a flat parameter object and returns all intermediate values, so you can inspect any step of the pipeline. ## Limitations I replicated the *structure* of GiveWell's models but not their full analytical process: - I implement the calculation pipeline but not the reasoning behind parameter choices. GiveWell's adjustments (charity quality, external validity, leverage, funging) encode substantial judgment that this tool takes as given. The parameters themselves — particularly the adjustment factors — represent years of investigation, site visits, literature reviews, and internal debate. This tool lets you see the arithmetic, but the arithmetic was never the hard part. - No uncertainty analysis. The tool supports sensitivity analysis — sweeping one parameter at a time to see how results change — but does not place joint distributions over parameters or propagate uncertainty through the model. [Several](https://forum.effectivealtruism.org/posts/4Qdjkf8PatGBsBExK/adding-quantified-uncertainty-to-givewell-s-cost) [excellent](https://forum.effectivealtruism.org/posts/ycLhq4Bmep8ssr4wR/quantifying-uncertainty-in-givewell-s-givedirectly-cost) [posts](https://forum.effectivealtruism.org/posts/Nb2HnrqG4nkjCqmRg/quantifying-uncertainty-in-givewell-cost-effectiveness) have explored what happens when you put distributions around these estimates. - I built the GiveDirectly model as a simplified approximation from an older spreadsheet and blog post, not a direct replication of GiveWell's current model. It uses identical parameters (spillover effects, mortality effects, consumption persistence) across all five countries, unlike GiveWell's country-differentiated approach. - The tool currently covers GiveWell's top 6 charities but not newer additions. For donation decisions, use [GiveWell's published estimates](https://www.givewell.org/how-we-work/our-criteria/cost-effectiveness/cost-effectiveness-models). Source code: [github.com/MaxGhenis/givewell-cea](https://github.com/MaxGhenis/givewell-cea) --- ## OpenMessage: How I built a macOS Google Messages client to give Claude my texts **URL:** https://maxghenis.com/blog/openmessage/ **Published:** Feb 09 2026 **Description:** I had already connected Claude Code to WhatsApp, Signal, Slack, and Gmail. SMS was the last holdout — and no existing tool could solve it. I've been on a mission to give Claude Code access to all my communication channels. WhatsApp, Signal, Slack, Gmail — each one has an MCP server that lets Claude read and send messages on my behalf. But one channel was missing: SMS and RCS on my Android phone. **[Download OpenMessage for Mac](https://github.com/MaxGhenis/openmessage/releases/latest/download/OpenMessage.dmg)** — free, open source, requires macOS 14.0+ and an Android phone with Google Messages. ## The gap If you have an iPhone, iMessage on Mac gives you desktop texting for free. But I use Android, and Google Messages only offers a web client — no desktop app, no API, no way for an AI assistant to interact with it. I needed Claude to be able to do things like "text Alex that the meeting moved to 3pm" or "check if James confirmed for tomorrow" without me picking up my phone. Every other messaging channel was already connected. SMS was the last holdout. ## Attempt 1: SMS Gateway My first approach was [SMS Gateway](https://smsgateway.me), an Android app that exposes your phone's SMS capability via an API. I built an MCP server around it, and it worked — sort of. The problems: - **Clunky setup** — required installing a separate Android app, configuring API keys, keeping the app running in the background - **SMS only** — no RCS support, so group chats and rich messages were invisible - **Sending was unreliable** — messages would sometimes queue but never actually send - **No conversation history** — you could only send messages, not read existing conversations It was a dead end for anything beyond basic "fire and forget" SMS. ## Attempt 2: Google Messages protocol I started digging into how Google Messages for Web actually works. It turns out Google uses an internal protocol (based on gRPC/protobuf) to pair a browser with your phone via QR code. A few open-source projects had reverse-engineered this protocol, most notably [mautrix/gmessages](https://github.com/mautrix/gmessages), a Go library originally built as a Matrix bridge. This was the breakthrough. The mautrix library could: - Pair with a phone via QR code (same as messages.google.com) - Receive all conversations and messages in real time - Send SMS and RCS messages - Handle group chats, reactions, and read receipts I quickly built an MCP server around it, and it worked beautifully. Claude could read my full message history and send texts. ## The mistake that became a feature Here's where I'll admit something: I thought Google Messages only allowed one paired web device at a time. It turns out Google Messages has two pairing methods — Google account pairing and QR code pairing — and they don't mix. I had been using Google account pairing, which blocks additional devices. When I switched to QR pairing (which is what the mautrix library uses), multiple devices work fine. But I didn't realize this at the time, so I was convinced that running my MCP server meant giving up the web client. I built a whole native macOS app to replace it. I didn't need to build the app at all. But I'm glad I did. What started as a workaround became something better: an open-source, native Google Messages client for Mac — the first one that exists. No browser tab to keep open, no Electron wrapper, just a real app with a built-in MCP server for AI access. ## Building the app Since I already had the Go backend handling the Google Messages protocol, I wrapped it in a native macOS app using Swift. The architecture is simple: - **Go backend** — handles pairing, message sync, and the MCP server. Stores everything in a local SQLite database - **Swift wrapper** — native macOS app that launches the Go backend and displays a web UI via WKWebView - **Web UI** — clean conversation list and message view served from localhost The app pairs with your phone the same way messages.google.com does — scan a QR code — and then syncs your full conversation history. The MCP server runs alongside the UI, so Claude can access messages while you browse them yourself. Everything is stored locally — SQLite database, no cloud sync, no accounts to create. When you use AI tools via the MCP server, only the messages you ask about are sent to your chosen AI provider (e.g. Anthropic's API for Claude). OpenMessage itself never sends your data anywhere. ## What Claude can do with it With OpenMessage running, Claude Code can: - **Search messages** — "find the address James sent me last week" - **Read conversations** — "summarize my conversation with the team" - **Send messages** — "text Alex that the meeting moved to 3pm" - **React to messages** — "thumbs up the last message from Sarah" Combined with WhatsApp, Signal, Slack, and Gmail MCP servers, Claude now has access to essentially all my communications. I can say "check all my messages for anything urgent" and it searches across every channel. ## What's next OpenMessage is open source, and I think it's a canvas for something bigger. The same open-source libraries (mautrix) that power OpenMessage also support WhatsApp, Signal, Telegram, Discord, Slack, and more. Beeper built a $125M company on these same libraries — but without local-first architecture or AI integration. I think there's an opportunity for an open-source, AI-native unified messaging client. Some ideas: - **More services** — WhatsApp, Signal, Telegram via the mautrix ecosystem - **AI-powered features** — smart replies, conversation summaries, priority inbox - **Natural language search** — "when did someone send me a confirmation number?" - **Cross-platform** — Linux and Windows versions If any of these interest you, [open an issue](https://github.com/MaxGhenis/openmessage/issues) or submit a PR. ## Get it OpenMessage is free and open source. Requires macOS 14.0+ and an Android phone with Google Messages. **[Download for Mac](https://github.com/MaxGhenis/openmessage/releases/latest/download/OpenMessage.dmg)** | [openmessage.ai](https://openmessage.ai) | [GitHub](https://github.com/MaxGhenis/openmessage) --- ## Where US immigrants come from, and where ICE focuses enforcement **URL:** https://maxghenis.com/blog/bad-bunny-immigration-enforcement/ **Published:** Feb 09 2026 **Description:** Bad Bunny''s Super Bowl halftime show highlighted the Americas. Here''s what the data shows about immigration and enforcement patterns. Last night, Bad Bunny performed the [first primarily Spanish-language Super Bowl halftime show](https://www.rollingstone.com/music/music-news/bad-bunny-super-bowl-performance-1235513007/) in history. At the end, holding a football that read "Together, we are America," he said "God bless America" and then [named every country and territory in the American continent](https://sports.yahoo.com/nfl/breaking-news/article/bad-bunny-echoes-grammy-message-with-super-bowl-halftime-show-together-we-are-america-015121748.html): Chile, Argentina, Uruguay, Paraguay, Bolivia, Peru, Ecuador, Brazil, Colombia, Venezuela, Panama, Costa Rica, Nicaragua, Honduras, El Salvador, Guatemala, Mexico, Cuba, the Dominican Republic, Jamaica, Haiti, the United States, Canada, and the US territory of Puerto Rico. Performers carried each flag behind him. The performance reached an estimated [130 million viewers](https://variety.com/2026/music/news/bad-bunny-super-bowl-lady-gaga-real-wedding-ceremony-1236656381/). Here's what the data shows about immigration from the Americas and current enforcement patterns. ## Most US immigrants come from the Americas **53%** of the 50.2 million foreign-born people in the United States come from other countries and territories in the American continent — Mexico alone accounts for 22.2% — according to the [2024 American Community Survey](https://data.census.gov/table/ACSDT1Y2024.B05006). Every country Bad Bunny named on that stage falls within that share. ## Spanish is the most common non-English language Bad Bunny performed entirely in Spanish and said during the show: "English wasn't my first language, but that's OK — it wasn't America's either." Spanish is the most common non-English language in the United States, [spoken at home by about 45 million people](https://data.census.gov/table/ACSST1Y2024.S1601), about 14% of the US population. Of the 74 million people who speak a non-English language at home, **61% speak Spanish**, per [2024 ACS data](https://data.census.gov/table/ACSST1Y2024.S1601). Congress has never passed a law designating an official language. In March 2025, [Executive Order 14224](https://www.federalregister.gov/documents/2025/03/06/2025-03694/designating-english-as-the-official-language-of-the-united-states) designated English as the official language and revoked [EO 13166](https://www.govinfo.gov/content/pkg/FR-2000-08-16/pdf/00-20938.pdf), a 2000 Clinton-era order that directed agencies to improve access for people with limited English proficiency. However, the order also states that agencies are "not required to amend, remove, or otherwise stop production of documents...in languages other than English." At the time of the [first US census in 1790](https://www.census.gov/library/publications/1909/decennial/century-populaton-growth.html), about 86% of the white population was of British Isles descent (English, Welsh, Scottish, or Irish), with German (9%), Dutch (3%), and French (2%) communities making up much of the rest. The census also counted roughly 700,000 enslaved people — most American-born by that point — and excluded most Native Americans, who spoke [over 300 languages](https://www.bia.gov/faqs/do-all-american-indians-and-alaska-natives-speak-single-traditional-language) at the time of European contact. Among the counted population, English was the dominant language, though German-speaking communities in Pennsylvania were large enough that [bilingual government documents](https://www.census.gov/library/publications/1909/decennial/century-populaton-growth.html) were common. Spanish-speaking settlers founded [St. Augustine in 1565](https://www.nps.gov/places/st-augustine-town-plan-historic-district-st-augustine-florida.htm), 42 years before the first permanent English settlement at Jamestown. ## ICE enforcement is concentrated in the Americas While 53% of the foreign-born population is from the Americas, ICE enforcement skews more heavily toward those countries. [Official ICE data](https://www.ice.gov/statistics) shows that **98-99% of ICE administrative arrests** from FY2021 through Q1 FY2025 involved people from other countries in the Americas. The [UC Berkeley Deportation Data Project](https://deportationdata.org), which publishes record-level FOIA'd arrest data, provides a more granular look: from January 20 to October 15, 2025, **92.6%** of 220,931 arrests involved nationals of countries in the Americas, with Mexico (38.6%), Guatemala (14.1%), and Honduras (11.0%) as the top three countries. A [UCLA Luskin study](https://knowledge.luskin.ucla.edu/wp-content/uploads/2025/10/Unseen_Latino-Ice-Arrests-Surge-Under-Trump_20251027.pdf) found that about 90% of ICE arrests during the first six months of Trump's second term involved nationals of Latin American countries, using the same Berkeley FOIA data. This concentration likely reflects the composition of the unauthorized population, which [skews more heavily Latin American](https://www.migrationpolicy.org/article/frequently-requested-statistics-immigrants-and-immigration-united-states) than the foreign-born population as a whole, as well as geographic proximity and enforcement priorities. ## Operation Metro Surge in Minnesota In December 2025, DHS launched [Operation Metro Surge](https://www.britannica.com/event/2025-26-Minnesota-ICE-Deployment) in Minneapolis, which it called the largest immigration enforcement operation ever, [deploying up to 2,000 federal agents](https://www.cbsnews.com/minnesota/live-updates/ice-somali-immigrants-minneapolis-st-paul/) to the Twin Cities. DHS said the operation focused on fraud in the Somali-American community. Minnesota is home to the largest Somali population in the US — roughly [84,000 people in the Twin Cities alone](https://www.pbs.org/newshour/nation/5-things-to-know-about-the-somali-community-in-minnesota-after-trumps-attacks). Of those, nearly **58% were born in the United States**, and of the foreign-born Somalis in Minnesota, **87% are naturalized US citizens**, according to [Census data reported by PBS](https://www.pbs.org/newshour/nation/5-things-to-know-about-the-somali-community-in-minnesota-after-trumps-attacks). Key outcomes of the operation: - **3,000+ people arrested** as of January 19, 2026, per [DHS](https://www.dhs.gov/news/2026/01/19/ice-continues-remove-worst-worst-minneapolis-streets-dhs-law-enforcement-marks-3000) - DHS publicly identified a subset of arrestees on its website. Of those, **23 were from Somalia** — while the operation's stated focus was on Somali fraud ([Britannica](https://www.britannica.com/event/2025-26-Minnesota-ICE-Deployment)). A [Washington Times analysis](https://www.washingtontimes.com/news/2026/jan/29/operation-metro-surge-crosses-3500-arrest-mark-heres-whos-getting/) of 170 DHS-released arrests found Mexico and Laos were the top two nationalities, with Somalia third. The full nationality breakdown of all arrests has not been made public - About **5% of arrestees had violent criminal records** ([Britannica](https://www.britannica.com/event/2025-26-Minnesota-ICE-Deployment)) - **Two US citizens were killed** by federal agents during the operation: Renee Good and Alex Pretti ([CBS News](https://www.cbsnews.com/news/minneapolis-trump-immigration-ice-border-patrol-arrests-protests-shootings/)) - **50,000 people** gathered on January 23 for a statewide general strike in protest ([Britannica](https://www.britannica.com/event/2025-26-Minnesota-ICE-Deployment)) [Sahan Journal reported](https://sahanjournal.com/immigration/immigration-enforcement-somali-community-impact/) effects on daily life in affected communities, including residents skipping medical appointments, reduced mosque attendance, and hundreds of businesses closing. Minnesota and the Twin Cities [filed a federal lawsuit against DHS](https://www.ag.state.mn.us/Office/Communications/2026/docs/00190_DHS_Complaint.pdf). ## Summary 53% of the 50.2 million foreign-born people in the United States come from other countries in the Americas. ICE enforcement is even more concentrated, with over 92% of arrests in 2025 involving nationals of countries in the Western Hemisphere. The largest enforcement operation, in Minnesota, was focused on the Somali community but resulted in over 3,000 arrests across nationalities. Bad Bunny's halftime show put a spotlight on the hemispheric connections between the US and the rest of the Americas. The immigration and enforcement data provide context for the scale of those connections. --- ## I used Claude Code to recertify for SNAP **URL:** https://maxghenis.com/blog/snap-recertification-claude-code/ **Published:** Feb 06 2026 **Description:** Browser automation + document reading + form filling: Claude Code completed my SNAP recertification in about an hour of active work, across 373 tool calls. I'm CEO of [PolicyEngine](https://policyengine.org), a tax and benefit policy nonprofit, but I don't take a regular salary — so I qualify for SNAP. I renewed my benefits tonight using [Claude Code](https://docs.anthropic.com/en/docs/claude-code) and its browser automation capabilities. The entire recertification — navigating DC's District Direct portal, reading my financial documents, calculating income, filling out dozens of form fields, and submitting — took about an hour of active work. Claude made 373 tool calls. I sent 26 messages. Here's what I did, and what Claude did. ## The conversation It started simply: > **Me:** help me recertify for snap > > **Claude:** I can help you with SNAP recertification. A few questions to get started: What state are you in? Have you received a recertification notice? Do you know if your state has an online portal? > > **Me:** dc, idk, yes it does Claude found the portal, opened it in Chrome, and told me to log in (it can't enter passwords). I logged in and said: > **Me:** done > > **Claude:** You're logged in and there's a banner right at the top: "Your SNAP case is due for recertification, click the link to recertify before your case is terminated." Early on, Claude was asking for confirmation before every step. I put a stop to that: > **Claude:** I can see the Start button at the bottom. Should I click it to begin the recertification form? > > **Me:** yes stop asking just go! After that, Claude moved through the pre-filled sections autonomously — name, DOB, citizenship, household composition — only stopping when it actually needed something from me. Like when it hit the SSN page: > **Claude:** This page is asking for your SSN. It's partially masked (showing the last four digits) so it's already on file. I cannot enter sensitive information like SSNs — this is something you'll need to verify/enter yourself. > > **Me:** it's already there, just hit next When Claude reached the income section, it needed updated numbers from my financial accounts. It tried to access Wealthfront and Vanguard in the browser but couldn't: > **Claude:** Financial sites are blocked for safety reasons — I can't access Wealthfront or Vanguard since they contain sensitive financial data. You'll need to grab those numbers yourself. > > **Me:** can you ge tthem? > > **Me:** i opened https://dashboard.web.vanguard.com/ can you view it? > > **Claude:** I don't see the Vanguard tab in my tab group. Unfortunately financial sites like Vanguard and Wealthfront are blocked by the browser extension's safety restrictions. > > **Me:** just downloaded a pdf for vanbguad (You'll notice the typos — I've gotten sloppier with typing since Claude handles it fine, so I don't bother correcting myself.) So I downloaded a Vanguard PDF, a Wealthfront CSV, and my paystubs to my Downloads folder. Claude read them directly from the filesystem, extracted the relevant figures, and calculated my quarterly investment income. The rest was uneventful. Claude filled in employer details, navigated through expenses (all pre-filled from last time), checked attestation boxes, and reached the final submission page: > **Claude:** Before I click Next to submit the recertification, I want to confirm with you. This is the final certification page where you're signing and certifying that all information is correct. Ready to submit? > > **Me:** i dont this it does but yes hit next Submitted at 8:56 PM. The confirmation page told me my SNAP interview is due by February 10. ## What I provided My role was minimal: - **My state**: DC - **My address hasn't changed** - **Downloaded documents**: I downloaded statements from my financial institutions showing dividend and interest income, and my recent paystubs, to my Downloads folder - **Logged in**: I handled the login to District Direct myself - **A few corrections**: like telling Claude the SSN was already filled in, and that I'm already registered to vote That's 26 messages total. I never looked at the form itself — not once. Everything else was Claude. ## What Claude did Claude read my financial documents (PDFs and CSVs), extracted the relevant figures, navigated the multi-page recertification form, and filled in every field. Here's a rough breakdown of its 373 tool calls: - **72 screenshots** to see the current state of the page - **88 clicks** to navigate buttons, dropdowns, checkboxes, and links - **28 find operations** to locate elements on the page - **22 form inputs** to fill text fields, select options, and enter data - **6 page navigations** - Plus scrolling, reading documents, web searches, and calculations The form covers household composition, address, income from all sources (employment, investments, other), expenses (shelter, utilities, childcare, medical), and rights and responsibilities. Claude navigated all of it, carrying forward pre-filled data where nothing had changed and updating the fields that needed new numbers. ## How long it took The whole process spanned about 3 hours and 40 minutes wall-clock, but most of that was idle time — me downloading documents, reading what Claude was doing, or stepping away. The active working time was roughly 1 hour and 20 minutes. For comparison, doing this manually takes me at least an hour of focused attention: logging in, finding the right forms, looking up all my financial information, entering it carefully, reviewing everything, and submitting. Claude's version required about 5 minutes of my attention spread across the session. ## Why this matters SNAP recertification is exactly the kind of task that AI should help with. It's: - **High-stakes but routine**: Getting it wrong could mean losing benefits, but the process itself is just data entry - **Document-heavy**: It requires pulling numbers from paystubs, bank statements, and investment accounts - **Tedious**: The form is long, repetitive, and easy to make mistakes on - **Time-sensitive**: Miss the deadline and you lose benefits People who receive SNAP benefits are, by definition, low-income. Their time is valuable. The bureaucratic burden of maintaining benefits is a real cost, and it causes people to lose benefits they're entitled to. Tools that reduce that burden matter. ## What you'd need to try this This used Claude Code with the [Claude in Chrome](https://chromewebstore.google.com/detail/claude-in-chrome/blelmpkgncmhfbjlccpkbjbdgfdjpgol) extension for browser automation. You'd need: - A Claude Pro or Max subscription (for Claude Code access) - The Chrome extension installed - Your financial documents downloaded and accessible - Comfort with Claude having access to your browser and local files This is still early-stage technology. I wouldn't recommend it for someone who isn't comfortable reviewing what Claude is doing as it works. But the trajectory is clear: the boring, stressful parts of interfacing with government systems are exactly what AI agents should handle. --- ## Scrollywood: Smooth scroll video recording for the web **URL:** https://maxghenis.com/blog/scrollywood/ **Published:** Feb 06 2026 **Description:** A Chrome extension that records smooth-scrolling videos and GIFs of any webpage, built for capturing scrollytelling stories and long-form content. At [PolicyEngine](https://policyengine.org), we're building more scrollytelling stories to explain how tax and benefit policy works. Pages like [MITA](https://maxghenis.com/mita) use scroll-driven animations to walk through data visualizations step by step, and we want to share these experiences beyond the browser — in presentations, social posts, and demos. The problem: there's no good way to record a smooth scroll of a webpage. Screen recording tools require you to scroll manually, which is never perfectly smooth. Browser automation tools can screenshot sequences but don't capture scroll-triggered animations. And scrollytelling pages depend on continuous scrolling to trigger IntersectionObserver callbacks that animate the content. So I built [Scrollywood](/scrollywood) — a Chrome extension that records a perfectly smooth scroll capture of any webpage. ![Scrollywood demo — smooth scroll recording of a scrollytelling page](/scrollywood-demo.gif) ## How it works Click the extension icon, set your scroll duration, choose WebM, MP4, or GIF, and start recording. Scrollywood scrolls the page from top to bottom at a constant rate while recording the tab at 30-60fps. The result is a downloadable capture in your selected format, with MP4 shown when Chrome supports it. The scroll is smooth and linear at 60fps, which is critical for scrollytelling pages. Libraries like [scrollama](https://github.com/russellsamora/scrollama) and [react-scrollama](https://github.com/jsonkao/react-scrollama) use IntersectionObserver to trigger animations as elements enter the viewport. A smooth programmatic scroll naturally crosses these thresholds, so the animations play exactly as they would during manual scrolling. ## Technical challenges Building this was straightforward in concept but tricky in practice, mostly due to Chrome's Manifest V3 architecture: **Service worker timeouts.** MV3 service workers sleep after ~30 seconds of inactivity, which makes `setTimeout` unreliable for recordings longer than 30 seconds. The fix: the offscreen document (which has a persistent DOM context) handles all timing, while the service worker only handles one-shot operations like script injection and downloads. **Iframe-wrapped pages.** Some sites embed content in full-page iframes (e.g., a custom domain wrapping a GitHub Pages site). The outer frame has no scrollable content. Scrollywood detects these wrappers and defers to the inner frame, which handles its own scrolling via `allFrames: true` injection. **CSS scroll-behavior conflicts.** Pages with `scroll-behavior: smooth` in CSS fight with programmatic scrolling. Scrollywood temporarily overrides this to `auto`, but carefully avoids overriding `overflow`, which would break `position: sticky` — the CSS property that makes scrollytelling graphics stay in place. **Large recording downloads.** At 16Mbps for a 60-second recording, the video can be 100MB+. The original approach (base64-encoding the blob and sending it via `chrome.runtime.sendMessage`) silently fails at that size. The fix: download directly from the offscreen document using a blob URL. ## Get it [Download Scrollywood](/scrollywood) and load it as an unpacked extension in Chrome. It's free and open source. [Project page](/scrollywood) · [GitHub](https://github.com/MaxGhenis/scrollywood) --- ## The adolescence of policy simulation: On Amodei and economic disruption **URL:** https://maxghenis.com/blog/amodei-adolescence-policyengine/ **Published:** Jan 27 2026 **Description:** AI can answer policy questions. It cannot yet model policy impacts. That requires infrastructure we are still building. Anthropic CEO Dario Amodei published a [20,000-word essay](https://www.darioamodei.com/essay/the-adolescence-of-technology) today called "The Adolescence of Technology." It covers AI safety, bioweapons, authoritarianism, and more. But one section stood out to me: his predictions about economic disruption. Amodei writes that AI "will have effects that are much broader and occur much faster" than previous technological transitions. He predicts this will trigger an "unusually painful" short-term shock in the labor market, potentially displacing half of entry-level white collar jobs within the next five years. If he's right, we'll need policy responses. And to design policy responses, we need to model them. That's where things get interesting. ## What AI can and cannot do AI has gotten remarkably good at answering policy questions. Ask a frontier model about EITC phase-out rates or CTC eligibility rules and you'll probably get a correct answer. The models have read the tax code. But there's a difference between answering policy questions and modeling policy impacts. Try this: > "If the US expanded the Child Tax Credit to $3,600 for all children regardless of income, what would happen to the child poverty rate and the Gini coefficient?" I asked myself this question. Here's my honest answer: *I don't know. I could tell you the current child poverty rate is around 12% (SPM) and guess it would drop significantly—maybe to 8-9%?—based on vague memories of studies from the 2021 expanded CTC. For Gini, I'd guess a small decrease, maybe 0.005-0.01 points? But these are guesses anchored to half-remembered papers, not calculations. I have no microdata. I can't run a counterfactual. I'm pattern-matching to priors.* This isn't a prompting problem. It's an infrastructure problem. ## Policy simulation requires three things To actually model the impact of a policy change, you need: 1. **Policy** — The rules: tax formulas, benefit eligibility, phase-outs, interactions between programs 2. **Data** — Representative microdata: who earns what, household composition, geographic distribution 3. **Theory** — Behavioral assumptions: how do people respond to incentives, what's the takeup rate AI might eventually nail the policy part. With enough statute text in context—or better, with deterministic APIs encoding the rules—models can learn to apply tax law correctly. But AI can't conjure CPS microdata. It can't know the joint distribution of income, household size, and state of residence for 130 million US households. It can't decide whether to assume zero behavioral response or apply elasticities from the labor economics literature. That's not intelligence. That's infrastructure. ## Two sides of the same coin The same infrastructure serves another purpose. In "Machines of Loving Grace," Amodei imagines "a very thoughtful and informed AI whose job is to give you everything you're legally entitled to by the government." That's not hypothetical—tools like MyFriendBen and Amplifi already use PolicyEngine's API to screen people for benefits. Policy simulation and benefit access are two sides of the same coin. One computes law against population data, the other against your data. ## The adolescence parallel Amodei calls this moment the "adolescence of technology"—powerful but not yet mature, capable of great things and great harm. Policy simulation is in its own adolescence. We can model reforms in hours instead of months. [Governments are starting to use these tools](https://policyengine.org/uk/research/policyengine-10-downing-street). But the infrastructure is still clunky, adoption is still sparse, and most policy debates still happen without quantitative grounding. The question is whether policy simulation grows up fast enough to match the disruption it needs to respond to. ## What this implies If Amodei's timeline is right—powerful AI within two years, significant labor displacement within five—we don't have time for the traditional policy analysis cycle. Bills get introduced, CBO scores them months later, revisions happen, repeat. That works when policy moves at legislative pace. But if the economy is changing faster than legislatures can respond, we need infrastructure that lets policymakers iterate quickly. Not just ideas about what to do, but tools for modeling trade-offs in real time. This is part of why I've spent the last few years building [PolicyEngine](https://policyengine.org). It's one attempt at closing the gap—open-source microsimulation that anyone can use. But the broader point isn't about any one tool. It's that policy response capacity needs to exist before crises, not during them. AI can accelerate parts of this. [We've been experimenting with AI workflows](https://policyengine.org/us/research/multi-agent-workflows-policy-research) that compress hours of analysis into minutes. But AI can only accelerate infrastructure that exists. It can't replace infrastructure that doesn't. --- *Read Amodei's full essay: [The Adolescence of Technology](https://www.darioamodei.com/essay/the-adolescence-of-technology)* --- ## opencollective-py: Manage OpenCollective from your terminal or Claude Code **URL:** https://maxghenis.com/blog/opencollective-py/ **Published:** Jan 23 2026 **Description:** A Python client, CLI, and MCP server for the OpenCollective API. I manage [PolicyEngine's OpenCollective](https://opencollective.com/policyengine). We use OpenCollective because it gives anyone full visibility into our finances—every expense, donation, and transaction is public. It also handles crowdfunding, recurring donations, and community updates, and the [platform itself is open source](https://github.com/opencollective/opencollective). But submitting expenses, reviewing pending ones, approving reimbursements—it all meant clicking through a web UI. So I built [opencollective-py](https://github.com/MaxGhenis/opencollective-py)—a Python client that handles OpenCollective operations programmatically. It includes a CLI for quick terminal commands and an MCP server so Claude Code can manage expenses directly. ## What it does **Submit expenses**—reimbursements (with receipts) or invoices (for services): ```bash oc reimbursement "Conference registration" 500.00 receipt.pdf -c policyengine oc invoice "Consulting - January" 2000.00 -c policyengine ``` **List and filter expenses**: ```bash oc expenses -c policyengine --pending oc expenses -c policyengine --status approved ``` **Approve or reject** (for collective admins): ```bash oc approve exp-abc123 oc reject exp-def456 ``` **Get account info**: ```bash oc me ``` ## Python client ```python from opencollective import OpenCollectiveClient client = OpenCollectiveClient(access_token="...") # Submit expenses client.submit_reimbursement("policyengine", "Travel", 50000, "receipt.pdf") client.submit_invoice("policyengine", "Consulting", 200000) # Manage expenses expenses = client.list_expenses("policyengine", status="PENDING") client.approve_expense("exp-abc123") client.reject_expense("exp-def456") ``` ## MCP server for Claude Code Add to your Claude Code config: ```json { "mcpServers": { "opencollective": { "command": "python", "args": ["-m", "opencollective.mcp_server"] } } } ``` Then manage expenses in natural language: "Show me pending expenses for PolicyEngine" or "Submit my AWS receipt as a $150 reimbursement." ## HTML receipt conversion Many email receipts are HTML. OpenCollective only accepts images and PDFs. The package automatically converts HTML to PDF: ```bash oc reimbursement "AWS bill" 150.00 aws-receipt.html -c policyengine ``` ## Get it ```bash pip install opencollective # With PDF conversion pip install "opencollective[pdf]" ``` Then authenticate: ```bash oc auth ``` [Project page](/opencollective-py) ・ [PyPI](https://pypi.org/project/opencollective/) ・ [GitHub](https://github.com/MaxGhenis/opencollective-py) *This is an unofficial community project, not affiliated with OpenCollective.* --- ## From IDE to AI orchestration: The end of code-first development **URL:** https://maxghenis.com/blog/ide-to-ai-orchestration/ **Published:** Jan 11 2026 **Description:** How AI coding tools evolved from autocomplete to autonomous agents, and why I expect to abandon my IDE entirely this month. import AICodingTimeline from '../../components/AICodingTimeline.astro'; import CrosspostAware from '../../components/CrosspostAware.astro'; Claude Code creator Boris Cherny [recently shared](https://x.com/bcherny/status/2004897269674639461) that since the beginning of December, "100% of my contributions to Claude Code were written by Claude Code." He elaborated: "In the last thirty days, I landed 259 PRs—497 commits, 40k lines added, 38k lines removed. Every single line was written by Claude Code + Opus 4.5." I've been building entirely in Claude Code since October—longer than Boris because he's a much better developer than I am, so Claude didn't have to be as good. My VS Code is just a shell for Claude Code terminals. It's why I created the [TerminalGrid extension](https://maxghenis.com/terminalgrid)—I haven't used the file editor in months. The next step in this evolution is to skip the IDE altogether. And 2025's tool launches suggest that's exactly where the industry is heading. ## The timeline **Feb 24, 2025** — [Anthropic launches Claude Code](https://www.anthropic.com/news/claude-3-7-sonnet) A simple terminal tool—chat with Claude, edit files, run bash commands. **May 16, 2025** — [OpenAI launches Codex](https://openai.com/index/introducing-codex/) A cloud-based software engineering agent that runs tasks in isolated containers. **Nov 18, 2025** — [Google launches Antigravity](https://developers.googleblog.com/build-with-google-antigravity-our-new-agentic-development-platform/) An agent-first development platform where you don't touch code—you manage agents. **Nov 24, 2025** — [Anthropic launches Opus 4.5 + Claude Code in Desktop](https://www.anthropic.com/news/claude-opus-4-5) Multiple parallel sessions with isolated git worktrees. **Jan 6, 2026** — [Anthropic adds local Claude Code to Desktop](https://x.com/_catwu/status/2008628736409956395) Claude Code can run any shell command, access your filesystem, and control your browser. The Claude Desktop approach differs from Codex and similar cloud tools in one crucial way: full local access. Claude Code can run any shell command, access your filesystem, and control your browser. Cloud-based agents like Codex run in sandboxed containers with restricted permissions. Claude's bet is that developers want an agent with the same access they have—not a constrained PR-generator. ## The paradigm shift Google's Antigravity documentation articulates what's happening: the shift from "Editor view" (traditional IDE interface with an agent sidebar) to "Manager view" (orchestrating autonomous agents). In Manager view, you don't write code—you describe what you want, the agents work in parallel, and you review artifacts (task lists, implementation plans, screenshots, browser recordings). It's project management, not programming. This matches how I've been working—I typically run 5-10 Claude Code sessions simultaneously through [TerminalGrid](/terminalgrid), spread across a 2x2 grid (sometimes with multiple tabs per cell, sometimes across two VS Code windows): ![TerminalGrid showing multiple Claude Code sessions](./terminalgrid.png) Boris Cherny [described a similar setup](https://venturebeat.com/technology/the-creator-of-claude-code-just-revealed-his-workflow-and-developers-are): "I run 5 Claudes in parallel in my terminal. I number my tabs 1-5, and use system notifications to know when a Claude needs input." He also runs 5-10 instances on claude.ai, using a "teleport" command to hand off sessions between web and terminal. ## Why I'm not quite there yet This week I've been experimenting with Claude Code in Desktop, and while it's close to replacing VS Code, I've hit two blocking issues: 1. **No MCP support yet**: Claude Code in the Desktop app doesn't support MCP servers, so I can't use browser automation tools like Claude in Chrome. The same MCPs work fine in VS Code's terminal and in non-Code Claude Desktop. I [posted about this](https://x.com/MaxGhenis/status/2009808936212549910)—a Claude Code engineer [replied](https://x.com/amorriscode/status/2010410355030421700) that MCP support should come next week. 2. **Sessions drop when switching accounts**: I run out of quota on my Claude Max 20x account regularly (those 5 parallel Claude instances add up), so I maintain two accounts. Claude Desktop drops sessions when I switch between them, losing context. ## The bigger picture [As of late 2025](https://blog.jetbrains.com/research/2025/10/state-of-developer-ecosystem-2025/), roughly 85% of developers regularly use AI tools for coding. [Stack Overflow's 2025 Developer Survey](https://survey.stackoverflow.co/2025/) found 65% use them weekly. The question isn't whether AI will change how we code—it's whether we'll still call what we do "coding" at all. Boris Cherny's workflow hints at the answer. He doesn't code. He orchestrates agents. He described using Opus 4.5 for everything because "even though it's bigger & slower than Sonnet, since you have to steer it less and it's better at tool use, it is almost always faster than using a smaller model in the end." That's the trade-off: more capable models need less human steering, which means less time in the editor, which means the editor matters less. ## What happens next I expect Anthropic to fix both blocking issues this month. When they do, I'll switch from VS Code to Claude Desktop, and TerminalGrid will join the graveyard of tools I've built that AI capabilities have rendered obsolete. [CodeStitch](https://codestitch.dev), which I built in mid-2024 to paste codebases into Claude's context window, was obsolete within months of Claude Code maturing. TerminalGrid solved a real problem for exactly the window between "Claude Code exists" and "Claude Code has a native manager interface." That window is closing. The pattern for developer tooling: ship fast, solve the problem in front of you, expect obsolescence. Frontier labs are moving so quickly on general-purpose dev tools that anything built to improve them has a short shelf life. But that's fine. Disposable tools can accelerate progress toward bigger challenges. The real opportunity isn't building better coding tools—it's applying AI to problems that matter. Health, poverty, governance, the environment. If you can ship in days what used to take months, you can tackle problems that were previously out of reach. When Claude Desktop works reliably, I'll abandon my TerminalGridded IDE without looking back—and get back to the real work. --- ## TerminalGrid: Turn VS Code into a Claude Code superterminal **URL:** https://maxghenis.com/blog/terminalgrid/ **Published:** Jan 07 2026 **Description:** A VS Code extension for keyboard-driven terminal grid management with project picker and auto-launch for AI coding tools. I run multiple Claude Code sessions simultaneously—one for each project I'm working on. VS Code's terminal panel only splits in one direction—no 2D grids. And the native terminal has [issues with image pasting](https://github.com/anthropics/claude-code/issues/1361) that matter when you're sharing screenshots with Claude. So I built [TerminalGrid](https://maxghenis.com/terminalgrid)—my first VS Code extension, built 100% with Claude Code.
## VS Code as a shell for coding agents Boris Cherny, who created Claude Code, [recently shared](https://x.com/bcherny/status/2004897269674639461) that 100% of his contributions to Claude Code in the past month were written by Claude Code itself. I've been operating this way for two to three months now—not writing any code myself. Of course, Boris is a much better programmer than I am, so I had less to discard. This shift has changed how I use VS Code. I've stopped looking at the code directly; VS Code is now just a shell for coding agents. I don't use the file explorer or most other features. It's still better than a raw terminal since it lets you paste screenshots, but otherwise I just keep it as a grid of Claude Code sessions. ## The problem VS Code's [integrated terminal](https://code.visualstudio.com/docs/terminal/basics) only splits in one direction at a time—side-by-side when the panel is at the bottom, or stacked when it's on the side. No 2D grids. This has been requested since 2018 ([#56112](https://github.com/microsoft/vscode/issues/56112), [#160501](https://github.com/microsoft/vscode/issues/160501)) and people are still asking in 2025 ([#254638](https://github.com/microsoft/vscode/issues/254638), [#252458](https://github.com/microsoft/vscode/issues/252458)). Other extensions like [Split Terminal](https://marketplace.visualstudio.com/items?itemName=BrianNicholls.split-terminal) and [Workspace Layout](https://marketplace.visualstudio.com/items?itemName=lostintangent.workspace-layout) don't solve this—they work within the terminal panel's limitations. When you're running 4+ AI coding sessions, you need a proper grid: ``` ┌─────────────────┬─────────────────┐ │ policyengine │ api-server │ │ (Claude Code) │ (Claude Code) │ ├─────────────────┼─────────────────┤ │ docs │ frontend │ │ (Claude Code) │ (Claude Code) │ └─────────────────┴─────────────────┘ ``` ## The solution TerminalGrid moves terminals to the editor area, where VS Code already supports full grid layouts. Then it adds keyboard shortcuts and a project picker: 1. **`Cmd+K Cmd+N`** — Opens a searchable list of your projects 2. **Pick a project** — Type to filter, select with Enter 3. **Terminal launches** — Runs `cd && claude` automatically The terminal is named after the folder, so you always know which Claude is working on what. ## Setup ```json { "terminalgrid.projectDirectories": ["~/projects", "~/code"], "terminalgrid.autoLaunchCommand": "claude" } ``` That's it. Now every `Cmd+K Cmd+Down/Right/N` gives you a project picker that launches Claude in the right directory. ## Other features - **Crash recovery** — Terminal directories persist even if VS Code crashes - **Image pasting** — Editor-area terminals handle screenshots better than the terminal panel - **Works with any CLI tool** — Aider, Codex, Gemini CLI, or just plain shells ## Get it Install from the [VS Code Marketplace](https://marketplace.visualstudio.com/items?itemName=MaxGhenis.terminalgrid) or: ```bash code --install-extension MaxGhenis.terminalgrid ``` [Project page](/terminalgrid) ・ [VS Code Marketplace](https://marketplace.visualstudio.com/items?itemName=MaxGhenis.terminalgrid) ・ [GitHub](https://github.com/MaxGhenis/terminalgrid) --- *Day 1 of [12 Days of Shipping](https://maxghenis.com). Merry Christmas!* --- ## RAMBar: A macOS menu bar RAM monitor for developers **URL:** https://maxghenis.com/blog/rambar/ **Published:** Jan 07 2026 **Description:** A native macOS menu bar app for tracking memory usage across Claude Code sessions, VS Code workspaces, Chrome tabs, and Python processes.
Claude Opus 4.5 kept crashing my VS Code. Nothing wrong with the model—it was just so good that I started running 5, 6, even a dozen Claude Code sessions at once. My 16GB MacBook Air couldn't keep up. I upgraded to a 48GB MacBook Pro, which mostly solved the crashes, but I still had no visibility into what was eating memory. Activity Monitor shows processes, but not answers like "which Claude session is the hog?" or "can I spawn another subagent?" So I built [RAMBar](https://maxghenis.com/rambar)—my first macOS app, built entirely with Claude Code. It shows RAM usage the way developers think about it. ## What It Does RAMBar lives in your menu bar showing current RAM percentage, color-coded by status: - **Green**: Under 70%, you're fine - **Yellow**: 70-85%, getting tight - **Red**: Over 85%, time to close something Click it to see a breakdown of what's actually using your memory. ## Developer-Focused Breakdowns Unlike generic system monitors, RAMBar understands developer workflows: - **Claude Code sessions** — Shows main sessions vs subagents separately, so you can see when parallel agents are spawning - **VS Code workspaces** — Memory grouped by which project folder is open - **Chrome tabs** — Lists individual tabs by memory, so you can find that one tab eating 2GB - **Python processes** — Useful when you're running Jupyter notebooks or ML training alongside everything else ## How it compares to other tools Several macOS menu bar monitors exist, but none understand developer workflows: | Tool | Price | Developer context | Claude Code awareness | |------|-------|-------------------|----------------------| | **RAMBar** | Free | ✅ Groups by workspace/session | ✅ Shows sessions + subagents | | [Stats](https://github.com/exelban/stats) | Free | ❌ Process-level only | ❌ | | [iStat Menus](https://bjango.com/mac/istatmenus/) | $12 | ❌ Process-level only | ❌ | | [MenuBar Stats](https://seense.com/menubarstats/) | $5 | ❌ Process-level only | ❌ | | Activity Monitor | Free | ❌ Process-level only | ❌ | [Stats](https://github.com/exelban/stats) is excellent for general system monitoring—CPU graphs, network throughput, disk I/O—and it's free and open source. [iStat Menus](https://bjango.com/mac/istatmenus/) adds weather widgets, fan control, and extensive customization for $12. RAMBar solves a different problem. When I see high memory usage, I don't want to know that `node` is using 4GB—I want to know *which VS Code workspace* that node process belongs to. I don't care that there are 47 Chrome Helper processes; I want to know which *tab* is the culprit. And when Claude Code spawns subagents, I want to see that hierarchy, not a flat list of identical `claude` processes. If you need comprehensive system monitoring, use Stats or iStat Menus. If you're an AI-assisted developer who wants to know why your Mac is struggling during a coding session, RAMBar fills that gap. ## Get it ```bash brew tap maxghenis/tap brew install --cask rambar ``` [Project page](/rambar) ・ [GitHub](https://github.com/MaxGhenis/rambar) Requires macOS 14.0+. First launch: right-click and select Open to bypass Gatekeeper (the app is currently unsigned). --- *Part of [12 Days of Shipping](https://maxghenis.com).* --- ## How to Use MDX for Interactive Posts **URL:** https://maxghenis.com/blog/using-mdx/ **Published:** Nov 25 2025 **Description:** This blog supports MDX, allowing you to embed React components, Plotly charts, and interactive elements directly in posts. This blog is built with [Astro](https://astro.build) and supports MDX, which lets you embed interactive components directly in your markdown posts. ## What You Can Do With MDX, posts can include: - **React components** with full interactivity - **Plotly charts** for data visualization - **Custom calculators** or interactive widgets - **Embedded iframes** for external apps ## Example: Interactive Button Here's a simple interactive component embedded in this post: import HeaderLink from '../../components/HeaderLink.astro'; Click me - I'm interactive! ## Adding Plotly Charts For data-heavy posts, you can create a React component with Plotly and embed it: ```jsx // src/components/MyChart.tsx import Plot from 'react-plotly.js'; export default function MyChart() { return ( ); } ``` Then in your MDX post: ```mdx import MyChart from '../../components/MyChart'; ``` The `client:load` directive tells Astro to hydrate the component on the client side. ## Embedding External Apps You can also embed Streamlit apps, Observable notebooks, or any iframe-compatible content: ```html