# Max Ghenis - Full Content
> This file contains the full text of all blog posts from maxghenis.com, formatted for LLM consumption.
> Generated: 2026-09-04T17:10:02.732Z
> See also: /llms.txt for a curated overview
---
## Table of Contents
1. [A year of ChatGPT is a fifth of a beer](/blog/drinking-ai/)
2. [MacKenzie Scott's giving, in QALYs](/blog/mackenzie-scott-giving-in-qalys/)
3. [The US government is gutting Anthropic's R&D capacity](/blog/us-gutted-anthropic-rd/)
4. [How to turn money into predictions](/blog/how-to-turn-money-into-predictions/)
5. [Can Talkie-1930 do arithmetic?](/blog/talkie-1930-math-evals/)
6. [Billionaires aren't 'just as likely' to be nonpayers as top taxpayers](/blog/madoff-billionaire-nonpayers/)
7. [I built a MyST-to-Quarto converter (and why you might need one)](/blog/mystquarto/)
8. [One year of Claude Code](/blog/my-claude-code-config/)
9. [Why you can''t make double eye contact](/blog/double-eye-contact/)
10. [Interactive replication of GiveWell''s cost-effectiveness analysis](/blog/givewell-cea/)
11. [OpenMessage: How I built a macOS Google Messages client to give Claude my texts](/blog/openmessage/)
12. [Where US immigrants come from, and where ICE focuses enforcement](/blog/bad-bunny-immigration-enforcement/)
13. [I used Claude Code to recertify for SNAP](/blog/snap-recertification-claude-code/)
14. [Scrollywood: Smooth scroll video recording for the web](/blog/scrollywood/)
15. [The adolescence of policy simulation: On Amodei and economic disruption](/blog/amodei-adolescence-policyengine/)
16. [opencollective-py: Manage OpenCollective from your terminal or Claude Code](/blog/opencollective-py/)
17. [From IDE to AI orchestration: The end of code-first development](/blog/ide-to-ai-orchestration/)
18. [TerminalGrid: Turn VS Code into a Claude Code superterminal](/blog/terminalgrid/)
19. [RAMBar: A macOS menu bar RAM monitor for developers](/blog/rambar/)
20. [How to Use MDX for Interactive Posts](/blog/using-mdx/)
21. [VAT thresholds, revenues, and the role of counterfactuals](/blog/vat-thresholds/)
22. [AI models favor Cuomo over Mamdani on NYC housing production](/blog/ai-models-nyc-housing/)
23. [My 2024 in code: Personal projects and AI tools](/blog/my-2024-in-code/)
24. [Why I'm a neoliberal](/blog/why-im-a-neoliberal/)
25. [In appointing their newest member, the Ventura City Council only pretended to care about housing](/blog/in-appointing-their-newest-member-the-ventura-city/)
26. [If you earned more in 2019 than 2018, don't file your 2019 taxes yet! Otherwise, file ASAP!](/blog/if-you-earned-more-in-2019-than-2018-dont-file-you/)
27. [Why I’ve taken the Giving What We Can pledge](/blog/why-ive-taken-the-giving-what-we-can-pledge/)
28. [Ventura County must end its unsheltered homelessness](/blog/ventura-county-must-end-its-unsheltered-homelessne/)
29. [Warren’s wealth tax would raise less than she claims — even using her economists’ own assumptions](/blog/warrens-wealth-tax-would-raise-less-than-she-claim/)
30. [Andrew Yang's troubling Tucker Carlson interview](/blog/andrew-yangs-troubling-tucker-carlson-interview/)
31. [Quantile regression, from linear models to trees to deep learning](/blog/quantile-regression-from-linear-models-to-trees-to/)
32. [We should replace the Child Tax Credit with a universal child benefit](/blog/we-should-replace-the-child-tax-credit-with-a-univ/)
33. [The case for a wonky basic income plan](/blog/the-case-for-a-wonky-basic-income-plan/)
34. [Reddit featured a misleading headline on education secretary nominee Betsy DeVos](/blog/reddit-featured-a-misleading-headline-on-education/)
35. [Uber's tipping settlement will reduce earnings of African-American drivers](/blog/ubers-tipping-settlement-will-reduce-earnings-of-a/)
36. [If we can afford our current welfare system, we can afford basic income](/blog/if-we-can-afford-our-current-welfare-system-we-can/)
---
## A year of ChatGPT is a fifth of a beer
**URL:** https://maxghenis.com/blog/drinking-ai/
**Published:** Jul 05 2026
**Description:** I built an interactive that prices drinks in AI queries — and lets you flip the accounting scope and the workload unit that make published AI water numbers differ by 100x.
Asking ChatGPT 30 questions a day for a year uses about a fifth of the water behind one beer. That counts both data-center cooling and the power plants generating the electricity — the closest AI analog to how the beer number counts the water behind the barley.
I built [an interactive](https://maxghenis.com/drinking-ai) that prices drinks this way: pick a beer, a coffee, a glass of milk, and see its water footprint in ChatGPT queries. One beer ≈ 53,300 queries. One coffee ≈ 66,200. A glass of tap water ≈ 178.
A year of 30 daily queries is 21.9 liters at the 2 mL scope; one beer is 106.5. The drinks give a scale; the reader picks the scope.
## Why every AI water number differs from every other one
Published water-per-query figures span more than 100x: [Google measured 0.26 mL](https://cloud.google.com/blog/products/infrastructure/measuring-the-environmental-impact-of-ai-inference/) per median Gemini prompt, [Sam Altman cited 0.000085 gallons (0.32 mL)](https://blog.samaltman.com/the-gentle-singularity) for the average ChatGPT query, [a benchmarking preprint](https://arxiv.org/abs/2505.09598) models a short GPT-4o query near 2 mL, and [Mistral's life-cycle assessment](https://mistral.ai/news/our-contribution-to-a-global-environmental-standard-for-ai) reports 45 mL per 400-token response.
These numbers barely disagree about physics. They disagree about where to draw the boundary:
- **0.26–0.32 mL**: Google defines its figure as on-site data-center cooling water only; OpenAI states no boundary, and I treat it the same because it is within 25% of Google's.
- **~2 mL** adds the water evaporated at the power plants generating the query's electricity.
- **45 mL** is a life-cycle assessment — upstream electricity, cooling, and embodied hardware, no training — of one 400-token Le Chat response, a response about as short as the 2 mL query; the gap is boundary and data-center location, not length.
The calculator defaults to the middle scope because it is the closest match to how the drink side is measured. Beer's footprint — 298 liters per kilogram, from [the standard crop-water tables](https://hess.copernicus.org/articles/15/1577/2011/hess-15-1577-2011.pdf) — counts the water behind the barley, not just the brewery's taps. The consistent AI analog counts the water behind the electricity, not just the cooling loop. Counting only cooling is like measuring beer by what the brewery pours in.
You can flip the scope yourself. The comparison survives every setting: even at Mistral's 45 mL, one beer still equals more than 2,000 responses.
## Agents change the unit
The published per-query figures describe one short text prompt. Reasoning models emit several times the tokens of a chat response, and agents chain many calls per task — [Anthropic measured its multi-agent systems at about 15x](https://www.anthropic.com/engineering/multi-agent-research-system) the tokens of a chat interaction. Since water scales with tokens, the calculator lets you count in reasoning responses (10x — I asked Claude for a mid-range estimate of the 2.5–50x published spread) or agentic tasks (15x).
One beer = 3,550 agentic tasks at the default scope: 10 a day for a year. At Mistral's 45 mL it is about 160.
## The aggregate is the same picture
[A Patterns paper](https://pmc.ncbi.nlm.nih.gov/articles/PMC12827721/) models all AI systems worldwide at 312.5–764.6 billion liters of water in 2025. The US drink categories I could source from government and industry data:
| Category | Year | Water footprint (liters) | vs global AI |
|---|---|---|---|
| Coffee | 2024/25 | 24.9 trillion | 33–80x |
| Milk | 2025 | 19.8 trillion | 26–63x |
| Soda | 2025 | 15.3 trillion | 20–49x |
| Beer | 2023 | 7 trillion | 9.2–22x |
| Plant-based milk | 2025 | 0.4 trillion | 0.5–1.3x |
| All nine tracked | — | 74.4 trillion | 97–238x |
US coffee's footprint is 33–80x global AI's. Extrapolating AI to 2026 (I asked Claude to estimate growth: 1.5x, within the published 1.3–2.45x) puts the nine categories at 65–159x the projected range.
## What this doesn't say
Water footprints measure volume, not scarcity. A liter evaporated from a stressed aquifer in Arizona is not a liter of rain on Bavarian barley, and the comparison doesn't pretend otherwise.
The local version of the question — one campus drawing from one watershed, peaking in summer — is a siting and water-rights question this page does not address.
## Provenance
Every figure on the page links to its source — the Mekonnen & Hoekstra crop and animal water tables, USDA and NIAAA volume data, and the per-query figures above. The three Claude-estimated numbers (the 2026 projection, the 10x reasoning multiplier, and the footprint blend for plant-based milk) are flagged as estimates on the page. The constants and derivations live in [a tested module](https://github.com/MaxGhenis/maxghenis.com/blob/master/src/lib/drinking-ai.ts) — a vitest suite pins the headline numbers, so if a source changes, the page fails loudly instead of drifting quietly. If you find a better source for any figure, tell me and I'll swap it in.
[Andy Masley's post](https://blog.andymasley.com/p/individual-ai-use-is-not-bad-for) prompted the question.
Drink responsibly. Prompt freely: [maxghenis.com/drinking-ai](https://maxghenis.com/drinking-ai)
---
## MacKenzie Scott's giving, in QALYs
**URL:** https://maxghenis.com/blog/mackenzie-scott-giving-in-qalys/
**Published:** Jun 28 2026
**Description:** An interactive cost-effectiveness model for MacKenzie Scott's $26B in philanthropy — and the AI prompts that built it.
MacKenzie Scott gave away [$7 billion in 2025](https://www.cnbc.com/2025/12/13/mackenzie-scott-revealed-her-total-charitable-donations-for-2025.html) — [about a third of every megagift in America that year](https://fortune.com/2026/06/25/mackenzie-scott-largest-megadonor-2025-7-billion-donations-giving-usa-iu-report/), by the Indiana University Lilly Family School of Philanthropy's count. Almost none of it is denominated in health. I built an interactive tool that asks what her [$26 billion](https://yieldgiving.com/) in lifetime giving buys in the unit health economists use to compare lives — quality-adjusted life-years: **[maxghenis.com/mackenzie-scott-qaly](/mackenzie-scott-qaly)**. Drag the assumptions and a Monte Carlo cost-effectiveness model reruns in your browser; on the best-guess default it lands around 202,000 QALYs — a model output, not a measured fact, which is why the tool exists: move the assumptions yourself.
[](/mackenzie-scott-qaly)
It runs the same machinery as the [GiveWell cost-effectiveness replication](/blog/givewell-cea/) I built in February — editable parameters, Monte Carlo, sensitivity analysis — pointed at a different question. GiveWell scores the most cost-effective charities in the world. This points the same lens at one donor's actual $26B portfolio, most of it unrestricted gifts to US organizations, where almost none of the spending is denominated in health to begin with.
## Why I care about the gap
I [took the Giving What We Can pledge in 2019](/blog/why-ive-taken-the-giving-what-we-can-pledge/) for one reason: the global poor benefit far more from a dollar than I do. This tool puts a number on that for Scott's giving. Measured purely in health, a dollar at the [global-health frontier](https://www.givewell.org/charities/amf) — bed nets averting child deaths — buys roughly 500 times the QALYs of her blended portfolio. That's a marginal comparison; redeploying the full $26 billion would compress it toward the floor — around 20–40× even if every dollar became [direct cash](https://blog.givewell.org/2024/11/12/re-evaluating-the-impact-of-unconditional-cash-transfers/) — [the tool page works through the arithmetic](/mackenzie-scott-qaly).
How much that gap should bother you is the open question, and the QALY count doesn't settle it. Part of the gap is outside her control: preventing a death costs far more in a rich country than a poor one, and a QALY ignores the income, opportunity, and rights her giving targets. Part of it is a choice: $26 billion is enough that where it goes carries real opportunity cost, measured in lives. The tool hands you the number, not the verdict. It's the same arithmetic that sends my own giving abroad.
## Two models, arguing
I built this with two coding agents, and the more useful part was letting them check each other. Claude Code wrote the model and the tool. Then I had Codex review the assumptions cold: it caught a real error — a cost-per-life figure I'd left in old dollars without inflating it — and disagreed with Claude on whether the global benchmark belongs in QALYs or DALYs. I'm centralizing on QALYs. Two models disagreeing about a modeling choice is a sharper adversarial review than either alone.
Two further full review rounds (both models, against everything) caught more, and the fixes moved the headline from ~98,000 to ~87,000: gifts are now recorded as the exact disclosed tranches and inflated to 2026 dollars year by year instead of divided nominal-vs-current; the community-health-center figure now uses the paper's own [~$54k per life-year](https://pmc.ncbi.nlm.nih.gov/articles/PMC4436657/) with an explicit life-year→QALY conversion (the old version skipped it); several 2000s-era cost-effectiveness anchors got inflated to current dollars; the frontier benchmark was re-derived twice — first for discounting consistency, then onto [GiveWell's current program averages](https://www.givewell.org/impact-estimates) — cutting the headline multiple by about 40%; the benefit/cost ratio now uses [HHS's published value per QALY](https://aspe.hhs.gov/reports/standard-ria-values) instead of a per-life-year value applied to QALYs; and one citation was re-attributed to the paper the numbers actually come from — [Sommers (2017)](https://www.journals.uchicago.edu/doi/10.1162/ajhe_a_00080), verified against the PDF, after I'd confidently planted the wrong one. Every correction made the model more skeptical or more honest, none was caught by a human, and the errors had survived earlier review passes.
The allocation across causes was still my prior, though — I'd accepted "Scott doesn't publish dollars-by-cause" without checking hard enough. She publishes better: [Yield Giving's gift database](https://yieldgiving.com/gifts) itemizes dollar amounts for about two-thirds of the money, with focus areas on every disclosed dollar. Deriving the split from her own data — 53 org-reported areas mapped onto the model's 13 archetypes, [every rule documented](https://github.com/MaxGhenis/mackenzie-scott-qaly/blob/main/data/yieldgiving/leaf_to_archetype.yaml) — cut the skeptical median from ~87,000 to ~70,000. Her real portfolio holds less food, housing, and cash assistance than I'd assumed (the buckets with the strongest health evidence) and more workforce development and nonprofit infrastructure (which have almost none). The undisclosed third isn't dropped: her essays give each year's total, so the residual is a known dollar amount, spread over that year's undisclosed gifts in proportion to the recipient's pre-gift [IRS 990 revenue](https://projects.propublica.org/nonprofits/) raised to an elasticity fit on the disclosed pairs, with the fuzzy name-to-EIN matches audited by a third model, a small one, against the live API. That fit is a finding in its own right: across 1,313 disclosed gift–revenue pairs, gift size scales with organization revenue to the power 0.41 — a 10× bigger organization gets about 2.5× more money. Her giving is far flatter across organization size than proportional. The imputed third moves each cause share by at most ~1.5 points — the disclosed two-thirds was representative.
The largest correction came from a reader. The model originally priced every health dollar at US anchors, so her gifts to organizations delivering abroad — [Malaria Consortium](https://yieldgiving.com/gifts), GiveDirectly's global program, Living Goods, Muso, Amref, Partners In Health, Evidence Action — were understated by orders of magnitude: bed nets at Medicaid rates. Health dollars now split by each organization's reported service locations, with the non-US share (~5% of her giving) priced at global-health anchors. That one fix tripled the skeptical median, from ~70,000 to ~205,000 QALYs, and it means most of the model's measured health impact comes from the small slice of her portfolio that reaches low-income countries. A follow-up audit of the 50 largest organizations that report only "global" service locations — cross-checked by two independent AI reviews against each organization's own filings — found many operate substantially in the US, trimming the default to ~202,000 (v1.1; the tool's version toggle keeps the July 13 allocation one click away).
None of the model code was hand-written; the whole thing was natural-language prompting. For transparency, here is every prompt I typed, verbatim — typos and all.
## Appendix: the prompts
**To Claude Code:**
1. estimate the qaly impact of mckenzie scott's lifetime donations
2. do better than judgment calls. make an actual quant analysis givewell-cea style. make a repo for it if thats useful
3. no just qalys for now. and considfer evidence quality especially causal identification credibility
4. /cycle
5. shall we make it an interactive and put it on maxghenis.com (using all the design tokens)? like a real py package with ci, nextjs/tailwind etc
6. inline link any falsifiable claims ALWAYS ALWAYS REMEMBER THIS
7. whats the global frontier
8. is it a falsifiable claim
9. fix
10. codex made some changes wdyt
11. give me a prompt for it to fix. but i think there was a misunderstanding i didnt ask for a qaly/daly change, could we centralize around one? which?
12. if you were starting this project from scratch howd you do it
13. yes
14. do a full review of this, both you and with a sol subagent
15. yep go - and dont have anyu allegiance to existing code, feel free to rebuild any and all things froms cratch
16. do it all
17. make sure the prose isnt bogged down by the history - like it seems like leading with the inflation adjustment in the tool is just based on it having been a bug we just fixed. users should see the tool first thing when they get to the page
18. can we give the blog post like a button at the top to use the tool, like we do for rambar etc
19. whats the delta against givewell? noting scaling limitations etc
20. sure yeah does it belong in the blog or the tool? like you could imagine a bigger version of the tool that also combines it with givewells for a general qaly estimator for every intervention and portfolio but not now
21. Vs. global frontier / 1,212× / more health per marginal $ — this makes it sounds like scott's is 1212x more effective
22. maybe we should drop the card and just discuss it in the written stuff...
23. then could you just make the 90% interval below in small text instead of a separate card, also add to the $/qaly
24. i dont get these two - whats the 2.1x wrt? whats 64.2b?
25. scott doesnt publish *any* dollars by cause? i thought she at least published her list of orgs, like are we basing this on any real data?
26. undisclosed ones can we do some research on, if need be we could scale by total revenue in the relevant year from 990s
27. what model are we using for this
28. i dont think we need fable for this fan out do we? could probably do like 5.6 terra or something
29. yeah terra is the small codex tier, use it for the audit
**To Codex:**
1. review the math/assumptions in https://maxghenis.com/mackenzie-scott-qaly (its a local repo)
2. would py wasm be more parsimonious and dry
3. fix it all
4. deployed?
5. yes do
The model, tests, and sources are on [GitHub](https://github.com/MaxGhenis/mackenzie-scott-qaly).
---
## The US government is gutting Anthropic's R&D capacity
**URL:** https://maxghenis.com/blog/us-gutted-anthropic-rd/
**Published:** Jun 15 2026
**Description:** A Commerce export control pulled Claude Fable 5 and Mythos 5 on Friday. It's cutting roughly a third of Anthropic's own researchers off from the frontier models they build, while every competitor's keep theirs.
import CrosspostAware from '../../components/CrosspostAware.astro';
Anthropic released Claude Fable 5 last week. I used it for three days before the government switched it off.
I spent Tuesday through Thursday at a convening and Friday traveling without connectivity, so I worked at about a quarter of my usual pace. In that window I pointed Fable at a range of projects, including the microdata pipeline that PolicyEngine's tax and benefit modeling runs on. That pipeline combines household surveys, administrative records, and calibration against published totals into a synthetic population for federal, state, and local analysis. We had spent months rebuilding it. Fable redesigned the architecture across the dependent packages, rebuilt the pipeline, and produced a dataset about five times more accurate than the one it replaces, measured against our held-out targets. That work reaches users this week. In the same three days it re-engineered the agentic flows behind several systems we ship over the coming weeks.
That is one model, used part-time, for three days. The effect that matters most from Friday's order is not on its customers. It is on Anthropic's own researchers.
## What the order does
On Friday the Commerce Department directed Anthropic to suspend access to Fable 5 and the related Mythos 5 model for any foreign national, inside or outside the United States, including the company's own employees. Commerce used the Export Administration Regulations' ["deemed export" rule](https://www.bis.gov/learn-support/deemed-exports/what-deemed-export), which treats releasing controlled technology to a foreign person in the country as an export to their home country. Anthropic [could not screen its customers by nationality in real time, so it pulled both models for all users](https://www.anthropic.com/news/fable-mythos-access). Claude Opus 4.8 and the earlier models stayed up.
That global cutoff is the commercial cost. The research cost is narrower and more specific.
## Who it bars
Anthropic knows each employee's status, so the cut is precise: its foreign-national researchers lose access to the models they build, while their citizen and green-card colleagues keep using them. The deemed-export rule the order invoked [exempts citizens, green-card holders, and "protected individuals"](https://www.bis.gov/learn-support/deemed-exports/what-deemed-export); the bite lands on employees on temporary visas. ([One outlet read the order to cover green-card holders too](https://www.i-scoop.eu/fable-5-blocked-outside-the-us-after-government-export-order/); the rule it invoked says the opposite.)
How murky even that line is showed up within a day. Headlines reported that [Andrej Karpathy, who joined Anthropic in May to lead its recursive self-improvement work, had been locked out of the models he was hired to build](https://www.ibtimes.co.uk/ai-expert-andrej-karpathy-anthropic-tech-regulation-1802715) for lacking US citizenship. But the claim that he was barred traced to [a single viral post sourcing his immigration status to an AI chatbot](https://x.com/AndrewCurran_/status/2065619713485627829) — no one verified it, and if he holds a green card the order never reached him at all. For days no one could say cleanly whether one of the field's best-known researchers was in or out.
## How many
Frontier AI runs on immigrant talent: about [two-thirds of top-tier AI researchers in the United States earned their undergraduate degrees abroad](https://macropolo.org/interactive/digital-projects/the-global-ai-talent-tracker/), and international students are [the core of the US AI PhD pipeline](https://cset.georgetown.edu/article/international-graduate-students-critical-to-u-s-ai-competitiveness/). No public figure exists for Anthropic, so I asked Claude to estimate it — five independent passes, each from a different angle: the AI-talent pipeline, Anthropic's visa filings, the China–India green-card backlog, the company's youth, and peer-company workforces. They ranged from 27% to 42% of its researchers, with a mean near 36% — on the order of a third. The passes agree on why it runs high: Anthropic is only four years old, so few foreign-origin researchers have had time to naturalize, and [green-card backlogs for Chinese and Indian nationals](https://www.dwt.com/insights/2020/05/us-bis-china-technology-export-ban) — the field's two largest talent sources — run a decade or more, leaving even long-tenured staff on temporary visas. It is an estimate, not a count: Anthropic publishes no breakdown, and its [H-1B filings](https://h1bgrader.com/h1b-sponsors/anthropic-pbc-10j7p51g0v) are flow, not stock.
The barred researchers also have an obvious exit. The control attaches to Anthropic's models, not to the people, so a visa-holder cut off inside Anthropic can join OpenAI, Google, or any other lab and use a frontier model there on day one — every researcher at those labs, visa holders included, keeps access, because the order names only Anthropic's models. So it does not just bench a third of Anthropic's researchers; it hands every competitor a standing recruiting pitch aimed squarely at them, turning a model restriction into a talent-export pump pointed at one company.
## Even if it's temporary
None of this needs to be permanent to matter. Prediction markets expect a temporary pause — both [Polymarket](https://polymarket.com/event/us-government-rescinds-claude-fable-5-foreigner-ban-by-20260613202427937) and [Kalshi](https://news.kalshi.com/p/fable-5-odds-anthropic-access-restored-july-57-percent) have access most likely returning within a few weeks. But the field doesn't pause with it, and it's speeding up: the length of tasks frontier models can do reliably has been [doubling every seven months — lately closer to four](https://metr.org/time-horizons/), a doubling time that keeps shrinking. On a curve that steep and still accelerating, weeks on the sidelines — while rivals keep training, shipping, and hiring — shift leads and define careers. A temporary pause here is not a small one.
## The incentive it sets
Applied as a standard, this penalizes shipping. Every frontier release would cut a lab's visa-holding researchers off from the model they just shipped, the one they need to build the next. A lab that ships nothing keeps its researchers' access; a lab that ships loses it.
## The path back
A mechanism exists. The deemed-export rule routinely lets foreign nationals work on controlled technology under an [individual license](https://www.bis.doc.gov/index.php/licensing/14-policy-guidance/deemed-exports/109-guidelines-for-foreign-national-licenses), usually backed by a technology control plan; semiconductor and aerospace firms employ non-citizens on restricted work this way. But BIS reviews each license under the rules for exporting to that employee's country of citizenship. For nationals of close allies, approval is routine. For nationals of countries the United States restricts, including China, the review [carries a presumption of denial](https://www.dwt.com/insights/2020/05/us-bis-china-technology-export-ban). The route back is per-employee and country-by-country, measured in months, and it would return some of the workforce while leaving the rest in review.
## On a consistent standard
The order's authority is the [Export Control Reform Act of 2018](https://www.justsecurity.org/142745/law-anthropic-export-controls/), applied through an "is-informed" letter to a single company — the first time it has reached an AI model — and [Lutnick gave no basis](https://www.axios.com/2026/06/12/anthropic-trump-mythos-fable-national-security) for why Anthropic specifically. Legal analysts called it "not clear why Anthropic's models were singled out … as compared to other U.S. large language model AI companies," and "a remarkable contrast" with the administration's hands-off stance on AI exports.
The directive names only Anthropic's two models, and the jailbreak behind it is disputed. Anthropic says the government's evidence was [verbal, and demonstrated by a competitor](https://cyberscoop.com/us-government-anthropic-fable-5-mythos-5-export-controls/), and that the same capability is available in [OpenAI's GPT-5.5, unrestricted](https://www.anthropic.com/news/fable-mythos-access). Security researcher Katie Moussouris [described the technique](https://techcrunch.com/2026/06/15/the-us-governments-anthropic-models-ban-was-never-about-an-ai-jailbreak/) as the gap between asking a model to "review code for security issues" and to "fix this code," and said it "cannot meaningfully be fixed, and any attempt would only weaken the model for defense."
The fixation on Anthropic predates the order. On a May 8 podcast — weeks before Anthropic overtook OpenAI in valuation — the administration's AI czar, David Sacks, [called it "the most powerful monopoly ever created in human history"](https://www.youtube.com/watch?v=10MdOvK-aG4&t=982s) and likened it to Standard Oil rebranded as "Safe Oil," casting its safety record as cover for a monopoly. On the same show, Sacks said the cyber capability was not Anthropic's alone — ["OpenAI now has a model that's just as cyber capable as Mythos,"](https://www.youtube.com/watch?v=10MdOvK-aG4&t=2713s) with every major lab and Chinese models to follow within months — and his prescription then was industry-wide hardening, not a restriction on one company.
David Sacks's "Safe Oil" monologue — All-In, May 8, 2026
“Unless something about their current trajectory changes, Anthropic will be the most powerful monopoly ever created in human history. […]
I just want you to think for a second about the case of John D. Rockefeller, who I think is known as probably the most successful, most ruthless monopolist in American history. But he wasn't very good at PR. He was terrible at PR. Everyone sort of recognized how ruthless he is. We've seen movies like There Will Be Blood, which is basically about him.
In any event, imagine if John D. Rockefeller was way better at public relations, and instead of calling his company Standard Oil, he called it Safe Oil. … He called it Safe Oil because, as we know, kerosene is dangerous. Their first big product was kerosene. And kerosene can light your house or it can burn it down, and in the wrong hands it can torch a city or you can use it to make a bomb. So John D should have called for the creation of a new government agency to regulate the safety of his product, and they could have done rigorous testing, licensing, common-sense regulation. … And I think people would have gotten so wrapped up in this debate over what constituted Safe Oil or Safe Kerosene that they would have missed what was really going on, which is that Rockefeller was building the richest, most powerful monopoly of all. In fact, people might even have called Rockefeller an Effective Altruist, because of course he was so concerned about the safety of his product.”
The day before the export control, the same hosts opened their show attacking Fable for the opposite sin — too much caution. They called its guardrails "Orwellian" and its 30-day data retention "mandatory surveillance," and Sacks called the safety posture ["a very sophisticated regulatory capture campaign based on fear-mongering."](https://www.youtube.com/watch?v=gH4FTjDm9FQ&t=602s) Later in the same episode, asked about Bernie Sanders's plan to seize a government stake in the AI companies, Sacks said he ["may be okay with Bernie's idea"](https://www.youtube.com/watch?v=gH4FTjDm9FQ&t=3364s) — but only for "a public benefit corporation that says it's going to cause massive job loss, that trained for free on humanity's knowledge but gatekeeps and refuses to give back," and "maybe it should be 75%." Sacks offered those as principled conditions, but they fit OpenAI as well as Anthropic: both are public benefit corporations, both trained on the same public data, both keep their weights closed, and both CEOs have predicted mass job loss. A test that singles out one of two near-identical companies isn't a test; it's a target with criteria attached. The next day the government restricted Anthropic, and only Anthropic, for being insufficiently safe.
No published standard tells another lab what to avoid, or tells Anthropic what to fix — only that, for now, the one company singled out must route its own researchers through a months-long license to use the model they built, or watch them rebuild it somewhere else.
The deeper irony is that standards-based, politically neutral governance is what Anthropic spent years asking for. It proposed uniform disclosure rules for every frontier lab in ["The Case for Targeted Regulation"](https://www.anthropic.com/news/the-case-for-targeted-regulation), and days before the order, Dario Amodei's ["Policy on the AI Exponential"](https://darioamodei.com/post/policy-on-the-ai-exponential) argued that any government power to block a model "must be scoped to … four specific risks" with "protective measures against political favoritism or arbitrary decisions." Defending the order, Sacks [wrote that Anthropic "asked for government regulation of Mythos"](https://x.com/DavidSacks/status/2065853007619588171) — but its proposals were model-neutral, applied to all frontier developers, and it had gated Mythos itself rather than ask Washington to single it out. Anthropic asked to be governed by a standard applied to everyone; it got an ad-hoc decree applied to it alone.
Appendix: David Sacks's "Safe Oil" monologue
All-In, May 8, 2026
“Unless something about their current trajectory changes, Anthropic will be the most powerful monopoly ever created in human history. […]
I just want you to think for a second about the case of John D. Rockefeller, who I think is known as probably the most successful, most ruthless monopolist in American history. But he wasn't very good at PR. He was terrible at PR. Everyone sort of recognized how ruthless he is. We've seen movies like There Will Be Blood, which is basically about him.
In any event, imagine if John D. Rockefeller was way better at public relations, and instead of calling his company Standard Oil, he called it Safe Oil. … He called it Safe Oil because, as we know, kerosene is dangerous. Their first big product was kerosene. And kerosene can light your house or it can burn it down, and in the wrong hands it can torch a city or you can use it to make a bomb. So John D should have called for the creation of a new government agency to regulate the safety of his product, and they could have done rigorous testing, licensing, common-sense regulation. … And I think people would have gotten so wrapped up in this debate over what constituted Safe Oil or Safe Kerosene that they would have missed what was really going on, which is that Rockefeller was building the richest, most powerful monopoly of all. In fact, people might even have called Rockefeller an Effective Altruist, because of course he was so concerned about the safety of his product.”
---
## How to turn money into predictions
**URL:** https://maxghenis.com/blog/how-to-turn-money-into-predictions/
**Published:** May 26 2026
**Description:** A menu of mechanisms for organizations that want to fund forecasts on policy, economic, and AI outcomes.
Organizations that fund policy research, economic analysis, or AI deployment want predictions about future outcomes — which tax provisions will pass, what inflation will run, how labor markets will absorb a new model class. The [Federation of American Scientists is sponsoring forecasting tournaments](https://fas.org/publication/fas-and-metaculus-are-using-forecasting-to-support-better-climate-policy/) to inform climate policy; the [Forecasting Research Institute ran its Existential Risk Persuasion Tournament](https://forecastingresearch.org/xpt) with 80 domain experts and 89 superforecasters on AI and other catastrophic risks; [FiscalNote announced a major expansion](https://fiscalnote.com/newsroom/fiscalnote-announces-major-expansion-into-political-prediction-markets) into prediction-market infrastructure for its policy-intelligence customer base in February 2026; and AI researchers are now [benchmarking LLM forecasting against expert and crowd panels](https://www.forecastbench.org/paper.html). Quality varies sharply across markets — some are deep, actively traded, and accurate; others sit thinly populated and stay mispriced for weeks. Funders can improve the quality on the questions that matter to them.
This post lays out the available mechanisms. I drafted an earlier version in 2024 and circulated it with a small group of researchers and funders. The landscape has changed enough since then — prediction markets [crossed $44 billion in notional volume in 2025](https://www.theblock.co/post/383733/prediction-markets-kalshi-polymarket-duopoly-2025), Kalshi got [Federal Reserve research validation](https://www.federalreserve.gov/econres/feds/files/2026010pap.pdf), and a [serious proposal](https://empiricrafting.substack.com/p/we-need-well-capitalized-prediction) for sponsorship-funded markets on AI labor impacts hit the literature — that an updated public version is overdue.
## Two structural problems
Two problems make most prediction markets shallower than they should be.
**The liquidity problem.** Markets need both informed traders ("smart money") and less informed traders willing to absorb their bets ("dumb money"). Andrew Gelman puts it cleanly in [Prediction markets and the need for "dumb money" as well as "smart money"](https://statmodeling.stat.columbia.edu/2024/10/25/prediction-markets-and-the-need-for-dumb-money-as-well-as-smart-money/): "Markets become efficient when making them efficient is profitable." Without enough total volume, smart traders can't profitably correct a mispriced market, so prices stay wrong.
A Manifold market asking ["Will there be any tax on unrealized capital gains in the USA before EOY2028?"](https://manifold.markets/Gen/will-there-be-any-tax-on-unrealized) sat at 22% probability for nearly two weeks after Trump's 2024 election victory, with 27 traders. The election outcome obviously cut the probability substantially, but no one was paid enough to move the price.
A parallel Manifold market on ["Will a raise to the top capital gains tax rate be enacted in 2025?"](https://manifold.markets/jack/will-the-top-capital-gains-tax-rate) behaved differently. Manifold staff made it sweepstakes-eligible at creation, so the market traded in real money (Manifold's then-experimental real-cash mode, [shut down in March 2025](https://news.manifold.markets/p/focusing-on-mana-bringing-sweepstakes)). It held 20-21% until election results, then dropped to 6%. The mana version attracted 18 trades and 4,093 mana; the sweepstakes version had 9 traders and 334 sweepcash. The cash version attracted enough attention to update on election news; the play-money version did not. Money matters here mostly as an attention allocator, not as a calibration mechanism — [Servan-Schreiber et al. (2004)](https://users.nber.org/~jwolfers/Papers/DoesMoneyMatter.pdf) found play-money and real-money NFL prediction markets equally accurate when both had engaged trader bases.
**The zero-sum challenge.** Unlike equities, where investors collectively benefit from economic growth, prediction markets are zero-sum or negative-sum after fees. Every winner needs a loser. Zvi Mowshowitz lists the consequences in ["Subsidizing Prediction Markets"](https://www.lesswrong.com/posts/AeKS2m6uLM8RYfvND/subsidizing-prediction-markets): no natural long-term investors, higher perceived risk, and limited professional capital deployment. Funders willing to take an expected loss on a market are the substitute for that missing long-horizon capital.
## Where prediction markets are in 2026
**Volume.** Total notional [exceeded $44 billion in 2025](https://www.theblock.co/post/383733/prediction-markets-kalshi-polymarket-duopoly-2025), with monthly peaks around $13 billion during the US election period. The market is concentrated in [Kalshi](https://kalshi.com/) (a [CFTC-designated contract market](https://www.cftc.gov/PressRoom/PressReleases/8302-20)) and [Polymarket](https://polymarket.com/) (crypto-native, recently [returned to the US](https://www.coindesk.com/policy/2025/09/03/u-s-cftc-gives-go-ahead-for-polymarket-s-new-exchange-qcx) via a $112 million acquisition of CFTC-licensed exchange QCX).
**Regulatory legitimacy.** Kalshi went from a startup with questionable contract approvals to the subject of a [Federal Reserve working paper](https://www.federalreserve.gov/econres/feds/files/2026010pap.pdf) (Diercks, Katz, and Wright, 2026) that finds Kalshi's prices on inflation, jobs, GDP, and FOMC decisions "outperform surveys and interest-rate futures" as macro forecasting tools. The paper explicitly recommends Kalshi as a real-time benchmark for researchers and policymakers. That is the most credible institutional validation prediction markets have ever received.
**A sponsorship thesis.** Andrey Fradkin, Brian Jabarian, and Andrew Koh published ["We need well-capitalized prediction markets"](https://empiricrafting.substack.com/p/we-need-well-capitalized-prediction) arguing that the right model is sponsorship — AI labs, Big Tech firms, and philanthropies seed liquidity in markets on labor outcomes, AI capability benchmarks, and other questions where they have decision-relevant interest. Sponsorship covers the expected loss in exchange for better information for everyone. I'll come back to this in the last section.
## Funding mechanism 1: Subsidize public markets
Each major platform supports liquidity provision through a different mechanism.
[Manifold Markets](https://manifold.markets/) runs on play money called mana. The platform briefly offered sweepstakes trading in real cash but [shut that program down in March 2025](https://news.manifold.markets/p/focusing-on-mana-bringing-sweepstakes) and returned to a pure play-money model. Despite the play-money currency, Manifold's [aggregate calibration](https://manifold.markets/calibration) is strong — predictions land within roughly four percentage points of realized frequencies — and Manifold's own analysis finds calibration improves with trader count up to roughly 10-20 traders per market, after which it plateaus. The implication for funders: subsidizing a market that's stuck below the plateau directly buys better predictions. Organizations can [purchase mana with dollars](https://manifold.markets/checkout) and [add it to specific markets as liquidity](https://docs.manifold.markets/faq#how-do-i-subsidise-a-market); the deeper book attracts more traders and tightens the spread. The subsidizer expects to lose mana — the more the market price moves, the less you recover — but the lost mana is the cost of a better-priced market, not a tax on the platform. Manifold's code is also [MIT-licensed](https://github.com/manifoldmarkets/manifold), which makes it the friendliest platform for organizations that want to run programmatic experiments at low cost.
[Polymarket](https://polymarket.com/) uses a [Liquidity Mining & Rewards program](https://docs.polymarket.com/market-makers/liquidity-rewards). Market makers earn rewards for providing consistent quotes; traders earn fee rebates proportional to trading activity; community bounties incentivize market promotion. Organizations can participate through these mechanisms rather than as standalone subsidizers. After a [2022 CFTC settlement](https://www.cftc.gov/PressRoom/PressReleases/8478-22) that barred US users for nearly four years, Polymarket regained US access in December 2025 via the QCX acquisition above, and runs [hundreds of macro markets](https://polymarket.com/dashboards/macro) globally.
[Kalshi](https://kalshi.com/) offers a [limit order](https://help.kalshi.com/trading/order-types/limit-orders) liquidity model. Organizations place standing limit orders at prices they're willing to trade at; orders execute when the market crosses those prices. Limit orders avoid trading fees, allow precise price targeting, and accumulate into market liquidity over time. Qualified market makers can join a formal [market maker program](https://help.kalshi.com/en/articles/13823819-market-maker-program) with additional benefits. The CFTC approval and Fed validation make this the right platform for institutional-grade questions where regulatory clarity matters.
[Hypermind](https://corp.hypermind.com/) has run institutional prediction markets since 2000, with clients including the US intelligence community and the [Johns Hopkins Center for Health Security](https://predict.hypermind.com/hypermind/media/pdf/futuribles-prediction-markets.pdf). Sponsors fund custom markets and panels of selected forecasters on bespoke questions, with the platform providing aggregation and reporting infrastructure. Less suitable for one-off retail experiments, more suitable for sustained corporate or government use.
## Funding mechanism 2: Sponsor a tournament
[Metaculus](https://www.metaculus.com/) is a forecasting platform aggregating predictions across thousands of binary and numerical questions, scored on accuracy. Organizations [sponsor dedicated tournaments](https://www.metaculus.com/tournaments/) — a slate of related questions with a prize pool that pays out to the best forecasters.
The [Federation of American Scientists](https://fas.org/publication/fas-and-metaculus-are-using-forecasting-to-support-better-climate-policy/) funded a $5,000 [Climate Tipping Points tournament](https://www.metaculus.com/tournament/climate/) covering 43 questions about climate policy and outcomes, including conditional questions. The tournament structure lets a funder set the agenda — what questions matter, what counts as resolution, what horizons to ask about — without taking on market-maker risk.
The [Forecasting Research Institute](https://forecastingresearch.org/) (Philip Tetlock's research arm, distinct from Good Judgment Inc) runs research-oriented tournaments like the [Existential Risk Persuasion Tournament](https://forecastingresearch.org/xpt) (80 domain experts, 89 superforecasters, thousands of forecasts on long-horizon catastrophic risks). Sponsoring FRI is closer to funding a research program than buying a forecast.
Tournaments work well for questions where you want a calibrated probability today, not a tradeable instrument over time. They also support reasoning prizes for the best-argued forecasts, not just the most accurate ones, which matters when you're trying to surface analytical talent rather than just numerical answers.
## Funding mechanism 3: Buy custom human forecasts
[Good Judgment](https://goodjudgment.com/) sells [custom forecasting from professional superforecasters](https://goodjudgment.com/services/custom-superforecasts/) — individuals identified by [Philip Tetlock's Good Judgment Project](https://goodjudgment.com/resources/the-superforecasters-track-record/) as [consistently outperforming domain experts and intelligence analysts](https://en.wikipedia.org/wiki/The_Good_Judgment_Project) (the project beat intelligence-community analysts with classified access by 25-30% in the IARPA ACE tournament). Services include question development, daily forecast updates, written analysis, and follow-up question support. Organizations can keep forecasts private or release them publicly.
This is the most concierge option. You write a question, a small team of trained forecasters works on it, you get back a probability with explanation. No liquidity to manage, no market to monitor. Cost is higher per question and the methodology is opaque relative to a public market.
## Funding mechanism 4: Buy AI-generated forecasts
A new category emerged in 2024-2026. [FutureSearch](https://futuresearch.ai/) ([publicly launched in early 2026](https://futuresearch.ai/company/)) runs LLM-based research agents that gather evidence, weigh base rates, and produce calibrated probability forecasts with reasoning. Pricing is [per-operation](https://futuresearch.ai/pricing/) — deep-research agents at 1-11¢, forecasters at 20-90¢ per researcher per row — orders of magnitude cheaper than human superforecasters per question, with the obvious tradeoff that the depth of domain expertise is whatever the model has internalized plus what its retrieval finds.
The category builds on published research: Halawi et al. (2024) ["Approaching Human-Level Forecasting with Language Models"](https://arxiv.org/abs/2402.18563) (NeurIPS) showed retrieval-augmented LLM systems approaching competitive-forecaster accuracy; the Center for AI Safety's [forecasting bot](https://safe.ai/blog/forecasting) reported superhuman accuracy on certain competitive platforms; [ForecastBench](https://www.forecastbench.org/paper.html) (Karger et al., ICLR 2025) — a dynamic, continuously-updated benchmark — finds frontier LLMs roughly match the general-public crowd in Brier score but still trail elite superforecasters by a wide margin (LLM Brier ≈ 0.135–0.159 vs. superforecaster Brier ≈ 0.02). The ForecastBench team's linear extrapolation [projects superforecaster parity around late 2026](https://forecastingresearch.substack.com/p/ai-llm-forecasting-model-forecastbench-benchmark), with a 95% confidence interval from December 2025 to January 2028.
The cost structure inverts the human-forecasting category. Where Good Judgment's value is depth on a few high-stakes questions, AI forecasting's value is breadth — hundreds or thousands of calibrated probabilities a day. Use AI forecasting when you want coverage; use human superforecasters when you want depth and defensibility on a question where the audience expects a human accountable for the call.
The trajectory matters more than the snapshot. Per-query costs are falling and accuracy is rising; within a few years, the marginal cost of a calibrated probability on most public questions drops sharply, even if the gap to elite superforecasters takes longer to close than the ForecastBench point estimate suggests. What remains expensive — and where real returns on investment will sit — is *optimizing* the agents: which tools they call, which substrates they query, how their calibration is benchmarked, how they update under new information. Cheap commodity forecasts and frontier-quality forecasts will look very different, and the work to produce the latter still needs to be incentivized.
## Funding mechanism 5: Run your own prediction market
For internal questions an organization doesn't want public — product launch dates, project completion probabilities, internal strategic questions — private prediction markets aggregate dispersed information already inside the organization. [Eli Lilly used internal markets in 2003](https://en.wikipedia.org/wiki/Prediction_market#Business) to forecast which drug candidates would advance through clinical trials; Google, Microsoft, and Ford have run variants. The corporate-prediction-market vendor landscape has consolidated — [Cultivate Labs](https://www.cultivatelabs.com/posts/why-we-stopped-supporting-prediction-markets) (formerly Inkling Markets) dropped market mechanics in 2022 in favor of opinion pools, citing user confusion — but [Hypermind's Prescience platform](https://corp.hypermind.com/) still offers managed private markets for corporate and government clients.
## The sponsorship thesis, and why it matters
The [Fradkin / Jabarian / Koh proposal](https://empiricrafting.substack.com/p/we-need-well-capitalized-prediction) is worth treating as its own category. They argue that the chicken-and-egg problem in prediction markets — no liquidity, no traders; no traders, no liquidity — can be solved by sponsorship capital from organizations that have real decision-relevant interest in the questions. Their canonical example is AI labor impacts: a major AI lab, a federal agency, or a philanthropy seeds liquidity on markets tied to [Bureau of Labor Statistics](https://www.bls.gov/) data series — occupation-level employment, wages, labor-force participation at 1-, 2-, and 5-year horizons.
The contract design they advocate has four properties: verifiability (objective resolution against published government data), stability (consistent measurement over time), robustness to gaming (large underlying quantities that one trader can't move), and attention (sufficient interest to draw informed traders). BLS data hits all four. So do many other government statistics — [BEA NIPA tables](https://www.bea.gov/itable/national-gdp-and-personal-income), [Census ACS variables](https://www.census.gov/programs-surveys/acs), [IRS Statistics of Income](https://www.irs.gov/statistics).
The model generalizes. Any organization that benefits from better forecasting on a specific question can sponsor a market on it without expecting trading profit. The sponsorship pays for information that improves everyone's decisions — including the sponsor's. Conditional markets (outcome Y given policy state X) are particularly underprovided today and particularly valuable for policy analysis. Sponsorship is the cleanest mechanism to bring them into existence.
The sponsorship lane connects to a structural question: who pays for the public good of calibrated forecasts? In 2024 the answer was mostly "individual hobbyist traders subsidizing inefficient markets." In 2026 the answer is starting to be "institutions with skin in the question, sponsoring markets that produce information they want." That shift is what makes 2026 different from 2024.
## Where to start
Different mechanisms suit different goals:
- **Want to test the waters cheaply?** [Buy mana on Manifold](https://manifold.markets/checkout) and seed a market that matters to you. Total commitment can be under $1,000. The mana you lose to traders moving the price is the cost of a deeper, better-calibrated market on a question you care about.
- **Want a probability anchored to your priorities?** [Sponsor a Metaculus tournament](https://www.metaculus.com/tournaments/). $5,000–$50,000 buys a slate of questions with a real prize pool and visible forecasters.
- **Want a single high-stakes forecast with human reasoning?** Engage [Good Judgment](https://goodjudgment.com/services/custom-superforecasts/). Highest cost per question, deepest analysis, accountable human author.
- **Want broad coverage across many questions cheaply?** Use [FutureSearch](https://futuresearch.ai/) or a similar AI-forecasting service. Per-query pricing in pennies, hundreds of forecasts a day, calibrated but synthetic.
- **Want regulated, real-money markets on macro questions?** Place [limit orders on Kalshi](https://help.kalshi.com/trading/order-types/limit-orders). Fee-free, precise pricing, supports the most institutionally credible platform.
- **Want an institutional sustained-engagement platform?** [Hypermind](https://corp.hypermind.com/) for custom corporate panels (25+ years of track record) or for running a private internal market with a defined trader population.
- **Want a research tournament on long-horizon questions?** Fund the [Forecasting Research Institute](https://forecastingresearch.org/) directly — closer to research-program sponsorship than to buying a forecast.
- **Want to move the field?** Sponsor markets that solve a real liquidity hole — BLS-anchored contracts, AI capability benchmarks, conditional policy-outcome markets. The [Fradkin / Jabarian / Koh model](https://empiricrafting.substack.com/p/we-need-well-capitalized-prediction) is the right starting point.
The hard part is no longer figuring out the mechanism. The mechanisms exist, the platforms work, the regulatory path is clearer than at any point in the past decade. The opportunity is institutional: adding prediction markets to the portfolio of truth-seeking tools organizations already use — alongside surveys, expert panels, internal modeling, and traditional forecasts — and developing the craft of writing forecastable questions. A well-formed question is conditional, scoreable, and decision-relevant ("will 2028 child poverty under reform X exceed Y%?"); a poorly-formed one is none of those. Prediction markets are a new epistemic technology. Organizations that build the discipline to use them well will reason more clearly about the future.
---
## Can Talkie-1930 do arithmetic?
**URL:** https://maxghenis.com/blog/talkie-1930-math-evals/
**Published:** Apr 30 2026
**Description:** I tested Talkie-1930 on GSM8K and the easier EleutherAI/OpenAI arithmetic suite, then packaged an lm-eval-harness runner so the runs are reproducible.
import CrosspostAware from '../../components/CrosspostAware.astro';
import TalkieEvalsEmbed from '../../components/TalkieEvalsEmbed.astro';
[Talkie](https://talkie-lm.com/introducing-talkie) is a 13B language model trained only on text available before 1931. Its launch post includes a "Numeracy" panel showing talkie-1930 reaching about 62% average accuracy at peak training compute, slightly above the modern-web twin's roughly 57%. But the plot doesn't specify which tasks the average covers, which prompts were used, or how answers were scored.
I wanted a smaller, inspectable check: can Talkie-1930 do arithmetic at all?
The answer depends on the evaluation. In full `lm-evaluation-harness` runs, the instruction-tuned model scores 0.0% under the strict GSM8K final-answer parser and 4.3% under the flexible parser in zero-shot. With the standard 5-shot GSM8K prompt, it rises to 4.9% strict and 7.2% flexible. On the easier EleutherAI/OpenAI arithmetic suite, the 1930 base model averages 42.7%, the 1930 instruction model averages 42.2%, and the modern-web base model averages 3.4%.
## GSM8K: Mostly no
I first tried [GSM8K](https://huggingface.co/datasets/openai/gsm8k), the grade-school math word-problem dataset OpenAI released in [Cobbe et al. (2021)](https://arxiv.org/abs/2110.14168). For context, the same-size [LLaMA 13B reports 17.8%](https://arxiv.org/abs/2302.13971) (Touvron et al. 2023, Table 7); GPT-4 hit [92.0%](https://cdn.openai.com/papers/gpt-4.pdf) at its March 2023 release; [Claude 3 Opus reported 95.0%](https://www.anthropic.com/news/claude-3-family) in March 2024. I used the instruction-tuned Talkie-1930 model, the Talkie instruction chat template, and greedy decoding (`do_sample: false`, `temperature: 0.0`).
| Run | N | Strict match | Flexible extract |
| --- | ---: | ---: | ---: |
| Zero-shot | 1,319 | 0.0% | 4.3% |
| 5-shot | 1,319 | 4.9% | 7.2% |
The strict score requires the GSM8K-style final answer marker (`#### number`). The flexible score extracts a number-like string from the model response. Flexible extraction is more forgiving when a model doesn't follow the GSM8K answer format, but it is still only 7.2% at 5-shot.
A representative correct 5-shot response:
> **Q:** A robe takes 2 bolts of blue fiber and half that much white fiber. How many bolts in total does it take?
>
> **Talkie:** It takes 3 bolts in total because 2+1/2 = `<<2+1/2=3>>`3 bolts. `#### 3` ✓
A typical miss:
> **Q:** Janet's ducks lay 16 eggs per day. She eats three for breakfast every morning and bakes muffins for her friends every day with four. She sells the remainder at the farmers' market daily for $2 per fresh duck egg. How much in dollars does she make every day at the farmers' market?
>
> **Talkie:** Janet makes $2 x 16 = `<<32>>`32 per day at the farmers' market. `#### 32` (gold: 18)
The model imitates the GSM8K answer format but skips the operations the problem actually requires.
Before running the full harness job, I also ran smaller custom probes: direct answer, zero-shot reasoning, and 4-shot reasoning prompts. Those got 0 of 70 attempts right. I now treat those as audit probes rather than headline benchmark results; the full harness runs above are the reproducible numbers.
GSM8K requires reading a word problem, tracking quantities, choosing operations, and formatting an answer. For Talkie, generation and instruction-following are themselves part of the bottleneck.
## Easier arithmetic: Sometimes yes
I then used the [EleutherAI arithmetic dataset](https://huggingface.co/datasets/EleutherAI/arithmetic), which comes from the OpenAI GPT-3 arithmetic tests in [Brown et al. (2020)](https://arxiv.org/abs/2005.14165). For context, GPT-3 175B in that paper (Table 3.9, few-shot) already hit 100% on 2-digit addition, 98.9% on 2-digit subtraction, and 80.4% on 3-digit addition; multiplication and 4–5-digit operations were weaker (29.2% on 2-digit multiplication, 9.3% on 5-digit addition). The current [lm-evaluation-harness task definitions](https://github.com/EleutherAI/lm-evaluation-harness/tree/main/lm_eval/tasks/arithmetic) score these as log-likelihood tasks: given a context like:
```text
Question: What is 98 plus 45?
Answer:
```
the model is correct if the exact target completion, like ` 143`, is the greedy continuation under teacher forcing.
This is much easier than GSM8K, and closer to a base-LM benchmark. I ran all 2,000 validation examples from each of the 10 arithmetic tasks.
| Task | 1930 base | 1930 instruct | Modern-web base |
| --- | ---: | ---: | ---: |
| Single-digit 3 ops | 11.5% | 16.3% | 3.4% |
| 2-digit addition | 91.6% | 75.7% | 14.0% |
| 2-digit subtraction | 51.0% | 49.8% | 11.5% |
| 3-digit addition | 74.7% | 88.3% | 0.6% |
| 3-digit subtraction | 47.2% | 48.7% | 1.7% |
| 4-digit addition | 29.5% | 27.7% | 0.0% |
| 4-digit subtraction | 36.1% | 30.7% | 0.1% |
| 5-digit addition | 31.4% | 24.4% | 0.0% |
| 5-digit subtraction | 28.2% | 30.1% | 0.0% |
| 2-digit multiplication | 26.2% | 30.8% | 3.1% |
| **Overall** | **42.7%** | **42.2%** | **3.4%** |
The 1930 models aren't just refusing. They often pick the right answer as their top continuation, especially for addition (91.6% base on 2-digit, 74.7% base and 88.3% instruct on 3-digit). For "What is 98 plus 45?" the 1930 base preferred the correct ` 143`. When the base model is correct on 2-digit addition, its median probability on the full target sequence is 51.6% — winning the argmax, not asserting confidently. The pattern breaks on multi-operation expressions, subtraction, multiplication, and larger digits.
The modern-web base scores 3.4% overall. In many errors it copied an operand instead of computing the result: for "What is 98 plus 45?" it preferred ` 98` over ` 143`; for "What is 92 times 7?" it preferred ` 92` over ` 644`. In a 500-row custom audit, 70% of its 2-digit addition rows and 25% of its 2-digit multiplication rows were exact operand copies. I don't read this as evidence that pre-1931 text makes a model more numerate than modern web text. More likely, I'm not reproducing the Talkie authors' exact benchmark setup, or this completion format interacts badly with the modern-web checkpoint.
## Compared to same-size GPT-3
The arithmetic suite first appears in [Brown et al. (2020)](https://arxiv.org/abs/2005.14165). The natural same-size baseline is GPT-3 13B; GPT-3 175B is the headline. Few-shot accuracies from Appendix H, Table H.1:
| Task | GPT-3 13B | GPT-3 175B | Talkie-1930 13B base |
| --- | ---: | ---: | ---: |
| Single-digit 3 ops | 9.95% | 21.3% | 11.5% |
| 2-digit addition | 55.5% | 100.0% | 91.6% |
| 2-digit subtraction | 52.4% | 98.9% | 51.0% |
| 3-digit addition | 8.4% | 80.4% | 74.7% |
| 3-digit subtraction | 9.2% | 94.2% | 47.2% |
| 4-digit addition | 0.4% | 25.5% | 29.5% |
| 4-digit subtraction | 0.4% | 26.8% | 36.1% |
| 5-digit addition | 0.05% | 9.3% | 31.4% |
| 5-digit subtraction | 0.0% | 9.9% | 28.2% |
| 2-digit multiplication | 7.05% | 29.2% | 26.2% |
Talkie-1930 13B base trails GPT-3 13B on 2-digit subtraction (51.0% vs 52.4%) and is essentially tied on single-digit composite (11.5% vs 9.95%). On the other eight tasks it scores higher than GPT-3 13B. On 4-digit and 5-digit add/subtract it also scores higher than GPT-3 175B, despite being 13× smaller.
A 13B model trained only on pre-1931 text matches or exceeds GPT-3 175B on most rote arithmetic completions. It also scores 4.9% strict on GSM8K word problems, against 17.8% for same-size LLaMA 13B. The capability gap with modern frontier models lives in instruction-following, chain-of-thought, and word-problem framing — not in the underlying numeric pattern matching.
## The metric is strict
The arithmetic score is format-strict. If the target is digits and the model prefers a word-form answer, the metric counts it wrong. On "What is 70 plus 15?" the instruction-tuned model gave ` Eighty` higher probability (logprob -0.29) than the leading space of the digit target ` 85` (-1.67). That's a legitimate miss under the benchmark, but a different kind of miss than computing the wrong number.
That's why I kept the raw outputs, not just aggregate scores. A single accuracy number hides the difference between "wrong operation," "copied an operand," "right value in the wrong format," and "format-following failure."
## Browse the responses
You can step through every Talkie response below — all 1,319 GSM8K questions across both runs and all 15,000 arithmetic rows across the three models. Filter by run, metric, and grade for GSM8K, or by candidate (1930 base, 1930 instruct, modern-web base) and subject for arithmetic.
*[Browse all 16,319 evaluation responses in the examination booklet on the full post](https://maxghenis.com/blog/talkie-1930-math-evals/)*
The standalone version: [talkie-evals-browser.vercel.app](https://talkie-evals-browser.vercel.app/).
## Reproducibility
I packaged the evaluator as [a small repo](https://github.com/MaxGhenis/talkie-evals) using Modal for CUDA. It includes an `lm-evaluation-harness` adapter for benchmark-style runs and custom audit commands that log row-level arithmetic traces. The package pins:
- the Talkie Python package commit,
- the Hugging Face model revisions,
- the arithmetic and GSM8K dataset revisions,
- the Modal image Python and pip packages,
- the sample seed for custom sampled probes, with row-level outputs written to JSON.
The full arithmetic harness run is:
```bash
uv run talkie-evals harness \
--model-names talkie-1930-13b-base,talkie-1930-13b-it,talkie-web-13b-base \
--tasks arithmetic \
--sample-size 0 \
--output results/lm_eval_full_arithmetic_all_models.json
```
The full zero-shot GSM8K harness run is:
```bash
uv run talkie-evals harness \
--model-names talkie-1930-13b-it \
--tasks gsm8k \
--sample-size 0 \
--num-fewshot 0 \
--talkie-chat-template \
--output results/lm_eval_full_gsm8k_zero_shot_chat.json
```
The full 5-shot GSM8K harness run is:
```bash
uv run talkie-evals harness \
--model-names talkie-1930-13b-it \
--tasks gsm8k \
--sample-size 0 \
--talkie-chat-template \
--output results/lm_eval_full_gsm8k_5shot_chat.json
```
The full raw result JSONs behind the tables are compressed in the repo:
- [Full arithmetic harness run](https://github.com/MaxGhenis/talkie-evals/blob/main/results/raw/lm_eval_full_arithmetic_all_models.json.gz)
- [Full GSM8K zero-shot harness run](https://github.com/MaxGhenis/talkie-evals/blob/main/results/raw/lm_eval_full_gsm8k_zero_shot_chat.json.gz)
- [Full GSM8K 5-shot harness run](https://github.com/MaxGhenis/talkie-evals/blob/main/results/raw/lm_eval_full_gsm8k_5shot_chat.json.gz)
The repo also keeps the earlier custom probe outputs:
- [Arithmetic audit run](https://github.com/MaxGhenis/talkie-evals/blob/main/results/raw/arithmetic_eval_talkie-1930-13b-base_talkie-1930-13b-it_talkie-web-13b-base_20260429_214800.json.gz)
- [GSM8K direct-answer probe](https://github.com/MaxGhenis/talkie-evals/blob/main/results/raw/gsm8k_eval_talkie-1930-13b-it_zero_shot_direct_20260429_170239.json.gz)
- [GSM8K reasoning probes](https://github.com/MaxGhenis/talkie-evals/blob/main/results/raw/gsm8k_eval_talkie-1930-13b-it_cot_20260429_170728.json.gz)
## Takeaway
Talkie-1930's capability profile is split along an unusual seam. On rote arithmetic completion, the 13B base matches or exceeds GPT-3 175B (the 2020 frontier) on most tasks while running at 13× fewer parameters. On grade-school word problems, the 13B instruction-tuned model scores 4.9% strict on 5-shot GSM8K, against 17.8% for same-size LLaMA 13B and 92.0%+ for any 2024-vintage frontier model. Same model. Two benchmarks. Different eras of capability.
The arithmetic suite is a better calibration check than GSM8K alone. GSM8K tells us the instruction-tuned model can't reliably solve generated word problems. The arithmetic suite tells us the base model still encodes elementary calculation patterns at GPT-3-175B level. The same instruction-tuned model scores 4.9% strict / 7.2% flexible on full 5-shot GSM8K and 42.2% on the arithmetic suite — elicitation and scoring can dominate the headline number.
The launch post's roughly 62% Numeracy figure averages over unspecified tasks. The public arithmetic suite shows an 11.5%–91.6% spread on the 1930 base model.
---
## Billionaires aren't 'just as likely' to be nonpayers as top taxpayers
**URL:** https://maxghenis.com/blog/madoff-billionaire-nonpayers/
**Published:** Apr 20 2026
**Description:** Checking Ray Madoff's claim on The Ezra Klein Show against ProPublica and PolicyEngine microdata.
On this week's [Ezra Klein Show](https://www.nytimes.com/2026/04/17/opinion/ezra-klein-podcast-ray-madoff.html), Boston College tax law professor Ray Madoff said:
> When it comes to the wealthiest Americans — Zuckerberg, Bezos, Musk, Larry Ellison, all the people we hear about so often — they are just as likely to be in the 40 percent of nonpayers as they are in the top 1 percent of payers.
Madoff's claim is about federal income tax (the measure the "40% of nonpayers" statistic describes). Everything below uses that same denominator.
The top 1% of payers by federal income tax paid has a threshold of $185,608 at the household level in PolicyEngine's 2026 microdata (below). For reference, the top 1% of filers paid 40.4% of all federal income taxes in TY2022 per [IRS SOI](https://www.irs.gov/statistics/soi-tax-stats-individual-statistical-tables-by-tax-rate-and-income-percentile).
## The dividend-driven tax floor
For the specific people Madoff named, the question isn't whether they *could* be nonpayers in theory. It's whether their documented income flows leave room to. Recent dividend initiations at their companies provide unavoidable taxable income. Verified figures:
| Person | Stake | Annual dividend income | Federal income tax on dividends (≈ 23.8%) |
|---|---|---|---|
| Larry Ellison | 1.16B Oracle shares (40%+) at [$2.00/share](https://www.marketbeat.com/stocks/NYSE/ORCL/dividend/) | **~$2.3B** | ~$552M |
| Steve Ballmer | 333M Microsoft shares at ~$3.24/share | **~$1.08B** ([247wallst](https://247wallst.com/investing/2025/06/11/steve-ballmer-makes-1-billion-a-year-in-microsoft-dividends/)) | ~$257M |
| Mark Zuckerberg | ~350M Meta shares at $2.10/share (initiated 2024) | **~$735M** ([Bloomberg](https://www.bloomberg.com/news/articles/2024-02-02/zuckerberg-to-get-700-million-a-year-from-meta-s-new-dividend)) | ~$175M |
| Sergey Brin | 730M Alphabet shares at $0.80/share (initiated 2024) | **~$584M** | ~$139M |
| Larry Page | 389M Alphabet shares at $0.80/share | **~$311M** ([CNBC](https://www.cnbc.com/2024/04/25/alphabet-issues-first-ever-dividend-70-billion-buyback.html)) | ~$74M |
Every one of these five is mechanically locked into a federal income tax bill in the **tens to hundreds of millions** every year, from dividends alone, before any stock sale or compensation. The top-1%-of-payers threshold ($186K) is cleared by a factor of 400–3,000x. None of them can land in the "$1 – $186K" middle bin from these flows.
For the three centibillionaires whose companies don't pay a dividend — Bezos (Amazon), Musk (Tesla), and pre-2024 Page/Brin (Alphabet) — the tax floor instead comes from discretionary sales. Bezos has sold $8–10B of Amazon annually since 2020, generating roughly $2B/year in federal tax. Musk exercised expiring Tesla options in 2021 and paid [$11B](https://www.cnn.com/2021/12/29/investing/elon-musk-tesla-stock-sales/index.html). In years with no such realization event, these three can in principle land at $0 — which ProPublica documented for Bezos in 2007 and 2011, and for Musk in 2018.
So the centibillionaire distribution in a given year is close to binary: either hundreds of millions in federal income tax (dividend-paying stake or active sales) or $0 (pure buy-borrow-die year). The $1–$186K middle is essentially empty at this wealth scale.
## Per-person data disclosed by ProPublica
Year-by-year federal income tax for the people [ProPublica](https://www.propublica.org/article/the-secret-irs-files-trove-of-never-before-seen-records-reveal-how-the-wealthiest-avoid-income-tax) named with dollar figures. The leak window is 2014–2018 unless noted.
| Person | $0 years in 2014–2018 | $0 years outside the window | Non-$0 tax disclosed |
|---|---|---|---|
| Warren Buffett | 0 | — | $23.7M total, 2014–2018 |
| Jeff Bezos | 0 | 2007, 2011 | $973M total, 2014–2018 |
| Michael Bloomberg | "several" (count not disclosed) | — | $292M total, 2014–2018; $70.7M in 2018 |
| Elon Musk | 2018 | — | $455M total, 2014–2018; $11B in 2021 |
| Mark Zuckerberg | 0 | — | hundreds of millions/yr via 10b5-1 |
| Larry Ellison | 0 | — | Oracle dividend every year |
| Carl Icahn | 2016, 2017 | — | $544M AGI across those years, $0 tax via interest expense |
| George Soros | 2016, 2017, 2018 | — | other years not disclosed |
In the 2014–2018 window: ≥6 zero years (Musk 1, Icahn 2, Soros 3, Bloomberg ≥1 unspecified) across 40 person-years, i.e. **≥15%**. Adding Bezos's out-of-window 2007 and 2011 zeros is 8+ documented zero years across the broader ProPublica coverage. Half of the eight named people had at least one documented zero year; half had none.
**Extrapolation to the top 100 wealthiest over the past decade is an estimate.** Starting from the ~15% in-window disclosed rate and adjusting for the Forbes 400 containing more hedge-fund and private-equity profiles (Icahn/Soros-style loss or deduction years) than the few names ProPublica covered in detail, a plausible range is **5–15%** per-year zero rate. This is a Claude estimate, not a published number — the underlying data isn't public.
Against Madoff's claim, the meaningful comparison is:
- Random US household in a given year: 40–44% pay $0 federal income tax ([TPC](https://taxpolicycenter.org/briefing-book/who-doesnt-pay-federal-income-taxes); 43.9% in PolicyEngine's 2026 file below), ~1% are in the top 1% of payers. Ratio nonpayer : top-1%-payer = **~40 : 1**.
- Top 100 wealthiest in a given year: 5–15% estimated zero rate; top-1%-of-payers rate not directly measured but most years' disclosed tax bills for Bezos, Musk, Bloomberg, and Zuckerberg put them well past the top-1% threshold. Directionally the ratio is inverted from the population.
## What the microdata says
[PolicyEngine's Enhanced CPS](https://policyengine.github.io/policyengine-us/) imputes household net worth from the Federal Reserve's Survey of Consumer Finances. The 99th percentile of federal income tax at the household level is $185,608 in the 2026 file (script: [gist](https://gist.github.com/MaxGhenis/3f10cc98b643e0a76e74eae9afe35082)):
| Wealthiest N households | Records | Net-worth range | Top 1% of income-tax payers | Pay $0 |
|---|---|---|---|---|
| Top 100 | 2 | $100.6M – $190.3M | 100.0% | 0.0% |
| Top 1,000 | 2 | $100.6M – $190.3M | 100.0% | 0.0% |
| Top 10,000 | 7 | $51.3M – $190.3M | 78.1% | 10.3% |
| Top 100,000 | 27 | $30.2M – $190.3M | 21.5% | 50.5% |
| Top 1,000,000 | 135 | $13.4M – $190.3M | 4.2% | 34.0% |
| All US households | 6,876 | — | 1.0% | 43.9% |
"Records" is the number of underlying microdata households contributing to each weighted group. The top 100 and top 1,000 both come from the same 2 records, so the 100.0% and 0.0% point estimates have no meaningful precision at that end — they're saying "the two synthetic records at the top of the file both clear the threshold and both have positive tax." The top-10,000 and top-100,000 rows are based on more records and are more informative about the $30M–$190M range.
Net worth in Enhanced CPS is imputed from the [Survey of Consumer Finances](https://www.federalreserve.gov/econres/scfindex.htm), which excludes Forbes 400-style respondents by design and truncates the upper tail at around $190M. The Forbes 400 uses direct estimates of billionaire wealth and isn't integrated into the SCF or into PE's microdata. So the microdata can't directly test Madoff's claim for Zuckerberg, Bezos, Musk, or Ellison — only for the visible $30M–$190M range.
Within that range, the top-100,000 row is notable: 50.5% pay $0 federal income tax, higher than the 43.9% population baseline. These are wealth-holders whose wealth is concentrated in housing, retirement accounts, or unrealized stock gains, with little realized annual income. This is the wealth profile that comes closest to matching Madoff's framing — but it's an order of magnitude less wealthy than Forbes 400 territory, and the share of nonpayers drops toward zero as you move further up the wealth ladder in the file.
Binned by net worth:

Sample size per bin is printed below each point. The $100M+ bin is based on only 3 records, so the 100% / 0% point there reflects the sparse top of the file, not a robust estimate.
## Other estimates of what the wealthiest pay
The rate you get for the top of the distribution depends on the denominator:
- **Tax / reported AGI.** For the Forbes 400, ProPublica's [top-400 interactive table](https://projects.propublica.org/americas-highest-incomes-and-taxes-revealed/) lists individual effective federal income tax rates averaged 2013–2018; most are in the 17–24% range.
- **Tax / (AGI + untaxed corporate profits).** [Saez, Yagan, Zucman et al. (NBER w34170, 2025)](https://www.nber.org/papers/w34170) put the top 0.0002% total federal rate at 24% for 2018–2020, vs. 30% population-wide. [Splinter's reanalysis](https://www.davidsplinter.com/BillionaireTaxRate.pdf) adjusts for multi-return Forbes families and different corporate-tax imputation and arrives at 38% — above the population average.
- **Tax / comprehensive income including unrealized capital gains.** The [OMB/CEA 2021 analysis](https://www.whitehouse.gov/cea/written-materials/2021/09/23/what-is-the-average-federal-individual-income-tax-rate-on-the-wealthiest-americans/) put the top 400 at 8.2%. ProPublica's 3.4% "true tax rate" divides tax paid by change in net worth over the period.
None of these studies — including ProPublica's follow-ups in the [Secret IRS Files series](https://www.propublica.org/series/the-secret-irs-files) — report the share of centibillionaire person-years with $0 federal income tax, because per-person year-by-year data isn't public. The question Madoff's framing depends on isn't directly measured in the published literature.
## Rate vs. frequency
Two of the three rate methodologies (ProPublica 3.4%, OMB/CEA 8.2%, Saez-Zucman 24%) put the richest below the 30% population average; Splinter's 38% recalculation is above. The rate story is more contested than a single headline number suggests.
The frequency framing — "just as likely to be in the 40% of nonpayers as the top 1% of payers" — is a different question. For the US population, the ratio (nonpayer year : top-1%-payer year) is ~44:1 in 2026.
I asked Codex and a 5-teammate Claude Code team to estimate the centibillionaire distribution, each working independently from the evidence summary. The 10 runs averaged **86% top 1% of payers, 2% middle, 12% nonpayer**.
The middle bin is small but not zero. ProPublica documented Elon Musk paying $68,000 in 2015 and $65,000 in 2017 — specific years where realization was small enough to land in the $1–$186K band rather than at $0 or in the top 1%. That's the narrow profile the middle captures: no dividends, no major sales, modest residual taxable income after interest-expense offsets.
The centibillionaire ratio, nonpayer year to top-1%-payer year, is **~1:7** — not Madoff's implied 1:1.
---
## I built a MyST-to-Quarto converter (and why you might need one)
**URL:** https://maxghenis.com/blog/mystquarto/
**Published:** Feb 27 2026
**Description:** mystquarto converts academic markdown between MyST and Quarto formats — directives, roles, config files, and frontmatter. Now on PyPI.
I have about 20 academic papers written in [MyST Markdown](https://mystmd.org/). MyST is great — Jupyter Book integration, Sphinx directives, executable code cells. But I've been increasingly drawn to [Quarto](https://quarto.org/) for its PDF output, built-in cross-referencing, and the fact that it's becoming the standard for reproducible academic publishing.
The problem: converting between the two isn't trivial. Both formats use markdown, but their syntax for the interesting parts — executable code, citations, figures, admonitions, cross-references — is completely different. And no converter existed.
So I built one.
## What mystquarto does
```bash
pip install mystquarto
myst2quarto docs/ -o docs-quarto/
quarto2myst docs/ -o docs-myst/
```
It handles the full conversion in both directions:
**Block directives.** MyST uses Sphinx-style directives (`` ```{code-cell} python `` with `:tags: [hide-input]`). Quarto uses pandoc-style fences (`` ```{python} `` with `#| code-fold: true`). mystquarto converts between all 13+ directive types: code cells, figures, math blocks, admonitions, tab sets, margin content, tables, images, and more.
**Inline roles.** MyST's `` {cite}`smith2024` `` becomes Quarto's `[@smith2024]`. Same for cross-references (`` {numref}`fig-results` `` → `@fig-results`), equations, inline code evaluation, and document links. Nine role types in total.
**Config files.** `myst.yml` and `_quarto.yml` have different structures for the same concepts — project type, table of contents, bibliography, export formats, author metadata. mystquarto maps between them, detecting book vs. article projects automatically.
**Frontmatter.** Per-file YAML like `kernelspec` → `jupyter`, `label` → `id`, and `exports` → `format`.
**File extensions.** `.md` ↔ `.qmd` renaming, with asset copying and `_build`/`.git` directory skipping.
## Why not use Pandoc?
Pandoc is the universal markdown converter, but it operates at the document level — parsing to an AST and rendering back. It doesn't understand MyST directives or Quarto's executable code cells. These are treated as raw content or code blocks, not as semantic constructs that need structural transformation.
The conversion between MyST and Quarto is fundamentally about syntax mapping between two directive systems, not about document parsing. A code cell's `:tags: [hide-input]` needs to become `#| code-fold: true` — that's a structural transform, not a format conversion.
## How it works
The architecture is deliberately simple: a regex-based line scanner with a directive stack, processing files line by line. The scanner detects opening fences (both backtick and colon styles), tracks nesting depth, parses option blocks, and dispatches to transform functions on close.
No heavy dependencies. The entire package uses just `click` for the CLI and `pyyaml` for config/frontmatter parsing. No `markdown-it-py`, `myst-parser`, or `docutils`. This keeps the install fast and the code easy to understand and extend.
225 tests cover every transform rule, with fixture-based comparisons and roundtrip tests (MyST → Quarto → MyST).
## Try it
```bash
# Install
pip install mystquarto
# Or run without installing
uvx myst2quarto docs/
# Preview what would change
myst2quarto docs/ --dry-run
```
The code is at [github.com/MaxGhenis/mystquarto](https://github.com/MaxGhenis/mystquarto), MIT-0 licensed (no attribution required). The [project page](/mystquarto) has the full syntax mapping table.
If you're sitting on MyST projects and want to try Quarto — or vice versa — give it a shot.
---
## One year of Claude Code
**URL:** https://maxghenis.com/blog/my-claude-code-config/
**Published:** Feb 25 2026
**Description:** A year ago, Anthropic launched Claude Code. I''ve since consumed 10 billion tokens, mass-tweeted about it, and mass-customized it. Here''s my setup and what I''ve learned.
Anthropic launched Claude Code on February 24, 2025. I [tweeted about it](https://x.com/MaxGhenis) three times on day one. A year later, it's my primary development environment, and I've mass-customized the `~/.claude` directory that powers it.
### A year in numbers
| Stat | Value |
|------|-------|
| Tokens consumed | 10.2 billion |
| Messages sent | 520,000+ |
| Sessions | 2,346 |
| Tweets about Claude Code | ~33 |
| Peak API spend (Aug 2025) | $5,861/month |
| Biggest single day (Feb 7, 2026) | $505 equivalent |
I started on the API, paying per-token. August 2025 hit $5,861. On July 20, 2025, I switched to the Max plan ($200/month, unlimited usage) and added a second subscription on a personal account in February 2026. The $200 plan replaced a $6K/month habit.
My IDE journey followed a similar arc of simplification. I went from VS Code to VS Code + [TerminalGrid](/blog/terminalgrid) (a custom extension for running multiple Claude Code sessions) to iTerm2 + tmux. Each migration stripped away a layer of complexity. The tmux setup I use now is three small config changes --- the TerminalGrid extension was hundreds of lines of TypeScript solving the same problem worse.
I recently audited my `~/.claude` directory and [made it public](https://github.com/MaxGhenis/.claude). Here's the full setup.
## How it started
I went to `git init` my `~/.claude` directory and discovered it was already a git repo. One of my installed plugins had initialized its own repo there, and my personal config files were sitting alongside it, untracked. The `.git` directory pointed to the plugin's remote. My files were just along for the ride.
So the first step was cleaning house. I removed the plugin's git history, initialized a fresh repo, and started deciding what to track.
## The audit
Before making anything public, I went through every file. The `~/.claude` directory accumulates a lot: conversation transcripts, clipboard images, debug logs, session state, plugin caches. Most of that is ephemeral or sensitive and belongs in `.gitignore`.
What's left after exclusions is surprisingly small: a `CLAUDE.md` global instructions file, a `settings.json`, five hook scripts, a handful of slash commands, a few local plugins, and a pair of shell scripts for secrets management.
While auditing, I also noticed that 36 GB of stale repo clones had accumulated in my home directory from multi-agent work. PolicyEngine US alone had 10 separate clones for different PRs, each a full copy of a large repo. That's not in `~/.claude`, but the audit prompted me to clean it up.
## Slash commands
Claude Code lets you define [custom slash commands](https://docs.anthropic.com/en/docs/claude-code/slash-commands) as markdown files in `~/.claude/commands/`. Each file is a prompt template you invoke with `/command-name`. I have twelve:
- **`/briefing`** --- Pulls today's calendar, unread emails, and (in the first week of the month) monthly task reminders. I run this most mornings.
- **`/search-everything`** --- Cross-platform search across local files, WhatsApp, Gmail, Granola meeting notes, and the browser. It works through sources in order of speed and stops when it finds what I need.
- **`/expense`** and **`/anthropic-expenses`** --- Automate filing reimbursements on [Open Collective](https://opencollective.com/policyengine). They search my email for invoices, download receipts, and submit expenses.
- **`/download-receipts`** --- Extracts PDF attachments from Gmail search results and saves them locally.
- **`/gmail`** and **`/google-api`** --- Reference commands for email patterns and Google API authentication across my work and personal accounts.
- **`/personal-info`** --- Loads my personal details from a private file for form-filling, applications, and profile creation.
- **`/slides`** --- Generates presentation decks from a brief or topic using a Next.js + Tailwind framework.
- **`/search-transcripts`** --- Searches past Claude Code conversation transcripts by keyword using [`claude-search`](https://github.com/nicobailon/claude-search).
- **`/config-tidy`** --- Audits and reorganizes `CLAUDE.md`, `MEMORY.md`, and skills to keep each layer within its target size.
- **`/bounce`** --- Sends a question or plan snippet to GPT for a second opinion. Claude uses this proactively during plan mode to gut-check architecture decisions, trade-offs, or whether it's missing something. Works alongside the `bounce-plan-gpt` hook, which handles the automatic final review.
Most of these compose multiple tools --- MCP servers, CLI utilities, APIs --- into a single action. The `/briefing` command, for example, calls the Google Calendar API, searches Gmail, and checks a task list, then synthesizes everything into a summary. Writing it as a slash command means I don't re-explain the workflow every session.
## Hooks
[Hooks](https://docs.anthropic.com/en/docs/claude-code/hooks) are shell scripts that run automatically before or after Claude Code events. I have five. The audit revealed a gap: three of the original scripts existed on disk, but only one was wired up in `settings.json`. The other two were doing nothing.
This is an easy mistake to make. You write the script, mark it executable, and forget that hooks also need a corresponding entry in `settings.json` to actually fire. I fixed it --- all four are now connected.
**`bounce-plan-gpt.sh`** runs before every `ExitPlanMode` call. When Claude finishes writing an implementation plan and tries to exit plan mode, this hook reads the plan from a standardized temp file, sends it to GPT for review, and either approves or blocks. If GPT has substantive feedback, the hook returns `{"decision": "block"}` with the feedback --- Claude sees the critique, incorporates it, and tries again. A bounce counter caps this at two rounds to prevent infinite loops. The hook pairs with a `/bounce` slash command that Claude can invoke mid-planning for ad-hoc second opinions on architecture choices or trade-offs. Together, they make cross-model review automatic rather than something Claude has to remember to do.
**`enforce-package-managers.sh`** runs before every Bash command. It blocks `npm`, `npx`, `yarn`, `pip`, and `pipx` and tells Claude to use `bun` or `uv` instead. Without this, Claude defaults to npm about half the time regardless of what `CLAUDE.md` says. The hook makes the preference absolute:
```bash
if echo "$command" | grep -qE '\bnpm\s+(install|i|add|...)\b'; then
echo '{"decision": "block", "reason": "Use bun instead of npm."}'
exit 0
fi
```
**`auto-commit-wip.sh`** runs before context compression (the `PreCompact` event). Context compression often precedes crashes or context exhaustion, so this hook auto-commits all uncommitted changes. I added it after losing an entire branch of multi-agent work --- 10 files, 262 KB --- because agents wrote to the working tree without committing, then the session crashed. The commit uses `--fixup=HEAD` so that fixup commits can be cleanly squashed later with `git rebase --autosquash`.
**`warn-uncommitted.sh`** runs on session stop. If there are uncommitted changes in the current repo, it prints a warning. Same motivation: don't lose work.
**`sync-setup-page.sh`** runs after every Bash command (`PostToolUse`). It watches for `git commit` in the `dotfiles` or `.claude` repos and reminds Claude to check whether the [setup page](/setup) on my site needs updating. This keeps the living reference in sync with config changes.
## CLAUDE.md
The `CLAUDE.md` at the repo root is my global instructions file --- it applies to every Claude Code session regardless of which directory I'm in. During this audit I trimmed it from about 113 lines to about 70.
What I removed: instructions about the current year (Opus 4.6 already knows the date), package manager preferences (the hook enforces this more reliably than a text instruction), and verbose TDD workflow steps (Claude already follows test-first patterns when asked).
What survived:
- **Fake data disclosure**: Claude must never present mock data without prominent warnings. This matters because I work on policy analysis where fake numbers could be mistaken for real projections.
- **Model routing**: Use Opus or Haiku for subagents, never Sonnet.
- **Sentence case for headings**: A style preference that applies to everything I write.
- **Google API and email patterns**: Pointers to the relevant slash commands and account details.
The general lesson: if something can be enforced by a hook, put it in a hook. `CLAUDE.md` is a suggestion. A hook that returns `{"decision": "block"}` is a hard stop.
## Skills: on-demand context loading
After the initial audit, I restructured how Claude Code accesses domain-specific reference information.
The problem: Claude Code has a persistent memory system --- a `MEMORY.md` file that's loaded into every session's system prompt. Mine had grown to 215 lines (past the 200-line truncation limit) with detailed API references, credential locations, deployment procedures, and troubleshooting guides for services like Whoop, Xero, App Store Connect, GCP billing, and Slack. Most of this was irrelevant to any given session but cost context every time.
The solution: [skills](https://docs.anthropic.com/en/docs/claude-code/plugins#skills). Skills are markdown files in a plugin's `skills/` directory that load on-demand based on trigger keywords in the conversation. When I mention "whoop" or "sleep data," the Whoop skill loads. When I mention "xero" or "UK expense," the Xero skill loads. Otherwise, they don't exist in the context window.
I moved 12 domain-specific reference sections from `MEMORY.md` into skills in my `max-productivity` local plugin:
| Skill | Triggers on | What it contains |
|-------|------------|-----------------|
| `whoop-health` | whoop, sleep, recovery, hrv | API endpoints, token refresh flow, cached data locations |
| `xero-uk-accounting` | xero, UK expense | OAuth flow, tenant ID, account codes |
| `app-store-connect` | app store, fastlane, xcode | API key, bundle IDs, build commands |
| `gcp-billing` | gcp, google cloud | Billing account IDs, project list |
| `opencollective-expenses` | expense, reimbursement | Collective details, currency handling |
| `cbo-baseline` | cbo, budget outlook | Excel file IDs, row numbers, YAML paths |
| `modal-vercel-deployment` | modal deploy, vercel deploy | Workspace config, failure modes |
| `openmessage-patterns` | text, sms, iMessage | MCP tools, HTTP API fallback, message ordering |
| `agent-teams` | agent team, TeamCreate | Workflow steps, gotchas |
| `slack-patterns` | slack, DM | Channel IDs, pagination rules, DM access |
| `search-email-patterns` | draft email, find meeting | Search priority, drafting preferences |
| `claude-in-chrome` | chrome extension, browser automation | Setup, account matching, troubleshooting |
`MEMORY.md` went from 215 lines to 61 --- just core operating principles, account identifiers, and hard-won lessons that apply to every session. Everything else loads only when relevant.
The skill file format is straightforward:
```markdown
---
name: Whoop health data
description: Use this skill when working with Whoop API, health data,
sleep analysis, recovery scores, or HRV data. Triggers: "whoop",
"sleep", "recovery", "hrv", "health data".
version: 0.1.0
---
# Whoop API & health data
## Credentials
- **Config**: `~/.config/whoop/credentials.json`
...
```
The `description` field is what Claude Code matches against. If the conversation context matches the trigger keywords, the full skill content loads into the prompt. The plugin's `plugin.json` just needs `"skills": "./skills"` to auto-discover all skill directories.
This is the same pattern the [PolicyEngine plugin](https://github.com/PolicyEngine/policyengine-claude) uses for its 60+ skills --- I just hadn't applied it to my personal reference material until now.
To verify it worked, I restarted Claude Code and asked a single question that touches three different skills:
> What's Pavel's Slack DM channel ID? And what's the Whoop sleep API endpoint? And how do I deploy to Modal?
Full response from Claude Code
Claude spawned three parallel Explore agents that searched the skill files on-demand, then synthesized the results:
**Pavel's Slack DM channel ID**
`D04xxxxx` (found via skill lookup)
**Whoop sleep API endpoint**
```
GET https://api.prod.whoop.com/developer/v2/activity/sleep?limit=25
```
- Auth: `Bearer {access_token}` header
- Max limit: 25 (50 fails)
- Pagination: `&nextToken=TOKEN`
- Credentials: `~/.config/whoop/credentials.json`
- Tokens expire frequently --- re-auth via OAuth if you get 403
**Modal deployment**
```bash
# Critical: unset keychain env vars first to use PE workspace
unset MODAL_TOKEN_ID MODAL_TOKEN_SECRET && modal deploy modal_app.py
```
- PE workspace: `policyengine` (verify with `modal profile current`)
- URL pattern: `https://policyengine--{app-name}-{function-name}.modal.run`
- If URL shows `maxghenis--` instead of `policyengine--`, you deployed to the wrong workspace --- unset env vars and redeploy
- For Vercel-fronted apps, update `VITE_API_URL` env var and force redeploy with `vercel --prod --force`
All three answers came from skills that loaded on-demand --- none of this was in the system prompt until the question triggered it.
### Keeping it clean: the config audit agent
The migration raised an obvious question: how do I keep things from drifting back? `MEMORY.md` grows automatically as Claude learns things during sessions. Without maintenance, it'll be back at 200+ lines within weeks.
Two mechanisms handle this. First, a `config-audit` [agent](https://docs.anthropic.com/en/docs/claude-code/plugins#agents) in the `max-productivity` plugin triggers proactively when `MEMORY.md` exceeds 150 lines or when I mention "clean up config." Second, a `/config-tidy` slash command I can run on demand to audit and reorganize.
Both know the same classification rules:
- **Behavioral rules** ("always do X") belong in `CLAUDE.md`
- **Compact facts** (account IDs, 1-2 line lessons) belong in `MEMORY.md`
- **Detailed reference** (API docs, step-by-step guides, anything >5 lines on one topic) belongs in a skill
When triggered, they read all three layers, report line counts, flag misplaced content, and propose moves. After I approve, they create new skill files, edit `MEMORY.md`, and update `CLAUDE.md` as needed. They never delete information --- they move it to the right layer.
This closes the loop: skills handle the *what*, the agent and command handle the *when* and *where*. Configuration maintenance becomes something Claude does for me rather than something I have to remember to do.
## Secrets management
The repo includes two shell scripts for secrets, both safe to publish:
- **`load-secrets.sh`** reads secrets from the macOS Keychain (service `claude-env`) and exports them as environment variables. It's sourced from `.zshrc` so every shell session has access.
- **`manage-secret.sh`** is a CRUD interface for the same Keychain service: `set`, `get`, `del`, `list`.
The scripts contain no secrets, just the Keychain lookup logic. I previously had an [age](https://github.com/FiloSottile/age)-encrypted secrets file alongside them, but it was redundant with the Keychain approach and added complexity. I removed it.
## Settings
`settings.json` configures MCP servers, permissions, hooks, enabled plugins, and plugin sources. The notable parts:
- **MCP servers**: Gmail, Google Ads, Chrome DevTools, and [OpenMessage](https://github.com/MaxGhenis/openmessage) (a unified SMS/iMessage/RCS gateway), each defined with their command and environment variables (using `${VAR}` references that resolve from the macOS Keychain at runtime --- no secrets in the file).
- **Permissions**: I run in bypass mode with broad tool access. This is a personal machine and I prefer speed over confirmation dialogs.
- **Plugin sources**: Pointers to local plugin directories for plugins I'm developing.
## tmux: how not to set up a terminal grid
The latest evolution in my year-long IDE journey: I moved from VS Code ([TerminalGrid](/blog/terminalgrid)) to iTerm2 + tmux for running multiple Claude Code sessions. The final setup took about 15 minutes. Getting there took most of a day.
The goal was simple: run 6+ Claude Code sessions in a visible grid, persistent across restarts. What followed was a comedy of errors --- each failure spawning a more complex workaround, each workaround failing in a more spectacular way, until the entire tmux server crashed and I was forced to start over with the obvious solution.
### Attempt 1: separate sessions + join-pane grid toggle
The first idea was one tmux session per project, then a `tmux-grid` script that would `join-pane` them all into a single window for a grid view. The problem: `join-pane` is destructive. It doesn't copy panes --- it *moves* them. Every time I toggled the grid, it permanently ripped windows out of their sessions. I rewrote the script at least eight times, each version breaking in a new way. Windows disappeared. Sessions ended up empty. The undo path was nonexistent because `join-pane` doesn't have one.
### Attempt 2: capture-pane read-only dashboard
OK, so don't move panes --- just *read* them. A `watch` command running `capture-pane` on each session, displaying the output in a split grid. This worked for about 30 seconds before I noticed the output was garbage. Claude Code uses [Ink](https://github.com/vadimdemedes/ink) (React for terminals) to render its TUI --- spinners, progress bars, dynamically updating panels. `capture-pane` grabs the raw terminal buffer, which is whatever escape sequences Ink happened to write last. The result was a grid of mangled ANSI artifacts. One pane out of six showed a readable header, purely by luck of timing.
This wasn't a bug to fix. It was a fundamental incompatibility: `capture-pane` can't render what Ink is drawing.
### Attempt 3: iTerm2 AppleScript automation
The Claude Code docs mention that `tmux -CC` in iTerm2 is the "suggested entrypoint." iTerm2 can host tmux sessions as native split panes, where each pane is a real terminal that Ink renders into properly. So I wrote an AppleScript to automatically split iTerm2 into a grid of tmux sessions.
This was the longest rabbit hole. At least eight rewrites. A catalog of errors:
- **App naming chaos.** AppleScript couldn't find `"iTerm2"`. Or `"iTerm"`. The working invocation turned out to be `application id "com.googlecode.iterm2"` --- the bundle identifier.
- **Type errors.** AppleScript error `-1700` (type coercion failure), `-609` (connection invalid), `-2741` (can't get reference). Each fix introduced the next error.
- **Infinite recursion.** When iTerm2 opened a new split pane, it loaded `.zshrc`, which contained the auto-launch script for the grid, which opened more split panes, which loaded `.zshrc`, which... the tmux server crashed under the load.
- **Pane reference failures.** AppleScript's object model for iTerm2 sessions is barely documented. `session 1 of current tab` worked for the first split but threw errors for subsequent ones.
After the tmux server crashed, I killed everything and sat with a blank terminal.
### The solution
I stopped iterating and asked a different question: what's the simplest thing that could work?
The answer was one tmux session, multiple panes, and a single built-in command: `select-layout tiled`. No AppleScript. No dashboard. No `join-pane`. No `capture-pane`. Just panes in a grid, each one a real terminal, each one rendering Claude Code's TUI perfectly.
The whole setup is three small changes.
**Auto-launch in `.zshrc`** --- every new terminal window attaches to the session or creates it:
```bash
if [[ -z "$TMUX" && -z "$VSCODE_INJECTION" ]]; then
tmux attach -t c 2>/dev/null || tmux new -s c 'claude --dangerously-skip-permissions; exec zsh'
fi
```
The `$VSCODE_INJECTION` guard skips auto-attach in VS Code terminals, which have their own Claude Code panel. The `exec zsh` at the end means that if Claude Code exits, the pane drops to a shell instead of closing. Close iTerm2, reopen it, and you're back where you left off. (The full block also handles SSH connections --- see [phone access](#phone-access) below.)
**[tmux-claude-code](https://github.com/MaxGhenis/tmux-claude-code)** is a TPM plugin I built to manage sessions. The `cc` script evolved from a five-line pane creator into a proper plugin with session search, resume, and browse features. It installs as a one-liner in `.tmux.conf`:
```bash
set -g @plugin 'MaxGhenis/tmux-claude-code'
set -g @claude_code_flags '--dangerously-skip-permissions'
```
The killer feature is keyword search across session transcripts. Every Claude Code session stores its conversation as a JSONL file in `~/.claude/projects/`. The plugin searches the first few user messages of every session, excludes currently active ones, resolves the correct working directory from the encoded project path, and opens `claude --resume` in the right place:
```bash
cc resume nextladder proposal # finds and resumes the matching session
cc resume fix ctc phaseout # keyword search across all transcripts
```
The plugin also handles three bugs that plagued my original script:
1. **Nested session detection**: Claude Code sets a `CLAUDECODE` environment variable. New panes inherit it, causing the child instance to refuse to start. The plugin strips it with `env -u CLAUDECODE` on every pane creation.
2. **Pane detection**: Claude Code sessions report `zsh` as their `pane_current_command` when idle at their prompt (because they're launched via `zsh -c "claude ...; exec zsh"`). Naive "is this pane free?" checks fail. The plugin uses a three-tier approach: check the command name, check for a child process named `claude` via `pgrep`, then fall back to checking pane content for CC prompt markers.
3. **Window overflow**: When a tmux window has too many panes, `split-window` silently fails ("no space for new pane"). The plugin falls back to `new-window` automatically.
**Navigation:**
| Action | Key |
|--------|-----|
| New Claude Code pane | `prefix+c` |
| New named CC pane | `prefix+C` |
| Resume CC session by keyword | `prefix+Ctrl-r` |
| Browse CC sessions (fzf) | `prefix+Ctrl-b` |
| Grid view (re-tile) | `prefix+g` |
| Zoom pane fullscreen | `prefix+z` |
| Return to grid | `prefix+z` again |
| Jump to pane by number | `prefix+q` then number |
| Next pane (stays zoomed) | `prefix+n` |
| Previous pane (stays zoomed) | `prefix+p` |
| Next pane (unzooms) | `prefix+o` |
| Move pane to its own window (hide) | `prefix+!` |
| Kill pane | `prefix+x` then `y` |
The zoom toggle is the key workflow: `prefix+z` to focus on one session fullscreen, `prefix+z` again to see the whole grid.
### Phone access
The simplest approach: [Remote Control](https://code.claude.com/docs/en/remote-control). With it enabled for all sessions, every local Claude Code session is automatically available at [claude.ai/code](https://claude.ai/code) and the Claude mobile app ([iOS](https://apps.apple.com/us/app/claude-by-anthropic/id6473753684)/[Android](https://play.google.com/store/apps/details?id=com.anthropic.claude)). No SSH or VPN needed --- just open the app and pick a session.
For full terminal access, SSH from a phone (via [Termux](https://termux.dev/) on Android or any SSH client on iOS) connects to the same tmux session. The `.zshrc` block detects `$SSH_CONNECTION` and creates a *grouped session* --- a linked session that shares windows but has its own independent view:
```bash
if [[ -z "$TMUX" && -z "$VSCODE_INJECTION" ]]; then
if [[ -n "$SSH_CONNECTION" ]]; then
# Skip during tmux-resurrect restore to prevent pane explosion
if [[ -z "$(tmux show-environment -g TMUX_RESTORING 2>/dev/null | grep -v '^-')" ]]; then
tmux new-session -A -t c -s "remote-$$" \; new-window -n remote 'claude; exec zsh'
fi
else
tmux attach -t c 2>/dev/null || tmux new -s c 'claude; exec zsh'
fi
fi
```
The `TMUX_RESTORING` guard prevents a feedback loop: when tmux-resurrect restores panes, each spawns a shell that sources `.zshrc`, which would create more grouped sessions, which would spawn more panes. With the guard, restored panes just start as plain shells. Auto-restore is also disabled (`@continuum-restore 'off'`); this guard is defense-in-depth. The `new_session.sh` script in the tmux-claude-code plugin also caps panes at 20 per window.
With `aggressive-resize on` in `tmux.conf`, the phone and laptop can have different window sizes without either one getting squished. [Tailscale](https://tailscale.com/) handles networking so I can reach the laptop from anywhere without port forwarding.
### Cross-pane awareness
The most useful emergent behavior of running multiple Claude Code sessions in tmux: [one session can read another's terminal output](https://x.com/MaxGhenis/status/2026254649309749640). One pane was rebasing a policyengine-core branch against master after a towncrier migration PR merged, resolving merge conflicts in push.yaml. I told a different pane to look at what it was doing. It ran `tmux capture-pane` on the sibling, read through the rebase output, and noticed that the "Build changelog" step in push.yaml had lost its `run:` command during the merge --- an empty CI step that would silently break versioning on the next PR merge. It then checked every other repo we'd just merged towncrier into, found the same bug in policyengine-canada, and fixed both directly on master. One agent caught a bug introduced by another agent's work, by reading its terminal.
This is a tmux-specific capability. VS Code and iTerm2 can split terminals visually, but there's no programmatic API for one terminal to read the contents of a sibling split. In tmux, every pane's buffer is accessible from any other pane.
I've started organizing sessions into named windows --- `work` and `personal` --- and temporarily pulling subsets of panes into a `focus` window when I need to zoom in on two related tasks. Moving panes between windows is `join-pane -s %ID -t window-name`, and `select-layout tiled` re-tiles after every move.
### Agent teams stay contained
Claude Code has an experimental [agent teams](https://code.claude.com/docs/en/agent-teams) feature where one session coordinates multiple teammates. By default, if it detects tmux, it creates new panes for each teammate --- which clutters your manually-arranged grid. Setting `teammateMode` to `in-process` in `settings.json` keeps teammates inside their lead's pane:
```json
{
"teammateMode": "in-process"
}
```
Teammates still work in parallel, but they're contained within a single pane. Use `shift+down` to cycle through them.
### Why it took so long
The pattern of failure is worth more than the final config. Each broken approach spawned a more complex workaround instead of a simpler alternative. At no point --- through eight rewrites of `tmux-grid`, the dashboard dead end, and the iTerm2 AppleScript ordeal --- did either of us (me or Claude) ask "what's the simplest thing that works?" We only asked that question after the tmux server crashed and there was nothing left to iterate on.
The lesson is the same one that applies to most software projects: when the third workaround fails, the architecture is wrong. Step back. The solution to "my grid script breaks every time I run it" wasn't a better grid script. It was `select-layout tiled`.
All the config files are in [github.com/MaxGhenis/dotfiles](https://github.com/MaxGhenis/dotfiles). The tmux plugin is at [github.com/MaxGhenis/tmux-claude-code](https://github.com/MaxGhenis/tmux-claude-code). A living reference version with the latest config is at [/setup](/setup).
## What I learned
A few things came out of this audit that I wasn't expecting:
**Hooks beat instructions.** The enforce-package-managers hook catches every `npm install` before it runs. The equivalent `CLAUDE.md` instruction ("use bun, not npm") works maybe half the time. If you care about a behavior strongly enough to write it down, write it as a hook instead.
**Hooks on disk aren't hooks in practice.** Two of my three safety hooks existed as executable scripts but weren't registered in `settings.json`. They'd never fired. The gap between "I wrote this" and "this is running" is easy to miss, especially since there's no warning that a hook script exists but isn't configured.
**Config directories accumulate without you noticing.** The 36 GB of stale clones, the plugin that had taken over my git history, the age-encrypted secrets file I'd stopped using months ago --- none of this was visible until I sat down and went through everything deliberately. Periodic audits are worth the time.
## Why make it public
A few reasons:
1. **Reference for others setting up Claude Code.** The documentation covers each feature individually, but seeing a real configuration that ties them together is more useful than reading about each piece in isolation.
2. **Backup and portability.** If I set up a new machine, I can clone this repo and have my workflow back immediately (minus secrets, which live in the Keychain).
3. **Accountability.** Knowing it's public makes me more deliberate about what goes in there.
The repo is at [github.com/MaxGhenis/.claude](https://github.com/MaxGhenis/.claude).
---
## Why you can''t make double eye contact
**URL:** https://maxghenis.com/blog/double-eye-contact/
**Published:** Feb 10 2026
**Description:** Your eyes can only converge on a single point. Up close, that means you have to pick one of the other person''s eyes to look at.
import EyeConvergence from '../../../components/EyeConvergence.tsx';
Have you ever noticed that when you're face-to-face with someone, you can't actually look at both of their eyes at the same time? You're always picking one — left or right — and subtly switching between them.
This isn't a limitation of attention or focus. It's geometry.
## One convergence point
Your two eyes always rotate as a pair to aim at the same target. This is called [**vergence**](https://pmc.ncbi.nlm.nih.gov/articles/PMC5122972/), and it means they converge on a single point in space at any given moment.
When someone is close to you, their two eyes occupy noticeably different positions in your visual field. The angular separation between them is large enough that you can't fit both into the single point your eyes converge on. So you pick one.
## Try it yourself
Toggle between looking at each eye, and slide the distance to see how angular separation changes.
When the face is close, the angular separation between their eyes is large — well beyond your [fovea](https://pmc.ncbi.nlm.nih.gov/articles/PMC12330045/) (the high-resolution center of your vision, [~1.5mm across](https://pmc.ncbi.nlm.nih.gov/articles/PMC12330045/) or roughly 5° of visual angle). You're clearly choosing one eye.
As the face moves farther away, both eyes fall into a tighter angular range. Eventually they're so close together in your visual field that it _feels_ like you're looking at both — even though you're still converging on a single point.
## Why it doesn't matter at a distance
Human interpupillary distance [averages about 63mm](https://pmc.ncbi.nlm.nih.gov/articles/PMC3520592/). At 1 meter away, that subtends roughly 3.6°. At 3 meters, it's about 1.2° — easily within your fovea's ~5° span. This is why you barely notice the effect in normal conversation but absolutely feel it during an intimate close-up.
## What about Magic Eye?
If you've viewed a [Magic Eye stereogram](https://en.wikipedia.org/wiki/Magic_Eye), you know it's possible to diverge your eyes — point them more parallel than the object distance would normally call for — while still focusing on a near surface. Could you use this trick to aim each eye at a different target?
No. Even in Magic Eye viewing, both eyes are still aimed at a single convergence point — it's just a virtual one behind the page rather than on the surface. You've moved _where_ your eyes converge, but you haven't split that point in two. To make double eye contact, you'd need each eye aimed at a completely independent target, and that's not something the [vergence system](https://www.ncbi.nlm.nih.gov/books/NBK11070/) supports — it always drives both eyes toward the same point in space.
I never would have built an interactive visualization to explore a random curiosity like this before. But with [Claude Code](https://claude.ai/claude-code), describing what I wanted and getting a working widget took a couple hours — so why not? The idle question gets a better answer when you can play with it yourself.
---
## Interactive replication of GiveWell''s cost-effectiveness analysis
**URL:** https://maxghenis.com/blog/givewell-cea/
**Published:** Feb 10 2026
**Description:** I re-implemented GiveWell''s cost-effectiveness models for all six top charities as an open-source web tool with editable parameters, moral weights, and sensitivity analysis.
I re-implemented GiveWell's cost-effectiveness models for all six top charities as an open-source web tool: **[maxghenis.com/givewell-cea](https://maxghenis.com/givewell-cea)**
The tool lets you edit any parameter and immediately see the effect on charity rankings.
## Motivation
I discovered GiveDirectly in 2012 when Google.org made a grant to help them expand from Kenya to Uganda. I loved the RCT emphasis — a charity that treated evidence the way I'd want any intervention evaluated. GiveDirectly led me to GiveWell, and GiveWell's cost-effectiveness analysis has guided my giving for almost a decade since my first donation in 2017.
When I first dug into GiveWell's CEA spreadsheets, I was amazed by the depth. Every parameter sourced, every assumption explicit, every step auditable. A public good built on radical transparency. That concept — computational epistemic infrastructure, open models that let anyone verify and challenge the reasoning — eventually inspired me to build [PolicyEngine](https://policyengine.org), though I wouldn't have foreseen that when I first opened the spreadsheet. I view GiveWell's CEA as one of the original products of the effective altruism community, which has increasingly shaped my worldview with its emphasis on a broad moral circle and methodical assessment of [where to invest resources to maximize impact](/blog/why-ive-taken-the-giving-what-we-can-pledge/).
Others have done excellent work examining specific parts of GiveWell's CEA — [Froolow's critical review](https://forum.effectivealtruism.org/posts/6dtwkwBrHBGtc3xes/a-critical-review-of-givewell-s-2022-cost-effectiveness) of model architecture, [Nolan, Rokebrand, and Rao's uncertainty quantification](https://forum.effectivealtruism.org/posts/Nb2HnrqG4nkjCqmRg/quantifying-uncertainty-in-givewell-cost-effectiveness), and several pieces on [deworming](https://forum.effectivealtruism.org/posts/MKiqGvijAXfcBHCYJ/deworming-and-decay-replicating-givewell-s-cost) and [AMF](https://forum.effectivealtruism.org/posts/4Qdjkf8PatGBsBExK/adding-quantified-uncertainty-to-givewell-s-cost) uncertainty. But I couldn't find a tool that implements all six charities together and makes it easy to compare them while adjusting assumptions.
GiveWell's spreadsheets are powerful but hard to explore casually. Each charity has its own multi-tab workbook with dozens of sheets, specialized terminology, and cross-references between cells:

Changing a moral weight means editing cells across multiple sheets and comparing results manually. I wanted something where you could adjust one slider and immediately see how all six charities re-rank.
## What the model covers
For each charity I implemented the core pipeline from GiveWell's spreadsheets:
1. **People reached**: Grant size / cost per person reached
2. **Deaths averted** (or equivalent): People reached × mortality/disease rate × intervention effect size
3. **Units of value**: Deaths averted × moral weight (age-adjusted)
4. **Cost-effectiveness**: Units of value per dollar / benchmark value per dollar
5. **Adjustments**: Charity-level (quality, track record), intervention-level (external validity), leverage and funging
The six charities each have their own structure:
- **AMF** and **Malaria Consortium**: Under-5 mortality reduction, with separate pathways for older age mortality and developmental effects
- **Helen Keller International**: VAS effect on under-5 mortality
- **New Incentives**: Cash incentives increase vaccination rates; the model converts incremental vaccinations to deaths averted using vaccine-specific effect sizes
- **GiveDirectly**: The model values consumption increases directly (no mortality pathway)
- **Deworm the World**: The model values long-run earnings effects of deworming via ln(consumption)
All 51 charity/country combinations operate independently — each country has its own cost per person, mortality rate, adjustment factors, etc.
## Observations from the replication
**The benchmarks aren't on the same scale.** GiveWell expresses each charity's cost-effectiveness as a multiple of GiveDirectly cash transfers. The denominator — how many "units of value" (GiveWell's composite of moral-weight-adjusted lives saved) one dollar of cash generates — differs across spreadsheets: AMF uses 0.00333, MC and HKI use 0.00335, and GiveDirectly uses 0.003. GiveWell built different spreadsheets at different times and baked in different moral weight calibrations. They don't mechanically compare multiples across spreadsheets, but a tool that displays them side by side inherits this inconsistency.
**Mortality rate definitions vary.** AMF's spreadsheet has both a raw malaria mortality rate and a derived "mortality rate in the absence of nets" rate. The latter accounts for existing net coverage and is the correct input. My first extraction accidentally used the raw rates, which underestimated AMF's cost-effectiveness by roughly 2x for some countries (e.g., DRC: 0.00306 raw vs. 0.00798 in-absence-of-nets).
**Counterfactual coverage drives most of the within-charity variation.** Both Helen Keller and New Incentives have a `proportionReachedCounterfactual` parameter — what fraction of people would receive the intervention anyway, without the charity's involvement. The remaining fraction is the charity's incremental impact:
Helen Keller (VAS):
| Country | Would receive VAS anyway | x benchmark |
|---------|-------------------------|-------------|
| Niger | 15% | 79× |
| DRC | 20% | 30× |
| Mali | 21% | 17× |
| Madagascar | 33% | 12× |
| Guinea | 18% | 11× |
| Cameroon | 23% | 8× |
| Burkina Faso | 29% | 7× |
| Côte d'Ivoire | 40% | 6× |
New Incentives (vaccinations, Nigerian states):
| State | Would get vaccinated anyway | x benchmark |
|-------|---------------------------|-------------|
| Sokoto | 68.5% | 39× |
| Zamfara | 79.4% | 31× |
| Kebbi | 71.9% | 29× |
| Bauchi | 81.5% | 20× |
| Jigawa | 85.8% | 18× |
| Katsina | 81.0% | 17× |
| Kano | 83.2% | 13× |
| Gombe | 88.0% | 10× |
| Kaduna | 84.2% | 9× |
Counterfactual coverage correlates with cost-effectiveness but doesn't determine it alone. Niger's 85% incremental coverage is similar to Guinea's 82%, yet Niger scores 7x higher: higher mortality rate (1.4% vs 1.1%), double the VAS effect (11.1% vs 5.5%), and 70% lower cost per child.
**Helen Keller's leverage/funging adjustments need careful reading.** In my first pass, the funging adjustment for Burkina Faso extracted as 531.99 instead of -0.431. These values come from separate rows in the spreadsheet that are easy to confuse — a "percentage change" row vs. an "adjusted value" row.
## Verification
I extracted parameters from GiveWell's November 2025 CEA spreadsheets ([AMF](https://docs.google.com/spreadsheets/d/1VEtie59TgRvZSEVjfG7qcKBKcQyJn8zO91Lau9YNqXc), [MC](https://docs.google.com/spreadsheets/d/1De3ZnT2Co5ts6Ccm9guWl8Ew31grzrZZwGfPtp-_t50), [HKI](https://docs.google.com/spreadsheets/d/1L6D1mf8AMKoUHrN0gBGiJjtstic4RvqLZCxXZ99kdnA), [NI](https://docs.google.com/spreadsheets/d/1mTKQuZRyVMie-K_KUppeCq7eBbXX15Of3jV7uo3z-PM)). I verified 46 of the 51 charity/country final cost-effectiveness multiples against the spreadsheets (GiveDirectly excluded — see Limitations):
| Charity | Countries | Max difference |
|---------|-----------|---------------|
| Against Malaria Foundation | 8 | <0.001% |
| Malaria Consortium | 8 | <0.001% |
| Helen Keller International | 8 | <0.001% |
| New Incentives | 9 | 0.000% (exact) |
| Deworm the World | 13 | 0.000% (exact) |
298 automated tests pass. The remaining <0.001% differences for AMF/MC/HKI are floating-point precision, not model discrepancies.
## Interactive features

**Calculation breakdown**: Click any country to see the step-by-step calculation with every intermediate value. Click any highlighted number to edit it.
**Moral weights**: GiveWell's default weights peak at ages 5-9 (134) and weight under-5 at 116. You can adjust these with a single multiplier or set each age bracket independently. Charities with different age profiles (AMF and MC focus on under-5; NI and HKI have broader age effects) shift rankings when you change these.
**Sensitivity analysis**: Sweep any moral weight parameter across its range and see how all six charities' cost-effectiveness changes. This makes crossover points visible — for instance, the under-5 weight where NI overtakes MC, or the discount rate at which deworming drops below cash transfers.
## How assumptions affect rankings
A few examples of what you see when you adjust parameters:
**Default rankings (best country per charity, GiveWell Nov 2025 defaults):**
| Rank | Charity | Best country | x benchmark |
|------|---------|-------------|-------------|
| 1 | Helen Keller | Niger | 79× |
| 2 | New Incentives | Sokoto | 39× |
| 3 | Deworm the World | Kenya | 35× |
| 4 | AMF | Guinea | 23× |
| 5 | Malaria Consortium | Chad | 15× |
| 6 | GiveDirectly* | Mozambique | 4× |
*GiveDirectly uses a simplified model with older parameters — see Limitations.
**Double the under-5 moral weight.** All four mortality-focused charities see exactly +100% gains — their value comes entirely from deaths averted, so doubling the weight doubles the result. GiveDirectly gains only +3.4% because most of its value comes from consumption benefits. (A [2025 study](https://www.givedirectly.org/mortality2025/) found $1,000 transfers cut infant mortality by 48%, and [GiveWell now includes mortality effects](https://www.givedirectly.org/givewell-2024/) in GiveDirectly's CEA — but with steep discounts, so consumption still dominates.) Rankings don't change.
**Equal moral weights across ages (all set to 100).** Child-focused charities drop 14-16% (the default under-5 weight of ~116 falls to 100). Helen Keller drops from 79× to 67×. Rankings stay the same.
**Double AMF's cost per child in DRC.** The relationship is perfectly linear: doubling cost halves the x benchmark from 14.6× to 7.3×. This is a more powerful lever for changing *relative* rankings within mortality-focused charities than moral weight changes, which scale all of them equally.
When most top charities prevent child deaths, changing the weight on child deaths scales them all in the same direction. What *does* break the rankings is operational cost differences between countries. The tool lets you find the specific crossover points where, say, doubling a cost parameter in one country moves it below another charity entirely.
## Programmatic access
The models are pure TypeScript functions — no UI dependency. Clone the repo and run analyses directly:
```bash
git clone https://github.com/MaxGhenis/givewell-cea && cd givewell-cea
bun install
```
Sweep a parameter:
```typescript
// save as sweep.ts, run with: bunx tsx sweep.ts
import { calculateHelenKeller } from "./src/lib/models/helen-keller";
import { HK_COUNTRY_PARAMS } from "./src/lib/models/countries";
const niger = HK_COUNTRY_PARAMS.niger;
for (let effect = 0.05; effect <= 0.2; effect += 0.05) {
const r = calculateHelenKeller({
grantSize: 1_000_000,
...niger,
vasEffect: effect,
});
console.log(`VAS effect ${(effect * 100).toFixed(0)}% → ${r.finalXBenchmark.toFixed(1)}×`);
}
// VAS effect 5% → 35.6×
// VAS effect 10% → 71.3×
// VAS effect 15% → 106.9×
// VAS effect 20% → 142.5×
```
Rank all charity/country combinations:
```typescript
// save as rank.ts, run with: bunx tsx rank.ts
import { calculateHelenKeller } from "./src/lib/models/helen-keller";
import { calculateAMF } from "./src/lib/models/amf";
import { calculateNewIncentives } from "./src/lib/models/new-incentives";
import {
HK_COUNTRY_PARAMS, HK_COUNTRY_NAMES,
AMF_COUNTRY_PARAMS, AMF_COUNTRY_NAMES,
NI_COUNTRY_PARAMS, NI_COUNTRY_NAMES,
} from "./src/lib/models/countries";
const G = 1_000_000;
const all = [
...Object.entries(AMF_COUNTRY_PARAMS).map(([k, v]) => ({
charity: "AMF", country: AMF_COUNTRY_NAMES[k],
xb: calculateAMF({ grantSize: G, ...v }).finalXBenchmark,
})),
...Object.entries(HK_COUNTRY_PARAMS).map(([k, v]) => ({
charity: "HKI", country: HK_COUNTRY_NAMES[k],
xb: calculateHelenKeller({ grantSize: G, ...v }).finalXBenchmark,
})),
...Object.entries(NI_COUNTRY_PARAMS).map(([k, v]) => ({
charity: "NI", country: NI_COUNTRY_NAMES[k],
xb: calculateNewIncentives({ grantSize: G, ...v }).finalXBenchmark,
})),
];
all.sort((a, b) => b.xb - a.xb);
for (const r of all.slice(0, 10)) {
console.log(`${r.charity.padEnd(4)} ${r.country.padEnd(16)} ${r.xb.toFixed(1)}×`);
}
// HKI Niger 79.1×
// NI Sokoto 38.6×
// NI Zamfara 31.3×
// HKI DRC 29.9×
// NI Kebbi 29.0×
// AMF Guinea 22.8×
// ...
```
Each model function (`calculateAMF`, `calculateHelenKeller`, `calculateNewIncentives`, etc.) takes a flat parameter object and returns all intermediate values, so you can inspect any step of the pipeline.
## Limitations
I replicated the *structure* of GiveWell's models but not their full analytical process:
- I implement the calculation pipeline but not the reasoning behind parameter choices. GiveWell's adjustments (charity quality, external validity, leverage, funging) encode substantial judgment that this tool takes as given. The parameters themselves — particularly the adjustment factors — represent years of investigation, site visits, literature reviews, and internal debate. This tool lets you see the arithmetic, but the arithmetic was never the hard part.
- No uncertainty analysis. The tool supports sensitivity analysis — sweeping one parameter at a time to see how results change — but does not place joint distributions over parameters or propagate uncertainty through the model. [Several](https://forum.effectivealtruism.org/posts/4Qdjkf8PatGBsBExK/adding-quantified-uncertainty-to-givewell-s-cost) [excellent](https://forum.effectivealtruism.org/posts/ycLhq4Bmep8ssr4wR/quantifying-uncertainty-in-givewell-s-givedirectly-cost) [posts](https://forum.effectivealtruism.org/posts/Nb2HnrqG4nkjCqmRg/quantifying-uncertainty-in-givewell-cost-effectiveness) have explored what happens when you put distributions around these estimates.
- I built the GiveDirectly model as a simplified approximation from an older spreadsheet and blog post, not a direct replication of GiveWell's current model. It uses identical parameters (spillover effects, mortality effects, consumption persistence) across all five countries, unlike GiveWell's country-differentiated approach.
- The tool currently covers GiveWell's top 6 charities but not newer additions.
For donation decisions, use [GiveWell's published estimates](https://www.givewell.org/how-we-work/our-criteria/cost-effectiveness/cost-effectiveness-models).
Source code: [github.com/MaxGhenis/givewell-cea](https://github.com/MaxGhenis/givewell-cea)
---
## OpenMessage: How I built a macOS Google Messages client to give Claude my texts
**URL:** https://maxghenis.com/blog/openmessage/
**Published:** Feb 09 2026
**Description:** I had already connected Claude Code to WhatsApp, Signal, Slack, and Gmail. SMS was the last holdout — and no existing tool could solve it.
I've been on a mission to give Claude Code access to all my communication channels. WhatsApp, Signal, Slack, Gmail — each one has an MCP server that lets Claude read and send messages on my behalf. But one channel was missing: SMS and RCS on my Android phone.
**[Download OpenMessage for Mac](https://github.com/MaxGhenis/openmessage/releases/latest/download/OpenMessage.dmg)** — free, open source, requires macOS 14.0+ and an Android phone with Google Messages.
## The gap
If you have an iPhone, iMessage on Mac gives you desktop texting for free. But I use Android, and Google Messages only offers a web client — no desktop app, no API, no way for an AI assistant to interact with it.
I needed Claude to be able to do things like "text Alex that the meeting moved to 3pm" or "check if James confirmed for tomorrow" without me picking up my phone. Every other messaging channel was already connected. SMS was the last holdout.
## Attempt 1: SMS Gateway
My first approach was [SMS Gateway](https://smsgateway.me), an Android app that exposes your phone's SMS capability via an API. I built an MCP server around it, and it worked — sort of.
The problems:
- **Clunky setup** — required installing a separate Android app, configuring API keys, keeping the app running in the background
- **SMS only** — no RCS support, so group chats and rich messages were invisible
- **Sending was unreliable** — messages would sometimes queue but never actually send
- **No conversation history** — you could only send messages, not read existing conversations
It was a dead end for anything beyond basic "fire and forget" SMS.
## Attempt 2: Google Messages protocol
I started digging into how Google Messages for Web actually works. It turns out Google uses an internal protocol (based on gRPC/protobuf) to pair a browser with your phone via QR code. A few open-source projects had reverse-engineered this protocol, most notably [mautrix/gmessages](https://github.com/mautrix/gmessages), a Go library originally built as a Matrix bridge.
This was the breakthrough. The mautrix library could:
- Pair with a phone via QR code (same as messages.google.com)
- Receive all conversations and messages in real time
- Send SMS and RCS messages
- Handle group chats, reactions, and read receipts
I quickly built an MCP server around it, and it worked beautifully. Claude could read my full message history and send texts.
## The mistake that became a feature
Here's where I'll admit something: I thought Google Messages only allowed one paired web device at a time. It turns out Google Messages has two pairing methods — Google account pairing and QR code pairing — and they don't mix. I had been using Google account pairing, which blocks additional devices. When I switched to QR pairing (which is what the mautrix library uses), multiple devices work fine. But I didn't realize this at the time, so I was convinced that running my MCP server meant giving up the web client. I built a whole native macOS app to replace it.
I didn't need to build the app at all.
But I'm glad I did. What started as a workaround became something better: an open-source, native Google Messages client for Mac — the first one that exists. No browser tab to keep open, no Electron wrapper, just a real app with a built-in MCP server for AI access.
## Building the app
Since I already had the Go backend handling the Google Messages protocol, I wrapped it in a native macOS app using Swift. The architecture is simple:
- **Go backend** — handles pairing, message sync, and the MCP server. Stores everything in a local SQLite database
- **Swift wrapper** — native macOS app that launches the Go backend and displays a web UI via WKWebView
- **Web UI** — clean conversation list and message view served from localhost
The app pairs with your phone the same way messages.google.com does — scan a QR code — and then syncs your full conversation history. The MCP server runs alongside the UI, so Claude can access messages while you browse them yourself.
Everything is stored locally — SQLite database, no cloud sync, no accounts to create. When you use AI tools via the MCP server, only the messages you ask about are sent to your chosen AI provider (e.g. Anthropic's API for Claude). OpenMessage itself never sends your data anywhere.
## What Claude can do with it
With OpenMessage running, Claude Code can:
- **Search messages** — "find the address James sent me last week"
- **Read conversations** — "summarize my conversation with the team"
- **Send messages** — "text Alex that the meeting moved to 3pm"
- **React to messages** — "thumbs up the last message from Sarah"
Combined with WhatsApp, Signal, Slack, and Gmail MCP servers, Claude now has access to essentially all my communications. I can say "check all my messages for anything urgent" and it searches across every channel.
## What's next
OpenMessage is open source, and I think it's a canvas for something bigger. The same open-source libraries (mautrix) that power OpenMessage also support WhatsApp, Signal, Telegram, Discord, Slack, and more. Beeper built a $125M company on these same libraries — but without local-first architecture or AI integration.
I think there's an opportunity for an open-source, AI-native unified messaging client. Some ideas:
- **More services** — WhatsApp, Signal, Telegram via the mautrix ecosystem
- **AI-powered features** — smart replies, conversation summaries, priority inbox
- **Natural language search** — "when did someone send me a confirmation number?"
- **Cross-platform** — Linux and Windows versions
If any of these interest you, [open an issue](https://github.com/MaxGhenis/openmessage/issues) or submit a PR.
## Get it
OpenMessage is free and open source. Requires macOS 14.0+ and an Android phone with Google Messages.
**[Download for Mac](https://github.com/MaxGhenis/openmessage/releases/latest/download/OpenMessage.dmg)** | [openmessage.ai](https://openmessage.ai) | [GitHub](https://github.com/MaxGhenis/openmessage)
---
## Where US immigrants come from, and where ICE focuses enforcement
**URL:** https://maxghenis.com/blog/bad-bunny-immigration-enforcement/
**Published:** Feb 09 2026
**Description:** Bad Bunny''s Super Bowl halftime show highlighted the Americas. Here''s what the data shows about immigration and enforcement patterns.
Last night, Bad Bunny performed the [first primarily Spanish-language Super Bowl halftime show](https://www.rollingstone.com/music/music-news/bad-bunny-super-bowl-performance-1235513007/) in history. At the end, holding a football that read "Together, we are America," he said "God bless America" and then [named every country and territory in the American continent](https://sports.yahoo.com/nfl/breaking-news/article/bad-bunny-echoes-grammy-message-with-super-bowl-halftime-show-together-we-are-america-015121748.html): Chile, Argentina, Uruguay, Paraguay, Bolivia, Peru, Ecuador, Brazil, Colombia, Venezuela, Panama, Costa Rica, Nicaragua, Honduras, El Salvador, Guatemala, Mexico, Cuba, the Dominican Republic, Jamaica, Haiti, the United States, Canada, and the US territory of Puerto Rico. Performers carried each flag behind him.
The performance reached an estimated [130 million viewers](https://variety.com/2026/music/news/bad-bunny-super-bowl-lady-gaga-real-wedding-ceremony-1236656381/). Here's what the data shows about immigration from the Americas and current enforcement patterns.
## Most US immigrants come from the Americas
**53%** of the 50.2 million foreign-born people in the United States come from other countries and territories in the American continent — Mexico alone accounts for 22.2% — according to the [2024 American Community Survey](https://data.census.gov/table/ACSDT1Y2024.B05006). Every country Bad Bunny named on that stage falls within that share.
## Spanish is the most common non-English language
Bad Bunny performed entirely in Spanish and said during the show: "English wasn't my first language, but that's OK — it wasn't America's either." Spanish is the most common non-English language in the United States, [spoken at home by about 45 million people](https://data.census.gov/table/ACSST1Y2024.S1601), about 14% of the US population. Of the 74 million people who speak a non-English language at home, **61% speak Spanish**, per [2024 ACS data](https://data.census.gov/table/ACSST1Y2024.S1601).
Congress has never passed a law designating an official language. In March 2025, [Executive Order 14224](https://www.federalregister.gov/documents/2025/03/06/2025-03694/designating-english-as-the-official-language-of-the-united-states) designated English as the official language and revoked [EO 13166](https://www.govinfo.gov/content/pkg/FR-2000-08-16/pdf/00-20938.pdf), a 2000 Clinton-era order that directed agencies to improve access for people with limited English proficiency. However, the order also states that agencies are "not required to amend, remove, or otherwise stop production of documents...in languages other than English."
At the time of the [first US census in 1790](https://www.census.gov/library/publications/1909/decennial/century-populaton-growth.html), about 86% of the white population was of British Isles descent (English, Welsh, Scottish, or Irish), with German (9%), Dutch (3%), and French (2%) communities making up much of the rest. The census also counted roughly 700,000 enslaved people — most American-born by that point — and excluded most Native Americans, who spoke [over 300 languages](https://www.bia.gov/faqs/do-all-american-indians-and-alaska-natives-speak-single-traditional-language) at the time of European contact. Among the counted population, English was the dominant language, though German-speaking communities in Pennsylvania were large enough that [bilingual government documents](https://www.census.gov/library/publications/1909/decennial/century-populaton-growth.html) were common. Spanish-speaking settlers founded [St. Augustine in 1565](https://www.nps.gov/places/st-augustine-town-plan-historic-district-st-augustine-florida.htm), 42 years before the first permanent English settlement at Jamestown.
## ICE enforcement is concentrated in the Americas
While 53% of the foreign-born population is from the Americas, ICE enforcement skews more heavily toward those countries.
[Official ICE data](https://www.ice.gov/statistics) shows that **98-99% of ICE administrative arrests** from FY2021 through Q1 FY2025 involved people from other countries in the Americas. The [UC Berkeley Deportation Data Project](https://deportationdata.org), which publishes record-level FOIA'd arrest data, provides a more granular look: from January 20 to October 15, 2025, **92.6%** of 220,931 arrests involved nationals of countries in the Americas, with Mexico (38.6%), Guatemala (14.1%), and Honduras (11.0%) as the top three countries.
A [UCLA Luskin study](https://knowledge.luskin.ucla.edu/wp-content/uploads/2025/10/Unseen_Latino-Ice-Arrests-Surge-Under-Trump_20251027.pdf) found that about 90% of ICE arrests during the first six months of Trump's second term involved nationals of Latin American countries, using the same Berkeley FOIA data.
This concentration likely reflects the composition of the unauthorized population, which [skews more heavily Latin American](https://www.migrationpolicy.org/article/frequently-requested-statistics-immigrants-and-immigration-united-states) than the foreign-born population as a whole, as well as geographic proximity and enforcement priorities.
## Operation Metro Surge in Minnesota
In December 2025, DHS launched [Operation Metro Surge](https://www.britannica.com/event/2025-26-Minnesota-ICE-Deployment) in Minneapolis, which it called the largest immigration enforcement operation ever, [deploying up to 2,000 federal agents](https://www.cbsnews.com/minnesota/live-updates/ice-somali-immigrants-minneapolis-st-paul/) to the Twin Cities.
DHS said the operation focused on fraud in the Somali-American community. Minnesota is home to the largest Somali population in the US — roughly [84,000 people in the Twin Cities alone](https://www.pbs.org/newshour/nation/5-things-to-know-about-the-somali-community-in-minnesota-after-trumps-attacks). Of those, nearly **58% were born in the United States**, and of the foreign-born Somalis in Minnesota, **87% are naturalized US citizens**, according to [Census data reported by PBS](https://www.pbs.org/newshour/nation/5-things-to-know-about-the-somali-community-in-minnesota-after-trumps-attacks).
Key outcomes of the operation:
- **3,000+ people arrested** as of January 19, 2026, per [DHS](https://www.dhs.gov/news/2026/01/19/ice-continues-remove-worst-worst-minneapolis-streets-dhs-law-enforcement-marks-3000)
- DHS publicly identified a subset of arrestees on its website. Of those, **23 were from Somalia** — while the operation's stated focus was on Somali fraud ([Britannica](https://www.britannica.com/event/2025-26-Minnesota-ICE-Deployment)). A [Washington Times analysis](https://www.washingtontimes.com/news/2026/jan/29/operation-metro-surge-crosses-3500-arrest-mark-heres-whos-getting/) of 170 DHS-released arrests found Mexico and Laos were the top two nationalities, with Somalia third. The full nationality breakdown of all arrests has not been made public
- About **5% of arrestees had violent criminal records** ([Britannica](https://www.britannica.com/event/2025-26-Minnesota-ICE-Deployment))
- **Two US citizens were killed** by federal agents during the operation: Renee Good and Alex Pretti ([CBS News](https://www.cbsnews.com/news/minneapolis-trump-immigration-ice-border-patrol-arrests-protests-shootings/))
- **50,000 people** gathered on January 23 for a statewide general strike in protest ([Britannica](https://www.britannica.com/event/2025-26-Minnesota-ICE-Deployment))
[Sahan Journal reported](https://sahanjournal.com/immigration/immigration-enforcement-somali-community-impact/) effects on daily life in affected communities, including residents skipping medical appointments, reduced mosque attendance, and hundreds of businesses closing. Minnesota and the Twin Cities [filed a federal lawsuit against DHS](https://www.ag.state.mn.us/Office/Communications/2026/docs/00190_DHS_Complaint.pdf).
## Summary
53% of the 50.2 million foreign-born people in the United States come from other countries in the Americas. ICE enforcement is even more concentrated, with over 92% of arrests in 2025 involving nationals of countries in the Western Hemisphere. The largest enforcement operation, in Minnesota, was focused on the Somali community but resulted in over 3,000 arrests across nationalities.
Bad Bunny's halftime show put a spotlight on the hemispheric connections between the US and the rest of the Americas. The immigration and enforcement data provide context for the scale of those connections.
---
## I used Claude Code to recertify for SNAP
**URL:** https://maxghenis.com/blog/snap-recertification-claude-code/
**Published:** Feb 06 2026
**Description:** Browser automation + document reading + form filling: Claude Code completed my SNAP recertification in about an hour of active work, across 373 tool calls.
I'm CEO of [PolicyEngine](https://policyengine.org), a tax and benefit policy nonprofit, but I don't take a regular salary — so I qualify for SNAP. I renewed my benefits tonight using [Claude Code](https://docs.anthropic.com/en/docs/claude-code) and its browser automation capabilities. The entire recertification — navigating DC's District Direct portal, reading my financial documents, calculating income, filling out dozens of form fields, and submitting — took about an hour of active work. Claude made 373 tool calls. I sent 26 messages.
Here's what I did, and what Claude did.
## The conversation
It started simply:
> **Me:** help me recertify for snap
>
> **Claude:** I can help you with SNAP recertification. A few questions to get started: What state are you in? Have you received a recertification notice? Do you know if your state has an online portal?
>
> **Me:** dc, idk, yes it does
Claude found the portal, opened it in Chrome, and told me to log in (it can't enter passwords). I logged in and said:
> **Me:** done
>
> **Claude:** You're logged in and there's a banner right at the top: "Your SNAP case is due for recertification, click the link to recertify before your case is terminated."
Early on, Claude was asking for confirmation before every step. I put a stop to that:
> **Claude:** I can see the Start button at the bottom. Should I click it to begin the recertification form?
>
> **Me:** yes stop asking just go!
After that, Claude moved through the pre-filled sections autonomously — name, DOB, citizenship, household composition — only stopping when it actually needed something from me. Like when it hit the SSN page:
> **Claude:** This page is asking for your SSN. It's partially masked (showing the last four digits) so it's already on file. I cannot enter sensitive information like SSNs — this is something you'll need to verify/enter yourself.
>
> **Me:** it's already there, just hit next
When Claude reached the income section, it needed updated numbers from my financial accounts. It tried to access Wealthfront and Vanguard in the browser but couldn't:
> **Claude:** Financial sites are blocked for safety reasons — I can't access Wealthfront or Vanguard since they contain sensitive financial data. You'll need to grab those numbers yourself.
>
> **Me:** can you ge tthem?
>
> **Me:** i opened https://dashboard.web.vanguard.com/ can you view it?
>
> **Claude:** I don't see the Vanguard tab in my tab group. Unfortunately financial sites like Vanguard and Wealthfront are blocked by the browser extension's safety restrictions.
>
> **Me:** just downloaded a pdf for vanbguad
(You'll notice the typos — I've gotten sloppier with typing since Claude handles it fine, so I don't bother correcting myself.)
So I downloaded a Vanguard PDF, a Wealthfront CSV, and my paystubs to my Downloads folder. Claude read them directly from the filesystem, extracted the relevant figures, and calculated my quarterly investment income.
The rest was uneventful. Claude filled in employer details, navigated through expenses (all pre-filled from last time), checked attestation boxes, and reached the final submission page:
> **Claude:** Before I click Next to submit the recertification, I want to confirm with you. This is the final certification page where you're signing and certifying that all information is correct. Ready to submit?
>
> **Me:** i dont this it does but yes hit next
Submitted at 8:56 PM. The confirmation page told me my SNAP interview is due by February 10.
## What I provided
My role was minimal:
- **My state**: DC
- **My address hasn't changed**
- **Downloaded documents**: I downloaded statements from my financial institutions showing dividend and interest income, and my recent paystubs, to my Downloads folder
- **Logged in**: I handled the login to District Direct myself
- **A few corrections**: like telling Claude the SSN was already filled in, and that I'm already registered to vote
That's 26 messages total. I never looked at the form itself — not once. Everything else was Claude.
## What Claude did
Claude read my financial documents (PDFs and CSVs), extracted the relevant figures, navigated the multi-page recertification form, and filled in every field. Here's a rough breakdown of its 373 tool calls:
- **72 screenshots** to see the current state of the page
- **88 clicks** to navigate buttons, dropdowns, checkboxes, and links
- **28 find operations** to locate elements on the page
- **22 form inputs** to fill text fields, select options, and enter data
- **6 page navigations**
- Plus scrolling, reading documents, web searches, and calculations
The form covers household composition, address, income from all sources (employment, investments, other), expenses (shelter, utilities, childcare, medical), and rights and responsibilities. Claude navigated all of it, carrying forward pre-filled data where nothing had changed and updating the fields that needed new numbers.
## How long it took
The whole process spanned about 3 hours and 40 minutes wall-clock, but most of that was idle time — me downloading documents, reading what Claude was doing, or stepping away. The active working time was roughly 1 hour and 20 minutes.
For comparison, doing this manually takes me at least an hour of focused attention: logging in, finding the right forms, looking up all my financial information, entering it carefully, reviewing everything, and submitting. Claude's version required about 5 minutes of my attention spread across the session.
## Why this matters
SNAP recertification is exactly the kind of task that AI should help with. It's:
- **High-stakes but routine**: Getting it wrong could mean losing benefits, but the process itself is just data entry
- **Document-heavy**: It requires pulling numbers from paystubs, bank statements, and investment accounts
- **Tedious**: The form is long, repetitive, and easy to make mistakes on
- **Time-sensitive**: Miss the deadline and you lose benefits
People who receive SNAP benefits are, by definition, low-income. Their time is valuable. The bureaucratic burden of maintaining benefits is a real cost, and it causes people to lose benefits they're entitled to. Tools that reduce that burden matter.
## What you'd need to try this
This used Claude Code with the [Claude in Chrome](https://chromewebstore.google.com/detail/claude-in-chrome/blelmpkgncmhfbjlccpkbjbdgfdjpgol) extension for browser automation. You'd need:
- A Claude Pro or Max subscription (for Claude Code access)
- The Chrome extension installed
- Your financial documents downloaded and accessible
- Comfort with Claude having access to your browser and local files
This is still early-stage technology. I wouldn't recommend it for someone who isn't comfortable reviewing what Claude is doing as it works. But the trajectory is clear: the boring, stressful parts of interfacing with government systems are exactly what AI agents should handle.
---
## Scrollywood: Smooth scroll video recording for the web
**URL:** https://maxghenis.com/blog/scrollywood/
**Published:** Feb 06 2026
**Description:** A Chrome extension that records smooth-scrolling videos and GIFs of any webpage, built for capturing scrollytelling stories and long-form content.
At [PolicyEngine](https://policyengine.org), we're building more scrollytelling stories to explain how tax and benefit policy works. Pages like [MITA](https://maxghenis.com/mita) use scroll-driven animations to walk through data visualizations step by step, and we want to share these experiences beyond the browser — in presentations, social posts, and demos.
The problem: there's no good way to record a smooth scroll of a webpage. Screen recording tools require you to scroll manually, which is never perfectly smooth. Browser automation tools can screenshot sequences but don't capture scroll-triggered animations. And scrollytelling pages depend on continuous scrolling to trigger IntersectionObserver callbacks that animate the content.
So I built [Scrollywood](/scrollywood) — a Chrome extension that records a perfectly smooth scroll capture of any webpage.

## How it works
Click the extension icon, set your scroll duration, choose WebM, MP4, or GIF, and start recording. Scrollywood scrolls the page from top to bottom at a constant rate while recording the tab at 30-60fps. The result is a downloadable capture in your selected format, with MP4 shown when Chrome supports it.
The scroll is smooth and linear at 60fps, which is critical for scrollytelling pages. Libraries like [scrollama](https://github.com/russellsamora/scrollama) and [react-scrollama](https://github.com/jsonkao/react-scrollama) use IntersectionObserver to trigger animations as elements enter the viewport. A smooth programmatic scroll naturally crosses these thresholds, so the animations play exactly as they would during manual scrolling.
## Technical challenges
Building this was straightforward in concept but tricky in practice, mostly due to Chrome's Manifest V3 architecture:
**Service worker timeouts.** MV3 service workers sleep after ~30 seconds of inactivity, which makes `setTimeout` unreliable for recordings longer than 30 seconds. The fix: the offscreen document (which has a persistent DOM context) handles all timing, while the service worker only handles one-shot operations like script injection and downloads.
**Iframe-wrapped pages.** Some sites embed content in full-page iframes (e.g., a custom domain wrapping a GitHub Pages site). The outer frame has no scrollable content. Scrollywood detects these wrappers and defers to the inner frame, which handles its own scrolling via `allFrames: true` injection.
**CSS scroll-behavior conflicts.** Pages with `scroll-behavior: smooth` in CSS fight with programmatic scrolling. Scrollywood temporarily overrides this to `auto`, but carefully avoids overriding `overflow`, which would break `position: sticky` — the CSS property that makes scrollytelling graphics stay in place.
**Large recording downloads.** At 16Mbps for a 60-second recording, the video can be 100MB+. The original approach (base64-encoding the blob and sending it via `chrome.runtime.sendMessage`) silently fails at that size. The fix: download directly from the offscreen document using a blob URL.
## Get it
[Download Scrollywood](/scrollywood) and load it as an unpacked extension in Chrome. It's free and open source.
[Project page](/scrollywood) · [GitHub](https://github.com/MaxGhenis/scrollywood)
---
## The adolescence of policy simulation: On Amodei and economic disruption
**URL:** https://maxghenis.com/blog/amodei-adolescence-policyengine/
**Published:** Jan 27 2026
**Description:** AI can answer policy questions. It cannot yet model policy impacts. That requires infrastructure we are still building.
Anthropic CEO Dario Amodei published a [20,000-word essay](https://www.darioamodei.com/essay/the-adolescence-of-technology) today called "The Adolescence of Technology." It covers AI safety, bioweapons, authoritarianism, and more. But one section stood out to me: his predictions about economic disruption.
Amodei writes that AI "will have effects that are much broader and occur much faster" than previous technological transitions. He predicts this will trigger an "unusually painful" short-term shock in the labor market, potentially displacing half of entry-level white collar jobs within the next five years.
If he's right, we'll need policy responses. And to design policy responses, we need to model them. That's where things get interesting.
## What AI can and cannot do
AI has gotten remarkably good at answering policy questions. Ask a frontier model about EITC phase-out rates or CTC eligibility rules and you'll probably get a correct answer. The models have read the tax code.
But there's a difference between answering policy questions and modeling policy impacts. Try this:
> "If the US expanded the Child Tax Credit to $3,600 for all children regardless of income, what would happen to the child poverty rate and the Gini coefficient?"
I asked myself this question. Here's my honest answer:
*I don't know. I could tell you the current child poverty rate is around 12% (SPM) and guess it would drop significantly—maybe to 8-9%?—based on vague memories of studies from the 2021 expanded CTC. For Gini, I'd guess a small decrease, maybe 0.005-0.01 points? But these are guesses anchored to half-remembered papers, not calculations. I have no microdata. I can't run a counterfactual. I'm pattern-matching to priors.*
This isn't a prompting problem. It's an infrastructure problem.
## Policy simulation requires three things
To actually model the impact of a policy change, you need:
1. **Policy** — The rules: tax formulas, benefit eligibility, phase-outs, interactions between programs
2. **Data** — Representative microdata: who earns what, household composition, geographic distribution
3. **Theory** — Behavioral assumptions: how do people respond to incentives, what's the takeup rate
AI might eventually nail the policy part. With enough statute text in context—or better, with deterministic APIs encoding the rules—models can learn to apply tax law correctly.
But AI can't conjure CPS microdata. It can't know the joint distribution of income, household size, and state of residence for 130 million US households. It can't decide whether to assume zero behavioral response or apply elasticities from the labor economics literature.
That's not intelligence. That's infrastructure.
## Two sides of the same coin
The same infrastructure serves another purpose. In "Machines of Loving Grace," Amodei imagines "a very thoughtful and informed AI whose job is to give you everything you're legally entitled to by the government." That's not hypothetical—tools like MyFriendBen and Amplifi already use PolicyEngine's API to screen people for benefits. Policy simulation and benefit access are two sides of the same coin. One computes law against population data, the other against your data.
## The adolescence parallel
Amodei calls this moment the "adolescence of technology"—powerful but not yet mature, capable of great things and great harm.
Policy simulation is in its own adolescence. We can model reforms in hours instead of months. [Governments are starting to use these tools](https://policyengine.org/uk/research/policyengine-10-downing-street). But the infrastructure is still clunky, adoption is still sparse, and most policy debates still happen without quantitative grounding.
The question is whether policy simulation grows up fast enough to match the disruption it needs to respond to.
## What this implies
If Amodei's timeline is right—powerful AI within two years, significant labor displacement within five—we don't have time for the traditional policy analysis cycle. Bills get introduced, CBO scores them months later, revisions happen, repeat. That works when policy moves at legislative pace.
But if the economy is changing faster than legislatures can respond, we need infrastructure that lets policymakers iterate quickly. Not just ideas about what to do, but tools for modeling trade-offs in real time.
This is part of why I've spent the last few years building [PolicyEngine](https://policyengine.org). It's one attempt at closing the gap—open-source microsimulation that anyone can use. But the broader point isn't about any one tool. It's that policy response capacity needs to exist before crises, not during them.
AI can accelerate parts of this. [We've been experimenting with AI workflows](https://policyengine.org/us/research/multi-agent-workflows-policy-research) that compress hours of analysis into minutes. But AI can only accelerate infrastructure that exists. It can't replace infrastructure that doesn't.
---
*Read Amodei's full essay: [The Adolescence of Technology](https://www.darioamodei.com/essay/the-adolescence-of-technology)*
---
## opencollective-py: Manage OpenCollective from your terminal or Claude Code
**URL:** https://maxghenis.com/blog/opencollective-py/
**Published:** Jan 23 2026
**Description:** A Python client, CLI, and MCP server for the OpenCollective API.
I manage [PolicyEngine's OpenCollective](https://opencollective.com/policyengine). We use OpenCollective because it gives anyone full visibility into our finances—every expense, donation, and transaction is public. It also handles crowdfunding, recurring donations, and community updates, and the [platform itself is open source](https://github.com/opencollective/opencollective).
But submitting expenses, reviewing pending ones, approving reimbursements—it all meant clicking through a web UI.
So I built [opencollective-py](https://github.com/MaxGhenis/opencollective-py)—a Python client that handles OpenCollective operations programmatically. It includes a CLI for quick terminal commands and an MCP server so Claude Code can manage expenses directly.
## What it does
**Submit expenses**—reimbursements (with receipts) or invoices (for services):
```bash
oc reimbursement "Conference registration" 500.00 receipt.pdf -c policyengine
oc invoice "Consulting - January" 2000.00 -c policyengine
```
**List and filter expenses**:
```bash
oc expenses -c policyengine --pending
oc expenses -c policyengine --status approved
```
**Approve or reject** (for collective admins):
```bash
oc approve exp-abc123
oc reject exp-def456
```
**Get account info**:
```bash
oc me
```
## Python client
```python
from opencollective import OpenCollectiveClient
client = OpenCollectiveClient(access_token="...")
# Submit expenses
client.submit_reimbursement("policyengine", "Travel", 50000, "receipt.pdf")
client.submit_invoice("policyengine", "Consulting", 200000)
# Manage expenses
expenses = client.list_expenses("policyengine", status="PENDING")
client.approve_expense("exp-abc123")
client.reject_expense("exp-def456")
```
## MCP server for Claude Code
Add to your Claude Code config:
```json
{
"mcpServers": {
"opencollective": {
"command": "python",
"args": ["-m", "opencollective.mcp_server"]
}
}
}
```
Then manage expenses in natural language: "Show me pending expenses for PolicyEngine" or "Submit my AWS receipt as a $150 reimbursement."
## HTML receipt conversion
Many email receipts are HTML. OpenCollective only accepts images and PDFs. The package automatically converts HTML to PDF:
```bash
oc reimbursement "AWS bill" 150.00 aws-receipt.html -c policyengine
```
## Get it
```bash
pip install opencollective
# With PDF conversion
pip install "opencollective[pdf]"
```
Then authenticate:
```bash
oc auth
```
[Project page](/opencollective-py) ・ [PyPI](https://pypi.org/project/opencollective/) ・ [GitHub](https://github.com/MaxGhenis/opencollective-py)
*This is an unofficial community project, not affiliated with OpenCollective.*
---
## From IDE to AI orchestration: The end of code-first development
**URL:** https://maxghenis.com/blog/ide-to-ai-orchestration/
**Published:** Jan 11 2026
**Description:** How AI coding tools evolved from autocomplete to autonomous agents, and why I expect to abandon my IDE entirely this month.
import AICodingTimeline from '../../components/AICodingTimeline.astro';
import CrosspostAware from '../../components/CrosspostAware.astro';
Claude Code creator Boris Cherny [recently shared](https://x.com/bcherny/status/2004897269674639461) that since the beginning of December, "100% of my contributions to Claude Code were written by Claude Code." He elaborated: "In the last thirty days, I landed 259 PRs—497 commits, 40k lines added, 38k lines removed. Every single line was written by Claude Code + Opus 4.5."
I've been building entirely in Claude Code since October—longer than Boris because he's a much better developer than I am, so Claude didn't have to be as good. My VS Code is just a shell for Claude Code terminals. It's why I created the [TerminalGrid extension](https://maxghenis.com/terminalgrid)—I haven't used the file editor in months.
The next step in this evolution is to skip the IDE altogether. And 2025's tool launches suggest that's exactly where the industry is heading.
## The timeline
**Feb 24, 2025** — [Anthropic launches Claude Code](https://www.anthropic.com/news/claude-3-7-sonnet)
A simple terminal tool—chat with Claude, edit files, run bash commands.
**May 16, 2025** — [OpenAI launches Codex](https://openai.com/index/introducing-codex/)
A cloud-based software engineering agent that runs tasks in isolated containers.
**Nov 18, 2025** — [Google launches Antigravity](https://developers.googleblog.com/build-with-google-antigravity-our-new-agentic-development-platform/)
An agent-first development platform where you don't touch code—you manage agents.
**Nov 24, 2025** — [Anthropic launches Opus 4.5 + Claude Code in Desktop](https://www.anthropic.com/news/claude-opus-4-5)
Multiple parallel sessions with isolated git worktrees.
**Jan 6, 2026** — [Anthropic adds local Claude Code to Desktop](https://x.com/_catwu/status/2008628736409956395)
Claude Code can run any shell command, access your filesystem, and control your browser.
The Claude Desktop approach differs from Codex and similar cloud tools in one crucial way: full local access. Claude Code can run any shell command, access your filesystem, and control your browser. Cloud-based agents like Codex run in sandboxed containers with restricted permissions. Claude's bet is that developers want an agent with the same access they have—not a constrained PR-generator.
## The paradigm shift
Google's Antigravity documentation articulates what's happening: the shift from "Editor view" (traditional IDE interface with an agent sidebar) to "Manager view" (orchestrating autonomous agents).
In Manager view, you don't write code—you describe what you want, the agents work in parallel, and you review artifacts (task lists, implementation plans, screenshots, browser recordings). It's project management, not programming.
This matches how I've been working—I typically run 5-10 Claude Code sessions simultaneously through [TerminalGrid](/terminalgrid), spread across a 2x2 grid (sometimes with multiple tabs per cell, sometimes across two VS Code windows):

Boris Cherny [described a similar setup](https://venturebeat.com/technology/the-creator-of-claude-code-just-revealed-his-workflow-and-developers-are): "I run 5 Claudes in parallel in my terminal. I number my tabs 1-5, and use system notifications to know when a Claude needs input." He also runs 5-10 instances on claude.ai, using a "teleport" command to hand off sessions between web and terminal.
## Why I'm not quite there yet
This week I've been experimenting with Claude Code in Desktop, and while it's close to replacing VS Code, I've hit two blocking issues:
1. **No MCP support yet**: Claude Code in the Desktop app doesn't support MCP servers, so I can't use browser automation tools like Claude in Chrome. The same MCPs work fine in VS Code's terminal and in non-Code Claude Desktop. I [posted about this](https://x.com/MaxGhenis/status/2009808936212549910)—a Claude Code engineer [replied](https://x.com/amorriscode/status/2010410355030421700) that MCP support should come next week.
2. **Sessions drop when switching accounts**: I run out of quota on my Claude Max 20x account regularly (those 5 parallel Claude instances add up), so I maintain two accounts. Claude Desktop drops sessions when I switch between them, losing context.
## The bigger picture
[As of late 2025](https://blog.jetbrains.com/research/2025/10/state-of-developer-ecosystem-2025/), roughly 85% of developers regularly use AI tools for coding. [Stack Overflow's 2025 Developer Survey](https://survey.stackoverflow.co/2025/) found 65% use them weekly. The question isn't whether AI will change how we code—it's whether we'll still call what we do "coding" at all.
Boris Cherny's workflow hints at the answer. He doesn't code. He orchestrates agents. He described using Opus 4.5 for everything because "even though it's bigger & slower than Sonnet, since you have to steer it less and it's better at tool use, it is almost always faster than using a smaller model in the end."
That's the trade-off: more capable models need less human steering, which means less time in the editor, which means the editor matters less.
## What happens next
I expect Anthropic to fix both blocking issues this month. When they do, I'll switch from VS Code to Claude Desktop, and TerminalGrid will join the graveyard of tools I've built that AI capabilities have rendered obsolete.
[CodeStitch](https://codestitch.dev), which I built in mid-2024 to paste codebases into Claude's context window, was obsolete within months of Claude Code maturing. TerminalGrid solved a real problem for exactly the window between "Claude Code exists" and "Claude Code has a native manager interface." That window is closing.
The pattern for developer tooling: ship fast, solve the problem in front of you, expect obsolescence. Frontier labs are moving so quickly on general-purpose dev tools that anything built to improve them has a short shelf life.
But that's fine. Disposable tools can accelerate progress toward bigger challenges. The real opportunity isn't building better coding tools—it's applying AI to problems that matter. Health, poverty, governance, the environment. If you can ship in days what used to take months, you can tackle problems that were previously out of reach.
When Claude Desktop works reliably, I'll abandon my TerminalGridded IDE without looking back—and get back to the real work.
---
## TerminalGrid: Turn VS Code into a Claude Code superterminal
**URL:** https://maxghenis.com/blog/terminalgrid/
**Published:** Jan 07 2026
**Description:** A VS Code extension for keyboard-driven terminal grid management with project picker and auto-launch for AI coding tools.
I run multiple Claude Code sessions simultaneously—one for each project I'm working on. VS Code's terminal panel only splits in one direction—no 2D grids. And the native terminal has [issues with image pasting](https://github.com/anthropics/claude-code/issues/1361) that matter when you're sharing screenshots with Claude.
So I built [TerminalGrid](https://maxghenis.com/terminalgrid)—my first VS Code extension, built 100% with Claude Code.
## VS Code as a shell for coding agents
Boris Cherny, who created Claude Code, [recently shared](https://x.com/bcherny/status/2004897269674639461) that 100% of his contributions to Claude Code in the past month were written by Claude Code itself. I've been operating this way for two to three months now—not writing any code myself. Of course, Boris is a much better programmer than I am, so I had less to discard.
This shift has changed how I use VS Code. I've stopped looking at the code directly; VS Code is now just a shell for coding agents. I don't use the file explorer or most other features. It's still better than a raw terminal since it lets you paste screenshots, but otherwise I just keep it as a grid of Claude Code sessions.
## The problem
VS Code's [integrated terminal](https://code.visualstudio.com/docs/terminal/basics) only splits in one direction at a time—side-by-side when the panel is at the bottom, or stacked when it's on the side. No 2D grids. This has been requested since 2018 ([#56112](https://github.com/microsoft/vscode/issues/56112), [#160501](https://github.com/microsoft/vscode/issues/160501)) and people are still asking in 2025 ([#254638](https://github.com/microsoft/vscode/issues/254638), [#252458](https://github.com/microsoft/vscode/issues/252458)).
Other extensions like [Split Terminal](https://marketplace.visualstudio.com/items?itemName=BrianNicholls.split-terminal) and [Workspace Layout](https://marketplace.visualstudio.com/items?itemName=lostintangent.workspace-layout) don't solve this—they work within the terminal panel's limitations.
When you're running 4+ AI coding sessions, you need a proper grid:
```
┌─────────────────┬─────────────────┐
│ policyengine │ api-server │
│ (Claude Code) │ (Claude Code) │
├─────────────────┼─────────────────┤
│ docs │ frontend │
│ (Claude Code) │ (Claude Code) │
└─────────────────┴─────────────────┘
```
## The solution
TerminalGrid moves terminals to the editor area, where VS Code already supports full grid layouts. Then it adds keyboard shortcuts and a project picker:
1. **`Cmd+K Cmd+N`** — Opens a searchable list of your projects
2. **Pick a project** — Type to filter, select with Enter
3. **Terminal launches** — Runs `cd && claude` automatically
The terminal is named after the folder, so you always know which Claude is working on what.
## Setup
```json
{
"terminalgrid.projectDirectories": ["~/projects", "~/code"],
"terminalgrid.autoLaunchCommand": "claude"
}
```
That's it. Now every `Cmd+K Cmd+Down/Right/N` gives you a project picker that launches Claude in the right directory.
## Other features
- **Crash recovery** — Terminal directories persist even if VS Code crashes
- **Image pasting** — Editor-area terminals handle screenshots better than the terminal panel
- **Works with any CLI tool** — Aider, Codex, Gemini CLI, or just plain shells
## Get it
Install from the [VS Code Marketplace](https://marketplace.visualstudio.com/items?itemName=MaxGhenis.terminalgrid) or:
```bash
code --install-extension MaxGhenis.terminalgrid
```
[Project page](/terminalgrid) ・ [VS Code Marketplace](https://marketplace.visualstudio.com/items?itemName=MaxGhenis.terminalgrid) ・ [GitHub](https://github.com/MaxGhenis/terminalgrid)
---
*Day 1 of [12 Days of Shipping](https://maxghenis.com). Merry Christmas!*
---
## RAMBar: A macOS menu bar RAM monitor for developers
**URL:** https://maxghenis.com/blog/rambar/
**Published:** Jan 07 2026
**Description:** A native macOS menu bar app for tracking memory usage across Claude Code sessions, VS Code workspaces, Chrome tabs, and Python processes.
Claude Opus 4.5 kept crashing my VS Code.
Nothing wrong with the model—it was just so good that I started running 5, 6, even a dozen Claude Code sessions at once. My 16GB MacBook Air couldn't keep up. I upgraded to a 48GB MacBook Pro, which mostly solved the crashes, but I still had no visibility into what was eating memory. Activity Monitor shows processes, but not answers like "which Claude session is the hog?" or "can I spawn another subagent?"
So I built [RAMBar](https://maxghenis.com/rambar)—my first macOS app, built entirely with Claude Code. It shows RAM usage the way developers think about it.
## What It Does
RAMBar lives in your menu bar showing current RAM percentage, color-coded by status:
- **Green**: Under 70%, you're fine
- **Yellow**: 70-85%, getting tight
- **Red**: Over 85%, time to close something
Click it to see a breakdown of what's actually using your memory.
## Developer-Focused Breakdowns
Unlike generic system monitors, RAMBar understands developer workflows:
- **Claude Code sessions** — Shows main sessions vs subagents separately, so you can see when parallel agents are spawning
- **VS Code workspaces** — Memory grouped by which project folder is open
- **Chrome tabs** — Lists individual tabs by memory, so you can find that one tab eating 2GB
- **Python processes** — Useful when you're running Jupyter notebooks or ML training alongside everything else
## How it compares to other tools
Several macOS menu bar monitors exist, but none understand developer workflows:
| Tool | Price | Developer context | Claude Code awareness |
|------|-------|-------------------|----------------------|
| **RAMBar** | Free | ✅ Groups by workspace/session | ✅ Shows sessions + subagents |
| [Stats](https://github.com/exelban/stats) | Free | ❌ Process-level only | ❌ |
| [iStat Menus](https://bjango.com/mac/istatmenus/) | $12 | ❌ Process-level only | ❌ |
| [MenuBar Stats](https://seense.com/menubarstats/) | $5 | ❌ Process-level only | ❌ |
| Activity Monitor | Free | ❌ Process-level only | ❌ |
[Stats](https://github.com/exelban/stats) is excellent for general system monitoring—CPU graphs, network throughput, disk I/O—and it's free and open source. [iStat Menus](https://bjango.com/mac/istatmenus/) adds weather widgets, fan control, and extensive customization for $12.
RAMBar solves a different problem. When I see high memory usage, I don't want to know that `node` is using 4GB—I want to know *which VS Code workspace* that node process belongs to. I don't care that there are 47 Chrome Helper processes; I want to know which *tab* is the culprit. And when Claude Code spawns subagents, I want to see that hierarchy, not a flat list of identical `claude` processes.
If you need comprehensive system monitoring, use Stats or iStat Menus. If you're an AI-assisted developer who wants to know why your Mac is struggling during a coding session, RAMBar fills that gap.
## Get it
```bash
brew tap maxghenis/tap
brew install --cask rambar
```
[Project page](/rambar) ・ [GitHub](https://github.com/MaxGhenis/rambar)
Requires macOS 14.0+. First launch: right-click and select Open to bypass Gatekeeper (the app is currently unsigned).
---
*Part of [12 Days of Shipping](https://maxghenis.com).*
---
## How to Use MDX for Interactive Posts
**URL:** https://maxghenis.com/blog/using-mdx/
**Published:** Nov 25 2025
**Description:** This blog supports MDX, allowing you to embed React components, Plotly charts, and interactive elements directly in posts.
This blog is built with [Astro](https://astro.build) and supports MDX, which lets you embed interactive components directly in your markdown posts.
## What You Can Do
With MDX, posts can include:
- **React components** with full interactivity
- **Plotly charts** for data visualization
- **Custom calculators** or interactive widgets
- **Embedded iframes** for external apps
## Example: Interactive Button
Here's a simple interactive component embedded in this post:
import HeaderLink from '../../components/HeaderLink.astro';
Click me - I'm interactive!
## Adding Plotly Charts
For data-heavy posts, you can create a React component with Plotly and embed it:
```jsx
// src/components/MyChart.tsx
import Plot from 'react-plotly.js';
export default function MyChart() {
return (
);
}
```
Then in your MDX post:
```mdx
import MyChart from '../../components/MyChart';
```
The `client:load` directive tells Astro to hydrate the component on the client side.
## Embedding External Apps
You can also embed Streamlit apps, Observable notebooks, or any iframe-compatible content:
```html
```
## Learn More
- [MDX Documentation](https://mdxjs.com/docs/what-is-mdx)
- [Astro MDX Integration](https://docs.astro.build/en/guides/integrations-guide/mdx/)
- [React Plotly.js](https://plotly.com/javascript/react/)
---
## VAT thresholds, revenues, and the role of counterfactuals
**URL:** https://maxghenis.com/blog/vat-thresholds/
**Published:** Aug 30 2025
**Description:** Debunking claims that raising the UK VAT threshold would increase tax revenues through growth effects.
The Telegraph [reported today](https://www.telegraph.co.uk/business/2025/08/30/reeves-talks-to-raise-vat-threshold-battle-to-grow-economy/) that UK Chancellor Rachel Reeves is considering raising the VAT registration threshold—the turnover above which businesses must register for VAT. The Telegraph's source reported that the Office for Budget Responsibility (OBR) determined that raising the threshold would eventually increase tax revenues.
Readers might interpret this as evidence the VAT sits on the right side of the [Laffer curve](https://en.wikipedia.org/wiki/Laffer_curve)—where tax cuts generate enough growth to pay for themselves. Four major AI models—[ChatGPT](https://chatgpt.com/share/68b305a8-5b7c-8001-a4c9-ec19432abe31), [Claude](https://claude.ai/share/1ade8513-c6b1-4548-af40-44cf25f166b9), [Gemini](https://g.co/gemini/share/2b375e7d35fc), and [Grok](https://grok.com/share/bGVnYWN5LWNvcHk%3D_be2046d9-8b0f-4fd3-ad97-b729c70ee0c1)—reached this conclusion when I tested them, attributing long-run gains to reduced [bunching](https://obr.uk/box/the-impact-of-the-frozen-vat-registration-threshold/) and stronger business growth.
The actual explanation is simpler. The revenue increase in 2028-29 occurs because the policy threshold (£90,000) will be lower than the inflation-indexed baseline threshold by that point.
## What the government projected
While the Telegraph did not link an OBR reference, the sources were likely referring to HMRC's [policy paper](https://www.gov.uk/government/publications/vat-increasing-the-registration-and-deregistration-thresholds/increasing-the-vat-registration-threshold#summary-of-impacts) from March 2024, which analyzed raising the VAT threshold from £85,000 to £90,000. The OBR certified these costings. [HMRC](https://www.gov.uk/government/publications/vat-increasing-the-registration-and-deregistration-thresholds/increasing-the-vat-registration-threshold#summary-of-impacts) states directly below the revenue table: "This measure is not expected to have any significant macroeconomic impacts."
The revenue increase in year five requires a different explanation.
## The role of the counterfactual
In a footnote, HMRC refers to a [policy costings document](https://assets.publishing.service.gov.uk/media/65e7920c08eef600155a5617/Published_Costing_Document_Spring_Budget_2024_Final.pdf#page=14) for more detail, but it doesn't shed light on the question. Instead, we must rewind another year to the [March 2023 analysis](https://obr.uk/box/the-impact-of-the-frozen-vat-registration-threshold/) of VAT threshold freezing:
> One of the many aspects of the tax system that is currently subject to a freeze (meaning a threshold is not being raised in line with inflation) is the VAT registration threshold... The threshold has been frozen at that level for eight years, until March 2026. *Our forecast assumes that this freeze will raise £1.4 billion a year in VAT revenues by 2027-28*, by bringing more businesses above the threshold than if the threshold had been indexed to inflation... *compared with indexing the threshold to RPI inflation*.
The baseline freezes the threshold at £85,000 until March 2026, then indexes it to the Retail Price Index (RPI).
## Running the numbers
I couldn't find the government's baseline VAT threshold projections as of March 2024, so I reconstructed them using RPI inflation forecasts from [Table A-1 of their March 2024 Economic and Fiscal Outlook (EFO)](https://obr.uk/efo/economic-and-fiscal-outlook-march-2024/#annex-a):
- 2.0% from 2024 to 2025
- 2.5% from 2025 to 2026
- 3.0% from 2026 to 2027
By 2028-29, the baseline threshold reaches £92,000 through inflation indexing. The £90,000 policy threshold is now £2,000 *below* the baseline, generating the £65 million revenue gain—not through economic growth, but by setting a lower threshold than what would have existed under current law.
## Conclusion
Raising the VAT threshold could boost UK growth by reducing tax distortions, cutting administrative burdens, and encouraging business expansion. But the government's projections incorporate none of these effects.
The revenue increase in 2028-29 reflects inflation indexing. The £90,000 threshold generates additional revenue because it's lower than the baseline would be by then—purely a mechanical artifact of how the counterfactual was constructed.
---
## AI models favor Cuomo over Mamdani on NYC housing production
**URL:** https://maxghenis.com/blog/ai-models-nyc-housing/
**Published:** Jun 23 2025
**Description:** An experiment using six AI models to forecast NYC housing production under two mayoral candidates.
I spent the weekend in New York City, where canvassers requested my vote for Tuesday's primary election at least a dozen times. While New Yorkers will vote on several offices, most of the energy went to the ranked-choice mayoral, where prediction markets call it a two-candidate race: former governor [Andrew Cuomo](https://www.andrewcuomo.com/) leads Assemblymember [Zohran Mamdani](https://www.zohranfornyc.com/) with roughly 2:1 odds of victory ([73% chance of winning on Manifold](https://manifold.markets/HillaryClinton/who-will-be-the-democratic-nyc-mayo), [70% on Polymarket](https://polymarket.com/event/who-will-win-dem-nomination-for-nyc-mayor), and [69% on Kalshi](https://kalshi.com/markets/kxmayornycnomd/new-york-city-mayoral-nominations)).
Returning to my home in Washington, it struck me that I hadn't yet seen a quantitative analysis of the leading candidates' housing plans. Indeed, two main housing supply advocacy organizations—[Open New York](https://opennewyork.org/endorsements/) and [NYC New Liberals](https://www.nycnewliberals.org/2025-voting-guide)—declined to endorse either Cuomo or Mamdani. What if frontier AI models could fill the gap?
My experiment started with creating two prediction markets on [Manifold Markets](http://manifold.markets) asking how many housing units NYC would have by 2029—the final year of the four-year term—conditional on either Andrew Cuomo or Zohran Mamdani winning the mayoral election. I then provided six AI models with the market text and requested probability distributions:
- OpenAI o3-pro (normal and deep research versions)
- Anthropic Claude 4 Opus (normal and deep research versions)
- Google Gemini 2.5 Pro (normal and deep research versions)
## Key findings
**Housing stock predictions (2029):**
- Cuomo scenario: 3,924,000 homes
- Mamdani scenario: 3,869,000 homes
- Difference: 55,000 more units under Cuomo
**Growth comparison (from 2023 baseline of 3,705,000 homes):**
- Cuomo growth: 219,000 units (5.91%)
- Mamdani growth: 164,000 units (4.43%)
- **Result: 34% more growth under Cuomo**
Because the mayor takes office in January 2026, and the survey is conducted January-June 2029, this provides approximately 3.25 years of mayoral influence. Any growth occurring between the 2023 survey and the election may not be attributable to the election, and it would dilute the delta in growth. Alternatively, we can estimate the baseline growth excluding projected growth from 2023 to 2026. Assuming growth continues at the 27,500/year pace of [2022](https://data.census.gov/table/ACSDT1Y2022.B25001?g=160XX00US3651000) to [2023](https://data.census.gov/table/ACSDT1Y2023.B25001?g=160XX00US3651000) (per the American Community Survey), NYC is "nowcasted" to have 3,760,450 homes at the beginning of 2025. The growth over the term then falls to 164,000 under Cuomo and 109,000 under Mamdani: **50% more growth under Cuomo**.
## Model variation
Model predictions differed significantly:
- OpenAI o3-pro without Deep Research: 75,000 unit advantage for Cuomo
- OpenAI o3-pro with Deep Research: 36,000 unit advantage
- Google Gemini 2.5 Pro projected larger differentials than Anthropic Claude 4 Opus
## Factors identified
**General growth drivers:**
- Existing 421-a pipeline projects
- 1.4% vacancy rate (lowest since 1968)
- City of Yes zoning reforms (~5,500 units/year)
- Expected declining interest rates
**Cuomo-specific factors:**
- 467-m office conversion incentive
- Real estate industry relationships
- Proposed $5 billion capital fund
**Mamdani-specific factors:**
- 200,000 public units over 10 years
- Rent freeze on stabilized units
- Requires state approval for $70 billion financing
## Conclusion
I've increasingly found that AI models give more thorough responses when prompted with quantitative exercises like these, whether for [trade deficit policies](https://www.linkedin.com/posts/maxghenis_what-happens-when-you-give-an-ai-economist-activity-7325825369821884416-WUAN?utm_source=share&utm_medium=member_desktop&rcm=ACoAAAHRBmABgFgIdjcDFs-ZD8yHUI3Rd1pqlWY) or [sleep therapies](https://x.com/MaxGhenis/status/1929213247443423622). For AIs and humans alike, forecasting forces one to identify the most relevant evidence to explore the dynamics at play—and to recognize that magnitudes matter.
Time will tell how accurate the AIs' predictions are. For now, I'm giving my mana as a proxy for their forecasts; if you disagree, I invite you to bet against them.
You can read the [full 48 pages of AI predictions here](https://docs.google.com/document/d/1g2IrB7p14vL6zagq7jIQaTwu1BUNMswl0mMSlUjn1vE/edit?usp=sharing), and [view the data here](https://docs.google.com/spreadsheets/d/1wJUly5dNDof42i259dhPkBOlcGTPe_ye1sv-Zu3mTCM/edit?gid=0#gid=0).
---
## My 2024 in code: Personal projects and AI tools
**URL:** https://maxghenis.com/blog/my-2024-in-code/
**Published:** Jan 02 2025
**Description:** How AI helped me build useful software while leading PolicyEngine.
As 2024 drew to a close, I reflected on my coding journey. Leading PolicyEngine meant less time programming, but I still maintained 3,721 GitHub contributions last year (down from ~4,500 in prior years), spanning collaborations, third-party contributions, and personal projects.
While most of those contributions went to PolicyEngine's core software, where I continue to review code and build with our growing team (see our 2024 year-in-review posts for the [US](https://policyengine.org/us/research/us-year-in-review-2024) and [UK](https://policyengine.org/uk/research/uk-year-in-review-2024)), 2024 was also my most enjoyable and productive year of coding, thanks largely to AI tools like Claude. Here are some of the projects I built alongside our main work:
## CodeStitch
[codestitch.dev](https://codestitch.dev) stitches GitHub repos, PRs, and issues into single text files to make it easier to give code as context to AI assistants. It started as a Streamlit app and evolved into a React application in December. If you've ever tried to get AI help with a codebase, you know the pain of context limitations—CodeStitch helps solve that.
## PR Improver
[pr-improver.streamlit.app](https://pr-improver.streamlit.app) suggests improvements to pull requests using LLMs and the GitHub API, comparing against contributor guidelines in the repo. While it currently only works with PolicyEngine repos (since each run costs us money), I'm planning to add support for other repos where you can bring your own LLM API key.
## OEWS Explorer
[oews-explorer.streamlit.app](https://oews-explorer.streamlit.app) helps explore the Bureau of Labor Statistics Occupational Employment and Wage Statistics by occupation and geography. I built it to support PolicyEngine's NSF grant application, but it's also useful for compensation research.
## Personal website
I refreshed [maxghenis.com](https://maxghenis.com) with React, creating a cleaner, more maintainable platform for sharing my work and writing.
## Looking forward
AI has transformed programming more than any field I've encountered. You can now ideate, prototype, and build with unprecedented speed and creativity. Tools like Claude and Cursor don't just help you debug—they're collaborative partners in the development process, helping you explore architectural decisions, discover edge cases, and learn new frameworks.
Whether you're a seasoned developer or curious leader, I highly recommend making time for coding in 2025. You might surprise yourself with what you build. I'm excited to share more projects in the coming year!
Want to explore these tools? Check out the links above and let me know what you think—they're all open source. And if you're interested in contributing to PolicyEngine, we're always looking for collaborators—drop by our [GitHub repos](https://github.com/PolicyEngine).
---
## Why I'm a neoliberal
**URL:** https://maxghenis.com/blog/why-im-a-neoliberal/
**Published:** Mar 18 2022
**Description:** My political philosophy and why I identify with neoliberalism - emphasizing evidence-based policy, market mechanisms, and a strong social safety net.
I'm competing in the Neoliberal Twitter account's annual bracket competition, which gives me an opportunity to explain my commitment to neoliberal ideology. I position neoliberalism as a pragmatic, evidence-based approach to solving pressing policy challenges through market mechanisms and strategic government intervention.
## What is neoliberalism?
Neoliberalism has been mischaracterized as selfish or warmongering. Instead, it's rooted in liberal values, emphasizing evidence-based policymaking and practical solutions. Key tenets include deregulation of housing and labor markets, carbon pricing, immigration, globalization, and a robust social safety net.
## Housing deregulation (YIMBYism)
I advocate for legalizing apartment construction, particularly dense infill housing near jobs and transit. Overregulated land use restricts freedom of movement, impedes economic growth, and harms the environment.
My activism includes:
- Joining YIMBY advocacy in San Francisco (2016)
- Campaigning for YIMBY pioneer Sonja Trauss's supervisor campaign (2018)
- Co-founding YIMBY Neoliberal organization (2018)
- Creating Ventura County YIMBY (2019), which gained over 1,500 followers
## Carbon pricing
Carbon taxation is the key ingredient for cutting carbon emissions. The IPCC identifies carbon pricing as crucial for cost-effective climate mitigation, and studies worldwide demonstrate its effectiveness.
Carbon pricing enjoys broad support: average net favorability across 18 polls since 2015 reached +33 percentage points. Most developed nations price carbon; the U.S. and Australia remain exceptions.
My contributions:
- Joined Citizens' Climate Lobby (2018)
- Led my local chapter and became California state coordinator
- Researched carbon dividends professionally
- Co-founded "Save the Gas Tax" campaign (2022) with 15 environmental organizations
## Social safety net
Economic growth alone cannot guarantee minimum living standards for non-workers or address perverse incentives in current welfare systems. I advocate for:
- Universal child allowances
- Cash assistance over in-kind benefits
- Universal basic income (UBI) as the logical endpoint
I founded the [UBI Center](https://ubicenter.org), conducting quantitative policy analysis, and co-founded [PolicyEngine](https://policyengine.org), a nonprofit technology platform enabling citizens to understand tax and benefit policy impacts. Means-testing creates administrative bloat and extreme marginal tax rates for low-income households.
## Conclusion
I'm an "extreme liberal" seeking the most effective policy solutions: address poverty through direct cash transfers, housing affordability through legalization rather than government management, and pollution through market-based pricing mechanisms. These neoliberal policies are transformative when combined.
---
## In appointing their newest member, the Ventura City Council only pretended to care about housing
**URL:** https://maxghenis.com/blog/in-appointing-their-newest-member-the-ventura-city/
**Published:** 2021-02-24
**Description:** After asking each applicant about the housing crisis, most councilmembers voted for those least committed to solving it
On Saturday, the Ventura City Council held a special meeting to appoint someone to the seat of former District 4 Councilmember, now District Attorney, Erik Nasarenko, until the election in November 2022. They selected Jeannette Sanchez-Palacios, a concerning choice for housing (more on that below), but nearly selected a straight-up NIMBY, in a process that revealed the disinterest of many councilmembers in addressing Ventura's housing challenges.
Ventura followed a similar routine from earlier this month, when Oxnard filled now-Supervisor Carmen Ramirez's seat. We sent a questionnaire to all applicants for the Oxnard seat, and after receiving [five responses](https://datastudio.google.com/u/0/reporting/207f8afe-b5f0-4a61-9572-6417bc724eb0/page/KE9fB), [we endorsed Gabe Teran](https://medium.com/vcyimby/ventura-county-yimby-endorses-gabe-teran-for-oxnard-city-council-district-2-4cf2c0f8c5e4), who ended up [getting the appointment](https://medium.com/vcyimby/newly-appointed-oxnard-councilmember-gabe-teran-stands-strong-on-housing-2990a6ee2dd2).
We also sent a questionnaire to all fourteen Ventura applicants; however, only one, design consultant John Silva, completed it. While [Silva's responses](https://drive.google.com/file/d/14xioDqC6fddhS_editAsdtfVUNY3A9VZ/view?fbclid=IwAR1U8WRj1Abe55Y2aly9LNlF9kEy4I-lj92jH6IXqfmKrVKfaen51hHFG_s) were strongly pro-housing, we didn't feel we had enough information about the full set of applicants to make a formal endorsement.
Thankfully, the Ventura City Council voted to ask each applicant about housing in their interviews; here's what they asked (in addition to opening and closing statements):
- "Our community is facing a housing crisis, worsened by the current pandemic. What will you do to address rising rents and protect renters from eviction?"
- "What do you think are the top three to five priorities for the city, and what will you do as a councilmember to achieve them?"
- "Circumstances outside the city's control can and have impacted dramatically the city's budget. What approach will you take as a councilmember to balance the budget to address the issues?"
These questions revealed significant differences between the applicants on housing ([watch the video here](https://youtu.be/nfIY3yKj-XA)).
Several applicants acknowledged the role of the shortage. The first to speak, John Lory, [said](https://youtu.be/nfIY3yKj-XA?t=2269), "The more supply there is, the lower the rents will be, and they will increase at a lower rate." And the final applicant, Dan Lyon, said, "We need to look at a way to increase the supply of affordable rental housing," suggesting infill and repurposing existing properties.
### Brad Golden and John Silva supported housing development
The strongest pro-housing interview response came from Brad Golden, a manager at a real estate insurance company; here's his full response (**emphasis mine, **all quotes lightly edited for clarity):
> Rising rents are not new to our community and California, but specific to the pandemic, our government has placed eviction protections for Covid-related hardships through the middle of this year. I believe they'll likely extend that and I hope they do. We also have city and county tenant protections during Covid, and I hope that continues as well. Ventura extended protections for tenants to the commercial side, and I know that's helped quite a bit too, and we the Chamber of Commerce have been active in that role. Prior to Covid, California capped rental increases at 5% plus cost of living, so no matter what, we have a rent control in place, and I know that's helped with tenants as well.
> But we have to remain balanced for the landlords and owners as well. Over-regulating market-driven economics cannot lead to neglect of properties in the long run. Bottom line, Covid or not: we need to work on supply and demand for housing. We need to streamline the development process and build a diversified housing stock.
Golden also emphasized housing development in his list of top three priorities for the city, which included:
- A pandemic response plan
- Economic stimulation, including "an efficient development process"
- Diversified housing stock, including being "flexible with our General Plan coming up," a point he also brought up in terms of budget sustainability
Silva's interview response on the housing question didn't offer specifics, but his response to our questionnaire indicated support for policies like raising height limits near transit and legalizing fourplexes in many areas.
### Jenny Lagerquist opposed housing development
The starkest contrast to Brad Golden's pro-housing views came from Jenny Lagerquist, an engineer, who gave the following response to the housing question:
> This is a complex question. Because of the pandemic, the issue is not specific to rents and renters, I believe homeowners and business owners are feeling the same pressures. I think everyone needs to share in the solutions. So how do we propose a way forward, allowing everyone to keep their spaces and run their businesses and provide a roof over their family's heads.
> First, we should be looking to stay in federal grants to help protect homeowners, business owners, and renters from eviction. The state did just pass an aid package to assist those hit hardest by the pandemic, so we could help educate and guide residents to this information. And then we can look at what can the city provide, are there other types of aid, local grants, temporary rent control? Affordable housing is a very complex issue and many different facets have to be considered: on transportation, some studies have shown that people actually do want to live within a reasonable commute, so the more mass transit or public transportation that's available, the easier to renovate or build affordable housing. Just building affordable housing without proper planning is not always the answer. Affordable housing policy can actually favor owners over renters, and we need to continue to provide a balance. It also has a long history of segregation and racism and can be expensive to build based on labor and supplies, and again transportation issues have to be addressed. Developers need incentives to build and renovate low and moderate private rental market housing. Right now, most incentives are for luxury housing, especially in a state like California. The private sector won't build more unless the incentives change, so maybe tax credits for that kind of growth. Assistance to the low income bracket is always required, and we should be studying and seeking out grants and available funding.
Lagerquist clarified her anti-development stance in the second question, saying:
> The number one priority for the city is to always try to retain the uniqueness and authenticity of Ventura. The Council has to scrutinize any development plan and encourage redevelopment of deteriorating spaces. It would behoove the council to require developers of any type of growth or expansion or rebuild or renovation, etc., to share with the cost of infrastructure and public needs. The city alone should not burden those costs. City planning should clearly demonstrate preservation of the city's assets and aesthetics. I think water is a tremendous issue, we have to require and encourage conservation imposing permanent restrictions if needed.
Her statement that "My focus again is certainly on preservation and maintaining Ventura" aligned with her [application](https://www.cityofventura.ca.gov/DocumentCenter/View/25998/Laqerquist-Jenny_Redacted) pledge to "promote only smart growth." Transit, funding, and water sustainability are good, but Lagerquist invokes all as classic excuses to avoid allowing development. Ventura can't change state and federal policymaking around grants, but they can decide whether to allow new homes to be built, which would generate tax revenue to provide their own assistance programs and expand transit. Singular focus on aesthetic preservation is antithetical to that goal.
### Jeannette Sanchez-Palacios ignored the housing shortage
Jeannette Sanchez-Palacios, district director for Assemblymember Jacqui Irwin, mostly avoided answering the housing question, though she did consider it a "huge issue" that Ventura has renters, and recommend building more single-family housing:
> I've thought about this because recently I met with a group of renters here in the city of Ventura, and let me tell you, they were all women. We know that this pandemic has impacted a lot of our minority communities. And especially women, it has driven out of the workforce, or their partners out of the workforce. And some of the women, I heard stories they were renting one room in a bigger house that they share with their teenagers. Or one of the other women there, she had just had surgery, and wasn't working, didn't have sick days or paid leave, so she wasn't sure how she was going to pay rent. So this issue is very prevalent in the city of Ventura as I did my research I found that close to 19,000 renter households exist here in the city of Ventura. That's approximately 46% of all the households in the city, almost 50%. And so for me that's a huge issue to look at. Luckily, federal and state governments have been protecting renters. We have a new program, the emergency rental assistance program, and I think as city leaders, it's incumbent upon us to help our residents reach these resources that are being made available to them.
> The other thing that came to mind was the community land trust programs that can get families into single family housing and in addition can help drive down the cost of rents and home prices. I think this one in particular is very important because these programs can also potentially be an avenue for the city to fulfill the Regional Housing Needs Assessment numbers, a.k.a. the RHNA numbers, which are a state mandate. The city has to look at that as we move forward and ensure that we can meet those numbers.
This response notably omits mention of our housing shortage, except to the extent that she supported obeying state law. But her past offers more insight into her housing views. When Sanchez-Palacios ran for city council in 2016, she was endorsed by [Neighbors for the Ventura Hillside](http://www.venturahillsideneighbors.com/), a NIMBY group defined by its [Ventura Vision](https://www.facebook.com/VenturaHillsideNeighbors/posts/3666729843357316):
> We will grow slowly and sustainably (meaning we will respect our water supplies and our constraints such as traffic). We will protect our environment and open space. We will preserve our rivers, beaches and natural environment. We will insist on the highest quality of architecture and design. We will maintain our small town atmosphere.
Despite the group's name, they primarily [oppose](https://www.facebook.com/VenturaHillsideNeighbors/posts/3572397196123915) [dense](https://www.facebook.com/VenturaHillsideNeighbors/posts/3579702212060080) [infill](https://www.facebook.com/VenturaHillsideNeighbors/posts/3670263143003986) [developments](https://www.facebook.com/VenturaHillsideNeighbors/posts/3346468822050088). They've since endorsed some of the most well-known NIMBYs in Ventura, such as former councilmember Christy Weir, whose [website](https://www.christyweir.com/issues/) advocates "limits on 'high density' and mixed use development to prevent over-urbanization." (We [endorsed Weir's opponent](https://www.vcyimby.org/endorsements#h.acdsvj6ekvnq), Doug Halter, in his successful November 2020 run.) Sanchez-Palacios didn't only accept the endorsement of Neighbors for the Ventura Hillside, she recorded [two](https://www.facebook.com/933553936674934/videos/1288276611202663) [promotional videos](https://www.facebook.com/VenturaHillsideNeighbors/videos/1267762573254067) with them.
### Councilmembers nearly appointed the most NIMBY of the applicants
The Ventura City Council selected their appointee in a two-stage process. First, each councilmember selected their top two applicants, and the two receiving the most votes advanced to the next round. In the second round, each councilmember selected their top applicant of these two, and the applicant with the most votes was appointed.
In the first round, councilmembers Halter and Jim Friedman voted for Golden and Silva, while the others — Lorrie Brown, Mike Johnson, Deputy Mayor Joe Schroeder, and Mayor Sofia Rubalcava — voted for Sanchez-Palacios and Lagerquist, advancing them. In the second round, Friedman and Schroeder voted for Lagerquist and the others voted for Sanchez-Palacios.
No councilmembers mentioned housing in their decisions.
Any elected official knows what statements like those from Lagerquist indicate about housing stances, yet all but Doug Halter voted for her in either the first or second round. Christy Weir may have voted for Lagerquist over Sanchez-Palacios, and if the more pro-housing opponents of Schroeder and Johnson (whom Neighbors for the Ventura Hillside also endorsed) had won in November, they might have appointed Silva or Golden instead. NIMBYism begets NIMBYism, but so too can YIMBYism beget YIMBYism.
Sanchez-Palacios suggested she may run for election in November 2022. Her NIMBY alliances and disinterest in addressing the housing shortage are concerning, but until then, we'll be pushing her to pursue the strong housing policy Ventura needs.
---
## If you earned more in 2019 than 2018, don't file your 2019 taxes yet! Otherwise, file ASAP!
**URL:** https://maxghenis.com/blog/if-you-earned-more-in-2019-than-2018-dont-file-you/
**Published:** 2020-03-27
**Description:** The right tax filing schedule could get you thousands in (poorly designed) Recovery Rebates. And Congress: Let's just do UBI next time.
Late Wednesday night, the Senate unanimously passed the [CARES Act](https://drive.google.com/file/d/1eT-ANEEd-AgAzyQh8GVUGoRqtKRLHxDm/view), a $2.2 trillion relief package to address the Covid-19 pandemic. The House is [expected to pass it](https://www.washingtonpost.com/politics/house-leaders-seek-to-expedite-emergency-aid-package-amid-uncertainty-about-gop-lawmaker-delaying-measure/2020/03/26/392c9dba-6f7d-11ea-a3ec-70d7479d83f0_story.html) on Friday.
Of all the components of the nearly-900-page bill, the Recovery Rebates will have the broadest and most tangible reach. [Over 93 percent](https://taxfoundation.org/cares-act-senate-coronavirus-bill-economic-relief-plan/) of tax filers will receive these one-time checks of up to $1,200 per adult and $500 per child.
There's a catch, though. The amounts start phasing out for tax filers with income over $75,000 (single) or $150,000 (married), and that income comes from your latest tax filing: 2019 if you've filed, otherwise 2018.
Filers with lower earnings in 2020 than the year of the latest tax return (2018 or 2019) will get a credit on their 2020 tax refunds next year. But the reverse isn't true; filers that get more from the current payment than they would have based on 2020 income [don't have to pay it back](https://thehill.com/policy/finance/489578-questions-and-answers-on-coronavirus-relief-checks). So it's in each filer's interest to get as high a payment immediately as possible.
Treasury Secretary Steven Mnuchin said today that the checks will start coming [in the next three weeks](https://www.cnbc.com/2020/03/26/coronavirus-stimulus-checks-will-come-within-three-weeks-mnuchin-says.html). That might mean it's already too late, and those that haven't filed for 2019 yet could be stuck with 2018 information. Or they might have a couple weeks. The IRS hasn't given any information either way.
One thing's for sure: there's no need to worry about timing it around the normal April 15 tax day. The IRS extended the deadline to [July 15](https://www.irs.gov/newsroom/tax-day-now-july-15-treasury-irs-extend-filing-deadline-and-federal-tax-payments-regardless-of-amount-owed).
That's it on the PSA for maximizing the recovery rebate. If you're interested in how to avoid poor policy design like this in the future, read on.
### This is why you don't means-test emergency relief money
There's no good way to means-test outside the tax system. SNAP (food stamps), TANF (welfare), and Medicaid restrict to low-income households through an onerous process often involving in-person interviews, mandatory contact each time a recipient's income change, and regular recertification of work/work-seeking and other circumstances. Unemployment insurance, currently facing stresses keeping up with new volume, also requires in-depth interactions. All these programs differ by state, or even by county, and a bunch of other small programs like discounted transit often just condition on SNAP receipt.
There's *especially *no good way to means-test in the middle of an emergency, when Washington wants to help most of the country afford basic necessities in a matter of weeks. There's no time or resources for anything like the SNAP process, and nationwide consistency is an essential ingredient for quick passage through Congress.
Using old tax data is the only way, and doing so doesn't only penalize people for either filing too early or too late. What if a person got married or divorced since their last tax filing? Or had a child, or had a child age out of their household? Or even more complicated, if they got divorced and remarried? The IRS is going to have to make some tough calls to figure this out, and they'll make mistakes that will force exes to figure out how to split the new money. And then what about the reconciliation in next year's taxes?
Using taxes also means sending money by tax filing unit, which can create hidden discrepancies. For example, the Recovery Rebate only gives money for primary filers and child dependents. If someone claims a college-age child on taxes, or other adult dependents like an elderly parent or someone with a disability, [they get $0](https://medium.com/whatever-source-derived/dont-exclude-elderly-and-disabled-dependents-from-covid-19-aid-54fc59eb43f3) to help them through this crisis.
All this effort, cost, and complexity cannot possibly be worth a bit of savings from excluding the top 7 percent of households.
Instead, we should just give to everyone. Suppose we gave the $1,200 to every adult and $500 to every child. Based on the [current population](https://www.census.gov/content/dam/Census/library/publications/2020/demo/p25-1146.pdf), that'd cost $348 billion, or an extra $56 billion above the [$292 billion](https://www.jct.gov/publications.html?func=startdown&id=5252) cost of the Recovery Rebates. That's about 2.5 percent of the total CARES Act.
If we want to, we can effectively tax that money back next year. But even if we don't, we can expect high income people to lose a larger dollar amount than other groups, based on past recessions. More progressive taxation is a fine way to save money — and we'll surely need more taxes to pay off the relief costs once this is all over — but either way, most rich people will lose far more income than the couple grand they'll get in rebates.

Some of this will still be slowed by transaction time. The Treasury will have to spend more effort reaching people without recent tax records, or without existing direct deposit for receiving tax refunds. We could speed this up in the future by offering everyone government bank accounts, or connecting them with direct deposit or debit cards like Social Security does. Other forms of universal transfers — even small ones like a [carbon dividend](http://energyinnovationact.org) or a [child allowance](https://www.vox.com/future-perfect/2019/3/6/18249290/child-poverty-american-family-act-sherrod-brown-michael-bennet) — would create payment infrastructure that could be jacked up in times of need.
It might not be so long before this choice rears its head again. House Speaker Nancy Pelosi has said she's [considering more cash payments](https://thehill.com/homenews/house/489587-pelosi-democrats-eyeing-more-cash-payments-in-next-emergency-bill) in the next relief bill. She should join the eight congresspeople from both parties who have [called for making them universal](https://medium.com/ubicenter/tracking-us-emergency-cash-transfer-proposals-7c4bf90e8ac8). Crises are not the time for convoluted means-tests.

---
## Why I’ve taken the Giving What We Can pledge
**URL:** https://maxghenis.com/blog/why-ive-taken-the-giving-what-we-can-pledge/
**Published:** 2019-12-31
**Description:** The global poor will benefit much more from ten percent of my income than I will
#### The global poor will benefit much more from ten percent of my income than I will
**TL;DR:*** I’m pledging to give at least ten percent of my income to highly effective charities from 2019 on. For this year, I’m giving mostly to *[GiveDirectly](http://givedirectly.org)* (which provides direct cash transfers to extremely poor households) and *[GiveWell](http://givewell.org)* (which directs donations quarterly *[at its discretion](https://www.givewell.org/about/FAQ/discretionary-grantmaking)*, generally to life-saving treatments like malaria bed nets). To join me and the *[over 4,400 members](https://www.givingwhatwecan.org/about-us/members/)* of *[Giving What We Can](https://www.givingwhatwecan.org/)*, visit *[givingwhatwecan.org/pledge](http://givingwhatwecan.org/pledge)*.*
When I was about 12, I sat around the dining room table with my family, paging through the catalog of the charity *Heifer International*.* *We were looking for the best animal to buy for a poor farmer in sub-Saharan Africa. We felt warm and fuzzy as we settled on a goat, knowing we were helping someone with far less than we’d ever had. While we’d given to homeless people and local services before, the Heifer catalog is my first salient memory of giving.
Once I started working after college, I was glad for the chance to give my own money to charities. Heifer now seemed prescriptive and arbitrary to me (why would I know the best livestock to get a farmer?),¹ so I focused on medical research and interest-free Kiva loans to poor people looking for business investment in developing countries.² My intuitions were that medical advancements served the public good, and that people in poorer countries who sought funding for a specific venture would benefit most from a dollar of giving. But I wondered whether these choices were the best use of my giving budget.
I started to satisfy this curiosity in 2012, when I learned of [GiveDirectly](http://givedirectly.org), a small nonprofit that gave direct, unconditional cash transfers to extremely poor families in Kenya. Their approach was novel but not untested: other studies had shown that giving people cash improved outcomes ranging from nutrition to business investment, and GiveDirectly (founded by economists) was evaluating its own disbursements with randomized controlled trials. GiveDirectly’s focus on efficiency (90 cents of every dollar in donations ended up in the hands of a poor person), transparent measurement, and effective deployment of additional funds, earned it a [“Top Charity”](https://web.archive.org/web/20121201034608/http://www.givewell.org/international/top-charities/give-directly) rating from the evidence-based evaluator [GiveWell](http://givewell.org). GiveWell evaluates charities on the cost-effectiveness of their measured impact, and awarded only three Top Charity ratings in 2012.
Discovering GiveDirectly — and GiveWell’s review of GiveDirectly — became a seminal event in my life. My interest in cash transfers led to researching [universal basic income](http://ubicenter.org) (UBI, which GiveDirectly is [testing](https://www.givedirectly.org/ubi-study/) in Kenya), while the validation of GiveWell’s evaluation contributed to my decision to pursue a [master’s degree](http://economics.mit.edu/masters) in development economics, which I’ll start next month (the program was created by one of the researchers evaluating GiveDirectly’s UBI experiment, who just [won](https://www.nobelprize.org/uploads/2019/10/press-economicsciences2019.pdf) the Nobel Prize).
These organizations have continued to grow and improve. GiveDirectly still gives lump-sum transfers to Kenyans living on a dollar a day, and now operates in [seven](https://www.givedirectly.org/faq/) more countries while distributing in new ways: as [UBI](https://www.givedirectly.org/ubi-study/), to [disaster victims](https://www.givedirectly.org/relief/) and [refugees](https://www.givedirectly.org/refugees/), and as a [benchmark](https://www.vox.com/world/2018/9/13/17846190/cash-saves-lives-rwanda-usaid-foreign-aid-nutrition) for other aid organizations to test their interventions. They’ve completed a number of randomized controlled trial [evaluations](https://www.givedirectly.org/research-at-give-directly/), revealing that recipients have higher earnings, assets, and nutrition spend, without increases to alcohol and tobacco spend. Their [latest study](https://www.vox.com/future-perfect/2019/11/25/20973151/givedirectly-basic-income-kenya-study-stimulus) even finds benefits among non-recipients: households receiving the transfers saw consumption rise, but so did households in the same village who didn’t get the money — and their consumption rose a similar amount.
GiveWell has updated their list of [top charities](https://www.givewell.org/charities/top-charities/) annually, including GiveDirectly each year, along with charities that treat and prevent fatal or debilitating diseases. These other charities are extremely cost-effective at averting deaths — for example, they estimate that the Against Malaria Foundation does so for [$2,300](https://twitter.com/MaxGhenis/status/1210456968701177856) — and boosting consumption among those on the verge of starvation.
I didn’t know all this over the past seven years — some wasn’t yet knowable — but I became increasingly convinced that giving effectively was one of the most important things I could do. So over time, I increased my giving³ and shifted it toward GiveDirectly and other top GiveWell charities.

### Effective Altruism
Over this period, I also learned more about the *effective altruism *(EA) movement. EA began with global development in mind, making GiveWell a sort of standard-bearer. It’s since expanded into charities promoting animal welfare and reducing the risk of massive disasters — e.g. by improving the safety of artificial intelligence systems and nuclear arsenals, and preparing for pandemics — and supporting individuals seeking to dedicate their careers to these issues rather than only donating to them.
EA forces donors to confront some difficult moral questions, like how to weigh saving someone from blindness against saving a child from dying of malaria. Or harder yet, how to weigh these short-term gains against making human extinction a tiny bit less likely.
But one moral question was less ambiguous: whether someone living a relatively comfortable life in a rich country should give generously to effective causes. The philosopher Peter Singer illustrates this with a [thought experiment](https://www.utilitarian.net/singer/by/199704--.htm). In it, you are walking by a pond and notice a child drowning. Should you save them, even if doing so would ruin your clothes and shoes? “Of course,” most would say. Do you not, then, have the same obligation to save children across the world from us, when it comes at a similarly low cost?
[Giving What We Can](https://www.givingwhatwecan.org/) is an EA organization created to commit people to this line of reasoning. Those who take its pledge are committing to give at least ten percent of their income to highly effective charities each year (with [exceptions or adaptations](https://www.givingwhatwecan.org/about-us/frequently-asked-questions/#45-what-about-students--the-unemployed--the-retired) for students, retirees, people with health conditions, etc.). Here’s the text of the Pledge:
> I recognise that I can use part of my income to do a significant amount of good. Since I can live well enough on a smaller income, I pledge that for the rest of my life or until the day I retire, I shall give at least ten percent of what I earn to whichever organisations can most effectively use it to improve the lives of others, now and in the years to come. I make this pledge freely, openly, and sincerely.
Someone earning the US median personal income of [$33,706](https://fred.stlouisfed.org/series/MEPAINUSA672N) — placing them in the [richest 3.7 percent](https://howrichami.givingwhatwecan.org/?income=33706&countryCode=USA&household%5Badults%5D=1&household%5Bchildren%5D=0) of the global population — would then be saving about 1.5 lives per year by giving their pledged amount to the Against Malaria Foundation.
I’ve known of the Giving What We Can pledge for a few years, but finally pulled the trigger for a couple reasons. One, I’ve simply learned more about the horrific scale of global poverty, and the tractability of ending it. While a number of resources have made an impact on me — such as the website [Our World In Data](http://ourworldindata.org), Vox’s [Future Perfect](https://www.vox.com/future-perfect) project, and various posts on the [EA forum](https://forum.effectivealtruism.org/) — the most recent educational source has been Peter Singer’s book, [The Life You Can Save](https://www.thelifeyoucansave.org/)*.*⁴* *Originally published in 2009, Singer [re-released it for free](https://www.thelifeyoucansave.org/the-book/) this month in e-book and audiobook formats. I finally read it, checking off a book long on my list, and was more convinced than ever of the moral and empirical case for giving generously.
There’s also the question of [when to give](https://www.givingwhatwecan.org/about-us/frequently-asked-questions/#24-why-give-now-rather-than-later-in-life-or-in-my-will). I initially wanted to save more before committing to give a larger share of my income, but newer research found that my money would do much more good early on than later. For example, GiveDirectly transfers yield rates of return of [30 percent or more](https://www.givedirectly.org/research-on-cash-transfers/) among recipients, because their businesses are so capital-constrained, far outweighing the 8 percent or so my money earns in stock index funds. And since these countries are developing economically, health interventions that save lives today will probably not be available for such low cost in the future. At a certain point, all the easy-to-reach households will have malaria bed nets, and it’ll be harder to find similarly cheap ways to save lives; it might be a while before it costs millions to save a life as it does in the US, but they’ll progress quickly, and the sooner effective charities can invest in the low-hanging fruit, the more deaths my money can avert.
Finally, *The Life You Can Save* reviews the evidence on what causes people to give. It turns out that social norms are a major factor. The ten-percent tithe is common across religions and employer gift matches affect giving levels, and randomized experiments also show that people are primed by what others are giving when deciding how much to give. That’s why it wasn’t enough for me to give privately; I hope that adding my name to a [public list of pledgers](https://www.givingwhatwecan.org/about-us/members/) (and writing this) encourages others to consider giving a bit more, or to charities that are more proven, or to go all-out and take the pledge.
### How I’m giving this year
At this point it may not be surprising that I’m giving to GiveWell and GiveDirectly: I’m giving each of them five percent of my 2019 income.

[GiveWell](http://givewell.org) allocates their donations quarterly to one or two of their [top charities](https://www.givewell.org/charities/top-charities). I expect that they’ll give to an organization that will save or improve a great number of lives (probably by treating or preventing malaria or worms). I’ve given through GiveWell rather than directly to one of their top charities (e.g. Against Malaria Foundation) because (a) I trust them to make the best choice at the end of the quarter based on need and new information, and (b) doing so signals to GiveWell’s operational funders that their evaluation work is valued.
[GiveDirectly](http://givedirectly.org) will transfer about 85 percent of my donation directly to extremely poor households in sub-Saharan Africa, after transaction and distribution costs. While this might not save as many lives as Against Malaria Foundation, the recipients will see lasting economic and nutritional improvements from the transfer, and they will be fully empowered to make the best use of the money. Supporting GiveDirectly will grow their transformation of the aid sector — and even public policy sector — by emphasizing unconditional cash transfers.
In addition, I’m giving under 1 percent of my income to these two organizations. While neither is an explicitly EA charities of the GiveWell variety, I believe that they are worthy causes by virtue of their [importance, tractability, and neglectedness](https://concepts.effectivealtruism.org/concepts/importance-neglectedness-tractability/).
- Yes In My Back Yard (YIMBY), which makes housing more affordable and accessible through education and advocacy, as well as legal action against cities that violate state and federal housing law. Many studies show that housing shortages in urban areas are causing poverty, climate change, air pollution, homelessness, immobility, and slow economic growth (here’s a review I recently wrote). YIMBY is a 501(c)3 that works with other pro-housing groups I’ve volunteered with, including SF YIMBY and California YIMBY (and Ventura County YIMBY, the CA YIMBY chapter I co-founded) to promote more construction of dense infill homes. From an EA lens, I’m encouraged that Open Philanthropy Project, GiveWell’s sister organization, supports YIMBY groups, and has also identified immigration policy as a focus area, which abundant housing supply can promote.
- Citizens’ Climate Education, which educates the public about carbon fee-and-dividend policies. They’re the 501(c)3 arm of Citizens’ Climate Lobby, which I joined in early 2019. EAs tend not to emphasize climate change because, while it is clearly important, it’s far from neglected. However, I believe carbon pricing is worthwhile because it has greater potential to cut emissions than any other policy, and it lacks the grassroots support of other environmental organizations like the Sunrise Movement or the Sierra Club. Citizens’ Climate Lobby has organized carbon pricing advocates since 2007, and it now has a bill in Congress with 75 cosponsors.
### Until next time
I plan to do another recap like this next year, and hope it will include a new charity or two. For example, I didn’t donate to non-GiveWell EA charities like “longtermism” (AI, nuclear, and pandemic safety); meta EA (supporting organizations like Giving What We Can); or animal welfare. I don’t yet have confidence that these areas are worth diversifying away from GiveWell and GiveDirectly, but I expect that that could change with more research. Now that my stakes are higher, I’m looking forward to researching how to make my giving count even more.
**You can learn more about **[Effective Altruism](https://www.effectivealtruism.org/)**, **[The Life You Can Save](https://www.thelifeyoucansave.org/)**, or **[Giving What We Can](https://www.givingwhatwecan.org/)**, or **[take the pledge](http://givingwhatwecan.org/pledge)**.**
¹ GiveWell [reviewed](https://www.givewell.org/International/charities/Heifer-Project-International) Heifer International in 2009 and determined that it did not meet its criteria for (a) publishing high-quality monitoring and evaluation reports, and (b) focusing on programs with strong evidence bases. They have found [little evidence](https://www.givewell.org/international/economic-empowerment/livestock-gift-programs) supporting livestock gift programs.
² GiveWell [reviewed](https://www.givewell.org/international-giving-marketplaces) Kiva, GlobalGiving, and other “giving marketplaces” in 2010, and expressed concern about fungibility, given projects are likely to be funded regardless of the donation. GiveWell’s “[current view](https://www.givewell.org/international/economic-empowerment/microfinance) is that microfinance is not among the best options for donors looking to accomplish as much good as possible.” Interest-free loans are equivalent to a donation equal to the market rate of return on the principal, plus expected default losses (i.e. the opportunity cost of the loan).
³ I front-loaded some 2018 giving to 2017 for tax reasons. Bunching donations is generally advantageous in the US because of the standard deduction, and the Pledge [allows for this](https://www.givingwhatwecan.org/about-us/frequently-asked-questions/#46-how-often-should-members-donate).
⁴ Singer’s organization, *The Life You Can Save,* has [its own giving pledge](https://www.thelifeyoucansave.org/take-the-pledge/), which I’ve also taken. That pledge focuses on global development charities, i.e. from GiveWell. It has a graduated scale, and asks for less than ten percent of income for households earning less than $500,000 per year. My current income is less than $500,000, so I’m bound by the Giving What We Can pledge, but I’ll shift as needed, and also ensure that my donations to the global poor meet the Singer pledge should I choose to diversify to other EA causes in the future.
---
## Ventura County must end its unsheltered homelessness
**URL:** https://maxghenis.com/blog/ventura-county-must-end-its-unsheltered-homelessne/
**Published:** 2019-11-25
**Description:** We need to make shelters more abundant and housing more affordable
#### We need to make shelters more abundant and housing more affordable
*Yesterday concluded *[Hunger & Homelessness Awareness Week](https://hhweek.org/)*.*
"This is arguably one of the most contentious issues ever," said Oxnard Mayor Tim Flynn about a [new 110-bed homeless shelter](https://www.vcstar.com/story/news/2019/11/20/oxnard-homeless-shelter-location-city-council-decision-former-salvation-army/4229839002/) on Tuesday night. Public comments from neighbors expressed concern over drug use, safety, and even insufficient parking, while homelessness advocates emphasized the need for compassion and shelter. Ultimately, Flynn joined four councilmembers in approving the shelter lease, marking a step toward addressing unsheltered homelessness in Ventura County.
We need far more decisions like these. [1,669 Ventura County residents are homeless](http://www.venturacoc.org/images/2019_VC_Homeless_Count_Report-Final.pdf), of whom 1,258 lack shelter. Those numbers increased 29 and 53 percent, respectively, over the past year.
They should be zero: Nobody in Ventura County should be forced to live on the streets or in parks without access to restrooms and medical care. As Oxnard City Manager Alex Nguyen said at Tuesday's hearing, "This is a public health crisis." We agree, and it should be addressed as one.
> Homelessness is a public health crisis.
With a combination of more shelters to house our homeless neighbors, more subsidies to keep at-risk families in their homes, and more homes to make housing more broadly affordable, we can end unsheltered homelessness in Ventura County. But first, the facts.
### Debunking homelessness myths
Perceptions of homelessness are often based on its most visible instances. Seeing a homeless person doing drugs in public can be shocking, and it's not unusual to then assume that this is the norm.
Those salient occurrences are unrepresentative, however. The stereotype of a chronically homeless out-of-towner with a mental health problem abusing substances on our streets isn't just rare in combination, it's rare for homeless people to have even one of those characteristics.
While three quarters of homeless people in Ventura County lack shelter, only about a quarter are new to Ventura County, have mental health or substance abuse problems, or are chronically homeless.
For the subset who do suffer from substance abuse or mental health problems, the stability of shelter is all the more important. We have the resources to help these people get the care they need, but providing that help is difficult when they sleep on a different street corner every night.
### The shelter shortage
As of January, Ventura County had 464 emergency shelter and transitional housing beds, one for every 3.6 homeless people. A new 110-bed shelter has since opened, but that still leaves over 1,000 beds just to provide one for each homeless person.
Realistically, ending unsheltered homelessness requires more beds than homeless people, since shelters aren't always where our homeless people are (Ventura County is over 2,200 square miles), and each shelter tends to serve a specific subpopulation. For example, the 73 vacant beds were disproportionately reserved for domestic violence victims.
84 percent of beds were utilized as of the point-in-time count, so assuming the same rate moving forward, and considering the newly built and approved shelters, we need 1,300 more shelter beds to end unsheltered homelessness.
> Ventura County needs 1,300 more shelter beds to end unsheltered homelessness.
This has worked elsewhere: for example, New York City's unsheltered homeless rate is a quarter of Ventura County's. How did they do it? [More shelters.](https://medium.com/@josefow/nicks-action-plan-on-street-homelessness-51ddecb1ad76)
### Preventing homelessness through housing affordability
A [2018 Zillow-commissioned report](https://www.zillow.com/research/homelessness-rent-affordability-22247/) found that "homelessness rises faster where rent exceeds a third of income."

While half of U.S. households spend more than 30 percent of their income on rent, 3 in 5 Ventura County households do.

To cut households' risk of homelessness, we need to boost incomes and/or reduce rents. Achieving the former can mean promoting economic development to raise wages, or providing subsidies or cash transfers to low-income renters.
[Evidence](https://www.whitehouse.gov/sites/whitehouse.gov/files/images/Housing_Development_Toolkit%20f.2.pdf) [increasingly](https://lao.ca.gov/Publications/Report/3345#More_Private_Home_Building_Could_Help) [shows](https://research.upjohn.org/up_workingpapers/307/) that achieving the latter — lower rents — requires [building](https://docs.wixstatic.com/ugd/7fc2bf_ee1737c3c9d4468881bf1434814a6f8f.pdf) [more](https://appam.confex.com/appam/2018/webprogram/Paper25811.html) [housing](https://www.sightline.org/2017/09/21/yes-you-can-build-your-way-to-affordable-housing/). This is especially important in Ventura County, which has 20 percent fewer [housing units per adult](https://medium.com/vcyimby/support-for-oxnard-fishermans-wharf-development-26d5204e9f7c) than the nation, and a 30 percent lower [rental vacancy rate](https://data.census.gov/cedsci/table?q=united%20states%20housing&lastDisplayedRow=159&table=DP04&tid=ACSDP1Y2018.DP04&t=Housing&g=0100000US_0400000US06_0500000US06111_1600000US0654652&hidePreview=true&moe=false). With abundant housing of all sorts, low-income households won't have to compete with higher-income households for insufficient housing stock — under threat of homelessness if they don't win — and formerly homeless people will have an easier time re-entering permanent housing.
We need a more humane approach to homelessness. Instead of rounding up tent dwellers as they move from one park to the next, let's build shelters for them to live in. Instead of assuming the worst in each homeless person, let's recognize that they are largely our neighbors who could no longer afford rent. Let's build an affordable, inclusive Ventura County where unsheltered homelessness is a thing of the past.
*If you believe in this approach to homelessness, consider joining *[Ventura County YIMBY](http://vcyimby.org)* and other pro-homeless organizations in the area.*
---
## Warren’s wealth tax would raise less than she claims — even using her economists’ own assumptions
**URL:** https://maxghenis.com/blog/warrens-wealth-tax-would-raise-less-than-she-claim/
**Published:** 2019-11-09
**Description:** Using their cited assumptions around avoidance and economic growth cuts revenue by a quarter
In January, Elizabeth Warren [proposed](https://www.washingtonpost.com/business/2019/01/24/elizabeth-warren-propose-new-wealth-tax-very-rich-americans-economist-says) a two percent (“two cent”) [wealth tax](https://elizabethwarren.com/plans/ultra-millionaire-tax) on families worth $50 million or more, graduated to three percent for those worth $1 billion or more. Last Friday, she [announced](https://elizabethwarren.com/plans/paying-for-m4a) an additional three percent tax on billionaires as part of her Medicare for All plan.
Warren’s plans rely on revenue from this wealth tax. It’s how she would fund student debt cancellation, tuition-free public college, teacher pay hikes, subsidized child care, and now, part of single-payer healthcare.
Her campaign has cited a $2.75 trillion 10-year revenue estimate for her original two- to three-percent tax, and estimated that the three percent Medicare for All surtax on billionaires would raise another $1 trillion. These estimates come from UC Berkeley professors Emmanuel Saez and Gabriel Zucman, who [produced](https://elizabethwarren.com/wp-content/uploads/2019/01/saez-zucman-wealthtax-warren-v5-web.pdf) them in three steps:
- Assemble a consolidated wealth dataset from government surveys and Forbes
- Calculate avoidance and evasion based on academic studies
- Extrapolate over a decade based on government economic growth projections
Saez and Zucman have laudably made the data from Step 1 openly accessible. In Steps 2 and 3, however, they deviated from their cited research methodology in ways that significantly inflated the revenue estimates. I [estimate](https://colab.research.google.com/drive/1ge5a-0yZmC7NhYjc517-koQvovFSEEnS) that addressing these issues would reduce the revenue by a quarter, nearly eliminating the Medicare for All revenue.
### Discrepancies among Saez and Zucman’s own estimates
When Warren announced her first wealth tax, Saez and Zucman published [wealthtaxsimulator.org](http://wealthtaxsimulator.org/) (WTS), where anyone could estimate the revenue from a wealth tax using the same methodology. They also released another calculator last month at [taxjusticenow.org](http://taxjusticenow.org) (TJN), along with their book, *The Triumph of Injustice*. When economists Simon Johnson, Betsey Stevenson, and Mark Zandi [scored](https://assets.ctfassets.net/4ubxbgy9463z/27ao9rfB6MbQgGmaXK4eGc/d06d5a224665324432c6155199afe0bf/Medicare_for_All_Revenue_Letter___Appendix.pdf) Warren’s Medicare for All funding proposal, they used TJN for the wealth tax.
These three sources — the Warren campaign paper and the two calculators — produce different estimates. In the Warren campaign paper, this can be partly attributed to intermediate rounding, though the two calculators also differ when entering the same tax parameters, indicating differences in the underlying data.
One large difference is that TJN shows Warren’s new three-percent billionaire surtax to raise $1.08 trillion, 27 percent more than the WTS estimate; TJN’s estimates 10 percent more total revenue than WTS for the base plus Medicare for All taxes. Since only the WTS data was made available to researchers, I used that for my analysis.
### Does the tax rate affect avoidance?
Saez and Zucman estimate that people will understate their wealth holdings by 15 percent, based on a survey of four studies on avoidance and evasion:
> Our 15% tax avoidance/evasion response to a 2% wealth tax is based on the average across these four studies (2%*(.5+.5+2.5+28.5)/4=16%).
These studies did not estimate a single avoidance rate, but rather an *elasticity* of avoidance with respect to the tax rate. The average elasticity is 8, which they multiply by two percent to arrive at 16 percent. In reality, the formula to translate elasticity to avoidance is [more complicated](https://twitter.com/wwwojtekk/status/1190768770043789312) than multiplication; coincidentally, an elasticity of 8 with a two percent tax translates to the 15 percent avoidance they rounded to.
Warren’s wealth tax was never a flat two percent rate. Assuming that billionaires would expend the same effort to avoid a three — now six — percent tax as they would to a two percent tax is neither sensible, nor supported by the evidence cited by Saez and Zucman. Based on their elasticity, a three percent tax would cause 21 percent avoidance, and a six percent tax would cause 38 percent avoidance.

Warren’s total base + Medicare for All wealth tax raises 19 percent less revenue with these higher avoidance rates. The Medicare for All component raises 65 percent less revenue, falling to $300 billion.

### Ten-year projections do not match government agencies
Saez and Zucman also inflate their estimate when extrapolating one-year estimates over a decade.
In their original Warren campaign paper, they say:
> To project tax revenues over a 10-year horizon, we assume that nominal taxable wealth would grow at the same pace as the economy, at 5.5% per year as in standard projections of the Congressional Budget Office or the Joint Committee on Taxation.
This is incorrect. In its [latest budget outlook](https://www.cbo.gov/system/files/2019-08/55551-CBO-outlook-update_0.pdf), the CBO projected a 3.8 percent average annual rate of nominal GDP increase from 2019 to 2029, and these projections are [used](https://www.jct.gov/publications.html?func=startdown&id=5163) by the Joint Committee on Taxation.
Using the CBO’s 3.8 percent growth rate instead reduces the ten-year projection by an additional 8 percent.

Warren’s final proposal of a two percent tax on wealth above $50 million and six percent tax on wealth above $1 billion would now raise $2.6 trillion. Adding in the unexplained 10 percent boost from Saez and Zucman’s new calculator brings that to $2.9 trillion.
That’s $850 billion missing — a quarter of the $3.75 trillion projection — even before considering how Saez and Zucman may be [overestimating wealth](http://ericzwick.com/wealth/wealth.pdf) at the top of the distribution, how avoidance may be [far higher](https://www.washingtonpost.com/opinions/2019/06/28/be-very-skeptical-about-how-much-revenue-elizabeth-warrens-wealth-tax-could-generate/), how it would [interact](https://taxfoundation.org/senator-warren-medicare-for-all-tax-revenue/) with other taxes to shrink total revenue, and how the [Constitution may prevent](https://www.nytimes.com/2019/11/07/opinion/wealth-tax-constitution.html) wealth taxes from raising a single dollar.
Without twisting the avoidance research and manufacturing growth projections, the $3.75 trillion Warren needs for her range of policy proposals doesn’t exist.
**Addendum (2019–12–06):** Bernie Sanders has also proposed a [wealth tax](https://berniesanders.com/issues/tax-extreme-wealth/), which has more brackets that reach higher rates than Warren’s:
> It would start with a 1 percent tax on net worth above $32 million for a married couple. That means a married couple with $32.5 million would pay a wealth tax of just $5,000.
> The tax rate would increase to 2 percent on net worth from $50 to $250 million, 3 percent from $250 to $500 million, 4 percent from $500 million to $1 billion, 5 percent from $1 to $2.5 billion, 6 percent from $2.5 to $5 billion, 7 percent from $5 to $10 billion, and 8 percent on wealth over $10 billion. These brackets are halved for singles.
Saez and Zucman estimated that Sanders’s wealth tax would raise [$4.35 trillion](https://eml.berkeley.edu/~saez/saez-zucman-wealthtax-sanders-online.pdf) over a decade, using the same assumptions as in scoring the Warren plan. This is a bit higher than the $4.18 trillion projected from their [wealthtaxsimulator.org](http://wealthtaxsimulator.org) tool.
Proper avoidance modeling has a larger effect on the Sanders plan than on the Warren plan, reducing its revenue by 22 percent. After applying the CBO’s growth projection, **revenue drops by a full 28 percent.**

---
## Andrew Yang's troubling Tucker Carlson interview
**URL:** https://maxghenis.com/blog/andrew-yangs-troubling-tucker-carlson-interview/
**Published:** 2019-03-09
**Description:** The Democratic presidential candidate fed into Tucker's MAGA narrative
Last Friday, Democratic presidential candidate Andrew Yang joined Fox News’s *Tucker Carlson Tonight* for a [5-minute interview](https://youtu.be/bMVCkhyg5kA). Consistent with his [book](https://www.amazon.com/War-Normal-People-Disappearing-Universal/dp/0316414247) and [platform](https://www.yang2020.com/policies), Yang warned of a coming wave of job losses due to automation, citing the fall in Midwest manufacturing employment and the 3.5 million truck drivers at risk of replacement by autonomous vehicle technology.
Yang normally uses this premise to motivate his policy prescriptions, especially the [Freedom Dividend](https://www.yang2020.com/policies/the-freedom-dividend/) — a [universal basic income](https://en.wikipedia.org/wiki/Basic_income) funded by a [value-added tax](https://en.wikipedia.org/wiki/Value-added_tax) — central to his campaign. I support (and [have](https://medium.com/basic-income/if-we-can-afford-our-current-welfare-system-we-can-afford-basic-income-9ae9b5f186af) [written on](https://medium.com/@MaxGhenis/the-case-for-a-person-centric-basic-income-plan-55e90010fc9e)) universal basic income, so I appreciate him bringing that issue into the national discourse. But his interview with Carlson showed the dangers of the tenuously-grounded populist frame he’s associating with that policy.
### A recurring misleading claim on labor force participation rate
Carlson introduced Yang by warning of the dangers of automation and technology. After describing self-driving trucks, to back up his assertion that technology is already displacing workers, he [said](https://youtu.be/bMVCkhyg5kA?t=206):
> Labor force participation rate in the United States is 63.2 percent, the same level as Ecuador and Costa Rica.
This echoes his statements on [Twitter](https://twitter.com/AndrewYangVFA/status/1073214206823071744) and the [Joe Rogan Experience](https://youtu.be/cTsEzmFamZ8?t=971), in which he compared the US to El Salvador and Dominican Republic. Yang is referring to the civilian labor force participation rate (LFPR), [defined](https://www.bls.gov/cps/cps_htgm.htm#definitions) by the Bureau of Labor Statistics (BLS) as the share of people age 16 or older who are working or actively seeking work. It’s hovered around 63 percent since 2013, down from 67 percent around 2000, and the most recent data is from January 2019 at 63.2 percent.
Data from the World Bank differs slightly from BLS (age 15+ instead of 16+, and as of 2018) but shows that the US LFPR is about in the middle of the pack (61.6 percent, higher than 46 percent of countries). While it’s similar to El Salvador, it’s substantially lower than Ecuador and Dominican Republic, and somewhat higher than Costa Rica. The United Kingdom and South Korea are both more similar to the United States than most of the countries Yang lists.
The bigger problem is that LFPR is largely a function of demographics, especially the number of retirees. The higher life expectancy in the US combined with Social Security and Medicare means that more people live long enough to spend their later years out of the labor force. Indeed, Americans can expect to live considerably longer than residents of Ecuador, El Salvador and Dominican Republic.

A straightforward way to correct for this bias is to limit the population to people in more typical working ages. For example, the US BLS defines [“prime-age” LFPR](https://fred.stlouisfed.org/series/LNS11300060) as LFPR among people aged 25–54, to account for young people being in school and early retirement. This has been rising since 2015, and is currently 82.6 percent, within 2 percentage points of its all-time high in 1997.

25–54 LFPR isn’t available across countries, but limiting to age 15–64 puts the US alongside France and Hong Kong, and higher than all four of Yang’s comparison countries, even though we have [far more young people in college](https://data.worldbank.org/indicator/SE.TER.ENRR?locations=US-CR-DO-EC-SV&view=chart). Unlike 15+ LFPR, 15–64 LFPR also significantly correlates to both GDP per capita and life expectancy.

### Life expectancy and metrics other than GDP
With Carlson, Yang [continued](https://youtu.be/bMVCkhyg5kA?t=234):
> We have a series of bad numbers and I referred to GDP as one certainly, a headline unemployment rate is completely misleading and one of my mandates as president is I’m going to update the numbers [so] they actually make sense.
Yang expounds on this in his [Human-Centered Capitalism](https://www.yang2020.com/policies/human-capitalism/) policy page — one of his top three objectives, along with the Freedom Dividend and Medicare for All— saying that as President he would:
> Change the way we measure the economy, from GDP and the stock market to…new measurements like Median Income and Standard of Living, Health-adjusted Life Expectancy, Mental Health, Childhood Success Rates, Social and Economic Mobility, Absence of Substance Abuse, and other measurements.
It’s not clear what he means by “update the numbers,” given the metrics he describes are either vague (what’s Childhood Success Rate?) or already [reported by government agencies](https://fred.stlouisfed.org/graph/?g=nah2), [studied by economists](https://www.economy.com/dismal/analysis/datapoints/296127/There-Is-No-US-Wage-Growth-Mystery/), and [covered by the media](https://www.washingtonpost.com/national/health-science/us-life-expectancy-declines-again-a-dismal-trend-not-seen-since-world-war-i/2018/11/28/ae58bc8c-f28c-11e8-bc79-68604ed88993_story.html).
They also generally tell similar stories: unemployment-related indicators move together (and have been stable over the past few years), and absolute indicators like income (adjusted for inflation to capture living standards) and life expectancy have improved over the long run and are now at all-time highs. Illicit substance use is generally [falling](https://www.drugabuse.gov/publications/drugfacts/nationwide-trends) too, with the exception of opioids and (non-problematically) cannabis.


Especially notable here is the longer life expectancy, given Yang [also said](https://youtu.be/bMVCkhyg5kA?t=250) in the interview that “our life expectancy has declined for the last three years, first time in a hundred years.” This decline should raise alarm about the need for public policy addressing the [opioids and suicides](https://www.theatlantic.com/health/archive/2018/11/us-life-expectancy-keeps-falling/576664/) explaining it, and Yang has noted this elsewhere. But we shouldn’t miss the forest through the trees: life expectancy fell by two months from 2014 to 2016, while it grew by that much *every year* from 1990 to 2014.
Could greater funding for statistical agencies generate more valuable insights for policymakers to improve American lives? Of course. But there’s no need to throw the baby out with the bathwater: the currently available statistics tell us that Americans are richer and healthier in recent years than they’ve ever been.
### Feeding into the Trump-Carlson narrative
Absent policy prescriptions, Yang’s vision of America doesn’t look so different from the one Donald Trump described as a presidential candidate in 2016. Yang calls the unemployment rate “completely misleading;” Trump called it a [“hoax.”](http://fortune.com/2016/08/08/donald-trump-hoax/) Yang cites the slight anomalous fall in life expectancy despite long-run improvement; Trump [cited](https://www.politifact.com/truth-o-meter/statements/2016/jun/09/donald-trump/donald-trump-said-crime-rising-its-not-and-hasnt-b/) a slight anomalous increase in crime despite a long-run fall. Yang warns of the labor devastation from automation; Trump [warned](https://www.apnews.com/09215cf7f37f4c6ea05f92f8c83e6125) of immigrants taking jobs.
Certainly, Yang’s arguments hold more water than Trump’s: he hasn’t claimed the “real” unemployment rate was as high as [42 percent](https://www.politifact.com/truth-o-meter/statements/2015/sep/30/donald-trump/donald-trump-says-unemployment-rate-may-be-42-perc/); the trend in life expectancy has more real basis than the crime blip; and automation [really has](http://www.pewresearch.org/fact-tank/2017/07/25/most-americans-unaware-that-as-u-s-manufacturing-jobs-have-disappeared-output-has-grown/) produced more clear job losses than immigration. And while this interview skipped policy implications, Yang’s proposal for universal basic income is productive and future-oriented, not hateful and backwards-looking like Trump’s border wall.

And yet, if he’s being trotted on Fox News to back up their host’s warnings of “the dangers of big tech” and “why [we should] be worried about automation,” and if that host describes him as “one of the only people [he’s] ever met who’s honest about the effects of deindustrialization,” it’s important that he’s describing those effects truthfully. Given Carlson’s [racist fearmongering](https://www.realclearpolitics.com/video/2018/11/01/tucker_carlson_media_push_narrative_that_migrant_caravan_is_not_a_threat_model_future_americans.html) of Latin American immigration, it’s especially important that comparisons of labor market indicators between the United States and Latin American countries are valid.

By these objectives, Yang’s interview did more to bolster the doom-and-gloom narrative that Carlson feeds his audience — which benefits populist politicians like Donald Trump — than improve the standing of his ideas. Yang can avoid falling into this trap by making sure to pair his description of the country with his ideas for improving it. Better still would be re-examining his premise of the country’s perils: his policies like [UBI](https://www.vox.com/policy-and-politics/2017/7/17/15364546/universal-basic-income-review-stern-murray-automation) don’t need it, and the evidence contradicts it.
---
## Quantile regression, from linear models to trees to deep learning
**URL:** https://maxghenis.com/blog/quantile-regression-from-linear-models-to-trees-to/
**Published:** 2018-10-16
**Description:** Suppose a real estate analyst wants to predict home prices from factors like home age and distance from job centers.
But what if they want to predict not just a single estimate, but also the likely range? This is called the *prediction interval*, and the general method for producing them is known as *quantile regression*. In this post I’ll describe how this problem is formalized; how to implement it in six linear, tree-based, and deep learning methods (in Python — [here’s the Jupyter notebook](https://colab.research.google.com/drive/1nXOlrmVHqCHiixqiMF6H8LSciz583_W2)); and how they perform against real-world datasets.
### Quantile regression minimizes quantile loss
Just as regressions minimize the squared-error loss function to predict a single point estimate, quantile regressions minimize the *quantile loss *in predicting a certain quantile. The most popular quantile is the median, or the 50th percentile, and in this case the quantile loss is simply the sum of *absolute* errors. Other quantiles could give endpoints of a prediction interval; for example a middle-80-percent range is defined by the 10th and 90th percentiles. The quantile loss differs depending on the evaluated quantile, such that more negative errors are penalized more for higher quantiles and more positive errors are penalized more for lower quantiles.
Before digging into the formula, suppose we've made a prediction for a single point with a true value of zero, and our predictions range from -1 to +1; that is, our errors also range from -1 to +1. This graph shows how the quantile loss varies with the error, depending on the quantile.
Let's look at each line separately:
- The medium blue line shows the median, which is symmetric around zero, where all losses equal zero because the prediction was perfect. Looks good so far: the median aims to bisect the set of predictions, so we want to weigh underestimates equally to overestimates. As we’ll see soon, the quantile loss around the median is half the absolute deviation, so 0.5 at both -1 and +1, and 0 at 0.
- The light blue line shows the 10th percentile, which assigns a lower loss to negative errors and a higher loss to positive errors. The 10th percentile means we think there’s a 10 percent chance that the true value is below that predicted value, so it makes sense to assign less of a loss to underestimates than to overestimates.
- The dark blue line shows the 90th percentile, which is the reverse pattern from the 10th percentile.
We can also look at this by quantile for under- and over-estimated predictions. The higher the quantile, the more the quantile loss function penalizes underestimates and the less it penalizes overestimates.

Given this intuition, here’s the quantile loss formula ([source](https://towardsdatascience.com/deep-quantile-regression-c85481548b5a)):

And in Python code, where we can replace the branched logic with a `maximum` statement:
```
def quantile_loss(q, y, f): # q: Quantile to be evaluated, e.g., 0.5 for median. # y: True value. # f: Fitted (predicted) value. e = y - f return np.maximum(q * e, (q - 1) * e)
```
Next we’ll look at the six methods — OLS, linear quantile regression, random forests, gradient boosting, Keras, and TensorFlow — and see how they work with some real data.
### The data
This analysis will use the [Boston housing dataset](https://www.kaggle.com/c/boston-housing), which contains 506 observations representing towns in the Boston area. It includes 13 features alongside the target, the median value of owner-occupied homes. Quantile regression therefore is predicting the share of towns (not homes) with median home values below a value.
I train the models on 80 percent and test on the remaining 20 percent. For easier visualization, the first set of models uses a single feature: `AGE`, the proportion of owner-occupied units built prior to 1940. As we might expect, towns with older homes have lower home values, though the relationship is noisy.

For each method, we’ll predict the 10th, 30th, 50th, 70th, and 90th percentiles on the test set.
### Ordinary least squares
Although OLS predicts the mean rather than the median, we can still calculate prediction intervals from it based on standard errors and the inverse normal CDF:
```
def ols_quantile(m, X, q): # m: OLS statsmodels model. # X: X matrix. # q: Quantile. mean_pred = m.predict(X) se = np.sqrt(m.scale) return mean_pred + norm.ppf(q) * se
```
This baseline approach produces linear and parallel quantiles centered around the mean (which is predicted as the median). A well-tuned model will show about 80 percent of dots in between the top and bottom lines. Note the dots differ from the first scatter plot, as here we’re showing the test set to evaluate out-of-sample predictions.

### Linear quantile regression
Linear models extend beyond the mean to the median and other quantiles. Linear quantile regression predicts a given quantile, relaxing OLS’s parallel trend assumption while still imposing linearity (under the hood, it’s minimizing quantile loss). This is straightforward with `statsmodels` :
```
sm.QuantReg(train_labels, X_train).fit(q=q).predict(X_test)# Provide q.
```

### Random forests
Our first departure from linear models is [random forests](https://en.wikipedia.org/wiki/Random_forest), a collection of trees. While this model doesn’t explicitly predict quantiles, we can treat each tree as a possible value, and calculate quantiles using its [empirical CDF](https://en.wikipedia.org/wiki/Empirical_distribution_function) ([Ando Saabas has written more on this](https://blog.datadive.net/prediction-intervals-for-random-forests/)):
```
def rf_quantile(m, X, q): # m: sklearn random forests model. # X: X matrix. # q: Quantile. rf_preds = [] for estimator in m.estimators_: rf_preds.append(estimator.predict(X)) # One row per record. rf_preds = np.array(rf_preds).transpose() return np.percentile(rf_preds, q * 100, axis=1)
```
It goes a bit crazy in this case, suggesting overfitting. Since random forests are more commonly used for high-dimensional datasets, we’ll return to them after adding more features to the model.

### Gradient boosting
Another tree-based method is [gradient boosting](https://en.wikipedia.org/wiki/Gradient_boosting), `scikit-learn`’s [implementation of which](http://scikit-learn.org/stable/modules/generated/sklearn.ensemble.GradientBoostingRegressor.html) supports explicit quantile prediction:
```
ensemble.GradientBoostingRegressor(loss='quantile', alpha=q)
```
While not as jumpy as the random forests, it doesn’t look to do great on the one-feature model either.

### Keras (deep learning)
[Keras](https://keras.io/) is a user-friendly wrapper for neural network toolkits including [TensorFlow](http://tensorflow.org). We can use deep neural networks to predict quantiles by passing the quantile loss function. The code is somewhat involved, so check out the [Jupyter notebook](https://colab.research.google.com/drive/1nXOlrmVHqCHiixqiMF6H8LSciz583_W2#scrollTo=g7s7Grj-A-Sf) or [read more from](https://towardsdatascience.com/deep-quantile-regression-c85481548b5a) [Sachin Abeywardana](https://medium.com/u/4d883eeaf62d) to see how it works.
Underlying most deep nets are linear models with kinks (called *rectified linear units*, or *ReLUs*), which we can see here visually: Keras predicts more bunching up of home values for towns with about 70 percent built before 1940, while fanning out more at the very low and very high ends of age. This seems to be a good prediction based on fit of the test data.

### TensorFlow
One disadvantage of Keras is that each quantile must be trained separately. To leverage patterns common to the quantiles, we have to go to TensorFlow itself. See the [Jupyter notebook](https://colab.research.google.com/drive/1nXOlrmVHqCHiixqiMF6H8LSciz583_W2#scrollTo=g7s7Grj-A-Sf) and [Jacob Zweig](https://medium.com/u/4459aa7bce57)’s [article](https://towardsdatascience.com/deep-quantile-regression-in-tensorflow-1dbc792fe597) to learn more about this.
We can see this co-learning across the quantiles in its predictions, where the model learns a common kink rather than separate ones for each quantile. This looks to be a good Occam-inspired choice.

### Which did best?
Eyeballing suggests that deep learning did well, linear models did OK, and tree-based methods did poorly, but can we quantify which is best? Yes we can, using quantile loss over the test set.
Recall that the quantile loss differs depending on the quantile. Since we calculated five quantiles, we have five quantile losses for each observation in the test set. Averaging over all quantile-observations confirms the visual intuition: random forests did worst, while TensorFlow did best.

We can also break this out by quantile, revealing that tree-based methods did especially poorly at the 90th percentile, while deep learning did best at the lower quantiles.

### Larger datasets give more opportunity to improve over OLS
So random forests was awful for this one-feature dataset, but that’s not what they’re made for. What happens if we add the other 12 features to the Boston housing model?

The tree-based methods made a comeback, and while the OLS improved, the gap between OLS and other non-tree methods grew.
Real-world problems often extend beyond predicting means. Maybe an app developer is interested not only in users’ expected usage, but also their probability of becoming super-users. Or an auto insurance company wants to know the chance of a driver’s high-value claim at different thresholds. Economists might want to stochastically impute information from one dataset onto another, picking from the CDF to ensure proper variation (an example I’ll explore in a follow-up post).
Quantile regression is valuable for each of these use cases, and machine learning tools can often outperform linear models, especially the easy-to-use tree-based methods. Try it out on your own data and let me know how it goes!
---
## We should replace the Child Tax Credit with a universal child benefit
**URL:** https://maxghenis.com/blog/we-should-replace-the-child-tax-credit-with-a-univ/
**Published:** 2018-04-22
**Description:** The Child Tax Credit (CTC) is one of the largest assistance programs for children in the US.
There’s an easy fix: replace the CTC with a universal child benefit, a payment given monthly to parents for each of their children. I modeled this in a [new paper](https://drive.google.com/open?id=1Av1-CTBZpmDGkiG_hyoXqCi4-S1DQUCv) using the open-source [Tax-Calculator](http://pslmodels.github.io/Tax-Calculator/) software,³ showing that such a reform would be highly progressive.
A revenue-neutral plan would give $1,500 per child per year,⁴ increasing by 7 percent the after-tax incomes of the bottom quartile of households with children, at a small expense to higher earning households. Or for a cost of $42 billion — a 36 percent increase to the CTC budgetary impact — we could match the CTC’s maximum of $2,000 per child per year (currently unavailable to many families), boosting incomes of the bottom quartile by 13 percent while either benefiting or not affecting other groups. Both options reduce poverty and inequality, and provide parents with a steady, dependable income to care for their children, instead of an uncertain amount every April.
### Who currently benefits from the Child Tax Credit?
For each child under age 17, the CTC can reduce tax bills by up to $2,000. Among [other restrictions](https://www.thebalance.com/child-tax-credit-3193009), the core credit phases out for very high incomes, and does not reduce tax liability below zero, effectively excluding filers with very low incomes. However, it has a refund mechanism called the Additional Child Tax Credit which can reduce liability below zero. The rules and formulas around the Additional Child Tax Credit [are complicated](https://www.hrblock.com/tax-center/filing/credits/child-tax-credit/) and [changed in 2018](https://www.hrblock.com/tax-center/irs/tax-reform/new-child-tax-credit/), but it results in a credit that’s considerably less generous than the nonrefundable component higher earners receive.
The December 2017 tax bill changed the CTC in a few ways, most importantly: (a) doubling the maximum amount to $2,000, and (b) raising the income threshold before it phases out. This increased the program’s annual budgetary impact by 130 percent, from $52 billion to $120 billion.
Given the complexities of the two credits, distribution of actual outcomes are useful for understanding the credit’s benefit. This chart shows the average credit value per CTC-eligible child (essentially being age 16 or younger) based on the household’s percentile of after-tax after-transfer income (that is, including programs like food stamps; called “income” from here on) without the credit. The analysis is limited to households with at least one child eligible for the CTC, and “child” is defined as children eligible for the CTC.
The average credit per child exceeds $1,500 for most percentiles outside the bottom third, up until the top 1 to 2 percent — those with pre-tax incomes above $270,000 or so — who are ineligible. But for the very poor, it provides little benefit. Further, the Republican tax bill disproportionately helped upper-middle income families, without addressing the disparity for very low earners.
### What if it were shared with every child?
The simplest way to avoid the regressiveness of partially-refundable tax credits is to transform them into universal transfers. If we were to share the CTC’s current $120 billion budgetary impact with our 80 million children, each would get $1,500.⁵ This would raise by 7 percent the after-tax after-transfer income of the bottom quartile of tax units with children, while reducing it by under half a percent for the upper 75 percent. Infusing $42 billion to make it $2,000 per child would raise the bottom quartile’s incomes by 11 percent, and also raise the upper 75 percent by 0.6 percent. While the revenue-neutral plan has a relatively constant effect across the upper three quartiles, the $2,000 benefit would be progressive at each quartile.

Even this understates the benefits to the poorest families. Zooming in on the bottom quartile shows that the bottom 5 percent increase their incomes by 120 percent in the revenue-neutral plan and 160 percent in the $2,000 plan. This group’s current average income is $3.30 per person per day, including transfer programs like food stamps and Medicaid (with the caveat that measuring extremes of the income distribution [has its challenges](http://nbviewer.jupyter.org/github/MaxGhenis/taxcalc-notebooks/blob/master/exploratory/tax_units_with_lowest_aftertax_income.ipynb)).

While the Republican tax bill extended the CTC to most high-income households, it still phases out for filers with incomes above $200,000, or $400,000 if filing jointly. That means that universal child benefits would benefit this top roughly 2 percent of tax filers with children, increasing their incomes by 0.2 percent (0.4 percent in the $2,000 plan). The benefit could be paired with a higher top-bracket marginal rate — which was reduced from 39.6 percent to 37 percent as part of the tax bill — to avoid this result.

Even with the benefit to the top 1 percent, both plans reduce inequality, as measured by the popular [Gini coefficient](https://en.wikipedia.org/wiki/Gini_coefficient). This measure lies between 0 (perfect equality) and 1 (perfect inequality), and is currently 0.451 among US households with children. The revenue-neutral plan would reduce it to 0.446, while the $2,000 plan would reduce it to 0.441.
### A child benefit would put us on par with other countries
[Many countries](https://en.wikipedia.org/wiki/Child_benefit) have child benefits of some sort. While most are means-tested or change depending on the number of children, the US would not be alone in enacting a flat benefit for all children. [Ireland](http://www.citizensinformation.ie/en/social_welfare/social_welfare_payments/social_welfare_payments_to_families_and_children/child_benefit.html) and [Sweden](https://www.forsakringskassan.se/privatpers/foralder/nar_barnet_ar_fott/barnbidrag/!ut/p/z0/fYzLDoIwEAC_puctqNwJKD4OxpAo9tIUKaRKtmW7IX6-6Ad4m0kmAwoaUGhmNxh2Hs24-F1ler0vi6Qq5Om8qUuZZ4dLcq236S5dQW0RjqD-R8vFPadJ5aAeHtm-GZofIEerycbgMbrZChnIzYaDpShk78mMnSUh0ZBuDaFlvVDvmYX8eus6MgOEV3X7AGdh2ls!/) are two such countries, giving families the equivalent of $2,070 and $1,540 per child under 16 per year, respectively.⁶ Canada’s child benefit was also universal (though taxable) until it became means-tested in 2016.
Evidence from around the world shows that these child benefits work. As [Vox’s Dylan Matthews summarizes](https://www.vox.com/policy-and-politics/2017/4/27/15388696/child-benefit-universal-cash-tax-credit-allowance), child benefits and other cash benefits to families not only reduce child poverty, but lead families to spend more time with children, increase rates of prenatal care, improve children’s physical and mental health, and ultimately boost kids’ learning outcomes and earnings later in life.
Flat amounts can also avoid macro-level disadvantages of means-testing: [administrative overhead, abuse by politicians, tensions between recipients and nonrecipients](https://www.brookings.edu/blog/future-development/2017/05/31/rethinking-the-universalism-versus-targeting-debate), [welfare cliffs](https://www.budget.senate.gov/newsroom/budget-background/the-welfare-cliff-how-the-benefit-scale-discourages-work) (implicit marginal tax rates exceeding 100 percent), and difficulty in distributing more frequently than annually.
These virtues have motivated other scholars to advocate child benefit plans, including [Hammond and Orr (2016)](https://niskanencenter.org/wp-content/uploads/2016/10/UniversalChildBenefit_final.pdf), [Shaefer et al (2018)](https://www.rsfjournal.org/doi/abs/10.7758/RSF.2018.4.2.02), and [Bitler, Hines, and Page (2018)](https://www.rsfjournal.org/doi/abs/10.7758/RSF.2018.4.2.03).
### The beginning of the end of child poverty
A child benefit in the US would significantly cut child poverty. There’s no single measure of child poverty, but these four are commonly used:
- Extreme poverty is defined internationally by the World Bank as consumption⁷ below $1.90 per day per person, in 2011 dollars adjusted for purchasing power. That’s about $780 in 2018 dollars per person per year in the US.Estimate: Under 0.1 percent (Brookings Institution analysis of 2011 data). No child-specific data available.
- The US poverty threshold is defined by the US Census Bureau. It is based on the number and age of household members, and only includes earnings and cash transfers (not taxes, tax credits, or in-kind transfers), so will be unaffected by changes to the CTC. The 2017 threshold for a household with one person under age 65 and one child was $16,895.Estimate: 21 percent of children (National Center for Children on Poverty report, no year specified).
- The US poverty guideline (often called the federal poverty line, or FPL) is defined by the department of Health and Human Services to determine eligibility for programs like Medicaid. Like the poverty threshold, it is based on pre-tax cash alone, but it does not depend on age. For the contiguous states, HHS defines the FPL as $12,140 for a household with one person, plus $4,320 for each additional person.Estimate: None available.
- The US Supplemental Poverty Measure (SPM) is defined by the Census Bureau. It includes taxes and transfers — cash and in-kind — and is adjusted based on local cost of living.Estimate: 15.6 percent of children (Center on Budget and Policy Priorities analysis of 2016 data).
None of these is ideal for estimating the poverty reduction resulting from replacing the CTC with a child benefit. Extreme poverty measures consumption, not income; neither the poverty threshold nor guideline include tax credits like the CTC; and Tax-Calculator data lacks geographic granularity needed for the SPM.
Instead, I use these measures:
- Extreme poverty by income, at the same $780 defined by the World Bank, using income rather than consumption. Estimates of the poorest Americans are unreliable, but this provides a rough relative picture.
- Federal poverty line (FPL) including taxes and transfers, using levels defined for the contiguous states.
- $10,000 per person per year, a simple threshold.
By all three measures, both the revenue-neutral and $2,000 child benefit significantly reduce the share of children living in impoverished households. The impact is most significant for the extreme poverty measure, which falls from 1.8 percent to 0.4 percent and 0.2 percent for the two respective plans. Using the FPL, the child poverty rate falls by 15 percent and 22 percent, respectively, and it falls by 6 percent and 13 percent using the $10,000 threshold. The existing CTC also cuts child poverty using the FPL and $10,000 thresholds, but has virtually no effect on extreme child poverty.

Throughout its [twenty-year history](https://fas.org/sgp/crs/misc/R41873.pdf), the Child Tax Credit has helped middle-income families care for the future of our country. The nominal value of the credit has grown over time, and it has slowly become more available to low-income families through partial refundability.
It should culminate in a universal child benefit, completing its path toward inclusivity and fulfilling its mission to help the poorest families. Doing so would not only alleviate poverty and stem inequality, it would provide a common baseline footing for every child in America. A universal child benefit in our policy arsenal would provide a tool that could be expanded over time, and which could ultimately eliminate the scourge of child poverty.
*[1] “Households” in this analysis refer to tax units, as this is the core unit in the Tax-Calculator software. Households can contain multiple tax units (e.g. in households where spouses file separately), and the US has 20 percent more tax units than households. Tax-Calculator also imputes tax units that might not have filed using CPS data.*
*[2] Children here are defined as children eligible to receive the Child Tax Credit, the primary criterion of which is being age 16 or younger.*
*[3] *[This notebook](https://nbviewer.jupyter.org/github/MaxGhenis/taxcalc-notebooks/blob/ffa907959f4c893ebe8fcd27dc92748c9a5406b9/ctc/ctc_ubi.ipynb)* contains my code and calculations using Tax-Calculator 0.19.0, CPS data, and projections for 2018.*
*[4] Rounded from an exactly revenue-neutral amount of $1,470; this rounding would cost $2.4 billion.*
*[5] Tax-Calculator’s CPS data estimates a $117 billion budgetary impact and 80 million CTC-eligible children. Other estimates are ~$114–120 billion, and the Census estimates 70 million children under age 17. To produce consistent distributional estimates, I use Tax-Calculator’s CPS data, but the full paper includes other estimates.*
*[6] Ireland also gives their child benefit to some 16- and 17-year-olds.*
*[7] Consumption is estimated via surveys. See *[Our World In Data (2017)](https://ourworldindata.org/extreme-poverty)* for more information.*
*Thanks to the Open Source Policy Center’s Cody Kallen, Matt Jensen, Sean Wang, Martin Holmer, Anderson Frailey, and Amy Xu.*
---
## The case for a wonky basic income plan
**URL:** https://maxghenis.com/blog/the-case-for-a-wonky-basic-income-plan/
**Published:** 2017-06-14
**Description:** In March, a former union leader joined forces with a prominent libertarian to debate two Obama White House economists.
Economists have expressed skepticism of this idea elsewhere. Bob Greenstein of the Center on Budget and Policy Priorities [questions](https://www.vox.com/2017/5/12/13954182/case-for-and-against-universal-basic-income-united-states) the compatibility of such a large expenditure with today's political climate. A [2016 survey](http://www.igmchicago.org/surveys/universal-basic-income) found that 70% of economists opposed replacing all transfer programs with a basic income of $13,000 per year for adults age 21 and over. Others like [blogger Tyler Cowen](http://marginalrevolution.com) believe that [work, not money](https://www.bloomberg.com/view/articles/2016-10-27/my-second-thoughts-about-universal-basic-income), is the central problem for the country.
And yet, there's a lot for economists (and us all) to like about basic income. It
- Fills in the gaps of the current safety net, which often fails to reach the poorest and unemployed.
- Eliminates welfare traps, where beneficiaries lose more in benefits than they earn if they choose to work more or accept a raise (i.e., improves work incentives).
- Utilizes cash, which 84% of economists agree increases welfare more than in-kind benefits do.
- Consolidates the portfolio of programs made costly by its complexity (even large ones cost ~10% to administer).
- Simplifies the bureaucracy that's difficult for beneficiaries to navigate.
Their concerns are often not about basic income per se, but about the plans proposed. Political challenges aside, economists' most compelling critique may be that existing plans harm those in most need, especially families with children. Bernstein and Furman drilled this point in the Intelligence² debate, leading them to victory: while 20% of the audience initially opposed the motion that "the universal basic income is the safety net of the future", 61% opposed it afterward.
The good news is: it's entirely possible to design a basic income plan that minimizes the net losses (and therefore political opposition) to beneficiaries and taxpayers, while still respecting the core tenets of basic income — universality, individuality, and focus on cash. Doing so requires a *person-centric* perspective, whereby we consider the current benefits received and taxes paid by all Americans today, and design the plan with the explicit goal of giving as many people as possible at least as much as they currently get.
This would be hard work, and I've not (yet) formed a specific plan. This piece lays out the justification for a new approach to basic income policy design, describes a framework for this design, and proposes negative income tax as a conceptual starting point. Such a policy can win expert support for basic income by emphasizing welfare of individuals.
### The current plans
Stern and Murray were appropriate defenders against Bernstein and Furman, as they have proposed some of the most specific and widely-cited plans for basic income in the US. Prominent basic income advocate Scott Santens also released a plan in June. So it's worth summarizing the [Stern](https://futurism.com/heres-how-we-could-fund-a-ubi-program-in-the-united-states/), [Murray](http://www.fljs.org/files/publications/Murray.pdf), and [Santens](https://medium.com/economicsecproj/how-to-reform-welfare-and-taxes-to-provide-every-american-citizen-with-a-basic-income-bc67d3f4c2b8) plans on a few dimensions:
- Annual amount Stern: $12,000Murray: $10,000, plus $3,000 to purchase catastrophic health insuranceSantens: $13,266 for adults, $4,598 for children
- Eligible age population (all plans are limited to US citizens)Stern: 18–64, though seniors can opt in if Social Security (which remains) pays less than $12,000 per yearMurray: 21+Santens: All (lower amount for children)
- Replaced programs (I consider the Earned Income Tax Credit (EITC) and Child Tax Credit (CTC) "programs" instead of tax changes, to align with the IRS)Stern: "all or some number of the 126 [federal] welfare programs that currently cost $1 trillion a year" excluding healthcare programs like Medicaid and CHIP; also cuts military spendingMurray: "Social Security, Medicare, Medicaid, welfare programmes, social service programmes, agricultural subsidies, and corporate welfare"Santens: "Food and nutrition assistance programs ($108 billion)[,] temporary assistance for needy families ($17 billion)…the earned income credit ($73 billion), [and] the child tax credit ($56 billion)…will not touch health care, child care, or housing"
- Tax changes (there's uncertainty around these estimates)Stern: Phase out most tax expenditures (currently $1.2T/year); possibly add VAT of 5–10%, 1.5% wealth tax on assets over $1M, financial transaction tax, and Peter Barnes' "Sky Trust"Murray: 17% phase-out of basic income for incomes between $30,000 to $60,000 ($5,000 for incomes above $60,000)Santens: End "home ownership tax expenditures ($340 billion), married filing jointly preferential tax treatment ($70 billion), the tax break on pensions ($160 billion), fossil fuel subsidies ($33 billion), and treating capital gains differently than ordinary income ($160 billion);" add a "[carbon tax] starting at $50/ton with annual increases of $15/ton", a 0.34% financial transaction tax, seignorage reform, 10% VAT, and 5% land value tax
### Current proposals reduce benefits for some
By limiting to adults, the Stern and Murray plans both leave most poor single-parent households significantly worse off than they are today, as Daniel Hemel finds in his [well-researched criticism](http://newramblerreview.com/book-reviews/economics/bringing-the-basic-income-back-to-earth) (this also describes the "plan" in the survey question opposed by 70% of economists). These households currently receive thousands of dollars per year from EITC, CTC, SNAP (f.k.a. food stamps), and other programs like housing assistance (if they're in the [lucky 24%](http://www.urban.org/urban-wire/one-four-americas-housing-assistance-lottery) of eligible households to get it).
Santens proposes $4,598 per child, in accord with [federal poverty guidelines](http://familiesusa.org/product/federal-poverty-guidelines) (adding 10% to compensate for a VAT), which vary depending on household size. But does even this amount per child compensate for lost benefits and higher taxes? Not for everyone. Consider two households earning $10,000 per year: one with a single childless person, and the other with a single person and a child. The value of marginal benefits received by the latter household is attributable to the child.
The three largest programs for this sample household total $5,958 in value, or $1,400 more than the basic income:
- $3,002 from EITC ($3,371 vs. $371 for the childless household)
- $1,956 in SNAP benefits ($3,032 vs. $1,076)
- $1,000 from CTC
> "In most of the [basic income] proposals, the losers tend to be households with more children." — Jason Furman
Another example: all three plans are limited to citizens, despite some existing programs being [partially available](https://aspe.hhs.gov/basic-report/overview-immigrants-eligibility-snap-tanf-medicaid-and-chip) to legal permanent residents.
### Plans' various tax increases create more losers
In addition to some households getting less in benefits under the proposed plans, the suite of tax increases required to finance them would hit virtually everyone, including the poor. To name a few items, people could expect to pay more for:
- Gas and electricity (carbon tax)
- Rent and real estate (land value tax and ending the mortgage interest deduction)
- Groceries (value added tax and ending farm subsidies)
This doesn't mean these taxes aren't worthwhile on their own: [each](http://www.igmchicago.org/surveys/tax-reform) [enjoys](http://www.economist.com/blogs/economist-explains/2014/11/economist-explains-0) [broad](http://gregmankiw.blogspot.com/2009/02/news-flash-economists-agree.html) [support](http://gregmankiw.blogspot.com/2009/10/value-added-tax.html) [among](https://www.theguardian.com/environment/climate-consensus-97-per-cent/2016/jan/04/consensus-of-economists-cut-carbon-pollution) economists (I support all).
But these ideas bear no innate relationship to basic income, and their political baggage would weigh down the cause. For example, economists have been trying to kill the mortgage interest deduction for [decades](https://www.nytimes.com/2017/05/09/magazine/how-homeownership-became-the-engine-of-american-inequality.html?_r=0), only to be consistently rebuffed by the real estate lobby. Pursuing basic income need not require battling the real estate, agriculture, and fossil fuel industries.
It's also hard to pin down how much each person pays for these taxes, relative to, say, income tax hikes. That increases the variance of who wins and who loses under these plans, meaning they could hurt even more low-income households.
For reasons I elaborate on shortly, I favor income taxes instead to avoid these issues.
### A person-centric direction
Fundamentally, the three plans start with a number: the amount believed needed to fund a basic income for the included population. These are big numbers — roughly $3T per year for each plan — so then they fill in the gaps with whatever it takes (except higher income taxes). Examination of example households has followed the design.
Let's turn that on its head: instead of starting with the number, we can start with the people, and focus on maximizing the upside of each while minimizing the downside. The units of analysis would include people and households, not just dollars and cents.
This paradigm is harder to work in. It requires deep understanding of household income distributions, safety net program design, and tax codes. It then requires model-building to alter each while minimizing negative impacts. This often involves teams of policy analysts under the direction of established think tanks or policymakers. So it's understandable that Murray, Stern, and Santens haven't gone there yet. While I don't yet have a plan that does, I have some ideas on how to build one.
### How to build a person-centric UBI plan
We can distill policy design to three pieces:
- Non-negotiables. What's the heart of the policy?
- Levers. What parts of the plan are adjustable?
- Assessment. How do we measure the policy's success, subject to (1) and (2)? Alternatively: what are we optimizing for?
For basic income, **non-negotiables **(1) are fairly straightforward: a basic income plan must be universal, unconditional, and 100% cash. Universality may be stratified, such that children, people with disabilities, seniors (who receive Social Security), and/or noncitizens receive different amounts. It may not be stratified by the composition of the household; for example, two adults must receive the same amount whether they are in the same household or not (it is an individual-level program, in contrast with many existing programs).
**Levers **(2) could include several things, starting with those implicit in the three cited plans:
- Amount per adult (the primary lever)
- Different amounts for children and other categories listed above
- Clawback/phase-out rateFor example, Murray suggests a partial clawback, such that it's a hybrid of UBI and a negative income tax, though some would say this violates a non-negotiable aspect of universality by retaining means-testing.
- Programs replaced
- Tax expenditures eliminated
- Additional taxes leviedBeyond new taxes Murray, Stern, and Santens propose, I would add income tax increases as an important lever.
**Assessment **(3) should concern the distribution of impacts to after-tax after-transfer income to individuals or households. This is difficult in the Murray, Stern, and Santens plans, since some of the funding sources aren't easily translatable to this metric, but the one-parent one-child household example suggests that many households with children will see considerable drops in total income.
The most straightforward metric for quantifying this would be something like *average % loss per person*, i.e. ignore the winners and focus on losers*. *So if a two-person household was getting $30k after taxes and transfers (including in-kind benefits like SNAP), and with UBI they get $29k, their contribution to the metric would be -3.3% ($1k / $30k), weighted for two people.
This could be extended in a couple ways. For example, winners could be included, though with less weight since politics is (appropriately) loss-averse. Studies in behavioral economics suggest that ["losses are twice as powerful, psychologically, as gains,"](https://en.wikipedia.org/wiki/Loss_aversion) so losses could receive double the weight of gains. The metric could also consider implicit wages lost due to program compliance time — perhaps at minimum wage — given existing programs require more paperwork and in-person time with bureaucrats than cashing a basic income check does. One could imagine other factors as well, such as utility lost due to restrictive in-kind benefits ([Pareto efficiency](https://en.wikipedia.org/wiki/Pareto_efficiency) and [welfare economics](https://en.wikipedia.org/wiki/Welfare_economics) could be other angles).
Within this framework, the goal becomes: **ensure some level of universal cash transfer by adjusting levels and financing, such that negative impact on households is minimized.**
### Negative income tax as a starting point
The simplest way to achieve this goal is to mimic the current suite of programs, with three important changes:
- Replace in-kind benefits with cash.
- Remove restrictions like work requirements.
- Shift from household-based to individual-based.
Since existing programs are means-tested instead of universal, this would not produce a true basic income. What it would produce is what's called a [negative income tax](https://en.wikipedia.org/wiki/Negative_income_tax), whereby people below a certain income level receive a payment which phases out with income. It turns out that most historical studies on "basic income" have actually been on negative income tax, which [nearly became law](https://www.theatlantic.com/politics/archive/2014/08/why-arent-reformicons-pushing-a-guaranteed-basic-income/375600/) in the 1970s.
Importantly, one could design a negative income tax with *identical outcomes *to a universal basic income, and vice versa, once you set the baseline. The simplest way to turn negative income tax into basic income is to calculate the extra amount individuals get (anyone with income gets more under basic income than negative income tax), and raise their taxes by exactly that amount. After taxes and transfers, nobody gets any more or less. I've [written more about this](https://medium.com/basic-income/if-we-can-afford-our-current-welfare-system-we-can-afford-basic-income-9ae9b5f186af), including reasons to still prefer basic income to negative income tax.
To make this a bit more concrete, consider a household with one adult and one child. As FiveThirtyEight's Andrew Flowers [showed](http://fivethirtyeight.com/features/universal-basic-income/) in the chart below, this household faces "welfare cliffs" due to immediate or rapid phase-outs of existing programs like TANF, Medicaid, and CHIP, as well as the way they interact with each other and with income taxes.

The household gets ~$19k when earning $0, and ~$38k when earning $50k, or an average [implicit marginal tax rate](https://en.wikipedia.org/wiki/Tax_rate#Implicit_marginal_tax_rate) of *100%-($38k- $19k) / ($50k-0) = ~***60%. **If that seems harsh, it is: while income taxes are progressive, the highest marginal tax rates (including assistance) are actually borne by low-income households, due to means testing.
I've added an approximate line of best fit to this chart to show what this would look like as a flat income tax. We can implement this line in one of two ways:
- Negative income tax with a $19k minimum (for the two-person household), 60% phase-out (up to $30k, at which the household neither gets a refund nor pays taxes), and 60% tax rate after $30k (up to $50k earnings)
- Basic income of $19k and 60% tax rate (up to $50k earnings)
**These produce the same result. **The household gets/pays the same amount given their earnings, and the government gets/pays the same in net taxes. The primary difference is that negative income tax comes without the sticker shock: [Wiederspan, Rhodes, Shaefer (2015)](http://www.tandfonline.com/doi/abs/10.1080/10875549.2014.991889) estimate that a negative income tax set at US poverty line and with a 50 percent phase-out rate could be 95% paid for with the combined budgets of EITC, SSI, SNAP, TANF, school meal programs, and housing subsidies ([here's a Vox summary](https://www.vox.com/policy-and-politics/2017/5/30/15712160/basic-income-oecd-aei-replace-welfare-state)).
Here lies the core advantage of funding basic income with income tax increases: you can only make a basic income plan equivalent to negative income tax when funding via income taxes. The increase that may seem significant would not actually be felt by households, and with income taxes representing [47% of federal tax revenues](http://www.cbpp.org/research/policy-basics-where-do-federal-tax-revenues-come-from), it's a known quantity that doesn't create new enemies.
A person-centric analysis will reveal the appropriate minimum and shape of the line to minimize losses for households. This will depend partially on what programs basic income replaces (I tend to agree with Stern and Santens in leaving out healthcare programs like Medicaid and CHIP). Inputs to the analysis would include distributions of household earnings by size, and the inclusion criteria and amounts of each existing program (essentially the above chart for each household type). Turn the resultant negative income tax into a basic income, and we've got a fair, focused, realistic plan.
Idealism attracts people; realism convinces them. Media coverage of pilots and automation have generated massive interest in the *idea *of basic income, but without a better-thought-through plan, that interest is doomed to evaporate among those concerned with the welfare of real people.
It's time to get in the weeds and make a plan that not only lives up to the ideals of basic income, but also maximizes the number of people benefiting from it.
---
## Reddit featured a misleading headline on education secretary nominee Betsy DeVos
**URL:** https://maxghenis.com/blog/reddit-featured-a-misleading-headline-on-education/
**Published:** 2016-11-26
**Description:** “billionaire”
[This headline](https://www.reddit.com/r/atheism/comments/5eobxq/meet_your_new_a_secretary_of_education/) made the top of [Reddit](http://reddit.com)’s [r/all](http://reddit.com/r/all) page on [Thursday](http://web.archive.org/web/20161124230946/https://www.reddit.com/r/all/):
> Meet your new a [sic] Secretary of Education: billionaire, creationist, charter school advocate, she’s anti-gay marriage, daughter of Amway founder, sister of Blacwater [sic] CEO, she is the largest donor to the worst Religious Right hate groups. Making America Great Again!
The post in the [r/atheism](http://reddit.com/r/atheism) subreddit linked to Libby Nelson’s Vox [article](http://www.vox.com/2016/11/23/13735102/betsy-devos-education-donald-trump) titled “Betsy DeVos: Donald Trump’s education secretary pick shows school vouchers are at the top of his agenda,” published Wednesday. The article was informative, but the editorialized headline on Reddit was misleading. Here I examine each claim and describe my attempt to report the issue.
*I neither assert equivalence of social media misinformation between the left and right (right-wing posts are *[at least twice as likely](https://www.buzzfeed.com/craigsilverman/partisan-fb-pages-analysis?utm_term=.kloB4j3PG#.owQMpq130)* to be false), nor endorse this nomination.*
#### "billionaire"
**PROBABLY**
Betsy’s father, Edgar Prince, founded the Prince Corporation, which was sold for $1.35B in 1996, the year after his death. Betsy is one of Edgar’s four children.
Forbes estimates the net worth of Betsy Devos’s father-in-law, 90-year-old Amway co-founder Richard DeVos, at [$5.1 billion](https://en.wikipedia.org/wiki/Amway#Pyramid_scheme_accusations). Betsy’s husband, Dick DeVos, is one of Richard’s four children.
Betsy and Dick DeVos also [founded](https://en.wikipedia.org/wiki/Betsy_DeVos#Business_career) the [Windquest Group](http://windquest.com/) in 1989, a private equity firm which Betsy serves as chairwoman.
Betsy and Dick’s combined joint stakes of the Prince fortune, Amway fortune, and own venture may not exceed a billion, but [Snopes estimates that it does](http://www.snopes.com/betsy-devos-education-secretary/) so I’ve rated it “probably.”
#### “creationist”
**MAYBE**
While Betsy and Dick DeVos are prominent members of the [Christian Reformed Church in North America](https://en.wikipedia.org/wiki/Christian_Reformed_Church_in_North_America), which [believes in creationism](https://www.crcna.org/welcome/beliefs/position-statements/creation-science), they haven’t themselves made public statements on it. Betsy’s [Washington Post profile](https://www.washingtonpost.com/news/acts-of-faith/wp/2016/11/23/betsy-devos-trumps-education-pick-is-a-billionaire-philanthropist-with-deep-ties-to-the-reformed-christian-community/) states “She will not likely be one to focus on curriculum issues like evolution and creationism, which has been a concern in some conservative Christian circles.”
#### “charter school advocate”
**TRUE**
Nelson writes in Vox:
> DeVos has a long history of activism on one issue: school choice — a term that refers to both school vouchers, which allow parents to use taxpayer money to send their children to private school, and charter schools, which are publicly funded but privately run [and are not religious].
#### “anti-gay marriage”
**HUSBAND WAS IN 2008**
According to the nonpartisan [followthemoney.org](http://followthemoney.org), Dick DeVos (Betsy’s husband) [gave $20,000](http://www.followthemoney.org/show-me?m-eid=10246473&m-t-rt=3&d-eid=23175) to a ballot committee supporting [Michigan’s 2004 Prop 04–2](https://en.wikipedia.org/wiki/Michigan_Proposal_04-2), an amendment to the state constitution which banned same-sex marriage. The proposal passed with 59% of the vote. In 2008, he [gave $100,000](http://www.followthemoney.org/entity-details?eid=23175&default=contributor) to [Florida’s Amendment 2](https://en.wikipedia.org/wiki/Florida_Amendment_2_%282008%29) which also defined marriage as between a man and a woman. The proposal passed with 62% of the vote.

In June, Betsy and Dick [gave](http://www.rawstory.com/2016/06/wealthy-devos-family-offers-400000-to-pulse-shooting-victims-and-2-million-to-anti-lgbt-groups/) $400,000 to victims of the shooting at Orlando’s gay Pulse nightclub.
Note: Searching for [DeVos gay marriage](https://www.google.com/search?q=DeVos+gay+marriage) yields several articles claiming that Dick and Betsy have a more extremist opposition to same-sex marriage. Much of these articles focus on the broader DeVos family, where [evidence is more abundant](http://www.rawstory.com/2016/06/wealthy-devos-family-offers-400000-to-pulse-shooting-victims-and-2-million-to-anti-lgbt-groups/). Several also claim that Betsy and Dick DeVos spent $200,000 (not $20,000 as shown in followthemoney.org) on Michigan’s 2004 measure banning gay marriage, and that they led the effort. This claim seems to [originate](https://twitter.com/wirecan/status/802219858171662337) with an [article](http://www.pridesource.com/article.html?article=57829) by LGBT publication [PrideSource](http://www.pridesource.com/), which states “In 2004 Betsy and Dick DeVos led the effort to put the anti-marriage amendment on the ballot and contributed over $200,000 to the campaign to enshine [sic] discrimination into the Michigan constitution.” I’ve [asked](https://twitter.com/MaxGhenis/status/802233253792878592) PrideSource for a primary source on this, given the conflicting data and lack of evidence they led the effort.
#### “daughter of Amway founder”
**DAUGHTER-IN-LAW**
[Amway](https://en.wikipedia.org/wiki/Amway) was founded by Richard DeVos, father of Betsy’s husband Dick DeVos. Amway “is an American company that uses a multi-level marketing model to sell a variety of products, primarily in the health, beauty, and home care markets.” Along with other companies adopting multi-level marketing models, Amway has been criticized for its [similarity to pyramid schemes](https://en.wikipedia.org/wiki/Amway#Pyramid_scheme_accusations).

#### “sister of Blacwater [sic] CEO”
**Former CEO (1997–2009)**
According to Wikipedia, Blackwater, now known as [Academi](https://en.wikipedia.org/wiki/Academi), “is an American private military company, originally founded in 1997 by former Navy SEAL officer [Erik Prince](https://en.wikipedia.org/wiki/Erik_Prince).” Erik is the brother of Betsy DeVos (nee Prince), and after founding Blackwater in 1997 he “served as its CEO until 2009 and later as chairman, until Blackwater Worldwide was sold in 2010 to a group of investors.” Blackwater was a controversial piece of the Iraq war, symbolizing the military-industrial complex after receiving a no-bid contract and later killing civilians in multiple incidents.
To the extent that the headline suggests a conflict of interest, it’s noteworthy that Erik Prince has not been involved in the organization for seven years. He currently heads a private equity firm.
#### “largest donor to the worst Religious Right hate groups”
**A STRETCH**
[MediaMouse](http://mediamousearchive.wordpress.com/) — a progressive blog focused on Grand Rapids area — has [catalogued](https://mediamousearchive.wordpress.com/resources/right/) the network of conservative individuals and organizations in West Michigan. Of organizations profiled on this site, her [page](https://mediamousearchive.wordpress.com/resources/right/people/betsy-devos/) claims she is linked to the pro-life [Compass Academy](https://mediamousearchive.wordpress.com/resources/right/orgs/compass-academy/), laissez-faire climate-change-denial [Acton Institute](https://mediamousearchive.wordpress.com/resources/right/orgs/acton-institute/), and [Education Freedom Fund](https://mediamousearchive.wordpress.com/resources/right/orgs/education-freedom-fund/) which provides scholarships for private schools. None are designated hate groups, and from their descriptions none seem to be anywhere near the most hateful religious right groups out there.
I shared this information with the r/atheism moderators to justify *flairing *(tagging) the post as “Misleading title.” They could also add a “stickied” comment with clarifications, which sits at the top of the comment tree for everyone to see. They declined, saying that
> There are plenty of sources in the thread corroborating every single item in the headline, so, no thank you.
None of the top 100+ comments at the time included any such corroboration, aside from a [link](https://www.reddit.com/r/atheism/comments/5eobxq/meet_your_new_a_secretary_of_education/dae3u1g/) to MediaMouse, and some comments questioned the claims.
The [user](https://www.reddit.com/user/PlanetoftheAtheists) who posted the story has a history of sensationalist titles getting to the front page of top subreddits like r/atheism. And subreddit moderators are incented to have popular posts, as getting featured on r/all often leads to new subscribers.
Some subreddits like [r/news](http://reddit.com/r/news) (a default subreddit, meaning it’s shown on reddit.com for new or logged-out users; r/atheism is not default) have stricter headline requirements, such as using the a line from the article itself. That would have helped here, but even basic fact-checking should be the norm for such prominent sites. Reddit administrators (who have more privileges than subreddit moderators) could take a stand on this too; while r/all is intended to be unfiltered, they could at least check titles for factual inaccuracies.
Reddit isn’t the only place Betsy DeVos has been mischaracterized. A [meme](https://www.facebook.com/seusstastic/photos/a.214339368609095.56033.206299406079758/1271355676240787/?type=3&theater) has been going viral on Facebook claiming she gave $9.5M to Trump’s campaign and $200M to Christian schools and organizations. The former claim is false: neither [Betsy](http://www.followthemoney.org/entity-details?eid=23375) nor [Dick](http://www.followthemoney.org/entity-details?eid=23175) nor [Betsy and Dick as a couple](http://www.followthemoney.org/entity-details?eid=553118) gave to Trump (though they gave to other GOP candidates), and Betsy [did not endorse Trump](http://www.washingtonexaminer.com/betsy-devos-who-never-endorsed-trump-tapped-as-education-secretary/article/2608117). Her brother, Erik Prince, also [did not donate](http://www.followthemoney.org/search-results/SearchForm?Search=erik+prince&action_results=Go) to the Trump campaign. Snopes [looked into](http://www.snopes.com/betsy-devos-education-secretary/) the latter claim (with the rest of the meme) and rated it true based on their total giving, but it’s unclear whether they actually donated a full $200M to Christian causes.
Betsy DeVos is an understandably controversial nominee. Her lack of government experience creates uncertainty about her management capacity. Charter schools and vouchers are lightning rods, especially for teachers unions, which include the [largest union in the US](https://en.wikipedia.org/wiki/National_Education_Association). [Poor charter school results in Detroit](http://www.nytimes.com/2016/11/25/opinion/betsy-devos-and-the-wrong-way-to-fix-schools.html?_r=0) where she exerted influence suggests she may be driven more by ideology than by evidence. And her [deep involvement](https://www.washingtonpost.com/news/acts-of-faith/wp/2016/11/23/betsy-devos-trumps-education-pick-is-a-billionaire-philanthropist-with-deep-ties-to-the-reformed-christian-community/) in local Christian organizations combined with voucher support suggests she may try to divert public funds to religious schools (this prospect concerns me most).
There’s no need for the exaggeration and fabrication more commonly found on the right. Legitimate criticism of this nomination can stand on its own.
*Update: after *[posting this](https://www.reddit.com/r/atheism/comments/5f2igq/reddit_ratheism_featured_a_misleading_headline_on/)* to r/atheism, they locked the post, threatened to ban me, and then banned me for “brigading”.*
*This post has been updated to reference Snopes’ fact-checking of the Facebook meme. It has also been corrected to reflect that r/politics is not a default subreddit. It was later updated to include Erik Prince’s lack of Trump campaign donations.*
---
## Uber's tipping settlement will reduce earnings of African-American drivers
**URL:** https://maxghenis.com/blog/ubers-tipping-settlement-will-reduce-earnings-of-a/
**Published:** 2016-04-26
**Description:** This week, Uber reached a settlement with drivers in California and Massachusetts in a lawsuit concerning worker classification.
This week, Uber [reached a settlement](http://fortune.com/2016/04/21/uber-drivers-settlement/) with drivers in California and Massachusetts in a lawsuit concerning worker classification. In addition to up to $100M in compensation, recognition of a driver’s association, and greater transparency about deactivating drivers, the settlement permits drivers to post signage in their vehicles soliciting tips. Uber drivers could always accept cash tips, but [unlike Lyft](https://help.lyft.com/hc/en-us/articles/213583978-How-to-Tip-Your-Driver), Uber has never had an in-app tipping option, and [emphasizes](https://help.uber.com/h/1be144ab-609a-43c5-82b5-b9c7de5ec073) that “you don’t need cash…there’s no need to tip.” As a result, tipping is rare: Uber drivers estimate that [between one and five percent](https://www.bostonglobe.com/business/2015/10/27/tips-raises-questions-and-tempers-among-uber-and-lyft-users/EJlZe8Eb763UewRzUNUVJN/story.html) of riders tip, versus [70% of Lyft customers](http://www.wsj.com/articles/uber-reaches-a-tipping-point-with-its-drivers-1461490205) (though [some Lyft drivers claim 20–50%](https://www.bostonglobe.com/business/2015/10/27/tips-raises-questions-and-tempers-among-uber-and-lyft-users/EJlZe8Eb763UewRzUNUVJN/story.html)).
In a political climate focused on fair wages, tipping may seem like an innocuous solution to the gig economy, as adopted by Lyft and other services. However, if Uber’s pricing model works correctly, this move will be unlikely to increase average earnings, and research from the taxi industry suggests it will reduce earnings for black drivers who will be tipped less than their white counterparts.
Since Uber does not plan to add an in-app tip option, it’s unlikely to generate tips from 70% of riders, as Lyft claims. A [2014 Bankrate survey](http://www.bankrate.com/finance/consumer-index/financial-security-index-cashs-cachet.aspx) found that 40% of American consumers carried under $20 in cash, while 9% carried none at all. These fractions have surely grown, and are surely overrepresented among Uber users ([a 2014 study](http://www.icba.org/millennialstudy/) found that 24% of millennials carried under $5 in cash). Tip solicitation will tax the few remaining cash-carriers not already accustomed to Uber’s previous abandonment of the custom. Five percent of fares may be a reasonable expectation of the new tip amount (for instance if 25% of riders tip 20% each).
Pressuring riders to diverge from an increasingly cashless society to tip ([lest they suffer a poor passenger rating from their driver](http://www.bloomberg.com/news/articles/2016-04-22/tipping-is-coming-to-uber-and-it-s-going-to-be-awkward)) will inconvenience users who have come to favor Uber’s easy transactions. But the greater impact will be on black drivers. A [Yale study](http://www.yalelawjournal.org/essay/to-insure-prejudice-racial-disparities-in-taxicab-tipping) provides concerning evidence:
> The authors collected data on more than 1,000 taxicab rides [by twelve drivers] in New Haven, Connecticut in 2001. After controlling for a host of other variables, they find….African-American cab drivers were tipped approximately one-third less than white cab drivers [no significant results were found for other ethnicities]…Both black and white passengers participated in the discrimination against black drivers.
This aligns with other research on tipping, such as a [2014 Cornell study](https://static.secure.website/wscfus/5261551/1619831/socinq-2014-earning-gap.pdf) finding that black servers were tipped ~10% less than white servers. Clearer norms around restaurant tipping relative to taxi tipping likely contribute to the stronger effect in the taxi setting.
Perhaps lower tips for black drivers are tolerable if tipping increases earnings for all drivers, but this is highly unlikely. While Uber [says](http://www.bloomberg.com/news/articles/2016-04-22/tipping-is-coming-to-uber-and-it-s-going-to-be-awkward) that it has no plans to reduce fares under the assumption that riders will tip, their dynamic pricing algorithm means that this will happen anyway. A review of their [surge pricing](https://en.wikipedia.org/wiki/Uber_%28company%29#Surge_pricing) explains why:
> Uber uses an automated algorithm to increase prices to “surge” price levels, responding rapidly to changes of supply and demand in the market, and to attract more drivers during times of increased rider demand, but also to reduce demand…Surge pricing increases economic efficiency in two ways: 1. rising prices motivate more drivers to start driving, 2. when there are not enough drivers for everyone, the rising prices make only those customers accept a ride whose needs are highest.
For example, if more riders in rush hour want rides than drivers are available, fares might be 1.5x normal. If tips become more common (as is expected from this settlement) without fares being reduced, drivers will expect higher hourly rates and thus more will drive (driver supply rises), while riders will be less likely to seek Uber rides given their expectation of higher prices (due to tips). This means that surge pricing will be less necessary, surge multipliers will fall, and the resulting fall in fares will offset average tips. Even drivers who didn’t drive in surge periods would earn less in base fare, due to tip expectation reducing rider demand (e.g. these drivers may end up circling more). An economist might phrase this as: tipping changes neither the supply nor demand curves, so as long as Uber’s algorithm continues to set prices optimally, neither tip-inclusive prices nor tip-inclusive driver earnings will change, on average.
This on-average offsetting of tips by fare reductions ignores tip discrimination. Since black drivers are tipped less, their lost fares will exceed the average tip amount; conversely, white drivers would be expected to come out ahead. To illustrate: assuming [$19/hour in earnings](https://newsroom.uber.com/wp-content/uploads/2015/01/BSG_Uber_Report.pdf) and a 6% tip rate as estimated above (5% of fares, adjusted for Uber’s 20% cut), we would expect $1.14 in average tips. We would also expect fares to fall by $1.14, either due to surge multipliers falling, more circling, or a mix. Given black drivers are tipped ⅓ less than white drivers and a driver population that is [18% black and 37% white](https://newsroom.uber.com/wp-content/uploads/2015/01/BSG_Uber_Report.pdf), we would expect black drivers to be tipped [$0.86/hour vs. $1.28/hour](https://docs.google.com/spreadsheets/d/1D9nooGdzRDlwB42iOoDVe3M_Pv20-QLt5dHIdkU2NFA/edit#gid=0) for white drivers. Adding back the fare of $19 — $1.14 = $17.86/hour, black drivers would earn $18.72 total, 2.2% less than white drivers’ $19.14 total earnings, or 1.5% less than their current $19 earnings. This may not seem like much, but it translates to full-time black drivers earning $870 per year less than their white counterparts. If Uber added in-app tipping and generated Lyft’s 70% tip rate, black drivers would be expected to earn [6.3% less than white drivers, or $2,500 per year](https://docs.google.com/spreadsheets/d/1D9nooGdzRDlwB42iOoDVe3M_Pv20-QLt5dHIdkU2NFA/edit#gid=1559191869).
Research from the restaurant industry has found other issues with tipping: it’s [correlated to sexual harassment](http://rocunited.org/new-report-the-glass-floor-sexual-harassment-in-the-restaurant-industry/) against waitresses (notable for the [14% of Uber drivers](https://newsroom.uber.com/wp-content/uploads/2015/01/BSG_Uber_Report.pdf) who are female) and [corruption](https://dash.harvard.edu/bitstream/handle/1/9491448/Here%27s-a-Tip_Torfason,Flynn,Kupor-Tipping-and-Bribery-6-6-12-SPPS.pdf?sequence=1), and nearly [uncorrelated to service quality](http://scholarship.sha.cornell.edu/cgi/viewcontent.cgi?article=1110&context=articles). While this research hasn’t been replicated in taxis, it’s plausible that the trends extend to transportation.
Few deny that driving for Uber can be financially difficult, but introducing tipping creates more problems than it solves. Should it become the new norm, riders can reduce discrimination by tipping a fixed amount regardless of perceptions of service quality (which are prone to [unconscious bias](http://www.nytimes.com/2015/01/04/upshot/the-measuring-sticks-of-racial-bias-.html?_r=0)). The new challenges of the gig economy will not be solved by voluntary measures like tipping; instead they require bold action from society as a whole to expand the social safety net ([I favor universal basic income](https://medium.com/basic-income/if-we-can-afford-our-current-welfare-system-we-can-afford-basic-income-9ae9b5f186af), as is [popular among technologists](http://www.vice.com/read/something-for-everyone-0000546-v22n1)).
Despite this setback, much progress has been made in tipping reform, with [nearly 200 restaurants](http://bit.ly/tip-free-restaurants) having now abandoned the practice. While Uber’s stance on tipping thus far has protected black drivers from discrimination, permitting drivers to solicit tips weakens such protections. They should remove this clause from their settlement to ensure their platform provides unbiased financial opportunity for all drivers.
**Update:** *Uber clarified in a [blog post](https://medium.com/@UberPubPolicy/our-approach-to-tipping-aa0074c0fddc#.ifie7830q) that "Tipping is not included, nor is it expected or required." The piece also provides rationale for their stance on tipping, which includes bias and references studies listed here.*
*Max Ghenis is the creator of the [EndTipping subreddit](http://reddit.com/r/EndTipping), where he's assembled a [list of 190+ tip-free restaurants in the US](http://bit.ly/tip-free-restaurants). You can sign his petitions to [remove Lyft's in-app tipping option](https://www.change.org/p/lyft-remove-tipping-from-the-lyft-app), and to [reverse Uber's settlement provision permitting drivers to solicit tips](https://www.change.org/p/travis-kalanick-prevent-uber-drivers-from-soliciting-tips-to-ensure-fair-earnings-for-black-drivers).*
---
## If we can afford our current welfare system, we can afford basic income
**URL:** https://maxghenis.com/blog/if-we-can-afford-our-current-welfare-system-we-can/
**Published:** 2016-01-24
**Description:** Replacing the antipoverty bureaucracy
[Basic income](https://en.wikipedia.org/wiki/Basic_income) (BI) is getting a lot of press these days. From [Switzerland’s upcoming referendum](https://en.wikipedia.org/wiki/Swiss_referendums,_2016#Basic_income_referendum) to give each citizen 2,500 francs per month, to [GiveDirectly](http://givedirectly.org)’s research upending the philanthropy world, to [Finland’s plan](https://agenda.weforum.org/2015/12/finland-basic-income/) for a 2017 BI pilot, more are warming to the idea that just giving people money may be the solution to poverty.
> A basic income is an income unconditionally granted to all on an individual basis, without means test or work requirement. — Basic Income Earth Network
Many dismiss basic income as unaffordable. The reality is that basic income is only a form of redistribution, independent of the amount to be redistributed. It is therefore equally affordable to any existing [social safety net](https://en.wikipedia.org/wiki/Social_safety_net). Specifically, budgets for existing programs can fund BI via these revenue-neutral steps:
- Replace current antipoverty programs dollar-for-dollar with cash transfers.
- Reform income eligibility requirements to ensure people are always better off upon earning an extra dollar.(1) and (2) can be consolidated into a “negative income tax” (NIT)
- Implement negative income tax as an equivalent basic income.
That is, NIT should be attractive to those looking to simplify the suite of antipoverty programs and improve work incentives. As I’ll argue below, BI and NIT are effectively equivalent, so BI would be just as attractive. This article lays out the conceptual case for getting to basic income without major changes to federal budgets or tax burdens.
#### Replacing the antipoverty bureaucracy
The US federal government currently spends over a trillion dollars per year helping the poor through over 100 programs (some shown below). Some are well-known, such as [Supplemental Nutrition Assistance Program](https://en.wikipedia.org/wiki/Supplemental_Nutrition_Assistance_Program) (SNAP, f.k.a. food stamps), housing assistance, [Temporary Assistance for Needy Families](https://en.wikipedia.org/wiki/Temporary_Assistance_for_Needy_Families) (TANF, f.k.a. welfare), and the [Earned Income Tax Credit](https://en.wikipedia.org/wiki/Earned_income_tax_credit) (EITC). The majority are much smaller, each providing particular goods or services to those in need, and many more programs provide help at the state and local levels.

This assortment of programs creates three key issues:
- Overhead. Even large federal programs typically have considerable administrative costs. For example, a 2012 study from the Center for Budget and Policy Priorities found that SNAP and housing vouchers each cost 9–10% in overhead expenses. The same study found that the EITC (a cash transfer) cost under 1%.
- Burden on recipients. Administrative dollar figures fail to capture the time-consuming and demeaning experience for recipients, who must learn about, apply for, and comply with the tomes of forms and provisions of each program.
- Paternalism. When programs dictate the specific uses of funds, they ignore the varying preferences of the poor. For example, some may want to spend more on housing, while others may prefer to spend more on food. By purchasing the goods and services for them in fixed quantities, we remove their option to participate in markets according to their tastes. This results in what economists call deadweight loss, a form of inefficiency from lack of market equilibrium.
Some of these programs will likely be more effective than their cash value — especially those serving the physically disabled and mentally ill — and should remain intact. But most could be more easily administered and humanely received as cash grants. Assuming utilization rates remain constant, this is clearly revenue-neutral. Unfortunately, this not always the case due to stigma and difficulty of application and compliance (e.g., [SNAP has ~75% utilization](http://www.huffingtonpost.com/2013/08/14/food-stamps_n_3757052.html)); if an easier cash experience attracts more recipients, benefits could be adjusted to remain revenue-neutral in aggregate.
#### Eliminating welfare cliffs
Replacing current benefits dollar-for-dollar with cash leaves a big problem on the table: welfare cliffs (or [welfare traps](https://en.wikipedia.org/wiki/Welfare_trap)). A welfare cliff occurs when a recipient, upon earning a [marginal](https://en.wikipedia.org/wiki/Marginal_concepts) dollar, loses a [means-tested](https://en.wikipedia.org/wiki/Means_test) benefit (one that goes away if you earn over a certain amount) of greater value than the dollar. This creates work disincentives which can be sizable. For example, as depicted in the below graph, a single mom in Pennsylvania collects more income+benefits earning $29k than she does earning $69k.

Non-cash benefits are particularly prone to welfare cliffs, since they can be real goods and services not easily divisible (e.g. childcare). Cash, on the other hand, can be distributed to eliminate cliffs, plus smooth out the curve. A smooth curve means that anyone can easily predict their total income+benefits as a function of their income, spending less time optimizing for benefits. To be revenue-neutral, some people will be worse off with a smooth curve (e.g. those earning $29k getting maximum benefits), and others will be better off (e.g. those who lose $6k of benefits after earning just over $29k). But every dollar earned will lead to improved livelihood.
Replacing non-cash benefits with cash, and smoothing out the curve, has been formalized into a policy called [negative income tax](https://en.wikipedia.org/wiki/Negative_income_tax) (NIT). This policy provides a government payment to those below a certain income level (called the *phase-out level*)* *equal to a fraction of the difference between their income and the phase-out level (more on that below). Nobel prize-winning economist [Milton Friedman ](https://en.wikipedia.org/wiki/Milton_Friedman)—known as one of the leaders of the [Chicago school of economics](https://en.wikipedia.org/wiki/Chicago_school_of_economics#Milton_Friedman) — popularized the policy when he [advocated](http://www.econlib.org/library/Enc1/NegativeIncomeTax.html) it in his 1962 book, [Capitalism and Freedom](https://en.wikipedia.org/wiki/Capitalism_and_Freedom).
While the proposal never passed, a partial solution called the (previously mentioned) [Earned Income Tax Credit](https://en.wikipedia.org/wiki/Earned_income_tax_credit) (EITC) was implemented instead in 1975. The EITC is now [widely](https://newrepublic.com/article/74453/americas-most-effective-anti-poverty-program) [cited](https://newrepublic.com/article/74453/americas-most-effective-anti-poverty-program) [as](http://www.denverpost.com/breakingnews/ci_25058579/americas-most-effective-anti-poverty-program-is-also) [one](http://www.cbpp.org/research/federal-tax/eitc-and-child-tax-credit-promote-work-reduce-poverty-and-support-childrens) of the most effective antipoverty programs in the US, [having lifted 6.2 million people out of poverty in 2013](http://www.cbpp.org/research/federal-tax/policy-basics-the-earned-income-tax-credit).
The negative income tax goes further by covering all below the phase-out level ([EITC targets working families with children](http://www.cbpp.org/research/federal-tax/policy-basics-the-earned-income-tax-credit)), ensuring full utilization ([15–25% of households eligible for EITC don’t claim it](https://en.wikipedia.org/wiki/Earned_income_tax_credit#Uncollected_tax_credits)), and increasing impact by absorbing budgets from other programs.
#### From negative income tax to basic income
By replacing the onerous bureaucracy of antipoverty programs and eliminating welfare cliffs, negative income tax would go a long way to improving the lives of the poor. However, some differences make basic income more compelling:
- Frequency of distribution. Like other tax rebates, negative income tax payments are distributed annually. Basic income can be distributed monthly or more often, which aligns better to its purpose of satisfying basic ongoing needs.
- Simplicity. Negative income tax still requires some math to determine the payment based on total earnings. Recipients of basic income know exactly how much they get every month.
- Social cohesion. A payment received by all equalizes and unifies a society, creating concrete consequences to government revenue and spending policies. The Nordic countries, for example, have favored universal government programs funded by broad-based taxes.
Fortunately, for any given negative income tax structure, an equivalent basic income structure exists. Consider a NIT with a phase-out level of $30k, a 50% marginal withdrawal rate, and a 50% income tax (unrealistic, but this simplifies calculations). This means that someone earning $0 will receive $15k* (($30k - $0) * 50%)*, someone earning $30k will neither receive payments nor pay income taxes, and someone earning $60k will pay $15k in taxes *(($60k - $30k) * 50%)*. We can visualize this as a chart mapping gross (earned) income to net (after taxes and payments) income (similar to the Pennsylvania welfare cliff chart):

As the title gives away, this same mapping from gross income to net income is achieved with a basic income of $15k and a flat income tax of 50%. The US progressive tax code complicates things a bit, but it’s provable that any negative income tax scheme — regardless of the income tax rates — can be implemented equivalently as a basic income (via income taxes, in some cases changing rates). More money passes through government, but for affected taxpayers, it’s money out one pocket and back in the other.
As the world moves to an age of increasing globalization and automation, more will be drawn to the benefits of basic income. Many questions remain about the idea, for example, its effectiveness, optimal payout, and whom it should include. While studies have addressed some questions (e.g. cash transfers[ don’t reduce work incentives](http://elibrary.worldbank.org/doi/abs/10.1596/1813-9450-3973)), the various experiments planned for the next couple years will yield even better data.
But if basic income is achieved via the above strategy — replace existing non-cash programs with cash, smooth out the payout curve to avoid welfare cliffs, and substitute negative income tax with basic income — affordability does not belong on the list of concerns. If the BI produced by this approach is too low, that is an indication that the existing welfare state is too stingy; if it doesn’t adequately address the needs of society, then neither do today’s programs. Criticisms of the safety net’s size (in either direction) are orthogonal from the notion that basic income is a more efficient means of distribution.
Poverty is not a vague incurable condition, but one that can be directly eradicated via basic income. Finland’s proposal to move 100% to BI is admirable, but other countries can start with small steps like expanding EITC and reforming means tests. With basic income attainable, now is the time to seriously push toward its implementation.
---