9 responses received
Synthesis generated 2026-09-07 13:22 UTC
► Commenting on this synthesis — via Hypothes.is
The annotation panel is open on the right side of this page. Highlight any text and click Annotate to add a comment, question, or pushback — visible to the rest of the team.
Free Hypothes.is account needed (~30 seconds). Use Reply to respond to existing notes, or highlight new text to add your own.
Underlined phrases have supporting quotes from respondents — hover to read.
Strategic Direction
Research focus and pivotal questions
Two broad clusters emerged from the responses, with meaningful tension between them.
AI × economics and labor markets attracted the most convergent support. Pacchiardi frames the AI–economics interface as the most relevant area for UJ focus"AI-economics interface (eg impacts on labour market... or macro-trends à la Ord) seems the most relevant area for the Unjournal to focus on"— Lorenzo Pacchiardi, echoed by Garavaglia (workforce skill shifts and regulatory frameworks), Du (labor welfare, privacy, organizational reform, AI literacy/inequality), Habermacher (realpolitik regulatory frameworks for labor impacts), and Kao (working papers plus semi-formal pieces like Trammell/Patel's "Capital in the 22nd Century" and Citrini's "2028 global intelligence crisis"). Tagat notes that there is already substantial ongoing labor-impact work"There is already a lot of ongoing work related to labor market impacts"— Anirudh Tagat — which, given UJ's curation mandate, is a pipeline advantage rather than a warning. DR agrees this is "very much in our wheelhouse" but flags the open question: does labor-market work reach the GCR-relevant / highest-global-impact issues we most want to elevate?
Governance, regulatory design, and risk parameters form the second cluster. Kalkar highlights gaps in regulatory interventions, comparative international governance, and risk-parameter estimation"There's a gap in regulatory interventions, comparative international governance analysis, and estimates of risk parameters for AI safety and governance"— Uma Kalkar, with pivotal questions on what mechanisms actually change frontier-lab behavior and how capability thresholds translate to regulatory triggers. Schwab pushes toward middle-power regulatory strategies and lessons from arms control / chem-bio prohibitions, especially Chinese cooperation patterns. Habermacher wants politically grounded (not "ivory tower") frameworks. DR notes: risk-parameter estimation is quantitative and comfortable for UJ; but the middle-power / IR-style work and Habermacher's "realpolitik" framing both need concrete examples — what are the papers, methods, and decision-relevant outputs? This may move UJ outside its empirical comfort zone.
Staying timely
Several complementary process ideas surfaced. Manheim urges a fast-track model: invite multiple reviewers, publish as soon as one submits"Fast track, 1+ evaluators, invites to more than 1 reviewer and move forward as soon as 1 is submitted"— David Manheim. Pacchiardi proposes pre-booking evaluator time before paper selection"'pre-booking' evaluator's time before deciding what paper to review so that you can choose a paper and are sure that someone will be able to look at that"— Lorenzo Pacchiardi — DR flags this as "a great idea," worth combining with giving evaluators limited choice among candidate papers. Pacchiardi also suggests crux-targeted 1-hour reviews using the evaluation explorer. Garavaglia suggests one evaluator plus AI-assisted briefing; Kao proposes ACX/Zvi-style commentary roundups clustering existing online reactions"ACX/Zvi Moshowitz style 'commentary roundups' that presents clusters of comments made by others online + light discussion"— Andrew Kao as a complement, not substitute. Kalkar cautions against further automation to protect quality standards. Tagat suggests an NBER-track-style pipeline for top-center AI-economics papers (though DR asks for clarification on how NBER is currently overlooked).
Who to bring in
Named individuals: Jonathan Prunty (Leverhulme CFI, Cambridge) and Marko Tesic (UK DSIT) for labor-market impacts (Pacchiardi). Institutional networks: RAND TASP Fellows and current/former GovAI Fellows (Kalkar); GovAI more broadly and possibly BlueDot (Manheim); ILO and AI developers for labor-welfare work (Du); AI-safety/catastrophic-risk funders including Schmidt Sciences to evaluate their commissioned work (Tagat); the existing AI Slack channel as a crowdsourcing venue (Tagat).
The Ord RL Scaling Evaluation
Should we still pursue it?
Respondent views were split before the latest development. Schwab and Habermacher felt the moment had likely passed"Moment passed."— Josephine Schwab; Manheim and Pacchiardi were lukewarm; Kalkar was the strongest advocate, arguing that the messier empirical picture creates a first-mover opportunity"Given that the empirical picture has become messier and more contested, it may actually make sense to be a 'first-mover' and evaluate it now"— Uma Kalkar.
This calculus has now changed materially. Toby Ord replied to DR on April 29, and the response substantially reopens the case:
- Ord endorses formal evaluation: he would be "very happy for there to be more formal analysis of these questions" (his own were "very quick analyses"). He won't participate, but does not object — a green light rather than an obstacle.
- He clarified his actual quantitative position, which is more nuanced than the "RL is hitting a wall" reading. In his framing, progress was ~60% compute / 40% other in the GPT-2/3/4 era; roughly half the compute boost is going away, leaving ~30% + 40% = 70% of prior speed, absent new compute-scaling forms or recursive self-improvement.
- This is a specific, evaluable quantitative claim not previously stated this clearly in public — and the gap between Ord's actual modest position and how it has been characterised in commentary is itself worth documenting.
- Ord intended a wrap-up essay for the EA Forum Scaling Series but won't write it — a concrete gap the UJ synthesis could directly fill.
The Schwab/Habermacher "moment passed" view should be updated in light of this: there is now (i) author endorsement, (ii) a sharpened quantitative claim to evaluate, and (iii) an identified publication gap.
If yes — what form?
Kalkar's proposal remains the strongest concrete template: a long-form EA Forum / PubPub post covering the full evidence base"a long-form EA Forum/ PubPub post. It should cover the full evidence base... with clear separation between what's empirically established, what's contested"— Uma Kalkar, with clear separation of established vs. contested claims and open research questions, circulated in draft to the GovAI community (and, given his endorsement, shared with Ord). This format also fits the wrap-up-essay role Ord vacated.
Points of apparent consensus
- AI × economics/labor is a rich, tractable pipeline for UJ evaluation — with the open question of whether it reaches the highest-GCR-impact work (DR).
- Broad opposition to expanding into technical AI safety evaluation, largely on comparative-advantage grounds (Alignment Journal, METR, Epoch, ARC Evals). DR qualifies: he takes the team-fit concern seriously, but is not yet convinced the space is fully covered — the Alignment Journal is scoped to alignment specifically, leaving other technical-safety areas potentially underserved.
- Faster turnaround is achievable via some mix of pre-booked evaluators, single-reviewer fast tracks, and AI-assisted briefing — without full automation.
- GovAI and RAND TASP networks are the natural recruiting pools for governance-side expertise.
Key tensions worth discussing
- Wheelhouse vs. GCR-relevance: the AI-economics cluster fits UJ's methods; the governance/risk-parameter cluster is arguably more mission-critical but demands less-quantitative expertise UJ hasn't built.
- Peer-reviewing model cards and technical reports (Manheim) drifts toward technical AI safety territory — the very expansion most respondents oppose. Can UJ reconcile these?
- Speed vs. rigor: single-reviewer fast tracks and commentary roundups risk diluting the credibility signal that is UJ's core product.
- Ord evaluation: the "moment passed" intuition now conflicts with author endorsement, a newly sharpened quantitative claim, and an identified publication gap — the balance has shifted toward proceeding.
Individual responses (9)
Josephine Schwab — ai_researcher 2026-04-25
David Manheim — field_specialist 2026-04-21
Lorenzo Pacchiardi — ai_researcher 2026-04-17
Otherwise, simply paying more allows people to drop other priorities.
Also agree that targeting very specific cruxes in a paper (eg highlighted from the evaluation explorer) could be more efficient, for instance requiring reviewers to be useful with 1 hour of work only
- Jonathan Prunty (Leverhulme Centre for the Future of Intelligence, University of Cambridge)
- Marko Tesic (DSIT, UK government)
Zhuoran Du 2026-04-17
AI developers are important
Uma Kalkar — ai_researcher 2026-04-17
Examples of possible RQs:
-- What governance mechanisms actually change frontier lab behavior, and under what conditions?
-- How do capability thresholds translate into tractable regulatory triggers?
-- What predicts adoption vs. resistance for cross-national diffusion of AI governance frameworks?
-- Consider the current/former GovAI Fellows
(These cohorts will already be specializing/focusing on elements of AI safety and governance research so it may be easier to get them to help review)
Would definitely circulate a first draft for edits across the GovAI community (Toby Ord included).
Pía Garavaglia — economist 2026-04-16
Anirudh Tagat — uj_team 2026-04-16
I think engage with funders of AI and catastrophic risk, alignment etc. (I think Schmidt is interested, but are only funding research right now) -- it might be useful to reach out to them to provide an open evaluation of the work that they are commissioning / providing grants to. This way we also get to engage with AI researchers working at the forefront (at least in economics, broadly).
Florian Habermacher — economist 2026-04-16
- politically grounded (I mean 'realpolitik' type not 'purely ivory tower abstract') Regulatory Frameworks for dealing with labor impacts (and maybe with social impacts more broadly)
Every Sat and Sun from 8 AM CEST until midnight CEST (most days)
Mon-Fri CET office hours (8 AM CEST until 6 PM CEST): variable availability (50% available 50% unavailable)
Andrew Kao — field_specialist 2026-04-15
But also worth considering evaluations of slightly less formal (but still high effort) pieces: as two examples, Phil Trammel and Dwarkesh Patel's post on Capital in the 22nd Century https://substack.com/@philiptrammell/p-182789127 and Citrini research's 2028 global intelligence crisis https://www.citriniresearch.com/p/2028gic
Separately, I wonder if something along the lines of ACX/Zvi Moshowitz style 'commentary roundups' that presents clusters of comments made by others online + light discussion could be useful. This would be as a complement, not substitute, to existing eval effort. For this, I think the relevant question is whether the typical reader of an evaluation is plugged into discourse enough to already know the prevailing sentiment/feedback towards a piece or not.