9 responses received
Synthesis generated 2026-07-21 13:18 UTC
► Commenting on this synthesis — via Hypothes.is
The annotation panel is open on the right side of this page. Highlight any text and click Annotate to add a comment, question, or pushback — visible to the rest of the team.
Free Hypothes.is account needed (~30 seconds). Use Reply to respond to existing notes, or highlight new text to add your own.
Underlined phrases have supporting quotes from respondents — hover to read.
Strategic Direction
Research focus and pivotal questions
Two broad clusters emerge from the responses, with a real tension between them that DR flags directly.
AI × economics and labor markets attracts the most convergent enthusiasm. Pacchiardi frames the AI-economics interface as most relevant for UJ"AI-economics interface (eg impacts on labour market [Anthropic's work] or macro-trends à la Ord) seems the most relevant area"— Lorenzo Pacchiardi, echoed by Garavaglia (workplace impacts, worker skills, regulatory frameworks), Du (labor welfare, privacy, organizational reform, AI literacy/inequality), and Habermacher (regulatory frameworks for labor impacts). Tagat notes substantial existing work in this area"There is already a lot of ongoing work related to labor market impacts"— Anirudh Tagat — which, given UJ's curation-not-production role, is a feature: a rich pipeline to evaluate. DR pushes back on framing this as a "warning" and confirms this cluster is very much in UJ's wheelhouse — but asks whether it really addresses GCR-relevant or highest-global-impact issues.
Governance, regulatory design, and risk parameters form the second cluster, more GCR-adjacent but demanding different expertise. Kalkar highlights gaps in regulatory interventions and comparative international governance analysis"gap in regulatory interventions, comparative international governance analysis, and estimates of risk parameters for AI safety and governance"— Uma Kalkar, with pivotal questions on what actually changes frontier lab behavior and how capability thresholds translate to regulatory triggers. Schwab pushes further into international relations: middle-power strategies, red-lines convergence, and lessons from arms control on Chinese cooperation. Habermacher wants "realpolitik"-grounded regulatory frameworks rather than ivory-tower abstractions.
DR's caveats matter here. He is comfortable with quantitative risk-parameter estimation but wants concrete examples of the regulatory-intervention and middle-power/IR work before committing — this would move UJ away from its empirical/quantitative comfort zone. He also asks Habermacher to specify what realpolitik-grounded research actually looks like, who produces it, and how it could inform funder/policymaker decisions.
Kao offers a third angle worth attention: extending evaluation to high-effort but less formal pieces like Trammell/Patel and Citrini Research"worth considering evaluations of slightly less formal (but still high effort) pieces"— Andrew Kao. Tagat additionally suggests engaging AI-risk funders (e.g., Schmidt) to openly evaluate commissioned work — a channel into frontier researchers.
Staying timely
Several converging process proposals emerged. Manheim wants aggressive fast-tracking: invite multiple reviewers and move forward as soon as one submits"Fast track, 1+ evaluators, invites to more than 1 reviewer and move forward as soon as 1 is submitted"— David Manheim. Garavaglia and Kao both accept single-evaluator outputs with AI-assisted briefing when a second reviewer is hard to secure.
Pacchiardi's pre-booking evaluator time before paper selection attracted DR's explicit endorsement — DR notes this is a great idea he has considered before"This seems like a great idea to me. Talked about it in the past, as well as giving evaluators some limited choice over which ones they would like to evaluate."— David Reinstein, potentially combined with giving reviewers limited choice among available pieces. Pacchiardi also floats crux-targeted 1-hour reviews keyed to specific claims (e.g., via the evaluation explorer).
Kao's commentary roundup format — ACX/Zvi-style clustering of existing online commentary with light discussion — could complement (not substitute) full evaluations. Kalkar cautions against over-automation: the LLM eval and prioritization tools are useful but not a substitute for humans"That tension of rigor vs. speed will always be there; I don't think I would suggest automating the process more for fear of possible impacting standards/quality"— Uma Kalkar. Tagat proposes an "NBER-track" pipeline targeting top research centers with faster turnarounds. DR requests clarification on what "NBER-overlooked" means here.
Who to bring in
Named individuals: Jonathan Prunty (Leverhulme CFI, Cambridge) and Marko Tesic (DSIT UK) for labor-market work (Pacchiardi). Networks: RAND TASP Fellows and current/former GovAI Fellows (Kalkar), BlueDot (Manheim), the AI Slack channel for crowdsourcing (Tagat). Institution types: ILO and AI developers directly (Du); AI-risk funders like Schmidt for commissioned-work evaluation (Tagat).
Technical AI Safety Expansion
Near-unanimous opposition. Manheim is "very much opposed to trying to compete on this""No, very much opposed to trying to compete on this"— David Manheim; Pacchiardi cites the Alignment Journal and traditional AI conferences as crowding the niche; Kalkar notes ARC Evals, METR, and Epoch already cover the space; Tagat recommends sticking to demonstrated domain expertise.
DR takes this seriously but is not fully convinced. He notes the Alignment Journal is moving toward a traditional-journal posture and is scoped only to alignment, leaving other technical AI safety research potentially uncovered by rapid, credible, expert evaluation. He also flags a tension: Manheim's proposal to peer-review model cards and technical reports would pull UJ into exactly the technical territory respondents want to avoid.
The Ord RL Scaling Evaluation
Should we still pursue it?
Views were split before new developments: Schwab said the "moment passed," Habermacher and Manheim were lukewarm, Pacchiardi remained on the fence, while Kalkar argued the messier empirical picture makes now a good "first-mover" window given Ord's reframing to 2027+.
Ord's April 29 email response substantially updates this calculus. Ord endorses more rigorous analysis ("I'd be very happy for there to be more formal analysis of these questions — mine were very quick analyses"), won't participate but does not object, and — most importantly — clarified his actual quantitative position for the first time: something like 60% compute / 40% other in the GPT2–4 era shifting to roughly 30% + 40% = 70% of old speed, with other drivers (algorithms, data, RL environments, recursive self-improvement) potentially compensating. He is not claiming RL has hit a fundamental limit — only that compute's relative contribution is declining.
This changes three things. First, there is now a specific, evaluable quantitative claim not previously stated this cleanly in public. Second, the gap between Ord's actual position and its "near effective limit" characterization is itself worth documenting. Third, Ord had intended a wrap-up essay for the EA Forum Scaling Series but won't write it — a gap a UJ synthesis could directly fill with author endorsement. The "moment passed" view (Schwab, Habermacher) should be reconsidered in light of this.
If yes — what form?
Kalkar's proposal remains the most concrete: a long-form EA Forum / PubPub post with clear separation between established, contested, and open"long-form EA Forum/ PubPub post. It should cover the full evidence base... with clear separation between what's empirically established, what's contested, and where further research is needed"— Uma Kalkar, with a draft circulated to the GovAI community including Ord. Ord's new quantitative framing should anchor the evaluable claim.
Points of apparent consensus
- Do not expand into technical AI safety evaluation as a primary strategic direction (with DR's caveat that some niche opportunities may exist beyond alignment-narrow work).
- AI × economics/labor is the most natural fit for UJ's existing capacity, though its GCR-relevance is an open question DR wants pressed.
- Some form of fast-track process is needed: single-evaluator publication when necessary, AI-assisted briefing, pre-booked evaluator slots.
- Tap existing networks (GovAI, RAND TASP, AI Slack) rather than building new outreach from scratch.
Key tensions worth discussing
- Wheelhouse vs. impact: AI × labor fits UJ's skills but may not touch the highest-stakes questions; governance/IR/middle-power work is more GCR-relevant but stretches UJ's quantitative-empirical comfort zone.
- Rigor vs. speed: Manheim's aggressive fast-tracking (publish on first submission) vs. Kalkar's caution against automation-driven quality erosion.
- Formality boundary: Manheim and Kao both push toward reviewing less-formal artifacts (model cards, technical reports, substack essays), but DR notes model-card review pulls UJ into technical AI safety territory respondents rejected.
- Ord evaluation timing: The "moment passed" intuition vs. Ord's fresh, endorsed, quantitatively specific clarification that creates a genuine synthesis opportunity.
- Specification gaps: DR flags that "regulatory interventions," "realpolitik frameworks," and "middle-power strategies" need concrete examples before UJ can judge fit — respondents should be pressed for exemplars.
Individual responses (9)
Josephine Schwab — ai_researcher 2026-04-25
David Manheim — field_specialist 2026-04-21
Lorenzo Pacchiardi — ai_researcher 2026-04-17
Otherwise, simply paying more allows people to drop other priorities.
Also agree that targeting very specific cruxes in a paper (eg highlighted from the evaluation explorer) could be more efficient, for instance requiring reviewers to be useful with 1 hour of work only
- Jonathan Prunty (Leverhulme Centre for the Future of Intelligence, University of Cambridge)
- Marko Tesic (DSIT, UK government)
Zhuoran Du 2026-04-17
AI developers are important
Uma Kalkar — ai_researcher 2026-04-17
Examples of possible RQs:
-- What governance mechanisms actually change frontier lab behavior, and under what conditions?
-- How do capability thresholds translate into tractable regulatory triggers?
-- What predicts adoption vs. resistance for cross-national diffusion of AI governance frameworks?
-- Consider the current/former GovAI Fellows
(These cohorts will already be specializing/focusing on elements of AI safety and governance research so it may be easier to get them to help review)
Would definitely circulate a first draft for edits across the GovAI community (Toby Ord included).
Pía Garavaglia — economist 2026-04-16
Anirudh Tagat — uj_team 2026-04-16
I think engage with funders of AI and catastrophic risk, alignment etc. (I think Schmidt is interested, but are only funding research right now) -- it might be useful to reach out to them to provide an open evaluation of the work that they are commissioning / providing grants to. This way we also get to engage with AI researchers working at the forefront (at least in economics, broadly).
Florian Habermacher — economist 2026-04-16
- politically grounded (I mean 'realpolitik' type not 'purely ivory tower abstract') Regulatory Frameworks for dealing with labor impacts (and maybe with social impacts more broadly)
Every Sat and Sun from 8 AM CEST until midnight CEST (most days)
Mon-Fri CET office hours (8 AM CEST until 6 PM CEST): variable availability (50% available 50% unavailable)
Andrew Kao — field_specialist 2026-04-15
But also worth considering evaluations of slightly less formal (but still high effort) pieces: as two examples, Phil Trammel and Dwarkesh Patel's post on Capital in the 22nd Century https://substack.com/@philiptrammell/p-182789127 and Citrini research's 2028 global intelligence crisis https://www.citriniresearch.com/p/2028gic
Separately, I wonder if something along the lines of ACX/Zvi Moshowitz style 'commentary roundups' that presents clusters of comments made by others online + light discussion could be useful. This would be as a complement, not substitute, to existing eval effort. For this, I think the relevant question is whether the typical reader of an evaluation is plugged into discourse enough to already know the prevailing sentiment/feedback towards a piece or not.