9 responses received
Synthesis generated 2026-07-20 13:18 UTC
► Commenting on this synthesis — via Hypothes.is
The annotation panel is open on the right side of this page. Highlight any text and click Annotate to add a comment, question, or pushback — visible to the rest of the team.
Free Hypothes.is account needed (~30 seconds). Use Reply to respond to existing notes, or highlight new text to add your own.
Underlined phrases have supporting quotes from respondents — hover to read.
Strategic Direction
Research focus and pivotal questions
Two broad clusters emerged, with meaningful tension between them on GCR-relevance versus UJ's existing comfort zone.
Cluster 1 — AI × economics/labor markets. Multiple respondents converged here as UJ's natural wheelhouse. Pacchiardi endorsed the AI-economics interface as most relevant"AI-economics interface (eg impacts on labour market [Anthropic's work] or macro-trends à la Ord) seems the most relevant area"— Lorenzo Pacchiardi, with Du, Garavaglia, Habermacher, and Kao echoing labor market impacts, workplace transformation, and skills/retraining regulation. Tagat noted there is already substantial ongoing work in this space"There is already a lot of ongoing work related to labor market impacts."— Anirudh Tagat — which, given UJ's evaluate-not-produce mandate, means a rich pipeline. DR notes: strong work exists here and it's very much in our wheelhouse, but he flags the open question of whether it truly gets at GCR-relevant issues or the highest global-impact questions. Habermacher pushed specifically for "realpolitik"-type regulatory frameworks for labor impacts; DR flagged this as too abbreviated — what methods, what producers, what decisions would it inform?
Cluster 2 — governance, regulation, and risk parameters. Kalkar identified gaps in regulatory interventions, comparative international governance, and quantitative risk-parameter estimation, offering concrete pivotal questions: what governance mechanisms actually change frontier lab behavior"What governance mechanisms actually change frontier lab behavior, and under what conditions?"— Uma Kalkar, and how capability thresholds translate into tractable regulatory triggers. Schwab pushed for middle-power regulatory strategies, cross-border incident cooperation, and lessons from arms-control precedents (chemical/biological prohibitions, Chinese ratification patterns). DR: the risk-parameters angle is quantitative and comfortable for us; the middle-power/IR angle would require moving beyond our current empirical comfort zone — needs concrete examples of the work before committing.
Cluster 3 — high-quality informal writing. Kao suggested extending evaluation to less formal but still high-effort pieces"slightly less formal (but still high effort) pieces: as two examples, Phil Trammel and Dwarkesh Patel's post on Capital in the 22nd Century"— Andrew Kao like Trammell/Patel on Capital in the 22nd Century or Citrini's 2028 intelligence-crisis analysis. Manheim complementarily proposed fast peer review of "non-peer-reviewable" documents (model cards, think-tank technical reports) — though DR notes this pulls toward the technical AI safety sphere the group was against.
On technical AI safety expansion, the consensus is negative — Manheim is "very much opposed to trying to compete," Pacchiardi cites the Alignment Journal and mainstream AI conferences as already covering the niche, Kalkar warns off given ARC Evals/METR/Epoch coverage, and Tagat advises sticking to domain expertise. DR takes this seriously but is not fully convinced: he notes the Alignment Journal is drifting toward traditional-journal norms and is limited to alignment specifically, leaving other technical AI safety subfields potentially underserved by rapid credible evaluation.
Staying timely
Several complementary process ideas surfaced. Manheim proposed a fast-track model with a single evaluator"Fast track, 1+ evaluators, invites to more than 1 reviewer and move forward as soon as 1 is submitted."— David Manheim — invite multiple, publish when the first submits. Kao endorsed this: multiple evaluators are valuable but shouldn't block publication. Pacchiardi's pre-booking evaluator time before paper selection"'pre-booking' evaluator's time before deciding what paper to review so that you can choose a paper and are sure that someone will be able to look at that in a timely manner"— Lorenzo Pacchiardi proposal was singled out by DR as "a great idea" worth revisiting, potentially paired with giving evaluators limited choice among candidates. Pacchiardi also proposed crux-targeted 1-hour reviews focused on specific evaluation-explorer claims. Garavaglia suggested one evaluator plus AI-assisted briefing; Kalkar endorsed LLM eval and prioritization tools with humans-in-the-loop but warned against over-automation. Kao floated ACX/Zvi-style commentary roundups as a complement to formal evaluations. Tagat suggested mimicking an NBER-track pipeline with faster turnaround (DR asked for clarification on the NBER framing).
Who to bring in
Specific names: Jonathan Prunty (Leverhulme CFI, Cambridge) and Marko Tesic (DSIT) for AI-labor (Pacchiardi). Networks: RAND TASP Fellows and current/former GovAI Fellows (Kalkar); GovAI folks and possibly BlueDot (Manheim); the International Labour Organisation for labor-welfare work (Du); AI-catastrophic-risk funders like Schmidt for evaluating their commissioned work (Tagat); the AI Slack channel for crowdsourcing (Tagat).
The Ord RL Scaling Evaluation
Should we still pursue it?
Initial responses were mixed-to-lukewarm: Schwab said the "moment passed," Habermacher thought it not top priority, Manheim was unsure UJ had a good place for a take, Pacchiardi remained on the fence. Kalkar was the strongest advocate for proceeding, arguing that the messier empirical picture creates a first-mover opportunity"Given that the empirical picture has become messier and more contested, it may actually make sense to be a 'first-mover' and evaluate it now."— Uma Kalkar, especially with Ord's reframing to 2027+.
New development (Apr 29): Ord's email response substantially changes the calculus. Ord explicitly endorses more formal analysis of these questions"I'd be very happy for there to be more formal analysis of these questions (mine were very quick analyses)"— Toby Ord, will not participate but does not object, and — critically — will not write the wrap-up essay he had intended for the EA Forum Scaling Series. That is a concrete gap UJ could fill.
More importantly, Ord clarified his actual quantitative claim, which is more nuanced than critics have characterized: he sees compute scaling as shifting from ~60% to ~30% of AI progress"speed was driven by something like 60% compute, 40% other... leaving it at 30% + 40% = 70% of its old speed unless other things... come in to boost things"— Toby Ord, yielding ~70% of prior speed absent new compensating factors — not a claim that RL has hit a fundamental limit. This reopens the "moment passed" verdict: there is now (1) an explicit endorsement, (2) a newly precise quantitative claim to evaluate, and (3) a documentable gap between Ord's actual position and its public reception — itself worth clarifying.
If yes — what form?
Kalkar proposed a long-form EA Forum / PubPub post covering the full evidence base with clear separation between what is empirically established, contested, and open — circulated in draft to the GovAI community including Ord. This maps well onto filling Ord's unwritten wrap-up-essay slot.
Points of apparent consensus
- Stay within AI governance and AI × economic/social impacts; do not attempt to compete on technical AI safety (with DR's caveat that some technical-safety subfields outside alignment may still be underserved).
- Timeliness improvements are worth serious investment: fast-track single-evaluator publication, pre-booked evaluator time, and AI-assisted briefing are all viable.
- Leverage existing specialized networks (GovAI, RAND TASP, AI Slack) rather than cold outreach.
- Ord himself endorsing rigorous analysis meaningfully strengthens the case for the RL scaling synthesis.
Key tensions worth discussing
- Wheelhouse vs. GCR-relevance: AI-labor work is where UJ has the deepest evaluator base, but DR's question — does this get at the highest-global-impact issues? — remains unresolved. Governance/risk-parameter work is more GCR-relevant but requires expertise UJ is still building.
- Speed vs. rigor: single-evaluator fast-track and AI briefing raise standards concerns (Kalkar); the group has not settled on where the quality floor sits.
- Model cards and technical reports: Manheim's proposal to peer-review these is attractive for timeliness and gap-filling, but pulls UJ toward the technical AI safety sphere respondents rejected — DR flagged this contradiction directly.
- Ord evaluation timing: "moment passed" (Schwab, Habermacher) versus "first-mover opportunity" (Kalkar) — Ord's new clarification and endorsement tilt this balance toward proceeding, but only if UJ can move quickly.
- Scope of "governance" work: DR's request for concrete examples of middle-power / realpolitik-regulatory / regulatory-intervention research before committing reflects unresolved uncertainty about whether these fit UJ's empirical-quantitative comfort zone.
Individual responses (9)
Josephine Schwab — ai_researcher 2026-04-25
David Manheim — field_specialist 2026-04-21
Lorenzo Pacchiardi — ai_researcher 2026-04-17
Otherwise, simply paying more allows people to drop other priorities.
Also agree that targeting very specific cruxes in a paper (eg highlighted from the evaluation explorer) could be more efficient, for instance requiring reviewers to be useful with 1 hour of work only
- Jonathan Prunty (Leverhulme Centre for the Future of Intelligence, University of Cambridge)
- Marko Tesic (DSIT, UK government)
Zhuoran Du 2026-04-17
AI developers are important
Uma Kalkar — ai_researcher 2026-04-17
Examples of possible RQs:
-- What governance mechanisms actually change frontier lab behavior, and under what conditions?
-- How do capability thresholds translate into tractable regulatory triggers?
-- What predicts adoption vs. resistance for cross-national diffusion of AI governance frameworks?
-- Consider the current/former GovAI Fellows
(These cohorts will already be specializing/focusing on elements of AI safety and governance research so it may be easier to get them to help review)
Would definitely circulate a first draft for edits across the GovAI community (Toby Ord included).
Pía Garavaglia — economist 2026-04-16
Anirudh Tagat — uj_team 2026-04-16
I think engage with funders of AI and catastrophic risk, alignment etc. (I think Schmidt is interested, but are only funding research right now) -- it might be useful to reach out to them to provide an open evaluation of the work that they are commissioning / providing grants to. This way we also get to engage with AI researchers working at the forefront (at least in economics, broadly).
Florian Habermacher — economist 2026-04-16
- politically grounded (I mean 'realpolitik' type not 'purely ivory tower abstract') Regulatory Frameworks for dealing with labor impacts (and maybe with social impacts more broadly)
Every Sat and Sun from 8 AM CEST until midnight CEST (most days)
Mon-Fri CET office hours (8 AM CEST until 6 PM CEST): variable availability (50% available 50% unavailable)
Andrew Kao — field_specialist 2026-04-15
But also worth considering evaluations of slightly less formal (but still high effort) pieces: as two examples, Phil Trammel and Dwarkesh Patel's post on Capital in the 22nd Century https://substack.com/@philiptrammell/p-182789127 and Citrini research's 2028 global intelligence crisis https://www.citriniresearch.com/p/2028gic
Separately, I wonder if something along the lines of ACX/Zvi Moshowitz style 'commentary roundups' that presents clusters of comments made by others online + light discussion could be useful. This would be as a complement, not substitute, to existing eval effort. For this, I think the relevant question is whether the typical reader of an evaluation is plugged into discourse enough to already know the prevailing sentiment/feedback towards a piece or not.