“If you want to be the best at anything, be the best at learning.”
New Projects
Uzbek SOEs
Some positive developments from the EA Global conference last month (thanks, Swapcard!). I met an economist from the World Bank who is currently working with Uzbekistan’s Ministry of Economy and Finance (MoEF) to help implement the World Bank’s (admittedly neoliberal) agenda for reforming Uzbekistan’s state-owned enterprises (SOEs).
The relevant context is that, as a former Soviet state, Uzbekistan has a weird number of state-owned industries (covering 42% of major economic sectors), including many that would normally be strange to run this way, such as agriculture and mining. The main issue with these SOEs is the lack of competition in sectors that would otherwise be quite competitive, leading to abysmal productivity, rent-seeking behaviour, and corruption.
While there are certainly ideological factors driving the push for liberalisation, Uzbekistan has good reasons to pursue reforms to this system.
The main original request from the World Bank/MoEF contact was to help “find which SOEs could have the greatest productivity increases from AI”, which I felt was a somewhat dubious thing to focus on. Instead, we worked out a better use case.
The problem, and the openness to technical solutions, provide a potentially promising opportunity to try one of the more ambitious AGF project streams I have in mind: AI policy assistants. Specifically, using large language models to help MoEF ministers synthesise large amounts of information on the SOEs (both quantitative and qualitative) and use it to make evidence-based decisions targeting SOE reform.
Nothing too complicated here: the goal is simply to synthesise all the quantitative and qualitative information about the SOEs (including the SOE reform reports) and output specific, actionable recommendations in line with World Bank research and within the MoEF’s legal mandate.
It seems that this (i.e., AI policy assistants, or at least using some form of large language models for decision assistance) is a direction many public-sector decision-making bodies are going in, and I want the AGF to be as close to the forefront of implementation as possible to ensure this is done right.
While the specific context of the SOEs is unique to Uzbekistan, the problem faced by the MoEF can be thought of more generally as an information synthesis → decision problem. There is a large corpus of quantitative data and qualitative records (reports, recommendations, statutes, and policy agendas) that must be understood in totality before decisions leveraging all the available information can be made.
Language models are certainly “there” in terms of their ability to retrieve information from documents (a pretty banal use case), but the gap lies in retrieving information in a useful way. For policymakers, this means ensuring that the retrieved information is turned into actionable recommendations in the form of a familiar policy briefing.
The key to success for these kinds of projects, in my experience, is bridging the gap between a “cool, technically interesting project” and something that actually influences decision-making. There is a kind of “tool hell” that developers can end up in when building technical solutions for governments as outsiders (I assume this problem also exists in the private sector): some flashy, potentially even user-friendly product is created, and absolutely no one uses it. This is much to the frustration of the developers, who are deeply convinced that their tool is the solution to every problem and will become infinitely useful if only one more feature is added.
The challenge here is not really technical; it is finding the right workflow in which this solution can actually be useful. This project seems like an exciting way to experiment with how large language models and RAG can be optimally incorporated into the policymaking process (and what better way to experiment than with a major national government department?).
Of course, I would also like to assess the efficacy of this approach by comparing it with the next-best alternative (i.e., manual, human-led SOE reviews) to determine whether the extra effort and technical debt are justified.
I will be hiring for this project, which will mark the first full-time AGF hire since we incorporated as a CIC last September (time to start thinking about HR stuff, yikes). If all goes well, the new hire and I will travel to Tashkent at the end of autumn to trial the tool and hand it over to the MoEF.
Southern Water
I am also in the process of scoping another project with Southern Water, specifically on predicting combined sewer overflows, or CSOs (I did some independent work on this back in April).
A simple project: just a prediction model using time-series data. The innovation will be to make better probabilistic predictions, which are needed for more confident decisions about when to act to avoid CSO events.
Model-Action Protocols (MAPs)
Underlying both the MoEF and Southern Water projects is something that I’ve been thinking about increasingly: linking models to actions. This is by far the biggest implementation gap in algorithmic governance. Prediction models and RAG are pretty standard/solved problems, but these tools are heavily underutilised in areas where they could have the greatest impact.
At first glance, this seems like a massive failure of adoption (and it is, in part, at least for less well-funded public-serving departments that chronically lag in technology adoption). But the more critical failure lies in linking model outputs to actions, particularly when those outputs can be quite technical or, in the case of RAG-type tools, extremely unstructured.
To address this, I’ve introduced “Model-Action Protocols” (MAPs) for every AGF project. These are essentially documents completed collaboratively with the partner organisation that explicitly link model outputs to a set of approved actions the organisation is comfortable taking based on those outputs. This is based partly on the Early Action Protocols used by the IFRC, such as the one linked to my White Nile flood prediction model.
This approach doesn’t seem particularly innovative, but it is absolutely critical for turning technical solutions into concrete action. Developers often have a completely different notion of what is useful from that of actual decision-makers, while decision-makers are often uncomfortable relying blindly on complicated models. You can have the most accurate and well-calibrated flood prediction model in the world, but without an action plan linked to its outputs, the best-case scenario is that the model’s prediction (likely distilled from a technically precise confidence statement, associated spatial prediction, and hazard analysis into some drivel like “model predicts high probability of flood!”) is included in a report to a minister, who will glance at it and call a meeting to understand what it says. Whether this process will ultimately have any bearing on a decision is fairly tenuous, and it is certainly less efficient than an automatic process. Provided these actions are agreed upon in advance and the model has reached an acceptable level of accuracy, linking specific actions to model outputs closes the implementation gap and enables actual automation.
The MAP also assigns a specific person to manage the model, which is critical for a successful handover.
Old Projects
BioScanCast Trial
The AGF (in partnership with the Biosecurity Forecasting Group) has officially launched a live trial of BioScanCast, our web-scraping, information-synthesising, and forecasting LLM tool. This means approximately 40 biosecurity experts will be competing in a live forecasting round against our model. Both the humans and our model will be forecasting the outcomes of 25 actively developing biosecurity risks, such as measles incidence in North America and mpox outbreaks, based on public health data and news reports.
We have a “baseline” model that just uses the Perplexity model to do live research without any additional custom scaffolding for the domain problem. This can be compared against BioScanCast, which has substantial additional robustness safeguards, including source-quality checks, custom dashboard scraping, and the option to access its own historical JSON output of scraped insights.
Advertisement for the forecast round.
The forecasts won’t resolve until the end of 2026, so this project is effectively paused until then as the forecasts run (our model executes automatically, so no further action needed from us at this point).
The goal of building something like this and running this trial is to establish a proof of concept for whether large language models can help reduce the cost of biosecurity surveillance. If our models are even just comparable to the median human output, this presents a strong argument that surveillance of public health reports can be partially automated at an extremely low cost. For reference, the human forecasters may spend hours per week monitoring all the sources needed to make their predictions. By comparison, the model runs automatically, can operate 24/7, and costs only a fraction of a cent to monitor each question.
The experts are also limited by their own domain: since we are forecasting such a wide variety of biosecurity threats, it’s likely that no single expert will have expertise in all the forecast questions. The model, on the other hand, can perform equally well across a variety of domains.
Thoughts on Not Attending the G7 This Year
For the first time since 2023, I’ve decided not to attend any of the G7/G20 summits this year. I made the decision following last year’s G7 Summit in Kananaskis, where I presented two policy briefings at Think7.
I was pretty on the fence about attending last year, but felt obligated to go because I was the lead author of the two policy briefings and wanted to do my best to represent my co-authors and ensure their work received maximum attention. Unfortunately, the entire summit was derailed by Trump leaving the summit early to bomb Iran. Considering the actual leaders’ summit is just two days long, this was devastating to any progress the G7 hoped to make that year, especially since most of the critical bilateral meetings with Trump (including Carney’s meeting to end the trade war with Canada) were planned for the day he left. This also diverted all the media attention, which normally focuses on the critical global issues the G7 is intended to address, towards speculation about what Trump was leaving to do.
The result of the Kananaskis Summit was no communiqué with specific commitments (as is typical), but six “joint statements” mostly consisting of vague, general statements of shared values agreed upon by all G7 leaders. None are particularly ambitious and, notably, none mention climate or the environment (which is an ideal collective action problem that the G7 is well-positioned to address).
It seems evident that any international forum in which the United States plays a crucial role is effectively handicapped at the moment. As such, I see no reason to waste time attending the Summits until they once again can be a productive forum. The commitments made by the G7 and G20 are not binding in any way, so they really only work if the leaders are willing to cooperate.
My colleagues at the G7 Research Group, however, still attended the Evian Summit.
Ominous photo one of my colleagues got at the press conference.
As a side note, it’s interesting to me that, although the vast majority of outputs from the Trump-era G7 summits have been pretty non-substantive, critical-mineral security (currently geopolitically dominated by China and Indonesia) has quietly moved to the top of the agenda for Western powers. This seems to be something the G7 countries are quietly panicking about and may be driving some of Trump’s Greenland obsession (a reactionary response to security briefings he is likely receiving). It is something to keep an eye on in the future: we will likely see increased stockpiling, expanded domestic extraction, and more diversified supply chains.
Etc.
Feeding the Chamber
I managed to help organise a diving trip to Portland with the Oxford scuba diving club. However, we had to divert from our planned two days in Weymouth (due to strong wind making boat dives impossible) and instead spend the first day at Swanage Pier, which is more protected. This diving spot is really shallow (about 3.5m), but has a ton of interesting marine life.
Due to the wind, visibility was not good (initially less than 1 m). Half of the group dropped out of the second dive at Swanage on the first day after being disappointed that they couldn’t see anything, but I persisted and stuck it out for another dive. This ended up being totally worth it: I saw a ton of stuff, including insane-looking spider crabs covered in seagrass, a Montagu’s blenny (not a tompot, unfortunately), tons of snakelocks anemones, some unfortunately named edible crabs, velvet crabs (which have bright red eyes), and, at the very end of the dive, a decently large cuttlefish. It was more like the cuttlefish found us: it clearly spotted us and deliberately came over to check us out, hovering right in front of us before moving on. It was a bit like being visited by a Cthulhu-headed alien.
Exciting.
The second day was slightly better, and we got out on the boat from Weymouth. The wind was still pretty strong, though, so we bounced pretty hard whenever the boat hit the waves. As a result, almost everyone got very seasick (except me and one of the more experienced divers). We ended up abandoning the original plan to find a wreck, and the only other non-seasick guy and I did a drift dive at a random location. This is when the current just kind of carries you along, so you end up somewhere different from where you started.
After getting back from the trip, I was really sore, especially in my elbows and knees. This was slightly concerning, since joint pain is a symptom of decompression illness, and I had technically dived more and for longer than anyone else on the trip. I wasn’t that concerned (I thought it was much more likely that I was bruised from the boat ride) until I was hit by a car on my way to work.
Hit by a van reversing at full speed during a three-point turn, right in the middle of the road (I was in the bike lane…). I was fine, but it broke my mudguard.
I have never in my life been properly hit by a car, so I was a bit concerned that there was something off with my reaction time (another DCI symptom). I ended up calling the closest hyperbaric chamber for advice, and they advised me to come in right away for treatment. The experience was a bit strange: the doctor on the phone said I “100%” had DCI and even said I was going to die if I didn’t see him. I thought this was odd, since the dives I did were very conservative and I had been bashed up pretty badly on the boat. But he was very insistent. He also told me not to tell any of the people I was diving with that I had been advised to go to the chamber, since “they would say I shouldn’t go”. Despite the strangeness, I went.
Five boring hours in this tube later: I had brought a stack of books to read, but it was very hard to focus with the oronasal oxygen mask I had to wear and the repeated tests they kept doing on me.
Interacting with the doctor in person was also quite strange. He ran a series of tests to look for neurological symptoms of DCI, but I kept passing them all—until I got to a balance test, where I had to stand on one foot with my eyes closed. I warned him that my balance is pretty bad at baseline, but he didn’t listen. So when I wobbled a bit, he wrote down “severe neurological symptoms of DCI”.
Fortunately, with a hyperbaric chamber, you can immediately tell whether you have DCI because the symptoms should resolve as the pressure increases, forcing nitrogen bubbles to shrink. In my case, my joint pain did not change at all, which meant it wasn’t caused by DCI (yay), but I still had to sit out the remaining five hours in the chamber.
When I got out of the chamber, he told me the secret to passing the balance test was to stabilise with my eyes open, then close them. I did this and obviously passed the test, so he said “my DCI was cured” by the chamber. My theory is that they need divers to use the chamber for DCI treatment so they can continue justifying its existence to the NHS.
Overall, it was a very strange experience. At least one benefit was that I got to feel fully oxygenated for basically the first time in my life (I have anaemia).
Usually I’m pretty sallow, but my face was noticeably flushed after the high-pressure oxygen.
Birthdays (cont.)
I played paintball as a belated birthday celebration and celebrated a friend’s birthday with a picnic.
Note: The mark on my hand is a 2 not a Z.
Whimsical sky at the Grandpont Nature Park.
Loops
I also noticed a fun maths/stats problem that some mystery person (likely one of my flatmates) had written on our blackboard. I managed to solve it (I think) by induction.
Induction proofs are really fun when you can find the pattern.