
A wave of AI slop is coming for climate science—the best defence is to retain and upskill our scientists
Artificial Intelligence (AI) and Machine Learning (ML) tools are fundamentally reshaping climate science. They have changed the way we code, simulate future climate, and gather and interpret climate observations. Used without sufficient scrutiny, these tools can inject errors into climate information and the decisions built on it. The solution, though, is not to resist AI, but to retain and upskill climate scientists so they can better integrate and interpret these tools.
By Taimoor Sohail (University of Melbourne) and Sanaa Hobeichi (UNSW Sydney). Monday 31 August, 2026.
AI is already transforming climate research workflows
Let’s start with the obvious – corporate LLMs like ChatGPT, Claude, and Google Gemini have integrated themselves into many of our research workflows, as sounding boards for research or grant ideas, or as tools to streamline emails – if you aren’t concerned about the privacy and IP considerations, that is. This year, coding agents have joined the army of programmers building our ocean and climate models. Many codebases on GitHub now count Claude Code or Codex among their top contributors. These developments bring their own threats to our scientific understanding and tooling. Do we really understand the minutiae of how Codex produced the analysis script that generated the figure in a paper draft? Who owns ideas that are spun up through a deep literature review with the help of Claude? These concerns are real, but they are not the focus of this piece. A less visible – and potentially more consequential – shift is the growing use of purpose-built AI/ML models within the systems that produce climate evidence and inform decisions. A useful way to understand this shift is to ask where AI sits in the chain that turns weather and climate information into action.
AI is entering the climate-to-impact pipeline
The process by which climate information is translated to quantified risks and impacts is central to climate research. Traditionally, this pipeline entailed climate scientists using well-established physical models, observing systems and services to connect knowledge of the atmosphere and ocean to real-world decisions. At its simplest, observations and physical models would produce weather forecasts and climate projections. These outputs would then be translated into hazards, combined with exposure and vulnerability, and used to estimate impacts, damages and losses. Such results inform warnings, operational planning and long-term adaptation. Insurers, renewable-energy operators, land and water managers, public-health agencies and emergency services all depend on different versions of this chain.

AI has, to varying degrees, entered this pipeline. For example, AI is useful in cases where physical simulations are too costly to repeat at the resolution or size required, where relationships are complex or only partially represented by conventional mathematics, or where large observational and model datasets contain patterns that are difficult to extract by other means. AI weather forecasting is already operational at leading centres: the European Centre for Medium-Range Weather Forecasts now runs its Artificial Intelligence Forecasting System alongside its physics-based forecasting system. AI-assisted quality control, data reconstruction, emulation and downscaling are established research areas and are increasingly used in practice. Even there, however, observational scarcity, validation across regions, rare extremes and future climates remains an active challenge.
Further down the pipeline, AI can be used to translate climate information to risks and impact. Such AI-based models of exposure, vulnerability, impacts and damages are less mature and highly sector-specific. A major reason is that impact data is often fragmented, inconsistent, sensitive or proprietary, particularly for health impacts and for indirect, long-term, and non-market losses.
Framing climate research as an information-to-impact pipeline exposes precisely why the implementation of AI needs to be done carefully, and with guardrails. AI has moved more quickly into operational use in weather forecasting, where data and verification systems are comparatively standardised, while downstream impact and loss models remain more experimental. The reliability of the overall impact pipeline therefore relies on accurate, accountable implementation of AI into each component of the chain. The weakest link can muddy the entire system.
AI brings long-standing scientific choices into sharp relief
Every quantitative study contains judgement calls. Researchers choose which data to include, how to treat missing values, how to re-grid information, which baseline and evaluation metrics to use, how to aggregate results across time and space, and how to represent uncertainty. These choices arise in physical modelling, statistics, epidemiology and economics as well as in AI. We should not imply that AI invented scientific judgement or the obligation to disclose it.
The AI-specific problem is that these choices can become more numerous, less visible and easier to propagate. Important decisions may sit several layers upstream of the final result: in data preprocessing, resampling, a loss function, a training and testing split, a default setting, or the way outputs are aggregated for a map. They may be discoverable only in a script, if they are documented at all. Once a dataset or trained model is shared, the same assumptions can travel rapidly into many downstream studies and services.
Take a recent paper one of us co-authored on quality-checking Antarctic Ocean salinity data using a neural network. The headline result is a genuine win for ML tools: comparing ship, float and seal-mounted sensor data using neural networks, we found that seal-mounted sensors carry a consistent salty bias that survives existing quality control. That is the kind of systematic observational bias these tools are good at catching. But getting to that result required building a synthetic dataset from a high-resolution ocean model, to account for the fact that ship, Argo and seal data are not collected in the same places or seasons. As part of that process, certain model grid cells had to be duplicated. We reported this ourselves in the paper: depending on the data source, between roughly 7% and 17% of synthetic profiles ended up duplicated this way. It would have been entirely possible not to notice the implications of this issue — nothing about the pipeline forces you to check whether you are altering the weighting or effective sample size through a resampling step three layers upstream of your headline number.
We raise this not to claim we always get it right – but because we have direct experience of how many small, individually defensible choices go into an ML pipeline, and how easily any one of them could go unexamined without the specialist statistical expertise and judgement needed to recognise that it requires scrutiny. That is the real risk: not an unscrupulous researcher tweaking inputs to get the answer they want, but a competent, well-intentioned team that simply doesn’t happen to look closely enough at one particular step in a long pipeline – a duplication scheme, a re-gridding choice, a train/test split – before publishing the result that depends on it.
A second example comes from recent work benchmarking AI methods for precipitation downscaling over Australia. How well a model appeared to perform changed depending on whether it was evaluated for total rainfall, spatial patterns, the seasonal cycle or long-term trends, and depending on the region being assessed. No model performed consistently well across every benchmark and Australian subregion. The lesson is not that the models failed. It is that selecting a model also means deciding which aspects of rainfall matter most: its suitability depends on the intended use, the metric, the spatial scale and the rainfall characteristic relevant to the decision.
Multiply this across a field that is rapidly adopting these tools, often without every judgement call being fully documented, and you get a slow drift — a gradual loosening of the link between a dataset and the assumptions baked into it. Climate science has a strong culture of sharing data, code and models. This accelerates discovery, but openness without provenance can also spread poorly understood assumptions: sharing code is not enough if the judgements embedded in a pipeline remain invisible. Researchers therefore have an obligation to make visible the methodological choices that could materially alter a result, its uncertainty, how estimated impacts are distributed across places or populations, or its fitness for a particular use. For material choices, researchers should explain the rationale, identify plausible alternatives and test whether their conclusions are sensitive to them. A downstream researcher building on that dataset or model may otherwise have no way to know which choices were scrutinised and which were defaults left untouched.
Even worse, the expertise needed to catch this kind of problem can be so specialised that identifying its source requires a concerted review effort. By the time that happens, the dataset or model may already be embedded in follow-up studies.
Not enough of us are ready for this shift
The key here is not to turn away from using these methods. The principles of scientific scrutiny are familiar, but AI can make modelling choices harder to see and easier to propagate. Instead, we need to retain experienced scientists and train more researchers in this space: people who can critically consider the physics and statistics underpinning an ML framework, understand its limitations and potential failure modes, and judge whether its data, architecture and evaluation are appropriate for the intended use. There are already ML-enhanced datasets and models in use. Can we, as a community, be confident that they have been rigorously vetted, clearly documented and used within their limits?
The job losses across the climate and environmental sector this year could hardly come at a worse time. If this is the risk – not deliberate misuse of AI, but well-intentioned use by researchers who haven’t been given the time, training, or institutional support to interrogate it properly – then losing experienced people now, just as the tools are becoming ubiquitous, takes us in the wrong direction. Retaining expertise and providing the time and structures for careful review are essential if AI is to strengthen, rather than weaken, climate science.