Search for AI in water management and the same article comes back a dozen times over. AI predicts demand. AI finds leaks. AI optimises treatment. AI decides when to irrigate. Every one of those statements is true, and together they are close to useless, because not one of them says where the line is.
The line is the part worth writing down. It is also the part almost nobody publishes, because it involves saying out loud that the fashionable tool is the wrong one for a large share of the work.
The Question That Decides It
Before reaching for either a solver or a learned model, ask one thing about the system in front of you: does its operating history contain the answer you need?
If it does, a model trained on that history can be very good indeed, often better than anything a physical model would give you for the same effort. If it does not, no volume of data rescues the situation, because the record describes a system that is about to stop existing.
That single question sorts most water problems correctly, and it is the reason our modelling practice is split the way it is.
Where AI Does Not Belong
Anything You Are About to Change
A new wellfield. A new trunk main. A new zone boundary. The next stage of pit dewatering. Each of these is a stress the system has never been put under, and that is precisely the case in which a model trained on past behaviour has nothing useful to say.
A model trained on historical hydrographs has never seen the wells you are proposing, so it has nothing to learn their drawdown from. A solver does not need to have seen them. That is the entire reason to build one. It is why groundwater modeling stays in MODFLOW and water distribution modeling stays in EPANET, however much telemetry the site happens to produce.
Anything Outside the Record
A design flood is normally larger than anything in the gauge record. A discharge consent asks what would happen under a load that has never been discharged. Both questions sit outside the training range by construction.
A learned model asked to extrapolate past its training range does not decline to answer. It answers fluently, it is wrong, and nothing in the output distinguishes that answer from a good one. This is why the hydraulics in our flood modeling work stay with HEC-RAS, and why process models carry the regulatory questions in water quality modeling.
Anything That Has to Be Defended Later
A design storm has to be reproducible by someone else, years later, regardless of how many storms happened to fall during your monitoring campaign. A drainage design a reviewer cannot reconstruct is not a design, it is a claim. The routing in stormwater and urban drainage work stays in SWMM for that reason, and the reason is procedural rather than technical.
Anything People Decided Rather Than Anything Data Revealed
Allocation rules, water rights and reservoir operating policy are not patterns waiting to be discovered. They are agreements. There is nothing to learn, because the rule is whatever the agreement says it is, and it changes the day the agreement changes. Basin allocation stays a rules-based simulation in WEAP for that reason, which is how we approach integrated water resources modeling.
Where Machine Learning Earns Its Place
In every case above the physical model keeps the question. The machine learning moves to the data feeding it. That is not a consolation prize. Most models that fail in the field do not fail on their hydraulics, they fail on their inputs.
Producing the Inputs a Solver Cannot Produce for Itself
EPANET requires the demand it is run against and cannot generate it. Pump scheduling, tank management and pressure control are only as good as that assumed demand, and a model trained on district metered area consumption with weather and calendar effects is a far better basis than a diurnal curve copied from a design manual. At basin scale, seasonal inflow forecasting from catchment rainfall and soil moisture signals helps an operator decide how much storage to hold back going into a dry season.
Downscaling and Bias Correction
Climate model output arrives on a grid far too coarse for a catchment decision, so downscaling learns the relationship between those large-scale fields and local station records. Forecast rainfall is bias corrected before it drives a flood model rather than argued about afterwards. Note the direction of travel in both cases: the learned model prepares the input, the physical model answers the question. That division is what our work with researchers and agencies in environment and research is built around.
Telling an Instrument Fault From a Real Event
A drifting sensor and a genuine hydraulic event look much the same in raw data. Water quality probes foul, and a fouled probe produces something that reads convincingly like a pollution incident. A model trained on each instrument's own behaviour separates the two, which is worth more than it sounds: it stops a fouled meter from being calibrated into a design as though it were truth.
The same method does different work in different places. On a process line it catches a fouling membrane or a drifting analyser well before a compliance sample does, which is how we approach monitoring for industrial clients. On a mine site it flags pore pressure readings that have started moving differently from the rest of their array. On a distribution network, anomaly detection on minimum night flow is how a burst gets found while it is still underground.
Why Gap Filling Has to Come Before Calibration
This is the part we would most like other people to steal, because getting it wrong is common and the damage is invisible.
Observation wells, SCADA feeds and gauge records lose stretches to power cuts, instrument faults and site access. Standard practice draws a straight line across the hole and carries on. The model is then calibrated against that repaired series.
Consider what has just happened. The interpolated stretch is now a calibration target, and it was never measured. The solver has no way of telling which points are observations and which are drawings, so it fits parameters that reproduce both with equal seriousness. The error does not stay where you left it either. It is absorbed into the fitted conductivity, or roughness, or storage coefficient, and from there it spreads into every prediction the calibrated model goes on to make.
The result is not a slightly worse model. It is a confidently wrong one, and the confidence is what makes it expensive. A model with visible gaps in its record invites scrutiny. A model calibrated against invented data looks finished.
So the reconstruction belongs before calibration, it belongs to a method you can describe and defend, and the reconstructed stretches should stay marked as reconstructed all the way through. Straight-line interpolation is also a model. It is simply a bad one that nobody writes down.
The One Setting Where the Learning Is the Main Event
There is an exception worth naming, because a rule with no exception usually means the rule has not been examined.
A commercial building produces a great deal of half-hourly meter data and very few questions that need a differential equation. That makes it the one setting where the machine learning is the main event rather than the support act. A model of each sub-meter's normal weekday, weekend and holiday pattern turns a leak into an obvious departure from expectation instead of a number somebody has to happen to notice. Cooling tower chemistry, cycles of concentration and pipe hydraulics remain arithmetic and physics, and we do not dress them up as anything else. That is the shape of the work on commercial buildings.
What to Ask Anyone Selling You AI for Water
If you are buying this rather than building it, five questions separate a considered approach from a rebranded one:
- Which part of this is learned and which part is solved? Anyone who cannot draw that line on their own system has not thought about it.
- What happens when the system changes? If the answer requires retraining on data that will not exist until after the change is built, the model cannot support the decision that authorises the change.
- Does my question sit inside the training range? Design events and proposed assets usually do not, and a learned model will not warn you when yours does not.
- How was missing data handled, and at what point? Ask specifically whether reconstruction happened before or after calibration. Silence here usually means a straight line.
- What was the model calibrated against, and what was the residual error? A model that has never reproduced a measured condition is an opinion with a mesh.
The Honest Version
Physics first, machine learning where it earns its place. It reads like a compromise and it is not one. It is a description of what the two tools are actually good at, arrived at by a team that holds a doctorate in each of the two disciplines rather than a strong opinion about one of them.
The uncomfortable half of that position is the half worth paying for. Being told which parts of your problem AI should not touch is more useful than a longer list of the parts it might. If you are scoping work of this kind, our water consulting and modeling services set out how we approach it, and the first thing we will do is tell you which side of the line your question falls on.
Where this ends up, in practice, is a system rather than a report: the model, the data pipeline that keeps it fed and the interface someone actually opens. That is what we mean by a water decision support system, and it is the form most of this work takes.
