When the National Academies surveyed state transportation departments in 2024 about artificial intelligence in pavement condition surveys, twelve of the thirty-eight agencies that answered the relevant question said they were unsure whether artificial intelligence was being used on their own data. Eight said it was. Eighteen said it was not, a distribution that explains more about the state of the field than any capability demonstration.
NCHRP Synthesis 636 reached 43 of the 52 agencies it approached, a response rate of 83 percent, and it found a technology that is genuinely in production, genuinely consequential, and largely invisible to the agencies buying it.
The place AI already touches the asset
Automated pavement condition surveys are the highway sector’s biggest computer vision deployment by mileage. A van drives the network, cameras and lasers record two and three-dimensional surface images, and software converts those images into cracking, rutting and faulting quantities that feed the pavement management system. Those quantities decide which projects enter the program, which is covered in how state DOTs decide which roads to rebuild.
The synthesis found the analysis concentrated in remarkably few hands. All 38 agencies that described their collection arrangements reported using one of the same five vendors, and 58 percent contracted the collection and the analysis together. Where artificial intelligence enters the distress identification, it enters through the vendor’s product.
The answers about validation are the ones that should trouble a chief engineer. Asked how the techniques had been developed, trained and evaluated, 15 of 27 agencies said they were unsure. Five of 27 had compared the machine results against a manual survey. Twenty-one of 27 were unsure whether the process complied with AASHTO R 85, the practice the synthesis treats as the reference for automated cracking protocols, and two said it did. Sixteen of 24 were unsure which models were being used, and no agency reported using deep learning, a claim that sits oddly against a market in which convolutional networks are the default tool for exactly this problem. The likelier explanation is that the agencies cannot say what their vendors run.
An agency in that position has outsourced a measurement, not a task. The distress numbers still carry legal and budgetary weight, they still feed federal condition reporting, and the trend line across years still depends on the vendor’s model version staying comparable to last year’s.
What agencies have built for themselves
A parallel National Academies study on machine learning at state transportation departments collected case studies from five agencies, and its value is that the failures are written down alongside the successes.
Nebraska’s work is the cleanest. Roughly 2.5 million geo-coded images from the forward-facing cameras on a pavement profiler van were used to detect and classify guardrail, reaching 97 percent accuracy on detection and 85 percent on sorting three guardrail types, validated against a manually labelled set of 1,500 images, at about $25,000 per project. A companion model located 1,303 marked crosswalks statewide. The recorded limitation is physical rather than statistical: the van sees only what lies in its path.
California’s litter detection pilot reached roughly 50 percent accuracy at identifying actual trash and a 70 percent match between the model’s severity rating and a human assessor’s. The case study also carries the most useful sentence in the collection, which is that even a simple result such as 80 percent accuracy in computer vision is not easy to explain to the people who have to act on it.
Missouri’s traffic management centre offers the only figure in the set that measures displacement rather than accuracy. The share of incidents first identified by machine learning rather than by an operator rose from 42.3 percent in November 2021 to 54.1 percent in October 2022, when its video analytics were running at 88 percent true incidents against 12 percent false positives. Delaware reported around 90 percent accuracy across incident types. It said it is working toward a sub-3 percent error in speed estimation, and it reported observing around 15 percent errors in traffic volume on arterials, where signal operations complicate the count.
Iowa’s entry is the one a procurement officer will recognize. The agency considered deep learning models for incident and anomaly detection and ultimately adopted simpler threshold-based methods, because of the cost of scaling the sophisticated version to statewide deployment. The accompanying observation is the arithmetic every operations manager eventually meets: a model that is 99 percent accurate across a statewide sensor network still generates an unacceptable number of false alarms, because the denominator is enormous and the alarms land on a small staff.
The most-quoted result in the field rests on twenty-four drive-throughs
Adaptive signal control is where transportation artificial intelligence gets its reputation, and the source of the reputation is a nine-intersection pilot in the East Liberty district of Pittsburgh. The SURTRAC system, developed at Carnegie Mellon University’s Robotics Institute, treats each intersection as an independent scheduler that plans the passage of currently approaching vehicles, then tells its downstream neighbours what to expect.
The published results are strong. Travel times fell 25.79 percent, average wait time 40.64 percent, stops 31.34 percent, and computed emissions 21.48 percent, weighted across four periods of the day. The baseline was not neglect: eight of the nine intersections had been running coordinated-actuated plans generated in a commercial offline optimisation package and installed in early 2011, with actuated free mode outside the peaks.
The evidence base is smaller than the numbers imply. Performance was measured by timed drive-through runs of the twelve highest-volume routes, three runs per period per scenario, twenty-four runs in total, with travel data captured by a GPS application on a phone and post-processed to isolate the evaluation routes. The before runs happened in March 2012 and the after runs in June 2012, three months later, because of a delay in obtaining the legal agreement to take control of the signals; June volumes ran about 5 percent higher. The paper also reports where the system lost: performance deteriorated on four of the twelve routes during the morning peak, three of them along one circle, which the authors read as evidence of an unbalanced bias in the pre-existing timing plan.
None of that makes the result wrong. It makes the result a pilot, published as a pilot, in a planning and scheduling conference, and it is the strongest published evaluation the approach has. The gap between that and the way the percentages now circulate is the gap this section exists to name, and the same pattern recurs across transportation technology coverage.
The oldest artificial intelligence on the freeway is a rule base
Washington State has metered ramps with fuzzy logic since the mid-1990s, when the Washington State Transportation Center at the University of Washington built the initial system. FHWA’s 2019 review of artificial intelligence for management and operations records that the deployment had reached more than 200 ramps, that all new ramp meters in the state use the method, and that no other method is considered. The reason the agency chose it is recorded too: earlier algorithms needed substantial code to estimate the effects of candidate metering rates, and calibrating their parameters was difficult. Only 5 percent of the original system’s code was the algorithm. The other 95 percent was interfacing with data inputs and field controllers.
That ratio has not changed for the newer techniques, and it is the reason the case studies above read as they do. The same FHWA report supplies the standard caution about the learned alternatives with an example that transfers directly to this work: a network trained on side views of ambulances may fail on ambulances seen from the front. Pattern recognition is good at what it was trained on.
Which points at the real hazard for infrastructure work. The training labels come from the practice the model is meant to replace, so the model inherits that practice’s blind spots and reports them back with a confidence figure attached. The uses in modelling rather than measurement are treated separately in digital twins for highways and bridges. In measurement, five of twenty-seven agencies had checked their machine distress results against a manual survey. Until that number rises, the honest description of artificial intelligence in pavement management is a measurement system nobody has audited, and its outputs are being used to schedule capital work on a national network.