Shipment Delay Prediction Model — Ground Logistics Operator, Chicago, USA
The Situation
A mid-size ground logistics operator running a 14-state distribution network across the Midwest and Southeast had watched its late delivery rate climb from 8% to 13% over 18 months. The causes were not mysterious. Winter weather on the northern routes, driver availability at 4 high-volume hubs, and capacity crunches at the Memphis and Atlanta sorting facilities accounted for most of it. The operations team could name the problems. They could not see them coming.
Dispatchers were working reactively. A shipment would fall behind at a hub, a driver would call in, weather would close a route, and the team would scramble to reroute or reallocate. By the time the scramble started, the delivery window was usually already gone. Customer penalty clauses on late deliveries were costing the network $1.4 million annually across its top 40 accounts.
Four years of shipment records existed: origin and destination, scheduled and actual delivery times, hub dwell times, driver assignment histories, route-level weather event logs, and vehicle maintenance flags. The data was in three separate systems and had never been used for anything beyond monthly reporting.
What Amorisoft Did
Amorisoft ran a 3-week data extraction process across the network's three operational systems. Four years of shipment records covering 2.3 million individual shipments were cleaned, linked at the shipment level, and structured into a training dataset.
The prediction target was binary: would a shipment arrive late relative to its committed delivery window. The model was trained to make this prediction at the 36-hour mark before scheduled delivery, chosen because dispatcher analysis showed 36 hours was the minimum time needed to execute a meaningful reroute or driver reallocation on most lane types.
Feature engineering drew on 38 variables. The 4 strongest predictors were hub dwell time at the originating facility relative to the lane average, weather forecast severity score for the primary route in the 48 hours following pickup, driver assignment recency on the specific lane, and vehicle age combined with maintenance flag status. Shipment weight and declared value, which the operations team had assumed would matter, ranked outside the top 15.
The model outputs a delay probability score for every active shipment at the 36-hour mark. Shipments scoring above 0.68 are flagged to the relevant dispatcher automatically through an integration with the network's existing dispatch platform. Dispatchers see the flag, the contributing risk factors, and a suggested action from a predefined playbook: reroute, reallocate driver, or notify customer proactively.
The suggested actions are not automated. Every intervention is a dispatcher decision. The model surfaces the risk and the options. The human makes the call.
Route-level calibration was applied after initial training because delay patterns on northern winter routes differed significantly from southeastern routes in ways the combined model was smoothing over. Calibrated route clusters brought overall prediction accuracy from 71% to 77%.
Results
Late delivery rates across the 14-state network dropped by 39% in the 5 months following full deployment. Delay predictions were confirmed accurate in 77% of flagged shipments on dispatcher review. Dispatchers intervened on 22% of flagged shipments, the remainder were monitored and resolved without action. Of the shipments where dispatchers intervened, 81% were delivered on time. The $1.4 million annual penalty exposure across the top 40 accounts dropped to $490,000 in the first full year post-deployment.
