EUROCONTROL PRC Data Challenge 2026 · ten European hubs
Taxi-Out, Measured
Every departure is logged twice: when it leaves its stand, and when it lifts off. This challenge erases the first, keeps the second, and asks for the gap between them.
DEPARTURES
Taxi-out · stand to wheels-up · ten hubs · 2025Each row is an airport, not a flight: its busiest departure runway, its typical (median) taxi-out and the time one departure in ten takes longer than, from 2.1 million 2025 departures. Status: de-icing where January 2026 brought snow, runway change where a runway's share of departures moved ten points or more from 2025, check records where off-blocks were copied from the timetable. The last row is the one the challenge is about.
One departure, seen from above
Follow it from its stand to the runway. The challenge erases the time at 1 and keeps the time at 3.
1 · Leaves the stand
??:??:??
Pushed back from its stand. This time is erased for every 2026 departure.
2 · Taxis and queues
? min
Taxi-out: the number each team must send, in seconds, for every departure.
3 · Wheels up
14:37:12
Take-off. This time stays in the file.
The task. For each of the departures from ten European airports in January and July 2026, the organisers erased moment 1 and kept moment 3. Teams send back the time between them. Everything else in the picture stays in the file with all its times: the parked aircraft, the arrivals, the queue, the runway in use. All of 2025, with both moments recorded, is there to learn from.
Not a forecast. These flights have long since flown. The question is how much the rest of the records give away about the moment each one left its stand, and some of those records turn out not to be measurements at all.
How a guess is marked
Each miss is squared: a guess 12 minutes out counts 144, one a minute out counts 1. The score is the side of the average square, in seconds, so lower is better, and a few big misses outweigh thousands of small ones.
Where we stand.
From stand to wheels-up
A departure's taxi-out starts when the aircraft leaves its parking position, pushed back by a tug or under its own power, and ends when its wheels leave the runway. In between it follows a taxiway to the runway, in winter perhaps stops at a de-icing pad, and waits at the holding point until the tower clears it to line up and take off.
At these ten airports the median is . The drive itself is a few minutes at 10 to 20 knots. The rest is waiting: for other aircraft on the taxiway, for a gap in arrivals when crossing a runway, and in the queue at the holding point, where departures leave one at a time with a wake-turbulence gap between them.
That queue is why the number matters. An airliner taxiing on its engines burns fuel for every minute it stands still, so minutes of queueing multiply into tonnes of fuel and CO2 across a day at a hub. Airports estimate taxi-out before each departure to plan when it will take off, and EUROCONTROL's Performance Review Commission tracks the excess over an unimpeded taxi as a measure of congestion. Its 2026 data challenge asks how well taxi-out can be recovered from the records.
What the data keeps, and the one thing it removes
The movements come from EUROCONTROL's Airport Operator Data Flow: every airport reports every arrival and departure, with its stand, runway and times. Taxi-out is not measured separately. It is take-off minus off-block: across all 2025 departures the subtraction matches the published figure to the second. The definition of the erased time is exact:
Actual off-block time: the date and time the aircraft has vacated the parking position (pushed back or on its own power).
EUROCONTROL-SPEC-175, Airport Operator Data Flow, ODI 11So the 2026 file blanks the off-block time together with taxi-out, and keeps everything else: every take-off time, every stand and runway, and arrivals with their in-block times. The prediction runs backwards. The take-off time is known, and the task is to reconstruct when the aircraft left its stand.
Is the off-block time not also in the flight plan? A copy of sorts is. Each departure is linked, where possible, to the record of EUROCONTROL's Network Manager, which handles every flight plan in European airspace, and that record keeps its own off-block time into 2026. But it is a separate measurement of the same moment, and it agrees with the airport's within a minute for only 9% (Istanbul) to 26% (Barcelona) of departures. A strong clue, not the answer.
Blanking one column also leaves the rest of the airport in view. For every 2026 departure the file shows which aircraft blocked in at the same stand and when, how many departures took off from the same runway in the minutes before and after, how many aircraft had pushed back without yet taking off, and which runways were in use that hour. Each of these reads the queue the departure was part of.
Ten airports, ten layouts
Taxi-out is first of all geography. Heathrow's two runways run either side of a crowded central terminal area, and its median taxi-out is the longest here. Schiphol's Polderbaan (36L) is several kilometres from the terminal, and departures on it take about seven minutes longer than on runway 24. Zurich chooses its runways by the clock, under noise rules and a German airspace ordinance, so the time of day largely decides the taxi route.
EUROCONTROL's own reference for an unimpeded taxi works at this level: the 10th percentile of taxi-out for each stand and runway pair over twelve months, valid only where at least ten flights fall at or below it. The model, a set of gradient-boosted decision trees trained on 2025, starts from the same reference and learns what adds to it.
A score that punishes the worst miss
The squared score has a name, root mean square error: square each miss in seconds, average, take the square root. A ten-minute miss costs a hundred times a one-minute miss. Here are the same predictions measured two ways, on the January and July 2025 flights held back from training to rehearse on, beside a guess that knows nothing about each flight:
Sort those held-back flights from worst prediction to best and add up their squared misses. If errors were spread evenly the curve would follow the diagonal.
The worst predictions cluster in one group: about fifteen departures in a thousand whose airport record never linked to a flight plan. They have no Network Manager off-block, no callsign, no airline. Most of them taxi like anyone else. A few are recorded as taxiing for hours.
The likeliest reason they are unlinked points at those few. Every linked record has its off-block within an hour of the flight plan's last planned off-block (the flight plan's window, below), so a record further than that from its plan stays unlinked. An off-block copied from the timetable for a flight that left hours late, or logged a day early, is exactly such a record. Two Paris departures of this kind carry about a quarter of all the error on their own: each took off on time, its off-block was logged about a day earlier, and nothing in the file flags it.
Most of those very large misses trace back to records that are not measurements.
Records that are not measurements
At Rome Fiumicino, the recorded off-block time of of departures equals the scheduled departure time to the minute, against 1 to 6% at the other airports. For those flights the published "taxi-out" is take-off minus schedule. A flight that left four hours late shows four hours of taxi-out.
Most of these records are harmless. The flag is commonest on flights the Network Manager saw leave within a few minutes of schedule, where take-off minus schedule is close to a normal taxi anyway. But on the rare delayed ones it creates the largest errors in the dataset: of Rome's departures with a flight-plan record and a recorded taxi-out over an hour are flagged rows.
The flag follows the data source more than the aircraft: some airlines have it on half their flights and others almost never, it peaks on late-evening schedules, and it is three times as common when the filed flight plan was never updated. The records look filled in from the timetable when no actual time was captured.
The 2026 file keeps the schedule and the take-off time, so the schedule gap is known for every 2026 departure. Whether a given record was filled in is not. The model hedges between the two readings:
where p is a classifier's estimate that the off-block was copied from the schedule. Take a flight that lifted off three hours after its scheduled time, where records like it turn out to be timetable copies three times in ten. Answer three hours, and seven times in ten the miss is nearly three hours. Answer a normal fifteen-minute taxi, and three times in ten the miss is nearly three hours, which the squared score still punishes hard. The blend, about an hour, is never right, but it costs the least on average.
Among flights with no flight-plan record the same pattern reaches at Rome, and a rarer one appears at several airports: an off-block logged a day before take-off, so the taxi-out reads about 24 hours plus a normal taxi.
At Rome the two patterns sort cleanly by how late the flight took off. The later it left, the likelier its off-block is the schedule, until past eight hours every one is. Between 14 and 26 hours late only two kinds are left, never a normal taxi: copied from the schedule, or logged a day early.
So for those flights the model stops guessing freely. Past six hours late it takes the share copied from the schedule straight from 2025's records, and between 14 and 26 hours late it weighs the schedule gap against a day plus a normal taxi.
The flight plan's two-hour window
A record that does link to a flight plan carries one more clock: the plan's last off-block time. On all such departures of 2025 the airport's off-block lies within an hour of it, never more than seconds away. The two tables look to have been joined on that window, which would also leave unlinked every record further out, the no-record flights above. Whatever the reason, it bounds the answer, and for the departures whose schedule falls outside the window, the off-block cannot have been copied from the schedule at all.
This bound was first published by another team in the challenge (EnioAguiar/prc-taxiout-2026), and confirmed here on the 2025 records before use.
Watching the push-back
Aircraft broadcast their position by radio, on the ground as well as in the air, and volunteers' ADS-B receivers record it; adsb.lol publishes their archive under an open licence. Where a receiver near the airport can see the stands, the trace shows the aircraft parked and then moving: the push-back itself, a median 35 seconds from the airport's record, against three minutes for the Network Manager's. That is visible for one 2025 departure in eight and one 2026 departure in five, best at Munich, Amsterdam and Heathrow, never at Istanbul or Paris. Most of the others are first heard already taxiing, which still shows the taxi took at least that long.
A small second model learns the first one's error from what the receivers saw, using predictions made for 2025 days the first model never trained on, and corrects the 2026 predictions with it. Scored day by day on days it never saw, it took the held-back error from 316.7 to 308.8 seconds; on the test months the upload carrying it gained 12 seconds.
The receivers have a blind spot that matters in winter. On a congested day an aircraft can stand for ten minutes or more away from its stand, at a de-icing pad or in a queue, and a trace that only sees "parked, then moving" dates the push-back from there. Where the airport's departures were taking over 20 minutes, the observed push-back fell a median two to four minutes after the airport's record, against a quarter of a minute on quiet days. Flights with a Network Manager record keep its off-block as an anchor; flights without one do not.
What moved the score
Each flight strip is one change to the model, scored on the held-back months and kept only when the gain held up under resampling (a paired bootstrap). They are grouped by idea, largest first; the number on each strip is the order it was added.
Tried and rejected
Each was scored against the model it was tried on, so its baseline is not always a step above. Most made the score worse. Two improved it: one by too little to trust, one in a way 2026 would not keep.
Why the steps that worked, worked
A squared score is minimised by the average outcome given what is known. When a flight's record can be one of several kinds, that average splits by kind, and each part is simpler to learn than the whole:
The copied part is known exactly, so the error sits in the probabilities. Getting Rome's right was worth more than any feature. The flight-plan window works for a geometric reason: moving a prediction onto a range that contains the truth can only bring it closer, flight by flight. Starting the trees from the Network Manager's own taxi time leaves them only the difference to learn. A second kind of boosted tree, CatBoost, averaged with the first, cancels part of each one's noise.
Queue counts from queueing theory, which predict taxi-out well in the research literature, added nothing here: the Network Manager's off-block already contains the outcome of the queue. The full survey is in the repository's RESEARCH.md.
January 2026 was a winter month
The held-back months stand in for the test months only as far as the two behave alike, and in 2026 they did not quite. Aircraft de-iced on the stand wait before off-block, which does not count. Aircraft de-iced at a remote pad do it during taxi-out, which does. The data shows which airports do which: in 2025 snow, taxi-out stretches at Paris, Munich and Istanbul, while at Heathrow the delay lands before off-block instead.
The held-back months could not show this. January 2025 was mild at most of these airports; January 2026 was not.
Amsterdam's first week of January 2026 shows what that does to the flights with no flight-plan record. From the 3rd to the 9th, the departures around them were taking 40 minutes to an hour from off-block to take-off, against about 15 on an ordinary day, and the test months hold over 200 such flights on those days. Their model, built from the flight's own few fields, had been predicting about 20 minutes.
Yet a flight with no record taxis like the flights around it: in 2025, outside Istanbul, 0.95 to 1.0 times the median of the departures at its airport in the same half hour, on quiet days and the worst alike. That median can be read for every 2026 flight from the take-off and Network Manager off-block times the file keeps, so where it exceeds 25 minutes, each no-record flight is now predicted at least 0.95 times it. It changed 431 flights and took the test score from 275.0 to 274.1 seconds.
Runway use moved as well. Heathrow ran easterly, taking off from 09R, far more often in 2026. Paris used 27R for departures in July 2025, apparently during works on the northern runways, and not at all in July 2026. The model learned its runway and queue effects from 2025's mix.
Cleared for take-off
Each upload is scored on the January and July 2026 departures, and the leaderboard keeps a team's best.
The final ranking takes one more file per team, uploaded once, that covers February and June 2026 as well. It is scored on January and July, on February and June, and on all four months together, alongside a review of the code and documentation.
Limits
- Winter. The held-back January was mild, so the model's handling of de-icing days rests on 2025's few snowy days and on the test months themselves.
- What the receivers miss. ADS-B dates the push-back for about one 2026 departure in five, and for none at Istanbul or Paris.
- Stands that opened in 2026. Munich's stands 103 to 108 and a handful of new Frankfurt positions (consistent with its Terminal 3 opening) never appear in 2025. They borrow a reference from neighbouring stands or their runway; only the 2026 score can say what that is worth.
- Rome. Still the airport with the largest error per flight. For flights with a flight-plan record, the classifier decides between schedule and normal taxi without a rule as clean as the no-record one.
- Day-early records. An off-block logged a day before an on-time take-off, like the two Paris flights, leaves nothing in the file to flag it.
References
- EUROCONTROL Performance Review Commission and OpenSky Network, PRC Data Challenge 2026: overview, data description and ranking rules. prc-data-challenge-2026.netlify.app
- EUROCONTROL, Additional taxi-out time performance indicator document, Edition 01.00, March 2023. Reference taxi time as the 10th percentile per stand and runway. ansperformance.eu
- EUROCONTROL, Airport Operator Data Flow, Data Specification, Edition 00-11. Definitions of AOBT, ATOT and the de-icing indicator. eurocontrol.int
- EUROCONTROL, Specification for Airport Collaborative Decision Making, Edition 1.0, January 2025. TOBT, TSAT, TTOT and pre-departure sequencing. eurocontrol.int
- EUROCONTROL, Network Manager DPI Implementation Guide, Edition 2.700. How A-CDM airports send off-block and taxi times to the Network Manager. eurocontrol.int
- EUROCONTROL, RECAT-EU wake turbulence categorisation and separation minima. eurocontrol.int
- Aviation Intelligence Portal, slot tolerance window (−5 to +10 minutes around a CTOT). ansperformance.eu
- Heathrow Airport, runway alternation and Aircraft De-icing Plan. heathrow.com
- Amsterdam Airport Schiphol, noise and runway combinations. Preferential runways 36L and 24 for take-off. schiphol.nl
- Flughafen Zürich, operating concepts. Runway use by time of day. flughafen-zuerich.ch
- Fraport, runway system and operating hours. Runway 18 West for southbound take-offs only. fraport.com
- adsb.lol, global history archive, 2025 and 2026: ADS-B traces from volunteer receivers, available under the Open Database License 1.0. Contains information from adsb.lol. github.com
- Iowa Environmental Mesonet, ASOS / METAR archive. Weather at the ten airports, January 2025 to July 2026. mesonet.agron.iastate.edu
- ICAO, Doc 4444 PANS-ATM: wake turbulence categories and departure separation; ICAO Doc 9432, Manual of Radiotelephony.
- I. Simaiakis and H. Balakrishnan, A queuing model of the airport departure process, Transportation Science, 2016. Taxi-out as unimpeded time, surface interaction and runway queue. mit.edu
- A. Swatowska and L. Gabagnou, A generalisable machine learning framework for taxi time prediction at A-CDM airports, SESAR Innovation Days 2025. CatBoost on 4.1 million flights at six airports. sesarju.eu
- L. Grinsztajn, E. Oyallon and G. Varoquaux, Why do tree-based models still outperform deep learning on tabular data?, NeurIPS 2022. arxiv.org
- L. Prokhorenkova et al., CatBoost: unbiased boosting with categorical features, NeurIPS 2018. arxiv.org
- A. Niculescu-Mizil and R. Caruana, Predicting good probabilities with supervised learning, ICML 2005. Calibration of boosted trees.
- Other teams' public write-ups, read for ideas and credited where used: EnioAguiar/prc-taxiout-2026 (the flight-plan window, and ADS-B ground tracks as a source of off-block times), radekacar/joyous-rainbow (linear leaves, tried and rejected here), skylinkapi/prc-data-challenge-2026-kind-mango (a similar 24-hour reading at Rome, found here independently). github.com