← SkySwap

How SkySwap works

SkySwap estimates disruption risk — the chance a flight is cancelled or arrives 15 or more minutes late — for U.S. domestic flights. It combines twelve months of government performance records with live conditions on the day of travel, and it publishes its validation results on this page so you can judge how well the model actually performs.

Data sources

SourceUsed forCoverage
U.S. Bureau of Transportation Statistics (BTS) On-Time Performance historical cancellations, delays, taxi times June 2025 – May 2026 · 151,682 flight-route records · 7,090 routes · 356 airports
NOAA Aviation Weather Center (METAR) live weather at departure, en route, and arrival live government feed
FAA National Airspace System status live ground stops and ground delay programs live government feed
NOAA/NWS point forecast + aviation TAF forecast outlook at departure and arrival for travel dates up to 7 days ahead government forecast; display context only — never an input to the score

The risk model

Every flight gets a base score from its historical record:

score = cancelPct × 3.0 + delayPct × 0.45 + avgDelay × 0.30

For travel dates up to 7 days ahead, SkySwap also shows a forecast outlook at the departure and arrival airports (NOAA/NWS point forecast, plus the aviation TAF within its ~30-hour window). This is display context only: forecasts are not validated the way the historical model is, so they never move the score. Beyond 7 days no airport-level government forecast exists, and the app says so rather than inventing precision.

The score is clamped to the range 4–100. On the day of travel only, a live adjustment is added on top of the base score:

live adjustment = (taxi-out − 15) × 0.8 + weatherImpact × 0.25

Per-flight history is blended toward the route average with weight n/(n+10) (empirical-Bayes shrinkage), so thin samples cannot produce inflated scores — a flight with only a handful of observed operations leans mostly on its route’s overall record.

Scores map to three bands: Low < 35, Moderate 35–59, High ≥ 60.

To be plain about it: this is SkySwap’s own model. It is not an industry standard, and no airline, regulator, or standards body has endorsed it.

Validation

The scoring method is validated out-of-sample against the most recent published BTS month the model has never seen. In the current experiment, the shipped model (trained on the 12 months through May 2026) was evaluated on 606,676 U.S. domestic flights from June 2026 (99.9% coverage). The overall June disruption base rate was 27.2%.

Two notes so these figures can be read correctly. First, this experiment is reproducible: the scoring replay ships in the repository (validate_holdout.cjs) and runs the app’s exact formula against the public BTS file for June 2026 — a month entirely outside the shipped model’s training window. The shipped dataset itself is likewise re-derivable: verify_data.cjs re-tallies the raw BTS files from scratch, checks every flight, route, and airport entry the app ships, and prints the raw record counts behind any single flight’s numbers. When the model is next rebuilt and absorbs June, the experiment is re-run against the next published month and these tables are refreshed. Second, absolute disruption rates vary with the season — the June 2026 base rate (27.2%) is well above the 20.9% measured in the previous run of this experiment on April 2026 data — so the percentages below describe June specifically. What the experiment validates is the ordering and separation of the score bands, which held in both months.

BandDisruptedCancelled
Low <35 25.6% 1.3%
Moderate 35–59 43.0% 5.0%
High ≥60 49.9% 9.0%
Score rangeFlightsDisruptedCancelled
4–14182,12417.3%0.6%
15–24271,16328.2%1.3%
25–3498,81133.9%2.6%
35–4435,67941.0%4.4%
45–5915,38747.5%6.5%
60–742,84449.8%8.5%
75–10066850.0%11.1%

Disruption rises monotonically with score across every range in the table. Flights scored High were disrupted at 1.8× the June base rate and cancelled at roughly 7× the rate of Low-scored flights.

What the data does not cover

The BTS On-Time Performance dataset is a legally mandated near-census, not a sample — but its scope is defined by regulation, and that scope has edges worth stating plainly.

Known limitations of the model

Independent assessment: the U.S. DOT Inspector General audited this dataset in 2024 (report AV2025003) and confirmed BTS verifies its accuracy through 117 automated checks, while noting that completeness and consistency checks could be strengthened. DOT relies on the same filings to penalise airlines for chronically delayed flights.