FernWX Lab Mode is a Research Preview, and can be disabled at any time. New sessions will default to the standard FernWX experience.FernWXLab
FernWX · Scientific methodsGeneral documentation

For National Weather Service science and operations officers

Forecast reconstruction
and observational verification.

FernWX retains successive forecasts and associated observations so that forecast evolution can be examined after an event. The analytical unit is a saved location, a valid time or event window, and the information captured by a specified earlier cutoff.

This note describes the archive, its comparison methods, and the limits relevant to retrospective case studies and forecast-process review. FernWX is an independent project, unaffiliated with NOAA or the NWS; official forecasts, warnings, and operational decisions remain with the responsible agencies.

Methods described

The retained record

Collection services retrieve provider products and preserve successive forecast records. Normalized values support comparison across sources; retained source payloads and identifiers support examination of the original product. A later collection does not overwrite an earlier forecast. A collection record is a sample of a provider product, however, and does not necessarily represent a new forecast issuance or model cycle.

Valid time
The instant or interval the forecast describes. Accumulations require their full period definition.
Issuance time
The provider’s product timestamp, where available. Its meaning depends on the source.
Capture time
When FernWX retrieved the product. This constrains whether that record was available in the archive at a historical cutoff.
Observation time
When a measurement or reported phenomenon occurred. Receipt time and correction time may be later.

NWS issuance timestamps can identify a product issuance. WeatherKit’s stored timestamp derives from its retrieval-side read time; Open-Meteo products without an issuance timestamp use retrieval time as a fallback. These clocks cannot be treated uniformly as numerical-model initialization times. Lead time measured from capture must be distinguished from lead time measured from issuance.

Repeated retrievals may contain identical values. Payload and normalized content hashes help distinguish repeated samples from changed content. A hash establishes identity of the retained representation; it does not establish measurement accuracy or independent authenticity of the source.

Observing systems and products

The archive brings together several kinds of meteorological evidence. Their spatial support, time support, and physical meaning remain source-specific. Availability varies by location, date, collection configuration, and retained coverage.

Source families and their interpretation in a case study
Source familyRecord and interpretation
Point forecastsNWS and retained provider forecasts, including WeatherKit and Open-Meteo. These are forecast products for a coordinate; they are not independent ensemble members or station measurements.
NWS text and hazard productsArea Forecast Discussions (AFDs), alerts, and associated issuance metadata. Office discussions provide regional reasoning; their prose is not a calibrated probability at every point in the forecast area.
Surface observationsOfficial aviation stations, mesonets, and other networks received through NWS, aviation, MADIS, and archival feeds. Station identity, observation time, variable, quality information, and siting class determine how a value can be used.
Upper-air observationsRadiosonde profiles obtained through the Iowa Environmental Mesonet (IEM), with NOAA’s Integrated Global Radiosonde Archive (IGRA) supporting historical coverage. Station and launch time define the profile.
Satellite and lightningGOES Advanced Baseline Imager (ABI) brightness temperatures; Geostationary Lightning Mapper (GLM) optical flashes; and, where collected, National Lightning Detection Network (NLDN) pulses. These are distinct measurements with different detection characteristics.
Event and environmental contextLocal Storm Reports, SPC outlooks, NHC advisory products, and USGS water observations. Reports, forecasts, and gauge measurements keep their original role; proximity alone does not establish a common storm or drainage basin.

One station can arrive through several feeds

Canonical surface records distinguish the physical station from the transport that delivered its report. The same aviation observation received through NWS, aviation, MADIS, or IEM is not an additional independent sample. Report variants and corrections retain their provenance. Association with a saved location expresses relevance, not co-location of the instrument with that coordinate.

MADIS quality flags describe data checks; they do not certify exposure or anemometer height. Official aviation, professional mesonet, specialized infrastructure, and community stations require different representativeness judgments. No universal adjustment to a 10 m wind is assumed.

Reconstruction at a cutoff

Forecast Moment holds the event window fixed while the reviewer changes one to four earlier capture cutoffs. For each provider, it selects the latest eligible, coherent hourly run captured by that cutoff, rejects inconsistent issuance clocks, and shows the retained hours overlapping the event. Missing hours remain missing; older runs do not fill gaps in the selected run.

Illustrative reconstructionSame event. Two information cutoffs.

Event window: 18:00–21:00 UTC

12:00 UTC cutoff
Inspect the latest eligible forecast already captured by 12:00 for the 18:00–21:00 event.
16:00 UTC cutoff
Repeat for that same event using records captured by 16:00. The forecast may have changed or remained identical.

A product issued at 11:30 but first captured at 12:10 is unavailable to the 12:00 reconstruction. This example contains no observed weather data.

The accompanying AFD is drawn from the location-relevant NWS office and must also have been captured by the selected cutoff. Original text and parsed sections preserve the discussion’s context. Later observations and reports can describe the outcome, but must be kept separate from what was available before the event.

This reconstructs the information retained by FernWX. It does not establish everything available to a forecast office, what an individual forecaster saw, or when a provider first made an uncaptured product public. A later historical import cannot retrospectively establish earlier availability in this archive.

Forecast verification

For a scalar variable with a comparable reference, FernWX reports signed error and absolute error for paired values. Mean absolute error (MAE) summarizes the magnitude of paired errors in the variable’s units.

Paired errorei = fi − oi

Mean absolute errorMAE = (1/N) ∑i=1N |fi − oi|

Here f is the forecast, o is the selected reference value, and N is the number of usable pairs. Positive signed error means the forecast exceeds the reference.

The reference changes the meaning of the score

Station reference
The forecast is compared with the selected station observation. Default NWS-reference comparisons require a station identifier. Station separation, elevation, exposure, and time offset still affect representativeness.
Provider self-reference
The forecast is compared with that provider’s later reported conditions. This measures consistency with that product; the reference is not necessarily an independent instrument observation.
Combined score
Where selected, the application weights station-reference MAE at 60% and self-reference MAE at 40%, requiring both components. These are application weights, not empirically calibrated estimates of forecast skill.

Pairing and sample selection

The single-target verification view defaults to forecast valid times within 30 minutes of the target and observations within 60 minutes, choosing the nearest available reference in time. Its response reports the tolerances and observation offset. Such a pair is not necessarily simultaneous and can be unsuitable for a rapidly changing convective event.

Recent source ranking uses the latest captured pre-valid hourly forecast for each provider and valid time, pairs the nearest reference within 60 minutes, and defaults to a 72-hour lookback. Because the selected lead can differ among providers, this is not a comparison at a common fixed lead. Five usable pairs permit a measured ranking; below that threshold, ordering falls back to a predefined source order. That minimum is an interface rule, not a statistical adequacy criterion.

A rigorous intercomparison needs a common set of valid times, explicit lead bins, the same reference definition, and coverage counts for each provider. Serial correlation and shared upstream inputs reduce sample independence. Reported standard errors do not by themselves establish a significant difference, and MAE alone does not establish skill relative to climatology or persistence.

Precipitation probability has no direct scalar observation in this verification method. Probability verification requires a specified event threshold and accumulation interval, matched occurrence data, and an appropriate probabilistic score. Probability of precipitation must not be interpreted as thunderstorm or lightning probability.

Environmental diagnostics

Radiosonde-derived quantities

Retained soundings support surface-based, mixed-layer, and most-unstable parcel calculations, including CAPE, CIN, LCL, LFC, and equilibrium level. The mixed-layer parcel uses the lowest 100 hPa; the most-unstable parcel is selected within the lowest 300 hPa. Other diagnostics include precipitable water, lapse rates, and bulk shear. Insufficient profile support produces an unavailable value rather than an assumed zero.

Thermodynamic calculations have numerical comparison tests against MetPy. That establishes agreement under tested inputs and assumptions, not independent validation of the observed atmosphere. Launch time, station distance, balloon drift, vertical coverage, and parcel choice remain part of the interpretation; a nearby ascent is not a contemporaneous vertical profile at the saved point.

Satellite and lightning context

ABI channel samples retain brightness temperatures, scan times, and quality information. A decrease in channel 13 brightness temperature at a fixed sampled location can support examination of cloud-top evolution, but is not a tracked updraft measurement. Advection, navigation, pixel selection, and parallax affect interpretation.

GLM records optical flashes; NLDN records collected here are pulses, with intracloud and cloud-to-ground classifications where supplied. Flash counts and pulse counts are not interchangeable. Their first retained detection is bounded by spatial coverage, cadence, and instrument sensitivity; it does not establish physical storm onset.

Sampling and interpretation

  • Collection is conditional. Cadence can increase during qualifying warnings. Some NLDN collection is gated by recent GLM activity and quota controls. These records do not form an unbiased, continuously sampled climatology.
  • Missing, zero, trace, and unknown differ. Missing reports cannot establish no precipitation, no lightning, or no severe weather. Variables within one report can have different observation times and accumulation periods.
  • Event boundaries have a definition. Warning-organized storm episodes use qualifying NWS warning intervals. Those intervals do not measure the physical beginning and end of a storm.
  • Reports require qualification. A measured gust, an estimated gust, and a damage report are different evidence. A reported peak-wind occurrence time can precede the enclosing report’s time. The highest sampled gust is not necessarily the event maximum.
  • Archive coverage is finite. Provider access, outages, collection start dates, retention, and historical storage availability bound a reconstruction. A displayed map layer is not evidence that its imagery was archived.
  • Case review has a limited inference. Retrospective associations and agreement among providers do not establish causation, calibrated confidence, warning performance, or general forecast superiority.

Reproducing a case review

The following procedure keeps the forecast question and the observed outcome comparable across reviewers. Historical coverage is limited to the retained record available for the selected location.

  1. Define the question. Record the location, event bounds, local time zone and UTC equivalents, variable, units, and any event threshold or accumulation period.
  2. Fix the information cutoffs. Use Forecast Moment in Memory to select a past window and earlier capture cutoffs. Compare the same valid hours and retain missing-hour and archive-coverage information.
  3. Inspect the selected evidence. Use the Lab reconstruction view to inspect run identifiers, issuance and capture clocks, canonical values, and coverage. Read the AFD that was available at each cutoff.
  4. Qualify the outcome. Record station identity and class, separation from the forecast point, observation time, quality flags, corrections, and report type. Keep later outcome evidence separate from the earlier information set.
  5. Preserve the selection. Save the reconstruction URL and exported JSON, including event bounds, cutoffs, view settings, and returned evidence. The API reference documents the same reconstruction response for reproducible retrieval. Account data require authorized access.

A saved URL reproduces the selection; an export preserves the evidence returned at review time. Later collection, historical imports, or corrections can change the records available to a subsequent read.