Research and validation

The technical detail behind Oceryn3D.

Specification, model design, how we measure skill and what is still to prove. Written for technical evaluators; for the overview, see the home page.

Specification

What one forecast contains.

Every forecast covers 24 levels and four variables, plus sea level, on the same grid as the data the model learned from.

Forecast specification
CoverageGlobal ocean, 75°S to 75°N
Grid0.1° (about 11 km at the equator), 3,600 × 1,499 cells
Depth24 levels from 0.5 m to about 1,000 m; land and seafloor masked per level
Output97 fields: eastward (u) and northward (v) current, potential temperature (T) and practical salinity (S) on 24 levels, plus sea-surface height
ForecastMean of the next five days, one step ahead
DataLong-term data from a global ocean model, five-day means: trained on 1981–2015, validated on 2016–2019, 2020–2023 held out for testing
ModelResidual U-Net, 13.4 million parameters; a global step is 119 tile evaluations on a single GPU
StatusResearch preview, interim validation

Model design

Two models, one ocean state.

The nowcast rebuilds today's 3D ocean from what can be measured at the surface. The forecast steps that state five days forward. Both use the same grid, levels and normalisation, so the nowcast's output is exactly the forecast's input.

Model 1

Oceryn3D Nowcast

Role
Rebuilds the 3D ocean now from surface measurements
Inputs
Sea-surface temperature, salinity and height, plus wind stress (τx, τy)
Output
Currents, temperature and salinity on 24 levels, plus sea level: 97 fields at 0.1°
Data
Satellite feeds planned; trains on global ocean model data today
Built and tested; training next

Model 2

Oceryn3D Forecast

Role
Steps the 3D ocean state five days forward
Method
Predicts how the state changes over the next five days, so it learns how the ocean moves, not only what it looks like
Output
The same 97 fields, as the mean of the next five days
Longer range
It forecasts every part of its own input except wind stress, so it can in principle be run again on its own output for 10, 15 or 20 days. That needs future wind forcing and its own validation.
Training; beats persistence (interim)

Why two models?

The forecast learns from 35 years of consistent 3D ocean data and never touches raw observations; only the nowcast does. A new satellite product means retraining the nowcast alone, and each model can be validated, and used, on its own: a 3D nowcast by itself serves drift and acoustics work.

Evidence for the nowcast

An earlier prototype, trained on a different ocean model (MOM5), reconstructed temperature and salinity in the top 500 m from surface fields alone, with R² of 0.98 and 0.96 on held-out years. That is feasibility evidence: model data in and out, and no currents. The new nowcast is built and tested on synthetic data; its training is next.

Validation results

Interim skill against persistence.

Persistence assumes the next five days look like the last five. It is the baseline every ocean forecast has to beat.

Scores by variable, validation years 2016–2019

Interim validation scores against persistence
VariableMSE ratioRMSE ratioSkill score
Eastward current (u)0.2670.520.73
Northward current (v)0.2120.460.79
Temperature0.7390.860.26
Salinity0.9220.960.08
Sea level0.7510.870.25

Ratio = model ÷ persistence (below 1 beats persistence). RMSE ratio = √(MSE ratio). Skill score = 1 − MSE ratio.

Not shown yet

  • Final scores on the held-out test years, 2020–2023
  • Skill when the forecast starts from the nowcast instead of global ocean model data
  • Comparison with independent observations: drifting buoys and Argo floats
See the validation plan

How these numbers were measured

Interim checkpoint after 54 of 100 training epochs (1 October 2026). Scored on a fixed sample of 36 forecast start dates spread across 2016–2019, years excluded from training. The starting state and the truth both come from the same global ocean model data, so these scores measure how well the model has learned the ocean's five-day evolution, not yet the skill of the full real-time chain. Errors are pooled over all ocean cells and 24 levels in normalised units, weighted by cell area, and divided once. Where checked so far, skill is highest where currents change most: the high-variability regions of western boundary currents, the Antarctic Circumpolar Current and the equator.

Validation plan

What we have shown, and what comes next.

Each step adds a harder test. Results are published with the method, including the ones that don't flatter the model.

  1. September 2026

    Four decades of 3D ocean data, ready to learn from

    3,119 global five-day fields from 1981 to 2023, standardised to 24 levels. Training uses 1981–2015, validation 2016–2019, and 2020–2023 is held back for the final test.

    Done
  2. 30 September 2026

    The forecast beats persistence

    Current errors well below persistence by the 21st training epoch, with no sign of overfitting: training and validation loss are equal.

    Done
  3. 1 October 2026

    Forecasting the full 3D state costs currents about 1%

    A control model trained on currents alone matched the main model within about 1%. The forecast keeps temperature, salinity and sea level, the state needed for longer forecasts.

    Done
  4. October 2026

    Training completes; one-time test on 2020–2023

    Error by variable, depth and region in physical units (m/s, °C, psu, m), against persistence, on years the model has never seen.

    In progress
  5. Next

    Nowcast training and the end-to-end chain

    Forecasts started from the nowcast, compared with forecasts started from global ocean model data and with persistence. If the chain loses the skill, the nowcast has failed, however good it looks alone.

    Planned
  6. Then

    Real-time inputs and independent checks

    Satellite temperature, salinity and altimetry with operational wind stress. Currents checked against drifting buoys; temperature and salinity against Argo floats.

    Planned

Research update

The forecast model has completed 54 of 100 training epochs and is still improving. Since epoch 21, the current error relative to persistence (MSE ratio) has fallen from 0.303 to 0.267 for the eastward component and from 0.239 to 0.212 for the northward.

A control model trained on currents alone came within about 1% of the main model, so we stopped it to save compute and kept the full 3D forecast. Next: finish training, then run the one-time test on 2020–2023.

u, MSE ratio
0.303 → 0.267
v, MSE ratio
0.239 → 0.212

See how it does on your waters.

A pilot scores Oceryn3D for your region on years it has never seen.

Request early access