Abstract. Canada, Manitoba and Saskatchewan have committed C$262.5M to reopen the Churchill trade corridor through Hudson Bay and Hudson Strait, water in which the Canadian Hydrographic Service (CHS) reports 15.8% of Arctic waters and 44.7% of key routes surveyed to modern standards. Under the 2,327 km Churchill–Atlantic route itself, 17% of the seabed carries a published sounding; the median survey year of those soundings is 1974. We train a heteroscedastic masked-completion network on the public CHS NONNA archive and score it on two sets of depths it never saw. (1) 109,044 research-cruise multibeam cells on the shelf (50–400 m) at locations where NONNA has no sounding: mean absolute error 4.9 m, against 13.4 m for the nearest published sounding and 5.5 m for the gravity-derived prior. (2) A temporal split using CHS's own dated Survey Index: a model trained only on soundings from before 2016 predicts the 6.46 million post-2016 soundings in the corridor to 13.3 m (18.8 m nearest sounding, 18.1 m gravity), with 72% of errors inside the model's 1σ. Because ships ground on the shallowest point and not the mean, we further train a hazard model on NONNA-10/100 pairs to predict the minimum depth within 500 m (error 2.6 m on held-out tiles versus 10.7 m if the archive depth is taken as the shoal) and hindcast six Transportation Safety Board groundings using only pre-incident soundings: the strike site falls in the top 10% of hazard among apparently safe water in 4 of 6 cases. We publish a calibrated uncertainty field, a σ-ranked survey plan (219 ship-days, C$40M), a sealed forecast of 1,314 predicted depths (SHA-256 on record), and the full code and validation cells. Limitations are stated in §8.
Hydrographic adequacy in the Canadian Arctic is a documented gap: the CHS, reporting through the IHO in December 2025, put 15.8% of Canadian Arctic waters and 44.7% of key routes at adequate survey standard [1]; in 2016 its Hydrographer-in-Charge for the Arctic put full charting "more than a decade" away [2]. Groundings on uncharted or poorly charted shoals recur — Hanseatic (1996), Clipper Adventurer (2010), Akademik Ioffe (2018), Thamesborg (2025) — and the TSB reports in each case identify the survey state of the water as a contributing factor [3–6].
The federal–provincial Churchill Plus commitment (C$262.5M, February 2026) funds port and rail, not charting. The question this report addresses is narrow: given only the public archive, how much of the corridor's seabed can be inferred with stated, tested uncertainty, and where should a survey ship go first? A second question follows from the first and is, to our knowledge, unaddressed in the completion literature: a completion model predicts mean depth, but a keel meets the shallowest point. We treat that quantity as a separate target.
Archive. CHS NONNA-100 (100 m) and NONNA-10 (10 m) non-navigational bathymetry, Open Government Licence – Canada, retrieved on 29 August 2026 through the CHS WCS endpoint. 437 national 100 m blocks (≈2,000×2,000 cells) and 3,601 10 m tiles; corridor subset of 94 blocks covering Hudson Bay, Hudson Strait, the Labrador approaches and Franklin Strait. NONNA is the published archive; CHS holds surveys not yet released to it. Throughout, "no published sounding" is the only claim made; "unsurveyed" is never inferred.
Physics prior. SRTM15+ V2.7 gravity-derived bathymetry (Scripps) [7], resampled to a 295 MB Canada raster; used as an input channel and as the always-reported baseline.
Independent multibeam. Gridded multibeam from the GMRT synthesis (Lamont-Doherty) [8], topo-mask layer, which returns measured cells only. In the corridor bounding box these derive from NCEI-archived research cruises (CCGS Amundsen 2003–2013, USCGC Healy, R/V Knorr, R/V Neil Armstrong, R/V Maria S. Merian; 36 surveys). Independence from NONNA is by provenance: research-cruise data are not CHS holdings.
Survey dates. The CHS Survey Index (DFO EGIS), 7,078 polygons with survey start/end and CATZOC grade, 1832–2016 [9]. Rasterized onto every block, it dates each sounding to the latest indexed survey covering it; a sounding outside every polygon is, by construction, post-2016. Nationally, 46.8% of NONNA-100 soundings and 46.5% of NONNA-10 soundings are post-index; in the corridor, 19,471,411 of 34,281,587 (56.8%).
Groundings. Positions, dates and drafts from TSB reports M96H0016, M10H0006, M12H0012, M14C0219, M18C0225 and the open M25C0241 occurrence [3–6]; the Thamesborg position is from AIS and press reporting and is flagged as such.
A hierarchical ConvNeXt-style U-Net with self-attention at the two coarsest scales, FiLM-conditioned on a learned resolution embedding (10 m / 100 m). Inputs per 256×256 patch: visible depth, visibility mask, gravity prior. All depth channels are normalized by the gravity patch statistics (mean, standard deviation floored at 5 m), so normalization is continuous across windows and the prediction is a residual over physics. Two heads: μ and log σ², trained with the heteroscedastic Gaussian negative log-likelihood on hidden cells plus an L1 term on visible cells. Masks mix harvested real NONNA gap patterns, random blocks and half-plane survey edges. Configurations: tiny (6.8M parameters) and small (34.8M; widths 96–768, depths 2/3/6/3). AdamW 2×10⁻⁴, cosine schedule, batch 16, bf16; 4,000 steps. Inference: overlapping windows (stride 128) with Hann blending; windows require ≥400 visible cells; predicted land (μ > −4 m) is dropped; inference is published only within 6 km of a sounding.
Where NONNA-10 and NONNA-100 overlap, the 10 m grid gives the sub-cell extreme statistics the 100 m grid hides. For each 100 m cell we compute the shallowest 10 m sounding within a 500 m radius (a 101×101 maximum filter on negative depths, requiring ≥50% coverage of the disc). 3,601 tiles mosaicked 4×4 give 475 groups of 800×800 cells. Along the Churchill route the shallowest point sits a median 6 m above the 100 m mean. The same backbone is trained to predict this extreme and its σ from the masked 100 m field and gravity, everywhere — the extreme is unknown even where a 100 m sounding exists. The hazard for a draft d is
Groups within 0.4° of any of the six grounding sites and the three geographic holdouts are excluded from training. Configurations: tiny (1,500 steps) and small (3,000 steps).
Test 1 (independent). Cells with a GMRT measured depth, no NONNA sounding, and a model fill within 60 cells (6 km) of a sounding; baselines are the nearest published sounding (via distance transform) and the gravity prior; stratified by distance and depth. Test 2 (temporal). The model is retrained with every post-index sounding blanked from the corpus (100 m and 10 m); it is then given the pre-2016 soundings of each corridor block and scored on the post-2016 ones. σ calibration. An isotonic map from raw σ to the 68th and 95th percentiles of |error| is fitted on Test 1 shelf cells and applied everywhere. Hindcast. For each grounding, the hazard model receives only soundings dated before the incident (all soundings for Thamesborg, since NONNA carries no post-2016 dates); the strike cell's P(shoal < draft), allowing 500 m of position slop, is ranked among "apparently safe" cells — mean-model depth deeper than twice the draft — within 25 km. A rank at or above the 90th percentile is counted as a landing.
| Distance to nearest sounding | n cells | SeabedNet (m) | Nearest sounding (m) | Gravity prior (m) |
|---|---|---|---|---|
| 0–0.5 km | 37,313 | 4.8 | 8.5 | 5.9 |
| 0.5–1 km | 31,657 | 4.6 | 12.9 | 5.5 |
| 1–2 km | 27,743 | 5.0 | 17.5 | 4.8 |
| 2–4 km | 12,151 | 5.5 | 20.8 | 5.8 |
| trend + natural-neighbour residual | 11.2 | |||
| trend + inverse-distance residual | 8.7 | |||
| GEBCO reference (not independent, §4.5) | 3.0 | |||
| all shelf 50–400 m | 109,044 | 4.9 | 13.4 | 5.5 |
| Depth band | n cells | SeabedNet (m) | Gravity prior (m) |
|---|---|---|---|
| coastal < 50 m (excluded, see §4.4) | 5,858 | 21.4 | 11.8 |
| shelf 50–400 m | 109,044 | 4.9 | 5.5 |
| slope 400–1000 m | 5,849 | 29.5 | 6.4 |
| basin > 1000 m | 265,934 | 10.9 | 6.3 |
On the shelf the model reduces error by 64% relative to the archive's own nearest sounding and the advantage grows with distance (5.5 vs 20.8 m at 2–4 km). Against the gravity prior the margin on these cells is small (4.9 vs 5.5 m): research cruises transit deep channels where gravity inversion already works, and the glacial shelf where it fails is where NONNA is dense and therefore absent from this test. Test 2 addresses that.
| Distance to nearest pre-2016 sounding | n cells | SeabedNet (m) | Nearest sounding (m) | Gravity (m) | Bias (m) | |err| ≤ 1σ |
|---|---|---|---|---|---|---|
| 0–0.5 km | 998,636 | 7.1 | 5.6 | 17.8 | -0.9 | 61% |
| 0.5–1 km | 910,868 | 10.8 | 11.8 | 17.9 | +0.2 | 69% |
| 1–2 km | 1,642,606 | 13.1 | 17.1 | 17.5 | +0.3 | 73% |
| 2–4 km | 2,543,006 | 15.8 | 25.4 | 18.2 | +2.6 | 76% |
| 4–8 km | 365,496 | 20.8 | 34.9 | 20.7 | +6.2 | 75% |
| trend + natural-neighbour residual | 16.8 | |||||
| trend + inverse-distance residual | 16.2 | |||||
| GEBCO reference (not independent, §4.5) | 6.2 | |||||
| all | 6,460,612 | 13.3 | 18.8 | 18.1 | +1.4 | 72% |
| Depth band | n cells | SeabedNet (m) | Nearest (m) | Gravity (m) | Bias (m) |
|---|---|---|---|---|---|
| 0–20 m | 118,374 | 13.9 | 21.1 | 11.1 | -6.3 |
| 20–50 m | 311,616 | 17.3 | 21.6 | 13.8 | -8.8 |
| 50–100 m | 672,069 | 13.9 | 17.4 | 15.3 | -2.4 |
| 100–200 m | 1,741,821 | 8.6 | 12.2 | 9.7 | +0.6 |
| 200–400 m | 2,346,496 | 12.0 | 17.2 | 17.2 | +2.2 |
| 400–9999 m | 1,254,668 | 21.1 | 31.1 | 34.7 | +6.3 |
Raw σ ranked error correctly (Spearman-like correlation 0.30 on Test 1 shelf cells) but was under-confident: 48% of errors inside 1σ where 68% is expected, 74% inside 2σ. The isotonic map (mean factor 1.9×) restores 68% / 95% coverage; the temporal test, scored with the raw σ of a differently trained model, independently shows 72% inside 1σ. All uncertainty on the atlas is the calibrated one. Inferred cells deeper than 400 m carry the gravity prior (Table 2).
| Raw σ decile (m) | n | |err| p50 (m) | |err| p90 (m) | inside 1σ |
|---|---|---|---|---|
| 0.3–2.1 | 38,697 | 5.0 | 14.1 | 13% |
| 2.1–3.0 | 38,749 | 4.1 | 12.3 | 32% |
| 3.0–3.8 | 38,720 | 4.2 | 12.4 | 42% |
| 3.8–4.9 | 38,719 | 3.7 | 11.7 | 56% |
| 4.9–5.9 | 38,894 | 4.2 | 11.3 | 60% |
| 5.9–7.2 | 38,908 | 4.8 | 12.2 | 64% |
| 7.2–10.3 | 38,716 | 5.3 | 15.5 | 69% |
| 10.3–17.8 | 38,808 | 6.9 | 21.7 | 76% |
| 17.8–29.7 | 38,751 | 9.1 | 29.2 | 85% |
| 29.7–135.6 | 38,705 | 19.3 | 76.1 | 79% |
5,858 Test 1 cells shallower than 50 m, at a median 131 m from a sounding and all at the shoreline of three blocks, disagree with every source: the model reads -20 m, the nearest CHS sounding -39 m, and gravity +9 m relative to GMRT. We attribute this to GMRT's coastal grid rather than measured swath and exclude the band from the headline pending inspection; it is reported here rather than removed.
Nearest-sounding is a floor. The classical gap fill is a gravity trend plus interpolated residuals (sounding minus SRTM15+ at the input soundings), interpolated by Delaunay natural-neighbour-style linear interpolation or by inverse-distance-squared over the 12 nearest residuals. On Test 1 these score 11.2 and 8.7 m against the model's 4.9 m; on Test 2, 16.8 and 16.2 m against 13.3 m (Tables 1 and 3, italic rows). GEBCO (NCEI global mosaic, GEBCO 2024/25, 15 arc-second) is reported as the chart-world reference and scores 3.0 m (Test 1) and 6.2 m (Test 2) on the 100 m grid; it is not an independent baseline because it ingests NONNA-100 through IBCAO v5 (the Arctic Seabed 2030 compilation) and the same NCEI cruise archive; its lower error on both tests is the residual of resampling data it already contains, not a prediction, and where it has no source soundings it reduces to SRTM15+, the gravity row. Gravity leakage. SRTM15+ V2.7 (released April 2025) inherits the cumulative NCEI multibeam archive of its point releases. Its error on the Test 1 cells is 5.5 m against 15.2 m on 282,193 ordinary NONNA-sounded shelf cells in the same blocks; the prior has almost certainly seen the cruise data. This favours the gravity baseline, which the model still edges on Test 1, and the model, which takes gravity as an input, inherits part of it. Test 2 is therefore the clean comparison: those NONNA shelf soundings are below SRTM15+'s resolution, and there gravity scores 18.1 m against the model's 13.3 m.
On held-out tile groups (around the grounding sites and the geographic holdouts), the 34.8M hazard model predicts the shallowest point within 500 m to 2.6 m mean absolute error where a 100 m sounding exists (10.7 m if that sounding is taken as the shoal; 12.1 m gravity) and 8.2 m where no sounding exists (11.3 m nearest field); 80% of errors inside 1σ, 95% inside 2σ. The 6.8M model: 4.0 / 9.4 m.
| Grounding | Pre-incident soundings in block | km to nearest | P(shoal < draft) at site | Hazard percentile | Mean-σ percentile | Nearest-sounding percentile |
|---|---|---|---|---|---|---|
| Thamesborg 2025-09-06 · TSB M25C0241 | 1,000,495 | 0.8 | 27% | 96 | 99 | 90 |
| Akademik Ioffe 2018-08-24 · TSB M18C0225 | 165,454 | 9.4 | 26% | 93 | 77 | 90 |
| Clipper Adventurer 2010-08-27 · TSB M10H0006 | 16,457 | 12.4 | 19% | 84 | 11 | 22 |
| Hanseatic 1996-08-29 · TSB M96H0016 | 148,672 | 0.3 | 94% | 99 | 18 | 91 |
| Nanny 2012 2012-10-25 · TSB M12H0012 | 12,501 | 6.3 | 95% | 95 | 74 | 98 |
| Nanny 2014 2014-10-14 · TSB M14C0219 | 36,222 | 0.0 | 62% | 79 | 98 | 82 |
The hazard percentile separates the four uncharted-shoal groundings from the baselines in a way neither the mean model's σ nor the nearest-sounding depth does: for Hanseatic the mean model's σ ranks the site at the 18th percentile (the area was densely sounded, so the mean model was confident) while the hazard field ranks it 99th. For Thamesborg the two agree (96th hazard, 99th σ), which is consistent with the court's finding that the ship was routed through CATZOC C water 1.7 km beyond the modern multibeam swath.
Table 6 lets the model see every published sounding around a site at inference; NONNA carries no dates after 2016, so a post-incident survey could be among them. Table 6b removes every sounding within 10 km of each site from the input as well (the hazard model was trained with all tile groups within 0.4° of the sites excluded); a 25 km hole is also shown, at which the site lies outside the 12 km inference envelope and the model returns no estimate. The two denominator choices are then varied: comparison window (10/25/50 km) and the apparently-safe filter (1.5×, 2×, 3× draft on the mean map). Skill: by construction 10% of apparently-safe cells lie above the 90th percentile, so with the 10 km hole the 4-of-6 result has binomial p = 0.0013 (all six) and 3 of 4, p = 0.0037, on the four uncharted-shoal groundings; with all soundings visible, 4 of 6 (p = 0.0013). Precision cannot be bounded from six incidents; corridor-wide 18% of cells exceed P = 5% and 26 of 1762 apparently-safe route-km are flagged.
| Grounding | all soundings | 10 km hole | 25 km hole | 10 km hole · 10 km window | 10 km hole · 50 km window | 10 km hole · safe 1.5× | 10 km hole · safe 3× |
|---|---|---|---|---|---|---|---|
| Thamesborg | 96 | 93 | — | 91 | 90 | 92 | 94 |
| Akademik Ioffe | 93 | 94 | — | 74 | 98 | 93 | 94 |
| Clipper Adventurer | 84 | 84 | — | 99 | 83 | 83 | 84 |
| Hanseatic | 99 | 95 | — | 86 | 95 | 86 | 99 |
| Nanny 2012 | 95 | 96 | — | 100 | 97 | 95 | 98 |
| Nanny 2014 | 79 | 60 | — | 33 | 70 | 53 | 73 |
| # | Lat | Lon | Inferred km² | mean σ (m) | peak σ (m) | Ship-days | C$M |
|---|---|---|---|---|---|---|---|
| 1 | 59.25°N | 59.75°W | 1,055 | 52.6 | 109.7 | 26.4 | 4.8 |
| 2 | 59.25°N | 60.25°W | 740 | 55.3 | 116.6 | 18.5 | 3.4 |
| 3 | 59.75°N | 60.25°W | 559 | 54.7 | 102.8 | 14.0 | 2.6 |
| 4 | 58.75°N | 59.25°W | 895 | 26.1 | 133.8 | 22.4 | 4.1 |
| 5 | 59.25°N | 59.25°W | 724 | 30.2 | 78.2 | 18.1 | 3.3 |
| 6 | 59.75°N | 60.75°W | 993 | 17.1 | 62.2 | 24.8 | 4.5 |
| 7 | 60.75°N | 65.25°W | 617 | 23.9 | 75.0 | 15.4 | 2.8 |
| 8 | 60.75°N | 64.75°W | 783 | 17.1 | 84.7 | 19.6 | 3.6 |
| 9 | 58.75°N | 58.75°W | 1,039 | 12.2 | 32.3 | 26.0 | 4.8 |
| 10 | 57.75°N | 56.75°W | 1,348 | 8.7 | 33.1 | 33.7 | 6.2 |
File forecast_2026-09-02.csv (1,314 cells: 314 on the found channel, the rest random inferred cells across the corridor) lists a predicted depth and calibrated 68% / 95% bands at cells that have no published sounding as of 29 August 2026. SHA-256 4d3f5dc8a5c2b50f3f72fba95384581fc778d6d453a3eec46aa2dc383c66be45, sealed 2026-09-02. Scoring rule: any later survey through these cells reports the mean absolute error and the fraction of soundings inside each band; we publish the score regardless of outcome. The depths were produced by the 34.8M-parameter v5-small completion model with the depth gate and calibrated σ of §4.3. The file is timestamped with OpenTimestamps (proof forecast_2026-09-02.csv.ots, three public calendars, Bitcoin-anchored) and its hash is in a GitHub commit of the validation repository, so the seal does not depend on our clock.
(i) NONNA is not the CHS holdings; a "no published sounding" cell may be surveyed. (ii) Test 1 under-represents the glacial shelf; Test 2 covers it but its labels are CHS's own later surveys, so datum and processing differences between survey eras are absorbed into the error. (iii) In water shallower than 50 m the mean-depth model is biased deep by 6–9 m (Table 4); the hazard model, not the mean model, is the appropriate product there. (iv) The hindcast has six cases, two of which are controls and two of which lie beyond the publication envelope; it is a benchmark, not a proof. (v) The Thamesborg position is not yet official. (vi) Survey-day rate and km²/day are assumptions from public contracts and stated arithmetic, not quotes. (ix) The gravity prior has very likely ingested the Test 1 cruises (§4.5). (vii) The physics prior itself ingests historical soundings, so the gravity baseline is not fully independent on surveyed water; on the cells of Test 1 it is. (viii) Nothing here is a chart. Every product carries "planning prior — not for navigation".
All scripts, result JSON files and the per-cell validation arrays are public at github.com/girardemilio3-svg/seabednet-validation; the atlas page and this report are generated from those files by build_atlas_v2.py, build_atlas_hazard.py and build_report.py, so no number is transcribed by hand. Data are public (CHS NONNA under the Open Government Licence – Canada; GMRT; CHS Survey Index; SRTM15+). Training the 34.8M completion model to 4,000 steps takes 17 minutes on one RTX 5090 and about 12 hours on an NVIDIA GB10; the hazard model 13 minutes. Model weights are available on request.
chs_edh_survey_index), 7,078 dated polygons with CATZOC.Version 1.0, 2 September 2026. Generated from result files; see the repository for the exact commit.