A review picks. A reviewer reads the finding she distrusts, re-runs the number she suspects, and leaves the rest alone. A whole-site re-test does not pick. On 20 September 2026 every derived dataset on this site was regenerated, every analysis re-run with the same scripts, and every page rewritten from the new artifacts. The claims did not know they were being tested. Nobody decided which ones deserved it. What came back is the cleanest audit this site has had, and it was not ordered.
The mechanism was a coordinate. From May until 19 September the site’s forecast and its climate findings had been computed for 17.5897°N, 101.4317°W, which is Los Llanitos, 44 km south-east of La Saladita; the constant was corrected to 17.8375°N, 101.7606°W and the rebuild followed the next day. That is the whole of the error. What matters is what the rebuild showed.
What a test with no author in it can see
The re-test was triggered by a gust. Ahead of the tropical low of 21 September the site showed 15.6 knots for the Monday while the cell that contains La Saladita showed 25.5, and that gap led back to the constant. The correction touched 92 site-point constants in the API and build scripts and 68 schema.org geo blocks, headers and iNaturalist links. Then every derived dataset, from birds to bathymetry to marine heatwaves, was regenerated at the new point in one pass by the same code that had produced it.
The beach station scored the move before any page did. Against 848 station days with at least 90% coverage, the model’s daily-high error fell from 3.05 to 2.34 °F, its daily-low error from 2.32 to 2.06 °F, its rain error from 2.80 to 2.26 mm. Gust error rose, 6.55 to 8.03 kt, which fits a sheltered anemometer reading low on calm days; on days when the station reads 15 kt or more, model and station agree to within about 2 kt. The temperature calibration told the same story in small: the old constants, +0.74 and −1.19 °F, had absorbed the 44 km offset as model bias, and the night-time term had the wrong sign. Refitted over 889 days they are +0.84 and −1.08.
What the correction could not reach
The storm findings came back from a second re-test, on 27 September, with the strong end still climbing and one claim gone. The strongest storm within 500 km of La Saladita in each decade rose from 120 kt (1959) to 185 kt (2015); the interpolated decade 90th percentile rose from 80 to 131 kt (r = 0.87, p = 0.005). The spread did not rise with them: decade standard deviation (p = 0.155) and interquartile range (p = 0.40) show no trend. The top is rising; the distribution is not widening, and the page that said it was now says it is not. The season arrives 2.30 days earlier per decade rather than 2.45 (p = 0.008), and peak wind within 500 km rises 4.8 kt per decade rather than 4.6. More storms inside 1,500 km, more storm-days, longer dwell. On NOAA’s current El Niño index the earlier start leans the same way in every phase and is significant in none of them, so the claim that it held within each phase is gone too.
The coordinate did not do this. The storm results come from HURDAT2, the hurricane best-track catalogue, computed as the distance from one point to every six-hourly fix inside a 500 or 1,500 km ring. Move the point 44 km and the rings barely notice, and through 20 September the storm pages did not change by a digit. Robust was the obvious word for that, and it was the wrong one. A finding that survives a coordinate correction is a finding whose input the correction never touched; the first re-test did not validate the storm results, it never reached them.
A paper reached them. Writing the intensity result up as a preprint meant mapping every number in it to a field in its artifact, and the mapping showed what the June script had computed. Its “decade 90th percentile” was the value at position ⌊0.9 n⌋ of each decade’s sorted years, and with ten or fewer years a decade that position is the last one: the headline compared the single strongest storm of the 1950s with the single strongest of the 2010s. Its satellite-era check, “rolling p90, r = 0.87”, was a rolling standard deviation. The storm table had been rebuilt from the current HURDAT2 file on 19 September, so every storm page was re-run on it with the estimator fixed.
That second audit is the same kind as the first. Nobody picked the estimator for review. A rule that every published number trace to a named field picked it, the way a gust picked the coordinate. The storm leg was out of reach of one accident and squarely inside the next.
What moved
Wave power. The June page headlined peak power climbing: +4.1 kW/m per decade, p=0.032. ERA5-Wave resolves the corrected site point to its nearest 0.5° node, 17.5°N, 102.0°W — about 45 km offshore to the south-west, not the break itself. At that offshore node the slope is +3.8 kW/m per decade at p=0.27. Same sign, no finding. The “24 to 81 kW/m” contrast the page drew between 1979 and 2024 does not exist here: 1979’s hardest hour was about 40 kW/m and the record hour is 232 kW/m, in 1993. Drop the 1993 hour and the trend comes back at p=0.004, which is exactly why it is headlined in neither direction now.
Warming. The climate page had surface warming at +0.12 °C per decade, p=0.001, R²=0.22. At the correct cell it is +0.278 °C per decade, p<0.001, R²=0.49. The published number was less than half the current one.
Rain. The rainfall-structure page had “one clear signal, four nulls”: wet-run clustering at +1.0 runs per decade, p=0.009, and nothing else. At the correct cell five of seven structural tests are significant. Clustering is +1.42 runs per decade at p=0.0009. The longest wet spell shortens 3.9 days per decade, p=0.00009. The longest dry spell shortens 7.2 days per decade, p=0.042, where the June text had said the dry spells were holding. The hardest single day of the year, the page’s “most interesting null” at +10 mm per decade and p=0.15, is +17.6 mm per decade at p=0.010. On the climate page the 99th-percentile daily rain went from +6.4 mm per decade at p=0.11 to +10.1 at p=0.017. Two of the five clear a Bonferroni correction for seven tests; three do not, and the pages say which.
Marine heatwaves. Already redone before the essays were revised. The June page had three cached years at the wrong cell. The rebuild re-probed the nearest OISST cell that is not land-masked, 17.875°N, 101.875°W, about 12 km west-south-west, and ran the full 44-year record: 129 events, all Moderate under the 1982–2011 baseline, +9.6 heatwave-days per decade at p=0.009. The all-Moderate result the June proxy gave was confirmed rather than overturned; the 90th-percentile gap here is about 1.1 °C.
Ground. The seismic page moved with everything else: the chance of a magnitude 6.5 or larger within 100 km went from 11% to 13%, nine qualifying events rather than eight, and the 1985 Pantla aftershock is 13 km away, not 33.
The essay that stood on three climate legs, the polite story and the dangerous story, now stands on two, with a third that leans the same way and cannot bear weight. The rain leg got stronger and changed shape. The storm leg was out of the coordinate’s reach, and the preprint narrowed it: the strongest storms are getting stronger, but the distribution around them is not spreading.
The move, read as an experiment
The natural-experiments page treats the correction the way it treats a hurricane year: sealed first, read second. The seal is post hoc and the page says so: it fixes the reading rule, not the author’s ignorance. Baseline: 58 settled forecast days from 18 July, when the station’s temperature record begins, to 13 September, at the old point. After: 10 days. Expectation: daily-high error falls by at least 1 °F once the model describes the right place. Falsifier: if the daily high and the night low do not fall together, the change is weather, not the cell.
Daily-high error fell from 2.48 to 1.58 °F: −0.90 °F, 95% interval −1.6 to −0.2, Welch p=0.01, permutation p=0.03. The verdict is INCONCLUSIVE, because the interval reaches across the 1 °F bar set in advance. The retrodesign gives it power 0.71 and an exaggeration factor of 1.19 if anyone called it significant. Ten days are ten days.
Night lows did not improve: −0.20 °F, interval −0.91 to +0.51, NULL_POWERED at power 0.09. The falsifier fired, and the station audit says it had to. Night-to-night spread in the measured low is 1.7 °F, and a 30-day mean of the station’s own lows beats every model on offer. There was nothing at night for a better cell to fix. Rain error went the other way, +1.32 mm with an interval from −2.4 to +5.05, inconclusive by design; rain here is convective and local at either point.
The practice that follows
Three rules, all already on the pages. First, the sealed verdict before the read: baseline, expectation, falsifier, bar, the author’s prediction with a confidence, and a hash of all of it, written down before the numbers are looked at. Every graded experiment on this site now carries one. It is the ordered audit made to behave like the accidental one: the reading rule is fixed before anyone knows what the reading will be. Second, a visible revision note at the top of every page the correction touched, dated 26 September 2026, giving the wrong-cell number beside the current one. Wave power, rainfall structure, climate trends, marine heatwaves, the essay; and, dated 27 September, the storm pages the estimator touched. Not a silent overwrite. Third, every number in a paper or on a page maps to a named field in its artifact, so a label that says 90th percentile cannot sit on code that takes the maximum.
The pages published on 9 June carried the wrong-cell numbers until 19 September, and the intensity page carried a maximum labelled as a percentile until 27 September. Each now says so in its first paragraph, and each says which of its numbers the cell could touch and which it could not. The next audit, ordered or not, should find that line already drawn.