Spectra Lab
Journey · real run, 9 Oct 2026

Why did my calibration fail? Symptom, cause and fix

Five common failures, each with a plot from a real run. One of our fixes did not work, and we show that too.

SymptomLikely causeFirst fix
Good in cross-validation, worse on new dataTest data too similar to training dataHold out a season, lot or range
Error rises with more componentsOverfitStop where the curve flattens
Many outlier flags on new dataSpectra shiftedInspect the flagged samples first
High values predicted lowShrinkage and driftMore reference samples; check slope
Fails on a second instrumentGain, offset, baselineSlope and bias correction, or standardization

It looked great in cross-validation, then did worse

SymptomCross-validation says 0.85; the unseen season says 1.00. In the published meat-fat run it was 2.64, then 5.02 outside the calibration range.
CauseThe test data resembled the training data. A new season or range shifts the spectra.
FixHold out a whole season, lot or range. Report that number.
Mango: cross-validation0.85Mango: unseen season1.00Meat fat: cross-validation2.64Meat fat: outside the range5.02

Source: Spectra Lab engine run, 9 Oct 2026, mango teaching subset (CC BY 4.0). Meat-fat figures are the published run on the SpectroScience Lab page.

I used too many components

SymptomThe error on new samples goes up as components are added.
CauseExtra components fit noise. Not seen here: from 10 to 15 components the unseen-season error moved from 1.00 to 1.01.
FixRead the curve and stop where it flattens. On small datasets the rise is much steeper.
11.5213579111315PLS componentsError, % dry matterUnseen seasonCross-validation

Source: Spectra Lab engine run, 9 Oct 2026, mango teaching subset (CC BY 4.0).

The outlier flags fired on a third of the fruit

Symptom33 of 100 exceed the T² limit and 4 exceed Q.
CauseA new season shifts every spectrum, so many samples sit far from the old calibration centre.
FixLook at the flagged fruit before you drop any. We are still refining how the Lab explains this.
1e-061e-050.000110100Hotelling T² (log scale)Q residual (log scale)T² limit 18.8Q limit 0.000112

Source: Spectra Lab engine run, 9 Oct 2026, mango teaching subset (CC BY 4.0).

Predictions run low for high values

SymptomBias -0.19; slope of predicted on reference 0.85. The model pulls high samples down.
CausePLS shrinks toward the mean. The next season is also slightly shifted.
FixWe tried a bias correction. It did not help: with 5 reference samples 1.00 became 1.03; with 50, 1.00 became 0.99.
00.250.50.751Typical error, % dry matter1.001.0351.001.01101.001.00200.991.00301.000.9950Before correctionAfter a bias correction

Source: Spectra Lab engine run, 9 Oct 2026, mango teaching subset (CC BY 4.0). Typical of 500 random draws of reference samples; scored on the rest.

It works on one instrument and fails on another

SymptomPredictions from a second instrument are off by a constant and scaled.
CauseInstruments differ in gain, offset and baseline. The model learned the first one.
FixCorrect slope and bias with a few samples measured on both. See the full journey.