Spectra Lab
Sample · mango dry matter

Entry 1: 600 fruit from three seasons, a fourth season held back

We want to predict dry matter in whole fruit from a scan. 600 fruit from three seasons build the model; 100 from a fourth season test it.

Entry 1 · Sample · 9 Oct 2026 · Lab engine run · page 1 of 8

What we measured

Dry matter by laboratory, 10.9 to 22.9 % in calibration, 12.0 to 23.8 % in the unseen season.

What we found

The two seasons overlap but the fourth runs higher on average.

What to do

Check that your own test set covers a different season, lot or range.

Reference

What the laboratory measured

Each bar is the share of fruit in a dry-matter range.

05101512.51517.52022.5Dry matter, reference value (%)Share of samples (%)Calibration, seasons 1 to 3Unseen season 4

Source: Spectra Lab engine run, 9 Oct 2026, mango teaching subset (CC BY 4.0).

Source

Where the data comes from

Anderson, Walsh, Subedi and Hayes (2020), Postharvest Biology and Technology 168; dataset licensed CC BY 4.0 (Mendeley Data, doi 10.17632/46htwnp833). Teaching subset: 600 calibration and 100 unseen-season spectra, 684 to 990 nm, 103 points.

Scan · 684 to 990 nm

Entry 2: every fruit gives a curve, and the curves look alike

The instrument records absorbance at 103 wavelengths. The line is the mean; the band is one standard deviation across fruit.

Entry 2 · Scan · 9 Oct 2026 · Lab engine run · page 2 of 8

What we measured

Absorbance at 103 wavelengths for each of 600 fruit.

What we found

The mean curve falls steeply at the short end and rises again at the long end; the band shows how much fruit differs at each wavelength.

What to do

Plot your own spectra before modelling. Odd curves here are easier to fix than odd predictions later.

Spectrum

What the instrument saw

-0.500.5700750800850900950Wavelength (nm)Absorbance, as stored in the file

Source: Spectra Lab engine run, 9 Oct 2026, mango teaching subset (CC BY 4.0).

Preprocessing

Entry 3: the same data gave 1.00 with a 1st derivative and 1.20 with SNV

Four pipelines, one model type, one unseen season. SNV was about 20 % worse than the 1st derivative here; on other data the order can change.

Entry 3 · Preprocessing · 9 Oct 2026 · Lab engine run · page 3 of 8

What we measured

Raw spectra and three corrections, each fitted on the calibration fruit only.

What we found

Derivatives remove the height and slope differences. The lowest unseen-season error came from the 1st derivative.

What to do

Compare on your own data. Do not carry a choice over from someone else's.

Before and after

Drag from raw to corrected

Corrected with
lowhigh700750800850900950Wavelength (nm)Scaled signal

Raw spectra: the curves differ mostly by height and slope.

Source: Spectra Lab engine run, 9 Oct 2026, mango teaching subset (CC BY 4.0). Twelve of the 600 calibration fruit, coloured from low to high dry matter.

Compare

Error by preprocessing

00.51Error, % dry matter0.871.06Raw spectra0.851.001st derivative0.891.032nd derivative1.011.20SNVCross-validation (RMSECV)Unseen season (RMSEP)

Source: Spectra Lab engine run, 9 Oct 2026, mango teaching subset (CC BY 4.0).

Numbers

All the numbers

Bias is predicted minus reference; SEP is the error after removing the bias.

PipelineComponentsRMSECVRMSEPBiasSEPRPD
Raw spectra110.8671.059-0.2061.0442.78
1st derivative100.8530.999-0.1870.9872.94
2nd derivative90.8921.028-0.0281.0332.81
SNV101.0051.197+0.0381.2022.42
Model

Entry 4: the error stops falling around 10 components, and the smoothing window barely matters

The Lab fits 1 to 15 PLS components and picks the lowest cross-validated error. The slider shows what each number does to the error on the unseen season.

Entry 4 · Model · 9 Oct 2026 · Lab engine run · page 4 of 8

What we measured

PLS with 1 to 15 components on each pipeline, and seven smoothing windows from 5 to 31 points.

What we found

Both errors are flat from about 9 components. The window changes the error by about 0.04.

What to do

Stop where the curve goes flat. A simpler model transfers better.

Components

Drag the number of components

Preprocessing
11.5213579111315PLS componentsError, % dry matterUnseen seasonCross-validation

10 components with 1st derivative.

Source: Spectra Lab engine run, 9 Oct 2026, mango teaching subset (CC BY 4.0).

Smoothing

Drag the smoothing window

11.021.0457911152131Smoothing window (points)Error, % dry matterUnseen season

Window 11 points: the Lab default.

Source: Spectra Lab engine run, 9 Oct 2026, mango teaching subset (CC BY 4.0). 1st derivative, engine Savitzky-Golay with a 2nd-order polynomial.

Validation

Entry 5: on a season the model never saw, the typical error is 1.00 % and the model under-predicts by 0.19

100 fruit from a season the model never saw: predicted dry matter against the laboratory value.

Entry 5 · Validation · 9 Oct 2026 · Lab engine run · page 5 of 8

What we measured

Predictions from the model for 100 fruit in the fourth season, against the laboratory values.

What we found

Most points sit near the dashed line. High values run low, which is why the slope is below 1.

What to do

Report the number from data the model never saw, with its bias.

Chart

Predicted against reference

Preprocessing
152012.51517.52022.5Reference dry matter (%)Predicted dry matter (%)dashed line: predicted equals reference

1st derivative: typical error 1.00 % dry matter, bias -0.19, error after removing bias 0.99, RPD 2.94, which the Lab bands as rough quantitative.

Source: Spectra Lab engine run, 9 Oct 2026, mango teaching subset (CC BY 4.0). 100 fruit from a season the model never saw. RPD here is the spread of the reference values divided by the error after removing bias.

Statistics

What the numbers say

StatisticValueWhat it tells you
RMSEP0.999 %typical error on new samples
Bias-0.187 %average predicted minus reference
SEP0.987 %error after removing the bias
RPD2.94spread of reference divided by SEP
Slope0.8551.000 would mean predictions scale perfectly with reference
R²0.886squared correlation, reference and predicted
Verdict

Entry 6: rough quantitative, RPD 2.94

The Lab sorts models by RPD into four bands. Whether the band is enough depends on what you need it for.

Entry 6 · Verdict · 9 Oct 2026 · Lab engine run · page 6 of 8

What we measured

RPD on the unseen season, with the Lab's bands: below 1 no predictive value, 1 to 2 screening only, 2 to 3 rough quantitative, 3 or more quantitative.

What we found

The model is 2.94, just under the quantitative line.

What to do

Put your own limit in and read the verdict.

Verdict

Where the model sits

2.94RPD
123

Source: Spectra Lab engine run, 9 Oct 2026, mango teaching subset (CC BY 4.0).

Your limit

Is that good enough for you?

Preprocessing

Your limit is 1.00; 1st derivative gives a typical error of 1.00 on a season it never saw, and an RPD of 2.94 (rough quantitative).

Source: Spectra Lab engine run, 9 Oct 2026, mango teaching subset (CC BY 4.0). Your limit is your own requirement; the Lab does not set it.

Outliers

Entry 7: 33 of 100 fruit exceed the T² limit; 4 exceed the Q limit

T² says a sample sits far from the calibration centre; Q says the model cannot rebuild its spectrum. A new season shifts spectra, so flags are information about the season.

Entry 7 · Outliers · 9 Oct 2026 · Lab engine run · page 7 of 8

What we measured

T² and Q for each unseen-season fruit against the limits set on the calibration fruit.

What we found

A third of the fruit are flagged on T². That tells us the season differs, not that the scans are bad.

What to do

Inspect the flagged samples before dropping any. We are still refining how the Lab explains this.

Chart

Each point is one unseen-season fruit

Dashed lines are the limits; red points are beyond at least one.

1e-061e-050.000110100Hotelling T² (log scale)Q residual (log scale)T² limit 18.8Q limit 0.000112

Source: Spectra Lab engine run, 9 Oct 2026, mango teaching subset (CC BY 4.0).

Limits and sources

Entry 8: what this report does not cover, and where the data came from

One dataset, one split, one model family.

Entry 8 · Limits · 9 Oct 2026 · Lab engine run · page 8 of 8

Limits

Six limits

  • One dataset, one split. Mango fruit, seasons 1 to 3 against season 4. Grain, dairy or feed would behave differently.
  • One model family. Partial least squares, 5-fold cross-validation, 1 to 15 components.
  • Not an accredited method. The report supports decisions; validate against your own requirements and standards before relying on a calibration.
  • Teaching subset. 600 + 100 spectra from the larger archive of 11,691.
  • Reproducible. The same files through the same engine give the same numbers; we re-ran it to check.
  • RPD here is the spread of the unseen-season reference values divided by the error after removing bias; the Lab's model overview computes RPD from cross-validation instead.
Sources

Attribution

Anderson, Walsh, Subedi and Hayes (2020), Postharvest Biology and Technology 168; dataset licensed CC BY 4.0 (Mendeley Data, doi 10.17632/46htwnp833). Teaching subset: 600 calibration and 100 unseen-season spectra, 684 to 990 nm, 103 points.

Engine: the Spectra Lab modelling engine (scikit-learn PLS). The published 1.01 comparison is the figure from the dataset paper for PLS on the full data.