Spectra Lab
Journey · real run, 9 Oct 2026

From scan to verdict: one sample, six dated entries

Follow a mango from the cuvette to a fit-for-purpose verdict. Every chart is a real run of the Lab engine.

Entry 1 · Sample · 9 Oct 2026

Start with what you want to predict

What we measured

600 whole mango fruit from three seasons; dry matter measured in the laboratory, from 10.9 to 22.9 % (mean 16.1).

What we found

The laboratory number is the truth the model learns from. The instrument never sees it directly.

What to do

Collect reference values for as many samples as you can before you scan.

lampfruit sampledetectorspectrum

Illustration, not data. Source: Spectra Lab engine run, 9 Oct 2026, mango teaching subset (CC BY 4.0).

Entry 2 · Scan

Each fruit gives a curve, not a number

What we measured

Every fruit scanned from 684 to 990 nm: 103 points.

What we found

The curves look alike. They differ in height, slope and a few small bends.

What to do

Look at your own spectra before you model them.

Corrected with
lowhigh700750800850900950Wavelength (nm)Scaled signal

Raw spectra: the curves differ mostly by height and slope.

Source: Spectra Lab engine run, 9 Oct 2026, mango teaching subset (CC BY 4.0). Twelve of the 600 calibration fruit, coloured from low to high dry matter.

Entry 3 · Preprocessing

Correct the curves, then compare the errors

What we measured

Four pipelines on the same model: raw, 1st derivative, 2nd derivative, SNV.

What we found

Error on the unseen season: 1.00 with 1st derivative, 1.20 with SNV.

What to do

Try several. The best choice depends on your data.

00.51Error, % dry matter0.871.06Raw spectra0.851.001st derivative0.891.032nd derivative1.011.20SNVCross-validationUnseen season

Source: Spectra Lab engine run, 9 Oct 2026, mango teaching subset (CC BY 4.0).

Entry 4 · Model

Choose how many components, and how much smoothing

What we measured

PLS with 1 to 15 components; a smoothing window from 5 to 31 points.

What we found

The error flattens at about 10 components. The smoothing window barely moves it.

What to do

Do not chase the last 0.01. Stop where the curve goes flat.

Preprocessing
11.5213579111315PLS componentsError, % dry matterUnseen seasonCross-validation

10 components with 1st derivative.

Source: Spectra Lab engine run, 9 Oct 2026, mango teaching subset (CC BY 4.0).

11.021.0457911152131Smoothing window (points)Error, % dry matterUnseen season

Window 11 points: the Lab default.

Source: Spectra Lab engine run, 9 Oct 2026, mango teaching subset (CC BY 4.0). 1st derivative, engine Savitzky-Golay with a 2nd-order polynomial.

Entry 5 · Validation

Test it on data the model has never seen

What we measured

100 fruit from a fourth season.

What we found

Typical error 1.00 % dry matter, bias -0.19, RPD 2.94.

What to do

Always hold out a whole season or lot, not random rows.

Preprocessing
152012.51517.52022.5Reference dry matter (%)Predicted dry matter (%)dashed line: predicted equals reference

1st derivative: typical error 1.00 % dry matter, bias -0.19, error after removing bias 0.99, RPD 2.94, which the Lab bands as rough quantitative.

Source: Spectra Lab engine run, 9 Oct 2026, mango teaching subset (CC BY 4.0). 100 fruit from a season the model never saw. RPD here is the spread of the reference values divided by the error after removing bias.

Entry 6 · Verdict

Rough quantitative: RPD 2.94 sits in the 2 to 3 band

What we measured

The Lab's bands: below 1 no predictive value, 1 to 2 screening only, 2 to 3 rough quantitative, 3 or more quantitative.

What we found

This model is 2.94, just under the quantitative line.

What to do

Compare the error with the limit you need, then decide.

2.94RPD
123
Preprocessing

Your limit is 1.00; 1st derivative gives a typical error of 1.00 on a season it never saw, and an RPD of 2.94 (rough quantitative).

Source: Spectra Lab engine run, 9 Oct 2026, mango teaching subset (CC BY 4.0). Your limit is your own requirement; the Lab does not set it.

Next step: calculators · waitlist · the NIR course.