Soil Analysis
Methodology
Where the data comes from, how the model is calibrated, what resolution you actually get, and what the system can and can't tell you.
What the model does
The model takes satellite imagery of a field and predicts the chemical and physical properties of the topsoil. It combines multispectral reflectance from Sentinel-2 with topographic, climatic, and geological context, then compares the resulting pattern against tens of thousands of laboratory soil samples used to train it. The output is a per-pixel prediction at 10-metre resolution for nitrogen, phosphorus, potassium, pH, organic matter, cation exchange capacity, and soil texture (sand, silt, clay).
Data sources
The model draws on several layers of open, globally available input data:
| Data | Source | Resolution |
|---|---|---|
| Satellite imagery | Sentinel-2 SR Harmonised (Copernicus) | 10 m |
| Elevation and terrain | FABDEM, ASTER GDEM, HAND | 250 m |
| Topographic indices | Slope, aspect, TWI (derived) | 10 m |
| Climate | BIOCLIM bioclimatic variables | Global |
| Geology | Bedrock age and parent material | Global |
| Landform | Terrain classification | Global |
For each analysis date, the system pulls 8 months of Sentinel-2 imagery backwards from the target date, builds a monthly cloud-masked composite, and feeds the full stack into the model. That temporal depth is what lets it see through partial cloud cover and seasonal vegetation change.
Resolution and what "10-metre" actually means
Source imagery is 10 metres per pixel — each pixel is a 10×10 m patch on the ground. To keep the report readable, pixels are aggregated into Voronoi cells within each cluster of the field: a small field gets fewer, larger cells; a large field gets many smaller ones, each showing one representative value per property in the data table. The underlying GeoTIFF (available in the export) keeps the full 10-metre grid for anyone who wants to work with it directly.
Regional calibration
Soil chemistry varies by region — the same reading can mean something different on Bulgarian chernozems, Hungarian alluvial soils, or the Mediterranean basin. The model handles this two ways: a single European model is trained on samples from across the continent and runs on every order, and regional calibration tables adjust the phosphorus and potassium classification thresholds based on the field's location (seeSoil properties for the Bulgarian and Hungarian scales). Supported regions today: Europe, including Ukraine.
Training data
The model trains against lab soil samples published through European open-data programs — notably LUCAS, the EU's pan-European soil monitoring dataset — supplemented with regional sampling campaigns. Siora also publishes some of its own validation datasets openly; seeOpen Soil Datasets for examples from Murcia (Spain) and Ukraine.
What the model is good at
- Spatial completeness — every part of the field gets a value, not just the few points where someone happened to sample.
- Repeatability — ordering the same field on the same date always returns the same answer.
- Affordability per hectare — a fraction of the cost of physical sampling at equivalent density.
- Historical analyses — any date back to 2020, useful for tracking change or auditing past management.
What the model is not good at
A few honest limitations:
- Not a substitute for a certified lab test where legal compliance or contractual documentation is required — it's a high-resolution prediction calibrated against lab data, not a lab test itself.
- Heavy vegetation hides the soil. Eight months of imagery helps mitigate this, but continuous dense canopy still produces less precise predictions than a field with bare-soil windows — pick dates near tillage, post-harvest, or fallow periods for the best results.
- Subsoil layers aren't captured — the model represents topsoil, not deeper horizons, drainage layers, or compaction zones.
- Results outside Europe aren't validated. The system may still run, but accuracy isn't guaranteed.
- No prescription or VRA output directly — the resolution supports variable-rate planning, but the system doesn't generate application maps or ISOXML files itself; build those from the shapefile export in your VRA tool of choice.
Validation
Siora compares its predictions against held-out lab samples not used in training, reporting accuracy per property and per region. Some of this work is published as open datasets so users can verify performance independently; specific accuracy figures are available on request.
Updating the model
The model improves continuously — new training data, refined calibration, expanded regional coverage — and updates apply to new orders going forward. Past orders keep their original values, so you can always reference exactly what you saw at the time.
Next steps
- Soil properties — how predictions become the values you see in the report.
- Open soil datasets — public datasets and validation work.
- Place an order — try it on one of your own fields.