Soil Spectral Ensemble Learning for Stable Organic Carbon Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional methods for determining soil organic carbon content are costly, time-consuming, and environmentally harmful, and single predictive models in soil spectroscopy have limited applicability and stability, making large-scale, accurate predictions challenging.

Innovation Solution

A spectrum-guided ensemble learning model is constructed by combining partial least squares regression, Cubist, and random forest models, using 10-fold cross-validation to determine optimal parameters, and integrating their predictions to form a more accurate soil organic carbon prediction model.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional laboratory physical and chemical analysis is used to determine soil organic carbon, then measurement precision is improved, but loss of time and cost increase significantly

Engineering Contradiction:
Improvesoil organic carbon measurement precisionVSAvoiddetermination cycle time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces traditional mechanical/chemical laboratory analysis systems with a spectral analysis system using visible-near infrared spectroscopy combined with ensemble learning algorithms. This substitution maintains measurement precision while dramatically reducing determination time and cost by using optical properties and computational models instead of lengthy chemical processes.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent creates a predictive model that copies the relationship between soil spectral characteristics and organic carbon content established from training data. Once the ensemble learning model is trained with laboratory-measured data, it can rapidly predict soil organic carbon content for new samples without repeating the full laboratory analysis process, thus reducing time and cost while maintaining precision.

Inventive Principle:
Principle #26Copying

2Area of stationary object

If field sampling is conducted for large area soil organic carbon estimation, then measurement coverage is improved, but cost and risk increase

Engineering Contradiction:
Improvesoil organic carbon estimation coverage areaVSAvoidsampling cost and risk
Core Design Contradiction:
Area of stationary objectVSEase of manufacture

Solution Approach 1:

The patent replaces physical field sampling and transportation with in-situ spectral measurement using portable visible-near infrared spectrometers. This allows direct measurement of soil organic carbon content at the field location, eliminating the need for sample collection, transportation, and laboratory processing, thereby reducing cost and risk while enabling large-area coverage.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent enables the soil itself to provide the measurement information through its spectral properties. The soil's interaction with visible-near infrared light contains the necessary information about organic carbon content, allowing the system to self-measure without requiring physical sample removal or complex field sampling procedures.

Inventive Principle:
Principle #25Self-service

3Device complexity

If single predictive models are used in soil spectral analysis, then model simplicity is maintained, but reliability and applicability are limited

Engineering Contradiction:
Improvepredictive model structureVSAvoidprediction model stability
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent combines multiple different predictive models (including partial least squares regression, random forest, and other machine learning models) into an ensemble learning system. Each model contributes its strengths, and their predictions are integrated to produce a final result that is more reliable and stable than any single model alone, while still maintaining reasonable computational complexity.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a composite predictive model structure that integrates multiple algorithms with different mathematical foundations and assumptions. This composite approach leverages the complementary strengths of linear models, tree-based models, and other techniques, resulting in a prediction system that is more robust across different soil types and conditions than any individual model.

Inventive Principle:
Principle #40Composite materials

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

The ensemble learning model achieves high prediction accuracy (R2=0.76, RMSE=9.55 g kg−1) and reduces costs while improving efficiency, offering a cost-effective and efficient method for large-scale soil organic carbon content prediction.

Implementation Method 1

soil visible-near infrared spectroscopy has the advantages of low cost, high portability and less external interference

Methodology Applied
Scientific EffectVisible-near infrared spectroscopy: Absorption Spectroscopy

Data Source

PatentUS12537075B2Method and device for spectral prediction of soil organic carbon based on spectrum-guided ensemble learning
Publication Date: 2026.01.27 ZJU HANGZHOU GLOBAL SCI & TECH INNOVATION CENT
  • US12537075B2 patent drawing

AI summary

Disclosed is a method and device for predicting soil organic carbon based on spectrum-guided ensemble learning. The method includes: obtaining a soil sample and a real organic carbon content and an original soil spectrum thereof, pre-processing the original soil spectrum to obtain a soil spectrum sample, constructing, based on the soil spectrum sample, a sample set, grouping the sample set into a first training set and a validation set, and using the real organic carbon content as a label; training, based on the first training set and the corresponding labels, a partial least squares regression model, a Cubist model and a random forest model to obtain carbon content predicted value sets of the three models; constructing, based on the carbon content predicted value sets of the three models and soil spectrum principal component data, a second training set, and training a second random forest model with the second training set and corresponding labels to obtain a spectrum-guided ensemble model. The method combines the advantages of different predictive models and can accurately predict the soil carbon content.