Cascaded Machine Learning for Crop Yield Estimation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for crop and grassland yield estimation, whether model-driven or data-driven, face inefficiencies and inaccuracies due to high costs, limited scalability, and complexity in handling large datasets, especially when dealing with diverse variables and spatial-temporal data from Remote Sensing and Geoinformation Systems.
Innovation Solution
A method involving a cascaded approach using Machine Learning Algorithms (MLA) where multiple models are trained and selected based on predefined error values, utilizing remote sensing data and socio-economic factors to estimate crop yields, allowing for high accuracy and scalability across various conditions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If model-driven approaches are used for yield estimation, then estimation accuracy can be improved through mechanistic modelling, but cost and complexity increase due to intensive calibration requirements and limited scalability
Solution Approach 1:
The patent replaces complex mechanical/calibration-based model-driven approaches with data-driven Machine Learning Algorithms that automatically learn patterns from data without requiring intensive manual calibration of biophysical processes, thus reducing complexity while maintaining accuracy
Solution Approach 2:
The patent changes the approach from using predefined mathematical equations with calibrated parameters to using Machine Learning models that automatically adjust parameters based on training data, enabling the system to adapt to different locations and conditions without manual recalibration
2Measurement precision
If data-driven approaches with large number of variables are used, then yield estimation accuracy can be improved, but model complexity increases and precision declines
Solution Approach 1:
The patent changes from classical statistical models to Machine Learning Algorithms that can handle large numbers of variables efficiently, transforming the mathematical approach to enable inclusion of multiple Remote Sensing and environmental variables without sacrificing model precision
Solution Approach 2:
The patent moves from two-dimensional statistical relationships to multi-dimensional Machine Learning models that can process and interpret complex interactions among multiple variables simultaneously, enabling comprehensive yield estimation without dimensional reduction
3Quantity of substance
If process-based models are used, then estimation can be performed with limited data, but scalability to other areas is restricted due to location-specific calibration
Solution Approach 1:
The patent creates a universal Machine Learning model that can be applied across different locations, crop types, and environmental conditions without requiring location-specific calibration, enabling the model to function universally while adapting to local conditions through training data
Solution Approach 2:
The patent uses Machine Learning to create a digital copy of agricultural systems that learns from training data and can predict outcomes for new locations without requiring physical replication or manual adaptation of the underlying models
Data Source
Figure 1~2
Figure 3~4
Figure 5~7
AI summary
The invention relates to the field of yield estimations based on Remote Sensing and/or Geoinformation Systems, particularly to a method for determining a yield estimation of crops of a field (200). The invention further relates to a system for determining a yield estimation of crops, to a use, to a program element, and to a computer-readable storage medium. The method comprises the steps of: providing a plurality of processed data (12) of the field (200); partitioning the plurality of processed data (12) into an L1 trainset (13) and an L1 testset (14), which is disjoint of the L1 trainset (13); training each L1 model (19m.i) of a plurality of L1 models (19) based on the entries of the L1 trainset (13); outputting, for each trained L1 model (19m.i), an L1 predicted yield value (19p.i), based on the L1 trainset (13), and comparing, for each L1 model (19m.i), the L1 predicted yield values (19p.i) of the L1 trainset (13) to the observed yield values (220); selecting an L1 subset of L1 models (19m.i) of the plurality of L1 models (19); training each L2 model (25.i) by the selected L1 subset of L1 models (19m.i); determining, for each trained L2 model (25.i), an L2 predicted yield value, based on the L1 testset (14); selecting an L2 subset of L2 models (25.i); determining the yield estimation of crops of the field (200) by applying the selected L1 subset of L1 models (19m.i) and the selected L2 subset of L2 models (25.i) on a time-series of unseen processed remote sensing data of the field (200).