Yield Data Outlier Detection for Accurate Harvest Maps
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Raw yield maps in agriculture often contain errors and inaccuracies, with up to 50% of observations being incorrect, which hinders precise agricultural management practices such as seeding schedules, irrigation, and fertilizer application.
Innovation Solution
A computer-based library and pipeline system for automatic outlier detection in yield data, employing discrete derivative analysis, local difference approaches, surface area calculations, and statistical spatial outlier detection to identify and remove erroneous data points, resulting in decontaminated yield maps.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual review and verification of harvested data is performed, then data accuracy is improved, but labor costs and processing time increase
Solution Approach 1:
The system performs automatic outlier detection and validation of harvested data using computational algorithms, eliminating the need for manual review. The processor automatically identifies outliers through statistical analysis and machine learning models, allowing the system to self-validate data quality without human intervention, thus reducing both time and labor costs while maintaining high accuracy
Solution Approach 2:
Manual mechanical review processes are replaced with automated electronic data processing systems. The patent employs computer-based algorithms, statistical models, and machine learning techniques to detect outliers and validate data, substituting human labor with automated computational mechanisms that operate faster and more consistently
2Reliability
If comprehensive data validation is performed, then data quality is improved, but computational resources and processing complexity increase
Solution Approach 1:
The data validation process is divided into distinct modular stages: initial data harvesting, outlier detection through statistical analysis, machine learning-based validation, and final quality assessment. Each stage processes specific aspects of data quality independently, allowing comprehensive validation without overwhelming computational complexity, as each module can be optimized and executed separately
Solution Approach 2:
The system performs preliminary outlier detection using statistical methods before applying more computationally intensive machine learning algorithms. By pre-identifying and filtering obvious outliers through simpler statistical tests, the system reduces the burden on complex validation algorithms, enabling comprehensive data quality checks while managing computational resources efficiently
Data Source
Figure 1
Figure 2(a)~2(b)
Figure 3
AI summary
In an embodiment, a method comprises determining, in received yield data, one or more passes, each pass including a plurality of observations. For each pass of the one or more passes, one or more discrete derivatives are determined, and based on the one or more discrete derivatives first outlier data is generated. First filtered data is generated by removing the first outlier data from the yield data. Furthermore, for each observation in the yield data, a plurality of nearest neighbor observations is determined, and used to determine a plurality of absolute differences in yield values. Based on the plurality of absolute differences, second outlier data is determined. Second filtered data is generated by removing the second outlier data from the first filtered data. Using a presentation layer of a computer system, a graphical representation of the second filtered data is generated and displayed on the computing system.