An ultra-high yield corn yield estimation method based on remote sensing time series curve characteristic parameters
By dividing the corn NDVI time series curve into growth and decline stages, fitting the logistic curve and processing the data, and constructing a spectral characteristic model for the entire growth period, the saturation problem of the corn yield estimation model in the high-yield stage in the existing technology is solved, and the accurate prediction of super-high-yield corn yield is achieved.
Patent Information
- Application Number
- CN202411482264.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-23
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-10-23
AI Technical Summary
Existing corn yield estimation models are prone to saturation in the high-yield stage, and most of them are based on single growth period predictions, lacking the comprehensive utilization of spectral characteristics throughout the entire growth period, resulting in inaccurate yield estimation.
The corn NDVI time series curve was divided into growth and decline stages and fitted into two logistic curves. Parameters with high correlation were selected as inputs to the machine learning regression model. The data were processed by linear interpolation and exponentially weighted moving average to construct a spectral characteristic model for the entire growth period.
It effectively overcomes the saturation limitations of the yield estimation model, achieves accurate estimation of super-high-yield corn yield, improves the stability and flexibility of the model, and is suitable for yield prediction of super-high-yield corn.
Smart Images

Figure CN119314062B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application discloses a high-yield corn yield estimation method based on remote sensing time sequence curve characteristic parameters and belongs to the technical field of corn yield estimation. BACKGROUND
[0002] Spectral feature analysis and yield modeling are performed on the high-yield corn to realize accurate monitoring of the growth of the high-yield corn and yield estimation of the field scale, ensure stable grain yield and increase, and effectively support agricultural production management and meet the urgent needs of the development of precision agriculture. At present, the methods for crop yield remote sensing estimation mainly include three types: 1. an empirical statistical model, which estimates the linear relationship between a prediction variable and crop yield in a given data set, is simple to calculate and has strong explanatory ability; 2. a data assimilation model, which integrates remote sensing information into a crop growth model to minimize the difference between remote sensing observation and model variables; and 3. a semi-mechanism model, which simplifies the parameter input relative to a mechanism model. Through comprehensive analysis of the foregoing research, the prior art does not well solve the problem that the existing corn yield estimation model is prone to saturation in the high-yield stage, and most of the models are based on a single growth period for yield prediction, and there are few studies on the use of time sequence data to comprehensively extract spectral features in the whole growth period for corn yield estimation. SUMMARY
[0003] The application aims to provide a high-yield corn yield estimation method based on remote sensing time sequence curve characteristic parameters to solve the problem of inaccurate corn yield estimation in the prior art.
[0004] A high-yield corn yield estimation method based on remote sensing time sequence curve characteristic parameters, which comprises the following steps: dividing the NDVI time sequence curve of corn into a growth stage and a decline stage, fitting the two stages into two Logistic curves respectively, selecting parameters with high correlation in the two Logistic curves as inputs of a machine learning regression model, and performing yield prediction.
[0005] The Logistic curve comprises:
[0006] ;
[0007] In the formula, is the upper limit value of the Logistic curve, representing the maximum NDVI value of the corn in the growth season, is the midpoint of the Logistic curve, corresponding to the middle period of the corn growth season, is the growth rate, determining the steepness of the Logistic curve near , and controlling the change rate of NDVI in the middle period of the growth season, It is the baseline offset value, which is used to adjust the vertical position of the Logistic curve and represents the basic NDVI value of corn in the non-growing season. is the fitted NDVI, is a natural constant, It is the cumulative day of the corn growing period.
[0008] The parameters of the two logistic curves include 、 、 、 、 、 、 、 ,in 、 、 、 The growth phase 、 、 、 , 、 、 、 The descending phase 、 、 、 .
[0009] The parameters with the highest correlation between the two logistic curves are 、 、 、 、 .
[0010] When using a machine learning regression model, the training set and validation set are divided by traversing k-fold cross validation to obtain the fold division.
[0011] Before performing NDVI time series curve division, the NDVI time series curve is smoothed.
[0012] Use linear interpolation to smooth the NDVI time series curve:
[0013] ; ;
[0014] Where, is the coordinate of the point to be interpolated on the NDVI time series curve, and are two known data points on the NDVI time series curve.
[0015] Use the exponentially weighted moving average method to smooth the NDVI time series curve:
[0016] ; ;
[0017] Where, yes The smoothed value at time, is the smoothing factor of the exponentially weighted moving average, It's time The original observation value of yes The smoothed value at time, Represents the window size of the exponentially weighted moving average.
[0018] Compared with the existing technology, the present invention has the following beneficial effects: the present invention fully utilizes the significant difference law of the time when the spectrum saturation is not reached during the growth period of corn, effectively overcomes the saturation limitation of the yield estimation model, and provides a stable and reliable model and method for accurate estimation of corn yield in the super-high-yield part. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a technical flow chart of the present invention;
[0020] Figure 2 Schematic diagram of logistic curve fitting;
[0021] Figure 3 Scatter plot for accuracy verification. DETAILED DESCRIPTION
[0022] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention are described clearly and completely below. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0023] A super-high-yield corn yield estimation method based on the characteristic parameters of remote sensing time series curves includes dividing the corn NDVI time series curve into a growth phase and a decline phase, fitting the two phases into two logistic curves respectively, and selecting the parameters with high correlation in the two logistic curves as the input of the machine learning regression model to predict the yield.
[0024] The logistic curve includes:
[0025] ;
[0026] Where, is the upper limit of the Logistic curve, indicating the maximum NDVI value reached by corn during the growing season. is the logistic curve The midpoint corresponds to the middle of the corn growing season. Is the growth rate, which determines the Logistic curve The steepness of the surrounding slopes controls the rate of change of NDVI in the middle of the growing season. It is the baseline offset value, which is used to adjust the vertical position of the Logistic curve and represents the basic NDVI value of corn in the non-growing season. is the fitted NDVI, is a natural constant, It is the cumulative day of the corn growing period.
[0027] The parameters of the two logistic curves include 、 、 、 、 、 、 、 ,in 、 、 、 The growth phase 、 、 、 , 、 、 、 The descending phase 、 、 、 .
[0028] The parameters with the highest correlation between the two logistic curves are 、 、 、 、 .
[0029] When using a machine learning regression model, the training set and validation set are divided by traversing k-fold cross validation to obtain the fold division.
[0030] Before performing NDVI time series curve division, the NDVI time series curve is smoothed.
[0031] Use linear interpolation to smooth the NDVI time series curve:
[0032] ; ;
[0033] Where, is the coordinate of the point to be interpolated on the NDVI time series curve, and are two known data points on the NDVI time series curve.
[0034] Use the exponentially weighted moving average method to smooth the NDVI time series curve:
[0035] ; ;
[0036] Where, yes The smoothed value at time, is the smoothing factor of the exponentially weighted moving average, It's time The original observation value of yes The smoothed value at time, Represents the window size of the exponentially weighted moving average.
[0037] The main drawbacks of existing technologies are: Spectral reflectance saturation in densely vegetated areas is one of the major contemporary challenges facing the remote sensing community. Existing remote sensing yield estimation models are prone to saturation during high-yield stages, leading to underestimation of these high-yield areas. Crop growth is an allometric process, and existing remote sensing yield estimation models are either limited to a single stage or exhibit inconsistent performance across stages. They often overlook the significant differences between crops at different growth stages, lack transferability across the entire growth period, and lack research on using time series data to comprehensively extract spectral signatures throughout the entire growth period for crop yield estimation.
[0038] The present invention comprehensively utilizes remote sensing image data and agricultural knowledge, constructs daily vegetation index time series data throughout the entire growth period of corn through linear interpolation and exponentially weighted moving average smoothing methods, analyzes in detail the differences and changes in the vegetation index time series curves of corn with different yields, divides the time series curves into a growth period and a decline period, and explores the key parameters hidden in the vegetation index time series curves of corn with different yields that can well describe the dynamic changes of the curves by fitting them into two logistic curves. Parameters with high correlation with yield are used as inputs to the regression method, and a yield estimation model is constructed that can solve the problem that common corn yield estimation models are easily saturated in the high-yield part and consider the spectral time series characteristics of corn throughout the growth period, thereby realizing spectral time series characteristic analysis and yield modeling research for super-high-yield corn.
[0039] Based on existing remote sensing platforms and sensors, spectral information is acquired and preprocessed throughout the corn growing season. The presence of clouds and cloud shadows can distort ground spectral information, severely impacting the accuracy of vegetation index time series reconstruction. Therefore, most time series applications of optical satellite data require accurate cloud and shadow masks. This paper uses dataset products and satellite data available throughout the corn growing season as an example. Cloud removal is performed using the dataset products' native bands. Clouds, cloud shadows, water, and snow are classified based on a set of top-of-atmosphere reflectance thresholds. A cloud mask is defined: pixels with a median value of 4 (shadow) and 5 (cloud) are considered cloud. Cloud removal is then performed across all images in the time series dataset. Cloud removal is performed using the surface reflectance dataset acquired from satellite data by combining multiple cloud detection metrics. First, the Normalized Difference Water Index (NDWI) is defined, and a cloud probability dataset is obtained. The data quality assessment (QA60) band is then used to assess the data quality. Pixels with an NDWI less than 0.2, a cloud probability less than 50, and a QA60 non-cloud marker are considered the final target pixels.
[0040] The estimation of vegetation characteristics is achieved by using the vegetation index VI derived from remote sensing data obtained from satellites, airborne or ground platforms. VI is calculated based on the red and near-infrared parts of the electromagnetic spectrum and has been widely used in qualitative and quantitative remote sensing monitoring of vegetation vitality and growth dynamics. The present invention takes the normalized difference vegetation index NDVI as an example. NDVI is an effective indicator for monitoring vegetation growth and changes. By analyzing the NDVI time series, the changes in corn growing season can be observed, including the beginning of growth, peak growth and withering stages, which are closely related to corn yield. The NDVI time series can also be used to extract the characteristics of the growing season, such as growth rate, growth duration and seasonal changes in growth. These characteristics can be used to analyze the growth of corn in different regions or different years, and then predict the changing trend of corn yield.
[0041] When processing vegetation index time series data, cloud detection algorithms may result in some null values. To correct these null values, interpolation and smoothing are commonly used techniques to fill missing values, smooth noise, or adjust the frequency of the data, thereby generating temporally continuous and spatially complete daily vegetation index time series data for the entire growth period of corn. This paper takes the use of linear interpolation and exponentially weighted moving average smoothing methods to construct a vegetation index time series as an example. Linear interpolation is a simple and commonly used interpolation method, suitable for processing time series data due to its simplicity, computational efficiency, and accuracy when processing steadily changing data. In linear interpolation, it is assumed that the data change between two adjacent data points is linear, and this linear relationship is then used to infer the value of the unknown point. The exponentially weighted moving average (EWMA) is a commonly used time series data smoothing method, often used to eliminate noise and fluctuations in the data to better observe the data trend. EWMA estimates the mean of the data by taking a weighted average of the time series data, assigning higher weights to recent data and lower weights to long-term data.
[0042] The present invention takes the NDVI time series curve as an example. The overall trend of the NDVI time series curve of corn with different yield levels is consistent, showing an increase before the end of the jointing stage and a decrease after the early stage of the waxy stage, and maintaining a high-level NDVI saturation platform between the end of the jointing stage and the early stage of the waxy stage. In the growth phase of the NDVI time series curve, the higher the yield level, the larger the NDVI value on the same date, the more intense the degree of increase after entering the jointing stage, and the larger the slope. In the high-level phase of the NDVI time series curve, corn is in the large trumpet stage and the silking stage, and performs flowering and pollination stages, when the NDVI value reaches a saturation level. In the declining phase of the NDVI time series curve, the higher the yield level, the larger the NDVI value on the same date, the more intense the degree of reduction after entering the maturity stage, and the larger the slope.
[0043] This invention uses remote sensing technology to predict corn yields, which is of great significance for ensuring food security, promoting the development of corn yield prediction models, and guiding agricultural production. However, most existing models are affected by spectral saturation effects, making it difficult to accurately estimate super-high-yield corn production, and they do not adequately consider the spectral time series characteristics of corn throughout its entire growth period. The key points of this invention are: comprehensively considering the spectral time series characteristics of corn throughout its growth period, detailed analysis of the dynamic differences in the corn vegetation index time series curves at different yield levels, and extracting characteristic parameters describing the curve limits, slope, inflection points, and offset by fitting a logistic curve. These parameters are then linked to yield to predict yield. The super-high-yield corn remote sensing monitoring model constructed in this invention combines the differences in the corn vegetation index time series curves and the spectral saturation law. It has certain mechanistic properties and can eliminate the problem of inaccurate estimation caused by not considering the differences between corn production stages. It can be used to predict super-high-yield corn yields and effectively address the problem of super-high-yield corn being easily underestimated. The method uses characteristic parameters that can effectively describe the slope and inflection points of the time series curve as input variables, and is highly reliable and flexible. By combining multi-source remote sensing data to achieve accurate estimation of corn yield in the super-high-yield stage, it will help promote the development of corn yield prediction models and achieve timely and accurate prediction of super-high-yield corn yields. This is of great significance for the scientific formulation of import and export decisions, grain market prices and trade, agricultural insurance assessment, and smart agriculture applications.
[0044] The technical process of the present invention is as follows Figure 1 As shown in the figure, data collection is completed by collating remote sensing data and measured yield data, and data preprocessing, time series data calculation, linear interpolation and exponentially weighted moving average are performed in sequence to complete the construction of vegetation index time series. Then, time series curve difference analysis, logistic curve fitting and characteristic parameter extraction, regression analysis and k-fold cross validation are performed in sequence to complete the characteristic parameter advance and modeling.
[0045] The logistic curve fitting schematic diagram of the present invention is as follows Figure 2 As shown, Figure 2 In this paper, NDVI was cut in the middle and fitted separately to obtain 8 fitting parameters. Then, correlation analysis was performed to select 5 fitting parameters for corn yield estimation. The accuracy verification scatter plot is shown in the figure below. Figure 3 As shown in the figure, the oblique line represents the straight line that the predicted yield and the measured yield are completely consistent. According to the scatter plot, it can be seen that the prediction results of the present invention are basically attached to both sides of the oblique line. After calculation, the square residual coefficient R 2 =0.72, root mean square error RMSE=1087.22kg / ha.
[0046] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents, and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for estimating super-high-yield corn yield based on characteristic parameters of remote sensing time series curves, characterized in that: This involves dividing the corn NDVI time series curve into a growth phase and a decline phase, fitting the two phases into two logistic curves, and selecting the parameters with high correlation in the two logistic curves as the input of the machine learning regression model to predict yield. The logistic curve includes: ; Where, is the upper limit of the Logistic curve, indicating the maximum NDVI value reached by corn during the growing season. It is the logistic curve The midpoint corresponds to the middle of the corn growing season. is the growth rate, which determines the logistic curve The steepness of the surrounding slopes controls the rate of change of NDVI in the middle of the growing season. It is the baseline offset value, which is used to adjust the vertical position of the Logistic curve and represents the basic NDVI value of corn in the non-growing season. is the fitted NDVI, is a natural constant, It is the accumulated days of corn growing period; The parameters of the two logistic curves include 、 、 、 、 、 、 、 ,in 、 、 、 The growth phase 、 、 、 , 、 、 、 The descending phase 、 、 、 ; The parameters with the highest correlation between the two logistic curves are 、 、 、 、 .
2. The method for estimating super-high-yield corn yield based on characteristic parameters of remote sensing time series curve according to claim 1, characterized in that: When using a machine learning regression model, the training set and validation set are divided by traversing k-fold cross validation to obtain the fold division.
3. The method for estimating super-high-yield corn yield based on characteristic parameters of remote sensing time series curve according to claim 2, characterized in that: Before performing NDVI time series curve division, the NDVI time series curve is smoothed.
4. The method for estimating super-high-yield corn yield based on characteristic parameters of remote sensing time series curve according to claim 3, characterized in that: Use linear interpolation to smooth the NDVI time series curve: ; ; Where, is the coordinate of the point to be interpolated on the NDVI time series curve, and are two known data points on the NDVI time series curve.
5. The method for estimating super-high-yield corn yield based on characteristic parameters of remote sensing time series curve according to claim 4, characterized in that: Use the exponentially weighted moving average method to smooth the NDVI time series curve: ; ; Where, yes The smoothed value at time, is the smoothing factor of the exponentially weighted moving average, It's time The original observation value of yes The smoothed value at time, Represents the window size of the exponentially weighted moving average.
Citation Information
Patent Citations
Crop yield estimation model capable of reconstructing VI time series curve based on Extreme mathematical model
CN107122739A
Machine Learning for Production Prediction
US20190024494A1