A method and system for predicting nitrogen content of multiple crops in a large-scale environment
By combining drone and satellite remote sensing imagery, vegetation indices were screened and an integrated model was constructed for bias correction, solving the accuracy problem of predicting nitrogen content of multiple crops in a large-scale environment and achieving improvements in high accuracy and applicability.
Patent Information
- Application Number
- CN202511564271.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-30
- Publication Date
- 2026-03-03
- Estimated Expiration
- 2045-10-30
AI Technical Summary
Existing technologies struggle to achieve high-precision prediction of nitrogen content in multiple crops on a large scale, especially in complex application scenarios involving multiple crop types. Neural network-based prediction models cannot meet accuracy requirements, and physical models have complex structures and high parameter uncertainties.
High-resolution remote sensing images were acquired using drones and combined with satellite remote sensing images. By constructing a statistical model of vegetation index and nitrogen content, highly correlated vegetation indices were selected, an integrated model was built, and bias correction was performed to achieve high-precision prediction.
It achieves high-precision prediction of nitrogen content in multiple crops under large-scale conditions, improves applicability in complex planting areas, and ensures the accuracy and reliability of prediction results through bias correction.
Smart Images

Figure CN121033698B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multi-crop nitrogen content estimation technology, and in particular to a method and system for predicting multi-crop nitrogen content under large-scale conditions. Background Technology
[0002] Changes in nitrogen content within crops directly affect chlorophyll content, cell structure, and various biochemical components in leaves. These physiological and biochemical alterations further lead to specific responses in the spectral reflectance characteristics of crop canopies across different wavelengths, including visible, near-infrared, and short-wave infrared. By constructing statistical models, physical models, and neural network-based prediction models relating crop canopy spectral characteristics (such as various vegetation indices) to measured nitrogen content, quantitative inversion estimation of nitrogen content in large-area crops can be achieved.
[0003] Compared to statistical models, physical models based on radiative transfer processes, while considering crop canopy structure and radiative transfer processes from a physical perspective, are complex in structure, have stringent requirements for the accuracy of input data, and suffer from parameter uncertainties in the inversion process, making them difficult to apply in practical production. While most neural network-based prediction models, as data-driven methods, avoid complex physical mechanism modeling, they rely on single or a few spectral features. These models may perform well under ideal, homogeneous field conditions, but in complex application scenarios involving large-scale, multi-crop environments, they cannot achieve high-precision prediction of nitrogen content across multiple crops. Summary of the Invention
[0004] Purpose of the invention: The purpose of this invention is to provide a method and system for predicting nitrogen content of multiple crops under large-scale conditions, so as to achieve accurate prediction of nitrogen content of multiple crop types under large-scale conditions.
[0005] Technical solution: To achieve the above objectives, the present invention provides a method for predicting nitrogen content in multiple crops on a large scale, comprising the following steps:
[0006] S1. Within a specific time period, for multiple target crops in the study area, acquire satellite remote sensing images covering the entire study area and the actual nitrogen content of all target crops, while using drones to acquire remote sensing images of different target crops respectively.
[0007] S2. For each target crop, based on the UAV remote sensing image data, calculate the vegetation index that is initially correlated with the measured nitrogen content, and select the target vegetation index related to the UAV remote sensing image from the initially correlated vegetation index.
[0008] S3. Construct a dataset in the form of "target vegetation index - measured nitrogen content" for each target crop based on UAV remote sensing images, which will be used to train the constructed crop nitrogen content prediction ensemble model.
[0009] S4. Based on the satellite remote sensing image, obtain the target vegetation index of each target crop and input it into the trained crop nitrogen content prediction ensemble model to obtain the initial prediction results of the nitrogen content of each target crop based on the satellite remote sensing image.
[0010] S5. Based on the deviation between the initial prediction results and the measured nitrogen content of each target crop, the initial prediction results are corrected to obtain the nitrogen content prediction results of multiple crops under a large-scale environment.
[0011] Preferably, the method for selecting the target vegetation index for UAV remote sensing imagery from the initial correlation vegetation index as described in S2 is as follows: first, perform collinearity test on all initial correlation vegetation indices, and then select the target vegetation index from the vegetation indices that pass the collinearity test.
[0012] Preferably, the correlation between any two initially correlated vegetation indices is calculated using the Pearson correlation coefficient. When the correlation coefficient between the two initially correlated vegetation indices is greater than a set correlation threshold, it is determined that there is a collinearity problem between the two initially correlated vegetation indices. The Pearson correlation coefficient is then used to calculate the correlation between the two vegetation indices with collinearity problems and the measured nitrogen content of the target crop. Vegetation indices with low correlation are eliminated, while vegetation indices with high correlation are considered to have passed the collinearity test. The vegetation indices that have passed the collinearity test are constructed into a variable set, and the target vegetation index is selected from the variable set using an iterative random regression forest model.
[0013] Preferably, the screening method is as follows: In the variable set, vegetation indices are sorted from high to low according to the Pearson correlation coefficient between the vegetation index and the measured nitrogen content of the target crop, and an empty set of the optimal feature set is constructed simultaneously; in descending order of the sorting, a vegetation index is selected from the variable set and added to the optimal feature set, and the optimal feature set is input into the iterative stochastic regression forest model to obtain the first coefficient of determination between the predicted nitrogen content of the target crop output by the stochastic regression forest model based on the optimal feature set and the measured nitrogen content of the target crop. If the first coefficient of determination obtained after adding the current vegetation index is greater than the first coefficient of determination obtained previously, then the currently added vegetation index is retained in the best feature set. Otherwise, the vegetation index is removed from the best feature set, and the selection process is repeated until all vegetation indices in the variable set have been input into the random regression forest model for regression and the first coefficient of determination is output. At this point, the features in the best feature set are the selected target vegetation indices.
[0014] Preferably, the crop nitrogen content prediction ensemble model in S3 consists of multiple target sub-models. The target sub-models are selected from the initial sub-models based on evaluation indicators, including mean squared error (MSE), root mean square error (RMSE), mean absolute error (MAE), and second coefficient of determination (R²). The initial sub-models include decision trees, random forests, XGBoost, LightGBM, support vector machines, and gradient boosting decision trees.
[0015] Preferably, the method for selecting target sub-models from initial sub-models based on evaluation metrics is as follows: the dataset is divided into a training set and a test set. The initial sub-models are trained independently using the training set. After training, the mean squared error (MSE), root mean square error (RMSE), mean absolute error (MAE), and second coefficient of determination (R²) of each initial sub-model are calculated on the test set. The initial sub-models with the highest accuracy in MSE, RMSE, MAE, and R² are defined as target sub-models A through D.
[0016] Preferably, the method for constructing the crop nitrogen content prediction integrated model from multiple target sub-models is as follows: calculate the weight values of the target sub-models A to D with respect to the evaluation indicators including mean square error (MSE), root mean square error (RMSE), mean absolute error (MAE), and the second coefficient of determination (R²).
[0017] For each evaluation index, the maximum weight value is selected to form the positive ideal solution PIS, and the minimum weight value is selected to form the negative ideal solution NIS, as expressed respectively:
[0018] , ;
[0019] Calculate the distances from the target sub-models A to D to the positive and negative ideal solutions respectively. , Calculate the relative proximity of target sub-models A through D based on their distances. Calculate the weights of the target sub-models A through D based on their relative tracking progress:
[0020] ,
[0021] Where i = 1~4, representing the target sub-models A~D respectively. , Let represent the distances from the i-th objective sub-model to the positive and negative ideal solutions, respectively. This represents the relative similarity of the i-th target sub-model. This represents the sum of the relative similarities between the target sub-models A to D;
[0022] Assign weights to the outputs of the target sub-models A through D respectively. The sum of the outputs of each objective sub-model multiplied by its corresponding weight is the output of the integrated crop nitrogen content prediction model composed of objective sub-models A to D. , is represented as:
[0023] ,
[0024] in, ~ These are the output results of the target sub-models A through D, respectively.
[0025] Preferably, the bias correction method described in S5 is as follows: calculating the average systematic bias (Bias) of the crop nitrogen content prediction ensemble model.
[0026] ,
[0027] in, Indicates the number of samples. Indicates the first One sample, y ' y represents the predicted value from the integrated model for crop nitrogen content prediction. true This represents the measured value of crop nitrogen content;
[0028] Predicted values from the ensemble model for crop nitrogen content prediction after correction Represented as: .
[0029] The present invention discloses a multi-crop nitrogen content prediction system for large-scale environments, comprising the following modules:
[0030] Data acquisition module: used to acquire satellite remote sensing images covering the entire study area and the actual nitrogen content of all target crops within a specific time period for multiple target crops in the study area, while using drones to acquire remote sensing images of different target crops separately;
[0031] Target vegetation index screening module: For each target crop, based on the UAV remote sensing image data, calculate the vegetation index that is initially correlated with the measured nitrogen content, and screen out the target vegetation index related to the UAV remote sensing image from the initially correlated vegetation index.
[0032] Predictive ensemble model building module: used to build a dataset in the form of "target vegetation index - measured nitrogen content" of each target crop based on UAV remote sensing imagery, which is used to train the constructed crop nitrogen content prediction ensemble model;
[0033] The target crop nitrogen content prediction module based on satellite remote sensing imagery is used to obtain the target vegetation index of each target crop based on satellite remote sensing imagery, input it into the trained crop nitrogen content prediction ensemble model, and obtain the initial prediction results of the nitrogen content of each target crop based on the satellite remote sensing imagery.
[0034] Target crop nitrogen content prediction and correction module: used to correct the deviation between the initial prediction results and the measured nitrogen content of each target crop, and obtain the nitrogen content prediction results of multiple crops in a large-scale environment.
[0035] Beneficial Effects: This invention has the following advantages: 1. This invention integrates high-resolution UAV data and large-scale satellite data. First, it uses UAV data to construct a high-precision nitrogen content prediction model, and then transfers the model to the satellite scale, achieving high-precision prediction of nitrogen content for multiple crops in a large-scale environment; 2. It selects target vegetation indices that are highly correlated with nitrogen content for different crops, solving the problem that existing methods rely on single or a few spectral features and are difficult to adapt to the physiological characteristics of multiple crops, significantly improving the applicability of the method in complex planting areas; 3. It introduces bias correction to correct the error of the preliminary prediction results at the satellite scale, effectively eliminating the bias caused by data source and model generalization, ensuring that the final prediction results are accurate and reliable. Attached Figure Description
[0036] Figure 1 This is a schematic diagram of the method flow of the present invention. Detailed Implementation
[0037] The technical solution of the present invention will be described in detail below with reference to the embodiments and accompanying drawings.
[0038] like Figure 1 As shown in this embodiment, a method for predicting nitrogen content in multiple crops under large-scale conditions is provided, including the following:
[0039] S1. Within a specific time period, for multiple target crops within the study area, acquire satellite remote sensing images covering the entire study area. At the same time, use drones to acquire remote sensing images of different target crops separately, and collect samples of all target crops for actual nitrogen content determination.
[0040] Optionally, each target crop is divided into several collection areas based on its spatial distribution characteristics, and each collection area contains multiple sets of collection points.
[0041] Optionally, the satellite remote sensing imagery is extracted from the Google Earth Engine data cloud platform, and the Sentinel-2 remote sensing imagery data is synchronized with the UAV remote sensing imagery and covers the same study area. The resolution of the Sentinel-2 remote sensing imagery can be set to 10m.
[0042] Furthermore, preprocessing was performed on the acquired UAV and satellite remote sensing image data. Specifically, during the UAV remote sensing image acquisition process, radiometric calibration was performed using a 27mm diameter diffuse reflectance standard white board (model JY-WS1). Based on the acquired image data and the reflectance data of the standard white board, radiometric correction was performed using an empirical linear method. After stitching multiple UAV remote sensing images into a complete image using Agisoft Metashape software, further radiometric correction was performed on the stitched image in ENVI 5.3 software to improve data accuracy. The stitched image was cropped according to the total study area, and then the data format was converted to a TIFF image file containing 6 bands.
[0043] For satellite remote sensing imagery, data from areas obscured by clouds is first excluded. This is achieved by using the QA60 quality assessment band for cloud masking, followed by removal of pixels covered by clouds and cirrus clouds. The de-clouded image data is then scaled for reflectance with a scaling factor of 0.0001 to convert the image data to surface reflectance within the range of 0 to 1. Finally, the image is cropped according to the total study area, and the data is converted to a TIFF image file containing 13 bands: coastal / aerosol B1 band, blue light B2 band (B), green light B3 band (G), red light B4 band (R), red-edge B5 band, red-edge B6 band (RE), red-edge B7 band (NIR), wide near-infrared B8 band, narrow near-infrared B8A band, water vapor B9 band, short-edge infrared B10 band, short-edge infrared B11 band, and short-edge infrared B12 band.
[0044] Furthermore, the average nitrogen content of the target crop within the same collection area was taken as the measured nitrogen content data for that area. The determination process was as follows: samples such as leaves, stems, and seeds were extracted, washed, dried (at 130°C to constant weight), pulverized and sieved to ensure uniformity, and the samples were accurately weighed using an analytical balance. The nitrogen content was determined using the Kjeldahl method.
[0045] S2. For different target crops, based on UAV remote sensing image data, calculate vegetation indices that are initially correlated with measured nitrogen content, and select target vegetation indices related to UAV remote sensing images from the initially correlated vegetation indices.
[0046] The primary vegetation indices include: Normalized Difference Vegetation Index (NDVI), Excess Green Index (EXG), Renormalized Difference Vegetation Index (RDVI), Atmospheric Resistance Vegetation Index (ARVI), Normalized Difference Greenness Index (NDGI), Pigment Specific Reflectance Index (PSRI), Visible Light Atmospheric Impedance Index (VARI), Difference Vegetation Index (DVI), Simple Pigment Ratio Index (SRPI), Normalized Pigment Chlorophyll Index (NPCI), Normalized Difference Red Edge Index (NDRE), Modified Chlorophyll Absorption and Reflectance Index (MCARI), Greenness Index (GI), Green Normalized Difference Vegetation Index (GNDVI), Soil Adjusted Vegetation Index (SAVI), Transformed Vegetation Index (TVI), Enhanced Vegetation Index (EVI), Ratio Vegetation Index (RVI), Green-Red Vegetation Index (GRVI), Normalized Difference Snow Index (NDSI), Chlorophyll Red Edge Stress Index (CRSI), Modified Simple Ratio Index (MSR), Terrestrial Chlorophyll Index (MTCI), Optimized Soil Adjusted Vegetation Index (OSAVI), and b index.
[0047] The calculation methods for each index are as follows:
[0048] , , , , , , , , , , , , , , , , , , , , , , , ;
[0049] In the formula, NIR represents the near-infrared band reflectivity, R represents the red band reflectivity, G represents the green band reflectivity, B represents the blue band reflectivity, and RE represents the red edge band reflectivity.
[0050] Among them, the b index can enhance the ability to distinguish targets such as water bodies or vegetation through normalization. It is often used for inversion of suspended sediment content or auxiliary assessment of ecological environment quality. Specifically, the b index is higher for clear water bodies, while the b index is lower for turbid water bodies (containing a large amount of sediment, with enhanced green / red reflection).
[0051] The method for selecting target vegetation indices from UAV remote sensing images based on initial correlation vegetation indices is as follows:
[0052] First, a collinearity test was performed on all initial correlation vegetation indices. Specifically, the Pearson correlation coefficient was used to calculate the correlation between any two initial correlation vegetation indices. The calculation formula is as follows:
[0053] ,
[0054] in, This represents the Pearson correlation coefficient between the p-th vegetation index and the q-th vegetation index within a certain collection area. Its value ranges from [-1, 1]. The closer the value is to 1, the stronger the linear relationship between the two indices and the more serious the collinearity problem. Indicates that within a certain collection area, at the th The p-th vegetation index value at each collection point; The table represents the number of data points within a certain data collection area. The qth vegetation index value at each collection point; This represents the average value of the p-th vegetation index across all collection points within a given collection area. This represents the average value of the q-th vegetation index across all collection points within a certain collection area. This represents the total number of collection points within a specific collection area. =1~ , indicating the first One collection point.
[0055] If the correlation coefficient between two vegetation indices exceeds the set correlation threshold (e.g., 0.8), it is considered that there is a collinearity problem between the two indices. In this case, the correlation between the two vegetation indices and the measured nitrogen content data of the target crop is tested. Specifically, the correlation between the two vegetation indices and the measured nitrogen content of the target crop is calculated using the Pearson correlation coefficient, and vegetation indices with low correlation are removed.
[0056] For the j-th vegetation index (denoted as X) j The correlation coefficient between the measured nitrogen content data of the target crop and the actual nitrogen content data (denoted as Y) is calculated using the following formula:
[0057] ,
[0058] in, This indicates that within a certain collection area, the first... Pearson correlation coefficient between a vegetation index and the measured nitrogen content Y of the target crop; Indicates that within a certain collection area, at the th The first collection point Individual vegetation index values; Indicates that within a certain collection area, the first... Measured nitrogen content values of the target crop at each collection point; Indicates that within a certain collection area, the first... The average of the vegetation indices across all collection points; This represents the average measured nitrogen content of the target crop at all sampling points within a specific sampling area. This indicates the total number of collection points within a certain collection area.
[0059] The higher the correlation coefficient between the vegetation index and the measured nitrogen content of the target crop, the higher the correlation between the vegetation index and the measured nitrogen content of the target crop, and the more suitable it is for subsequent analysis.
[0060] After removing all initially correlated vegetation indices through collinearity detection, the remaining vegetation indices are constructed into a variable set, and target vegetation indices are selected from this set. Optionally, an iterative random regression forest model is used to select the target vegetation indices, as follows:
[0061] 1. Variable clustering: Pearson correlation coefficients between all vegetation indices and the measured nitrogen content of the target crop. Sort by size from highest to lowest;
[0062] 2. Construct the optimal empty feature set;
[0063] 3. Following the order from highest to lowest, select one remotely sensed vegetation from the variable set and add it to the optimal feature set. Then, input the optimal feature set into the iterative stochastic regression forest model and obtain the first coefficient of determination between the predicted nitrogen content of the target crop output by the stochastic regression forest model based on the optimal feature set and the measured nitrogen content of the target crop. , The closer the value of is to 1, the better the effect of the current input optimal feature set;
[0064] 4. If the result is obtained after adding the current vegetation index Greater than any of the previous ones If the newly added index is selected, it will be retained in the best feature set; otherwise, the vegetation index will be removed from the best feature set.
[0065] 5. Repeat steps 3-4 until all indices in the variable set have been input into the random regression forest model for regression and output. At this point, the feature in the optimal feature set is the selected target vegetation index.
[0066] S3. Construct a dataset in the form of "target vegetation index - measured nitrogen content from UAV remote sensing images" for the same target crop, which will be used to train the constructed crop nitrogen content prediction ensemble model.
[0067] The predictive ensemble model consists of multiple sub-models. The initial sub-models include decision trees, random forests, XGBoost, LightGBM, support vector machines, and gradient boosting decision trees. The target sub-model with the best predictive performance is selected from the initial sub-models to form the predictive ensemble model. The selection process is as follows:
[0068] The dataset is divided into training and test sets. The training set is then input into each of the initial sub-models, allowing for independent training of each sub-model. During training, GridSearchCV is used to optimize the hyperparameters of the initial sub-models. After optimization, the predictive performance of all initial sub-models is comprehensively evaluated on the test set using evaluation metrics (Mean Squared Error MSE, Root Mean Squared Error RMSE, Mean Absolute Error MAE, and Second Coefficient of Determination R²). The evaluation metric calculation formula for any given initial sub-model is as follows:
[0069] , , , ;
[0070] in, This represents the target crop nitrogen content value predicted by the initial sub-model. This represents the mean nitrogen content of the target crop predicted by the initial sub-model. This represents the measured nitrogen content value of the target crop. This represents the average nitrogen content of the target crop as measured in actual measurements. This represents the total number of samples in the test set. Indicates the first in the test set One sample.
[0071] Four target sub-models were selected from the initial sub-models based on four evaluation metrics, and weights were assigned to each of the four target sub-models to form a predictive ensemble model. The specific process is as follows:
[0072] We select the initial sub-models with the highest accuracy in MSE, RMSE, MAE, and R², respectively. Among them, for MSE, RMSE, and MAE, we select the initial sub-models with the smallest calculation results and define them as target sub-model A, target sub-model B, and target sub-model C, respectively. For R², we select the initial sub-model with the largest calculation result and define it as target sub-model D.
[0073] The TOPSIS method is used to evaluate the overall performance of the target sub-models A to D, and corresponding weights are assigned to them based on the evaluation results. Specifically, the weight of each target sub-model with respect to the evaluation indicators MSE, RMSE, MAE, and R² is calculated using the following formula:
[0074] ,
[0075] in, This represents the weight of the i-th objective sub-model with respect to the j-th evaluation index. Let i represent the j-th evaluation index in the i-th sub-model, i=1~4, representing the target sub-models A~D respectively, and j=1~4, representing the evaluation indices MSE, RMSE, MAE, and R² respectively.
[0076] A total of 16 weight values were obtained through calculation, with each evaluation index comprising 4 weight values. For each evaluation index, the maximum and minimum weight values were selected from the calculation results. The 4 evaluation indicators corresponding to the maximum weight value constitute the positive ideal solution (PIS), and the 4 evaluation indicators corresponding to the minimum weight value constitute the negative ideal solution (NIS), as shown below:
[0077] ,
[0078] .
[0079] Calculate the distances from the target sub-models A to D to the positive and negative ideal solutions respectively, using the following formulas:
[0080] ,
[0081] .
[0082] The relative proximity of target sub-models A through D is calculated based on the distance. The relative proximity measures the relative position of the target sub-model between the "positive ideal solution" and the "negative ideal solution". The relative proximity value is between [0, 1]. The larger the value, the closer the target sub-model is to the positive ideal solution and the better its performance. The formula for calculating the relative proximity is as follows:
[0083] .
[0084] The weights of the target sub-models A through D are calculated based on their relative tracking progress, using the following formulas:
[0085] ;
[0086] Where i = 1~4, representing the target sub-models A~D respectively. , Let represent the distances from the i-th objective sub-model to the positive and negative ideal solutions, respectively. This represents the relative similarity of the i-th target sub-model. This represents the sum of the relative similarities between the target sub-models A to D.
[0087] Assign weights to the outputs of the target sub-models A through D respectively. The sum of the product of each target sub-model's output and its corresponding weight is the output of the prediction ensemble model consisting of target sub-models A through D. , is represented as:
[0088] ,
[0089] In the formula, ~ These are the output results of the target sub-models A through D, respectively.
[0090] S4. Obtain the target vegetation index of all target crops in the satellite remote sensing image, input it into the trained crop nitrogen content prediction ensemble model, and obtain the nitrogen content prediction of each target crop.
[0091] Optionally, target vegetation index data based on Sentinel-2 remote sensing images can be obtained in different acquisition areas based on the blue light B2 band (B), green light B3 band (G), red light B4 band (R), red edge B6 band (RE), red edge B7 band (NIR) in Sentinel-2 remote sensing image data and the target vegetation index calculation formula.
[0092] By inputting the target vegetation index based on Sentinel-2 remote sensing imagery into the target crop nitrogen content prediction ensemble model, the prediction results of multi-target crop nitrogen content under large-scale environment are obtained.
[0093] S5. Based on the deviation between the predicted and measured nitrogen content values of each target crop, the nitrogen content data predicted by the integrated crop nitrogen content prediction model is corrected for deviation. The specific process is as follows: A "predicted value y" is constructed based on a one-to-one correspondence between the sampling coordinate location and the remote sensing index data coordinate location. ’ -Measured value y true "Sample-based calibration set S" cal The average systematic bias (Bias) of the crop nitrogen content prediction ensemble model is calculated based on the calibration set, using the following formula:
[0094] ,
[0095] in, This indicates the number of samples in the calibration set. Indicates the first in the correction set One sample.
[0096] when A value greater than 0 indicates that the nitrogen content data predicted by the crop nitrogen content prediction ensemble model is generally too high, with the magnitude of the overestimation being the absolute value of the Bias; when... When the value is less than 0, it indicates that the nitrogen content data predicted by the crop nitrogen content prediction integrated model is generally too low, and the degree of overestimation is the absolute value of Bias.
[0097] use The data is uniformly corrected to the nitrogen content data predicted by the crop nitrogen content prediction ensemble model to obtain the corrected nitrogen content data. The correction formula is as follows:
[0098] .
[0099] S6. Real-time acquisition of satellite remote sensing images of multiple crop types under large-scale conditions, calculation of target vegetation indices for all target crops in the satellite remote sensing images, input into the trained crop nitrogen content prediction ensemble model to obtain nitrogen content prediction values, and then deviation correction of the nitrogen content prediction values to obtain nitrogen content prediction results for multiple crops under large-scale conditions.
[0100] This invention leverages the significantly higher resolution but smaller coverage of remote sensing imagery acquired by unmanned aerial vehicles (UAVs) compared to satellite-acquired imagery. First, a high-precision prediction model for nitrogen content in target crops is trained using the ultra-high resolution UAV-acquired remote sensing imagery. Then, this model is used to predict nitrogen content based on satellite remote sensing imagery covering multiple target crops, allowing the model to adapt to the distribution of satellite remote sensing imagery data. This enables nitrogen content prediction for different target crops based on satellite remote sensing imagery. Finally, bias corrections are applied to the prediction results, achieving accurate nitrogen content prediction for different crop types on a large scale.
Claims
1. A method for predicting nitrogen content in multiple crops under large-scale environments, characterized in that, Includes the following steps: S1. Within a specific time period, for multiple target crops in the study area, acquire satellite remote sensing images covering the entire study area, collect samples of all target crops for actual nitrogen content determination, and use drones to acquire remote sensing images of different target crops respectively. S2. For each target crop, based on the UAV remote sensing image data, calculate the vegetation index that is initially correlated with the measured nitrogen content, and select the target vegetation index related to the UAV remote sensing image from the initially correlated vegetation index. S3. Construct a dataset in the form of "target vegetation index - measured nitrogen content" for each target crop based on UAV remote sensing images, which will be used to train the constructed crop nitrogen content prediction ensemble model. S4. Based on the satellite remote sensing image, obtain the target vegetation index of each target crop and input it into the trained crop nitrogen content prediction ensemble model to obtain the initial prediction results of the nitrogen content of each target crop based on the satellite remote sensing image. S5. Based on the deviation between the initial prediction results of nitrogen content of each target crop and the measured nitrogen content, the initial prediction results are corrected for deviation to obtain the nitrogen content prediction of multiple crops under a large-scale environment. The method for bias correction is as follows: calculate the average systematic bias (Bias) of the crop nitrogen content prediction ensemble model. , in, Indicates the number of samples. Indicates the first One sample, y ' y represents the predicted value from the integrated model for crop nitrogen content prediction. true This represents the measured value of crop nitrogen content, and the predicted value from the ensemble model for crop nitrogen content prediction after correction. Represented as: .
2. The method for predicting nitrogen content of multiple crops under large-scale environments according to claim 1, characterized in that, The method described in S2 for selecting target vegetation indices for UAV remote sensing images from initial correlation vegetation indices is as follows: first, perform collinearity testing on all initial correlation vegetation indices, and then select the target vegetation indices from the vegetation indices that pass the collinearity test.
3. The method for predicting nitrogen content of multiple crops in a large-scale environment according to claim 2, characterized in that, The correlation between any two initially correlated vegetation indices is calculated using the Pearson correlation coefficient. If the correlation coefficient between the two initially correlated vegetation indices is greater than a set correlation threshold, it is determined that there is a collinearity problem between the two initially correlated vegetation indices. The Pearson correlation coefficient is then used to calculate the correlation between the two vegetation indices with collinearity problems and the measured nitrogen content of the target crop. Vegetation indices with low correlation are removed, while vegetation indices with high correlation are considered to have passed the collinearity test. The vegetation indices that have passed the collinearity test are constructed into a variable set, and the target vegetation index is selected from the variable set using an iterative random regression forest model.
4. The method for predicting nitrogen content of multiple crops in a large-scale environment according to claim 3, characterized in that, The screening method is as follows: In the variable set, vegetation indices are sorted from high to low according to the Pearson correlation coefficient between the vegetation index and the measured nitrogen content of the target crop, and an empty set of the optimal feature set is constructed simultaneously; in descending order of the sorting, a vegetation index is selected from the variable set and added to the optimal feature set, and the optimal feature set is input into an iterative stochastic regression forest model to obtain the first coefficient of determination between the predicted nitrogen content of the target crop output based on the optimal feature set and the measured nitrogen content of the target crop. ; If the first coefficient of determination obtained after adding the current vegetation index is greater than the previously obtained first coefficient of determination, then the currently added vegetation index is retained in the optimal feature set; otherwise, the vegetation index is removed from the optimal feature set. The selection process is repeated until all vegetation indices in the variable set have been input into the random regression forest model for regression and the first coefficient of determination is output. At this point, the features in the optimal feature set are the selected target vegetation indices.
5. The method for predicting nitrogen content of multiple crops under large-scale environments according to claim 1, characterized in that, The crop nitrogen content prediction ensemble model described in S3 consists of multiple objective sub-models. These objective sub-models are selected from the initial sub-models based on evaluation indicators, including mean squared error (MSE), root mean square error (RMSE), mean absolute error (MAE), and second coefficient of determination (R²). The initial sub-models include decision trees, random forests, XGBoost, LightGBM, support vector machines, and gradient boosting decision trees.
6. The method for predicting nitrogen content of multiple crops in a large-scale environment according to claim 5, characterized in that, The method for selecting target sub-models from initial sub-models based on evaluation metrics is as follows: Divide the dataset into training and test sets. Train the initial sub-models independently using the training set. After training, calculate the mean squared error (MSE), root mean square error (RMSE), mean absolute error (MAE), and second coefficient of determination (R²) for each initial sub-model on the test set. Select the initial sub-models with the highest accuracy in MSE, RMSE, MAE, and R² as target sub-models A through D.
7. The method for predicting nitrogen content of multiple crops in a large-scale environment according to claim 6, characterized in that, The method for constructing the crop nitrogen content prediction integrated model from multiple target sub-models is as follows: calculate the weight of each target sub-model A to D with respect to the evaluation indicators, including mean square error (MSE), root mean square error (RMSE), mean absolute error (MAE), and the second coefficient of determination (R²). For each evaluation index, the maximum weight value is selected to form the positive ideal solution PIS, and the minimum weight value is selected to form the negative ideal solution NIS, as expressed respectively: , ; Calculate the distances from the target sub-models A to D to the positive and negative ideal solutions respectively. , Calculate the relative proximity of target sub-models A through D based on their distances. Calculate the weights of the target sub-models A through D based on their relative tracking progress: , Where i = 1~4, representing the target sub-models A~D respectively. , Let represent the distances from the i-th objective sub-model to the positive and negative ideal solutions, respectively. This represents the relative similarity of the i-th target sub-model. This represents the sum of the relative similarities between the target sub-models A to D; Assign weights to the outputs of the target sub-models A through D respectively. The sum of the outputs of each objective sub-model multiplied by its corresponding weight is the output of the integrated crop nitrogen content prediction model composed of objective sub-models A to D. , is represented as: , In the formula, ~ These are the output results of the target sub-models A through D, respectively.
8. A multi-crop nitrogen content prediction system for large-scale environments, characterized in that, Includes the following modules: Data acquisition module: Used to acquire satellite remote sensing images covering the entire study area for multiple target crops within a specific time period, and to collect samples of all target crops for actual nitrogen content determination. At the same time, it uses drones to acquire remote sensing images of different target crops respectively. Target vegetation index screening module: For each target crop, based on the UAV remote sensing image data, calculate the vegetation index that is initially correlated with the measured nitrogen content, and screen out the target vegetation index related to the UAV remote sensing image from the initially correlated vegetation index. Predictive ensemble model building module: used to build a dataset in the form of "target vegetation index - measured nitrogen content" of each target crop based on UAV remote sensing imagery, which is used to train the constructed crop nitrogen content prediction ensemble model; The satellite remote sensing image-based target crop nitrogen content prediction module is used to obtain the target vegetation index of each target crop based on the satellite remote sensing image, input it into the trained crop nitrogen content prediction ensemble model, and obtain the initial prediction results of the nitrogen content of each target crop based on the satellite remote sensing image. Target crop nitrogen content prediction and correction module: used to correct the deviation between the initial prediction results and the measured nitrogen content of each target crop, and obtain the nitrogen content prediction of multiple crops under large-scale environment; The method for bias correction is as follows: calculate the average systematic bias (Bias) of the crop nitrogen content prediction ensemble model. , in, Indicates the number of samples. Indicates the first One sample, y ' y represents the predicted value from the integrated model for crop nitrogen content prediction. true This represents the measured value of crop nitrogen content, and the predicted value from the ensemble model for crop nitrogen content prediction after correction. Represented as: .
Citation Information
Patent Citations
Wheat nitrogen content measuring and calculating method based on hyperspectral remote sensing image of unmanned aerial vehicle
CN117074340A
Wheat LAI fine inversion method fusing unmanned aerial vehicle and satellite optical remote sensing vegetation index set
CN118736448A