A wheat scab prediction method and system based on multispectral remote sensing data and XGBoost algorithm
By combining multispectral remote sensing data and the XGBoost algorithm, the problems of insufficient timeliness and spatial coverage in predicting wheat scab disease severity in traditional methods have been solved, achieving high-precision field-level disease monitoring and prediction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGQIU NORMAL UNIVERSITY
- Filing Date
- 2026-05-12
- Publication Date
- 2026-07-24
AI Technical Summary
Existing technologies are insufficient for rapid, real-time prediction of wheat scab disease severity levels over a wide area with high spatial resolution. Traditional methods are time-consuming, labor-intensive, and have limited coverage. Existing models have failed to fully achieve disease level monitoring at the large-area field level.
A prediction method based on multispectral remote sensing data and the XGBoost algorithm was adopted. By acquiring multispectral remote sensing images, multispectral vegetation indices were calculated, a feature database was constructed, and the hyperparameters of the XGBoost model were optimized using a differential optimization algorithm to generate a field-level spatial distribution map of Fusarium head blight disease severity.
It enables large-scale, high-precision, field-scale prediction and spatial mapping of wheat scab disease, improving prediction accuracy and supporting real-time decision-making by grassroots agricultural extension workers and growers.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This invention relates to the field of agricultural remote sensing data processing technology, specifically to a method and system for predicting wheat scab based on multispectral remote sensing data and the XGBoost algorithm. Background Technology
[0002] Fusarium head blight is one of the major diseases affecting wheat during its growth period. It directly reduces the thousand-grain weight, leading to yield losses and, in severe cases, total crop failure, posing a potential threat to national food security. The pathogen infects the protein structure of wheat, reducing its processing performance and edible quality, resulting in economic losses. The vomitoxin and other toxins produced by infected wheat grains are biotoxic; consumption by humans and animals can cause vomiting, diarrhea, immunosuppression, and even teratogenic and carcinogenic effects, seriously threatening human health. Due to its explosive and irreversible nature, the disease not only increases pesticide application costs but also poses a significant challenge to the stable development of regional agricultural economies.
[0003] With current climate change, altered atmospheric conditions, and changes in planting systems, the affected area and severity of wheat scab in some regions are showing an increasing trend. In particular, preventing wheat scab is urgent and necessary in Xinyang and Zhumadian cities in Henan Province. Xinyang and Zhumadian are located in the Huaihe, Jianghuai, and Huanghuai river basins, a transitional zone between warm temperate and subtropical climates with a humid climate. The wheat heading and flowering period often coincides with prolonged periods of overcast and rainy weather in these areas, providing suitable natural conditions for the infection and outbreak of scab fungus. The long-term practice of straw return to the field in these areas has led to the accumulation of scab fungal perithecia in the soil year after year. Without early and scientific prevention, this can easily trigger an epidemic of wheat scab, directly threatening food security and the economic benefits of farmers.
[0004] Relying solely on manual field surveys for wheat scab prevention will lead to delays and irreparable losses. Therefore, there is an urgent need to develop detailed disease severity prediction methods for wheat scab, as well as methods for rapid, large-scale, field-level disease severity prediction.
[0005] Since previous studies have mainly considered wheat scab to be related to meteorological factors, traditional prevention and monitoring of wheat scab have mainly relied on manual field surveys, chemical control, and predictive models constructed using meteorological elements.
[0006] Traditional methods have some effect on the prevention of wheat scab, but they also have shortcomings. Field surveys involve plant protection personnel sampling and observing the symptoms on wheat ears, but this method is time-consuming, labor-intensive, has limited coverage, and often only detects symptoms after they appear, easily missing the optimal window for control. Chemical control and variety breeding can effectively reduce the incidence rate, but these methods involve indiscriminate spraying of chemical pesticides, which can increase economic expenditures. Variety breeding requires long-term field trials. Predictive models built using meteorological elements are limited by limited meteorological station data and low spatial resolution.
[0007] Domestically, some researchers have invented a wheat scab prediction method based on the PSO-BA algorithm, as well as a wheat scab risk prediction system and method using deep learning and random forests. While these methods achieve wheat scab prediction, they do not incorporate multi-source remote sensing data and fail to obtain a spatial distribution map of wheat scab disease severity in raster image form. Other researchers have invented a wheat scab detection method based on UAV remote sensing imagery. Although UAV imagery has high resolution, it cannot achieve spatial prediction of wheat scab disease severity over large areas.
[0008] Internationally, the forecasting and prediction of wheat scab started early and has developed into an integrated and information-based forecasting system. Meteorological data is the core foundation of foreign forecasting models, which comprehensively utilize biological monitoring, cropping system analysis, and advanced algorithmic models. International research institutions also apply satellite remote sensing and UAV multispectral imaging technology to wheat scab forecasting. However, these methods have not yet been able to fully achieve large-area, field-level forecasting of wheat scab severity.
[0009] Real-time prediction of wheat scab disease severity is of great strategic significance for ensuring national food security and sustainable agricultural development. Although current research has made some progress in understanding the disease's epidemiological mechanisms and evolution, it still faces scientific bottlenecks: traditional wheat scab disease severity surveys mainly rely on manual field surveys, which are slow, require significant manpower and financial investment, and cannot quickly and in real-time predict the severity of wheat scab over large areas. Relying solely on meteorological stations and low-spatial-resolution meteorological factors such as precipitation, temperature, and humidity makes it difficult to achieve high spatial resolution disease severity prediction over large areas. Existing models still have shortcomings in multi-source remote sensing information integration and high-resolution raster mapping, making it difficult to predict wheat scab disease severity from point to large area. While existing research has introduced algorithmic optimization, it has not yet fully achieved large-area field-level disease severity monitoring.
[0010] In recent years, the development of remote sensing monitoring technology and machine learning algorithms has provided important data and technical support for large-area, field-level, and high-precision early identification and spatial mapping of diseases, making up for the shortcomings of traditional methods in terms of timeliness and spatial coverage.
[0011] Traditional manual sampling and ground-based meteorological station monitoring are insufficient to cover the vast and fragmented farmland areas of major production regions. Landsat 8 / 9, a multispectral remote sensing satellite, offers advantages such as short revisit periods, rich spectral bands, high spatial resolution, and free access, enabling large-area, plot-level monitoring. The continuous development of machine learning algorithms provides algorithmic support for modeling with high predictive accuracy.
[0012] In summary, there is an urgent need for a method that can combine multispectral remote sensing big data and efficient optimization algorithms to achieve large-area, high-precision, field-scale prediction and spatial mapping of wheat scab disease. Summary of the Invention
[0013] To overcome the problems in the existing technology, the purpose of this invention is to provide a wheat scab prediction method based on multispectral remote sensing data and the XGBoost algorithm, so as to realize large-area, high-precision, field-scale wheat scab disease prediction and spatial mapping.
[0014] To achieve the above objectives, this invention provides a method for predicting wheat scab based on multispectral remote sensing data and the XGBoost algorithm, comprising the following steps:
[0015] Step S1: Obtain field wheat scab disease sample data in the target area and construct a disease sample database;
[0016] Step S2: Acquire multispectral remote sensing images covering the target area, preprocess the multispectral remote sensing images, calculate various multispectral vegetation indices representing the physiological state of vegetation based on the multispectral remote sensing images, and generate a multispectral vegetation index image.
[0017] Step S3: Based on the disease sample database, extract the vegetation index attribute values of each sample point from the corresponding multispectral vegetation index image to construct a feature database for model training and prediction.
[0018] Step S4: Use the differential optimization algorithm to optimize the hyperparameters of the XGBoost model to obtain the optimal XGBoost prediction model;
[0019] Step S5: Input the multispectral vegetation index image into the optimal XGBoost prediction model to perform pixel-by-pixel prediction of Fusarium head blight disease level and generate a spatial distribution map of Fusarium head blight disease level at the field level in the target area.
[0020] Furthermore, in step S1, obtaining field wheat scab disease sample data in the target area includes:
[0021] Step S11: Randomly set up field sampling points within the target area;
[0022] Step S12: Record the number of wheat ears infected with Fusarium head blight at each sampling point, and classify the disease severity into multiple levels based on the infection rate;
[0023] Step S13: Record the disease severity, longitude, and latitude of each sampling point to form a disease sample database.
[0024] Furthermore, in step S2, the multispectral remote sensing imagery includes Landsat 8 and Landsat 9 (Landsat 8 / 9) multispectral images; the preprocessing includes radiometric calibration, atmospheric correction, band stitching, administrative boundary cropping, and geographic projection conversion.
[0025] The multispectral vegetation indices include Normalized Difference Vegetation Index (NDVI), Normalized Difference Infrared Index (NDII), Normalized Difference Water Index (NDWI), Ratio Vegetation Index (RVI), Greenness Vegetation Index (GNDVI), and Fvédération Và Création (FVC).
[0026] Furthermore, the formula for the Normalized Difference Vegetation Index (NDVI) is as follows:
[0027]
[0028] Among them, band4 is the surface reflectance in the red band, and band5 is the surface reflectance in the near-infrared band.
[0029] The formula for the Normalized Difference Infrared Index (NDII) is expressed as follows:
[0030]
[0031] Among them, band5 is the surface reflectance of the near-infrared band, and band6 is the surface reflectance of the short-wave infrared band 1.
[0032] The formula for the Normalized Difference Water Index (NDWI) is as follows:
[0033]
[0034] Among them, band5 is the surface reflectance of the near-infrared band, and band7 is the surface reflectance of the short-wave infrared band 2.
[0035] The formula for the Ratio Vegetation Index (RVI) is as follows:
[0036]
[0037] Among them, band5 is the surface reflectance of the near-infrared band, and band6 is the surface reflectance of the short-wave infrared band 1.
[0038] The formula for the Greenness Vegetation Index (GNDVI) is as follows:
[0039]
[0040] Among them, band3 is the surface reflectance in the green light band, and band5 is the surface reflectance in the near-infrared band.
[0041] The formula for vegetation cover (FVC) is as follows:
[0042] ;
[0043] in, This represents the normalized vegetation index value for the current pixel. and These represent the entire NDVI raster layer. The maximum and minimum values.
[0044] Furthermore, the method for obtaining the target region is as follows:
[0045] Acquire multispectral images of early crop growth and land classification data for the same year covering the target area;
[0046] Calculate the vegetation index in the early stage of crop growth, reclassify the land classification data, and extract the crop attribute categories;
[0047] Set a vegetation index threshold, and extract areas whose vegetation index is greater than the threshold and belong to the crop category as crop planting area images.
[0048] Furthermore, in step S4, the step of optimizing the hyperparameters of the XGBoost model using a differential optimization algorithm includes:
[0049] Step S41: Define the set of XGBoost hyperparameters to be optimized, including the maximum depth of the tree. Learning rate eta, minimum loss reduction threshold gamma required for node splitting, and subsample ratio used when training each tree;
[0050] Step S42: Generate an initial population. Each individual in the population is a hyperparameter vector X, where the j-th parameter value of the individual is X. i,j upper and lower bounds of parameter search and Randomly generated from the given information, expressed by the formula:
[0051] ;
[0052] Where random(0,1) is a uniformly distributed random number between 0 and 1. and These represent the lower and upper bounds of the values of each hyperparameter, respectively.
[0053] Step S43: Using the root mean square error (RMSE) of 10-fold cross-validation as the fitness function, calculate the RMSE for each set of parameters, where the RMSE formula is expressed as:
[0054] ;
[0055] Where N represents the total number of samples, g is the sample order index, and true g Pred represents the measured disease level of the g-th sample point. g This indicates that the XGBoost model predicts the severity of the disease at the g-th sample point.
[0056] Step S44: Perform differential mutation operation on the current population to generate candidate parameter combinations V. candidate For each target vector X m Its mutation vector formula is expressed as:
[0057] ;
[0058] Among them, X r1 X r2 X r3 Let F be three distinct parent parameter vectors randomly selected from the current population, where r1, r2, and r3 are the corresponding individual indices and none of them are equal to the current individual index m. weight V is the variation weighting coefficient. candidate For candidate parameter combinations generated through mutation;
[0059] Step S45: Apply boundary constraints to the mutated candidate parameter combinations and determine whether to update the individual parameters using greedy selection. If the RMSE of the candidate parameter combination is less than the RMSE of the current individual parameter combination, the candidate parameter combination is retained for the next generation; otherwise, the original parameter combination is retained. The formula is as follows:
[0060] ;
[0061] in, This represents the parameter combination of the m-th individual in the t-th iteration. This represents the parameter state of the m-th individual in the (t+1)-th generation. RMSE() represents the calculation of the root mean square error. "Otherwise" means that the RMSE value predicted by the new parameter combination is greater than or equal to the RMSE value predicted by the original parameter combination.
[0062] Step S46: Repeat steps S44 and S45 until the preset iteration termination condition is met. Select the parameter combination that obtains the minimum RMSE during the iteration process as the optimal hyperparameter and construct the optimal XGBoost prediction model.
[0063] Furthermore, step S4 also includes
[0064] Step S47: This step evaluates the accuracy of the optimal XGBoost prediction model. The RMSE (Recognition Coefficient of Performance) R² is used to evaluate the model's prediction performance. The R² formula is as follows:
[0065] ;
[0066] Where, true mean R² represents the average of all measured wheat scab disease severity levels. R² ranges from 0 to 1, with values closer to 1 indicating higher model prediction accuracy. N represents the total number of samples, and g is the sample point index. g Pred represents the measured disease level of the g-th sample point. g This indicates that the XGBoost model predicts the severity of the disease at the g-th sample point.
[0067] Furthermore, step S4 also includes:
[0068] Step S48: Calculate the feature importance score by extracting the total gain of each feature in the XGBoost model when splitting across all decision tree nodes. The importance score of the k-th feature is... k The formula is expressed as:
[0069] ;
[0070] Where trees represent all decision trees in the XGBoost model, and Feature k Let Gain be the k-th independent variable, representing the value of the feature used by the model at the decision tree node. k The improvement in the accuracy of Fusarium head blight disease severity prediction brought about by the splitting process.
[0071] Furthermore, the step of inputting the multispectral vegetation index image into the optimal XGBoost prediction model for predicting the severity of Fusarium head blight includes, in step S5:
[0072] Step S51: Multiply the multispectral vegetation index image with the crop planting area image to synthesize a multi-band input image;
[0073] Step S52: Use the multi-band input image as the input feature of the optimal XGBoost prediction model, perform disease level prediction pixel by pixel for each field pixel in the wheat planting area, reclassify the Fusarium head blight disease level prediction results, and finally generate a spatial distribution map of Fusarium head blight disease level at the field level in the target area.
[0074] The present invention also provides a wheat scab prediction system based on Landsat multispectral data and the XGBoost algorithm, comprising:
[0075] The sample data acquisition module is used to acquire field wheat scab disease sample data in the target area and construct a disease sample database.
[0076] The remote sensing data acquisition and preprocessing module is used to acquire multispectral remote sensing images covering the target area, preprocess the multispectral remote sensing images, calculate various multispectral vegetation indices representing the physiological state of vegetation based on the images, and generate multispectral vegetation index images.
[0077] The feature extraction module is used to extract the vegetation index attribute values of each sample point from the corresponding multispectral vegetation index image based on the disease sample database, and to construct a feature database for model training and prediction.
[0078] The model optimization and prediction module is used to optimize the hyperparameters of the XGBoost model using a differential optimization algorithm to obtain the optimal XGBoost prediction model. The multispectral vegetation index image is then input into the optimal XGBoost prediction model to perform pixel-by-pixel prediction of Fusarium head blight disease level. The prediction results of Fusarium head blight disease level are then reclassified to generate a spatial distribution map of Fusarium head blight disease level at the field level in the target area.
[0079] This invention introduces the XGBoost algorithm into the field of wheat scab prediction. Utilizing the tree-structure regularization mechanism and nonlinear fitting capability of the XGBoost algorithm, it analyzes the complex nonlinear mapping relationship between multidimensional spectral vegetation index features and complex disease levels, improving prediction accuracy compared to traditional linear regression or simple threshold models. Furthermore, the DO algorithm is employed for global adaptive optimization of the XGBoost model's hyperparameters. This adaptive optimization of hyperparameters enhances the model's sensitivity in capturing extreme disease levels. The resulting field-level spatial distribution map of scab disease levels can directly serve grassroots agricultural extension workers and growers, promoting sustainable agricultural development. Attached Figure Description
[0080] Figure 1 Flowchart of the wheat scab prediction method based on multispectral remote sensing data and XGBoost algorithm of the present invention;
[0081] Figure 2 The wheat scab disease severity level and corresponding remote sensing image pixel value at the sampling points in this invention;
[0082] Figure 3 The present invention is used to calculate the NDVI in January in wheat distribution areas;
[0083] Figure 4 This invention is used to calculate the crop planting area of the current year in the wheat distribution area;
[0084] Figure 5 The wheat planting area in this invention;
[0085] Figure 6 The Landsat 8 / 9 multispectral vegetation index in the form of a raster image is used in this invention;
[0086] Figure 7 The parameter optimization process based on the DO algorithm in this invention;
[0087] Figure 8 The prediction performance of the DO-XGBoost model in this invention;
[0088] Figure 9 The ranking of the importance of independent variables in the DO-XGBoost model of this invention;
[0089] Figure 10 Spatial distribution map of wheat scab severity in Xinyang City and Zhumadian City in this invention. Detailed Implementation
[0090] Example 1
[0091] This embodiment uses Xinyang City and Zhumadian City in southern Henan Province as the target study area, and conducts wheat scab disease prediction based on Landsat 8 / 9 multispectral remote sensing data and the XGBoost algorithm. The specific steps are as follows.
[0092] S1: Obtain field wheat scab disease sample data in the target area and construct a disease sample database;
[0093] The method for determining the target region is as follows:
[0094] First, acquire early-stage multispectral images of crops and land classification data for the same year covering the target area.
[0095] Specifically, download the January Landsat-8 / 9 multispectral remote sensing imagery from the USGS official data website, and download the MODIS land classification data for that year from the NASA Earth Data website. In January, the main grain crop grown in Xinyang City and Zhumadian City is wheat. The multispectral reflectance information of green wheat is obvious, while the information of other crops and deciduous forests is not obvious. Therefore, the January imagery can be used to identify wheat-growing areas.
[0096] Preprocessing was performed on the January Landsat 8 and Landsat 9 (Landsat 8 / 9) multispectral remote sensing image data, including radiometric calibration and cosine atmospheric correction in R. Each image in the same band from January was stitched together and cropped using administrative boundaries in shapefile format for Xinyang City and Zhumadian City, resulting in cropped bands 4 and 5. All bands underwent projection transformation to the WGS84 geographic coordinate system with a spatial resolution of 0.0005 degrees.
[0097] The Normalized Difference Vegetation Index (NDVI) for January is calculated using the following formula:
[0098] ;
[0099] Among them, band4 is the fourth band of Landsat 8 / 9, that is, the surface reflectance in the red light band, and band5 is the fifth band, that is, the surface reflectance in the near-infrared band.
[0100] Then, the vegetation index in the early stage of crop growth is calculated, and the land classification data is reclassified to extract crop attribute categories.
[0101] Specifically, the MODIS land classification data is reassigned, with pixels having an attribute value of 12 set to 100, indicating a crop land type, and other attribute values set to 0, indicating non-crop land. The data is then projected to a geographic coordinate system with a spatial resolution of 0.0005 degrees.
[0102] Finally, a vegetation index threshold is set, and areas with a vegetation index greater than the threshold and belonging to the crop category are extracted as crop planting area images.
[0103] Specifically, based on the NDVI data for January and the reassigned MODIS land classification data, the wheat planting areas for Xinyang City and Zhumadian City in that year were calculated.
[0104] By comparing the calculated wheat planting area with the wheat planting area officially reported by Xinyang City and Zhumadian City, the optimal NDVI threshold was determined to be 0.292. Wheat is in its early growth stage in January, and the winter vegetation index is generally lower than the May vegetation index. Areas with an NDVI greater than 0.292 in January and a reassigned land classification attribute value of 100 were designated as wheat planting areas. The calculated wheat planting area in this example is 10924.25 square kilometers, which is close to the 10930.73 square kilometers inferred from official information, indicating that the NDVI threshold is reasonable. The wheat planting area attribute value was assigned as 1, and non-wheat planting areas were assigned the null value NA, resulting in a crop planting area image. In this embodiment, the crop planting area image is the wheat planting area image.
[0105] Step S11: Randomly set up field sampling points within the target area.
[0106] In this embodiment, sampling points were randomly arranged in the main wheat-producing areas of Xinyang City and Zhumadian City. The potential severely affected areas of wheat scab are mainly the areas along the Huai River in Xinyang and Zhumadian City.
[0107] Step S12: Record the number of wheat ears infected with Fusarium head blight at each sampling point, and classify the disease into multiple levels based on the infection rate.
[0108] Specifically, May is the critical period for the occurrence of wheat scab. In May, the number of diseased ears of wheat with scab was checked and recorded at each sampling point. To highlight the damage caused by wheat scab and facilitate timely control, based on field surveys and the national standard "Technical Specification for Monitoring and Forecasting of Wheat Scab" (GB / T15796-2011), the severity of wheat scab was finely divided into six levels: Level 0, Level 1, Level 2, Level 3, Level 4, and Level 5. The basis for calculating the severity level is the diseased ear rate, calculated as: Number of diseased ears / Total number of ears at the sampling point × 100, in percentage (%). Level 0: No diseased ears. Level 1: Diseased ear rate less than 10%. Level 2: Diseased ear rate 10%–20%. Level 3: Diseased ear rate 20%–30%. Level 4: Diseased ear rate 30%–40%. Level 5: Diseased ear rate 40%–100%.
[0109] Step S13: Record the disease severity, longitude, and latitude of each sampling point to form disease sample data.
[0110] Preliminary processing was performed on the disease severity, longitude, and latitude data of all sampling points. This preliminary processing included removing outliers, missing values, and sampling points less than 50 meters from houses or roads. The processed sampling point data was stored in CSV format to establish a wheat scab sampling point database.
[0111] Step S2: Acquire multispectral remote sensing images (lacking period descriptions) covering the target area, preprocess the multispectral remote sensing images, calculate various multispectral vegetation indices representing the physiological state of vegetation based on the images, and generate a multispectral vegetation index image.
[0112] Specifically, download the May Landsat-8 / 9 multispectral remote sensing imagery from the USGS official data website. Preprocess the May imagery, including radiometric calibration, cosine atmospheric correction, and band-by-band mosaicking. Crop the imagery using the administrative boundaries of Xinyang City and Zhumadian City in shapefile format to obtain cropped bands 1, 2, 3, 4, 5, 6, and 7. Project the imagery into the WGS84 geographic coordinate system with a spatial resolution of 0.0005 degrees.
[0113] Based on the above wavelengths, six vegetation indices were calculated: NDVI, Normalized Difference Infrared Index (NDII), Normalized Difference Water Index (NDWI), Ratio Vegetation Index (RVI), Greenness Vegetation Index (GNDVI), and FVC (Further Vegetation Coverage). The formulas for each vegetation index are as follows:
[0114] The formula for the Normalized Difference Vegetation Index (NDVI) is as follows:
[0115] ;
[0116] Among them, band4 is the surface reflectance in the red band, and band5 is the surface reflectance in the near-infrared band.
[0117] The formula for the Normalized Difference Infrared Index (NDII) is expressed as follows:
[0118] ;
[0119] Among them, band5 is the surface reflectance in the near-infrared band, and band6 is the surface reflectance in the shortwave infrared band 1.
[0120] The formula for the Normalized Difference Water Index (NDWI) is as follows:
[0121] ;
[0122] Among them, band5 is the surface reflectance in the near-infrared band, and band7 is the surface reflectance in the short-wave infrared band 2.
[0123] The formula for the Ratio Vegetation Index (RVI) is as follows:
[0124] ;
[0125] Among them, band5 is the surface reflectance in the near-infrared band, and band6 is the surface reflectance in the shortwave infrared band 1.
[0126] The formula for the Greenness Vegetation Index (GNDVI) is as follows:
[0127] ;
[0128] Among them, band3 is the surface reflectance in the green light band, and band5 is the surface reflectance in the near-infrared band.
[0129] The formula for vegetation cover (FVC) is as follows:
[0130] ;
[0131] in, This represents the normalized vegetation index value for the current pixel. and These represent all pixels in the entire NDVI raster layer. The maximum and minimum values.
[0132] Step S3: Based on the disease sample database, extract the vegetation index attribute values of each sample point from the corresponding multispectral vegetation index image to construct a feature database for model training and prediction.
[0133] A unified geospatial reference coordinate system is established. In this embodiment, the WGS84 geospatial coordinate system is used to ensure spatial location matching of multi-source data. The wheat scab sample point database constructed in step S1 is read, and the spatial location information of each sample point, namely longitude and latitude, is extracted as a coordinate set, along with wheat scab disease severity data.
[0134] For each sample point coordinate, pixel attribute values for the corresponding spatial location were extracted from the NDVI, NDII, NDWI, RVI, GNDVI, and FVC raster layers. A comprehensive feature database was constructed, with each row representing an independent sample point. The columns contained, in order, longitude, latitude, wheat scab disease severity level, and the corresponding six multispectral remote sensing vegetation indices. The wheat scab disease severity level was the dependent variable of the model, and the six vegetation indices were the independent variables. This feature database was exported and stored as a CSV file.
[0135] Step S4: Use the differential optimization algorithm to optimize the hyperparameters of the XGBoost model to obtain the optimal XGBoost prediction model.
[0136] Step S41: In the R programming environment, read the sample attribute data in CSV format as described above. Define the set of XGBoost hyperparameters to be optimized, including the maximum tree depth (max).depth The learning rate eta, the minimum loss reduction threshold gamma required for node splitting, and the proportion of subsamples used when training each tree.
[0137] Step S42: Randomly generate an initial population of 10 within the defined boundaries. Within the upper and lower bounds of the parameters [U... min U max An initial population is randomly generated from [a set of elements], where the j-th parameter value of the i-th individual is X. i,j The formula is expressed as:
[0138] ;
[0139] Where X represents a parameter vector containing max depth X is a set of four hyperparameters: eta, gamma, subsample, and random(0,1) is a uniformly distributed random number between 0 and 1. i,j For the j-th parameter value of the i-th individual, U min U max Indicates the search range for the parameters.
[0140] Step S43: Calculate the root mean square error (RMSE) for each parameter group using 10-fold cross-validation. This method involves randomly dividing all samples into 10 equal subsets. In each iteration, 9 subsets are used for model training, and the remaining subset is used for independent validation, calculating the validation set RMSE. After 10 iterations, the average of the 10 validation set RMSEs is used to measure the generalization error of the current parameter combination. Finally, the parameter combination with the smallest average validation set RMSE is selected as the optimal parameters for the model. The RMSE formula is expressed as:
[0141] ;
[0142] Where N represents the total number of samples, g is the sequential index of the sample, and true g and pred g and represent the measured wheat scab disease severity level and the disease severity level predicted by the XGBoost model for the g-th sample point, respectively.
[0143] Step S44: Perform differential mutation operation on the current population to generate candidate parameter combinations V. candidat .
[0144] In this embodiment, three individuals whose indices are not equal to the current individual's index m are randomly selected from the population, and candidate offspring parameter combinations V are generated using a difference vector. candidate The formula is expressed as:
[0145] ;
[0146] Among them, F weight V represents the variation weights, used to control the strength of the influence of the difference information on the reference vector. candidat F represents the candidate parameter combinations generated through mutation. weight The larger the value of F, the larger the search step size. weight The smaller the value, the more refined the search. This example uses a median value of 0.5. r1 X r2 X r3 These are three randomly selected parent parameter vectors, where r1, r2, and r3 are the indices of three individuals randomly selected from the current population. (X) r2 -X r3 ) is a difference vector, representing the distribution direction of the current population in the search space. A large difference indicates a large variation, which is beneficial for global search.
[0147] Step S45: Apply boundary constraints to the mutated parameters and determine whether to update the individual based on the RMSE value. Specifically, if the RMSE of the candidate parameter combination is less than the RMSE of the current individual parameter combination, the candidate parameter combination is retained for the next generation; otherwise, the original parameter combination is retained. The selection rule formula is expressed as follows:
[0148] ;
[0149] in, This represents the parameter combination of the m-th individual in the current iteration of the t-th generation, i.e., a set of XGBoost hyperparameters; This represents the parameter state of the m-th individual in the next iteration of generation t+1. RMSE() represents the calculation of the root mean square error. "Otherwise" means that the RMSE value predicted by the new parameter combination is greater than or equal to the RMSE value predicted by the original parameter combination.
[0150] Step S46: Repeat steps S44-S45 a total of 10 times. In each round, poor parameters are continuously replaced with better ones. The final selected parameters are the optimal parameters obtained during the 10 rounds of optimization. The optimal parameters obtained in this embodiment are: max depth =5, eta=0.0839, gamma=0, subsample=0.50.
[0151] Step S47: Evaluate the XGBoost prediction model optimized by the DO algorithm based on RMSE and the coefficient of determination R².
[0152] In this embodiment, the RMSE of the 10-fold cross-validation was 0.53, and the R² was 0.91, indicating that the model has high goodness of fit and generalization ability, and strong interpretability of the data. The R² formula is expressed as:
[0153] ;
[0154] Where, true mean R² represents the average of all measured wheat scab disease severity levels. R² ranges from 0 to 1, with values closer to 1 indicating higher model prediction accuracy. N represents the total number of samples, and g is the sample point index. g Pred represents the measured disease level of the g-th sample point. g This indicates that the XGBoost model predicts the severity of the disease at the g-th sample point.
[0155] Step S48: Use the varImp function in R to extract the relative importance of the six independent variables. The importance score for the k-th feature is... k The formula is expressed as:
[0156] ;
[0157] Among them, importance k The importance score for the k-th feature, where trees represent the decision trees in the XGBoost model. k It is the k-th independent variable, where k is the index. Gain is the gain, i.e., the gain when the model uses a Feature. k After the decision was made, how much did the model improve its accuracy in predicting the severity of Fusarium head blight?
[0158] The optimal parameters are input into the XGBoost prediction model to construct the optimal DO-XGBoost prediction model.
[0159] Step S5: Input the multispectral vegetation index image into the optimal XGBoost prediction model to perform pixel-by-pixel prediction of Fusarium head blight disease level, reclassify the prediction results of Fusarium head blight disease level, and generate a spatial distribution map of Fusarium head blight disease level at the field level in the target area.
[0160] Step S51: Multiply the multispectral vegetation index image with the crop planting area image to synthesize a multi-band input image.
[0161] Specifically, six TIFF-formatted multispectral vegetation index raster images (NDVI, NDII, NDWI, RVI, GNDVI, and FVC) were sequentially read. The six vegetation index raster images were then multiplied separately from the wheat-grown area image, assigning null values to pixels in non-wheat-grown areas and retaining only the vegetation index information for wheat-grown areas. The six multiplied single-band images were then combined into a single multi-band image.
[0162] Step S52: Use the multi-band image as the input feature of the optimal DO-XGBoost model to predict the severity of Fusarium head blight on a pixel-by-pixel basis for wheat planting areas in the entire Xinyang City and Zhumadian City.
[0163] Based on the disease severity classification criteria for wheat Fusarium head blight in this example, the disease severity prediction results were reclassified. Pixels with a predicted value of 0 were classified as level 0; predicted values in the interval (0, 1) were classified as level 1; those in the interval (1, 2) as level 2; those in the interval (2, 3) as level 3; those in the interval (3, 4) as level 4; and pixels with predicted values greater than 4 were classified as level 5. This resulted in a spatial distribution map of Fusarium head blight severity levels with high spatial resolution at the field level, providing reliable technical support for large-scale, high-precision, and quantitative spatial prediction of wheat Fusarium head blight.
[0164] Example 2:
[0165] This embodiment provides a wheat scab prediction system based on Landsat multispectral data and the XGBoost algorithm, including:
[0166] The sample data acquisition module is used to acquire field wheat scab disease sample data in the target area and construct a disease sample database.
[0167] The remote sensing data acquisition and preprocessing module is used to acquire multispectral remote sensing images covering the target area, preprocess the multispectral remote sensing images, calculate various multispectral vegetation indices representing the physiological state of vegetation based on the images, and generate a multispectral vegetation index image.
[0168] The feature extraction module is used to extract the vegetation index attribute values of each sample point from the corresponding multispectral vegetation index image based on the disease sample database, and to construct a feature database for model training and prediction.
[0169] The model optimization and prediction module is used to optimize the hyperparameters of the XGBoost model using a differential optimization algorithm to obtain the optimal XGBoost prediction model. The multispectral vegetation index image is then input into the optimal XGBoost prediction model to perform pixel-by-pixel prediction of Fusarium head blight disease level. The prediction results of Fusarium head blight disease level are then reclassified to generate a spatial distribution map of Fusarium head blight disease level at the field level in the target area.
Claims
1. A method for predicting wheat scab based on multispectral remote sensing data and the XGBoost algorithm, characterized in that, Includes the following steps: Step S1: Obtain field wheat scab disease sample data in the target area and construct a disease sample database; Step S2: Acquire multispectral remote sensing images covering the target area, preprocess the multispectral remote sensing images, calculate various multispectral vegetation indices representing the physiological state of vegetation based on the multispectral remote sensing images, and generate a multispectral vegetation index image. Step S3: Based on the disease sample database, extract the vegetation index attribute values of each sample point from the corresponding multispectral vegetation index image to construct a feature database for model training and prediction. Step S4: Use the differential optimization algorithm to optimize the hyperparameters of the XGBoost model to obtain the optimal XGBoost prediction model; Step S5: Input the multispectral vegetation index image into the optimal XGBoost prediction model to perform pixel-by-pixel prediction of Fusarium head blight disease level and generate a spatial distribution map of Fusarium head blight disease level at the field level in the target area.
2. The wheat scab prediction method based on multispectral remote sensing data and XGBoost algorithm according to claim 1, characterized in that, In step S1, obtaining field wheat scab disease sample data in the target area includes: Step S11: Randomly set up field sampling points within the target area; Step S12: Record the number of wheat ears infected with Fusarium head blight at each sampling point, and classify the disease severity into multiple levels based on the infection rate; Step S13: Record the disease severity, longitude, and latitude of each sampling point to form a disease sample database.
3. The wheat scab prediction method based on multispectral remote sensing data and XGBoost algorithm according to claim 1, characterized in that, In step S2, the multispectral remote sensing imagery includes Landsat 8 and Landsat 9 (Landsat 8 / 9) multispectral images; the preprocessing includes radiometric calibration, atmospheric correction, band stitching, administrative boundary cropping, and geographic projection conversion. The multispectral vegetation indices include Normalized Difference Vegetation Index (NDVI), Normalized Difference Infrared Index (NDII), Normalized Difference Water Index (NDWI), Ratio Vegetation Index (RVI), Greenness Vegetation Index (GNDVI), and Fvédération Và Création (FVC).
4. The wheat scab prediction method based on multispectral remote sensing data and XGBoost algorithm according to claim 3, characterized in that, The formula for the Normalized Difference Vegetation Index (NDVI) is as follows: ; Among them, band4 is the surface reflectance in the red band, and band5 is the surface reflectance in the near-infrared band. The formula for the Normalized Difference Infrared Index (NDII) is expressed as follows: ; Among them, band5 is the surface reflectance of the near-infrared band, and band6 is the surface reflectance of the short-wave infrared band 1. The formula for the Normalized Difference Water Index (NDWI) is as follows: ; Among them, band5 is the surface reflectance of the near-infrared band, and band7 is the surface reflectance of the short-wave infrared band 2. The formula for the Ratio Vegetation Index (RVI) is as follows: ; Among them, band5 is the surface reflectance of the near-infrared band, and band6 is the surface reflectance of the short-wave infrared band 1. The formula for the Greenness Vegetation Index (GNDVI) is as follows: ; Among them, band3 is the surface reflectance in the green light band, and band5 is the surface reflectance in the near-infrared band. The formula for vegetation cover (FVC) is as follows: ; in, This represents the normalized vegetation index value for the current pixel. and These represent the entire NDVI raster layer. The maximum and minimum values.
5. The wheat scab prediction method based on multispectral remote sensing data and XGBoost algorithm according to claim 1, characterized in that, The method for obtaining the target region is as follows: Acquire multispectral images of early crop growth and land classification data for the same year covering the target area; Calculate the vegetation index in the early stage of crop growth, reclassify the land classification data, and extract the crop attribute categories; Set a vegetation index threshold, and extract areas with a vegetation index greater than the threshold that belong to the crop category as crop planting area images.
6. The wheat scab prediction method based on multispectral remote sensing data and XGBoost algorithm according to claim 1, characterized in that, In step S4, the step of optimizing the hyperparameters of the XGBoost model using a differential optimization algorithm includes: Step S41: Define the set of XGBoost hyperparameters to be optimized, including the maximum depth of the tree. Learning rate eta, minimum loss reduction threshold gamma required for node splitting, and subsample ratio used when training each tree; Step S42: Generate an initial population. Each individual in the population is a hyperparameter vector X, where the j-th parameter value of the individual is X. i,j upper and lower bounds of parameter search and Randomly generated from the given information, expressed by the formula: ; Where random(0,1) is a uniformly distributed random number between 0 and 1. and These represent the lower and upper bounds of the values of each hyperparameter, respectively. Step S43: Using the root mean square error (RMSE) of 10-fold cross-validation as the fitness function, calculate the RMSE for each group of parameters, where the RMSE formula is expressed as: ; Where N represents the total number of samples, g is the sample order index, and true g Pred represents the measured disease level of the g-th sample point. g This indicates that the XGBoost model predicts the severity of the disease at the g-th sample point. Step S44: Perform differential mutation operation on the current population to generate candidate parameter combinations V. candidate For each target vector X m Its mutation vector formula is expressed as: ; Among them, X r1 X r2 X r3 Let F be three distinct parent parameter vectors randomly selected from the current population, where r1, r2, and r3 are the corresponding individual indices and none of them are equal to the current individual index m. weight V is the variation weighting coefficient. candidate For candidate parameter combinations generated through mutation; Step S45: Apply boundary constraints to the mutated candidate parameter combinations and determine whether to update the individual parameters using greedy selection. If the RMSE of the candidate parameter combination is less than the RMSE of the current individual parameter combination, the candidate parameter combination is retained for the next generation; otherwise, the original parameter combination is retained. The formula is as follows: ; in, This represents the parameter combination of the m-th individual in the t-th iteration. This represents the parameter state of the m-th individual in the (t+1)-th generation. RMSE() represents the calculation of the root mean square error. "Otherwise" means that the RMSE value predicted by the new parameter combination is greater than or equal to the RMSE value predicted by the original parameter combination. Step S46: Repeat steps S44 and S45 until the preset iteration termination condition is met. Select the parameter combination that obtains the minimum RMSE during the iteration process as the optimal hyperparameter and construct the optimal XGBoost prediction model.
7. The wheat scab prediction method based on multispectral remote sensing data and XGBoost algorithm according to claim 6, characterized in that, Step S4 also includes Step S47: This step evaluates the accuracy of the optimal XGBoost prediction model. The RMSE (Recognition Coefficient of Performance) R² is used to evaluate the model's prediction performance. The R² formula is as follows: ; Where, true mean R² represents the average of all measured wheat scab disease severity levels. R² ranges from 0 to 1, with values closer to 1 indicating higher model prediction accuracy. N represents the total number of samples, and g is the sample point index. g Pred represents the measured disease level of the g-th sample point. g This indicates that the XGBoost model predicts the severity of the disease at the g-th sample point.
8. The wheat scab prediction method based on multispectral remote sensing data and XGBoost algorithm according to claim 6, characterized in that, Step S4 further includes: Step S48: Calculate the feature importance score by extracting the total gain of each feature in the XGBoost model when splitting across all decision tree nodes. The importance score of the k-th feature is... k The formula is expressed as: ; Where trees represent all decision trees in the XGBoost model, and Feature k Let Gain be the k-th independent variable, representing the value of the feature used by the model at the decision tree node. k The improvement in the accuracy of Fusarium head blight disease severity prediction brought about by the splitting process.
9. A method for predicting wheat scab based on multispectral remote sensing data and the XGBoost algorithm according to claim 5 or 6, characterized in that, The step of inputting the multispectral vegetation index image into the optimal XGBoost prediction model to predict the severity level of Fusarium head blight includes, in step S5: Step S51: Multiply the multispectral vegetation index image with the crop planting area image to synthesize a multi-band input image; Step S52: Use the multi-band input image as the input feature of the optimal XGBoost prediction model, perform disease level prediction pixel by pixel for each field pixel in the wheat planting area, reclassify the Fusarium head blight disease level prediction results, and finally generate a spatial distribution map of Fusarium head blight disease level at the field level in the target area.
10. A wheat scab prediction system based on Landsat multispectral data and the XGBoost algorithm, characterized in that, include: The sample data acquisition module is used to acquire field wheat scab disease sample data in the target area and construct a disease sample database. The remote sensing data acquisition and preprocessing module is used to acquire multispectral remote sensing images covering the target area, preprocess the multispectral remote sensing images, calculate various multispectral vegetation indices representing the physiological state of vegetation based on the images, and generate multispectral vegetation index images. The feature extraction module is used to extract the vegetation index attribute values of each sample point from the corresponding multispectral vegetation index image based on the disease sample database, and to construct a feature database for model training and prediction. The model optimization and prediction module is used to optimize the hyperparameters of the XGBoost model using a differential optimization algorithm to obtain the optimal XGBoost prediction model. The multispectral vegetation index image is then input into the optimal XGBoost prediction model to perform pixel-by-pixel prediction of Fusarium head blight disease level. The prediction results of Fusarium head blight disease level are then reclassified to generate a spatial distribution map of Fusarium head blight disease level at the field level in the target area.