Wheat scab identification method based on CARS-Ridge algorithm and new index
Through the integration of CARS-Ridge algorithm and the new index WFItwo and WFIthree, the problem of inefficiency of wheat gibberellia identification model is solved, high-precision disease identification and pesticide spraying guidance are achieved, and wheat yield and quality are improved.
Patent Information
- Application Number
- CN202310631473.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-31
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2043-05-31
AI Technical Summary
The existing wheat gibberellia identification model has low operating efficiency and low accuracy, making it difficult to accurately capture weak signals in disease spectral information. The lack of specificity of traditional vegetation indexes leads to inaccurate identification.
The CARS-Ridge algorithm is used to combine the new index WFItwo and WFIthree to build the optimal wheat gibberellia identification model through data dimensionality reduction and modeling, and integrate the new vegetation index and optimal model to improve the recognition accuracy.
It greatly improves the accuracy and stability of wheat gibberellosis identification, can accurately judge the distribution and incidence of field disease infection, guide pesticide spraying, reduce waste and environmental pollution, and improve wheat quality and yield.
Smart Images

Figure CN117132853B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of crop disease remote sensing monitoring, in particular to a wheat scab identification method based on a CARS-Ridge algorithm fused with a novel index. Background Art
[0002] Wheat, one of China's most important grain crops, has long been plagued by the impact of fusarium head blight on its yield and quality. This disease directly disrupts wheat grain development, causing grain shrinkage and weight loss, resulting in yield losses of 10-70% in endemic years. In recent years, fusarium head blight has become increasingly severe with global warming. Therefore, accurately identifying the development of fusarium head blight is crucial for successful disease control.
[0003] Currently, spectral reflectance remains a promising candidate for wheat fusarium head blight identification. Although hyperspectral data contains many spectral bands, not all bands are sensitive to the monitoring target. With limited training sample data, the accuracy of disease identification tends to improve initially and then decrease as the number of bands analyzed increases, and the model's operating efficiency also gradually decreases. Furthermore, fusarium head blight primarily affects the wheat ear. Compared to the large-scale identification of wheat leaves for rust and powdery mildew, the signal from the wheat ear that can be acquired by the sensor is weaker, especially in the early stages of the disease. Therefore, while traditional vegetation indices derived from mathematical transformations of spectral reflectance are a simple and effective tool for detecting crop diseases, their lack of disease specificity often results in low model accuracy. Currently, it is impossible to quantitatively describe or identify wheat fusarium head blight.
[0004] In order to solve the defects of low efficiency and low accuracy of the wheat fusarium scabra identification model, the most important part of this invention is how to use advanced algorithms to effectively mine the most sensitive bands from a large amount of spectral information and develop new indexes to accurately capture weak signals in the spectral information of wheat diseases and improve the accuracy and stability of the disease identification model. Summary of the Invention
[0005] In order to solve the defects of low operating efficiency and low precision of wheat fusarium scabra identification model, the purpose of the present invention is to provide a wheat fusarium scabra identification method based on the CARS-Ridge algorithm fused with a new index, which greatly improves the precision of existing disease inversion and improves the accuracy of wheat fusarium scabra identification.
[0006] To achieve the above object, the present invention adopts the following technical solution: a method for identifying wheat scab based on the CARS-Ridge algorithm and a new index, the method comprising the following steps in order:
[0007] (1) Data acquisition: Obtain hyperspectral data of wheat fusarium canopy and calculate the incidence rate of each sample plot;
[0008] (2) Data dimensionality reduction: CARS, PCA and SPA algorithms were used to reduce the dimensionality of the wheat fusarium canopy hyperspectral data, and three different feature sets were obtained;
[0009] (3) Modeling: Using RF, PLSR, and Ridge algorithms, we obtained nine wheat fusarium head blight recognition models, namely, CARS-RF, CARS-PLSR, CARS-Ridge, PCA-RF, PCA-PLSR, PCA-Ridge, SPA-RF, SPA-PLSR, and SPA-Ridge.
[0010] (4) Determine the optimal model: The optimal model was determined by performing a ten-fold cross-validation on the results of the nine wheat scab recognition models;
[0011] (5) Development of new indices: Based on 12 vegetation indices, two new indices, namely WFI, were constructed. two and WFI three ;
[0012] (6) Constructing the optimal wheat fusarium head blight identification model: integrating the new index with the optimal model to obtain the optimal wheat fusarium head blight identification model.
[0013] The step (1) specifically includes the following steps:
[0014] (1a) Spectral data measurement: Hyperspectral data of wheat scab canopy were obtained. Disease occurrence was divided into two types: natural occurrence and artificial inoculation. All wheat canopy spectral reflectances were measured using a FieldSpec4 spectrometer.
[0015] (1b) Statistical morbidity: The morbidity of each plot is calculated to quantitatively estimate the severity of wheat infection. The wheat morbidity is carried out simultaneously with the canopy spectrum collection. The DER, the ratio of the number of diseased wheat ears in the plot to the total number of surveyed ears, is used as an important evaluation indicator of the severity of wheat infection.
[0016] In step (2), the dimensionality reduction of hyperspectral data using the CARS algorithm specifically includes the following steps:
[0017] (2a) Performing Monte Carlo model sampling: Sampling variables through Monte Carlo MC in an iterative and competitive manner. Setting the Monte Carlo MC sampling to run N times and obtain N subset variables. In each sampling run, randomly select 80% of the samples in the subset variables to establish the Monte Carlo model.
[0018] (2b) Wavelength selection by exponential decreasing function EDF: The exponential decreasing function EDF is used to force the wavelength to decrease and remove the variables with small absolute values of regression coefficients;
[0019] (2c) Adaptive reweighted sampling is performed: variables with larger absolute values of regression coefficients are retained to form a new feature set;
[0020] (2d) Iterative loop: The new feature set is input into the Monte Carlo model and the root mean square error of the ten-fold cross validation is calculated. The subset corresponding to the minimum root mean square error recorded in N runs is the optimal wavelength combination screened out. The optimal wavelength combination constitutes the feature set determined by the CARS algorithm;
[0021] The exponentially decreasing function EDF is defined as:
[0022] r i =ae -ki
[0023] Where i represents the number of sampling runs; in the Nth sampling run, only two wavelengths are retained, r1 = 1, r N =2 / P; a and k are constants, a is the value of the exponentially decreasing function EDF at i=0, k is the rate of decrease of the exponentially decreasing function EDF, and the calculation formulas of a and k are as follows:
[0024]
[0025]
[0026] The number of wavelengths P is 501, and the total number N of sampling runs is set to 50.
[0027] In step (3), modeling by the Ridge algorithm specifically refers to:
[0028] Assuming the linear regression model y=xβ+ε, the objective function of the regression parameter β is:
[0029]
[0030] in, x is the independent variable, i.e., the canopy spectral reflectance of the sample plot; y is the dependent variable, i.e., the incidence rate; β and ε are the regression coefficient and error, respectively; n is the number of samples; λ is the regularization parameter, whose value is determined by the ridge trace method to balance the variance and bias of the linear regression model; the bias of the linear regression model increases with the increase of λ, while the variance increases in the opposite direction;
[0031] The objective function of the regression coefficient β is equivalent to:
[0032]
[0033] Where x i and y i Represents the i-th independent variable and dependent variable respectively, n represents the number of samples, α represents the regression coefficient vector, which is used to fit the unknown parameters of the linear relationship between the independent variable and the dependent variable; according to the β-λ curve in the Cartesian coordinate system, that is, the ridge curve, at the point where β tends to be stable, find the optimal λ value.
[0034] The step (4) specifically includes the following steps:
[0035] (4a) Three different feature sets obtained from the CARS, PCA, and SPA algorithms were used as input variables and combined with the RF, PLSR, and Ridge algorithms to construct a total of nine wheat fusarium head blight recognition models: CARS-RF, CARS-PLSR, CARS-Ridge, PCA-RF, PCA-PLSR, PCA-Ridge, SPA-RF, SPA-PLSR, and SPA-Ridge. The three different input feature sets were evenly divided into ten parts, nine for modeling and one for validation, and the results were repeated ten times. The ten-fold cross-validation method was used to evaluate the recognition ability of each model for wheat fusarium head blight.
[0036] (4b) The accuracy and performance of each model are determined by the coefficient of determination R 2 and root mean square error RMSE are used to evaluate the two statistical indicators, R 2 The calculation formulas for RMSE are:
[0037]
[0038]
[0039] Where n is the number of samples, t i is the actual value, d i is the model prediction value; the smaller the RMSE value, the more accurate the model; R 2 The value of is between 0 and 1, the closer to 1 the better, indicating the degree to which the model explains the changes in the dependent variable; when R 2 When R is 1, the model fits the data perfectly; when R 2 When it is 0, the model cannot explain the changes in the dependent variable;
[0040] (4c) In all models, R 2 The closer the model is to 1 and the closer the RMSE is to 0, the better its ability to identify diseases is. According to the model operation results, the CARS-Ridge model is determined to be the optimal model.
[0041] The step (5) specifically includes the following steps:
[0042] (5a) First, the correlation between the full spectrum band and the ratio of diseased wheat ears in the sample to the total number of wheat ears in the survey (DER) was calculated according to the calculation formula of the Pearson correlation coefficient. The correlation result was between 0 and 1. A fixed threshold was selected to divide the bands. The wavelengths with a correlation coefficient greater than 0.7 were selected as the primary wavelengths. Assuming (x1, y1), (x2, y2),…, (x n ,y n ) represent the values of two variables X and Y, n is the number of samples, and the calculation formula of Pearson correlation coefficient is:
[0043]
[0044] Where R is the Pearson correlation coefficient, x i and y i denote the i-th independent variable and dependent variable respectively;
[0045] is the covariance, which measures the correlation between variables X and Y, and are the mean values of the samples, namely:
[0046]
[0047]
[0048] (5b) The correlation between the vegetation index and the proportion of diseased wheat ears in the sample to the total number of wheat ears in the survey (DER) was calculated according to the calculation formula of the Pearson correlation coefficient. Similarly, the vegetation index with a correlation coefficient with DER of more than 0.7 was selected as the primary vegetation index according to the threshold division; the 12 vegetation indices were sensitive pigment index SIPI, anthocyanin reflectance index ARI, transformed chlorophyll absorption reflectance index TCARI, modified chlorophyll absorption ratio index MCARI, normalized difference vegetation index NDVI, modified simple ratio index MSR, ratio vegetation structure index RVSI, plant senescence reflectance index PSRI, triangular vegetation index TVI, nitrogen reflectance index NRI, photochemical reflectance index PRI and physiological reflectance index PHRI;
[0049] (5c) The correlation between the primary wavelength and the primary vegetation index was calculated according to the calculation formula of the Pearson correlation coefficient. The threshold division was used to further divide the primary wavelengths with a correlation coefficient greater than 0.7 with the primary vegetation index. The wavelengths that were finally retained were used to construct two new indices WFI. two and WFI three :
[0050]
[0051]
[0052] Among them, R 687 、R 659 and R 760 is the wavelength that is ultimately retained.
[0053] The step (6) specifically refers to: using the new index and the feature set determined based on the CARS algorithm as inputs of the optimal model to construct an optimal wheat fusarium head blight recognition model.
[0054] From the above technical solution, it can be seen that the beneficial effects of the present invention are as follows: First, the present invention evaluates the application potential of hyperspectral data in wheat scab identification by reducing data dimension and combining with new index construction, and proposes CARS-Ridge algorithm and WFI two and WFI three The development of the new index has determined the most accurate wheat fusarium head blight identification model, that is, the optimal wheat fusarium head blight identification model; second, the present invention greatly improves the accuracy of existing disease inversion and overcomes the defect of inaccurate wheat fusarium head blight identification. The present invention can be used as the basis for regional wheat disease remote sensing monitoring to help accurately judge the infection distribution and occurrence degree of wheat diseases in the field. In practical applications, it can help accurately guide field pesticide spraying, avoid pesticide spraying waste and environmental pollution, and is of great significance to improving wheat quality and yield; third, the present invention can provide method support for the detection of wheat fusarium head blight, and can also provide valuable reference for the diagnosis of diseases of other crops in the field. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 is a flow chart of the method of the present invention;
[0056] Figure 2 It is a scatter plot of the measured data and predicted data of the 9 models in the present invention;
[0057] Figure 3 Diagram of the methodology developed for the novel index in this invention;
[0058] Figure 4 This is the fitting result diagram after the CARS-Ridge model and the new exponential fusion. DETAILED DESCRIPTION
[0059] like Figure 1 As shown, a wheat scab identification method based on the CARS-Ridge algorithm and the new index is provided, which includes the following steps in sequence:
[0060] (1) Data acquisition: Obtain hyperspectral data of wheat fusarium canopy and calculate the incidence rate of each sample plot;
[0061] (2) Data dimensionality reduction: CARS, PCA and SPA algorithms were used to reduce the dimensionality of the wheat fusarium canopy hyperspectral data, and three different feature sets were obtained;
[0062] (3) Modeling: Using RF, PLSR, and Ridge algorithms, we obtained nine wheat fusarium head blight recognition models, namely, CARS-RF, CARS-PLSR, CARS-Ridge, PCA-RF, PCA-PLSR, PCA-Ridge, SPA-RF, SPA-PLSR, and SPA-Ridge.
[0063] (4) Determine the optimal model: The optimal model was determined by performing a ten-fold cross-validation on the results of the nine wheat scab recognition models;
[0064] (5) Development of new indices: Based on 12 vegetation indices, two new indices, namely WFI, were constructed. two and WFI three ;
[0065] (6) Constructing the optimal wheat fusarium head blight identification model: integrating the new index with the optimal model to obtain the optimal wheat fusarium head blight identification model.
[0066] The step (1) specifically includes the following steps:
[0067] (1a) Measured spectral data: Experiments were conducted in 2018, 2019, and 2022 to obtain valid hyperspectral data of wheat fusarium head blight canopy. The experimental sites were mainly located in Baihu and Guohe Towns in Anhui Province and the Xiaotangshan Experimental Base in Beijing. Disease occurrence was divided into natural occurrence and artificial inoculation. All wheat canopy spectral reflectance was measured using a FieldSpec4 spectrometer;
[0068] (1b) Statistical morbidity: The morbidity of each plot is calculated to quantitatively estimate the severity of wheat infection. The wheat morbidity is carried out simultaneously with the canopy spectrum collection. According to the "National Standard for Wheat Fusarium Scab Survey", the ratio of the number of diseased wheat ears in the plot to the total number of surveyed ears DER is used as an important evaluation indicator for the severity of wheat infection.
[0069] In step (2), the dimensionality reduction of hyperspectral data using the CARS algorithm specifically includes the following steps:
[0070] (2a) Performing Monte Carlo model sampling: Sampling variables through Monte Carlo MC in an iterative and competitive manner. Setting the Monte Carlo MC sampling to run N times and obtain N subset variables. In each sampling run, randomly select 80% of the samples in the subset variables to establish the Monte Carlo model.
[0071] (2b) Wavelength selection by exponential decreasing function EDF: The exponential decreasing function EDF is used to force the wavelength to decrease and remove the variables with small absolute values of regression coefficients;
[0072] (2c) Adaptive reweighted sampling: further refine the wavelength selection and retain the variables with larger absolute values of regression coefficients to form a new feature set;
[0073] (2d) Iterative loop: The new feature set is input into the Monte Carlo model and the root mean square error of the ten-fold cross validation is calculated. The subset corresponding to the minimum root mean square error recorded in N runs is the optimal wavelength combination screened out. The optimal wavelength combination constitutes the feature set determined by the CARS algorithm;
[0074] The exponentially decreasing function EDF is defined as:
[0075] r i =ae -ki
[0076] Where i represents the number of sampling runs; in the Nth sampling run, only two wavelengths are retained, r1 = 1, r N =2 / P; a and k are constants, a is the value of the exponentially decreasing function EDF at i=0, k is the rate of decrease of the exponentially decreasing function EDF, and the calculation formulas of a and k are as follows:
[0077]
[0078]
[0079] The number of wavelengths P is 501, and the total number N of sampling runs is set to 50.
[0080] In addition to the CARS algorithm, the PCA algorithm can reduce the dimensionality of the dataset, enhance interpretability, and minimize information loss. The SPA algorithm uses vector projection analysis to identify the smallest set of redundant variables in spectral data, thereby reducing collinearity between variables and the number of modeling variables. Therefore, in this paper, the PCA algorithm and the SPA algorithm were used to compare with the CARS algorithm, and a total of three different feature sets were obtained.
[0081] In step (3), modeling by the Ridge algorithm specifically refers to:
[0082] Assuming the linear regression model y=xβ+ε, the objective function of the regression parameter β is:
[0083]
[0084] in, x is the independent variable, i.e., the canopy spectral reflectance of the sample plot; y is the dependent variable, i.e., the incidence rate; β and ε are the regression coefficient and error, respectively; n is the number of samples; λ is the regularization parameter, whose value is determined by the ridge trace method to balance the variance and bias of the linear regression model; the bias of the linear regression model increases with the increase of λ, while the variance increases in the opposite direction;
[0085] The objective function of the regression coefficient β is equivalent to:
[0086]
[0087] Where x i and y i Represents the i-th independent variable and dependent variable respectively, n represents the number of samples, α represents the regression coefficient vector, which is used to fit the unknown parameters of the linear relationship between the independent variable and the dependent variable; according to the β-λ curve in the Cartesian coordinate system, that is, the ridge curve, at the point where β tends to be stable, find the optimal λ value.
[0088] In addition to the Ridge algorithm, the RF algorithm is also a useful algorithm that integrates the performance of the decision tree algorithm for classification or prediction of variable values. The PLSR algorithm is a common method for multivariate data analysis and is increasingly used in research fields such as bioinformatics, machine learning, and chemometrics. Therefore, to demonstrate the effectiveness of the Ridge algorithm, the present invention also uses the RF algorithm and the PLSR algorithm for modeling.
[0089] The step (4) specifically includes the following steps:
[0090] (4a) Three different feature sets obtained from the CARS, PCA, and SPA algorithms were used as input variables and combined with the RF, PLSR, and Ridge algorithms to construct a total of nine wheat fusarium head blight recognition models: CARS-RF, CARS-PLSR, CARS-Ridge, PCA-RF, PCA-PLSR, PCA-Ridge, SPA-RF, SPA-PLSR, and SPA-Ridge. The three different input feature sets were evenly divided into ten parts, nine for modeling and one for validation, and the results were repeated ten times. The ten-fold cross-validation method was used to evaluate the recognition ability of each model for wheat fusarium head blight.
[0091] (4b) The accuracy and performance of each model are determined by the coefficient of determination R 2 and root mean square error RMSE are used to evaluate the two statistical indicators, R 2 The calculation formulas for RMSE are:
[0092]
[0093]
[0094] Where n is the number of samples, t i is the actual value, d i is the model prediction value; the smaller the RMSE value, the more accurate the model; R 2 The value of is between 0 and 1, the closer to 1 the better, indicating the degree to which the model explains the changes in the dependent variable; when R 2 When R is 1, the model fits the data perfectly; when R 2 When it is 0, the model cannot explain the changes in the dependent variable;
[0095] (4c) In all models, R 2 The closer the model is to 1 and the closer the RMSE is to 0, the better its ability to identify diseases is. According to the model operation results, the CARS-Ridge model is determined to be the optimal model.
[0096] like Figure 3 As shown, the step (5) specifically includes the following steps:
[0097] (5a) First, the correlation between the full spectrum band and the ratio of diseased wheat ears in the sample to the total number of wheat ears in the survey (DER) was calculated according to the calculation formula of the Pearson correlation coefficient. The correlation result was between 0 and 1. A fixed threshold was selected to divide the bands. The wavelengths with a correlation coefficient greater than 0.7 were selected as the primary wavelengths. Assuming (x1, y1), (x2, y2),…, (x n ,y n ) represent the values of two variables X and Y, n is the number of samples, and the calculation formula of Pearson correlation coefficient is:
[0098]
[0099] Where R is the Pearson correlation coefficient, x i and y i denote the i-th independent variable and dependent variable respectively;
[0100] is the covariance, which measures the correlation between variables X and Y, and are the mean values of the samples, namely:
[0101]
[0102]
[0103] (5b) The correlation between the vegetation index and the proportion of diseased wheat ears in the sample to the total number of wheat ears in the survey (DER) was calculated according to the calculation formula of the Pearson correlation coefficient. Similarly, the vegetation index with a correlation coefficient with DER of more than 0.7 was selected as the primary vegetation index according to the threshold division; the 12 vegetation indices were sensitive pigment index SIPI, anthocyanin reflectance index ARI, transformed chlorophyll absorption reflectance index TCARI, modified chlorophyll absorption ratio index MCARI, normalized difference vegetation index NDVI, modified simple ratio index MSR, ratio vegetation structure index RVSI, plant senescence reflectance index PSRI, triangular vegetation index TVI, nitrogen reflectance index NRI, photochemical reflectance index PRI and physiological reflectance index PHRI;
[0104] (5c) The correlation between the primary wavelength and the primary vegetation index was calculated according to the calculation formula of the Pearson correlation coefficient. The threshold division was used to further divide the primary wavelengths with a correlation coefficient greater than 0.7 with the primary vegetation index. The wavelengths that were finally retained were used to construct two new indices WFI. two and WFI three :
[0105]
[0106]
[0107] Among them, R 687 、R 659 and R 760 is the wavelength that is ultimately retained.
[0108] Universal analysis of the new index: the new index, 12 vegetation indices and the ratio of diseased wheat ears in the sample to the total number of surveyed ears (DER) were fitted to obtain the linear fitting equations of each index and DER. 2 The superiority and necessity of the new index were evaluated by the value of R and RMSE; the universality of the new index was further verified by independent data of wheat fusarium canopy.
[0109] The step (6) specifically refers to: using the new index and the feature set determined based on the CARS algorithm as inputs of the optimal model to construct an optimal wheat fusarium head blight recognition model.
[0110] Example 1
[0111] In this study, Experiment 1 was conducted in Baihu Town, Lujiang County, Hefei City, Anhui Province in 2018, with Shengxuan 6 as the primary cultivar. Experiment 2 was conducted in Guohe Town, Lujiang County, Hefei City, Anhui Province in 2019, with Yangmai 25 as the primary cultivar. Scab naturally occurred in both study areas. Experiment 3 was conducted in 2022 at the Xiaotangshan Experimental Base in Beijing, with Yangmai 25 and Zhengmai 9023 as the primary cultivars. The occurrence of wheat scab was artificially inoculated. Wheat plants were grown under controlled laboratory conditions and sprayed with a mixture of pathogens at varying concentrations during the flowering stage for three consecutive days. The laboratory temperature was maintained at 20-25°C and the humidity at 70% to ensure successful inoculation.
[0112] All wheat canopy spectral reflectance measurements were collected using a FieldSpec 4 spectrometer. This spectrometer has a 25° field of view and covers the spectral range from 350 to 2500 nm. Spectral resolution is 3 nm in the 350-1000 nm and 10 nm in the 1000-2500 nm range, respectively. Data collection took place between 10:00 AM and 2:00 PM, ensuring clear, cloudless, and windless weather. To maximize the collection of effective wheat canopy information and minimize the impact of background noise on spectral reflectance, the sensor field of view completely covered the top of the wheat plot, and the sensor probe was always pointed vertically downward during the acquisition process. Spectral data were collected from wheat during the grain-filling stage. Ten measurements were collected each time, and the average of these ten measurements was used as the final spectral reflectance of the sample. Before each measurement, a 40 cm x 40 cm standard white plate was used to correct for changes in illumination. In Experiment 1, canopy reflectance measurements were collected for a total of 198 wheat samples: 58 samples from Experiment 1, 100 samples from Experiment 2, and 40 samples from Experiment 3. Table 1 summarizes the basic description of all experimental information:
[0113] Table 1
[0114]
[0115]
[0116] Most current studies on wheat fusarium head blight can qualitatively distinguish between healthy and infected individuals, or simply grade the severity of the disease. However, quantitative research on the severity of wheat infection is still very limited. Therefore, in the present invention, the incidence of each sample point is counted and used to quantitatively estimate the severity of wheat infection. The wheat disease severity survey is carried out simultaneously with the canopy spectrum collection. According to the "National Standard for Survey Specifications of Wheat Fusarium Head Blight", the diseased ear rate is used as an important assessment indicator of the degree of wheat disease. The ratio of the number of diseased wheat ears in the sample plot to the total number of ears surveyed, DER, ranges from 0-100%. Based on this, the diseased ear rate of each sample plot is finally calculated.
[0117] Three different algorithms, CARS, PCA and SPA, were used to directly reduce the dimensionality of hyperspectral data and obtain different sets of sensitive bands. Three different algorithms, RF, PLSR and Ridge, were selected for modeling.
[0118] Hyperspectral data across the entire wavelength range has high dimensionality and redundancy, which increases computational complexity. This highly redundant data can also lead to unstable convergence in prediction or classification models. In many cases, using optimal wavelengths rather than the full wavelength range can yield better classification results. Therefore, identifying sensitive information across a large number of wavelengths is a crucial challenge in hyperspectral remote sensing research. This paper employs three algorithms, CARS, SPA, and PCA, to reduce the dimensionality of hyperspectral data.
[0119] In order to accurately detect the severity of the disease and obtain more information, different machine learning methods are applied in model development. In this invention, RF, PLSR and Ridge algorithms are used to build the model.
[0120] Based on the optimized parameters in RF, PLSR and Ridge regression models, the present invention uses three different feature sets obtained by three dimensionality reduction algorithms, CARS, PCA and SPA, as input variables, combined with three modeling algorithms, RF, PLSR and Ridge, to construct a total of 9 models. The classification model was tested on ten prediction sets to compare the performance of different models in identifying wheat scab. Figure 2 As shown, in general, the performance of different models is not the same. The R 2 The values for PCA-RF and PCA-PLSR were poorly evaluated and could not be used as evaluation methods for wheat fusarium head blight severity. While PCA-Ridge, SPA-RF, SPA-PLSR, SPA-Ridge, and CARS-RF could reflect the severity of wheat fusarium head blight infection to a certain extent, they were not optimal for this application. CARS-Ridge and CARS-PLSR performed well, but CARS-Ridge showed the greatest potential for identifying canopy fusarium head blight.
[0121] The huge amount of hyperspectral data will limit the processing capacity of the software. In addition to using algorithms to screen out sensitive bands, the remote sensing field often uses mathematical combinations of bands to construct vegetation indices. These indices are usually used as representative subsets of all band variables for the target problem to reduce the "curse of dimensionality". Although these vegetation indices can characterize the changes in crop physiology and biochemistry under disease stress, the external manifestations of wheat infected with ergot are reflected in the ear, and the disease information at the top of the wheat ear that can be obtained by remote sensing means is relatively weak. Therefore, traditional vegetation indices do not show good robustness in the identification of ergot. Many studies on ergot have adopted similar methods, and almost all studies use aggregated traditional vegetation indices as sensitive features for identifying diseases, which greatly limits the effectiveness of model disease identification. The present invention constructs a new spectral index based on the spectral response of crop diseases and traditional vegetation indices. The new index shows great potential in the study of quantification of the severity of wheat ergot, and also shows good robustness.
[0122] The new index was fitted with published vegetation indices and wheat disease severity to evaluate the superiority and necessity of the new index, and an independent data set was used to further verify the universality of the new index.
[0123] (1) Table 2 compares the predictive ability of all indicators for wheat diseases, including the new indicators proposed in this paper and the indicators in the published literature. Among all 14 indicators, WFI two (R 2 =0.604, RMSE=19.579) and WFI three (R 2 =0.641, RMSE=18.652) has the best predictive ability. Most of these published indicators use green light, red edge and near infrared bands. Although they can reflect the health status of wheat as the disease develops to a certain extent, none of them shows the same prediction ability as WFI. two and WFI three The results show that when using WFI two and WFI three In the case of the formula, the optimal wavelengths of 687 nm, 760 nm, and 659 nm may be crucial for improving the high predictive ability of the model.
[0124] Table 2 Analysis of the identification results of wheat scab by all indicators
[0125]
[0126] The feasibility and universality of the new index were evaluated using wheat canopy hyperspectral data obtained in 22 years. The results are shown in Table 3. two (R2 =0.4955, RMSE=16.759) and WFI three (R 2 =0.5103, RMSE =16.511) had the highest recognition accuracy. The selected vegetation index is the most commonly used index in crop disease diagnosis and has been widely used in the identification of diseases such as wheat rust and powdery mildew, and has also been applied to wheat scab. Analysis of the two experimental data shows that although most published vegetation indices have good application potential in wheat scab identification, they are not as good as WFI. two and WFI three Accurate and reliable.
[0127] Table 3 Analysis of the universality results of the new index
[0128]
[0129] In the early stage of this invention, the optimal wavelength based on the CARS algorithm and the optimal modeling method based on the Ridge algorithm were determined, proving the necessity of the CARS-Ridge algorithm in the identification of wheat scab. two and WFI three Compared with the good fusarium identification ability of the traditional vegetation index, the new index is considered to be combined with the optimal wavelength of the hyperspectral extracted by the CARS algorithm and modeled using the Ridge algorithm to further improve the model accuracy, such as Figure 4 As shown, it can be seen that compared with Figure 2 The CARS-Ridge model recognition results (R 2 =0.892, RMSE=10.252), WFI two and WFI three Integration into the CARS-Ridge model can more effectively detect the degree of wheat DER infection, with an accuracy rate 2% higher than that of the CARS-Ridge model (R 2 =0.914, RMSE=9.114). Using CARS-Ridge and WFI two and WFI three The complementary information can capture the details of the disease more accurately, and the fitting line of the scatter plot is closer to the 1:1 line. These results further prove that the CARS-Ridge algorithm and WFI are related to wheat scab identification. two and WFI three Reliability and stability of the new index.
[0130] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions merely illustrate the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A wheat scab identification method based on the CARS-Ridge algorithm and a new index, characterized by: The method comprises the following steps in sequence: (1) Data acquisition: Obtain hyperspectral data of wheat fusarium canopy and calculate the incidence rate of each sample plot; (2) Data dimensionality reduction: CARS, PCA and SPA algorithms were used to reduce the dimensionality of the acquired wheat fusarium canopy hyperspectral data, and three different feature sets were obtained; (3) Modeling: Using RF, PLSR, and Ridge algorithms, we obtained nine wheat fusarium head blight recognition models, namely, CARS-RF, CARS-PLSR, CARS-Ridge, PCA-RF, PCA-PLSR, PCA-Ridge, SPA-RF, SPA-PLSR, and SPA-Ridge. (4) Determine the optimal model: The optimal model was determined by performing a ten-fold cross-validation on the results of the nine wheat scab recognition models; (5) Development of new indices: Based on 12 vegetation indices, two new indices, namely WFI, were constructed. two and WFI three ; (6) Constructing the optimal wheat scab recognition model: integrating the new index with the optimal model to obtain the optimal wheat scab recognition model; The step (5) specifically includes the following steps: (5a) First, the correlation between the full spectrum band and the ratio of diseased wheat ears in the sample to the total number of wheat ears in the survey (DER) was calculated according to the calculation formula of the Pearson correlation coefficient. The correlation result was between 0 and 1. A fixed threshold was selected to divide the bands. The wavelengths with a correlation coefficient greater than 0.7 were selected as the primary wavelengths. Assuming (x1, y1), (x2, y2),…, (x n ,y n ) represent the values of two variables X and Y, n is the number of samples, and the calculation formula of Pearson correlation coefficient is: Where R is the Pearson correlation coefficient, x i and y i denote the i-th independent variable and dependent variable respectively; is the covariance, which measures the correlation between variables X and Y, and are the mean values of the samples, namely: (5b) The correlation between the vegetation index and the proportion of diseased wheat ears in the sample to the total number of wheat ears in the survey (DER) was calculated according to the calculation formula of the Pearson correlation coefficient. Similarly, the vegetation index with a correlation coefficient with DER of more than 0.7 was selected as the primary vegetation index according to the threshold division; the 12 vegetation indices were sensitive pigment index SIPI, anthocyanin reflectance index ARI, transformed chlorophyll absorption reflectance index TCARI, modified chlorophyll absorption ratio index MCARI, normalized difference vegetation index NDVI, modified simple ratio index MSR, ratio vegetation structure index RVSI, plant senescence reflectance index PSRI, triangular vegetation index TVI, nitrogen reflectance index NRI, photochemical reflectance index PRI and physiological reflectance index PHRI; (5c) The correlation between the primary wavelength and the primary vegetation index was calculated according to the calculation formula of the Pearson correlation coefficient. The threshold division was used to further divide the primary wavelengths with a correlation coefficient greater than 0.7 with the primary vegetation index. The wavelengths that were finally retained were used to construct two new indices WFI. two and WFI three : Among them, R 687 、R 659 and R 760 is the wavelength that is ultimately retained.
2. The wheat scab identification method based on the CARS-Ridge algorithm fusion novel index according to claim 1, characterized in that: The step (1) specifically includes the following steps: (1a) Spectral data measurement: Hyperspectral data of wheat scab canopy were obtained. Disease occurrence was divided into two types: natural occurrence and artificial inoculation. All wheat canopy spectral reflectances were measured using a FieldSpec4 spectrometer. (1b) Statistical morbidity: The morbidity of each plot is calculated to quantitatively estimate the severity of wheat infection. The wheat morbidity is carried out simultaneously with the canopy spectrum collection. The DER, the ratio of the number of diseased wheat ears in the plot to the total number of surveyed ears, is used as an important evaluation indicator of the severity of wheat infection.
3. The wheat scab identification method based on the CARS-Ridge algorithm fusion novel index according to claim 1, characterized in that: In step (2), the dimensionality reduction of hyperspectral data using the CARS algorithm specifically includes the following steps: (2a) Performing Monte Carlo model sampling: Sampling variables through Monte Carlo MC in an iterative and competitive manner. Setting the Monte Carlo MC sampling to run N times and obtain N subset variables. In each sampling run, randomly select 80% of the samples in the subset variables to establish the Monte Carlo model. (2b) Wavelength selection by exponential decreasing function EDF: The exponential decreasing function EDF is used to force the wavelength to decrease and remove the variables with small absolute values of regression coefficients; (2c) Adaptive reweighted sampling is performed: variables with larger absolute values of regression coefficients are retained to form a new feature set; (2d) Iterative loop: The new feature set is input into the Monte Carlo model and the root mean square error of the ten-fold cross validation is calculated. The subset corresponding to the minimum root mean square error recorded in N runs is the optimal wavelength combination screened out. The optimal wavelength combination constitutes the feature set determined by the CARS algorithm; The exponentially decreasing function EDF is defined as: r i =ae -ki Where i represents the number of sampling runs; in the Nth sampling run, only two wavelengths are retained, r1 = 1, r N =2 / P; a and k are constants, a is the value of the exponentially decreasing function EDF at i=0, k is the rate of decrease of the exponentially decreasing function EDF, and the calculation formulas of a and k are as follows: The number of wavelengths P is 501, and the total number N of sampling runs is set to 50.
4. The wheat scab identification method based on the CARS-Ridge algorithm fused with a novel index according to claim 1, characterized in that: In step (3), modeling by the Ridge algorithm specifically refers to: Assuming the linear regression model y=xβ+ε, the objective function of the regression parameter β is: in, x is the independent variable, i.e., the canopy spectral reflectance of the sample plot; y is the dependent variable, i.e., the incidence rate; β and ε are the regression coefficient and error, respectively; n is the number of samples; λ is the regularization parameter, whose value is determined by the ridge trace method to balance the variance and bias of the linear regression model; the bias of the linear regression model increases with the increase of λ, while the variance increases in the opposite direction; The objective function of the regression coefficient β is equivalent to: Where x i and y i Represents the i-th independent variable and dependent variable respectively, n represents the number of samples, α represents the regression coefficient vector, which is used to fit the unknown parameters of the linear relationship between the independent variable and the dependent variable; according to the β-λ curve in the Cartesian coordinate system, that is, the ridge curve, at the point where β tends to be stable, find the optimal λ value.
5. The wheat scab identification method based on the CARS-Ridge algorithm fusion novel index according to claim 1, characterized in that: The step (4) specifically includes the following steps: (4a) Three different feature sets obtained from the CARS, PCA, and SPA algorithms were used as input variables and combined with the RF, PLSR, and Ridge algorithms to construct a total of nine wheat fusarium head blight recognition models: CARS-RF, CARS-PLSR, CARS-Ridge, PCA-RF, PCA-PLSR, PCA-Ridge, SPA-RF, SPA-PLSR, and SPA-Ridge. The three different input feature sets were evenly divided into ten parts, nine for modeling and one for validation, and the results were repeated ten times. The ten-fold cross-validation method was used to evaluate the recognition ability of each model for wheat fusarium head blight. (4b) The accuracy and performance of each model are determined by the coefficient of determination R 2 and root mean square error RMSE are used to evaluate the two statistical indicators, R 2 The calculation formulas for RMSE are: Where n is the number of samples, t i is the actual value, d i is the model prediction value; the smaller the RMSE value, the more accurate the model; R 2 The value of is between 0 and 1, the closer to 1 the better, indicating the degree to which the model explains the changes in the dependent variable; when R 2 When R is 1, the model fits the data perfectly; when R 2 When it is 0, the model cannot explain the changes in the dependent variable; (4c) In all models, R 2 The closer the model is to 1 and the closer the RMSE is to 0, the better its ability to identify diseases is. According to the model operation results, the CARS-Ridge model is determined to be the optimal model.
6. The wheat scab identification method based on the CARS-Ridge algorithm fused with a novel index according to claim 1, characterized in that: The step (6) specifically refers to: using the new index and the feature set determined based on the CARS algorithm as inputs of the optimal model to construct an optimal wheat fusarium head blight recognition model.