Rice smut extraction method and system based on hyperspectral data band selection
By combining genetic algorithm with partial least squares, correlation coefficient and inter-class instability index method to optimize the selection of hyperspectral data bands, the problem of local optimal combination of characteristic bands was solved, and the accuracy and speed of rice smut monitoring were improved.
Patent Information
- Application Number
- CN202211482287.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-24
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2042-11-24
AI Technical Summary
In the monitoring of crop diseases and insect pests, the characteristic band combinations of existing hyperspectral remote sensing images are often locally optimal, making it difficult to further improve monitoring accuracy and being affected by noise data.
A genetic algorithm combined with the partial least squares method was used to perform preliminary feature band selection, and the correlation coefficient method and the inter-class instability index method were used to further optimize the feature bands, reduce the dimension of the hyperspectral data, denoise it, and improve the monitoring accuracy.
By optimizing the characteristic band combination, reducing the data dimension and eliminating noise, the accuracy of rice smut monitoring and the speed of model monitoring were improved, ensuring that the prediction accuracy was not reduced.
Smart Images

Figure CN115830444B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing hyperspectral image data dimensionality reduction, and in particular to a rice smut extraction method and system based on hyperspectral data band selection. Background Art
[0002] During its growth and development, rice plants are exposed to various fungi, which can lead to severe declines in quality and yield. Rice false smut (RFS) is a late-stage fungal disease caused by the ascomycete pathogen Villosiclava virens, which primarily targets the rice panicle. RFS depletes the entire rice panicle, resulting in a severe decline in rice quality and significant losses. RFS reduces thousand-grain weight and seed germination by up to 35%. After rice planting, the pathogen remains viable in the soil and infects seedlings. Although RFS is primarily concentrated in small areas around the original source, indiscriminate spraying of pesticides over the entire field is a common practice, increasing prevention costs and causing environmental pollution. To minimize the economic losses and environmental pollution caused by pesticides, it is necessary to accurately assess the distribution and severity of RFS. An automated, non-destructive, rapid, sensitive, and selective method is needed to rapidly detect plant diseases, reduce pesticide and fertilizer use, and support sustainable agricultural production. Remote sensing technology has shown unique advantages in the field of crop disease and pest stress monitoring due to its accuracy, speed, large area and non-destructiveness.
[0003] In recent years, with the rapid development of the drone industry, drone-based clothing remote sensing has played a significant role in crop pest and disease monitoring due to its high spatial resolution, timely data acquisition, and low cost. Therefore, drone-based hyperspectral photogrammetry is an effective method for rapid and accurate monitoring of small- and medium-scale crop pests and diseases.
[0004] Many researchers have used drone remote sensing imagery to monitor crop pests and diseases. For example, high-spatial-resolution aerial imagery was used to monitor the extent of yellow leaf spot disease in banana crops, achieving excellent accuracy when employing support vector machine methods. Using multispectral cameras and drones to acquire time-series aerial multispectral imagery, and studying crop spectral data from different periods, an effective method for monitoring early-stage crop pests and diseases was developed. In addition to using high-resolution and multispectral imagery to monitor crop pests and diseases, hyperspectral imagery is also an important tool for detecting crop pests and diseases. Some researchers have used hyperspectral imaging technology to identify tomato yellow leaf curl disease, using spectral characteristic parameters and characteristic spectral bands, such as first-order derivative reflectance spectra and absolute reflectance difference spectra, to accurately monitor crop disease status. Other researchers have also selected characteristic spectral bands within hyperspectral bands to reduce the dimensionality of hyperspectral data, achieving higher accuracy.
[0005] In research using hyperspectral remote sensing imagery to monitor crop pests and diseases, healthy crops are often inoculated with the relevant pests to achieve a more uniform infection, allowing for more detailed study of the various stages of disease progression. However, in nature, pests and diseases typically infect crops from a few locations and gradually spread to nearby healthy crops, resulting in more complex disease patterns. Therefore, to develop more accurate crop pest and disease monitoring models, in addition to actively inoculating crops with pests and diseases and studying the changes in spectral characteristic curves at various stages of disease progression, it is also necessary to focus on studying the infection patterns of crops under natural conditions. Existing studies have used machine learning classification methods such as support vector machines (SVMs) and histogram analysis for feature selection, while others have used neural networks (CNNs) to build models across the entire spectrum or genetic algorithms to extract feature bands. Many researchers have used random forest (RF) models to develop crop pest and disease detection models, achieving high accuracy. In summary, spectral analysis and calculation of spectral parameters from hyperspectral imagery are the primary methods for pest and disease monitoring, and the use of machine learning models and deep learning networks is a current research hotspot.
[0006] Furthermore, using hyperspectral remote sensing imagery for crop pest and disease monitoring and reducing data dimensionality through band selection of hyperspectral data are also approaches to improving pest and disease monitoring accuracy. Some studies have used genetic algorithms combined with partial least squares methods to screen hyperspectral feature bands, guided regularized random forests (GRRFs) for hyperspectral feature band screening, or combined genetic algorithms with support vector machines (SVMs) to select hyperspectral bands. However, the methods used in these studies only provide preliminary selection of preferred feature bands, and the resulting feature band combinations are locally optimal. To further improve the accuracy of crop pest and disease monitoring, further selection of feature bands from hyperspectral bands is necessary to achieve even more optimal feature band combinations. Summary of the Invention
[0007] In view of this, the purpose of the present invention is to propose a rice smut extraction method and system based on the selection of UAV hyperspectral data bands. The method utilizes the correlation and inter-class separability of hyperspectral bands, selects feature bands based on genetic algorithm combined with partial least squares, and further selects feature bands by using correlation coefficient method and inter-class instability index method, so as to obtain a better feature band combination, thereby achieving the purpose of reducing the dimension of hyperspectral data and denoising, and ultimately improving the monitoring accuracy of rice smut.
[0008] In order to achieve the above object, the present invention provides the following technical solutions:
[0009] The rice smut extraction method based on hyperspectral data band selection provided by the present invention comprises the following steps:
[0010] Acquire UAV hyperspectral image data;
[0011] Furthermore, the acquisition of hyperspectral image data and preprocessing are performed according to the following steps:
[0012] Acquire hyperspectral images of the study area to obtain images that are hyperspectral images obtained by multiple photogrammetry measurements over a period of time;
[0013] Perform filtering and normalization operations on hyperspectral image data.
[0014] Conduct field measurements and obtain the location distribution information of healthy rice and rice with smut disease in rice areas;
[0015] Comparing hyperspectral image data with field measurements of healthy and smut-affected rice locations yielded spectral reflectance data at the sampling points.
[0016] Constructing a prediction model, wherein the prediction model uses partial least squares and sampling point spectral reflectance data to screen out preferred characteristic bands;
[0017] A characteristic band optimization threshold is determined, and an optimal characteristic band combination is selected from the spectral reflectance data of the sampling points according to the characteristic band optimization threshold. The optimal characteristic band combination is used as rice smut prediction data.
[0018] Furthermore, the characteristic waveband preferred threshold value is obtained by calculating a correlation coefficient to obtain a correlation coefficient threshold value; or / and the characteristic waveband preferred threshold value is obtained by calculating an inter-class instability index to obtain an inter-class instability index threshold value.
[0019] Furthermore, the method further comprises the following steps:
[0020] The evaluation criteria for the prediction model are established using the confusion matrix and accuracy, precision, recall, and F1 score according to the binary classification situation. The matrix is established as follows:
[0021] Confusion Matrix
[0022]
[0023] Among them, TP is the number of predicted positive examples and the actual positive examples, FP is the number of actual negative examples but predicted positive examples; TN is the number of predicted negative examples and the actual negative examples, FN is the number of actual negative examples but predicted positive examples;
[0024] The accuracy, precision, recall and F1 score are calculated according to the following formula:
[0025]
[0026]
[0027]
[0028]
[0029] Among them, accuracy represents the matrix and accuracy; precision represents the precision rate; recall represents the recall rate; and F1 represents the F1 score.
[0030] Furthermore, the correlation coefficient is calculated according to the following steps:
[0031]
[0032] Where r is the Pearson correlation coefficient, and are the means of variables X and Y, Xi and Yi are the element values of variables X and Y respectively; n is the number of elements of the variable.
[0033] Furthermore, the correlation coefficient threshold is performed according to the following steps:
[0034] The preferred hyperspectral band is obtained by comparing the correlation coefficient with the preset threshold, and the prediction accuracy of the preferred hyperspectral band is calculated using the prediction model. The correlation coefficient threshold with the highest prediction accuracy is obtained as the correlation coefficient threshold of the monitoring model.
[0035] Furthermore, the inter-class instability index is performed according to the following steps:
[0036]
[0037] In the formula, ISIC i is the inter-class instability index of the two classes of samples at the i-th band, Δwithin,i and Δbetween,i are the intra-class deviation and inter-class deviation respectively, S 1,i is the standard deviation of the first type of samples in the i-th band, S 2,i is the standard deviation of the second type of samples in the i-th band, m 1,i is the mean of the first type of samples in the i-th band, m 2,i is the mean of the second type of samples in the i-th band.
[0038] Furthermore, the inter-class instability index threshold is determined by the following steps:
[0039] The threshold interval is set according to the inter-class instability index, and a certain threshold is used to select the preferred feature band within the threshold interval. The inter-class instability index threshold is obtained based on the preferred feature band and the prediction model.
[0040] Furthermore, the hyperspectral data normalization process is performed according to the following formula:
[0041]
[0042] In the formula, x i is the spectral reflectance value of a certain band, x max and x min are the maximum and minimum values of a band respectively.
[0043] The rice smut extraction system based on hyperspectral data band selection includes a memory, a processor, and a computer program stored in the memory and runnable on the processor, characterized in that the above method is implemented when the processor executes the program.
[0044] The beneficial effects of the present invention are:
[0045] The present invention provides a rice smut extraction method and system based on hyperspectral data band selection. The method uses a genetic algorithm combined with a partial least squares method to preliminarily select hyperspectral bands, and then uses a correlation coefficient method and an inter-class instability index method to further optimize the preliminarily selected feature bands. Compared with traditional band optimization and hyperspectral data dimensionality reduction methods, the method used in the present invention optimizes a better feature band combination, has a smaller number of feature bands, and has higher prediction accuracy after feature band modeling. Therefore, while ensuring that the prediction accuracy is not reduced, the data dimension is further reduced and noise data is eliminated, thereby improving the accuracy of rice smut monitoring and the model monitoring speed.
[0046] Other advantages, objects, and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art upon examination of the following description or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to make the purpose, technical solutions and beneficial effects of the present invention more clear, the present invention provides the following drawings for illustration:
[0048] Figure 1 This is an information map of the geographical location of the study area and the occurrence of rice false smut.
[0049] Figure 2 Comparison of rice spectral data before and after filtering.
[0050] Figure 3 This is the workflow diagram for band optimization.
[0051] Figure 4 Flowchart for genetic algorithm calculation.
[0052] Figure 5 This is the result of genetic algorithm band screening.
[0053] Figure 6 This is the calculation result of the correlation coefficient of the genetic algorithm screening characteristic bands.
[0054] Figure 7 This is the calculation result of the inter-class instability index of each characteristic band screened by the genetic algorithm. DETAILED DESCRIPTION
[0055] The present invention will be further described below with reference to the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it. However, the embodiments are not intended to limit the present invention.
[0056] Example 1
[0057] like Figure 1 As shown, the rice smut extraction method based on the selection of UAV hyperspectral data bands provided in this embodiment includes the following steps:
[0058] Acquire UAV hyperspectral images of the study area and perform preprocessing operations such as radiometric correction, geocoding, and geometric correction. The resulting images are hyperspectral images obtained from multiple UAV photogrammetry measurements over a period of time.
[0059] Perform filtering and normalization operations on UAV hyperspectral image data;
[0060] Conduct field measurements and obtain the location distribution information of healthy rice and rice with smut disease in rice areas;
[0061] Comparing UAV hyperspectral images with field measurements of healthy rice and rice smut locations yielded spectral reflectance data for a series of sampling points.
[0062] According to the spectral reflectance data of the sampling points, the dataset is divided into a training dataset and a test dataset with a ratio of 7:3;
[0063] Table: Dataset division
[0064]
[0065] In the table, yes and no indicate whether the disease occurs, and m1-m2 are the sampling area numbers;
[0066] Based on the binary classification, the confusion matrix and accuracy, precision, recall, and F1-score are used to establish the evaluation criteria for the prediction results. These are used as the evaluation criteria for estimating the prediction results of the following four models.
[0067] Table: Confusion Matrix
[0068]
[0069] The accuracy, precision, recall and F1-score are calculated according to the following formula:
[0070]
[0071]
[0072]
[0073]
[0074] Where, accuracy represents the matrix and accuracy; precision represents the precision rate; recall represents the recall rate; F1 represents the score; TP is the number of predicted positive examples and the actual positive examples, FP is the number of actual negative examples but predicted positive examples; TN is the number of predicted negative examples and the actual negative examples, and FN is the number of actual negative examples but predicted positive examples.
[0075] Constructing a genetic algorithm prediction model, wherein the prediction model uses a partial least squares sieve to calculate the spectral reflectance data of the sampling points and select the optimal characteristic bands;
[0076] In this embodiment, the monitoring model is constructed based on the genetic algorithm. Figure 4 As shown, the genetic algorithm calculation flow chart and steps are as follows:
[0077] (1) Coding: Using binary coding method, “0” means not selected; “1” means selected.
[0078] (2) Generation of the initial population: Randomly generate N initial strings to form the initial population.
[0079] (3) Fitness function. The fitness function indicates the quality of an individual or solution. For feature selection problems, the construction of the fitness function is very important. It is mainly based on the separability criterion of the category and the classification ability of the feature.
[0080] (4) The individual with the highest fitness, that is, the best individual in the population, is unconditionally copied to the next generation population, and then the parent population is subjected to genetic operators such as selection, crossover and mutation to reproduce the other n-1 gene strings of the next generation population.
[0081] (5) If the set number of generations is reached, the best chromosome is returned and used as the basis for feature selection, and the algorithm ends. Otherwise, return to (4) to continue the next generation of reproduction.
[0082] Using genetic algorithm and taking RMSE calculated by partial least squares as fitness function, a hyperspectral band selection model was constructed and appropriate model initialization parameters were set.
[0083] Based on the genetic algorithm combined with the partial least squares algorithm, the frequency of each band selection is obtained by twenty operations, and three appropriate frequency benchmarks are selected to obtain the preliminary preferred characteristic bands. The prediction accuracy of the model is established based on the characteristic bands selected by each benchmark, and the characteristic band combination with the highest prediction accuracy is obtained;
[0084] According to the following formula, the Pearson correlation coefficient between the characteristic bands selected by the genetic algorithm is calculated;
[0085]
[0086] In the formula, r is the Pearson correlation coefficient, and are the means of variables X and Y, respectively. i and Y i are the element values of variables X and Y respectively; n is the number of variable elements.
[0087] First, a threshold value between 0.8 and 1.0 was selected as the criterion for optimizing the hyperspectral bands based on the correlation coefficient;
[0088] In this embodiment, each time a threshold is selected, bands that do not meet the threshold conditions are eliminated from the preliminary preferred bands, and the preferred bands are used to build a monitoring model and calculate the prediction accuracy;
[0089] Finally, the correlation coefficient threshold with the highest prediction accuracy is obtained;
[0090] According to the threshold value obtained above, the bands whose correlation coefficients are greater than the threshold condition in the characteristic bands selected by the genetic algorithm are eliminated to obtain the characteristic band combination selected by the correlation coefficient method.
[0091] According to the following formula, the inter-class instability index of the characteristic band selected by the genetic algorithm is calculated;
[0092]
[0093] In the formula, ISIC i is the inter-class instability index of the two classes of samples at the i-th band, Δwithin,i and Δbetween,i are the intra-class deviation and inter-class deviation respectively, S 1,i is the standard deviation of the first type of samples in the i-th band, S 2,i is the standard deviation of the second type of samples in the i-th band, m 1,i is the mean of the first type of samples in the i-th band, m 2,i is the mean of the second type of samples in the i-th band.
[0094] The selection is made within the interval where the inter-class instability index is concentrated, a certain step size is set, and a certain threshold is used to select the preferred band within the threshold interval. Then, the selected preferred band is used to establish a random forest prediction model, and the accuracy is evaluated to determine the inter-class instability index threshold.
[0095] First, set the step size larger and narrow the threshold selection interval based on the prediction accuracy calculated by the selected threshold;
[0096] In this embodiment, a smaller step size is set to obtain the optimal inter-class instability index threshold;
[0097] Finally, after comparing the inter-class instability index with the threshold, the bands with inter-class instability index greater than the threshold are eliminated from the characteristic bands obtained by the genetic algorithm, and the hyperspectral characteristic band combination optimized by the inter-class instability index method is obtained;
[0098] In this embodiment, multiple models are used to combine the above-extracted feature bands, and models are established and predicted according to the feature bands extracted by genetic algorithm, feature bands extracted by correlation coefficient, feature bands extracted by inter-class instability index, and feature bands extracted by correlation coefficient + inter-class instability index.
[0099] In this embodiment, the rice smut data is represented by the extracted preferred characteristic bands. The preferred characteristic bands are used to extract the rice smut data, that is, the spectral reflectance data of the characteristic bands are used to establish a model to predict the rice smut diseased area.
[0100] Example 2
[0101] like Figure 3 As shown, the rice smut extraction method based on the selection of UAV hyperspectral data bands provided in this example belongs to the hyperspectral band optimization method, which includes the following steps:
[0102] Step 1: Preprocessing of UAV hyperspectral data
[0103] The UAV hyperspectral data has hundreds of bands, and the ground objects are imaged in a unified manner at an altitude of more than 100 meters. Sometimes, due to the fact that the signal-to-noise ratio of the instrument is not in the optimal working state, or due to the combined effect of interference factors such as dark current, there is a certain amount of noise in the spectral reflectance of different bands, resulting in the reflectance of adjacent bands showing a jagged feature. In this embodiment, the Savitzky-Golay convolution smoothing method using a moving window and a quadratic polynomial is used to smooth and denoise the hyperspectral data (see Figure 2 ).
[0104] Step 2: Normalization of UAV hyperspectral data
[0105] Normalize the hyperspectral data of each band (refer to the formula below)
[0106]
[0107] In the formula, x i is the spectral reflectance value of a certain band, x max and x min are the maximum and minimum values of a band respectively.
[0108] Step 3: Binary encode the UAV hyperspectral band; "0" indicates that the band is not selected; "1" indicates that the band is selected; this is used for logical operations in the genetic algorithm.
[0109] The UAV hyperspectral reflectance data and the smut occurrence area data are input into the population initialization module, and then step 4 is entered to perform the next calculation; the rice smut in this embodiment is rice false smut.
[0110] Step 4: Select the RMSE calculated from the prediction results of the partial least squares model as the fitness function. The greater the individual's fitness, the greater the probability that the individual will be inherited to the next generation, and vice versa;
[0111] Step 5: Set the genetic algorithm population size, genetic algorithm iteration times, mutation probability, crossover probability, etc.;
[0112] If the population size is too small, inbreeding can occur, leading to pathological genes and congenital loss of effective alleles. Even with a high mutation probability, competitive genes may still be eliminated. Furthermore, a high mutation probability can be extremely destructive to the existing population. Furthermore, random errors in genetic operators hinder the correct propagation of effective patterns in small populations, preventing population evolution from producing the expected number according to the pattern theorem. If the population size is too large, convergence is difficult, resources are wasted, and robustness is reduced. If the mutation probability is too small, population diversity decreases too quickly, leading to a rapid loss of effective genes that is difficult to repair. A high mutation probability can maintain population diversity but also increases the likelihood that optimal solutions will be eliminated. Similar to the mutation probability, a high crossover probability can easily destroy existing solutions, increase randomness, and easily miss the optimal individuals. A low crossover probability cannot effectively update the population. If the genetic algorithm iteration rate is too low, the algorithm will not converge easily. If the number of iterations is too large, the algorithm or population will mature prematurely, and further evolution will only increase time expenditure and waste resources. When the fitness of the optimal band combination no longer improves, or the number of iterations of the genetic algorithm reaches the preset number of iterations, the operation is terminated. After repeated experiments and tests, the initial population size is set to 30, the crossover probability is set to 0.5, the mutation probability is set to 0.01, and the maximum iteration is now 100.
[0113] Step 6: Use genetic algorithm to select the initial optimal characteristic band after 20 operations (see Figure 5 )
[0114] exist Figure 5 In the variable selection frequency chart, there are three horizontal lines, indicating that characteristic bands with frequencies greater than the line value were selected for modeling. The position of the horizontal line is determined based on model accuracy, and the optimal band combination with the highest accuracy is selected as the band screening result. As shown in the figure, the number of characteristic bands selected based on the top horizontal line is 8, the number of characteristic bands selected based on the middle horizontal line is 18, and the number of characteristic bands selected based on the bottom horizontal line is 42. Modeling was performed using the characteristic bands selected based on the three horizontal lines. When modeling was performed using 8 characteristic bands, the model's prediction accuracy was 76.91%; when modeling was performed using 18 characteristic bands, the model's prediction accuracy was 83.44%; and when modeling was performed using 42 characteristic bands, the model's prediction accuracy was 83.44%. The 18 characteristic bands selected based on the middle horizontal line (accounting for 6.59% of the total spectral bands) were selected for modeling analysis and subsequent band screening.
[0115] Step 7: Calculate the characteristic bands selected by the genetic algorithm to obtain the optimal judgment threshold of the selected characteristic band; specifically, the following two methods are used:
[0116] Step 71: Calculate the Pearson correlation coefficient for the characteristic bands selected by the genetic algorithm according to the following formula
[0117]
[0118] A coefficient value of 1 means that X and Y can be well described by a straight line equation, all data fall well on a straight line, and Y increases as X increases. A coefficient value of -1 means that all data points fall on a straight line, and Y decreases as X increases. A coefficient value of 0 means that there is no linear relationship between the two variables. In this method, the correlation coefficient between the bands selected by the genetic algorithm is first calculated, and a threshold value between 0.8 and 1.0 is selected as the standard for the correlation coefficient to select hyperspectral bands. Each time a threshold is selected, the bands that do not meet the threshold conditions are eliminated from the preliminary selected bands, and the selected bands are used to construct a monitoring model and calculate the prediction accuracy. Finally, the correlation coefficient threshold that gives the highest prediction accuracy is obtained.
[0119] Step 72: Calculate the inter-class instability index for the characteristic bands selected by the genetic algorithm according to the following formula
[0120]
[0121] The inter-class instability index is an important indicator to characterize the separability of each category in the band. According to the size of the inter-class instability index, it can be directly judged whether a certain band is conducive to more accurate classification of samples. z,i The smaller it is, the closer the spectral reflectance of each sample in the same category is, and the smaller the discreteness of the data is. Therefore, the smaller the intra-class deviation Δwithin,i is, the more conducive it is to sample classification; when the absolute value of the mean difference within each category |m z,i -m j,i A larger | indicates greater variability in the spectral reflectance of samples across different categories, leading to better separability of the spectral data. Therefore, a larger inter-class deviation Δbetween,i facilitates sample classification. Therefore, bands with smaller inter-class instability indices are more conducive to improved classification accuracy. Therefore, it is important to select bands with smaller inter-class instability indices for monitoring and eliminate bands with larger indices to further reduce data dimensionality and improve model efficiency and accuracy. The most important factor in selecting hyperspectral feature bands using the inter-class instability index method is the choice of threshold. To find the optimal threshold, a series of thresholds must be selected. Initially, a larger step size can be set. Based on the prediction accuracy calculated from the selected threshold, the threshold selection interval can be narrowed. Finally, a smaller step size is set to obtain the optimal threshold. This method effectively reduces the time required to find the optimal threshold. Using prediction accuracy as an evaluation metric, after successively selecting a series of thresholds within the threshold interval, the inter-class instability index is compared with the threshold to select the optimal hyperspectral band for establishing a rice false smut prediction model and calculate prediction accuracy.
[0122] Step 8: Use random forest, support vector machine, gradient boosting tree and multi-layer perceptron to build models for the selected preferred feature bands, and predict the model prediction results according to the optimal judgment threshold and the following formula (see Table 1-4 for the results)
[0123]
[0124]
[0125]
[0126]
[0127] The prediction model in this embodiment can also use random forest, support vector machine, gradient boosting tree and multi-layer perceptron to establish models according to the preferred feature bands.
[0128] Step 9: Summarize the optimal band selection method based on the prediction results and obtain the following results
[0129] (1) After the feature bands were optimized, the accuracy of the four models calculated using the optimized bands was above 80%. Among them, the gradient boosting tree model and random forest model used the final 13 optimized bands to build models, with prediction accuracies of 85.62% and 84.10% respectively. The random forest and gradient boosting tree models had the best monitoring accuracy.
[0130] (2) The sensitive bands for monitoring rice smut are between 698nm-750nm and 974nm-984nm.
[0131] (3) When using only one method to select the optimal band, the correlation coefficient method is better than the inter-class instability index method in selecting the optimal characteristic band.
[0132] Table 1: Random Forest (RF) Model
[0133]
[0134]
[0135] When all five bands were removed simultaneously, the resulting optimal feature band combination reduced the data size by 27.8% compared to the original optimal feature bands, while maintaining significant improvements in accuracy, precision, recall, and F1 score compared to both the pre-band removal and single-band removal methods. For rice smut monitoring, the goal is to accurately identify diseased rice areas, so accuracy and precision are key. Using the correlation coefficient and inter-class instability index to remove bands increased accuracy and precision by 2.22% and 1.70%, respectively.
[0136] Table 2: Support Vector Machine (SVM) Model
[0137]
[0138] Among them, the use of the correlation coefficient method to eliminate two bands can further reduce the amount of data while ensuring that the model prediction accuracy is not reduced; however, after using the inter-class instability index method to eliminate three bands, evaluation indicators such as accuracy, precision, recall, and F1 score all decrease; when five bands are eliminated at the same time, the accuracy and F1 score remain unchanged, the precision rate decreases by 1.13%, and the recall rate increases by 2.17%.
[0139] Table 3: Gradient Boosted Tree (GBDT) Model
[0140]
[0141]
[0142] The accuracy, precision, recall, F1 score and other evaluation indicators of the selected bands screened by the correlation coefficient method and the inter-class instability index method have all been improved. Among them, the accuracy rate increased by 2.4% and the recall rate increased by 1.99% when the five bands screened by the correlation coefficient and the inter-class instability index were simultaneously eliminated.
[0143] Table 4: Multilayer Perceptron (MLP) Model
[0144]
[0145] Similar to the verification results of the RF model and GBDT model, the correlation coefficient method and the inter-class instability index method were used to eliminate bands, and the evaluation indicators such as accuracy, recall rate, and F1 score were improved. However, the accuracy of the bands filtered by eliminating the correlation coefficient and the inter-class instability index alone decreased slightly.
[0146] The above embodiments are merely preferred embodiments for the purpose of fully illustrating the present invention, and the scope of protection of the present invention is not limited thereto. Equivalent substitutions or modifications made by those skilled in the art based on the present invention are within the scope of protection of the present invention. The scope of protection of the present invention shall be subject to the claims.
Claims
1. A rice smut extraction method based on hyperspectral data band selection, characterized by: The following steps are involved: Acquire hyperspectral image data and perform preprocessing; Conduct field measurements and obtain the location distribution information of healthy rice and rice with smut disease in rice areas; Comparing hyperspectral image data with field measurements of healthy and smut-affected rice locations yielded spectral reflectance data at the sampling points. Constructing a prediction model, wherein the prediction model uses partial least squares and sampling point spectral reflectance data to screen out preferred characteristic bands; Determining a characteristic band threshold, and selecting an optimal characteristic band combination from the spectral reflectance data of the sampling points according to the characteristic band threshold, wherein the optimal characteristic band combination is used as rice smut prediction data; The characteristic band threshold is obtained by calculating the correlation coefficient to obtain the correlation coefficient threshold and the characteristic band threshold is obtained by calculating the inter-class instability index to obtain the inter-class instability index threshold; The correlation coefficient is calculated according to the following steps: in, is the Pearson correlation coefficient, and are the means of variables X and Y, respectively. and are the element values of variables X and Y respectively; n is the number of elements of the variable; The correlation coefficient threshold is performed according to the following steps: The preferred hyperspectral band is obtained by comparing the correlation coefficient with the preset threshold, and the prediction accuracy of the preferred hyperspectral band is calculated using the prediction model, and the correlation coefficient threshold with the highest prediction accuracy is obtained as the correlation coefficient threshold of the monitoring model; The inter-class instability index is performed according to the following steps: In the formula, It is in The inter-class instability index of the two types of samples at the band, and are intra-class bias and inter-class bias, respectively. For the first type of samples, The standard deviation of the band, For the second type of samples, The standard deviation of the band, For the first type of samples, The mean of the bands, For the second type of samples, The mean of the bands; The inter-class instability index threshold is determined by the following steps: The threshold interval is set according to the inter-class instability index, and a certain threshold is used to select the preferred feature band within the threshold interval. The inter-class instability index threshold is obtained based on the preferred feature band and the prediction model.
2. The rice smut extraction method based on hyperspectral data band selection according to claim 1, characterized in that: The acquisition of hyperspectral image data and preprocessing are performed according to the following steps: Acquire hyperspectral images of the study area to obtain images that are hyperspectral images obtained by multiple photogrammetry measurements over a period of time; Perform filtering and normalization operations on hyperspectral image data.
3. The rice smut extraction method based on hyperspectral data band selection according to claim 1, characterized in that: The following steps are also included: Based on the binary classification, the confusion matrix and the accuracy, precision, recall and F1 scores are used to establish the evaluation criteria of the prediction model. The accuracy, precision, recall and F1 scores are calculated according to the following formula: in, Representation matrix and accuracy; Indicates the accuracy rate; represents the recall rate; represents the F1 score; TP is the number of predicted positive examples and the actual positive examples, FP is the number of actual negative examples but predicted positive examples; TN is the number of predicted negative examples and the actual negative examples, and FN is the number of actual negative examples but predicted positive examples.
4. The rice smut extraction method based on hyperspectral data band selection according to claim 1, characterized in that: The hyperspectral data normalization process is performed according to the following formula: In the formula, is the spectral reflectance value of a certain band, and are the maximum and minimum values of a band respectively.
5. A rice smut extraction system based on hyperspectral data band selection, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 4 is implemented.