Method for improving agricultural pest recognition or classification model performance based on multispectral image

By adding feature scaling-based spectral mapping module and multi-scale feature alignment channel attention mechanism in agricultural pest recognition or classification models, the feature selection and model performance improvement problems of multispectral images in agricultural pest recognition and classification are solved, and more efficient disease recognition and classification effects are achieved.

CN119992337APending Publication Date: 2025-05-13CHANGCHUN UNIV OF SCI & TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510132380.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-11-28
Filing Date
2025-02-06
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

When using multispectral images for agricultural pest identification and classification, there are problems such as subjectivity and limitations in feature selection, strict requirements for input data format and number of channels in the deep learning model, and limitations in classic channel attention methods when processing multispectral images.

Method used

Add a feature scaling-based spectral mapping module to agricultural pest identification or classification models, use the GDAL remote sensing image processing library to read multi-spectral images, and optimize spectral feature representation through feature scaling processing. Combined with deep learning models, critical quantitative analysis of spectral bands is performed, and the importance of each band is evaluated through channel stacking and culling methods. Using the channel attention mechanism of multi-scale feature alignment, we focus on rich feature information on multi-dimensional spectral bands, and integrate complementary information of different spectral bands.

Benefits of technology

Optimizing spectral feature representations through the spectral mapping module, the deep learning model shows good performance in the disease level classification task. The channel attention mechanism significantly improves the generalization performance of the model and can more effectively utilize the feature information of multispectral images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992337A_ABST
    Figure CN119992337A_ABST
Patent Text Reader

Abstract

The invention relates to a method for improving agricultural pest recognition or classification model performance based on a multispectral image, and relates to the technical field of agricultural pest model performance improvement. The invention discloses a method for improving the performance of an agricultural pest recognition or classification model based on a multispectral image. A spectrum mapping module is added in an agricultural disease and insect pest identification or classification model, a spectrum key quantitative analysis method based on a deep learning model is adopted, and the utilization rate of the model for each spectrum band characteristic is improved by using a channel attention mechanism, so that the model shows good performance in disease grade classification; in addition, an effective technical support is provided for intelligent disease monitoring and management in the agricultural field, practical application of multispectral image processing is promoted, and an important foundation is laid for development of a disease detection technology in the future.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of agricultural pest and disease model performance improvement, and in particular to a method for improving the performance of agricultural pest and disease recognition or classification models based on multispectral images. Background Art

[0002] With the rapid development of artificial intelligence technology, disease detection technology based on machine learning has made significant progress (Das S, Pattanayak S, Behera P R. Application of machine learning: a recent advancement in plant diseases detection [J]. Journal of Plant Protection Research, 2022: 122-135-122-135.). In 2018, Hossain proposed a method for diagnosing diseases from tea pictures based on a support vector machine (SVM) classifier with an accuracy rate of over 90%, achieving accurate and automated tea disease detection (Hossain, Selim, et al. "Recognition and detection of tea leaf's diseases using support vector machine." 2018 IEEE 14th International Colloquium on Signal Processing & Its Applications (CSPA). IEEE, 2018.). Devi et al. proposed a peanut disease detection method that combines Harris corner detection, HOG feature extraction, and KNN classifier. The performance of the model was greatly improved by optimizing feature extraction, with an accuracy of 97.67% (Devi KS, Srinivasan P, Bandhopadhyay S. H2K–A robust and optimum approach for detection and classification of groundnut leaf diseases [J]. Computers and Electronics in Agriculture, 2020, 178: 105749.). However, the performance of these early machine learning methods was greatly affected by the subjectivity and limitations of feature selection, and it was difficult to extract the best and stable features from a large number of features, resulting in poor generalization performance of the model.

[0003] With the emergence of deep learning methods, deep learning technologies such as Convolutional Neural Networks (CNN) have been introduced into agricultural disease detection. CNN has significantly reduced agricultural manual intervention with its automatic feature extraction and strong generalization capabilities, and has greatly improved the accuracy and efficiency of disease detection (Latif, Ghazanfar, et al. "Deep learning utilization in agriculture: Detection of rice plant diseases using an improved CNN model." Plants 11.17 (2022): 2230.). For example, Mohanty et al. proposed a disease detection method based on crop images captured by smartphones, which uses portable mobile phone devices for real-time detection of crop diseases. This innovation not only improves the efficiency of disease detection, but also effectively reduces the need for manual intervention (Mohanty SP, Hughes DP, Salathé M. Using deep learning for image-based plant disease detection [J]. Frontiers in plant science, 2016, 7: 1419.). Agarwal et al. developed a CNN-based leaf disease detection method that achieved an average accuracy of 91.2% in the classification of 9 diseases and 1 healthy category, further verifying the effectiveness of deep learning methods in improving the accuracy of disease recognition (Agarwal, Mohit, et al. "ToLeD: Tomato leaf disease detection using convolution neural network." Procedia Computer Science 167 (2020): 293-301.). These methods all use deep learning to extract features from RGB images, successfully decouple artificial features, and enhance generalization performance. However, the information dimension of RGB images is limited, and it is difficult to fully capture the subtle features of diseases in different spectral bands.In particular, when the image quality degrades (such as blur, increased noise) or the lighting conditions change, the performance of CNN-based deep learning models is often significantly affected (Hu, Chengsong, et al. "Influence of image quality and light consistency on the performance of convolutional neural networks for weed mapping." Remote Sensing 13.11 (2021): 2140.).

[0004] Compared with traditional RGB images, the use of multispectral images for automated disease recognition and detection in agriculture has many advantages (Singh V, Sharma N, Singh SA review of imaging techniques for plant disease detection [J]. Artificial Intelligence in Agriculture, 2020, 4: 229-242.). Multispectral crop images captured by multispectral cameras not only have richer spectral features, but also more stable image quality. In 2018, Duarte-Carvajalino et al. used the principal component analysis (PCA) method to reconstruct multispectral images and achieved high detection accuracy in potato disease detection, confirming the effectiveness of multispectral images in the field of agricultural disease detection (Duarte-Carvajalino, Julio M., et al. "Evaluating late blight severity in potato crops using unmanned aerial vehicles and machine learning algorithms." Remote Sensing 10.10 (2018): 1513.). Alnaggar et al. studied the effect of early detection of rice blast based on RGB images and multispectral images including red, green and near-infrared bands combined with deep learning methods. The experimental results showed that compared with using only RGB images, the combination of RGB images and multispectral images can significantly improve the F1 accuracy of rice blast detection, indicating the importance of multispectral images in improving disease detection performance (Alnaggar YA, Sebaq A, Amer K, et al. Rice plant disease detection and diagnosis using deep convolutional neural networks and multispectral imaging [C] / / International Conference on Model and Data Engineering. Cham: Springer Nature Switzerland, 2022: 16-25.). Although multispectral images have obvious advantages in disease detection, deep learning models have strict requirements on the format and number of channels of input data.Existing convolutional neural networks (CNNs) are mainly based on processing three-channel RGB images, while multispectral images usually contain three to ten bands, which makes it impossible for existing models to effectively model the rich information in each band. In addition, studies have shown that different crops, varieties and diseases show diversity in spectral characteristics, which leads to different sensitivities of band reflectance to disease severity (Feng, Wei, et al. "Canopy vegetation indices from in situ hyperspectral data to assess plant water status of winter wheat under powdery mildew stress." Frontiers in Plant Science 8 (2017): 1219.). How to effectively select and use representative spectral bands to effectively improve the performance of the model still needs further study.

[0005] In order to further improve the expressiveness of image features, existing studies have introduced attention mechanisms to improve the model. This mechanism significantly enhances the model's ability to identify defects by selectively emphasizing important features and suppressing irrelevant features. For example, Bastidas et al. proposed a deep learning model for multispectral images, Channel Attention Networks (CAN). By applying a soft attention mechanism on each channel, CAN significantly improved the semantic segmentation performance of multispectral images (Bastidas AA, Tang H. Channel attention networks [C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition workshops. 2019: 0-0.). Fan et al. proposed a parallel network model CMAP-Net based on the channel attention mechanism, which greatly improved the classification performance of the deep learning model when processing multispectral images (FanX, Li X, Yan C, et al. Converging Channel Attention Mechanisms with Multilayer Perceptron Parallel Networks for Land Cover Classification[J]. Remote Sensing, 2023, 15(16): 3924.). Although improving the network by introducing the channel attention mechanism can effectively improve the performance of the model, this targeted composite network structure is relatively complex, and there is still a lack of lightweight channel attention modules that can effectively process multispectral images.

[0006] In summary, the main defects of existing disease detection technology are as follows:

[0007] (1) When using machine learning methods to process multispectral images, the manual feature engineering required is rather cumbersome, and the performance of the machine learning model directly depends on the manual feature extraction process. Due to the limitations of human intervention, the generalization performance of existing machine learning methods for intelligent automatic identification of pests and diseases is poor.

[0008] (2) Deep learning methods are used to extract features from RGB images. Due to the limited information dimension of RGB images, it is difficult to fully capture the subtle features of the disease in different spectral bands. Although multispectral images have the characteristics of stable image quality and rich spectral data, deep learning models have strict requirements on the format and number of channels of input data. Existing convolutional neural networks (CNNs) mainly process three-channel RGB images, while multispectral images usually contain three to ten bands. When faced with hyperspectral images with too many bands, due to the limitations of convolutional neural networks, the excessive channel dimensions of image data will greatly affect the performance of the model, resulting in the inability of existing models to effectively model the rich information in each band. In addition, there is currently a lack of quantitative analysis methods for the importance of spectral bands in the context of rice blast, and there is insufficient understanding of the key bands in the task of rice blast severity classification.

[0009] (3) The existing classic channel attention method has certain limitations when dealing with multispectral images. The channel identifiers generated at a single angle are easily affected by noise information when capturing features, and the fully connected layer is used to model the relationship between channels, which is relatively simple and may not be able to capture the relationship between complex spectral bands, thus affecting the model performance. Summary of the invention

[0010] The present invention aims to solve the technical problems in the prior art and provide a method for improving the performance of agricultural pest and disease recognition or classification models based on multispectral images.

[0011] In order to solve the above technical problems, the technical solutions of the present invention are as follows:

[0012] A method for improving the performance of agricultural pest identification or classification models based on multispectral images, comprising the following steps:

[0013] Step 1, adding a spectral mapping module based on feature scaling to the agricultural pest identification or classification model;

[0014] The spectral mapping module uses the GDAL remote sensing image processing library to read multispectral images in TIF format in a lossless manner, and then performs mapping transformation based on the spectral reflectance of feature scaling. Through normalized feature scaling processing, the spectral feature representation of the multispectral image is optimized;

[0015] Step 2: Quantitative analysis of spectral band criticality based on deep learning model;

[0016] Combined with the deep learning model, the importance of each spectral band optimized in step 1 in the disease grade classification task is quantitatively evaluated to identify the spectral bands that have a significant impact on the model performance;

[0017] Step 3: Use the channel attention mechanism to improve the model's utilization of the features of each spectral band;

[0018] The channel attention method with multi-scale feature alignment focuses on the rich spectral feature information on multi-dimensional spectral bands and integrates the complementary information on different spectral bands to achieve more accurate attention score allocation, thereby improving the generalization performance of the model.

[0019] In the above technical solution, in step 2, the channel stacking and channel elimination methods are used to evaluate the importance of each spectral band in the disease grade classification task.

[0020] In the above technical solution, further, step 2 includes the following steps:

[0021] First, the spectral bands of the multispectral image are extracted. For the five spectral bands obtained by extraction, these five bands are stacked into three-channel images in sequence in channel stacking, and the input of RGB images is simulated. Feature extraction and classification performance evaluation are performed through a deep learning model, so as to obtain the influence and contribution of each band on the classification performance of the model. In channel elimination, one band is specified to be eliminated at a time for the five extracted bands, aiming to evaluate the relationship and influence between the band and the remaining bands. These two spectral key experiments are used to quantitatively evaluate the contribution and relationship of different spectral bands in the disease grade classification task.

[0022] In the above technical solution, step 3 specifically includes the following steps:

[0023] Step 3.1: Perform global average pooling and maximum pooling operations on the input multispectral image features to integrate spatial information and obtain two different channel identification vectors C M and C A ,This channel identification vector represents the initial identification vector of each spatial information that has not been calibrated;

[0024] C M =MaxPool(X) (1)

[0025] C A =AvgPool(X) (2)

[0026] X represents the original multispectral image features, MaxPool(X) represents average pooling of X, and AvgPool(X) represents maximum pooling of X;

[0027] Step 3.2, use a fully connected layer as the mapping matrix W E , the two channel identification vectors C in step 3.1 M and C AMap them to the same high-dimensional space, so that two different channel identification vectors generate two matrices Q and K after the same mapping;

[0028] Q=MLP(C M ) (3)

[0029] K=MLP(C A ) (4)

[0030] MLP(C M ) and MLP(C A ) represent the channel identification vector C M and C A Mapping is performed;

[0031] Step 3.3, transpose and multiply the two new matrices Q and K obtained after mapping in step 3.2, and calculate a matrix ACM with the number of channels in rows and columns, which is used to represent the similarity and weight distribution between each channel;

[0032] ACM=Q·K T =MLP(C M )·(MLP(C A )) T (5)

[0033] T stands for transpose;

[0034] Step 3.4, scale the matrix ACM obtained in step 3.3, and then transpose the scaled matrix to obtain ACM′;

[0035]

[0036] Among them, E dim is the dimension of the mapping layer, T represents transpose, and ACM′ represents the attention coefficient matrix;

[0037] Step 3.5, the attention coefficient matrix ACM′ is mapped to a one-dimensional feature vector through the fully connected layer, and activated by the Sigmoid activation function, and then multiplied by the original multispectral image feature X to obtain a new feature vector OutPut Feature after channel attention calibration;

[0038] OutPut Feature=X·Sigmoid(MLP(ACM′)) (7).

[0039] In the above technical solution, the channel attention method of multi-scale feature alignment in step 3 is replaced by SeNet, ECANet, and CBAM channel attention methods.

[0040] The beneficial effects of the present invention are:

[0041] The method of the present invention for improving the performance of agricultural pest and disease identification or classification model based on multispectral images, by adding a spectral mapping module to the agricultural pest and disease identification or classification model, adopts the GDAL remote sensing image processing library, and reads the multispectral image in TIF format in a lossless manner. This module optimizes the representation of spectral feature information through spectral reflectance mapping based on feature scaling, which not only ensures the integrity of the data, but also solves the problem of feature degradation that may be encountered in the model during training. Accurate spectral representation becomes the basis for all subsequent analysis and feature extraction, and is the key to the model's effective processing of multispectral data. This module enables all deep learning models to process multispectral images in TIF format with the help of this method, bringing extensive progress to the application of multispectral images in agriculture.

[0042] The method of improving the performance of agricultural pest identification or classification models based on multispectral images of the present invention adopts a spectral critical quantitative analysis method based on a deep learning model. Through this method, we combine the deep learning model to evaluate the importance of each spectral band in the task of disease grade classification to identify the bands that have a significant impact on model performance. This innovative method can provide a deeper understanding of the spectral features related to disease grade classification, thereby enhancing the understanding of model optimization and data features.

[0043] Furthermore, the method of the present invention uses channel stacking and channel elimination methods to explore the nature of the effect of each band on the model, which can quantify the specific contribution of each spectral band to the deep learning model in automated disease classification, thereby deeply understanding the importance of these bands and their interrelationships. In addition, this method provides an important reference for model optimization and helps to improve classification performance. By evaluating the changes in model performance after eliminating specific bands, we can further clarify the redundancy of each band and its unique value. This mechanism not only enhances the understanding of spectral features, but also lays a solid foundation for subsequent research and applications.

[0044] The method of the present invention for improving the performance of agricultural pest and disease recognition or classification models based on multispectral images uses a channel attention mechanism to improve the model's utilization of the features of each spectral band. Aiming at the problems existing in the classic channel attention mechanism, a multi-scale feature aligned channel attention method, SpectralAttention, is proposed. This method can effectively focus on the rich feature information on multi-dimensional spectral bands, and integrate complementary information on different spectral bands to achieve more accurate attention score allocation, effectively improve the generalization performance of the model, and improve the flexibility and efficiency of the deep learning model in processing multispectral images.

[0045] The method of the present invention for improving the performance of agricultural pest and disease recognition or classification models based on multispectral images, by adding a spectral mapping module to the agricultural pest and disease recognition or classification model, adopting a spectral critical quantitative analysis method based on a deep learning model and utilizing a channel attention mechanism to improve the model's utilization of the features of each spectral band, enables the model to exhibit good performance in disease grade classification, and also provides an effective technical support for intelligent disease monitoring and management in the agricultural field, promotes the practical application of multispectral image processing, and lays an important foundation for the future development of disease detection technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] The present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0047] Figure 1 This is a technical roadmap for the method of the present invention for improving the performance of agricultural pest and disease identification or classification models based on multispectral images.

[0048] Figure 2 Schematic diagram of the channel stacking process of the present invention.

[0049] Figure 3 Schematic diagram of the channel rejection process of the present invention.

[0050] Figure 4 Schematic diagram of the channel attention method process for multi-scale feature alignment of the present invention.

[0051] Figure 5 Loss curves for Swin Transformer and Vision Transformer before and after adding the spectral mapping module.

[0052] Figure 6 This is a bar chart of the results after each model channel is eliminated.

[0053] Figure 7 It is a line graph of the experimental results after single channel stacking.

[0054] Figure 8 Schematic diagram of the process of generating an attention matrix by mapping through two different fully connected layers to generate the same identifier using the same pooling operation.

[0055] Fig. 9 Schematic diagram of the process of generating an attention matrix by mapping through two different fully connected layers to generate different identifiers using different pooling operations.

[0056] Fig.10 This is the framework diagram of Swin Transformer.

[0057] Fig.11This is the framework diagram of Vision Transformer. DETAILED DESCRIPTION

[0058] The inventive concept of the present invention is: the core of the present invention is to optimize the performance of agricultural pest identification or classification models, especially the accuracy and efficiency of rice blast disease grade classification. Compared with traditional RGB images, multispectral images have the advantages of rich spectral information, wide spectral range, and high image quality. However, this characteristic of multispectral images also brings challenges: the existing deep learning model has strict requirements on the number of channels of the image, and when the number of channels increases, more noise may also be introduced, affecting the performance of the model. For example, traditional multispectral image processing methods mostly use reconstruction or dimensionality reduction to pre-process the bands in the multispectral image, screen out the key bands, and then reconstruct them into three-channel RGB images or fuse them into 3-channel images. At present, there is a lack of an effective method to effectively model multispectral data while retaining rich spectral information and wide spectral bands. In view of the limitations of deep learning methods in processing multispectral images, the present invention proposes a spectral mapping module based on feature scaling, which can optimize the representation of spectral reflectance data and avoid the problem of feature degradation caused by uneven distribution of spectral reflectance. In addition, when facing high-spectral images with up to 150 or even thousands of bands, the choice of whether to use one-dimensional convolution for dimensionality reduction specified in the module can effectively reduce the dimensionality of multi-band feature information, thereby maximizing the performance of the deep learning model. Existing spectral criticality analysis methods under disease backgrounds mostly use PCA methods or mathematical analysis methods to extract key bands. However, deep learning models are automated for image feature extraction, which results in the inability of mathematical analysis methods to effectively adapt to the feature extraction and data mining capabilities of different models, resulting in additional computational overhead and affecting the performance of the model. Existing methods lack quantitative analysis methods for spectral importance, and lack the ability to adaptively extract key features. The present invention combines deep learning models and disease multispectral images to propose a spectral criticality quantitative analysis method under a specified disease background. At the same time, the spectral importance quantitative analysis method of channel stacking and channel elimination can deeply understand the spectral features under the specified disease background and optimize the classification effect. This method can effectively reflect the nature of the role of different spectral bands on different models, thereby helping to deeply understand the spectral features and optimize the classification performance. The traditional channel attention method SeNet only uses the global average pooling operation to generate a spatial identifier. This approach is relatively simple and one-sided, and cannot fully represent the information of the features at each spatial position. The CBAM method, on the other hand, passes the two spatial identifiers obtained by the global average pooling operation and the maximum pooling operation through the same Sequential layer, which contains three fully connected layers. This undoubtedly increases the complexity of the calculation. In addition, the CBAM method sums the two new spatial identifiers obtained after three fully connected layers, which may destroy the relationship between the spatial identifier vectors themselves.Most importantly, the present invention proposes a lightweight channel attention method for effectively processing multispectral images. The method is based on a multi-scale feature alignment strategy, generates channel identifiers at two different angles, significantly enhances the model's ability to capture fine features, can effectively integrate complementary information from different spectral bands, and adaptively enhances the ability to mine features on key spectral bands, thereby improving the overall performance of the model. Through comparative experiments with classic channel attention methods (such as Squeeze-and-Excitation Networks, SeNet) and (Convolutional Block Attention Module, CBAM), the effectiveness of the channel attention mechanism based on multi-scale feature fusion is verified. Specifically, the present invention introduces global average pooling and maximum pooling operations to generate two different spatial identifiers, and the dimensional values ​​of these two identifiers represent spatial information on different channels. During neural network training, the weights on the corresponding channels of these two identifiers should be similar and show consistency. Therefore, we calculate the similarity of the two spatial identifiers through a mapping matrix and inner product operation, and reduce the obtained attention coefficient matrix to a one-dimensional feature vector through a fully connected layer, so that the model can help us find a suitable weight coefficient vector and multiply it with the original feature to achieve the calibration function of the channel coefficient. This design can capture the spatial relationship between features more comprehensively, is more suitable for the complex characteristics of multispectral data, and improves the accuracy and effectiveness of channel attention calculation.

[0059] The technical roadmap of the method of improving the performance of agricultural pest identification or classification model based on multispectral images of the present invention can be found in Figure 1, this method significantly improves the classification accuracy and effect through three innovative improvements. First, by adding a spectral mapping module based on feature scaling to the existing agricultural pest and disease identification or classification model, the module uses the GDAL remote sensing image processing library to read multispectral images in TIF format in a lossless manner. This module optimizes the representation of spectral feature information through spectral reflectance mapping based on feature scaling, which not only ensures the integrity of the data, but also solves the problem of feature degradation that may be encountered in the model training process. Accurate spectral representation becomes the basis for all subsequent analysis and feature extraction, and is the key to the effective processing of multispectral data by the model. This module enables all deep learning models to process multispectral images in TIF format with the help of this method, which brings extensive progress to the application of multispectral images in agriculture. Secondly, due to the different sensitivities of different spectral bands to disease characteristics, in order to enhance the understanding of the key bands in the disease grade classification task of rice blast, the present invention proposes a spectral criticality quantitative analysis method based on a deep learning model. Through this method, we combine the deep learning model to evaluate the importance of each spectral band in the rice blast disease grade classification task to identify the bands that have a significant impact on the model performance. This innovative method can provide a deeper understanding of the spectral features related to the classification of rice blast grades, thereby enhancing the understanding of model optimization and data features. Finally, the present invention proposes to use the channel attention mechanism to improve the model's utilization of the features of each spectral band. In view of the problems existing in the classical channel attention mechanism, a multi-scale feature alignment channel attention method, SpectralAttention, is proposed. This method can effectively focus on the rich feature information on multi-dimensional spectral bands, and integrate the complementary information on different spectral bands, so as to achieve more accurate attention score allocation, effectively improve the generalization performance of the model, and improve the flexibility and efficiency of the deep learning model in processing multispectral images. In summary, the present invention proposes a spectral mapping module based on feature scaling, which ensures the effective extraction of multi-dimensional features through normalized feature scaling processing, solves the problem of information loss in feature extraction, and effectively improves the model's ability to utilize multispectral data and overall performance. Aiming at the problem of insufficient quantitative analysis methods for spectral importance under specific disease backgrounds, the present invention uses channel stacking and elimination methods to verify the correlation of spectral bands in the classification of rice blast disease grades, deeply understands spectral features under specified disease backgrounds, and optimizes classification effects. This paper proposes a multi-scale feature alignment channel attention method, which can effectively focus on the rich feature information on multi-dimensional spectral bands and integrate the complementary information on different spectral bands to achieve more accurate attention score allocation, effectively improve the generalization performance of the model, and improve the flexibility and efficiency of deep learning models in processing multispectral images. This innovative method provides stronger feature extraction capabilities for multispectral image analysis and promotes research progress in related fields.

[0060] The present invention is described in detail below with reference to the accompanying drawings.

[0061] A method for improving the performance of agricultural pest identification or classification models based on multispectral images, comprising the following steps:

[0062] Step 1, adding a spectral mapping module based on feature scaling to the agricultural pest identification or classification model;

[0063] The spectral mapping module uses the GDAL remote sensing image processing library to read multispectral images in TIF format in a lossless manner, and then performs mapping transformation based on the spectral reflectance of feature scaling. Through normalized feature scaling processing, the spectral feature representation of the multispectral image is optimized;

[0064] Step 2: Quantitative analysis of spectral band criticality based on deep learning model;

[0065] Combined with the deep learning model, the importance of each spectral band optimized in step 1 in the disease grade classification task is quantitatively evaluated to identify the spectral bands that have a significant impact on the model performance;

[0066] Step 3: Use the channel attention mechanism to improve the model's utilization of the features of each spectral band;

[0067] The channel attention method with multi-scale feature alignment focuses on the rich spectral feature information on multi-dimensional spectral bands and integrates the complementary information on different spectral bands to achieve more accurate attention score allocation, thereby improving the generalization performance of the model.

[0068] The various steps of the method of the present invention are described in more detail below.

[0069] First, step 1 is introduced in more detail. In view of the limitations of deep learning models in processing multispectral images in TIF format, the present invention proposes a spectral mapping module based on feature scaling. Although multispectral images contain richer spectral bands and spectral information, the data characteristics limit their storage and processing methods. The traditional method of converting multispectral images into RGB images and then storing them in JPEG or PNG format is convenient, but there is a problem of data loss, and the true physical meaning corresponding to each pixel value in the multispectral image cannot be retained. The pixel value in the multispectral image represents the true spectral reflectance of the material or substance at the corresponding position. At the same time, multispectral images usually contain 3 to 10 bands, which has more channels than three-channel RGB images, thus bringing additional challenges to the image reading and processing of deep learning models.

[0070] To solve this problem, the method of the present invention adds a spectral mapping module based on feature scaling to the agricultural pest identification or classification model, and uses the GDAL remote sensing image processing library to read multispectral images in TIF format in a lossless manner. This module optimizes the representation of spectral feature information through spectral reflectance mapping based on feature scaling, which not only ensures the integrity of the data, but also solves the feature degradation problem that the model may encounter during training. Accurate spectral representation becomes the basis for all subsequent analysis and feature extraction, and is the key to the model's effective processing of multispectral data. This module enables all deep learning models to process multispectral images in TIF format with the help of this method, bringing extensive progress to the application of multispectral images in agriculture.

[0071] Secondly, step 2 is introduced in more detail. Since different spectral bands have different sensitivities to disease characteristics, in order to enhance the understanding of key bands in disease (such as rice blast) grade classification tasks, the present invention proposes a spectral criticality quantitative analysis method based on a deep learning model. Through this method, we combine the deep learning model to evaluate the importance of each spectral band in the disease (such as rice blast) grade classification task to identify the bands that have a significant impact on model performance. This innovative method can provide a deeper understanding of the spectral features related to rice blast grade classification, thereby enhancing the understanding of model optimization and data characteristics.

[0072] Specifically, we use channel stacking and channel elimination methods to explore the nature of the effect of each band on the model. This method first extracts each band of the multispectral image through band extraction, and then stacks it into a three-channel image to simulate the input of the RGB image. Then, a deep learning model is used to achieve comprehensive feature extraction and classification performance evaluation ( Figure 2 ). Then, we will conduct a band elimination experiment, by eliminating the specified spectral bands, and analyze the impact of eliminating specific bands on the classification results ( Figure 3 ), and then quantitatively evaluate the contribution and mutual relationship of different bands in the task of disease (such as rice blast) grade classification. In this specific implementation: first, extract each spectral band of the multispectral image; for the five spectral bands obtained by extraction, stack these five bands in sequence into a three-channel image in channel stacking, simulate the input of RGB image, and perform feature extraction and classification performance evaluation through a deep learning model, so as to obtain the influence and contribution of each band on the classification performance of the model; in channel elimination, for the five extracted bands, specify to eliminate one band at a time, in order to evaluate the mutual relationship and influence between this band and the remaining bands; through these two spectral key experiments, quantitatively evaluate the contribution and mutual relationship of different spectral bands in the task of disease grade classification.

[0073] Through this method, we can quantify the specific contribution of each spectral band to the deep learning model in automated disease classification, so as to gain a deeper understanding of the importance of these bands and their interrelationships. In addition, the method proposed in this invention provides an important reference for model optimization and helps to improve classification performance. By evaluating the changes in model performance after removing specific bands, we can further clarify the redundancy of each band and its unique value. This mechanism not only enhances the understanding of spectral features, but also lays a solid foundation for subsequent research and applications.

[0074] Finally, step 3 is introduced in more detail. The present invention proposes to use the channel attention mechanism to improve the model's utilization of the features of each spectral band. In view of the problems existing in the classic channel attention mechanism, a multi-scale feature aligned channel attention method, SpectralAttention (hereinafter referred to as SpectralAtt), is proposed. This method can effectively focus on the rich feature information on multi-dimensional spectral bands, and integrate the complementary information on different spectral bands to achieve more accurate attention score allocation, effectively improve the generalization performance of the model, and improve the flexibility and efficiency of the deep learning model in processing multi-spectral images. Figure 4 This is the specific implementation process of SpectralAtt.

[0075] 1. For the input multispectral image features, global average pooling and maximum pooling operations are performed to integrate spatial information and obtain two different channel identification vectors C M and C A This channel identification vector represents the initial identification vector of each spatial information that has not been calibrated.

[0076] C M =MaxPool(X) (1)

[0077] C A =AvgPool(X) (2)

[0078] X represents the original multispectral image features, MaxPool(X) represents average pooling of X, and AvgPool(X) represents maximum pooling of X;

[0079] AvgPool (average pooling) and MaxPool (maximum pooling) are commonly used pooling operations in convolutional neural networks (CNNs). They are mainly used to reduce the dimension of feature maps and improve the computational efficiency of the model while retaining important feature information. The average pooling operation reduces the size of the feature map by calculating the average value of all pixels in a specified area (such as 2x2 or 3x3). This operation can remove noise from the feature map and retain global valid information as much as possible, but this may lead to the loss of key features. MaxPool selects the maximum value as the output in the specified area, thereby reducing the size of the feature map. This operation can effectively retain important features in the image, especially edges and corners. The advantage is that it can retain important features, but the disadvantage is that it will ignore global information.

[0080] 2. Use a fully connected layer as the mapping matrix W E This is to map the two spatial description vectors into the same high-dimensional space, so that two different spatial description vectors (i.e., channel identification vectors) can generate two matrices Q and K after the same mapping, which can better align the spatial description relationship and effectively reduce the computational complexity.

[0081] Q=MLP(C M ) (3)

[0082] K=MLP(C A ) (4)

[0083] MLP(C M ) and MLP(C A ) represent the channel identification vector C M and C A Mapping is performed;

[0084] 3. Transpose and multiply the two new matrices Q and K obtained after mapping to calculate a matrix ACM with the number of channels in rows and columns, which is used to represent the similarity and weight distribution between the channels.

[0085] ACM=Q·K T =MLP(C M )·(MLP(C A )) T (5)

[0086] T stands for transpose;

[0087] 4. Scale the resulting ACM matrix to make it satisfy a distribution with a variance of 1, which helps to keep the attention coefficient stable. Then we need to perform a transposition to ensure that the independence of the row vectors is not destroyed during the dimensionality reduction process.

[0088]

[0089] Among them, E dim is the dimension of the mapping layer, T represents transpose, and ACM′ represents the attention coefficient matrix;

[0090] 5. The attention coefficient matrix ACM′ is mapped to a one-dimensional feature vector through a fully connected layer to find a suitable channel weight vector, and activated by a Sigmoid activation function. This effectively reduces the computational complexity while introducing nonlinearity, allowing the model to calibrate the channel weight more flexibly, and then multiply it with the original multispectral image feature X to obtain a new feature vector after channel attention calibration. This design can fully capture the spatial relationship between features, is more suitable for the complex characteristics of multispectral data, and significantly improves the accuracy and effectiveness of channel attention calculation.

[0091] OutPut Feature=X·SigMoid(MLP(ACM′)) (7)

[0092] By introducing the channel attention mechanism, the weight distribution between different channels can be dynamically learned, thereby enhancing the representation ability of features. This method can help the network better focus on and utilize important information in the input features, improving the performance and generalization ability of the model.

[0093] The experimental results of using the above-mentioned method of improving the performance of agricultural pest identification or classification model based on multispectral images of the present invention to improve the performance of existing models are as follows:

[0094] 1. Spectral mapping module ablation experiment

[0095] like Figure 5 As shown in the Loss curve graphs in Vision Transformer and Swin Transformer, we can see that after adding the spectral mapping module, the model effectively avoids the problem of feature degradation.

[0096] like Fig.10 This is the framework diagram of Swin Transformer. Before the image is input into the deep learning network, we will use the spectral mapping module to convert it into a tensor that the deep learning network can normally handle for subsequent feature extraction and processing. Then we will add a spectral attention module at the beginning of the architecture to calibrate the channel weights of the original image data, and then add spectral attention modules after the Swin Transformer Block.

[0097] In Vision Transformer, its framework is as follows Fig.11 As shown, we also pass the Spectral Mapping Module before the image is input into the deep learning network, and then pass it through the Spectral Attenion module before it is divided into patches by the convolution kernel.

[0098] 2. Experimental results of quantitative analysis of channel criticality

[0099] Figure 6 The bar chart of the experimental results of using multispectral images after channel elimination as training data on different models shows that after removing the red band, the performance and various evaluation indicators on all models have dropped significantly, while the performance will be slightly improved when the red edge band is removed. This is because the wavelength of the red edge band is included in the wavelength of the red band, so when removing this band, the noise feature is removed, while the effective feature is retained on the red band. Therefore, we can roughly believe that the red band has the worst correlation and the red edge band has the strongest correlation. In addition, the other three bands show different correlations on different models, so it is impossible to find a unified correlation ranking.

[0100] In the rice blast disease grade classification task, the red band data was stacked three times to form a three-channel image and used for model training. It was found that good performance was achieved on four different models (see Figure 7 ). However, when the red band data is used alone for training, its performance even exceeds the results of the experiment after stacking other channels. This shows that among the five channels, the red band data plays an irreplaceable role in the task, and the information contained in it cannot be replaced by the other four channels combined. On the contrary, removing the blue band or the near-infrared band is equivalent to removing some sub-components. In this case, the neural network can make up for this loss by performing repeated abstraction and feature extraction in the remaining channels. Therefore, even if the blue band or the near-infrared band is removed, the impact on the model performance will not be too great.

[0101] In summary, the information contained in the red band plays a unique and important role in the rice blast disease grade classification task, while the information in other channels can be complemented and features extracted through neural network learning, and has relatively little impact on model performance.

[0102] 3. Experimental results analysis after the improvement of channel attention mechanism

[0103] Table 1 Experimental results of adding different channel attention mechanisms for comparative experiments

[0104]

[0105]

[0106] As shown in Table 1, we can see that the special channel attention mechanism proposed in the present invention obtains different effects when added to different models. At the same time, in order to verify the effectiveness of the present invention, the present invention is compared with other classic attention methods. The experimental results are shown in the above table. On the four models of VGG, ResNet, EfficientNet and VisionTransformer, the best effect is achieved after adding the SpectralAttention module proposed in the present invention. Compared with the SeNet method, the CBAM method performs better on the EfficientNet and Swin Transformer models, but both cannot achieve the effect of SpectralAttention.

[0107] 4. SpectralAttention module comparison experiment

[0108] In order to systematically evaluate and verify the innovation and superiority of the proposed SpectralAttention module, this paper conducted ablation experiments on the internal structure of the SpectralAtt module based on the multispectral rice blast dataset. These experiments aim to deeply analyze the specific impact of different pooling strategies and fully connected layer configurations on model performance. Specifically, we constructed two schemes to explore the sensitivity and effectiveness of the internal mechanism of SpectralAttention. In the first scheme (such as Figure 8 As shown in , we replace the two differential pooling operations in SpectralAttention with a single, identical pooling operation to generate repeated spatial identifiers. These identical identifiers are then fed into two independent fully connected layers for mapping to simulate the generation process of the attention matrix. This setting aims to analyze the impact of the diversity of pooling operations on the model's feature extraction capabilities. The second solution (as shown in Fig. 9 ) further explored the flexibility of the fully connected layer architecture by retaining the different pooling operations in the original SpectralAttention to generate unique spatial identifiers, but then sending each identifier to its own independent fully connected layer for processing. This setting aims to evaluate the importance of the independent mapping ability of the fully connected layer for the construction of the final attention matrix and its efficiency in synergy with the pooling operation.

[0109] In order to further explore the design rationality of the SpectralAttention module and its superiority over other attention mechanisms, we conducted an ablation experiment (Table 2). The ablation experiment results show that different pooling operations are used inside the SpectralAtt module to generate two independent spatial identifiers. The attention matrix constructed by the spatial identifiers mapped by the same fully connected layer achieves the best effect. This not only ensures the computational complexity and efficiency, but also reflects the advantages of the mapping layer. That is, the attention coefficient matrix we expect to obtain must be able to align the weights in the same dimension in the same mapping layer. The use of two different pooling operations avoids the errors caused by more noise features or fewer key features in a certain spectral band. Calculating the correct attention coefficient matrix through the difference between two different spatial identifiers will undoubtedly achieve better results.

[0110] Table 2 Performance results of ablation experiments on various models

[0111]

[0112]

[0113] In summary, the method of the present invention, by integrating the spectral mapping module, the spectral critical quantitative analysis method and the channel attention mechanism, enables the model to not only show good performance in the grade classification of rice blast, but also provides an effective technical support for intelligent disease monitoring and management in the agricultural field, promotes the practical application of multispectral image processing, and lays an important foundation for the future development of disease detection technology.

[0114] The method of the present invention, the spectral mapping module based on feature scaling is not limited to processing multispectral image data with three to ten channels, but can also process hyperspectral image data with more than 10 channels. Existing alternatives usually use dimensionality reduction or reconstruction to convert multispectral images in TIF format into three-channel RGB images, while the spectral mapping module can perform band dimensionality reduction through one-dimensional convolution operation, thereby automating the processing of multispectral and hyperspectral images, thereby exerting the potential of multispectral and hyperspectral images.

[0115] The method of the present invention uses mathematical analysis methods such as spectral analysis or principal component analysis, but does not rely on quantitative analysis methods of deep learning models. There are certain limitations. The results obtained by classical mathematical analysis methods may not be effectively adapted to deep learning models, thereby limiting the model performance.

[0116] In addition to the channel attention method proposed in the present invention, channel attention methods such as SeNet, ECANet, and CBAM can also perform weight adaptive calibration on image feature channels for deep learning models to achieve the purpose of realizing the channel attention mechanism.

[0117] Obviously, the above embodiments are merely examples for the purpose of clear explanation, and are not intended to limit the implementation methods. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation methods here. The obvious changes or modifications derived therefrom are still within the scope of protection of the invention.

Claims

1. A method for improving the performance of agricultural pest recognition or classification models based on multispectral images, characterized in that: The following steps are involved: Step 1, adding a spectral mapping module based on feature scaling to the agricultural pest identification or classification model; The spectral mapping module uses the GDAL remote sensing image processing library to read multispectral images in TIF format in a lossless manner, and then performs mapping transformation based on the spectral reflectance of feature scaling. Through normalized feature scaling processing, the spectral feature representation of the multispectral image is optimized; Step 2: Quantitative analysis of spectral band criticality based on deep learning model; Combined with the deep learning model, the importance of each spectral band optimized in step 1 in the disease grade classification task is quantitatively evaluated to identify the spectral bands that have a significant impact on the model performance; Step 3: Use the channel attention mechanism to improve the model's utilization of the features of each spectral band; The channel attention method with multi-scale feature alignment focuses on the rich spectral feature information on multi-dimensional spectral bands and integrates the complementary information on different spectral bands to achieve more accurate attention score allocation, thereby improving the generalization performance of the model.

2. The method according to claim 1, characterized in that In step 2, channel stacking and channel elimination methods are used to evaluate the importance of each spectral band in the disease grade classification task.

3. The method according to claim 2, characterized in that Step 2 includes the following steps: First, the spectral bands of the multispectral image are extracted. For the five spectral bands obtained by extraction, these five bands are stacked into three-channel images in sequence in channel stacking, and the input of RGB images is simulated. Feature extraction and classification performance evaluation are performed through a deep learning model, so as to obtain the influence and contribution of each band on the classification performance of the model. In channel elimination, one band is specified to be eliminated at a time for the five extracted bands, aiming to evaluate the relationship and influence between the band and the remaining bands. These two spectral key experiments are used to quantitatively evaluate the contribution and relationship of different spectral bands in the disease grade classification task.

4. The method according to claim 1, characterized in that: Step 3 specifically includes the following steps: Step 3.1: Perform global average pooling and maximum pooling operations on the input multispectral image features to integrate spatial information and obtain two different channel identification vectors C M and C A ,This channel identification vector represents the initial identification vector of each spatial information that has not been calibrated; C M =MaxPool(X) (1) C A =AvgPool(X) (2) X represents the original multispectral image features, MaxPool(X) represents average pooling of X, and AvgPool(X) represents maximum pooling of X; Step 3.2, use a fully connected layer as the mapping matrix W E , the two channel identification vectors C in step 3.1 M and C A Map them to the same high-dimensional space, so that two different channel identification vectors generate two matrices Q and K after the same mapping; Q=MLP(C M ) (3) K=MLP(C A ) (4) MLP(C M ) and MLP(C A ) represent the channel identification vector C M and C A Mapping is performed; MLP represents the fully connected layer used. The same fully connected layer is used here to map the channel identification vector through the same neural network layer to confirm that it is correct. Step 3.3, transpose and multiply the two new matrices Q and K obtained after mapping in step 3.2, and calculate a matrix ACM with the number of channels in rows and columns, which is used to represent the similarity and weight distribution between each channel; ACM=Q·K T =MLP(C M )·(MLP(C A )) T (5) T stands for transpose; Step 3.4, scale the matrix ACM obtained in step 3.3, and then transpose the scaled matrix to obtain ACM′; Among them, E dim is the dimension of the mapping layer, T represents transpose, and ACM′ represents the attention coefficient matrix; Step 3.5, the attention coefficient matrix ACM′ is mapped to a one-dimensional feature vector through the fully connected layer, and activated by the Sigmoid activation function, and then multiplied by the original multispectral image feature X to obtain a new feature vector OutPut Feature after channel attention calibration; OutPut Feature=X·SigMoid(MLP(ACM′)) (7).

5. The method according to claim 1, characterized in that The channel attention method of multi-scale feature alignment in step 3 is replaced by SeNet, ECANet, and CBAM channel attention methods.

Citation Information

Cited By

  • Disease and insect disease image multispectral imaging and deep learning intelligent detection system

    CN120451798A