Wheat scab spore recognition method based on yolov5-eca-asff
By introducing ECA and ASFF modules into the YOLOv5 network, the problem of identifying wheat scab spores in complex backgrounds was solved, and efficient and accurate detection of wheat scab spores was achieved.
Patent Information
- Application Number
- CN202310444199.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-24
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2043-04-24
AI Technical Summary
Existing technologies struggle to efficiently identify wheat scab spores in complex environments, especially in situations with multiple targets and complex backgrounds, where detection accuracy and speed are limited.
A method based on Yolov5-ECA-ASFF is adopted. By adding an ECA module to the end of the residual block of the backbone network CSPNet and combining it with the ASFF module to perform feature fusion in the Neck layer, a wheat scab spore identification model is constructed. The ECA module enhances channel features and the ASFF module performs adaptive feature fusion to make full use of features at different scales.
It improves the accuracy, recall, and average precision of identifying wheat scab spores, achieving rapid and accurate detection that is superior to traditional methods.
Smart Images

Figure CN116524255B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of wheat scab spore identification technology, specifically a wheat scab spore identification method based on Yolov5-ECA-ASFF. Background Technology
[0002] Spore image recognition is the application of image recognition algorithms in the detection of fungal spores in agricultural pests and diseases. Its main purpose is to accurately locate target pathogenic spores in images. With the rapid development of image processing technology and artificial intelligence algorithms, spore image recognition has become a new research hotspot and has achieved certain results. Spore image recognition technology can be divided into machine learning and deep learning methods based on different algorithm principles. In the process of disease control, a large number of samples are needed to train the classifier to achieve the classification and judgment of spore types.
[0003] In machine learning research on spore recognition, Yang et al. proposed a distance-transform Gaussian filtering algorithm based on a dataset of microscopic images of rice spores. This algorithm, using a combination of texture and shape features and a decision tree fusion matrix method, separated adherent spores and then selected four shape features (area, perimeter, ellipticity, complexity) and three texture features (entropy, homogeneity, contrast) for classification using a decision tree model, achieving a detection accuracy of up to 94%. However, due to the limited number of extracted features and the singular target, this method lacks practical applicability. Wang et al., based on a dataset of 600 images of *Botrytis cinerea*, *Effectivesia cocovenenans*, and *Effectivesia xanthogenae* spores, extracted 90 features using image preprocessing methods such as mean filtering, Gaussian filtering, Otsu's method binarization, morphological operations, and masking operations. They established an SVM (Support Vectorachine) classification model, achieving an average accuracy of 91.68%. However, this method did not perform dimensionality reduction optimization on a large number of features, resulting in relatively low detection accuracy and limited detection speed. Based on a dataset of diffraction fingerprint images of potato diseases, Wang et al. used 13 selected diffraction fingerprint features in an SVM classification model. The test results showed that the support vector machine model had an average accuracy of 92.72% in identifying three fungal spores in greenhouse crops. The accuracy of this method was improved after feature optimization. However, the accuracy of this method was not verified when other microorganisms were present in the air.
[0004] While traditional machine learning has achieved some success in spore detection, it is only suitable for situations where the target is simple and has obvious features and a simple background. However, for recognition tasks with diverse targets and complex and variable backgrounds, it is difficult to use machine learning to extract surface features to achieve good detection performance. Therefore, deep learning is needed to extract a large number of features from complex targets through convolutional structures to complete the recognition task.
[0005] With the improvement of computer hardware performance, deep learning has developed rapidly due to its advantages of low cost and high efficiency, and various neural networks have been applied to the detection of small targets in microscopic images. For example, Jubayer et al. used a pre-trained Yolov5 algorithm to detect mold on food surface microscopic images, achieving an accuracy of 98.10%. However, it was not tested in complex backgrounds, and the network's feature extraction capability for small targets was not verified. Wang et al. proposed a lightweight Yolov4 network based on the ECA fusion attention mechanism, which efficiently implemented local cross-channel interaction using one-dimensional convolution to extract the dependencies between channels, improving the recognition of small targets. However, its detection performance varied greatly in different scenarios. Qiu et al. proposed a Yolov5 network combined with adaptive feature fusion ASFF to improve feature scale invariance and enhance the detection effect of small targets in road traffic detection, addressing the problems of dense, multi-scale, and small targets. However, it still had false detections when the target was occluded. To address the challenges of detecting very small targets and uncertain features in UAVs, Dadboud et al. employed YOLOv5's mosaic data augmentation technique, achieving relatively good detection performance. However, its robustness is poor, and its detection effect is poor in complex scenarios.
[0006] Deep learning, compared to machine learning, possesses strong learning capabilities, requires less feature engineering, and is highly portable. However, for small object detection, it is characterized by limited available features and high requirements for localization accuracy, and most network architectures are not optimized for small object detection.
[0007] However, traditional research methods mostly involve pre-segmenting the target region, further extracting multi-dimensional features, and selecting the best features. These selected features are then combined with various classifiers such as SVM to achieve target detection. Deep learning, through its complex convolutional neural network structure, adaptively extracts the optimal features for different targets without requiring redundant feature extraction or region pre-segmentation. It can simultaneously perform detection and classification tasks and can detect microscopic images with more complex backgrounds with high accuracy. However, deep learning has high requirements for datasets, requiring a large number of samples to prevent poor fitting and low accuracy. Furthermore, the complexity of the structure and the computational cost of convolutions mean that network training requires significant computing power. Existing work often only makes single structural modifications to improve the network's feature extraction or multi-dimensional semantic information extraction capabilities. While this improves detection performance, the network's detection capabilities in various scenarios remain poor.
[0008] Therefore, how to achieve effective and rapid detection and identification of wheat scab spores has become an urgent technical problem to be solved. Summary of the Invention
[0009] The purpose of this invention is to address the shortcomings of existing technologies in detecting and identifying small targets such as wheat scab spores, and to provide a method for identifying wheat scab spores based on Yolov5-ECA-ASFF to solve the above problems.
[0010] To achieve the above objectives, the technical solution of the present invention is as follows:
[0011] A method for identifying wheat scab spores based on Yolov5-ECA-ASFF includes the following steps:
[0012] Establishment of spore image dataset: Collect and gather images of Fusarium head blight spores, preprocess and augment the images to construct a spore image dataset;
[0013] Constructing a wheat Fusarium head blight spore recognition model: Based on the Yolov5 network, a wheat Fusarium head blight spore recognition model was constructed by integrating the ECA and ASFF modules;
[0014] Training of the wheat Fusarium head blight spore recognition model: Input the spore image dataset into the wheat Fusarium head blight spore recognition model for training;
[0015] Obtaining wheat Fusarium head blight spores to be identified: Obtain wheat Fusarium head blight spores to be identified and pre-treat them;
[0016] Obtaining the identification results of wheat Fusarium head blight spores to be identified: The pre-processed wheat Fusarium head blight spores to be identified are input into the trained wheat Fusarium head blight spore identification model to obtain the wheat Fusarium head blight spore identification results.
[0017] The establishment of the spore image dataset includes the following steps:
[0018] A total of 10,000 images were selected from the collected images, including images of spores of four types of wheat scab and images of spores of four types of scab fungi mixed with spores of four other fungi, as the target recognition dataset.
[0019] The dataset was expanded from 10,000 images to 20,000 images through data augmentation;
[0020] The augmented image data was manually labeled using Labelme software. After labeling, the dataset with JSON file labels was converted into the label file type required by YOLOv5 to construct the spore image dataset.
[0021] The construction of the wheat Fusarium head blight spore recognition model includes the following steps:
[0022] A wheat Fusarium head blight spore recognition model was constructed based on YOLOv5, and the model was configured with four layers:
[0023] The first layer is Input, which is used as the input end for adaptive scaling of the image. It integrates the K-means genetic algorithm to adaptively calculate the optimal anchor box value of the dataset, thereby enhancing the ability to detect small objects.
[0024] The second layer is the Backbone. The backbone structure of the Backbone network includes the Focus, CSP module and SPP module. The Focus integrates the width and height information of the input image in the channel and further performs feature mapping on the base layer. The CSP structure is divided into two parts and fused using a cross-stage hierarchical structure. The SPP module at the end is used for three structures: convolutional layer, pooling layer and selection filter.
[0025] The third layer is the Neck, which includes an upsampling structure of FPN+PAN. The FPN layer is combined with the PAN module at the end for upsampling.
[0026] The fourth layer is the Output layer. The Output layer is used to evaluate whether the algorithm is accurate in detecting and locating targets. It uses the GIOU loss function to calculate the accuracy. The GIOU loss function removes the remaining prediction results after dividing the GIOU_Loss contribution value by the maximum value and outputs the result with the highest classification probability. At the same time, it generates bounding boxes and predicts the type of targets within the bounding boxes.
[0027] Configure the ECA attention module:
[0028] The ECA module replaces the fully connected layers of the original structure with a fast one-dimensional convolution method to obtain non-linear feature information generated across channels.
[0029] A one-dimensional convolution is defined with a kernel size of k, which represents the coverage range of cross-channel information, meaning that the current channel and its k neighboring channels participate in the prediction of channel attention.
[0030] There is a mapping relationship between k and the dimension C of the total channels, and when the dimension C of the total channels is determined, a one-dimensional convolution kernel k is obtained by adaptive calculation.
[0031] Projecting the two-dimensional convolution kernel onto the spatial domain yields a nonlinear function relating to the distribution of image feature points. This function can be used to quickly determine the optimal solution, and its mapping relationship is linear, as shown in the formula:
[0032] Φ(C) = ak - b,
[0033] However, its linear mapping relationship is too simplistic, and the number of channels in convolutional networks is generally set to a power of 2. Therefore, the linear function is extended to nonlinear functions, and its calculation formula is as follows:
[0034] C=Φ(K=2(ak-b)
[0035] From the above, specifying the number of channels C, we obtain the formula:
[0036]
[0037] In the above formula: x odd Let b = 1 and a = 2 be the odd numbers closest to x.
[0038] The second-layer backbone of the wheat scab spore recognition model was improved by adding an ECA module to the end of its CSP residual block;
[0039] Configure the ASFF feature fusion module:
[0040] The ASFF feature fusion module maps and fuses scale features at various levels, adaptively learns spatial weight parameters, obtains a new weighting method, and achieves the fusion of features from different layers.
[0041] In the three-layer structure of ASFF, Level 1, Level 2, and Level 3 are the three feature layers of the PANet module. After feature fusion of the three layers, the results of the three feature layers Level 1, Level 2, and Level 3 are output by X(1), X(2), and X(3) to perform convolution calculation:
[0042] Multiply X(1), X(2), and X(3) by the weight parameters α(3), β(3), and γ(3) respectively, and sum them to obtain the feature output ASFF3 after feature fusion. The formula for this process is as follows:
[0043]
[0044] in, This represents the new feature map obtained through ASFF. These represent the weight parameters of the three feature layers, which are then satisfied using the Softmax function. These represent the features of layers 1, 2, and 3, respectively.
[0045] Upsampling or downsampling methods are used to ensure that the feature structure of the output of each level (Level 1, Level 2, Level 3) is the same, while keeping the number of channels constant.
[0046] In ASFF3 at the bottom, the feature information of Level 1 and Level 2 is first compressed to the same number of channels as Level 3 through 1×1 convolution. Then, it is upsampled at four times and two times to obtain the same dimension as Level 3. Finally, the accumulation operation is performed.
[0047] The output of the second-layer Neck layer FPN+PAN structure of the wheat scab spore identification model is combined with the ASFF feature fusion module to map and fuse the scale features of each layer, and adaptively learn the spatial weight parameters to obtain a new weighting method, thus realizing the fusion of features from different layers.
[0048] The training of the wheat Fusarium head blight spore recognition model includes the following steps:
[0049] Set up a PyTorch neural network training environment with Python 3.8 and CUDA 11.6.
[0050] The image input size is set to 640×640, the confidence threshold is set to 0.5, the initial learning rate is set to 0.001, the weight decay coefficient is set to 0.0005, the model training batch size is set to 16, and the training iteration period is set to 100 epochs.
[0051] Input the spore image dataset into the wheat scab spore recognition model, complete the training, and generate a weight file;
[0052] The first layer of the input terminal stitches the data using Mosaic and random scaling, cropping, and arrangement methods, and then scales the image data to a standard size of 640×640 before sending it into the detection network.
[0053] The second backbone layer first uses its focus to slice the image, sending the W and H feature information into the channel space; the CSP module expands the input channels and performs convolution calculations to extract features; the ECA module calculates the number of channels C from the CSP features using a non-linear function mapping, from which a new convolution kernel k is obtained, achieving adaptive weight adjustment of the CSP features. This results in a double-downsampled feature map without information loss; finally, the SPP module uses pyramid pooling to transform the multi-size feature map output by ECA into a fixed-size feature map and feature vector required by the Neck layer.
[0054] The third layer, Neck, first upsamples the images at resolutions of 76×76, 38×38, and 19×19 from high to low using an FPN structure to obtain semantic information of the target spores. Then, it downsamples the images at resolutions of 19×19, 38×38, and 76×76 from low to high using a PAN structure to obtain coordinate information of the target spores. Finally, the ASFF module compresses the feature information of 38×38 and 76×76 resolutions into the same number of channels as 19×19 through convolution calculation, and then upsamples it to make the outputs of the three layers at the same dimension. The obtained weight parameters are used as feature coefficients for output, realizing the fusion of high and low resolutions and making full use of feature information at various scales.
[0055] The fourth layer Output uses GIOU_Loss as the bounding box loss function and calculates the accuracy. After every five training iterations, the accuracy is calculated and a weight file is generated. This process is repeated for one hundred iterations. Afterward, the average accuracy and various evaluation metrics are generated, and the optimal weight file is retained in all weight files for Fusarium head blight spore detection and identification.
[0056] The generated weight file enables rapid and accurate detection and identification of Fusarium head blight spores.
[0057] Beneficial effects
[0058] The present invention provides a method for identifying wheat scab spores based on YOLOv5-ECA-ASFF. Compared with existing technologies, this method adds an ECA spatial attention mechanism to the end of the CSPNet residual block of the YOLOv5s backbone network to enhance the channel features of the input feature map. Furthermore, it introduces an ASFF module with an adaptive feature fusion mechanism at the end of its Neck feature extraction network. This fully utilizes the high-level information and low-level features of the input image to construct a YOLOv5-ECA-ASFF network, effectively achieving rapid and accurate detection and identification of wheat scab spores.
[0059] This invention selected the classic single- and double-stage target recognition networks Faster-RCNN, Yolov4, and Yolov5 for comparative experiments on a wheat scab microspore dataset. The results show that the network model proposed in this invention, while maintaining a relatively constant detection speed, achieves accuracy (P), recall (R), mean average precision (mAP), and F1-score of 98.57%, 96.5%, 98.4%, and 97.4% respectively for identifying Fusarium graminearum, the causal agent of wheat scab. These overall evaluation parameters are higher than those of the mainstream target recognition networks. Attached Figure Description
[0060] Figure 1This is a sequence diagram of the method of the present invention;
[0061] Figure 2 This is a diagram of the culture medium for wheat scab spores in the existing technology;
[0062] Figure 3 Image of wheat scab spores in existing technology;
[0063] Figure 4 This is an enhanced image of wheat scab spore data related to the present invention;
[0064] Figure 5 This is a diagram of the Yolov5s network structure involved in this invention;
[0065] Figure 6 This is a structural diagram of the ECA attention mechanism involved in this invention;
[0066] Figure 7 This is a diagram of the ASFF feature fusion module involved in this invention;
[0067] Figure 8 This is a diagram of the Yolov5-ECA-ASFF network structure involved in this invention;
[0068] Figure 9 The detection graph is from the original Yolov5 network;
[0069] Figure 10 This is a detection graph for the two-stage Faster-RCNN network. Detailed Implementation
[0070] To provide a better understanding of the structural features and effects achieved by the present invention, a detailed description is provided below, accompanied by preferred embodiments and accompanying drawings:
[0071] like Figure 1 As shown, the present invention provides a method for identifying wheat scab spores based on Yolov5-ECA-ASFF, comprising the following steps:
[0072] The first step is to establish a spore image dataset: collect and gather images of Fusarium wilt spores, preprocess and augment the images to construct a spore image dataset.
[0073] In practical applications, the first step is to culture the spores of the wheat scab fungus: the inoculum is placed on a solid culture medium and then placed in a constant temperature incubator. After a certain growth period, spore suspensions and solid culture media containing spores are prepared for taking microscopic images of the fungus. (See figure) Figure 2After screening a total of 46 bacterial colonies, eight fungal spores were selected for imaging: four Fusarium species that mainly cause Fusarium head blight in wheat, and four common fungal spores. These four Fusarium species are Fusarium graminearum, Fusarium moniliforme, and Fusarium truncatum, which cause Fusarium head blight in wheat. Additionally, common field spores such as *Anthracnose spores*, *Gynostemma pentaphyllum*, and *Trichophyton mentagrophytes*, which share similarities in morphology, color, and size with Fusarium head blight spores, as well as dust particles, were introduced to increase the complexity of the dataset. Next, the microscopic images of Fusarium head blight spores were acquired: a diluted spore suspension was aspirated using a narrow-mouthed dropper and dropped onto the center of a glass slide. A new dropper tip was used to gently scrape solid culture medium and mix it into the liquid on the slide, resulting in 500 slides containing various spores. The prepared slides were placed on the microscope stage. Twenty randomly selected images from each slide, each containing multiple fungal spores, were photographed using the microscope. First, a low-power lens was used to locate the suitable target area. Then, a high-power lens (10×40) was used to focus and expose the target area. Once the system was stable and the images were clear, microscopic images of the wheat scab fungal spores were acquired using the image acquisition system. (See attached image.) Figure 3 This experiment collected three types of spore microscopic images: pure Fusarium head blight spores, mixed fungal spores, and spores under different lighting conditions. A total of 10,010 spore microscopic images were collected. The Fusarium head blight spore dataset was evaluated and verified by plant protection experts to ensure its authenticity and validity. Most images contain multiple types of spores and have high background complexity to ensure the proposed method has high generalization performance and robustness. The microscope used was a Zeiss Axio Vert. A1 laboratory inverted microscope for data acquisition.
[0074] The establishment of the spore image dataset includes the following steps:
[0075] (1) Select 10,000 images from the collected images, including spore images of 4 types of wheat scab and mixed spore images of 4 types of scab fungi, as the target recognition dataset.
[0076] (2) The dataset was expanded from 10,000 images to 20,000 images through data augmentation. The augmentation effect is shown in [link to data]. Figure 4 .
[0077] (3) The augmented image data was manually labeled using Labelme software. After labeling, the dataset with JSON file labels was converted into the label file type required by YOLOv5 to construct the spore image dataset.
[0078] The second step is to construct a wheat Fusarium head blight spore recognition model: Based on the Yolov5 network, the ECA and ASFF modules are integrated to construct a wheat Fusarium head blight spore recognition model. The specific steps are as follows:
[0079] (1) A wheat scab spore recognition model was constructed based on YOLOv5. The structure of the YOLOv5 model is shown below. Figure 5 The wheat scab spore recognition model is set up with four layers:
[0080] The first layer is the Input layer. The Input layer can adaptively scale the image and integrates the K-means genetic algorithm to adaptively calculate the optimal anchor box value of the dataset, thereby enhancing the ability to detect small objects.
[0081] The second layer is the Backbone. The Backbone network backbone consists of Focus, CSP, ECA, and SPP modules. Focus integrates information such as the width and height of the input image into the channels. When no information is lost, its computational cost is only 0.44 times that of a regular convolution. The ECA structure is shown below. Figure 6 Further, by performing feature mapping on the base layer, the CSP structure is divided into two parts, which are then fused using a cross-stage hierarchical structure to address the issue of recurring gradient information after optimization of the convolutional neural network backbone. The ECA module calculates the number of channels C for the CSP features using a non-linear function mapping, from which a new convolutional kernel k is obtained, achieving adaptive weight adjustment of the CSP features. This results in a double-downsampled feature map without information loss. The final SPP module determines whether the convolutional layer, pooling layer, and selection filter need to be applied, transforming the multi-size feature map output by ECA into a fixed-size feature map and feature vector required by the Neck layer, enhancing the robustness of the network structure in terms of spatial layout and detecting diverse target shapes.
[0082] The third layer is the Neck, which consists of an upsampling / downsampling structure of FPN, PAN, and ASFF. The FPN structure upsamples images at resolutions from high to low (76×76, 38×38, 19×19) to obtain semantic information of the target spores. Then, the PAN structure downsamples images at resolutions from low to high (19×19, 38×38, 76×76) to obtain coordinate information of the target spores. Finally, the ASFF module compresses the feature information from 38×38 and 76×76 resolutions into the same number of channels as 19×19 through convolution, and then upsamples it to ensure the outputs of all three layers are on the same dimension. The resulting weight parameters are used as feature coefficients to achieve the fusion of high and low resolutions. The ASFF structure is shown below. Figure 7 It makes full use of feature information at various scales, and each backbone layer integrates the features extracted by different detection layers to obtain more effective information to be transmitted to the prediction layer.
[0083] The fourth layer is Output. Output is used to evaluate whether the algorithm is accurate in detecting and locating targets. It uses the GIOU loss function to calculate the accuracy. The GIOU loss function removes the remaining prediction results after dividing the GIOU_Loss contribution value by the maximum value, and outputs the result with the highest classification probability. At the same time, it generates bounding boxes and predicts the type of targets within the bounding boxes.
[0084] (2) Configure the ECA attention module:
[0085] The ECA module replaces the fully connected layers of the original structure with a fast one-dimensional convolution method to obtain non-linear feature information generated across channels.
[0086] A one-dimensional convolution is defined with a kernel size of k, which represents the coverage range of cross-channel information, meaning that the current channel and its k neighboring channels participate in the prediction of channel attention.
[0087] There is a mapping relationship between k and the dimension C of the total channels, and when the dimension C of the total channels is determined, a one-dimensional convolution kernel k is obtained by adaptive calculation.
[0088] Projecting the two-dimensional convolution kernel onto the spatial domain yields a nonlinear function relating to the distribution of image feature points. This function can be used to quickly determine the optimal solution, and its mapping relationship is linear, as shown in the formula:
[0089] Φ(C) = ak - b,
[0090] However, its linear mapping relationship is too simplistic, and the number of channels in convolutional networks is generally set to a power of 2. Therefore, the linear function is extended to nonlinear functions, and its calculation formula is as follows:
[0091] C=Φ(K=2 (ak-b)
[0092] From the above, specifying the number of channels C, we can obtain the formula:
[0093]
[0094] In the above formula: x odd Let x be the odd number closest to x, b = 1, a = 2.
[0095] (3) Improve the second layer Backbone of the wheat scab spore recognition model by adding an ECA module to the end of its CSP residual block.
[0096] The second layer of the wheat scab spore identification model is improved. Its Focus, CSP and SPP structures perform simple cross-layer fusion for feature extraction of small targets, without paying attention to detailed features. Therefore, this invention adds an ECA module to the end of the CSPNet residual block of the backbone network to enhance the channel features of the input feature map and perform adaptive weight adjustment for detailed features, thereby improving the network's feature extraction capability for different channels.
[0097] (4) Configure the ASFF feature fusion module:
[0098] The ASFF feature fusion module maps and fuses scale features at various levels, adaptively learns spatial weight parameters, obtains a new weighting method, and achieves the fusion of features from different layers.
[0099] In the three-layer structure of ASFF, Level 1, Level 2, and Level 3 are the three feature layers of the PANet module. After feature fusion of the three layers, the results of the three feature layers Level 1, Level 2, and Level 3 are output by X(1), X(2), and X(3) to perform convolution calculation:
[0100] Multiply X(1), X(2), and X(3) by the weight parameters α(3), β(3), and γ(3) respectively, and sum them to obtain the feature output ASFF3 after feature fusion. The formula for this process is as follows:
[0101]
[0102] in, This represents the new feature map obtained through ASFF. These represent the weight parameters of the three feature layers, which are then satisfied using the Softmax function. These represent the features of layers 1, 2, and 3, respectively.
[0103] Upsampling or downsampling methods are used to ensure that the feature structure of the output of each level (Level 1, Level 2, Level 3) is the same, while keeping the number of channels constant.
[0104] In ASFF3 at the bottom, the feature information of Level 1 and Level 2 is first compressed to make the number of channels the same as Level 3 through 1×1 convolution. Then, it is upsampled at four times and two times to obtain the same dimension as Level 3. Finally, the accumulation operation is performed.
[0105] (5) The output of the second layer Neck layer FPN+PAN structure of the wheat scab spore identification model is combined with the ASFF feature fusion module to map and fuse the scale features of each layer, and adaptively learn the spatial weight parameters to obtain a new weighting method, thereby realizing the fusion of features of different layers.
[0106] This paper improves upon the YOLOv5s Neck layer, which uses an FPN+PAN structure to output multi-layer features for small targets. However, this method simply converts feature maps to the same size and then accumulates them, failing to fully utilize the characteristics at various scales. Therefore, this invention combines an ASFF module at the output of the FPN+PAN structure in the Neck layer. By mapping and fusing the scale features of each level and adaptively learning the spatial weight parameters, a new weighting method is obtained, achieving the fusion of features from different layers. The new network structure is described in [link to new network structure]. Figure 8 This allows the network to fully utilize both high-level information and low-level features of the image.
[0107] The third step is training the wheat Fusarium head blight spore recognition model: The spore image dataset is input into the wheat Fusarium head blight spore recognition model for training. The training of the wheat Fusarium head blight spore recognition model includes the following steps:
[0108] (1) Set up a PyTorch neural network training environment with Python=3.8 and CUDA=11.6.
[0109] (2) Set the image input size to 640×640, the confidence threshold to 0.5, the initial learning rate to 0.001, the weight decay coefficient to 0.0005, the model training batch size to 16, and the training iteration period to 100.
[0110] (3) Input the spore image dataset into the wheat scab spore recognition model, complete the training and generate the optimal weight file.
[0111] A1) The first layer Input terminal stitches the data using Mosaic and random scaling, random cropping, and random arrangement methods, then scales the image data to a standard size of 640×640 before sending it into the detection network;
[0112] A2) The second backbone layer first uses its focus to slice the image, sending the W and H feature information into the channel space; the CSP module expands the input channels and performs convolution calculations to extract features; the ECA module calculates the number of channels C from the CSP features using a non-linear function mapping, from which a new convolution kernel k is obtained, achieving adaptive weight adjustment of the CSP features. This results in a double-downsampled feature map without information loss; finally, the SPP module uses pyramid pooling to transform the multi-size feature map output by ECA into a fixed-size feature map and feature vector required by the Neck layer.
[0113] A3) The third layer, Neck, first upsamples the images at resolutions of 76×76, 38×38, and 19×19 from high to low using the FPN structure to obtain the semantic information of the target spores. Then, it downsamples the images at resolutions of 19×19, 38×38, and 76×76 from low to high using the PAN structure to obtain the coordinate information of the target spores. Finally, the ASFF module compresses the feature information of 38×38 and 76×76 resolutions into the same number of channels as 19×19 through convolution calculation, and then upsamples it to make the outputs of the three layers at the same dimension. The obtained weight parameters are used as feature coefficients for output, realizing the fusion of high and low resolutions and making full use of feature information at various scales.
[0114] A4) The fourth layer Output uses GIOU_Loss as the bounding box loss function and calculates the accuracy. After every five training iterations, the accuracy is calculated and a weight file is generated. Repeat the above steps. After one hundred iterations of training, the average accuracy and various evaluation metrics are generated, and the best weight file is retained in all weight files for Fusarium head blight spore detection and identification.
[0115] (4) Use the generated weight file to achieve rapid and accurate detection and identification of Fusarium head blight spores.
[0116] Step 4: Obtaining wheat scab spores to be identified: Obtain wheat scab spores to be identified and pre-treat them.
[0117] Step 5: Obtaining the identification results of wheat Fusarium head blight spores to be identified: Input the pre-processed wheat Fusarium head blight spores to be identified into the trained wheat Fusarium head blight spore identification model to obtain the wheat Fusarium head blight spore identification results.
[0118] To verify the accuracy of the detection and identification of wheat scab spores in this invention, the obtained weight file was used to test 2001 complex wheat scab spore images in the test set.
[0119] This invention uses precision (P), F1-score, recall (R), and mAP50 (mean Average Precision) as evaluation metrics for each model. P and R represent the accuracy of the detection algorithm in positive and correct samples, respectively. However, P and R cannot directly assess detection accuracy. Therefore, mAP50 and the F1 index are used to evaluate the detection algorithm's capabilities. Precision refers to the probability that a detected positive sample is actually a positive sample, while recall refers to the probability that an actual positive sample is found. mAP50 is the average AP value for all categories of detection results when the predicted bounding box and the ground truth bounding box intersect, with an IoU (Intersection Over Union) threshold of 0.5. Higher mAP50 and F1 index indicate higher network accuracy. The six evaluation metrics range from 0 to 1 (smaller values indicate poorer segmentation). The calculation methods for the six metrics are as follows:
[0120]
[0121]
[0122]
[0123]
[0124]
[0125]
[0126] Where A and B are the sizes of the predicted bounding box and the ground truth bounding box, respectively. TP (True Positives) refers to the number of objects actually detected in the dataset. FP (False Positives) refers to the number of falsely detected objects by the detection model. FN (False Negatives) refers to the number of objects missed by the detection model.
[0127] Meanwhile, to verify the effectiveness of the detection algorithm proposed in this invention, this study added two feature fusion mechanisms, ASFF and BiFPN, to the Yolov5s backbone. The comparative experimental results are shown in Table 1.
[0128] Table 1 Comparative Experimental Results of Different Feature Fusion Mechanism Modules
[0129]
[0130]
[0131] The results show that both feature fusion mechanisms improved the detection results. ASFF adaptively learns the spatial weights for fusion of three layers of feature maps after the same scaling operation, which is superior to traditional cascaded multi-level feature fusion methods, with its mAP improving by 2.8% compared to Yolov5s. BiFPN adopts a weighted bidirectional feature pyramid approach, first achieving efficient bidirectional cross-scale connections and then weighted feature map fusion, thus improving its mAP by 2.3% compared to Yolov5s. However, due to the bidirectional nature of BiFPN and the fact that it only removes one node compared to PANet, BiFPN has a higher computational cost, resulting in a 6-frame decrease in detection speed compared to ASFF. Therefore, the comparative experimental results show that the network model combining the ASFF attention fusion mechanism is the best in all performance aspects.
[0132] Then, ablation experiments were conducted using the channel attention mechanism ECA, the hybrid domain attention mechanism CA, and the CBAM. Comparative experiments were performed, building and comparing Yolov5-CA-ASFF, Yolov5-ECA-ASFF, and Yolov5-CBAM-ASFF, with results shown in Table 2. The results show that all three attention mechanisms improved the feature detection capability of fungal spores, but with increased computational load, longer processing time, and a decrease in detection rate. In terms of accuracy, ECA and CBAM differed by only one percentage point. However, because the ECA module avoids some cross-channel interaction strategies, it uses an adaptive one-dimensional convolution kernel method to determine the coverage of some cross-channel interactions, resulting in a significantly better detection rate than CBAM. Therefore, the ECA attention mechanism introduced in this paper is more suitable for improving the detection capability of wheat scab spore insertion networks.
[0133] Table 2. Comparative experimental results of different attention mechanism modules
[0134]
[0135] After the two comparative experiments mentioned above, the network model YOLOv5-ECA-ASFF for fungal spore detection was obtained. However, the robustness of the network has not yet been verified. Therefore, to further verify the performance of YOLOv5-ECA-ASFF in environments with multiple fungi and dust particle interference, a dataset was selected that incorporated mixed spores with certain similarities in morphology, color, and size to wheat scab spores. The pure scab spore dataset was designated as dataset A, and the mixed spore dataset as dataset B. Robustness tests and performance comparisons of the network were conducted using these two datasets. The improved network structure was compared with YOLOv4 and YOLOv5s on the mixed datasets, and the detection results are shown in Table 3. The detection comparison results for normal lighting data and low lighting data are shown in Table 3. Figure 9 and Figure 10 Even with high dataset complexity, the improved network still shows significant improvement over other methods across various detection metrics, and retains a certain advantage in spore detection performance.
[0136] Table 3 shows the comparative experimental results of Yolov4, Yolov5, and Yolov5-ECA-ASFF in two different scenarios: dataset A (pure Fusarium head blight spores) and dataset B (mixed spores).
[0137]
[0138] Because wheat scab spores are small targets, this invention proposes a Yolov5-ECA-ASFF network structure to improve performance in detecting small targets. To evaluate the performance of the detection model among mainstream networks, this section compares the proposed network structure with single-order YOLO series networks and the representative two-order Faster-RCNN object detection network under the same model training environment and parameter configuration. Comparative experiments are conducted using four network models—Yolov4, Yolov5s, Yolov5-ECA-ASFF, and Faster-RCNN—on the same dataset. The metrics include precision, recall, AP, F1 score, weight parameters, and detection rate. The improved Yolov5-ECA-ASFF network is primarily compared with Yolov4 and Yolov5s. The detection results for wheat scab spores using these networks are shown in Table 4.
[0139] Table 4. Comparison of experimental results between classic single- and double-order networks and improved networks.
[0140]
[0141] As shown in Table 4, compared with other methods, the improved network structure significantly improves both mAP and F1 score while maintaining the detection rate. Using the same dataset for training and testing, compared with YOLOv4 and YOLOv5s, the proposed method shows improvements in P (percentage difference) of 5.9% and 6.7%, R (percentage difference) of 4.4% and 4.0%, and mAP of 4.8% and 5.6%, respectively. It outperforms YOLOv4 in both model parameters and detection rate. The proposed detection algorithm, with an increase of 1.8MB of model parameters, only reduced the detection rate by one frame, indicating that the proposed method is superior in accuracy, detection speed, and portability. Therefore, the wheat scab spore detection model proposed in this invention is reliable. The Faster-RCNN algorithm, as a two-stage detection method, outperforms YOLOv4 in detection accuracy. Compared with the original YOLOv5s algorithm, the algorithm in this invention achieves better performance in mAP values across different IoU threshold ranges.
[0142] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention. The scope of protection claimed by the appended claims and their equivalents is defined.
Claims
1. A method for identifying wheat scab spores based on YOLOv5-ECA-ASFF, characterized in that, Includes the following steps: 11) Establishment of spore image dataset: Collect and collect images of Fusarium head blight spores, preprocess and augment the images to construct a spore image dataset; 12) Constructing a wheat Fusarium head blight spore recognition model: Based on the Yolov5 network, a wheat Fusarium head blight spore recognition model was constructed by integrating the ECA and ASFF modules; The construction of the wheat Fusarium head blight spore recognition model includes the following steps: 121) A wheat Fusarium head blight spore recognition model was constructed based on Yolov5, and the model was set to four layers: The first layer is Input, which is used as the input end for adaptive scaling of the image. It integrates the K-means genetic algorithm to adaptively calculate the optimal anchor box value of the dataset, thereby enhancing the ability to detect small objects. The second layer is the Backbone. The backbone structure of the Backbone network includes the Focus, CSP module and SPP module. The Focus integrates the width and height information of the input image in the channel and further performs feature mapping on the base layer. The CSP structure is divided into two parts and fused using a cross-stage hierarchical structure. The SPP module at the end is used for three structures: convolutional layer, pooling layer and selection filter. The third layer is the Neck, which includes an upsampling structure of FPN+PAN. The FPN layer is combined with the PAN module at the end for upsampling. The fourth layer is the Output layer. The Output layer is used to evaluate whether the algorithm is accurate in detecting and locating targets. It uses the GIOU loss function to calculate the accuracy. The GIOU loss function removes the remaining prediction results after dividing the GIOU_Loss contribution value by the maximum value and outputs the result with the highest classification probability. At the same time, it generates bounding boxes and predicts the type of targets within the bounding boxes. 122) Configure the ECA attention module: The ECA module replaces the fully connected layers of the original structure with a fast one-dimensional convolution method to obtain non-linear feature information generated across channels. A one-dimensional convolution is defined with a kernel size of k, which represents the coverage range of cross-channel information, meaning that the current channel and its k neighboring channels participate in the prediction of channel attention. There is a mapping relationship between k and the dimension C of the total channels, and when the dimension C of the total channels is determined, a one-dimensional convolution kernel k is obtained by adaptive calculation. Projecting the two-dimensional convolution kernel onto the spatial domain yields a nonlinear function relating to the distribution of image feature points. This function can be used to quickly determine the optimal solution, and its mapping relationship is linear, as shown in the formula: , However, its linear mapping relationship is too simplistic, and the number of channels in convolutional networks is generally set to a power of 2. Therefore, the linear function is extended to nonlinear functions, and its calculation formula is as follows: , From the above, specifying the number of channels C, we obtain the formula: , In the above formula: Let b be the odd number closest to x, and b=1, a=2. 123) Improve the second-layer backbone of the wheat scab spore recognition model by adding an ECA module to the end of its CSP residual block; 124) Configure the ASFF feature fusion module: The ASFF feature fusion module maps and fuses scale features at various levels, adaptively learns spatial weight parameters, obtains a new weighting method, and achieves the fusion of features from different layers. In setting up the three-layer structure of ASFF, , , These are the three feature layers of the PANet module, which are then fused through a three-layer feature fusion process. , , The result is from , , Output the features of the three feature layers and perform convolution calculations: Will , , Multiply by the weight parameters respectively , , Then sum them up to obtain the feature output after feature fusion. The formula for this process is as follows: , in, This represents the new feature map obtained through ASFF. , , These represent the weight parameters of the three feature layers, which are then satisfied using the Softmax function. , , , These represent the features of layers 1, 2, and 3, respectively. Ensure by using upsampling or downsampling methods , , The feature structure of each layer's output is the same, and the number of channels remains unchanged; The bottom one First, , The feature information is compressed by using 1×1 convolution to make it consistent with the feature information. The same result was obtained by upsampling at four times and two times the original value. In the same dimension, the summation operation is performed at the end; 125) For the output of the second layer Neck layer FPN+PAN structure of the wheat scab spore recognition model, the ASFF feature fusion module is combined to map and fuse the scale features of each layer, and adaptively learn the spatial weight parameters to obtain a new weighting method, thereby realizing the fusion of features of different layers. 13) Training the wheat Fusarium head blight spore recognition model: Input the spore image dataset into the wheat Fusarium head blight spore recognition model for training; 14) Obtaining wheat Fusarium head blight spores to be identified: Obtain wheat Fusarium head blight spores to be identified and pre-treat them; 15) Obtaining the identification results of wheat Fusarium head blight spores to be identified: Input the pre-processed wheat Fusarium head blight spores to be identified into the trained wheat Fusarium head blight spore identification model to obtain the wheat Fusarium head blight spore identification results.
2. The method for identifying wheat scab spores based on Yolov5-ECA-ASFF according to claim 1, characterized in that, The establishment of the spore image dataset includes the following steps: 21) Select 10,000 images from the collected images, including spore images of four types of fungi that cause wheat scab and mixed spore images of four types of fungi that cause wheat scab, as the target recognition dataset; 22) Expand the dataset from 10,000 images to 20,000 images through data augmentation; 23) The augmented image data was manually labeled using Labelme software. After labeling, the dataset with JSON file labels was converted into the label file type required by Yolov5 to construct the spore image dataset.
3. The method for identifying wheat scab spores based on Yolov5-ECA-ASFF according to claim 1, characterized in that, The training of the wheat Fusarium head blight spore recognition model includes the following steps: 31) Set up a PyTorch neural network training environment with Python 3.8 and CUDA 11.
6. 32) Set the image input size to 640×640, the confidence threshold to 0.5, the initial learning rate to 0.001, the weight decay coefficient to 0.0005, the model training batch size to 16, and the training iteration period to 100 epochs. 33) Input the spore image dataset into the wheat scab spore recognition model, complete the training, and generate a weight file; 331) The first layer Input terminal stitches the data using Mosaic and random scaling, random cropping, and random arrangement methods, and then scales the image data to a standard size of 640×640 before sending it into the detection network; 332) The second backbone layer first uses its focus to slice the image, sending the W and H feature information into the channel space; the CSP module expands the input channels and performs convolution calculations to extract features; the ECA module calculates the number of channels C for the CSP features through a non-linear function mapping, from which a new convolution kernel k can be obtained, realizing adaptive weight adjustment of the CSP features, and obtaining a double-downsampled feature map without information loss; finally, the SPP module transforms the multi-size feature map output by ECA into a fixed-size feature map and feature vector required by the Neck layer through pyramid pooling; 333) The third layer, Neck, first upsamples the images at resolutions of 76×76, 38×38, and 19×19 from high to low using the FPN structure to obtain the semantic information of the target spores. Then, it downsamples the images at resolutions of 19×19, 38×38, and 76×76 from low to high using the PAN structure to obtain the coordinate information of the target spores. Finally, the ASFF module compresses the feature information of 38×38 and 76×76 resolutions into the same number of channels as 19×19 through convolution calculation, and then upsamples it so that the outputs of the three layers are in the same dimension. The obtained weight parameters are used as feature coefficients for output, realizing the fusion of high and low resolutions and making full use of feature information at various scales. 334) The fourth layer Output uses GIOU_Loss as the loss function for the bounding box and calculates the accuracy. After every five training iterations, the accuracy is calculated and a weight file is generated. Repeat the above steps. After one hundred iterations of training, the average accuracy and various evaluation indicators are generated, and the best weight file is retained in all weight files for the detection and identification of Fusarium head blight spores. 34) Use the generated weight file to achieve rapid and accurate detection and identification of Fusarium head blight spores.
Citation Information
Patent Citations
Hyperspectral nondestructive testing method for wheat scab
CN112446298A
Wheat wheat head blight identification method and system based on convolutional neural network
CN114187233A