A method, system, and storage medium for identifying forestry pests
By improving the Yolov8n network, introducing DA_SCConv and EMCA modules, optimizing the backbone structure and detection process, the problems of high error rate and low efficiency of forestry pest recognition in the prior art are solved, and higher detection accuracy and real-time performance are achieved.
Patent Information
- Application Number
- CN202510472745.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-04-16
AI Technical Summary
The prior art has problems such as high error rate, low efficiency and inability to achieve comprehensive monitoring in forest pest identification, especially in the case of a wide variety of forest pests and a complex forest environment.
By improving the Yolov8n network, the DA_SCConv module and EMCA module are introduced to optimize the backbone structure and detection process, including spatial adaptive attention unit, spatial reconstruction unit SRU, hollow convolution and channel reconstruction unit CRU, to enhance the feature expression and detection accuracy of the model.
It improves the accuracy and real-time nature of forestry pest detection, reduces the impact of complex background on detection results, and reduces the complexity and computational cost of the model.
Smart Images

Figure CN119992076B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of agricultural disaster prevention, and particularly relates to a method, a system, and a storage medium for identifying forestry pests. Background Art
[0002] Forestry pests refer to various pests that cause harm to forest vegetation. They not only directly affect the growth and health of trees but may also trigger a series of ecological problems. The early identification and accurate judgment of pests are crucial for implementing effective control measures. Therefore, it is particularly important to develop efficient and accurate pest identification methods. Traditional forestry pest identification methods mainly rely on manual monitoring. These methods usually involve experts' on-site inspections of forest areas for observation, recording, and analysis. Although this method can identify pest species to a certain extent, due to the large variety of pest species, significant individual differences, and the complexity of the forest environment, manual identification often faces problems of high error rates and low efficiency. In addition, the manual identification process is time-consuming and laborious, especially in vast forest areas, where it is difficult to achieve comprehensive monitoring, resulting in many potential pests not being discovered in a timely manner.
[0003] With the rapid development of computer vision and deep learning technologies, especially the wide application of convolutional neural networks (CNNs) in the field of image recognition, pest identification methods have gradually shifted towards automation. The YOLO series of models, as one of the current advanced object detection algorithms, have been widely used in various visual recognition tasks due to their efficient real-time detection capabilities. Yolov8, as a relatively new version of this series, has demonstrated significant improvements in accuracy and speed, becoming an important tool in object recognition research. However, despite the certain advantages of Yolov8 in performance, it still faces challenges in forestry pest identification. First, due to the large variety of pest species, the model may experience confusion during classification. In addition, although the datasets used come from forestry pest control projects on the Internet with white backgrounds, which reduces the impact of complex backgrounds on recognition, the number of parameters of the Yolov8 model is relatively large, which will lead to efficiency problems in actual deployment. At the same time, dense pest features may affect the detection accuracy in high-density situations. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the purpose of the present invention is to provide a method for identifying forestry pests, aiming to solve the technical problems mentioned in the background art.
[0005] To achieve the above purpose, the present invention is implemented through the following technical solutions:
[0006] A method for identifying forestry pests includes the following steps:
[0007] Obtain a dataset related to forestry pests and divide the dataset into a training set, a validation set, and a test set according to a ratio;
[0008] Build a neural network model based on the Yolov8n network, and introduce the DA_SCConv module and the EMCA module into the Yolov8n network to optimize the Yolov8n network;
[0009] Among them, the optimized Yolov8n network sequentially includes a backbone structure, the DA_SCConv module, the SPPF module, and the EMCA module. The DA_SCConv module includes a spatial adaptive attention unit, a spatial reconstruction unit SRU, a dilated convolution, and a channel reconstruction unit CRU;
[0010] Input the training set into the optimized Yolov8n network for training;
[0011] Use the test set to evaluate the trained neural network model to obtain a detection result;
[0012] The step of inputting the training set into the optimized Yolov8n network for training specifically includes:
[0013] Input the training set as an initial feature map into the optimized Yolov8n network, and the initial feature map is processed by the backbone structure to obtain a new feature map;
[0014] Input the new feature map into the spatial adaptive attention unit to make the features in the new feature map focus on a preset area in space;
[0015] Input the new feature map into the spatial reconstruction unit SRU to distinguish useful information and redundant information in the new feature map;
[0016] Process the new feature map processed by the spatial reconstruction unit SRU through dilated convolution to increase the receptive field;
[0017] Input the new feature map processed by the dilated convolution into the channel reconstruction unit CRU to reorganize and filter unnecessary features in the new feature map according to channel features;
[0018] Perform a residual connection between the new feature map processed by the DA_SCConv module and the initial feature map to optimize the feature expression of the new feature map;
[0019] Input the optimized new feature map into the SPPF module to obtain an updated feature map;
[0020] Input the updated feature map into the EMCA module for global average pooling and global maximum pooling processing, and use adaptive convolution to capture the features of each channel in the Yolov8n network to complete the training of the Yolov8n network.
[0021] According to one aspect of the above technical solution, the backbone structure is Conv-Conv-C2f-Conv-C2f-Conv-C2f-Conv-C2f, where Conv represents convolution and C2f represents a feature fusion module.
[0022] According to one aspect of the above technical solution, the initial feature map is processed by the backbone structure to obtain a new feature map, which specifically includes:
[0023] The initial feature map sequentially passes through the backbone structure composed of alternating Conv and C2f, and an intermediate feature map is obtained after being processed by the fourth C2f;
[0024] The intermediate feature map passes through a 1×1 convolutional layer to obtain a new feature map, and the new feature map ∈ (W×H×C).
[0025] According to one aspect of the above technical solution, inputting the new feature map into a spatial adaptive attention unit to focus the features in the new feature map on a preset area in space specifically includes:
[0026] Input the new feature map into the spatial adaptive attention unit;
[0027] Flatten the new feature map through the spatial adaptive attention unit to divide each spatial position into an independent token;
[0028] Process the token through Transformer and label it as X flattened and add a positional encoding P to obtain a positional representation X pos ;
[0029] ;
[0030] Input the X pos into the Transformer encoder for self-attention calculation;
[0031] ;
[0032] where Q represents a query vector, K represents a key vector, C represents a value vector, and T represents a transpose;
[0033] Obtain the attention matrix-weighted V and output the feature representation A spatital ;
[0034] ;
[0035] Restore the feature representation to the spatial dimensions W×H×C.
[0036] According to one aspect of the above technical solution, inputting the new feature map into the spatial reconstruction unit SRU to distinguish useful information and redundant information in the new feature map specifically includes:
[0037] Input the new feature map into the spatial reconstruction unit SRU and normalize the new feature map through group normalization;
[0038] Map the weight values of the normalized new feature map to the interval (0, 1) through the Sigmoid activation function;
[0039] Set a weight threshold T. In the new feature map, regard the feature information with a weight value higher than the weight threshold as useful information and label it as high feature information W1, and regard the feature information with a weight value lower than the weight threshold as redundant information and label it as low feature information W2;
[0040] Adopt a cross-reconstruction operation method to multiply the high feature information W1 by the new feature map to obtain the first feature X1 W , and multiply the low feature information W2 by the new feature map to obtain the second feature X2 W ;
[0041] For the first feature X1 W Perform feature splitting to obtain a first sub-feature X 11 W and a second sub-feature X 12 W , and for the second feature X2 W Perform feature splitting to obtain a third sub-feature X 21 W and a fourth sub-feature X 22 W ;
[0042] For the first sub-feature X 11 W and the fourth sub-feature X 22 W Add them to obtain the enhanced high-information feature X W1 , and for the second sub-feature X 12 W and the third sub-feature X 21 W Add them to obtain the compressed low-information feature X W2 ;
[0043] Add the enhanced high-information feature X W1 and the compressed low-information feature X W2 to obtain the spatial feature output X W and reconstruct the spatial dimension of the new feature map through the spatial feature output X W
[0044] According to one aspect of the above technical solution, the new feature map after the dilated convolution processing is input into the channel reconstruction unit CRU, and unnecessary features in the new feature map are reorganized and filtered according to channel features, specifically including:
[0045] Divide the spatial feature output X W into a first channel feature map with a ratio of αC and a second channel feature map with a ratio of (1−α)C according to channels;
[0046] Extract features from the first channel feature map through the stereo matching model GWCNet and the optical flow model PWCNet respectively, and add and combine them into a first initial channel feature Y1;
[0047] Compress the second channel feature map through the optical flow model PWCNet, and perform a residual connection with the second channel feature map before compression to obtain a second initial channel feature Y2;
[0048] Perform pooling processing on the first initial channel feature Y1 and the second initial channel feature Y2 respectively to generate a first weight vector S1 and a second weight vector S2;
[0049] Normalize the first weight vector S1 and the second weight vector S2 through the Softmax function respectively to obtain a first channel weight β1 and a second channel weight β2;
[0050] Multiply the first channel weight β1 and the first initial channel feature Y1 channel by channel to obtain a first intermediate feature, and multiply the second channel weight β2 and the second initial channel feature Y2 channel by channel to obtain a second intermediate feature;
[0051] Add the first intermediate feature and the second intermediate feature to obtain an output channel feature Y, and use the output channel feature Y to reconstruct the channel dimension of the new feature map.
[0052] According to one aspect of the above technical solution, the new feature map after being processed by the DA_SCConv module is connected with the initial feature map in a residual manner to optimize the feature expression of the new feature map, specifically including:
[0053] The new feature map reconstructed based on the spatial dimension and the new feature map reconstructed based on the channel dimension are used to calculate the new feature map processed by the DA_SCConv module through element-wise multiplication;
[0054] ;
[0055] Among them, SRU(X1) is the new feature map reconstructed based on the spatial dimension, CRU(X2) is the new feature map reconstructed based on the channel dimension, and DA_SCConv(X) is the new feature map processed by the DA_SCConv module;
[0056] The new feature map processed by the DA_SCConv module is connected with the initial feature map through residual connection to retain the original information of the new feature map processed by the DA_SCConv module, thereby optimizing the feature expression of the new feature map.
[0057] According to one aspect of the above technical solution, inputting the updated feature map into the EMCA module for global average pooling and global max pooling processing, and using adaptive convolution to capture the features of each channel in the Yolov8n network to complete the training of the Yolov8n network, specifically including:
[0058] Input the updated feature map into the EMCA module for global average pooling processing and global max pooling processing respectively to obtain the first processing result y avg and the second processing result y max ;
[0059] Fuse the first processing result and the second processing result to obtain the pooling result y;
[0060] ;
[0061] Capture the interaction between the features of each channel and its adjacent channels in the Yolov8n network through adaptive convolution operation;
[0062] ;
[0063] ;
[0064] Among them, w i is the interaction feature, σ is the sigmoid activation function, k is the size of the convolution kernel, w j is the weight of the convolution kernel, γ and b are hyperparameters, C represents the number of channels, φ(C) is the mapping of C, |x| odd represents the odd number closest to X, i represents the serial number of the current channel, and j represents the adjacent channel index;
[0065] The interaction features obtained through the adaptive convolution operation are subjected to residual connection with the updated feature map to complete the training of the Yolov8n network;
[0066] During the process of training the Yolov8n network using the training set, the key hyperparameters in the Yolov8n network are adjusted and optimized using the validation set, and the key hyperparameters include the learning rate, batch size, and weight decay coefficient.
[0067] The present invention also provides a forestry pest identification system, including:
[0068] An acquisition module: used to acquire a dataset related to forestry pests and divide the dataset into a training set, a validation set, and a test set according to a ratio;
[0069] An optimization module: used to establish a neural network model based on the Yolov8n network, and introduce a DA_SCConv module and an EMCA module into the Yolov8n network to optimize the Yolov8n network; wherein the optimized Yolov8n network sequentially includes a backbone structure, the DA_SCConv module, an SPPF module, and the EMCA module, and the DA_SCConv module includes a spatial adaptive attention unit, a spatial reconstruction unit SRU, a dilated convolution, and a channel reconstruction unit CRU;
[0070] A training module: used to input the training set into the optimized Yolov8n network for training;
[0071] The training module is specifically used for:
[0072] Taking the training set as an initial feature map and inputting it into the optimized Yolov8n network, and the initial feature map is processed by the backbone structure to obtain a new feature map;
[0073] Inputting the new feature map into the spatial adaptive attention unit to make the features in the new feature map focus on a preset area in space;
[0074] Inputting the new feature map into the spatial reconstruction unit SRU to distinguish useful information from redundant information in the new feature map;
[0075] Processing the new feature map processed by the spatial reconstruction unit SRU through dilated convolution to increase the receptive field;
[0076] Inputting the new feature map processed by the dilated convolution into the channel reconstruction unit CRU to reorganize and filter unnecessary features in the new feature map according to channel features;
[0077] Perform a residual connection between the new feature map processed by the DA_SCConv module and the initial feature map to optimize the feature representation of the new feature map;
[0078] Input the optimized new feature map into the SPPF module to obtain an updated feature map;
[0079] Input the updated feature map into the EMCA module for global average pooling and global max pooling processing, and use adaptive convolution to capture the features of each channel in the Yolov8n network to complete the training of the Yolov8n network;
[0080] Detection module: used to evaluate the trained neural network model using the test set to obtain detection results.
[0081] The present invention also provides a storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above-mentioned forest pest identification method is implemented.
[0082] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0083] By improving the Yolov8n network, the original modules are replaced with the DA_SCConv module and the EMCA module before and after the SPPF module in the backbone structure. Specifically, the DA_SCConv module includes a spatial adaptive attention unit, a spatial reconstruction unit SRU, a dilated convolution, and a channel reconstruction unit CRU. The spatial adaptive attention unit can automatically focus the features on the spatially important regions. The spatial reconstruction unit SRU can analyze the information in the spatial dimension of the dataset, separate the useful information from the redundant spatial features. The dilated convolution can obtain a larger receptive field and capture richer context information. The channel reconstruction unit CRU can perform recombination and filtering processing according to the channel features to suppress unnecessary features and enhance the representation ability of the key channels. Therefore, the DA_SCConv module can perform CNN compression on the spatial redundancy and channel redundancy between the features in the dataset to reduce the redundant features, lower the complexity of the model, and improve the computational performance. Then, the final output is connected with the initial input feature map through a residual connection. Through the residual connection, the original information in the input features can be retained to form a richer and more concise feature representation. Then, through the extremely lightweight channel attention module EMCA, global average pooling and global max pooling are performed on the dataset, and the results of global average pooling and global max pooling are fused. Adaptive convolution is used to capture the interaction between each channel and its adjacent channels. The features obtained through the adaptive convolution are connected with the input features through a residual connection. In the neck structure (Neck) of the Yolov8n network of the present invention, the EMCA module is added to replace a part of the original neck structure. Specifically, the EMCA module is also added to the upper layer processing at the output end. Therefore, the EMCA module can effectively capture the long-range dependence relationships in the dataset, reduce the influence of complex backgrounds on the detection results, and improve the accuracy and real-time performance of forest pest detection. Description of the Drawings
[0084] Figure 1 It is a flowchart of the forest pest recognition method in the first embodiment of the present invention;
[0085] Figure 2 For Figure 1 The detailed flowchart at step S30 in
[0086] Figure 3 It is a model framework diagram of the improved YOLOv8 in the first embodiment of the present invention;
[0087] Figure 4 It is a structural block diagram of the forest pest recognition system in the second embodiment of the present invention;
[0088] Figure 5 It is a structural block diagram of the electronic device in the third embodiment of the present invention;
[0089] The following specific embodiments will further illustrate the present invention in conjunction with the above-mentioned drawings. Specific Embodiments
[0090] To facilitate the understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant drawings. Several embodiments of the present invention are shown in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided to make the disclosure of the present invention more thorough and comprehensive.
[0091] It should be noted that when an element is referred to as being "fixed to" another element, it can be directly on the other element or there can also be an intermediate element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be an intermediate element at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are for illustrative purposes only.
[0092] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used in the description of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0093] Please refer to Figures 1 to 3 , which shows a method for identifying forestry pests in the first embodiment of the present invention, including the following steps:
[0094] S10. Obtain a data set related to forestry pests and divide the data set into a training set, a validation set, and a test set according to a ratio;
[0095] S20. Establish a neural network model based on the Yolov8n network, and introduce a DA_SCConv module and an EMCA module into the Yolov8n network to optimize the Yolov8n network;
[0096] The optimized Yolov8n network sequentially includes a backbone structure, the DA_SCConv module, an SPPF module, and the EMCA module. The DA_SCConv module includes a spatial adaptive attention unit, a spatial reconstruction unit SRU, a dilated convolution, and a channel reconstruction unit CRU;
[0097] S30. Input the training set into the optimized Yolov8n network for training;
[0098] S40. Use the test set to evaluate the trained neural network model to obtain a detection result;
[0099] The step of inputting the training set into the optimized Yolov8n network for training specifically includes:
[0100] S31. Input the training set as an initial feature map into the optimized Yolov8n network. After being processed by the backbone structure, a new feature map is obtained;
[0101] S32. Input the new feature map into the spatial adaptive attention unit to make the features in the new feature map focus on a preset area in space;
[0102] S33. Input the new feature map into the spatial reconstruction unit SRU to distinguish useful information and redundant information in the new feature map;
[0103] S34. Process the new feature map processed by the spatial reconstruction unit SRU through dilated convolution to increase the receptive field;
[0104] S35. Input the new feature map processed by the dilated convolution into the channel reconstruction unit CRU to reorganize and filter unnecessary features in the new feature map according to channel features;
[0105] S36. Perform a residual connection between the new feature map processed by the DA_SCConv module and the initial feature map to optimize the feature expression of the new feature map;
[0106] S37. Input the optimized new feature map into the SPPF module to obtain an updated feature map;
[0107] S38. Input the updated feature map into the EMCA module for global average pooling and global max pooling processing, and use adaptive convolution to capture the features of each channel in the Yolov8n network to complete the training of the Yolov8n network.
[0108] It can be understood that the present invention improves the Yolov8n network by replacing the original modules with a DA_SCConv module and an EMCA module before and after the SPPF module in the backbone structure. Specifically, the DA_SCConv module includes a spatial adaptive attention unit, a spatial reconstruction unit SRU, a dilated convolution, and a channel reconstruction unit CRU. The spatial adaptive attention unit can automatically focus the features on the spatially important regions. The spatial reconstruction unit SRU can analyze the information in the spatial dimension of the dataset, separate the useful information from the redundant spatial features. The dilated convolution can obtain a larger receptive field and capture richer context information. The channel reconstruction unit CRU can perform recombination and filtering processing according to the channel features to suppress unnecessary features and enhance the representation ability of the key channels. Therefore, the DA_SCConv module can perform CNN compression on the spatial redundancy and channel redundancy between the features in the dataset to reduce the redundant features, thereby reducing the complexity of the model and improving the computational performance. Then, the final output is connected with the initial input feature map through a residual connection. Through the residual connection, the original information in the input features can be retained to form a richer and more concise feature representation. Then, through the extremely lightweight channel attention module EMCA, global average pooling and global max pooling are performed on the dataset, and the results of the global average pooling and global max pooling are fused, and adaptive convolution is used to capture the interaction between each channel and its adjacent channels. The features obtained through the adaptive convolution are connected with the input features through a residual connection. The present invention adds an EMCA module to replace a part of the original neck structure in the neck structure (Neck) of the Yolov8n network. Specifically, an EMCA module is also added to the processing of the upper layer at the output end. Therefore, the EMCA module can effectively capture the long-range dependencies in the dataset, reduce the influence of complex backgrounds on the detection results, and improve the accuracy and real-time performance of forest pest detection.
[0109] Specifically, in this embodiment, the backbone structure is Conv-Conv-C2f-Conv-C2f-Conv-C2f-Conv-C2f;
[0110] The step S31 specifically includes:
[0111] The initial feature map sequentially passes through the backbone structure composed of the Conv and the C2f in an alternating manner, and an intermediate feature map is obtained after the fourth C2f processing;
[0112] The intermediate feature map is passed through a 1×1 convolutional layer to obtain a new feature map, and the new feature map ∈ (W×H×C), where Conv represents convolution and C2f represents a feature fusion module.
[0113] It can be understood that the function of the backbone structure is to perform convolution operations and feature fusion on the training set, and obtain an intermediate feature map (20*20*1024) after preprocessing. Then, a new feature map is obtained through a 1×1 convolutional layer. The new feature map has three dimensions (W×H×C), where W represents the width, H represents the height, and C represents the number of channels. The backbone structure is an existing network model and will not be elaborated here.
[0114] Further, the specific steps of step S32 include:
[0115] Input the new feature map into the spatial adaptive attention unit;
[0116] Flatten the new feature map through the spatial adaptive attention unit to divide each spatial position into an independent token;
[0117] Process the token through the Transformer and label it as X flattened and add the position encoding P to obtain the position representation X pos ;
[0118] ;
[0119] Input the X pos into the Transformer encoder for self-attention calculation;
[0120] ;
[0121] where Q represents the query vector, K represents the key vector, C represents the value vector, and T represents the transpose;
[0122] Obtain the attention matrix weighted V and output the feature representation A spatital ;
[0123] ;
[0124] Restore the feature representation to the spatial dimension W×H×C.
[0125] It can be understood that through the above steps of processing, the spatial adaptive attention unit can automatically focus the features on the important regions in space, preparing for subsequent processing work.
[0126] Further, the specific steps of step S33 include:
[0127] Input the new feature map into the spatial reconstruction unit SRU and standardize the new feature map through group normalization;
[0128] Map the weight values of the normalized new feature map to the interval (0, 1) through the Sigmoid activation function;
[0129] Set a weight threshold T. In the new feature map, regard the feature information with a weight value higher than the weight threshold as useful information and label it as high feature information W1, and regard the feature information with a weight value lower than the weight threshold as redundant information and label it as low feature information W2;
[0130] Adopt a cross-reconstruction operation method, multiply the high feature information W1 by the new feature map to obtain the first feature X1 W and multiply the low feature information W2 by the new feature map to obtain the second feature X2 W ;
[0131] Perform feature splitting on the first feature X1 W to obtain the first sub-feature X 11 W and the second sub-feature X 12 W , and perform feature splitting on the second feature X2 W to obtain the third sub-feature X 21 W and the fourth sub-feature X 22 W ;
[0132] Add the first sub-feature X 11 W and the fourth sub-feature X 22 W to obtain the enhanced high-information feature X W1 , and add the second sub-feature X 12 W and the third sub-feature X 21 W to obtain the compressed low-information feature X W2 ;
[0133] Add the enhanced high-information feature X W1 and the compressed low-information feature X W2 to obtain the spatial feature output X W , and reconstruct the spatial dimension of the new feature map through the spatial feature output X W .
[0134] It can be understood that the new feature map can be grouped and normalized by the spatial reconstruction unit SRU, and then the new feature map is standardized, so as to form a more consistent distribution among different feature groups. Then, the normalized new feature map passes through the Sigmoid activation function to map the weight value to between 0 and 1. Then, a threshold T is set to divide the new feature map into high feature information W1 and low feature information W2. Among them, when the weight value at a certain position is higher than T, it is considered that the feature information at this position is rich and belongs to the high feature information W1, while when the weight value is lower than T, it is considered that the feature at this position is mainly redundant information and belongs to the low feature information W2. Then, cross-reconstruction operation is adopted to obtain the enhanced high-information feature X W1 and the compressed low-information feature X W2 , finally, through the concatenation (Concat) operation, the enhanced high-information feature X W1 and the compressed low-information feature X W2 are added together to obtain the spatial feature output X W , the spatial feature output X W can reflect the useful information and redundant information in the new feature map, and this spatial feature output X W reconstructs the spatial dimension of the new feature map.
[0135] Further, the step S35 specifically includes:
[0136] Dividing the spatial feature output X W into a first channel feature map with a ratio of αC and a second channel feature map with a ratio of (1−α)C according to channels;
[0137] Extracting the features in the first channel feature map through the stereo matching model GWCNet and the optical flow model PWCNet respectively, and adding and merging them into a first initial channel feature Y1;
[0138] Compressing the second channel feature map through the optical flow model PWCNet, and performing a residual connection with the second channel feature map before compression to obtain a second initial channel feature Y2;
[0139] Performing pooling processing on the first initial channel feature Y1 and the second initial channel feature Y2 respectively to generate a first weight vector S1 and a second weight vector S2;
[0140] Normalizing the first weight vector S1 and the second weight vector S2 through the Softmax function respectively to obtain a first channel weight β1 and a second channel weight β2;
[0141] Multiply the first channel weight β1 and the first initial channel feature Y1 channel by channel to obtain a first intermediate feature, and multiply the second channel weight β2 and the second initial channel feature Y2 channel by channel to obtain a second intermediate feature;
[0142] Add the first intermediate feature and the second intermediate feature to obtain an output channel feature Y, and use the output channel feature Y to reconstruct the channel dimension of the new feature map.
[0143] It can be understood that the channel reconstruction unit CRU can divide the X output by the SRU W into two feature maps according to channels, perform partition processing, then use the stereo matching model GWCNet and the optical flow model PWCNet to extract features in the first channel feature map, and add and combine them into the first initial channel feature Y1. Thus, the training set is trained in the stereo space. Then, only use the optical flow model to compress the second channel feature map, and perform a residual connection with the second channel feature map before compression to obtain the second initial channel feature Y2. By extracting two initial channel features, key data can be retained. Then, pooling processing on the two can reduce the data dimension, greatly reducing the data calculation amount. At the same time, two weight vectors are obtained, and then the two weight vectors are normalized to obtain the weights of the two channels. By multiplying the two channel weights and the initial channels channel by channel respectively, the intermediate features in the initial channels can be obtained. Finally, adding the two intermediate features can obtain the output channel feature Y. The output channel feature Y can be used to reorganize and filter the unnecessary features in the new feature map, while enhancing the representation ability of the key channels, thereby obtaining a new feature map after reconstruction of the channel dimension.
[0144] Further, the step S36 specifically includes:
[0145] Based on the new feature map reconstructed in the spatial dimension and the new feature map reconstructed in the channel dimension, calculate the new feature map processed by the DA_SCConv module through element-wise multiplication;
[0146] ;
[0147] where SRU(X1) is the new feature map reconstructed in the spatial dimension, CRU(X2) is the new feature map reconstructed in the channel dimension, and DA_SCConv(X) is the new feature map processed by the DA_SCConv module;
[0148] The new feature map processed by the DA_SCConv module is subjected to residual connection with the initial feature map to retain the original information of the new feature map processed by the DA_SCConv module, thereby optimizing the feature expression of the new feature map.
[0149] It can be understood that the new feature map after reconstruction in the spatial dimension and the new feature map after reconstruction in the channel dimension are multiplied element by element to obtain the new feature map processed by the DA_SCConv module. At this time, the training of the DA_SCConv module ends. Then, the new feature map processed by the DA_SCConv module is subjected to residual connection with the initial feature map, which can solve the problems of gradient disappearance and gradient explosion, and can also help the model converge faster, improve the stability of the network, thereby retaining the original information on the new feature map, and further optimizing the feature expression of the new feature map.
[0150] Then, the optimized new feature map is processed by the existing SPPF module to obtain the updated feature map.
[0151] Further, the step S38 specifically includes:
[0152] The updated feature map is respectively input into the EMCA module for global average pooling processing and global max pooling processing to respectively obtain the first processing result y avg and the second processing result y max ;
[0153] The first processing result and the second processing result are fused to obtain the pooling result y;
[0154] ;
[0155] The interaction between the features of each channel and its adjacent channel features in the Yolov8n network is captured through adaptive convolution operations;
[0156] ;
[0157] ;
[0158] where, w i is the interaction feature, σ is the sigmoid activation function, k is the size of the convolution kernel, w j is the weight of the convolution kernel, γ and b are hyperparameters, C represents the number of channels, φ(C) is the mapping of C, |x| odd represents the odd number closest to X, i represents the serial number of the current channel, and j represents the adjacent channel index;
[0159] The interaction features obtained through the adaptive convolution operation are subjected to residual connection with the updated feature map to complete the training of the Yolov8n network.
[0160] It can be understood that the last step of training is to perform pooling on the updated feature map through the EMCA module to reduce the dimensionality of the data, which enables the model to better extract important features and reduce redundant information. Then, through the adaptive convolution operation, the standard convolution can be simply and effectively modified, and the size of the convolution kernel can be changed according to the learnable local pixel features, thereby optimizing the model and capturing the interaction between each channel and its adjacent channels, and useful channel features can be retained. The purpose of the residual connection is to solve the problems of gradient vanishing and gradient explosion, and at the same time, it can also help the model converge faster and improve the stability of the network. After that, the results can be output, and thus the training of the model is completed.
[0161] Preferably, during the process of training the Yolov8n network using the training set, the key hyperparameters in the Yolov8n network are adjusted and optimized using the validation set to improve the convergence speed and detection accuracy of the model. The key hyperparameters include the learning rate, batch size, and weight decay coefficient. At the same time, the cosine annealing learning rate strategy is adopted to adjust the learning rate at different stages of training to better adapt to the characteristics of the dataset and the training process of the model.
[0162] Furthermore, a variety of data augmentation techniques are applied to the forestry pest dataset, including random cropping, rotation, scaling, color jittering, blurring, etc., so that the model can adapt to complex actual application scenarios and improve the robustness of the model to multiple pest types.
[0163] Furthermore, after the training is completed, the test set is used to evaluate the model. Five models, namely Yolov4, YOLOv4_tiny, SSD_MobileNet, Yolov4_MF, and Yolov8, are used as benchmarks and compared with Yolov8_SCM, as shown in the following table:
[0164]
[0165] It should be noted that Figure 3 In it, If shortcut means if the skip connection, detect means the detection head of the model, Bottleneck means the bottleneck layer, add means the fusion operation, Bn means batch normalization, SiLU means the activation function, Split means the feature map segmentation, MaxPool means the max pooling, Bbox.Loss means the bounding box loss, and Cls.Loss means the classification loss.
[0166] In summary, the forestry pest identification method in the above embodiments of the present invention can perform CNN compression on the spatial redundancy and channel redundancy between features in the dataset to reduce redundant features, thereby reducing the complexity of the model and improving the computing performance. It can also effectively capture the long-distance dependencies in the dataset, reduce the influence of complex backgrounds on the detection results, and improve the accuracy and real-time performance of forestry pest detection.
[0167] Please refer to Figure 4 , which shows the forestry pest identification system in the second embodiment of the present invention, including:
[0168] Acquisition module 11: used to acquire a dataset related to forestry pests and divide the dataset into a training set, a validation set, and a test set according to a ratio;
[0169] Optimization module 12: used to establish a neural network model based on the Yolov8n network and introduce a DA_SCConv module and an EMCA module into the Yolov8n network to optimize the Yolov8n network; where the optimized Yolov8n network sequentially includes a backbone structure, the DA_SCConv module, an SPPF module, and the EMCA module, and the DA_SCConv module includes a spatial adaptive attention unit, a spatial reconstruction unit SRU, a dilated convolution, and a channel reconstruction unit CRU;
[0170] Training module 13: used to input the training set into the optimized Yolov8n network for training;
[0171] The training module 13 is specifically used for:
[0172] Taking the training set as an initial feature map and inputting it into the optimized Yolov8n network, and obtaining a new feature map after the initial feature map is processed by the backbone structure;
[0173] Inputting the new feature map into the spatial adaptive attention unit to make the features in the new feature map focus on a preset area in space;
[0174] Inputting the new feature map into the spatial reconstruction unit SRU to distinguish useful information and redundant information in the new feature map;
[0175] Processing the new feature map processed by the spatial reconstruction unit SRU through dilated convolution to increase the receptive field;
[0176] Inputting the new feature map processed by the dilated convolution into the channel reconstruction unit CRU to reorganize and filter unnecessary features in the new feature map according to channel features;
[0177] Perform a residual connection between the new feature map processed by the DA_SCConv module and the initial feature map to optimize the feature representation of the new feature map;
[0178] Input the optimized new feature map into the SPPF module to obtain an updated feature map;
[0179] Input the updated feature map into the EMCA module for global average pooling and global maximum pooling processing, and use adaptive convolution to capture the features of each channel in the Yolov8n network to complete the training of the Yolov8n network;
[0180] Detection module 14: Used to evaluate the trained neural network model using the test set to obtain detection results.
[0181] The present invention also proposes an electronic device. Please refer to Figure 5 , showing the electronic device in the third embodiment of the present invention, including a memory 10, a processor 20, and a computer program 30 stored on the memory 10 and executable on the processor 20. When the processor 20 executes the computer program 30, the above-mentioned forest pest identification method is implemented.
[0182] Among them, the memory 10 includes at least one type of storage medium. The storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. The memory 10 can be an internal storage unit of the electronic device in some embodiments, such as the hard disk of the electronic device. The memory 10 can also be an external storage device in other embodiments, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 10 can also include both the internal storage unit of the electronic device and the external storage device. The memory 10 can be used not only to store application software and various types of data installed in the electronic device, but also to temporarily store data that has been output or will be output.
[0183] Among them, the processor 20 can be an Electronic Control Unit (ECU, also known as a vehicle computer), a Central Processing Unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips in some embodiments, and is used to run the program code stored in the memory 10 or process data, such as executing an access restriction program, etc.
[0184] It should be noted thatFigure 5 The structure shown does not constitute a limitation on the electronic device. In other embodiments, the electronic device may include fewer or more components than shown, or combine certain components, or have a different component arrangement.
[0185] An embodiment of the present invention also provides a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the forest pest identification method as described above is implemented.
[0186] Those skilled in the art can understand that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a defined sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0187] More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection part (electronic device) having one or more wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.
[0188] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0189] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0190] The above-described embodiments merely represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the appended claims.
Claims
1. A method for identifying forest pests, characterized in that: The steps include: Acquire a data set related to forest pests, and divide the data set into a training set, a validation set, and a test set in proportion; A neural network model is established based on the Yolov8n network, and a DA_SCConv module and an EMCA module are introduced into the Yolov8n network to optimize the Yolov8n network; The optimized Yolov8n network includes a backbone structure, a DA_SCConv module, an SPPF module, and an EMCA module in sequence, wherein the DA_SCConv module includes a spatial adaptive attention unit, a spatial reconstruction unit SRU, a hole convolution, and a channel reconstruction unit CRU; Input the training set into the optimized Yolov8n network for training; Using the test set to evaluate the trained neural network model to obtain a detection result; Inputting the training set into the optimized Yolov8n network for training specifically includes: Input the training set as an initial feature map into the optimized Yolov8n network, and obtain a new feature map after the initial feature map is processed by the backbone structure; Inputting the new feature map into a spatial adaptive attention unit so that the features in the new feature map are focused on a preset area in space; Inputting the new feature map into the spatial reconstruction unit SRU to distinguish useful information from redundant information in the new feature map; The new feature map processed by the spatial reconstruction unit SRU is processed by a dilated convolution to increase the receptive field; Inputting the new feature map after the hole convolution processing into the channel reconstruction unit CRU, and reorganizing and filtering the unnecessary features in the new feature map according to the channel features; Performing a residual connection between the new feature map processed by the DA_SCConv module and the initial feature map to optimize the feature expression of the new feature map; Inputting the optimized new feature map into the SPPF module to obtain an updated feature map; The updated feature map is input into the EMCA module for global average pooling and global maximum pooling processing, and adaptive convolution is used to capture the features of each channel in the Yolov8n network to complete the training of the Yolov8n network.
2. The forest pest identification method according to claim 1, characterized in that: The backbone structure is Conv-Conv-C2f-Conv-C2f-Conv-C2f-Conv-C2f, where Conv represents convolution and C2f represents a feature fusion module.
3. The forest pest identification method according to claim 2, characterized in that: The initial feature map is processed by the backbone structure to obtain a new feature map, which specifically includes: The initial feature map is sequentially processed by the backbone structure formed by the interlacing of the Conv and the C2f, and an intermediate feature map is obtained after the fourth C2f processing; The intermediate feature map is passed through a 1×1 convolution layer to obtain a new feature map, where the new feature map ∈ (W×H×C).
4. The forest pest identification method according to claim 3, characterized in that: The step of inputting the new feature map into a spatial adaptive attention unit so that the features in the new feature map are focused on a preset area in space specifically includes: Inputting the new feature map into a spatially adaptive attention unit; Flattening the new feature map by the spatial adaptive attention unit to divide each spatial position into an independent token; The token is processed by Transformer and marked as X flattened , and add the position code P to get the position representation X pos ; ; The X pos Input to Transformer encoder for self-attention calculation; ; Among them, Q represents the query vector, K represents the key vector, C represents the value vector, and T represents the transpose; Get the attention matrix weight V and output feature representation A spatital ; ; The feature representation is restored to the spatial dimension W×H×C.
5. The forest pest identification method according to claim 4, characterized in that: The step of inputting the new feature map into the spatial reconstruction unit SRU to distinguish useful information from redundant information in the new feature map specifically includes: Inputting the new feature map into the spatial reconstruction unit SRU, and normalizing the new feature map by group normalization; The weight value of the standardized new feature map is mapped to the interval (0, 1) through the Sigmoid activation function; A weight threshold T is set. In the new feature graph, feature information with a weight value higher than the weight threshold is regarded as useful information and marked as high feature information W1, and feature information with a weight value lower than the weight threshold is regarded as redundant information and marked as low feature information W2; Using the cross-reconstruction operation method, the high feature information W1 is multiplied by the new feature map to obtain the first feature X1 W , and multiplying the low feature information W2 by the new feature map to obtain the second feature X2 W ; The first feature X1 W Perform feature splitting to obtain the first sub-feature X 11 W and the second sub-feature X 12 W , and the second feature X2 W Perform feature splitting to obtain the third sub-feature X 21 W and the fourth sub-feature X 22 W ; The first sub-feature X 11 W and the fourth sub-feature X 22 W Add together to get the enhanced high information feature X W1 , and the second sub-feature X 12 W and the third sub-feature X 21 W Add to get the compressed low information feature X W2 ; The enhanced high information feature X W1 and the compressed low information feature X W2 Add to obtain the spatial feature output X W , and output X through the spatial feature W The spatial dimension of the new feature map is reconstructed.
6. The forest pest identification method according to claim 5, characterized in that: The new feature map after the hole convolution processing is input into the channel reconstruction unit CRU, and the non-essential features in the new feature map are reorganized and filtered according to the channel features, specifically including: Output the spatial feature X W Divide into the first channel feature map with a ratio of αC and the second channel feature map with a ratio of (1−α)C according to the channel; The features in the first channel feature map are extracted by the stereo matching model GWCNet and the optical flow model PWCNet respectively, and added and merged into the first initial channel feature Y1; The second channel feature map is compressed by the optical flow model PWCNet, and a residual connection is performed with the second channel feature map before compression to obtain a second initial channel feature Y2; Performing pooling processing on the first initial channel feature Y1 and the second initial channel feature Y2 respectively to generate a first weight vector S1 and a second weight vector S2 respectively; Normalizing the first weight vector S1 and the second weight vector S2 respectively by using a Softmax function to obtain a first channel weight β1 and a second channel weight β2 respectively; Multiplying the first channel weight β1 and the first initial channel feature Y1 channel by channel to obtain a first intermediate feature, and multiplying the second channel weight β2 and the second initial channel feature Y2 channel by channel to obtain a second intermediate feature; The first intermediate feature and the second intermediate feature are added to obtain an output channel feature Y, and the output channel feature Y is used to reconstruct the channel dimension of the new feature map.
7. The forest pest identification method according to claim 6, characterized in that: The performing residual connection between the new feature map processed by the DA_SCConv module and the initial feature map to optimize the feature expression of the new feature map specifically includes: Based on the new feature map reconstructed in the spatial dimension and the new feature map reconstructed in the channel dimension, the new feature map processed by the DA_SCConv module is calculated by element-by-element multiplication; ; Among them, SRU (X1) is the new feature map after reconstruction of the spatial dimension, CRU (X2) is the new feature map after reconstruction of the channel dimension, and DA_SCConv (X) is the new feature map after processing by the DA_SCConv module; A residual connection is performed between the new feature map processed by the DA_SCConv module and the initial feature map to retain the original information of the new feature map processed by the DA_SCConv module, thereby optimizing the feature expression of the new feature map.
8. The forest pest identification method according to claim 7, characterized in that: The updated feature map is input into the EMCA module for global average pooling and global maximum pooling, and adaptive convolution is used to capture the features of each channel in the Yolov8n network to complete the training of the Yolov8n network, specifically including: The updated feature maps are respectively input into the EMCA module for global average pooling processing and global maximum pooling processing to obtain the first processing results y avg and the second processing result y max ; Fusing the first processing result and the second processing result to obtain a pooling result y; ; The interaction between each channel feature and its adjacent channel features in the Yolov8n network is captured through adaptive convolution operations; ; ; Among them, w i is the interaction feature, σ is the sigmoid activation function, k is the size of the convolution kernel, and w j is the weight of the convolution kernel, γ and b are hyperparameters, C represents the number of channels, φ(C) is the mapping of C, |x| odd represents the odd number closest to X, i represents the sequence number of the current channel, and j represents the index of the adjacent channel; Performing a residual connection between the interactive features obtained by the adaptive convolution operation and the updated feature map to complete the training of the Yolov8n network; In the process of training the Yolov8n network using the training set, the validation set is used to adjust and optimize key hyperparameters in the Yolov8n network, wherein the key hyperparameters include learning rate, batch size, and weight decay coefficient.
9. A forestry pest identification system, characterized in that: include: Acquisition module: used to acquire a data set related to forest pests, and divide the data set into a training set, a validation set, and a test set in proportion; Optimization module: used to establish a neural network model based on the Yolov8n network, and introduce the DA_SCConv module and the EMCA module into the Yolov8n network to optimize the Yolov8n network; wherein the optimized Yolov8n network includes a backbone structure, the DA_SCConv module, the SPPF module, and the EMCA module in sequence, and the DA_SCConv module includes a spatial adaptive attention unit, a spatial reconstruction unit SRU, a hole convolution, and a channel reconstruction unit CRU; Training module: used for inputting the training set into the optimized Yolov8n network for training; The training module is specifically used for: Input the training set as an initial feature map into the optimized Yolov8n network, and obtain a new feature map after the initial feature map is processed by the backbone structure; Inputting the new feature map into a spatial adaptive attention unit so that the features in the new feature map are focused on a preset area in space; Inputting the new feature map into the spatial reconstruction unit SRU to distinguish useful information from redundant information in the new feature map; The new feature map processed by the spatial reconstruction unit SRU is processed by a dilated convolution to increase the receptive field; Inputting the new feature map after the hole convolution processing into the channel reconstruction unit CRU, and reorganizing and filtering the unnecessary features in the new feature map according to the channel features; Performing a residual connection between the new feature map processed by the DA_SCConv module and the initial feature map to optimize the feature expression of the new feature map; Inputting the optimized new feature map into the SPPF module to obtain an updated feature map; Input the updated feature map into the EMCA module for global average pooling and global maximum pooling processing, and use adaptive convolution to capture the features of each channel in the Yolov8n network to complete the training of the Yolov8n network; Detection module: used to evaluate the trained neural network model using the test set to obtain a detection result.
10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the forestry pest identification method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
YOLOv8 vehicle identification method, system and device and storage medium
CN118470494A
Track foreign matter detection method based on YOLO algorithm
CN119478380A