Forestry pest identification method and system, and storage medium

By introducing the DA_SCConv module and EMCA module into the YOLOv8n network, the network structure is optimized, and the problems of high error rate and low efficiency in forestry pest recognition in the prior art are solved, and higher detection accuracy and real-time performance are achieved.

CN119992076AActive Publication Date: 2025-05-13JIANGXI AGRICULTURAL UNIVERSITY
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510472745.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-05-13
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

The prior art has problems such as high error rate, low efficiency and inability to achieve comprehensive monitoring in forestry pest identification. Especially when using the YOLOv8 model, it faces confusion caused by a wide variety of pests and efficiency problems caused by large model parameters.

Method used

By introducing the DA_SCConv module and EMCA module in the YOLOv8n network, the network structure is optimized, including the spatial adaptive attention unit, the spatial reconstruction unit SRU, the hollow convolution and channel reconstruction unit CRU, as well as global average pooling and global maximum pooling processing, to reduce redundant features and improve computing performance.

Benefits of technology

It improves the accuracy and real-time nature of forestry pest detection, reduces the impact of complex background on detection results, and reduces the complexity and computational cost of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992076A_ABST
    Figure CN119992076A_ABST
Patent Text Reader

Abstract

The invention provides a forestry pest identification method and system, and a storage medium. The method comprises the following steps: acquiring a data set; the DASCConv and the EMCA are introduced into the Yolov8n, so that the Yolov8n is obtained; inputting the training set into Yolov8n for training; evaluating the model by using the test set; inputting the training set into Yolov8n training comprises the following steps: taking the training set as an initial feature map, and processing the initial feature map to obtain a new feature map; inputting the new feature map into the SRU, and distinguishing useful information from redundant information; inputting the processed new feature map into the CRU so as to recombine and filter unnecessary features of the new feature map; performing residual connection on the new feature map and the initial feature map; and inputting the feature map into the EMCA for processing, and capturing the features of the channel. According to the method, useful information and redundant information are separated, the calculation performance is improved, the long-distance dependency relationship in the data is captured, and the precision of forestry pest detection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of agricultural disaster prevention, and in particular to a forestry pest identification method, system, and storage medium. Background Art

[0002] Forest pests refer to various pests that cause harm to forest vegetation. They not only directly affect the growth and health of trees, but may also cause a series of ecological problems. The early identification and accurate judgment of pests are crucial for the implementation of effective prevention and control measures. Therefore, it is particularly important to develop efficient and accurate pest identification methods. Traditional forest pest identification methods mainly rely on manual monitoring. These methods usually involve experts conducting field visits to forest areas to observe, record and analyze. Although this method can identify pest species to a certain extent, due to the wide variety of pests, large individual differences, and the complexity of the forest environment, manual identification often faces high error rates and low efficiency. In addition, the manual identification process is time-consuming and labor-intensive, especially in vast forest areas, where it is difficult to achieve comprehensive monitoring, resulting in many potential pests not being discovered in time.

[0003] With the rapid development of computer vision and deep learning technology, especially the widespread application of convolutional neural networks (CNN) in the field of image recognition, pest identification methods are gradually shifting towards automation. As one of the current advanced target detection algorithms, the YOLO series of models has been widely used in various visual recognition tasks with its efficient real-time detection capabilities. As a newer version of the series, Yolov8 has demonstrated significant improvements in accuracy and speed, becoming an important tool in target recognition research. However, despite its certain advantages in performance, Yolov8 still faces challenges in forestry pest identification. First, due to the wide variety of pests, the model may be confused when classifying. In addition, although the dataset used comes from a forestry pest control project on the Internet, the background is all white, which reduces the impact of complex background on recognition, but the Yolov8 model has a relatively large number of parameters, which will lead to efficiency problems in actual deployment. At the same time, dense pest features may affect the accuracy of detection in high-density situations. Summary of the invention In view of the deficiencies in the prior art, the purpose of the present invention is to provide a forestry pest identification method, aiming to solve the technical problems mentioned in the background technology.

[0004] In order to achieve the above object, the present invention is implemented by the following technical solutions: A method for identifying forest pests comprises the following steps: Acquire a data set related to forest pests, and divide the data set into a training set, a validation set, and a test set in proportion; A neural network model is established based on the Yolov8n network, and a DA_SCConv module and an EMCA module are introduced into the Yolov8n network to optimize the Yolov8n network; The optimized Yolov8n network includes a backbone structure, a DA_SCConv module, an SPPF module, and an EMCA module in sequence, wherein the DA_SCConv module includes a spatial adaptive attention unit, a spatial reconstruction unit SRU, a hole convolution, and a channel reconstruction unit CRU; Input the training set into the optimized Yolov8n network for training; Using the test set to evaluate the trained neural network model to obtain a detection result; The inputting the training set into the optimized Yolov8n network for training specifically includes: Input the training set as an initial feature map into the optimized Yolov8n network, and obtain a new feature map after the initial feature map is processed by the backbone structure; Inputting the new feature map into a spatial adaptive attention unit so that the features in the new feature map are focused on a preset area in space; Inputting the new feature map into the spatial reconstruction unit SRU to distinguish useful information from redundant information in the new feature map; The new feature map processed by the spatial reconstruction unit SRU is processed by a dilated convolution to increase the receptive field; Inputting the new feature map after the hole convolution processing into the channel reconstruction unit CRU, and reorganizing and filtering the unnecessary features in the new feature map according to the channel features; Performing a residual connection between the new feature map processed by the DA_SCConv module and the initial feature map to optimize the feature expression of the new feature map; Inputting the optimized new feature map into the SPPF module to obtain an updated feature map; The updated feature map is input into the EMCA module for global average pooling and global maximum pooling processing, and adaptive convolution is used to capture the features of each channel in the Yolov8n network to complete the training of the Yolov8n network.

[0005] According to one aspect of the above technical solution, the backbone structure is Conv-Conv-C2f-Conv-C2f-Conv-C2f-Conv-C2f, where Conv represents convolution and C2f represents a feature fusion module.

[0006] According to one aspect of the above technical solution, the initial feature map is processed by the backbone structure to obtain a new feature map, which specifically includes: The initial feature map is sequentially processed by the backbone structure formed by the interlacing of the Conv and the C2f, and an intermediate feature map is obtained after the fourth C2f processing; The intermediate feature map is passed through a 1×1 convolution layer to obtain a new feature map, where the new feature map ∈ (W×H×C).

[0007] According to one aspect of the above technical solution, the step of inputting the new feature map into a spatial adaptive attention unit so that the features in the new feature map are focused on a preset area in space specifically includes: Inputting the new feature map into a spatially adaptive attention unit; Flattening the new feature map by the spatial adaptive attention unit to divide each spatial position into an independent token; The token is processed by Transformer and marked as X flattened , and add the position code P to get the position representation X pos ; ; The X pos Input to Transformer encoder for self-attention calculation; ; Among them, Q represents the query vector, K represents the key vector, C represents the value vector, and T represents the transpose; Get the attention matrix weight V and output feature representation A spatital ; ; The feature representation is restored to the spatial dimension W×H×C.

[0008] According to one aspect of the above technical solution, the inputting the new feature map into the spatial reconstruction unit SRU to distinguish useful information from redundant information in the new feature map specifically includes: Inputting the new feature map into the spatial reconstruction unit SRU, and normalizing the new feature map by group normalization; The weight value of the standardized new feature map is mapped to the interval (0, 1) through the Sigmoid activation function; A weight threshold T is set. In the new feature graph, feature information with a weight value higher than the weight threshold is regarded as useful information and marked as high feature information W1, and feature information with a weight value lower than the weight threshold is regarded as redundant information and marked as low feature information W2; Using the cross-reconstruction operation method, the high feature information W1 is multiplied by the new feature map to obtain the first feature X1 W , and multiplying the low feature information W2 by the new feature map to obtain the second feature X2 W ; The first feature X1 W Perform feature splitting to obtain the first sub-feature X 11 W and the second sub-feature X 12 W , and the second feature X2 W Perform feature splitting to obtain the third sub-feature X 21 W and the fourth sub-feature X 22 W ; The first sub-feature X 11 W and the fourth sub-feature X 22 W Add together to get the enhanced high information feature X W1 , and the second sub-feature X 12 W and the third sub-feature X 21 W Add to get the compressed low information feature X W2 ; The enhanced high information feature X W1 and the compressed low information feature X W2 Add to obtain the spatial feature output X W , and output X through the spatial feature W The spatial dimension of the new feature map is reconstructed.

[0009] According to one aspect of the above technical solution, the new feature map after the hole convolution processing is input into the channel reconstruction unit CRU, and the non-essential features in the new feature map are reorganized and filtered according to the channel features, specifically including: Output the spatial feature X W Divide into the first channel feature map with a ratio of αC and the second channel feature map with a ratio of (1−α)C according to the channel; The features in the first channel feature map are extracted by the stereo matching model GWCNet and the optical flow model PWCNet respectively, and added and merged into the first initial channel feature Y1; The second channel feature map is compressed by the optical flow model PWCNet, and a residual connection is performed with the second channel feature map before compression to obtain a second initial channel feature Y2; Performing pooling processing on the first initial channel feature Y1 and the second initial channel feature Y2 respectively to generate a first weight vector S1 and a second weight vector S2 respectively; Normalizing the first weight vector S1 and the second weight vector S2 respectively by using a Softmax function to obtain a first channel weight β1 and a second channel weight β2 respectively; Multiplying the first channel weight β1 and the first initial channel feature Y1 channel by channel to obtain a first intermediate feature, and multiplying the second channel weight β2 and the second initial channel feature Y2 channel by channel to obtain a second intermediate feature; The first intermediate feature and the second intermediate feature are added to obtain an output channel feature Y, and the output channel feature Y is used to reconstruct the channel dimension of the new feature map.

[0010] According to one aspect of the above technical solution, the new feature map processed by the DA_SCConv module is residually connected with the initial feature map to optimize the feature expression of the new feature map, specifically including: Based on the new feature map reconstructed in the spatial dimension and the new feature map reconstructed in the channel dimension, the new feature map processed by the DA_SCConv module is calculated by element-by-element multiplication; ; Among them, SRU (X1) is the new feature map after reconstruction of the spatial dimension, CRU (X2) is the new feature map after reconstruction of the channel dimension, and DA_SCConv (X) is the new feature map after processing by the DA_SCConv module; A residual connection is performed between the new feature map processed by the DA_SCConv module and the initial feature map to retain the original information of the new feature map processed by the DA_SCConv module, thereby optimizing the feature expression of the new feature map.

[0011] According to one aspect of the above technical solution, the updated feature map is input into the EMCA module for global average pooling and global maximum pooling, and adaptive convolution is used to capture the features of each channel in the Yolov8n network to complete the training of the Yolov8n network, specifically including: The updated feature maps are respectively input into the EMCA module for global average pooling processing and global maximum pooling processing to obtain the first processing results y avg and the second processing result y max ; Fusing the first processing result and the second processing result to obtain a pooling result y; ; The interaction between each channel feature and its adjacent channel features in the Yolov8n network is captured through adaptive convolution operations; ; ; Among them, w i is the interaction feature, σ is the sigmoid activation function, k is the size of the convolution kernel, and w j is the weight of the convolution kernel, γ and b are hyperparameters, C represents the number of channels, φ(C) is the mapping of C, |x| odd represents the odd number closest to X, i represents the sequence number of the current channel, and j represents the index of the adjacent channel; Performing a residual connection between the interactive features obtained by the adaptive convolution operation and the updated feature map to complete the training of the Yolov8n network; In the process of training the Yolov8n network using the training set, the validation set is used to adjust and optimize key hyperparameters in the Yolov8n network, wherein the key hyperparameters include learning rate, batch size, and weight decay coefficient.

[0012] The present invention also provides a forestry pest identification system, comprising: Acquisition module: used to acquire a data set related to forest pests, and divide the data set into a training set, a validation set, and a test set in proportion; Optimization module: used to establish a neural network model based on the Yolov8n network, and introduce the DA_SCConv module and the EMCA module into the Yolov8n network to optimize the Yolov8n network; wherein the optimized Yolov8n network includes a backbone structure, the DA_SCConv module, the SPPF module, and the EMCA module in sequence, and the DA_SCConv module includes a spatial adaptive attention unit, a spatial reconstruction unit SRU, a hole convolution, and a channel reconstruction unit CRU; Training module: used for inputting the training set into the optimized Yolov8n network for training; The training module is specifically used for: Input the training set as an initial feature map into the optimized Yolov8n network, and obtain a new feature map after the initial feature map is processed by the backbone structure; Inputting the new feature map into a spatial adaptive attention unit so that the features in the new feature map are focused on a preset area in space; Inputting the new feature map into the spatial reconstruction unit SRU to distinguish useful information from redundant information in the new feature map; The new feature map processed by the spatial reconstruction unit SRU is processed by a dilated convolution to increase the receptive field; Inputting the new feature map after the hole convolution processing into the channel reconstruction unit CRU, and reorganizing and filtering the unnecessary features in the new feature map according to the channel features; Performing a residual connection between the new feature map processed by the DA_SCConv module and the initial feature map to optimize the feature expression of the new feature map; Inputting the optimized new feature map into the SPPF module to obtain an updated feature map; Input the updated feature map into the EMCA module for global average pooling and global maximum pooling processing, and use adaptive convolution to capture the features of each channel in the Yolov8n network to complete the training of the Yolov8n network; Detection module: used to evaluate the trained neural network model using the test set to obtain a detection result.

[0013] The present invention also provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, the forestry pest identification method as described above is implemented.

[0014] Compared with the prior art, the present invention has the following beneficial effects: By improving the Yolov8n network, the DA_SCConv module and EMCA module are used to replace the original modules before and after the SPPF module of the backbone structure. Specifically, the DA_SCConv module includes a spatial adaptive attention unit, a spatial reconstruction unit SRU, a hole convolution and a channel reconstruction unit CRU. The spatial adaptive attention unit can automatically focus the features on the important areas in space. The spatial reconstruction unit SRU can analyze the information of the spatial dimension of the data set and separate the useful information from the redundant spatial features. The hole convolution can obtain a larger receptive field and capture richer contextual information. The channel reconstruction unit CRU can reorganize and filter according to the channel features to suppress unnecessary features and enhance the representation ability of the key channels. Therefore, the DA_SCConv module can perform CNN compression on the spatial redundancy and channel redundancy between the features in the data set to reduce the redundant features to reduce the model. The complexity of the model is reduced and the computing performance is improved; then the final output is residually connected with the initial input feature map. Through the residual connection, the original information in the input feature can be retained to form a richer and more concise feature expression; then the data set is processed by global average pooling and global maximum pooling through an extremely lightweight channel attention module EMCA, and the results of global average pooling and global maximum pooling are fused, and the interaction of each channel and its adjacent channels is captured using adaptive convolution; the features obtained by adaptive convolution are residually connected with the input features. The present invention adds an EMCA module to the neck structure (Neck) of the Yolov8n network to replace part of the original neck structure. Specifically, the EMCA module is also added to the upper layer processing at the output end; therefore, the EMCA module can effectively capture the long-distance dependency in the data set, reduce the impact of complex background on the detection results, and improve the accuracy and real-time performance of forestry pest detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is a flow chart of the forestry pest identification method in the first embodiment of the present invention; Figure 2 for Figure 1 Detailed flow chart of step S30 in FIG. Figure 3 This is a model framework diagram of the improved YOLOv8 in the first embodiment of the present invention; Figure 4 It is a structural block diagram of a forest pest identification system in a second embodiment of the present invention; Figure 5 is a structural block diagram of an electronic device in a third embodiment of the present invention; The following specific implementation manner will further illustrate the present invention in conjunction with the above-mentioned drawings. DETAILED DESCRIPTION

[0016] In order to facilitate the understanding of the present invention, the present invention will be described more fully below with reference to the relevant drawings. Several embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive.

[0017] It should be noted that when an element is referred to as being "fixed to" another element, it may be directly on the other element or there may be a central element. When an element is considered to be "connected to" another element, it may be directly connected to the other element or there may be a central element at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are for illustrative purposes only.

[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which the present invention belongs. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.

[0019] See also Figures 1 to 3 , shown is a forestry pest identification method in the first embodiment of the present invention, comprising the following steps: S10, obtaining a data set related to forestry pests, and dividing the data set into a training set, a validation set, and a test set in proportion; S20, establishing a neural network model based on the Yolov8n network, and introducing a DA_SCConv module and an EMCA module into the Yolov8n network to optimize the Yolov8n network; The optimized Yolov8n network includes a backbone structure, a DA_SCConv module, an SPPF module, and an EMCA module in sequence, wherein the DA_SCConv module includes a spatial adaptive attention unit, a spatial reconstruction unit SRU, a hole convolution, and a channel reconstruction unit CRU; S30, inputting the training set into the optimized Yolov8n network for training; S40, using the test set to evaluate the trained neural network model to obtain a detection result; The inputting the training set into the optimized Yolov8n network for training specifically includes: S31, inputting the training set as an initial feature map into the optimized Yolov8n network, and obtaining a new feature map after the initial feature map is processed by the backbone structure; S32, inputting the new feature map into a spatial adaptive attention unit so that the features in the new feature map are focused on a preset area in space; S33, inputting the new feature map into the spatial reconstruction unit SRU to distinguish useful information from redundant information in the new feature map; S34, processing the new feature map processed by the spatial reconstruction unit SRU through a dilated convolution to increase the receptive field; S35, inputting the new feature map after the hole convolution processing into the channel reconstruction unit CRU, and reorganizing and filtering the unnecessary features in the new feature map according to the channel features; S36, performing a residual connection between the new feature map processed by the DA_SCConv module and the initial feature map to optimize the feature expression of the new feature map; S37, inputting the optimized new feature map into the SPPF module to obtain an updated feature map; S38, input the updated feature map into the EMCA module for global average pooling and global maximum pooling processing, and use adaptive convolution to capture the features of each channel in the Yolov8n network to complete the training of the Yolov8n network.

[0020] It can be understood that the present invention improves the Yolov8n network by using the DA_SCConv module and the EMCA module to replace the original modules before and after the SPPF module of the backbone structure. Specifically, the DA_SCConv module includes a spatial adaptive attention unit, a spatial reconstruction unit SRU, a hole convolution and a channel reconstruction unit CRU. The spatial adaptive attention unit can automatically focus the features on the important areas in space. The spatial reconstruction unit SRU can analyze the information of the spatial dimension of the data set and separate the useful information from the redundant spatial features. The hole convolution can obtain a larger receptive field and capture richer contextual information. The channel reconstruction unit CRU can reorganize and filter according to the channel features to suppress unnecessary features and enhance the representation ability of the key channels. Therefore, the DA_SCConv module can perform CNN compression on the spatial redundancy and channel redundancy between the features in the data set to reduce redundant features. To reduce the complexity of the model and improve the computing performance; then the final output is residually connected with the initial input feature map. Through the residual connection, the original information in the input feature can be retained to form a richer and more concise feature expression; then the data set is processed by global average pooling and global maximum pooling through an extremely lightweight channel attention module EMCA, and the results of global average pooling and global maximum pooling are fused, and adaptive convolution is used to capture the interaction of each channel and its adjacent channels; the features obtained by adaptive convolution are residually connected with the input features. The present invention adds an EMCA module to the neck structure (Neck) of the Yolov8n network to replace part of the original neck structure. Specifically, the EMCA module is also added to the upper layer of the output end; therefore, the EMCA module can effectively capture the long-distance dependency in the data set, reduce the impact of complex background on the detection results, and improve the accuracy and real-time performance of forestry pest detection.

[0021] Specifically, in this embodiment, the backbone structure is Conv-Conv-C2f-Conv-C2f-Conv-C2f-Conv-C2f; The step S31 specifically includes: The initial feature map is sequentially processed by the backbone structure formed by the interlacing of the Conv and the C2f, and an intermediate feature map is obtained after the fourth C2f processing; The intermediate feature map is passed through a 1×1 convolution layer to obtain a new feature map, where the new feature map ∈ (W×H×C) where Conv represents convolution and C2f represents a feature fusion module.

[0022] It can be understood that the role of the backbone structure is to perform convolution operations and feature fusion on the training set, obtain the intermediate feature map (20*20*1024) after preprocessing, and then obtain the new feature map through a 1×1 convolution layer. The new feature map has three dimensions (W×H×C), W represents width, H represents height, and C represents the number of channels. The backbone structure is an existing network model and will not be described in detail here.

[0023] Furthermore, the step S32 specifically includes: Inputting the new feature map into a spatially adaptive attention unit; Flattening the new feature map by the spatial adaptive attention unit to divide each spatial position into an independent token; The token is processed by Transformer and marked as X flattened , and add the position code P to get the position representation X pos ; ; The X pos Input to Transformer encoder for self-attention calculation; ; Among them, Q represents the query vector, K represents the key vector, C represents the value vector, and T represents the transpose; Get the attention matrix weight V and output feature representation A spatital ; ; The feature representation is restored to the spatial dimension W×H×C.

[0024] It can be understood that through the above steps, the spatial adaptive attention unit can automatically focus the features on spatially important areas to prepare for subsequent processing.

[0025] Furthermore, the step S33 specifically includes: Inputting the new feature map into the spatial reconstruction unit SRU, and normalizing the new feature map by group normalization; The weight value of the standardized new feature map is mapped to the interval (0, 1) through the Sigmoid activation function; A weight threshold T is set. In the new feature graph, feature information with a weight value higher than the weight threshold is regarded as useful information and marked as high feature information W1, and feature information with a weight value lower than the weight threshold is regarded as redundant information and marked as low feature information W2; Using the cross-reconstruction operation method, the high feature information W1 is multiplied by the new feature map to obtain the first feature X1 W , and multiplying the low feature information W2 by the new feature map to obtain the second feature X2 W ; The first feature X1 W Perform feature splitting to obtain the first sub-feature X 11 W and the second sub-feature X 12 W , and the second feature X2 W Perform feature splitting to obtain the third sub-feature X 21 W and the fourth sub-feature X 22 W ; The first sub-feature X 11 W and the fourth sub-feature X 22 W Add together to get the enhanced high information feature X W1 , and the second sub-feature X 12 W and the third sub-feature X 21 W Add to get the compressed low information feature X W2 ; The enhanced high information feature X W1 and the compressed low information feature X W2 Add to obtain the spatial feature output X W , and output X through the spatial feature W The spatial dimension of the new feature map is reconstructed.

[0026] It can be understood that the new feature map can be grouped and normalized by the spatial reconstruction unit SRU, and then the new feature map can be standardized to form a more consistent distribution between different feature groups. Then the normalized new feature map is mapped to a weight value between 0 and 1 through the Sigmoid activation function, and then a threshold T is set to divide the new feature map into high feature information W1 and low feature information W2. When the weight value of a certain position is higher than T, it is considered that the feature information of the position is rich and belongs to the high feature information W1, and the weight value is lower than T, it is considered that the feature of the position is mainly redundant information and belongs to the low feature information W2. Then, the cross reconstruction operation is used to obtain the enhanced high information feature X W1 and compressed low-information feature X W2 Finally, through the concatenation operation, the enhanced high information feature X W1 And the compressed low information feature XW2 Add to obtain the spatial feature output X W , spatial feature output X W It can reflect the useful information and redundant information in the new feature map. The spatial feature output X W The spatial dimension of the new feature map is reconstructed.

[0027] Furthermore, the step S35 specifically includes: Output the spatial feature X W Divide into the first channel feature map with a ratio of αC and the second channel feature map with a ratio of (1−α)C according to the channel; The features in the first channel feature map are extracted by the stereo matching model GWCNet and the optical flow model PWCNet respectively, and added and merged into the first initial channel feature Y1; The second channel feature map is compressed by the optical flow model PWCNet, and a residual connection is performed with the second channel feature map before compression to obtain a second initial channel feature Y2; Performing pooling processing on the first initial channel feature Y1 and the second initial channel feature Y2 respectively to generate a first weight vector S1 and a second weight vector S2 respectively; Normalizing the first weight vector S1 and the second weight vector S2 respectively by using a Softmax function to obtain a first channel weight β1 and a second channel weight β2 respectively; Multiplying the first channel weight β1 and the first initial channel feature Y1 channel by channel to obtain a first intermediate feature, and multiplying the second channel weight β2 and the second initial channel feature Y2 channel by channel to obtain a second intermediate feature; The first intermediate feature and the second intermediate feature are added to obtain an output channel feature Y, and the output channel feature Y is used to reconstruct the channel dimension of the new feature map.

[0028] It can be understood that the channel reconstruction unit CRU can convert the X output by SRU into WDivide into two feature maps according to the channel, perform partition processing, and then use the stereo matching model GWCNet and the optical flow model PWCNet to extract features in the first channel feature map, and add and merge them into the first initial channel feature Y1, so that the training set is trained in the stereo space, and then compress the second channel feature map only through the optical flow model, and perform residual connection with the second channel feature map before compression to obtain the second initial channel feature Y2. By extracting the two initial channel features, the key data can be retained, and then pooling the two can reduce the dimension of the data, greatly reducing the amount of data calculation, and at the same time obtain two weight vectors, and then normalize the two weight vectors to obtain the weights of the two channels, and multiply the two channel weights by the initial channel channel by channel to obtain the intermediate features in the initial channel, and finally add the two intermediate features to obtain the output channel feature Y, which can be used to reorganize and filter the non-essential features in the new feature map, and enhance the representation ability of the key channel, thereby obtaining a new feature map reconstructed by the channel dimension.

[0029] Furthermore, the step S36 specifically includes: Based on the new feature map reconstructed in the spatial dimension and the new feature map reconstructed in the channel dimension, the new feature map processed by the DA_SCConv module is calculated by element-by-element multiplication; ; Among them, SRU (X1) is the new feature map after reconstruction of the spatial dimension, CRU (X2) is the new feature map after reconstruction of the channel dimension, and DA_SCConv (X) is the new feature map after processing by the DA_SCConv module; A residual connection is performed between the new feature map processed by the DA_SCConv module and the initial feature map to retain the original information of the new feature map processed by the DA_SCConv module, thereby optimizing the feature expression of the new feature map.

[0030] It can be understood that the new feature map reconstructed by the spatial dimension and the new feature map reconstructed by the channel dimension can be multiplied element by element to obtain the new feature map processed by the DA_SCConv module. At this time, the training of the DA_SCConv module is completed. Then the new feature map processed by the DA_SCConv module is residually connected with the initial feature map, which can solve the problems of gradient disappearance and gradient explosion. At the same time, it can also help the model converge faster and improve the stability of the network, thereby retaining the original information on the new feature map, and optimizing the feature expression of the new feature map.

[0031] Then, the optimized new feature map is processed by the existing SPPF module to obtain an updated feature map.

[0032] Furthermore, the step S38 specifically includes: The updated feature maps are respectively input into the EMCA module for global average pooling processing and global maximum pooling processing to obtain the first processing results y avg and the second processing result y max ; Fusing the first processing result and the second processing result to obtain a pooling result y; ; The interaction between each channel feature and its adjacent channel features in the Yolov8n network is captured through adaptive convolution operations; ; ; Among them, w i is the interaction feature, σ is the sigmoid activation function, k is the size of the convolution kernel, and w j is the weight of the convolution kernel, γ and b are hyperparameters, C represents the number of channels, φ(C) is the mapping of C, |x| odd represents the odd number closest to X, i represents the sequence number of the current channel, and j represents the index of the adjacent channel; The interactive features obtained by the adaptive convolution operation are residually connected with the updated feature map to complete the training of the Yolov8n network.

[0033] It is understandable that the last step of training is to pool the updated feature map through the EMCA module to reduce the dimension of the data, which enables the model to better extract important features and reduce redundant information. Then, through the adaptive convolution operation, the standard convolution can be simply and effectively modified, and the size of the convolution kernel can be changed according to the learnable local pixel features, thereby optimizing the model and capturing the interaction between each channel and its adjacent channels. Useful channel features can be retained. The purpose of the residual connection is to solve the problems of gradient disappearance and gradient explosion, and it can also help the model converge faster and improve the stability of the network. After that, the results can be output, thus completing the training of the model.

[0034] Preferably, in the process of training the Yolov8n network using the training set, the validation set is used to adjust and optimize the key hyperparameters in the Yolov8n network to improve the convergence speed and detection accuracy of the model. The key hyperparameters include learning rate, batch size, and weight decay coefficient. At the same time, a cosine annealing learning rate strategy is adopted to adjust the learning rate at different stages of training to better adapt to the characteristics of the data set and the training process of the model.

[0035] Furthermore, a variety of data enhancement techniques are applied to the forestry pest dataset, including random cropping, rotation, scaling, color jittering, blurring, etc., so that the model can adapt to complex practical application scenarios and improve the robustness of the model to various pest types.

[0036] Furthermore, after the training is completed, the model is evaluated using the test set, and the five models of Yolov4, YOLOv4_tiny, SSD_MobileNet, Yolov4_MF and Yolov8 are used as benchmarks to compare with Yolov8_SCM, as shown in the following table:

[0037] It should be noted that Figure 3 In the figure, If shortcut means if there is a jump connection, detect means the detection head of the model, Bottleneck means the bottleneck layer, add means the fusion operation, Bn means batch normalization, SiLU means the activation function, Split means feature map segmentation, MaxPool means maximum pooling, Bbox.Loss means bounding box loss, and Cls.Loss means classification loss.

[0038] In summary, the forest pest identification method in the above embodiments of the present invention can perform CNN compression on the spatial redundancy and channel redundancy between the features in the data set to reduce redundant features to reduce the complexity of the model and improve computing performance. It can also effectively capture the long-distance dependencies in the data set, reduce the impact of complex background on the detection results, and improve the accuracy and real-time performance of forest pest detection. Please refer to Figure 4 , shown is a forest pest identification system in a second embodiment of the present invention, comprising: Acquisition module 11: used to acquire a data set related to forestry pests, and divide the data set into a training set, a validation set, and a test set in proportion; Optimization module 12: used to establish a neural network model based on the Yolov8n network, and introduce the DA_SCConv module and the EMCA module into the Yolov8n network to optimize the Yolov8n network; wherein the optimized Yolov8n network includes a backbone structure, the DA_SCConv module, the SPPF module, and the EMCA module in sequence, and the DA_SCConv module includes a spatial adaptive attention unit, a spatial reconstruction unit SRU, a hole convolution and a channel reconstruction unit CRU; Training module 13: used for inputting the training set into the optimized Yolov8n network for training; The training module 13 is specifically used for: Input the training set as an initial feature map into the optimized Yolov8n network, and obtain a new feature map after the initial feature map is processed by the backbone structure; Inputting the new feature map into a spatial adaptive attention unit so that the features in the new feature map are focused on a preset area in space; Inputting the new feature map into the spatial reconstruction unit SRU to distinguish useful information from redundant information in the new feature map; The new feature map processed by the spatial reconstruction unit SRU is processed by a dilated convolution to increase the receptive field; Inputting the new feature map after the hole convolution processing into the channel reconstruction unit CRU, and reorganizing and filtering the unnecessary features in the new feature map according to the channel features; Performing a residual connection between the new feature map processed by the DA_SCConv module and the initial feature map to optimize the feature expression of the new feature map; Inputting the optimized new feature map into the SPPF module to obtain an updated feature map; Input the updated feature map into the EMCA module for global average pooling and global maximum pooling processing, and use adaptive convolution to capture the features of each channel in the Yolov8n network to complete the training of the Yolov8n network; Detection module 14: used to evaluate the trained neural network model using the test set to obtain a detection result.

[0039] The present invention also provides an electronic device, see Figure 5 , shown is an electronic device in the third embodiment of the present invention, including a memory 10, a processor 20, and a computer program 30 stored in the memory 10 and executable on the processor 20, and the processor 20 implements the above-mentioned forestry pest identification method when executing the computer program 30.

[0040] Among them, the memory 10 includes at least one type of storage medium, and the storage medium includes a flash memory, a hard disk, a multimedia card, a card-type memory (for example, an SD or DX memory, etc.), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 10 can be an internal storage unit of an electronic device, such as a hard disk of the electronic device. In other embodiments, the memory 10 can also be an external storage device, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. Further, the memory 10 can also include both an internal storage unit of the electronic device and an external storage device. The memory 10 can not only be used to store application software and various types of data installed in the electronic device, but also can be used to temporarily store data that has been output or is to be output.

[0041] Among them, in some embodiments, the processor 20 can be an electronic control unit (Electronic Control Unit, abbreviated as ECU, also known as a vehicle computer), a central processing unit (Central Processing Unit, CPU), a controller, a microcontroller, a microprocessor or other data processing chip, used to run the program code stored in the memory 10 or process data, such as executing access restriction programs, etc.

[0042] It should be pointed out that Figure 5 The structure shown does not constitute a limitation on the electronic device. In other embodiments, the electronic device may include fewer or more components than those shown in the figure, or combine certain components, or arrange the components differently.

[0043] The embodiment of the present invention further provides a readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the forestry pest identification method as described above is implemented.

[0044] Those skilled in the art will appreciate that the logic and / or steps represented in the flowchart or otherwise described herein, for example, may be considered as an ordered list of executable instructions for implementing logical functions, and may be specifically implemented in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For purposes of this specification, "computer-readable medium" may be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.

[0045] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or, if necessary, processing in another suitable manner, and then stored in a computer memory.

[0046] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or a combination thereof: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0047] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0048] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.

Claims

1. A method for identifying forest pests, characterized in that: The steps include: Acquire a data set related to forest pests, and divide the data set into a training set, a validation set, and a test set in proportion; A neural network model is established based on the Yolov8n network, and a DA_SCConv module and an EMCA module are introduced into the Yolov8n network to optimize the Yolov8n network; The optimized Yolov8n network includes a backbone structure, a DA_SCConv module, an SPPF module, and an EMCA module in sequence, wherein the DA_SCConv module includes a spatial adaptive attention unit, a spatial reconstruction unit SRU, a hole convolution, and a channel reconstruction unit CRU; Input the training set into the optimized Yolov8n network for training; Using the test set to evaluate the trained neural network model to obtain a detection result; The inputting the training set into the optimized Yolov8n network for training specifically includes: Input the training set as an initial feature map into the optimized Yolov8n network, and obtain a new feature map after the initial feature map is processed by the backbone structure; Inputting the new feature map into a spatial adaptive attention unit so that the features in the new feature map are focused on a preset area in space; Inputting the new feature map into the spatial reconstruction unit SRU to distinguish useful information from redundant information in the new feature map; The new feature map processed by the spatial reconstruction unit SRU is processed by a dilated convolution to increase the receptive field; Inputting the new feature map after the hole convolution processing into the channel reconstruction unit CRU, and reorganizing and filtering the unnecessary features in the new feature map according to the channel features; Performing a residual connection between the new feature map processed by the DA_SCConv module and the initial feature map to optimize the feature expression of the new feature map; Inputting the optimized new feature map into the SPPF module to obtain an updated feature map; The updated feature map is input into the EMCA module for global average pooling and global maximum pooling processing, and adaptive convolution is used to capture the features of each channel in the Yolov8n network to complete the training of the Yolov8n network.

2. The forest pest identification method according to claim 1, characterized in that: The backbone structure is Conv-Conv-C2f-Conv-C2f-Conv-C2f-Conv-C2f, where Conv represents convolution and C2f represents a feature fusion module.

3. The forest pest identification method according to claim 2, characterized in that: The initial feature map is processed by the backbone structure to obtain a new feature map, which specifically includes: The initial feature map is sequentially processed by the backbone structure formed by the interlacing of the Conv and the C2f, and an intermediate feature map is obtained after the fourth C2f processing; The intermediate feature map is passed through a 1×1 convolution layer to obtain a new feature map, where the new feature map ∈ (W×H×C).

4. The forest pest identification method according to claim 3, characterized in that: The step of inputting the new feature map into a spatial adaptive attention unit so that the features in the new feature map are focused on a preset area in space specifically includes: Inputting the new feature map into a spatially adaptive attention unit; Flattening the new feature map by the spatial adaptive attention unit to divide each spatial position into an independent token; The token is processed by Transformer and marked as X flattened , and add the position code P to get the position representation X pos ; ; The X pos Input to Transformer encoder for self-attention calculation; ; Among them, Q represents the query vector, K represents the key vector, C represents the value vector, and T represents the transpose; Get the attention matrix weight V and output feature representation A spatital ; ; The feature representation is restored to the spatial dimension W×H×C.

5. The forest pest identification method according to claim 4, characterized in that: The step of inputting the new feature map into the spatial reconstruction unit SRU to distinguish useful information from redundant information in the new feature map specifically includes: Inputting the new feature map into the spatial reconstruction unit SRU, and normalizing the new feature map by group normalization; The weight value of the standardized new feature map is mapped to the interval (0, 1) through the Sigmoid activation function; A weight threshold T is set. In the new feature graph, feature information with a weight value higher than the weight threshold is regarded as useful information and marked as high feature information W1, and feature information with a weight value lower than the weight threshold is regarded as redundant information and marked as low feature information W2; Using the cross-reconstruction operation method, the high feature information W1 is multiplied by the new feature map to obtain the first feature X1 W , and multiplying the low feature information W2 with the new feature map to obtain the second feature X2 W ; The first feature X1 W Perform feature splitting to obtain the first sub-feature X 11 W and the second sub-feature X 12 W , and the second feature X2 W Perform feature splitting to obtain the third sub-feature X 21 W and the fourth sub-feature X 22 W ; The first sub-feature X 11 W and the fourth sub-feature X 22 W Add together to get the enhanced high information feature X W1 , and the second sub-feature X 12 W and the third sub-feature X 21 W Add to get the compressed low information feature X W2 ; The enhanced high information feature X W1 and the compressed low information feature X W2 Add to obtain the spatial feature output X W , and output X through the spatial feature W The spatial dimension of the new feature map is reconstructed.

6. The forest pest identification method according to claim 5, characterized in that: The new feature map after the hole convolution processing is input into the channel reconstruction unit CRU, and the non-essential features in the new feature map are reorganized and filtered according to the channel features, specifically including: Output the spatial feature X W Divide into the first channel feature map with a ratio of αC and the second channel feature map with a ratio of (1−α)C according to the channel; The features in the first channel feature map are extracted by the stereo matching model GWCNet and the optical flow model PWCNet respectively, and added and merged into the first initial channel feature Y1; The second channel feature map is compressed by the optical flow model PWCNet, and a residual connection is performed with the second channel feature map before compression to obtain a second initial channel feature Y2; Performing pooling processing on the first initial channel feature Y1 and the second initial channel feature Y2 respectively to generate a first weight vector S1 and a second weight vector S2 respectively; Normalizing the first weight vector S1 and the second weight vector S2 respectively by using a Softmax function to obtain a first channel weight β1 and a second channel weight β2 respectively; Multiplying the first channel weight β1 and the first initial channel feature Y1 channel by channel to obtain a first intermediate feature, and multiplying the second channel weight β2 and the second initial channel feature Y2 channel by channel to obtain a second intermediate feature; The first intermediate feature and the second intermediate feature are added to obtain an output channel feature Y, and the output channel feature Y is used to reconstruct the channel dimension of the new feature map.

7. The forest pest identification method according to claim 6, characterized in that: The performing residual connection between the new feature map processed by the DA_SCConv module and the initial feature map to optimize the feature expression of the new feature map specifically includes: Based on the new feature map reconstructed in the spatial dimension and the new feature map reconstructed in the channel dimension, the new feature map processed by the DA_SCConv module is calculated by element-by-element multiplication; ; Among them, SRU (X1) is the new feature map after reconstruction of the spatial dimension, CRU (X2) is the new feature map after reconstruction of the channel dimension, and DA_SCConv (X) is the new feature map after processing by the DA_SCConv module; A residual connection is performed between the new feature map processed by the DA_SCConv module and the initial feature map to retain the original information of the new feature map processed by the DA_SCConv module, thereby optimizing the feature expression of the new feature map.

8. The forest pest identification method according to claim 7, characterized in that: The updated feature map is input into the EMCA module for global average pooling and global maximum pooling, and adaptive convolution is used to capture the features of each channel in the Yolov8n network to complete the training of the Yolov8n network, specifically including: The updated feature maps are respectively input into the EMCA module for global average pooling processing and global maximum pooling processing to obtain the first processing results y avg and the second processing result y max ; Fusing the first processing result and the second processing result to obtain a pooling result y; ; The interaction between each channel feature and its adjacent channel features in the Yolov8n network is captured through adaptive convolution operations; ; ; Among them, w i is the interaction feature, σ is the sigmoid activation function, k is the size of the convolution kernel, and w j is the weight of the convolution kernel, γ and b are hyperparameters, C represents the number of channels, φ(C) is the mapping of C, |x| odd represents the odd number closest to X, i represents the sequence number of the current channel, and j represents the index of the adjacent channel; Performing a residual connection between the interactive features obtained by the adaptive convolution operation and the updated feature map to complete the training of the Yolov8n network; In the process of training the Yolov8n network using the training set, the validation set is used to adjust and optimize key hyperparameters in the Yolov8n network, wherein the key hyperparameters include learning rate, batch size, and weight decay coefficient.

9. A forestry pest identification system, characterized in that: include: Acquisition module: used to acquire a data set related to forest pests, and divide the data set into a training set, a validation set, and a test set in proportion; Optimization module: used to establish a neural network model based on the Yolov8n network, and introduce the DA_SCConv module and the EMCA module into the Yolov8n network to optimize the Yolov8n network; wherein the optimized Yolov8n network includes a backbone structure, the DA_SCConv module, the SPPF module, and the EMCA module in sequence, and the DA_SCConv module includes a spatial adaptive attention unit, a spatial reconstruction unit SRU, a hole convolution, and a channel reconstruction unit CRU; Training module: used for inputting the training set into the optimized Yolov8n network for training; The training module is specifically used for: Input the training set as an initial feature map into the optimized Yolov8n network, and obtain a new feature map after the initial feature map is processed by the backbone structure; Inputting the new feature map into a spatial adaptive attention unit so that the features in the new feature map are focused on a preset area in space; Inputting the new feature map into the spatial reconstruction unit SRU to distinguish useful information from redundant information in the new feature map; The new feature map processed by the spatial reconstruction unit SRU is processed by a dilated convolution to increase the receptive field; Inputting the new feature map after the hole convolution processing into the channel reconstruction unit CRU, and reorganizing and filtering the unnecessary features in the new feature map according to the channel features; Performing a residual connection between the new feature map processed by the DA_SCConv module and the initial feature map to optimize the feature expression of the new feature map; Inputting the optimized new feature map into the SPPF module to obtain an updated feature map; Input the updated feature map into the EMCA module for global average pooling and global maximum pooling processing, and use adaptive convolution to capture the features of each channel in the Yolov8n network to complete the training of the Yolov8n network; Detection module: used to evaluate the trained neural network model using the test set to obtain a detection result.

10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the forestry pest identification method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Crop disease identification method, system and device and storage medium

    CN115965874A

  • YOLOv8 vehicle identification method, system and device and storage medium

    CN118470494A

  • Track foreign matter detection method based on YOLO algorithm

    CN119478380A

  • High-resolution remote sensing image double-branch segmentation method

    CN119494961A

  • Method for image motion deblurring, apparatus, electronic device and medium therefor

    US20240404025A1