Dense tiny pest image detection method based on multi-kernel attention network
Through the method based on multi-core attention network, the problem of difficulty in detecting dense micro pest images in the prior art is solved, and high-precision and high-efficiency pest detection is achieved, which is suitable for real-time detection needs in the agricultural field.
Patent Information
- Application Number
- CN202510104704.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-23
AI Technical Summary
The prior art is difficult to effectively detect dense micro-pest images, especially in the agricultural field. It is difficult for models to obtain sufficient features for accurate identification, and the computing resources are consumed, making it difficult to meet the real-time detection needs.
The dense micro-pest image detection method based on multi-core attention network is adopted, and the data set is generated through the dense pest area focus extraction module, multi-scale features are extracted using cross-stage local networks, and a small-scale pest dual-channel multi-core feature extraction model is constructed for training to achieve accurate pest detection.
It significantly improves the detection accuracy of dense micro-pest images, reduces the consumption of computing resources, improves the detection speed and accuracy, and can effectively identify micro-pests in complex agricultural environments.
Smart Images

Figure CN120014554A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of pest image recognition, and in particular to a method for detecting dense tiny pest images based on a multi-core attention network. Background Art
[0002] In modern agriculture, pests such as aphids are serious harm to crops. They suck plant sap, spread viruses, and secrete honeydew, causing the leaves of crops to turn yellow or even the whole plant to die, seriously affecting the yield and quality of crops. The rise of convolutional neural networks has promoted the advancement of target detection technology, resulting in two-stage and one-stage detection methods, but there are still many challenges in the application of agricultural pest detection.
[0003] Especially in the recognition of dense micro-images, the existing methods have many drawbacks. First of all, from the perspective of the image's own characteristics, the resolution of dense micro-images such as agricultural pests is usually low, and the small number of pixels makes a large number of key details lost. For example, the limb texture and morphological characteristics of aphids are difficult to present clearly, which makes it difficult for the model to obtain sufficient and effective features for accurate recognition. At the same time, in the images taken in the field, the scale of pests varies greatly, and due to the dense distribution, the proportions between different pests are imbalanced. The model must not only deal with extremely small individuals, but also take into account large-scale targets. This places extremely high demands on feature capture capabilities, and a little carelessness will lead to loss of one while focusing on the other. In addition, pests are densely arranged on the plants, and mutual occlusion occurs frequently. The information of the occluded part cannot be obtained by the model, which further increases the difficulty of recognition. The receptive field of some existing classic deep learning models is difficult to adapt to dense micro-images such as agricultural pests. If the receptive field is too small, it cannot capture enough contextual information, resulting in inaccurate recognition of micro-pest individuals; if the receptive field is too large, it will introduce too much irrelevant background information, interfering with the model's judgment. In addition, in order to effectively extract the features of densely packed tiny pest images, a larger-scale model may be required, but this will undoubtedly greatly increase the consumption of computing resources and reduce the computing speed. In agricultural scenarios where real-time pest detection is required, it is difficult to meet actual application needs. Summary of the invention
[0004] The purpose of the present invention is to solve the defect that it is difficult to detect dense tiny pest images in the prior art, and to provide a dense tiny pest image detection method based on a multi-core attention network to solve the above problem.
[0005] In order to achieve the above object, the technical solution of the present invention is as follows:
[0006] A dense tiny pest image detection method based on a multi-core attention network includes the following steps:
[0007] Generation of dense tiny pest image dataset: Using the dense pest area focus extraction module, the dense pest area is focused and extracted to generate a dense tiny pest image dataset;
[0008] Extraction of multi-scale features of pest images: Extract multi-scale features by a cross-stage local network;
[0009] Construct a dual-channel multi-core feature extraction model for small-scale pests;
[0010] Training of small-scale pest dual-channel multi-core feature extraction model: Input the dense tiny pest image dataset into the small-scale pest dual-channel multi-core feature extraction model for training;
[0011] Acquisition of dense tiny pest image detection results: Obtain the pest image to be detected and input it into the cross-stage local network to obtain the multi-scale feature map of the pest image, and input the multi-scale feature map of the pest image into the trained small-scale pest dual-channel multi-core feature extraction model to obtain the pest image detection result.
[0012] The generation of the dense tiny pest image dataset comprises the following steps:
[0013] Obtain images of wheat infested by pests, annotate and preprocess the images, and construct a pest dataset;
[0014] Construct a focused extraction module for dense pest areas;
[0015] Constructing a dense tiny pest image dataset: Input the pest dataset into the dense pest area focus extraction module to extract dense pest area images, and mix the dense pest area images with the pest dataset to form a dense tiny pest image dataset.
[0016] The extraction of multi-scale features of the pest image comprises the following steps:
[0017] Construct a cross-stage local network, which includes a focusing structure, a convolutional normalized activation structure, and a cross-stage local layer structure. The cross-stage local layer structure decomposes the residual block stack and introduces a spatial pyramid pooling structure before the last cross-stage local layer structure to extract features.
[0018] The input image obtained from the dense tiny pest image dataset first enters the focusing structure for slicing operation to generate multiple low-resolution feature maps and stack them to quadruple the number of input channels. Then, the multi-channel feature maps are spliced, combined and stacked together, and cross-processed by the convolution normalization activation structure and the cross-stage local layer structure to enhance the small-scale pest features, and obtain the pest small-scale feature map.
[0019] The focusing structure extracts the pixel value of every other pixel from the input image, generating four independent feature layers in this way. These four feature layers are then stacked together to integrate the width and height information of the image into the channel dimension, thereby increasing the number of input channels by four times.
[0020] The convolution normalization activation structure adds a batch normalization layer and a SiLU activation function after the convolution layer. The formula of the SiLU activation function f(x) is:
[0021] f(x) = x·sigmoid(x)
[0022] Among them, the formula of the sigmoid function is expressed as:
[0023]
[0024] Among them, x represents the input value, and e represents the natural constant;
[0025] The cross-stage local layer structure is divided into two parts: one part directly performs forward propagation to retain relatively original feature information, including basic features at different scales; the other part undergoes multiple repeated 1×1 basic convolutions and 3×3 depthwise separable convolution operations to further extract and refine features and mine more representative features at different scales. These two parts of information are merged at the end of the network, so that features at different levels and scales can be integrated to enrich the diversity of features, form a feature map containing multi-scale information, and obtain a multi-scale feature map of pest images.
[0026] The construction of a small-scale pest dual-channel multi-core feature extraction model comprises the following steps:
[0027] The small-scale pest dual-channel multi-core feature extraction model is set up to include a dual-channel feature pyramid module and a multi-core attention network module;
[0028] A dual-channel feature pyramid module is set up, which consists of a top-down information flow and a bottom-up information flow;
[0029] Among them, the top-down information flow starts from the higher layers of the network and gradually transfers information to the lower layers. First, the higher-level feature maps are upsampled and then fused with the feature maps of the middle or bottom layers to enhance the semantic information in the lower-level feature maps. The formula is expressed as:
[0030] P5=Conv(f3)
[0031]
[0032] Among them, f i(i=3,2,1) represent the three feature maps of top-down information flow input, P i (i=5, 4, 3) represent the three feature maps output by the top-down information flow, Represents a fusion operation, Conv represents a convolution operation, CSPLayer represents a cross-stage local layer structure, and UpSample represents an upsampling operation;
[0033] Bottom-up information flow: Starting from the lower-level feature map, the information is gradually passed to the higher-level feature map. In the bottom-up information flow, the lower-level feature map is downsampled and fused with the higher-level feature map. The formula is expressed as:
[0034]
[0035] Among them, P i (i=5, 4, 3) represent the three feature maps output by the top-down information flow, They represent the final output feature maps obtained from the bottom-up information flow, represents a fusion operation, CSPLayer represents a cross-stage local layer structure, and Downsample represents a downsampling operation;
[0036] Construct a multi-core attention network module, which includes an aggregation module, an extraction module, and a reconstruction module.
[0037] Set the aggregation module, which uses a standard depth-wise separable convolution with a kernel size of k×k, where k is the set kernel size;
[0038] Set up the extraction module, which uses a multi-branch deep strip convolution layer. In each branch, a 1×j depthwise separable strip convolution and a j×1 depthwise separable strip convolution are used to simulate a j×j standard depthwise separable convolution, where the kernel size j in different branches is different;
[0039] Set the reconstruction module, which uses a 1×1 standard convolution to reconstruct the relationship between different channels;
[0040] First, the feature information from the dual-channel feature pyramid network is output as a local information feature map through the aggregation module. Secondly, the multi-scale features of the local information feature map are extracted through the extraction module to obtain multi-scale feature maps of different channels. Finally, the multi-scale feature map is reconstructed through the reconstruction module to reconstruct the relationship between different channels and output the final feature map.
[0041] The multi-core attention network module formula is expressed as:
[0042]
[0043] Among them, P refers to the input feature, Att and Out represent the attention map and output results respectively, Conv 1×1 Represents a 1×1 convolution operation, DW-Conv is the abbreviation of depthwise separable convolution, Scale i Where i takes the value of 0, 1, 2, or 3, and Scale i When 1, 2, or 3 is selected, it indicates branch options, and Scale 0 indicates identity mapping.
[0044] The training of the small-scale pest dual-channel multi-core feature extraction model includes the following steps:
[0045] Set the bounding box regression loss function:
[0046] The bounding box regression loss function L of the small-scale pest dual-channel multi-core feature extraction model is defined as:
[0047]
[0048] Among them, IoU represents the intersection-over-union ratio of the predicted bounding box and the true bounding box area, b and b gt denote the center points of the predicted bounding box and the true bounding box, respectively, and ρ 2 (b,b gt ) is the squared Euclidean distance between the center points of the predicted bounding box and the true bounding box, c is the diagonal length of the minimum enclosing rectangle of the predicted bounding box and the true bounding box, α is the weight coefficient used to balance the impact of the aspect ratio, and v is a measure of the consistency of the aspect ratio between the predicted bounding box and the true bounding box. The calculation formula is:
[0049]
[0050] w and h are the width and height of the predicted bounding box, w gt and h gt is the width and height of the true bounding box, and the formula for the weight coefficient α is:
[0051]
[0052] The formula for intersection over union (IoU) is:
[0053]
[0054] bbox is the area size of the predicted bounding box, bbox gt is the area of the true bounding box, ∩ represents the intersection operation, and ∪ represents the union operation;
[0055] The multi-scale feature map of the pest image from the cross-stage local layer structure is input into the dual-channel feature pyramid module for training.
[0056] According to the top-down information flow, first, the multi-scale feature map of the pest image at the top layer is processed by a two-dimensional convolution to obtain the feature map P5, P5 is then upsampled to expand the image size of the feature map, and then it is spliced with the multi-scale feature map of the pest image at the middle layer to obtain the feature map P4 through a two-dimensional convolution operation through a cross-stage local layer structure, P4 is then upsampled to expand the image size of the feature map, and then it is spliced with the multi-scale feature map of the pest image at the bottom layer to obtain the feature map P3 through a cross-stage local layer structure. According to the bottom-up information flow, first, the feature map P3 at the bottom layer is directly output to obtain the feature map At the same time, the feature map P3 is downsampled and spliced with the feature map P4 through the cross-stage local layer structure to obtain the output feature map Then repeat the downsampling and splicing of feature map P5 through the cross-stage local layer structure to obtain the output feature map
[0057] Each output feature map After training with the multi-core attention network module, it first passes through a standard depthwise separable convolution, and then performs multi-scale fusion with the original feature map through three depthwise separable convolution branches with different convolution kernel sizes to obtain the attention map Att. The attention map Att is then multiplied with the feature map P to obtain the final training output Out.
[0058] Output Out calculates the center point of the pest target in the pest image through classification training, calculates the major and minor axes of the pest target in the pest image through regression training, and determines the pest target detection frame according to the center point and major and minor axes of the pest target;
[0059] Through continuous training, the predicted bounding box is continuously made close to the real bounding box according to the bounding box regression loss function L, and the gradient back propagation algorithm is used to adjust the weight of the model. Finally, the trained model learns to detect pests in pest images.
[0060] The construction of the concentrated pest area focus extraction module comprises the following steps:
[0061] Set the focus factor to determine the specified pixel size to be extracted;
[0062] According to the width W and height H of the input pest image, the number of cut row blocks and column blocks is obtained:
[0063]
[0064] Among them, num rows Indicates the number of row blocks, num cols Indicates the number of column blocks,
[0065] The focus extraction algorithm of the dense pest area focus extraction module is set, and its expression is as follows:
[0066] x_start i =i×focusfactor(i=1,2,3...focusfactor),
[0067] x_end i =min((i+1)×focusfactor,H)(i=1,2,3...focusfactor),
[0068] y_start j =j×focusfactor(j=1,2,3...focusfactor),
[0069] y_end j =min((j+1)×focusfactor, W)(j=1, 2, 3...focusfactor)
[0070] image_patch=cropImage(x_start i ,y_start j , x_end i , y_end j )(
[0071] =1,2,3...focusfactor, j=1,2,3...focusfactor)
[0072] Among them, x_start i 、x_end i ,y_start j 、y_end j They represent the x-axis starting point coordinate, x-axis end point coordinate, y-axis starting point coordinate, and y-axis end point coordinate respectively. min means taking the minimum number. cropImage means focusing and cropping operations based on the four coordinates of the image.
[0073] Determine whether there are pest instances in the focus area. If yes, save the image and annotation information; if no, discard it.
[0074] The construction of a dense tiny pest image dataset comprises the following steps:
[0075] Get pest images from the pest dataset;
[0076] Input the pest image into the dense pest area focus extraction module to obtain a number of dense pest images according to the set focus factor size;
[0077] According to the density of pests, set the density level k i The formula is as follows:
[0078]
[0079] m represents the number of images, n i represents the number of pests in the i-th image, max(n1, n2…n i )(i=1,2,3…m) means taking the number of pests in the picture with the largest number of pests among all the images,
[0080] Set multiple intervals and set a density level for each interval. If k i If it falls within a certain interval, the pest density is the density level of the interval, until the pest density affects the setting of the focus factor;
[0081] The obtained dense pest images are mixed with the original pest dataset to form a dense tiny pest image dataset.
[0082] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, a dense tiny pest image detection method based on a multi-core attention network can be implemented.
[0083] A computer device, characterized in that it includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes, a dense tiny pest image detection method based on a multi-core attention network can be implemented.
[0084] Beneficial Effects
[0085] The present invention discloses a method for detecting dense tiny pest images based on a multi-core attention network, which significantly improves the detection accuracy of dense tiny pest images compared with the prior art. It uses a dense pest area focusing extraction module to extract dense pest areas, and effectively extracts multi-scale features of pests through a cross-stage local network. It is combined with the proposed small-scale pest dual-channel multi-core feature extraction model to more accurately identify tiny pests and accelerate the detection speed and accuracy of pests.
[0086] The present invention constructs a concentrated pest area focus extraction module, which can accurately extract concentrated pest area images from pest images with complex backgrounds, generate more targeted and adaptive data sets, and classify the pest density so that the model can be optimized according to different pest density scenarios, effectively enhancing the detection capability in complex environments in actual agricultural production.
[0087] The present invention improves the efficiency and accuracy of pest detection by constructing a dual-channel multi-core feature extraction model for small-scale pests. Through the top-down and bottom-up bidirectional information flow design of the dual-channel feature pyramid module, efficient fusion of multi-scale features is achieved, information loss is reduced, and the detection performance of the model for pests of different sizes is comprehensively improved. Through the multi-core attention network module, richer and more detailed multi-scale features of pests can be accurately captured, and at the same time, the extracted multi-scale features can be optimized and integrated to output more representative feature maps, thereby more accurately identifying tiny pests. BRIEF DESCRIPTION OF THE DRAWINGS
[0088] Figure 1 is a method sequence diagram of the present invention;
[0089] Figure 2 A model structure diagram of a small-scale pest dual-channel multi-core feature extraction model of the present invention;
[0090] Figure 3 This is a model structure diagram of the multi-core attention network module of the present invention;
[0091] Figure 4 This is a result diagram of pest disease images detected using the method of the present invention. DETAILED DESCRIPTION
[0092] In order to have a further understanding and recognition of the structural features and the effects achieved by the present invention, a preferred embodiment and accompanying drawings are used for detailed description as follows:
[0093] like Figure 1 As shown, the dense tiny pest image detection method based on a multi-core attention network described in the present invention comprises the following steps:
[0094] The first step is to generate a dense tiny pest image dataset: use the dense pest area focus extraction module to focus on and extract the dense pest area to generate a dense tiny pest image dataset.
[0095] The collection process uses a specially designed collection device, which consists of a front macro lens camera, a mobile data transmission terminal, and a retractable carbon fiber stand. The collected pests include a variety of pests, most of which are common pests in the fields. At the same time, in order to address the problem of the small number of rare pest species collected, data enhancement processing is performed on them, including data rotation and data flipping, to expand the proportion of rare pest species in the pest dataset. The annotation information includes the center point coordinates, width, height, and pest category of the pest bounding box. Severely occluded pests with a visible part less than a certain ratio are not labeled. The annotation file is saved in a special format to construct a pest dataset.
[0096] (1) Obtain images of wheat infested by pests, annotate and preprocess the images, and construct a pest dataset;
[0097] (2) Construct a focused extraction module for dense pest areas.
[0098] By observing the pest dataset, we can see that the distribution of pests is often clustered, possibly in clusters or bands, and scattered appearance is less common. In addition, the background in the pest dataset often looks complicated, which brings many difficulties to the detection of pests. In order to reduce the influence of cluttered background during the detection process and enable the model to pay more attention to the pests that need to be detected, thereby improving the detection accuracy, it is necessary to build a dense pest area focus extraction module to solve such problems.
[0099] The construction of the concentrated pest area focus extraction module includes the following steps:
[0100] A1) Setting the focus factor to determine the specified pixel size to be extracted;
[0101] A2) According to the width W and height H of the input pest image, obtain the number of cutting row blocks and column blocks:
[0102]
[0103] Among them, num rows Indicates the number of row blocks, num cols Indicates the number of column blocks,
[0104] A3) Setting the focus extraction algorithm of the dense pest area focus extraction module, the expression is as follows:
[0105] x_start i =i×focusfactor(i=1,2,3...focusfactor),
[0106] x_end i =min((i+1)×focusfactor, H)(i=1,2,3...focusfactor),
[0107] y_start j =j×focusfactor(j=1,2,3...focusfactor),
[0108] y_end j =min((j+1)×focusfactor, W)(j=1,2,3…focusfactor)
[0109] image_patch=cropImage(x_starti ,y_start j , x_end i , y_end j )(
[0110] =1,2,3…focusfactor, j=1,2,3…focusfactor)
[0111] Among them, x_start i 、x_end i ,y_start j 、y_end j They represent the x-axis starting point coordinate, x-axis end point coordinate, y-axis starting point coordinate, and y-axis end point coordinate respectively. min means taking the minimum number. cropImage means focusing and cropping operations based on the four coordinates of the image.
[0112] A4) Determine whether there is a pest instance in the focus area. If yes, save the image and annotation information; if no, discard it.
[0113] (3) Constructing a dense micro-pest image dataset: Input the pest dataset into the dense pest area focus extraction module to extract the dense pest area image, and mix the dense pest area image with the pest dataset to form a dense micro-pest image dataset. Since the denser the pest distribution and the smaller the pest scale, the greater the detection difficulty, in order to strengthen the solution to the dense and small-scale problems of pests, it is necessary to reconstruct this pest dataset. First, obtain pest images from the pest dataset, and input the pest images into the dense pest area focus extraction module to obtain a series of dense pest images based on the set focus factor size. The obtained dense pest images are mixed with the original pest dataset to form a dense micro-pest image dataset. The reconstructed dense micro-pest image dataset can well enhance the detection of small-scale dense pests.
[0114] Building a dense dataset of tiny pest images includes the following steps:
[0115] B1) Obtain pest images from the pest dataset;
[0116] B2) inputting the pest image into a dense pest area focus extraction module to obtain a number of dense pest images according to the set focus factor size;
[0117] B3) Classify pests according to their density and set density level k i The formula is as follows:
[0118]
[0119] m represents the number of images, n irepresents the number of pests in the i-th image, max(n1, n2…n i )(i=1,2,3…m) means taking the number of pests in the picture with the largest number of pests among all the images,
[0120] Set multiple intervals and set a density level for each interval. If k i If it falls within a certain interval, the pest density is the density level of the interval, until the pest density affects the setting of the focus factor;
[0121] B4) Mixing the obtained dense pest images with the original pest dataset to form a dense tiny pest image dataset.
[0122] The second step is to extract multi-scale features of pest images: extract multi-scale features by a cross-stage local network. During the design process, it is necessary to ensure that the network can effectively extract multi-scale features and improve detection accuracy, while also controlling the amount of calculation to avoid excessive consumption of computing resources and low operating efficiency due to overly complex structures. Therefore, a cross-stage local network is constructed to extract multi-scale features.
[0123] (1) Construct a cross-stage local network, which includes a focusing structure, a convolutional normalized activation structure, and a cross-stage local layer structure. The cross-stage local layer structure decomposes the residual block stack and introduces a spatial pyramid pooling structure before the last cross-stage local layer structure to extract features;
[0124] The input image obtained from a dense tiny pest image dataset first enters the focusing structure for slicing operation to generate multiple low-resolution feature maps and stack them to quadruple the number of input channels. After that, the multi-channel feature maps are spliced, combined and stacked together, and cross-processed by the convolutional normalization activation structure and the cross-stage local layer structure to enhance the small-scale pest features and obtain a small-scale feature map of pests.
[0125] (2) The focusing structure extracts the pixel value of every other pixel from the input image. In this way, four independent feature layers are generated. These four feature layers are then stacked together to integrate the width and height information of the image into the channel dimension, thereby increasing the number of input channels by four times.
[0126] (3) Convolutional normalization activation structure adds a batch normalization layer and SiLU activation function after the convolution layer. The formula of SiLU activation function f(x) is:
[0127] f(x) = x·sigmoid(x)
[0128] Among them, the formula of the sigmoid function is expressed as:
[0129]
[0130] Among them, x represents the input value and e represents the natural constant.
[0131] Batch normalization can ensure that the training speed is not affected when the network depth increases, and accelerate the convergence of the model. The SiLU activation function allows the network to better learn image features. In this process, the different scale features in the feature map are further extracted and enhanced, some tiny details and features of larger areas are enhanced, and the differences between features of different scales are more obvious.
[0132] (4) The cross-stage local layer structure is divided into two parts: one part directly performs forward propagation, retains relatively original feature information, and contains basic features at different scales; the other part undergoes multiple repeated 1×1 basic convolutions and 3×3 depthwise separable convolution operations to further extract and refine features and mine more representative features at different scales. These two parts of information are merged at the end of the network, so that features at different levels and scales can be integrated to enrich the diversity of features, form a feature map containing multi-scale information, and obtain a multi-scale feature map of pest images. This fusion not only integrates the original features and processed features, but also combines the features obtained under different receptive fields, thereby more comprehensively covering the multi-scale information in the image.
[0133] The third step is Figure 2 As shown, a dual-channel multi-core feature extraction model for small-scale pests is constructed.
[0134] (1) A small-scale pest dual-channel multi-core feature extraction model is set up, including a dual-channel feature pyramid module and a multi-core attention network module.
[0135] (2) Setting up a dual-channel feature pyramid module, which consists of a top-down information flow and a bottom-up information flow;
[0136] Among them, the top-down information flow: starting from the higher layers of the network and gradually transferring information to the lower layers, first upsampling the higher-level feature maps, and then fusion with the feature maps of the middle or bottom layers to enhance the semantic information in the lower-level feature maps. This fusion helps to supplement fine details and improve the ability to capture small objects and fine features. The formula is expressed as:
[0137] P5=Conv(f3)
[0138]
[0139] Among them, f i (i=3, 2, 1) represent the three feature maps of top-down information flow input, P i(i=5, 4, 3) represent the three feature maps output by the top-down information flow, Represents a fusion operation, Conv represents a convolution operation, CSPLayer represents a cross-stage local layer structure, and UpSample represents an upsampling operation;
[0140] Bottom-up information flow: Starting from the lower-level feature map, the information is gradually passed to the higher-level feature map. In the bottom-up information flow, the lower-level feature map is downsampled and fused with the higher-level feature map. The formula is expressed as:
[0141]
[0142] Among them, P i (i=5, 4, 3) represent the three feature maps output by the top-down information flow, They represent the final output feature maps obtained from the bottom-up information flow, represents a fusion operation, CSPLayer represents a cross-stage local layer structure, and Downsample represents a downsampling operation.
[0143] (3) Figure 3 As shown, a multi-core attention network module is constructed. The multi-core attention network module includes an aggregation module, an extraction module, and a reconstruction module.
[0144] Set the aggregation module, which uses a standard depth-wise separable convolution with a kernel size of k×k, where k is the set kernel size;
[0145] Set up the extraction module, which uses a multi-branch deep strip convolution layer. In each branch, a 1×j depthwise separable strip convolution and a j×1 depthwise separable strip convolution are used to simulate a j×j standard depthwise separable convolution, where the kernel size j in different branches is different;
[0146] Set the reconstruction module, which uses a 1×1 standard convolution to reconstruct the relationship between different channels;
[0147] First, the feature information from the dual-channel feature pyramid network is output as a local information feature map through the aggregation module. Secondly, the multi-scale features of the local information feature map are extracted through the extraction module to obtain multi-scale feature maps of different channels. Finally, the multi-scale feature map is reconstructed through the reconstruction module to reconstruct the relationship between different channels and output the final feature map.
[0148] The multi-core attention network module formula is expressed as:
[0149]
[0150] Among them, P refers to the input feature, Att and Out represent the attention map and output results respectively, Conv 1×1 Represents a 1×1 convolution operation, DW-Conv is the abbreviation of depthwise separable convolution, Scale i Where i takes the value of 0, 1, 2, or 3, and Scale i When 1, 2, or 3 is selected, it indicates branch options, and Scale 0 indicates identity mapping.
[0151] The fourth step is to train the small-scale pest dual-channel multi-core feature extraction model: input the dense tiny pest image dataset into the small-scale pest dual-channel multi-core feature extraction model for training.
[0152] The training of the small-scale pest dual-channel multi-core feature extraction model includes the following steps:
[0153] (1) Setting the bounding box regression loss function:
[0154] The bounding box regression loss function L of the small-scale pest dual-channel multi-core feature extraction model is defined as:
[0155]
[0156] Among them, IoU represents the intersection-over-union ratio of the predicted bounding box and the true bounding box area, b and b gt denote the center points of the predicted bounding box and the true bounding box, respectively, and ρ 2 (b,b gt ) is the squared Euclidean distance between the center points of the predicted bounding box and the true bounding box, c is the diagonal length of the minimum enclosing rectangle of the predicted bounding box and the true bounding box, α is the weight coefficient used to balance the impact of the aspect ratio, and v is a measure of the consistency of the aspect ratio between the predicted bounding box and the true bounding box. The calculation formula is:
[0157]
[0158] w and h are the width and height of the predicted bounding box, w gt and h gt is the width and height of the true bounding box, and the formula for the weight coefficient α is:
[0159]
[0160] The formula for intersection over union (IoU) is:
[0161]
[0162] bbox is the area size of the predicted bounding box, bbox gt is the area of the true bounding box, ∩ represents the intersection operation, and ∪ represents the union operation.
[0163] (2) The multi-scale feature map of the pest image generated from the cross-stage local layer structure is input into the dual-channel feature pyramid module for training.
[0164] According to the top-down information flow, first, the multi-scale feature map of the pest image at the top layer is processed by a two-dimensional convolution to obtain the feature map P5, P5 is then upsampled to expand the image size of the feature map, and then it is spliced with the multi-scale feature map of the pest image at the middle layer to obtain the feature map P4 through a two-dimensional convolution operation through a cross-stage local layer structure, P4 is then upsampled to expand the image size of the feature map, and then it is spliced with the multi-scale feature map of the pest image at the bottom layer to obtain the feature map P3 through a cross-stage local layer structure. According to the bottom-up information flow, first, the feature map P3 at the bottom layer is directly output to obtain the feature map At the same time, the feature map P3 is downsampled and spliced with the feature map P4 through the cross-stage local layer structure to obtain the output feature map Then repeat the downsampling and splicing of feature map P5 through the cross-stage local layer structure to obtain the output feature map
[0165] (3) Each output feature map After training with the multi-core attention network module, it first passes through a standard depthwise separable convolution, and then performs multi-scale fusion with the original feature map through three depthwise separable convolution branches with different convolution kernel sizes to obtain the attention map Att. The attention map Att is then multiplied with the feature map P to obtain the final training output Out.
[0166] The output Out is calculated by classification training to obtain the center point of the pest target in the pest image, and the major and minor axes of the pest target in the pest image are calculated by regression training. The pest target detection frame is determined according to the center point and major and minor axes of the pest target.
[0167] (4) Through continuous training, the predicted bounding box is continuously made close to the real bounding box according to the bounding box regression loss function L, and the weight of the model is adjusted using the gradient back propagation algorithm. Finally, the trained model learns to detect pests in pest images.
[0168] The fifth step is to obtain the detection results of dense tiny pest images: obtain the pest image to be detected and input it into the cross-stage local network to obtain the multi-scale feature map of the pest image, and input the multi-scale feature map into the trained small-scale pest dual-channel multi-core feature extraction model to obtain the pest image detection results.
[0169] Here, a computer-readable storage medium is also provided, on which a computer program is stored. When the computer program is executed by a processor, a dense micro-pest image detection method based on a multi-core attention network can be implemented. A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes, a dense micro-pest image detection method based on a multi-core attention network can be implemented.
[0170] Depend on Figure 4 As shown, the method of the present invention detects pests in 4 randomly selected pest images, and it can be seen that each pest can be accurately detected, indicating that the method of the present invention has good detection accuracy for dense and small-scale pests in complex environments.
[0171] Table 1 Comparison of the method of the present invention and the most advanced detection method in dense micro-pest data sets
[0172]
[0173]
[0174] In Table 1, the best results are shown in bold. As can be seen from Table 1, compared with the detection results of other methods, the method of the present invention can greatly improve the detection accuracy, far exceeding other models in the average accuracy of pests, and the number of parameters of the model is also lightweight and compact, which proves that the present invention can be applied to actual agricultural scenarios.
[0175] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions only describe the principles of the present invention. The present invention may be subject to various changes and improvements without departing from the spirit and scope of the present invention. These changes and improvements fall within the scope of the present invention. The scope of protection claimed by the present invention is defined by the attached claims and their equivalents.
Claims
1. A dense micro-pest image detection method based on a multi-core attention network, characterized in that: The following steps are involved: 11) Generation of dense tiny pest image dataset: Using the dense pest area focus extraction module, the dense pest area is focused and extracted to generate a dense tiny pest image dataset; 12) Extraction of multi-scale features of pest images: Extract multi-scale features by a cross-stage local network; 13) Construct a dual-channel multi-core feature extraction model for small-scale pests; 14) Training of small-scale pest dual-channel multi-core feature extraction model: Input the dense tiny pest image dataset into the small-scale pest dual-channel multi-core feature extraction model for training; 15) Obtaining the detection results of dense tiny pest images: Obtain the pest image to be detected and input it into the cross-stage local network to obtain the multi-scale feature map of the pest image, and input the multi-scale feature map of the pest image into the trained small-scale pest dual-channel multi-core feature extraction model to obtain the pest image detection results.
2. According to claim 1, a dense tiny pest image detection method based on a multi-core attention network is characterized in that: The generation of the dense tiny pest image dataset comprises the following steps: 21) Obtain images of wheat infested by pests, annotate and preprocess the images, and construct a pest dataset; 22) Construct a focused extraction module for densely populated pest areas; 23) Constructing a dense micro-pest image dataset: Input the pest dataset into the dense pest area focus extraction module to extract dense pest area images, and mix the dense pest area images with the pest dataset to form a dense micro-pest image dataset.
3. According to claim 1, a dense tiny pest image detection method based on a multi-core attention network is characterized in that: The extraction of multi-scale features of the pest image comprises the following steps: 31) constructing a cross-stage local network, which includes a focusing structure, a convolutional normalized activation structure, and a cross-stage local layer structure. The cross-stage local layer structure decomposes the residual block stack and introduces a spatial pyramid pooling structure before the last cross-stage local layer structure to extract features; The input image obtained from the dense tiny pest image dataset first enters the focusing structure for slicing operation to generate multiple low-resolution feature maps and stack them to quadruple the number of input channels. Then, the multi-channel feature maps are spliced, combined and stacked together, and cross-processed by the convolution normalization activation structure and the cross-stage local layer structure to enhance the small-scale pest features, and obtain the pest small-scale feature map. 32) The focusing structure extracts the pixel value of every other pixel from the input image, generating four independent feature layers in this way, and then stacking these four feature layers together to integrate the width and height information of the image into the channel dimension, thereby increasing the number of input channels by four times; 33) Convolutional normalization activation structure adds a batch normalization layer and SiLU activation function after the convolution layer. The formula of SiLU activation function f(x) is: f(x)=x·sigmodi(x) Among them, the formula of the sigmoid function is expressed as: Among them, x represents the input value, and e represents the natural constant; 34) The cross-stage local layer structure is divided into two parts: one part directly performs forward propagation to retain relatively original feature information, including basic features at different scales; the other part undergoes multiple repeated 1×1 basic convolutions and 3×3 depthwise separable convolution operations to further extract and refine features and mine more representative features at different scales. These two parts of information are merged at the end of the network, so that features at different levels and scales can be integrated to enrich the diversity of features, form a feature map containing multi-scale information, and obtain a multi-scale feature map of pest images.
4. According to claim 1, a dense tiny pest image detection method based on a multi-core attention network is characterized in that: The construction of a small-scale pest dual-channel multi-core feature extraction model comprises the following steps: 41) Setting a small-scale pest dual-channel multi-core feature extraction model includes a dual-channel feature pyramid module and a multi-core attention network module; 42) Setting a dual-channel feature pyramid module, the dual-channel feature pyramid module consists of a top-down information flow and a bottom-up information flow; Among them, the top-down information flow starts from the higher layers of the network and gradually transfers information to the lower layers. First, the higher-level feature maps are upsampled and then fused with the feature maps of the middle or bottom layers to enhance the semantic information in the lower-level feature maps. The formula is expressed as: P5=Conv(f3) Among them, f i (i=3,2,1) represent the three feature maps of top-down information flow input, P i (i=5,4,3) represent the three feature maps output by the top-down information flow, Represents a fusion operation, Conv represents a convolution operation, CSPLayer represents a cross-stage local layer structure, and UpSample represents an upsampling operation; Bottom-up information flow: Starting from the lower-level feature map, the information is gradually passed to the higher-level feature map. In the bottom-up information flow, the lower-level feature map is downsampled and fused with the higher-level feature map. The formula is expressed as: Among them, P i (i=5,4,3) represent the three feature maps output by the top-down information flow, They represent the final output feature maps obtained from the bottom-up information flow, represents a fusion operation, CSPLayer represents a cross-stage local layer structure, and Downsample represents a downsampling operation; 43) Construct a multi-core attention network module, which includes an aggregation module, an extraction module and a reconstruction module. Set the aggregation module, which uses a standard depth-wise separable convolution with a kernel size of k×k, where k is the set kernel size; Set up the extraction module, which uses a multi-branch deep strip convolution layer. In each branch, a 1×j depthwise separable strip convolution and a j×1 depthwise separable strip convolution are used to simulate a j×j standard depthwise separable convolution, where the kernel size j in different branches is different; Set the reconstruction module, which uses a 1×1 standard convolution to reconstruct the relationship between different channels; First, the feature information from the dual-channel feature pyramid network is output as a local information feature map through the aggregation module. Secondly, the multi-scale features of the local information feature map are extracted through the extraction module to obtain multi-scale feature maps of different channels. Finally, the multi-scale feature map is reconstructed through the reconstruction module to reconstruct the relationship between different channels and output the final feature map. The multi-core attention network module formula is expressed as: Among them, P refers to the input feature, Att and Out represent the attention map and output results respectively, Conv 1×1 Represents a 1×1 convolution operation, DW-Conv is the abbreviation of depthwise separable convolution, Scalei i Where i takes the value of 0, 1, 2, or 3, and Scale i When 1, 2, or 3 is selected, it indicates branch options, and Scale 0 indicates identity mapping.
5. The method for detecting dense tiny pest images based on a multi-core attention network according to claim 1, characterized in that: The training of the small-scale pest dual-channel multi-core feature extraction model includes the following steps: 51) Set the bounding box regression loss function: The bounding box regression loss function L of the small-scale pest dual-channel multi-core feature extraction model is defined as: Among them, IoU represents the intersection-over-union ratio of the predicted bounding box and the true bounding box area, b and b gt denote the center points of the predicted bounding box and the true bounding box, respectively, and ρ 2 (b,b gt ) is the squared Euclidean distance between the center points of the predicted bounding box and the true bounding box, c is the diagonal length of the minimum enclosing rectangle of the predicted bounding box and the true bounding box, α is the weight coefficient used to balance the impact of the aspect ratio, and v is a measure of the consistency of the aspect ratio between the predicted bounding box and the true bounding box. The calculation formula is: w and h are the width and height of the predicted bounding box, w gt and h gt is the width and height of the true bounding box, and the formula for the weight coefficient α is: The formula for intersection over union (IoU) is: bbox is the area size of the predicted bounding box, bbox gt is the area of the true bounding box, ∩ represents the intersection operation, and ∪ represents the union operation; 52) The multi-scale feature map of the pest image from the cross-stage local layer structure is input into the dual-channel feature pyramid module for training. According to the top-down information flow, first, the multi-scale feature map of the pest image at the top layer is processed by a two-dimensional convolution to obtain the feature map P5, P5 is then upsampled to expand the image size of the feature map, and then it is spliced with the multi-scale feature map of the pest image at the middle layer to obtain the feature map P4 through a two-dimensional convolution operation through a cross-stage local layer structure, P4 is then upsampled to expand the image size of the feature map, and then it is spliced with the multi-scale feature map of the pest image at the bottom layer to obtain the feature map P3 through a cross-stage local layer structure. According to the bottom-up information flow, first, the feature map P3 at the bottom layer is directly output to obtain the feature map At the same time, the feature map P3 is downsampled and spliced with the feature map P4 through the cross-stage local layer structure to obtain the output feature map Then repeat the downsampling and splicing of feature map P5 through the cross-stage local layer structure to obtain the output feature map 53) Each output feature map After training with the multi-core attention network module, it first passes through a standard depthwise separable convolution, and then passes through three depthwise separable convolution branches with different convolution kernel sizes to perform multi-scale fusion with the original feature map to obtain the attention map Att. The attention map Att is then multiplied with the feature map P to obtain the final training output Out. Output Out calculates the center point of the pest target in the pest image through classification training, calculates the major and minor axes of the pest target in the pest image through regression training, and determines the pest target detection frame according to the center point and major and minor axes of the pest target; 54) Through continuous training, the predicted bounding box is continuously made close to the real bounding box according to the bounding box regression loss function L, and the weight of the model is adjusted using the gradient back propagation algorithm. Finally, the trained model learns to detect pests in pest images.
6. The method for detecting dense tiny pest images based on a multi-core attention network according to claim 2, characterized in that: The construction of the concentrated pest area focus extraction module comprises the following steps: 61) Setting the focus factor focusfactor to determine the specified pixel size to be extracted; 62) According to the width W and height H of the input pest image, obtain the number of cutting row blocks and column blocks: Among them, num rows Indicates the number of row blocks, num cols Indicates the number of column blocks, 63) Set the focus extraction algorithm of the dense pest area focus extraction module, and its expression is as follows: x_start i =i×focusfactor(i=1,2,3…focusfactor), x_end i =min((i+1)×focusfactor,H)(i=1,2,3…focusfactor), y_start j =j×focusfactor(j=1,2,3…focusfactor), y_end j =min((j+1)×focusfactor,W)(j=1,2,3…focusfactor) image_patch=cropImage(x_start i ,y_start j ,x_end i ,y_end j )(=1,2,3…focusfactor,j=1,2,3…focusfactor) Among them, x_start i 、x_end i ,y_start j 、y_end j They represent the x-axis starting point coordinate, x-axis end point coordinate, y-axis starting point coordinate, and y-axis end point coordinate respectively. min means taking the minimum number. cropImage means focusing and cropping operations based on the four coordinates of the image. 64) Determine whether there is a pest instance in the focus area. If yes, save the image and annotation information; if no, discard it.
7. The method for detecting dense tiny pest images based on a multi-core attention network according to claim 2, characterized in that: The construction of a dense tiny pest image dataset comprises the following steps: 71) Obtain pest images from the pest dataset; 72) Inputting the pest image into a dense pest area focus extraction module to obtain a number of dense pest images according to the set focus factor size; 73) According to the density of pests, set the density level k i The formula is as follows: m represents the number of images, n i represents the number of pests in the i-th image, max(n1,n2…n i )(i=1,2,3…m) means taking the number of pests in the picture with the largest number of pests among all the images, Set multiple intervals and set a density level for each interval. If k i If it falls within a certain interval, the pest density is the density level of the interval, until the pest density affects the setting of the focus factor; 74) The obtained dense pest images are mixed with the original pest dataset to form a dense tiny pest image dataset.
8. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by the processor, a dense tiny pest image detection method based on a multi-core attention network as described in any one of claims 1 to 7 can be implemented.
9. A computer device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, a dense tiny pest image detection method based on a multi-core attention network as described in any one of claims 1 to 7 can be implemented.
Citation Information
Patent Citations
Insect pest image detection method based on channel attention mechanism
CN113487576A
Fourier lamination microscopic imaging reconstruction method, device and equipment
CN116579924A
Cerebral aneurysm detection model establishment method and device, equipment and storage medium
CN116740533A
Gesture recognition method based on multi-core dynamic attention mechanism
CN117152838A
Automatic driving target detection system for small objects
CN118279566A