An Adaptive Feature Fusion Detection Method for Similar Pest Images under the Lamp
By building a joint dual-channel attention mechanism and a pyramid feature network of jump convolution modules, the identification problem caused by high pest shape similarity and similar sizes is solved, and real-time pest detection with high accuracy is achieved.
Patent Information
- Application Number
- CN202111577178.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-22
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-12-22
AI Technical Summary
The prior art is difficult to effectively identify pests with similar appearances and similar sizes, resulting in poor classification accuracy, and the existing methods ignore high-level semantic information or simply superimpose low-level high-level features, resulting in inflexible recognition effects.
Adaptive feature fusion detection method based on similar pest images under lamps is adopted, and a non-local similar pest detection network with a combined dual-channel attention mechanism is constructed, and a equilibrium size jump convolution module pyramid feature network is combined to perform feature fusion and optimization to output similar pest detection results.
It realizes effective identification of pests with high appearance similarity and small size difference, has real-time detection capabilities, and improves classification accuracy.
Smart Images

Figure CN114241317B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of agricultural insurance image processing, and in particular to an adaptive feature fusion detection method for similar pest images under a lamp. Background Art
[0002] There are a wide variety of agricultural pests, and the shapes of two species are very similar, and the differences within the species are extremely weak, making it difficult to distinguish them through rough feature extraction; for the same species, due to different postures, backgrounds, and shooting angles, there are large intra-class differences. Therefore, in order to successfully perform fine-grained classification on two similar categories, the most important thing is to find the discriminative region blocks in the image that can distinguish these two species, and be able to represent the features of these discriminative region blocks well to achieve a better purpose of identifying categories. For pest targets with high similarity, small size and proximity, current methods only consider the low-level feature maps in the feature pyramid as their local features and ignore the high-level semantic information, resulting in good localization effects for pest targets but poor classification accuracy. On the other hand, simply superimposing the information of both low-level and high-level feature maps will make the local features of pests chaotic and lack pertinence, affecting the recognition effect and making the intelligent detection method not flexible enough.
[0003] Fine-grained image classification, also known as sub-category image classification, is a very popular research topic in the fields of computer vision, pattern recognition, etc. in recent years. Its purpose is to make a more detailed sub-category division of images belonging to the same basic category (such as cars, dogs, flowers, birds, etc.). However, due to the subtle inter-class differences between sub-categories, compared with ordinary image classification tasks, fine-grained image classification is more difficult. The goal of fine-grained image classification is to distinguish different sub-categories under the same common category. Due to the large inter-class similarity of the dataset, fine-grained image classification is more challenging than traditional image classification. In previous work, component-based methods and attention-based methods have been dedicated to mining the salient regions in images while ignoring the subtle differences used to distinguish easily confused classes. Summary of the Invention
[0004] The purpose of the present invention is to provide an adaptive feature fusion detection method for similar pest images under a lamp, which has a good detection and recognition effect on pests with similar shapes and can achieve real-time detection.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions: An adaptive feature fusion detection method for similar pest images under a lamp, the method includes the following steps in sequence:
[0006] (1) Obtain a dataset of similar pest images and perform preprocessing;
[0007] (2) Extract the similarity features of pests;
[0008] (3) Construct a non-local similar pest detection network with a joint dual-channel attention mechanism, and input the pest similarity features extracted in step (2) into the non-local similar pest detection network;
[0009] (4) Design a jumping convolutional module pyramid feature network with a balanced size, take the output of the non-local similar pest detection network as the input of the jumping convolutional module pyramid feature network, and the jumping convolutional module pyramid feature network outputs a reference detection box that meets the conditions;
[0010] (5) Optimize the center points of the selected reference detection boxes, and use the non-maximum suppression method for duplicate removal optimization;
[0011] (6) Output the similar pest detection results.
[0012] The specific steps of step (1) include the following steps:
[0013] (1a) Obtain a similar pest image dataset, segment the image data in the similar pest image dataset, and perform size normalization preprocessing on all the segmented single-class pest images;
[0014] (1b) Randomly select 100 of all the single-class pest images that have undergone size normalization preprocessing for similarity statistics.
[0015] The specific steps of step (2) include the following steps:
[0016] (2a) Extract the pest similarity texture:
[0017] In the RGB color space, use the Sobel operator to calculate each of the three color channels, and return two vectors corresponding to the horizontal and vertical directions in each channel: vector a (Rx, Gx, Bx) and vector b (Ry, Gy, By), and calculate through formulas (1) to (4):
[0018]
[0019] a·b = R x ·R y +G x ·G y +B x ·B y (3)
[0020]
[0021] θ is the cosine angle between vectors a and b;
[0022] Measure the similarity color of pests: Quantize the results obtained from the R, G, and B channels into images of 64 colors respectively. Use 4 different primitives to detect textures in the image C(x, y) to obtain the texture primitive image T(x, y). Finally, describe the texture features with a multi-primitive histogram according to T(x, y). The definitions of the multi-primitive histogram are shown in formulas (5) and (6):
[0023] H(T(P1)) = N{θ(P1) = v1 ∩ θ(P2) = v2 ||P1 - P2|| = D}(5)
[0024] H(T(P1)) = N{θ(P1) = w1 ∩ θ(P2) = w2 ||P1 - P2|| = D}(6)
[0025] Among them, P1 = (x1, y1) and P2 = (x2, y2) represent two adjacent pixel points with a distance of D in the original image. The corresponding pixel points in the primitive image are T(P1) = w1 and T(p2) = w2 respectively. In the texture direction matrix θ(x, y), the directions of points P1 and P2 are θ(P1) = v1 and θ(P2) = v2 respectively; N represents the number of times v1 and v2 appear together, represents the number of times w1 and w2 appear together. v1 and v2 are the four test primitives of the image C(x, y), and w1 and w2 are the corresponding pixel points in the texture primitive image T(x, y); H(T(P1)) represents the number of times the same edge direction appears simultaneously under a certain color background; H(T(P1)) represents the number of times the same color appears under a certain edge direction; The calculation formula for the texture feature vector fv of the image is as follows:
[0026]
[0027] Among them represents connection. The similarity of images I1 and I2 is defined as follows:
[0028] S I (I1, I2) = ||f v (I1) - f v (I2)|| -1
[0029] (2b) Extract the similarity size of pests:
[0030] Given an RGB image with the shape of H×W and the i-th pest bounding box Hi×Wi, the target relative size ORSi of this pest object is defined as:
[0031]
[0032] Calculate the ORS of the c-th category in the entire dataset in the following way:
[0033]
[0034] where M is the number of pest objects, and the function sgn(·) indicates that the category of the i-th pest is the c-th category, which is defined as:
[0035]
[0036] Finally, obtain the ORS distribution map of pest species for all categories.
[0037] The specific steps of step (3) include the following steps:
[0038] (3a) Set the input of the non-local similarity pest detection network as the candidate region of the pest image and the output as the pest image feature result;
[0039] (3b) Set the first layer of the non-local similarity pest detection network to extract multi-scale local features of the pest target through the method of max pooling: according to the position and size of each pest feature region, extract the corresponding local features of the pest target in the similar pest feature map;
[0040] (3c) Set the second layer of the non-local similarity pest detection network to calculate the feature weights through a dual-channel attention module: according to the size of each pest feature region, calculate the corresponding feature weights of each pest feature region in the feature maps of each scale through the channel attention function;
[0041] (3d) Set the third layer of the adaptive similarity pest detection network as the non-local feature fusion of similar pest targets: fuse the similar pest features with the weights calculated by the tg non-local feature function;
[0042] Calculate the function output of each pest feature region, introduce the non-local feature fusion function, and construct a feature fusion that combines the features of the second layer of the extracted features with the features of the adjacent upper and lower layers. The formula of the non-local feature fusion function is as follows:
[0043]
[0044] In the formula, the input is x, the output is y, i and j respectively represent a certain spatial position of the input, and X i is a vector, the dimension of which is the same as the number of channels of X, f is a function that calculates the similarity relationship between any two points, and g is a mapping function that maps a point into a vector.
[0045] The specific steps of step (4) include the following steps:
[0046] (4a) Add a skip convolution module to the pyramid feature network. Given a set of filters K with shape (C, C, kh, kw), the filter set K consists of K1, K2, K3, K4, and K5, which are divided into two branches and perform different functions through K1, K2, K3, K4, and K5. In the skip convolution module, feature transformation is performed on the original scale and the small size after downsampling. For a given X, it is divided into two parts, one part is X1 and the other part is X2; C is the number of channels, K1 to K5 are filters; kh is the height, and kw is the width;
[0047] (4b) Use max pooling to reduce the scale:
[0048] M1 = MaxPool r ()X1(T)1 = MaxPool r M1
[0049] Among them, r is the downsampling rate and stride of the pooling process; MaxPool() is the max pooling function;
[0050] (4c) T1 is used as the input of filters K2 and K3, and the features are restored to the original scale after the upsampling operation. The result is:
[0051] X1' = Up(F2(T1)) = Up(T1 × K2)
[0052] X1” = Up(F3(X1')) = Up(X1' × K3)
[0053] Among them, Up() is the upsampling function, and X1’ is the result obtained by two upsamplings of the product operation of filter K2 and T1;
[0054] Among them, F2(T1)) = T1 × K2, F3(X1') = X1' × K3 is a simplified form of convolution. Then, the calibration operation is expressed as:
[0055] Y1' = F4(X1) ⊙ Sigmoid(X1”)
[0056] Among them, Y1’ is the dot product form of the value of the feature vector X1” passing through the activation function Sigmoid and X1 passing through the convolution function;
[0057] F4(X1) = X1 × K4, Sigmoid(·) is an activation function. Calculate the final result of the skip convolution module as:
[0058] Y1 = F5(Y1')
[0059] F5(Y1')) = Y1' × K2
[0060] Among them, Y1 is the vector value of the feature vector Y1' after passing through the convolution function;
[0061] (4d) Y2 = F(X2) = X2 × K1
[0062] Among them, Y2 is the vector value obtained by multiplying the feature vector X2 and the filter K1;
[0063] (4e) Sum Y1 and Y2 to obtain the final result Y, completing the construction of the skip convolution module pyramid feature network with the skip convolution module added;
[0064] (4f) Output pest classification and reference detection boxes:
[0065] Calculate the classification labels according to the Softmax regression model, sort them according to the confidence level, and select the values whose probability exceeds the confidence interval from largest to smallest in turn; then use the loss function to train the non-local similarity pest detection network and output the reference detection boxes with classification labels.
[0066] The specific steps of step (5) include the following steps:
[0067] (5a) Filter redundant boxes: Select the bounding box with the largest score for filtering;
[0068] Select the boxes with an intersection over union greater than 0.6 as the reference detection boxes for regression, and use the non-maximum suppression method for screening. The specific steps include:
[0069] (5a1) Sort the scores of all boxes, and select the highest score and its corresponding box:
[0070] (5a2) Traverse the remaining boxes. If the overlapping area IOU between the box that meets the intersection over union greater than the specified condition and the current highest-score box is greater than a certain threshold, delete the box;
[0071] (5a3) Continue to select the box with the highest score from the unprocessed boxes and repeat the above process until there are no detection boxes that meet the conditions;
[0072] (5b) Perform center point optimization:
[0073] (5b1) Define the center point. Add a single-layer branch in parallel with the regression branch to predict the "centrality" of a position. Given the regression targets l, t, r, and b of a position, the center target is defined as:
[0074]
[0075] Among them, l, t, r, and b are the distance values from the left, top, right, and bottom ends to the coordinates of the center point of the image respectively;
[0076] (5b 2) Participate in the training using the binary cross - entropy (BCE) loss and the sigmoid activation function, and jointly calculate the loss to ensure that the calculated detection boxes are close to the center position. The cross - entropy function used is:
[0077]
[0078] where M is the number of categories; y ic is the sign function, taking values of 0 or 1; if the label category of sample i is equal to C, it takes the value of 1, otherwise 0; p ic is the predicted probability that the observed sample i belongs to category c.
[0079] As can be seen from the above technical solutions, the beneficial effects of the present invention are as follows: First, it can effectively identify a pest data set with high similarity in appearance and small difference in size; Second, it introduces a dual - channel attention mechanism, uses a non - local module to reconstruct the feature extraction method, and adds a skip convolution module to reconstruct the feature pyramid; Third, the present invention is a one - stage anchor - free detector and can achieve real - time detection effects. Brief Description of the Drawings
[0080] Figure 1 is the flowchart of the method of the present invention;
[0081] Figure 2 is the proportion of 24 types of pests in the picture;
[0082] Figure 3 is the schematic diagram of the skip correction convolution module;
[0083] Figure 4 is the schematic diagram of predicting a 4D ((l, t, r, b)) vector for a certain position point;
[0084] Figure 5 is the schematic diagram of the detection results of YOLO3 (the first column), Faster RCNN (the second column), FCOS (the third column) and the present DACFA (the fourth column). Detailed Embodiments
[0085] As Figure 1 shown, a method for adaptive feature fusion detection of similar pest images based on under - lamp images, the method includes the following steps in sequence:
[0086] (1) Obtain a data set of similar pest images and perform pre - processing;
[0087] (2) Extract the similarity features of pests;
[0088] (3) Construct a non - local similar pest detection network with a joint dual - channel attention mechanism, and input the pest similarity features extracted in step (2) into the non - local similar pest detection network;
[0089] (4) Design a jump convolution module pyramid feature network with balanced dimensions, take the output of the non-local similarity pest detection network as the input of the jump convolution module pyramid feature network, and the jump convolution module pyramid feature network outputs reference detection boxes that meet the conditions;
[0090] (5) Optimize the center points of the selected reference detection boxes and use non-maximum suppression method for duplicate removal optimization;
[0091] (6) Output the similar pest detection results.
[0092] The specific steps of step (1) include the following steps:
[0093] (1a) Obtain a similar pest image dataset, segment the image data in the similar pest image dataset, and perform size normalization preprocessing on all the segmented single-class pest images;
[0094] (1b) Randomly select 100 single-class pest images that have undergone size normalization preprocessing for similarity statistics.
[0095] The specific steps of step (2) include the following steps:
[0096] (2a) Extract pest similarity textures:
[0097] In the RGB color space, use the Sobel operator to calculate each of the three color channels separately, and return two vectors corresponding to the horizontal and vertical directions in each channel: vector a (Rx, Gx, Bx) and vector b (Ry, Gy, By), and calculate through formulas (1) to (4):
[0098]
[0099] a·b = R x ·R y +G x ·G y +B x ·B y (3)
[0100]
[0101] θ is the cosine angle between vectors a and b;
[0102] Measure the similarity color of pests: Quantize the results obtained from the R, G, and B channels into images of 64 colors respectively, perform texture detection in the image C(x, y) with 4 different primitives respectively to obtain the texture primitive image T(x, y), and finally describe the texture features with a multi-primitive histogram according to T(x, y). The definitions of the multi-primitive histogram are shown in formulas (5) and (6):
[0103] H(T(P1)) = N{θ(P1) = v1 ∩ θ(P2) = v2 ||P1 - P2|| = D}(5)
[0104] H(T(P1)) = N{θ(P1) = w1 ∩ θ(P2) = w2 ||P1 - P2|| = D}(6)
[0105] Among them, P1 = (x1, y1) and P2 = (x2, y2) represent two adjacent pixel points with a distance of D in the original image. Their corresponding pixel points in the primitive image are T(P1) = w1 and T(p2) = w2 respectively. In the texture direction matrix θ(x, y), the directions of points P1 and P2 are θ(P1) = v1 and θ(P2) = v2 respectively; N represents the number of times v1 and v2 appear together, represents the number of times w1 and w2 appear together. v1 and v2 are the four test primitives of the image C(x, y), and w1 and w2 are the corresponding pixel points in the texture primitive image T(x, y); H(T(P1)) represents the number of times the same edge directions appear simultaneously under a certain color background; H(T(P1)) represents the number of times the same color appears under a certain edge direction; The calculation formula for the texture feature vector fv of the image is as follows:
[0106]
[0107] Among them represents connection. The similarity of images I1 and I2 is defined as follows:
[0108] S I (I1, I2) = ||f v (I1) - f v (I2)|| -1
[0109] There is another extraction method for the texture of pest similarity:
[0110] The similarity measures of the dataset can be divided into texture gray-level and color measurement methods. Gray-level methods generally use hash algorithms, which are common methods for describing image similarity. In terms of recognition effect, the dHash algorithm is better than the aHash algorithm but not as good as the pHash algorithm. The perceptual hash algorithm is used to analyze the similarity of 24 categories of pests extracted at different scales. First, 100 randomly selected classified intercepted images are selected from each category, and the selected images are scaled to 32×32 and 64×64 for calculating their respective similarities. The confusion matrices calculated are shown in Table 1 and Table 2:
[0111] Table 1 Data table of phash similarity algorithm for 17 subcategories of similar pests with size (32*32).
[0112]
[0113] Table 2 Data table of phash similarity algorithm for 7 subcategories of similar pests with size (32*32):
[0114]
[0115] (2b) Extract the similarity size of pests:
[0116] Given an RGB image with the shape of H×W and the i-th pest bounding box Hi×Wi, the target relative size ORSi of this pest object is defined as:
[0117]
[0118] The ORS of the c-th category in the entire dataset is calculated in the following way:
[0119]
[0120] where M is the number of pest objects, and the function sgn(·) indicates that the category of the i-th pest is the c-th category, which is defined as:
[0121]
[0122] Finally, the ORS distribution maps of pest species in all categories are obtained.
[0123] Step (3) specifically includes the following steps:
[0124] (3a) Set the input of the non-local similar pest detection network as the candidate region of the pest image and the output as the feature result of the pest image;
[0125] (3b) Set the first layer of the non-local similarity pest detection network to extract multi-scale local features of pest targets by the method of max pooling: According to the position and size of each pest feature region, extract the corresponding local features of pest targets in the similarity pest feature map;
[0126] (3c) Set the second layer of the non-local similarity pest detection network to calculate feature weights through a dual-channel attention module: According to the size of each pest feature region, calculate the corresponding feature weights of each pest feature region in the feature maps of each scale through the channel attention function:
[0127] (3d) Set the third layer of the adaptive similarity pest detection network to non-local feature fusion of similar pest targets: Use the weights calculated by the tg non-local feature function to fuse the similar pest features:
[0128] Calculate the function output of each pest feature region, introduce the non-local feature fusion function, and construct a feature fusion that combines the features of the second layer of the extracted features with the adjacent upper and lower layers. The formula of the non-local feature fusion function is as follows:
[0129]
[0130] In the formula, the input is x, the output is y, i and j respectively represent a certain spatial position of the input, and X i is a vector, the dimension is the same as the number of channels of X, f is a function that calculates the similarity relationship between any two points, and g is a mapping function that maps a point into a vector.
[0131] The specific steps of step (4) are as follows:
[0132] (4a) Add a skip convolution module to the pyramid feature network. Given a set of filter sets K with the shape of (C, C, kh, kw), the filter set K is composed of K1, K2, K3, K4, K5, and is divided into two branches. Perform different functions through K1, K2, K3, K4, K5. In the skip convolution module, perform feature transformation on the original scale and the small size after downsampling. For the given X, it is divided into two parts, one part is X1, and the other part is X2; C is the number of channels, K1 to K5 are filters; kh is the height, and kw is the width;
[0133] (4b) Use max pooling to reduce the scale:
[0134] M1 = MaxPool r ()X1(T)1 = MaxPool r M1
[0135] Among them, r is the downsampling rate and stride of the pooling process; MaxPool() is the max pooling function;
[0136] (4c) T1 is used as the input of filters K2 and K3, and after the upsampling operation, the features are restored to the original scale, and the result is:
[0137] X1' = Up(F2(T1)) = Up(T1 × K2)
[0138] X1” = Up(F3(X1')) = Up(X1' × K3)
[0139] where Up() is the upsampling function, and X1’ is the result obtained by two upsamplings of the product operation of filter K2 and T1;
[0140] where F2(T1)) = T1 × K2, F3(X1') = X1' × K3 is a simplified form of convolution. Then, the calibration operation is expressed as:
[0141] Y1' = F4(X1) ⊙ Sigmoid(X1”)
[0142] where Y1’ is the dot product form of the value of the feature vector X1” passing through the activation function Sigmoid and X1 passing through the convolution function;
[0143] F4(X1) = X1 × K4, Sigmoid(·) is an activation function, and the final result of calculating the skip convolution module is:
[0144] Y1 = F5(Y1')
[0145] F5(Y1')) = Y1' × K2
[0146] where Y1 is the vector value of the feature vector Y1’ passing through the convolution function vector;
[0147] (4d) Y2 = F(X2) = X2 × K1
[0148] where Y2 is the vector value of the product of the feature vector X2 and filter K1;
[0149] (4e) Summing Y1 and Y2 to obtain the final result Y, completing the construction of the skip convolution module pyramid feature network with the skip convolution module added;
[0150] As Figure 3 shown, the skip convolution module enables each spatial position to adaptively encode the context from distant regions, which is also the difference between it and the traditional FPN convolution.
[0151] (4f) Output pest classification and reference detection box:
[0152] Calculate the classification labels according to the Softmax regression model, sort them according to the confidence level, and select the values whose probability exceeds the confidence interval from largest to smallest in turn; then use the loss function to train the non-local similarity pest detection network, and output the reference detection boxes with classification labels.
[0153] The specific steps of step (5) are as follows:
[0154] (5a) Filter redundant boxes: Select the bounding box with the largest score for filtering;
[0155] Select the boxes with an intersection over union (IoU) greater than 0.6 as the reference detection boxes for regression, and use the non-maximum suppression method for screening. The specific steps include:
[0156] (5a1) Sort the scores of all boxes, and select the highest score and its corresponding box:
[0157] (5a2) Traverse the remaining boxes. If the overlapping area (IoU) between the box that meets the IoU greater than the specified condition and the current highest-score box is greater than a certain threshold, delete the box;
[0158] (5a3) Continue to select a box with the highest score from the unprocessed boxes, and repeat the above process until there are no detection boxes that meet the conditions;
[0159] (5b) Perform center point optimization:
[0160] (5b1) Construct the center point definition as follows. Add a single-layer branch, parallel to the regression branch, to predict the "centerness" of a location. Given the regression targets l, t, r, and b at a location, the center target is defined as:
[0161]
[0162] Among them, l, t, r, and b are the distance values from the left, top, right, and bottom to the center point coordinates of the image, as Figure 4 shown. Specifically, a simple and effective strategy is proposed to suppress these detected low-quality bounding boxes without introducing any hyperparameters. Add a single-layer branch, parallel to the regression branch, to predict the "centerness" of a location.
[0163] (5b2) Use the binary cross-entropy (BCE) loss and the sigmoid activation function to participate in the training, and jointly calculate the loss to ensure that the calculated detection boxes are close to the center position. The cross-entropy function used is:
[0164]
[0165] Among them, M is the number of categories; y icis a sign function, taking values of 0 or 1; if the label category of sample i is equal to C, then the value is 1, otherwise it is 0; p ic is the predicted probability that the observed sample i belongs to category c.
[0166] like Figure 2 As shown, the ORS of all pest objects is no larger than 1%, which indicates that all pests in the work are of small size. In addition, most categories hold an ORS close to 0.5%, which means similar scale difficulty in similar pest dataset tasks.
[0167] In order to test the effectiveness of the present invention, Faster R-CNN, Fcos and YOLO3 are selected for comparison with the present invention. The pest detection results are shown in Table 3. It can be observed that the present method is superior to Faster R-CNN and YOLO3. The mAP of the present method can reach 45%, which is 14.2% higher than YOLO3 and 3.1% higher than Faster R-CNN. For extremely special pests ("21" and "23" categories), the detection accuracy is lower than that of other types of pests. However, the present method is still significantly better than YOLO3 and Faster R-CNN, thanks to the feature fusion module. The detection results are shown in Table 4:
[0168] Table 3 Comparison results of DACFA with other detection methods
[0169]
[0170] Table 4 AP50 and all categories of pests using different detection methods on similar pest datasets (unit: %).
[0171]
[0172]
[0173] In order to directly observe the advantages of the pest detection method proposed in the present invention, some visualized pest detection results of the present invention, YOLO3 and FasterR-CNN are given, such as Figure 5 As shown, this method has higher accuracy and fewer misses than other methods.
[0174] In summary, the present invention has the function of effectively identifying pest datasets with high appearance similarity and small size difference; introducing a dual-channel attention mechanism, using a non-local module to reconstruct the feature extraction method, and adding a jump convolution module to reconstruct the feature pyramid; the present invention is a one-stage anchor-free detector that can achieve real-time detection effects.
Claims
1. An adaptive feature fusion detection method for similar pest images under the lamp, characterized in that: The method includes the following steps in sequence: (1) Obtain a dataset of similar pest images and perform preprocessing; (2) Extract the similarity features of pests; (3) Construct a non-local similar pest detection network with a joint dual-channel attention mechanism, and input the pest similarity features extracted in step (2) into the non-local similar pest detection network; (4) Design a jumping convolution module pyramid feature network with a balanced size, use the output of the non-local similar pest detection network as the input of the jumping convolution module pyramid feature network, and the jumping convolution module pyramid feature network outputs reference detection boxes that meet the conditions; (5) Optimize the center points of the selected reference detection boxes and use the non-maximum suppression method for duplicate removal optimization; (6) Output the detection results of similar pests.
2. The adaptive feature fusion detection method for similar pest images under the lamp according to claim 1, wherein: Step (1) specifically includes the following steps: (1a) Obtain a dataset of similar pest images, segment the image data in the dataset of similar pest images, and perform size normalization preprocessing on all single-class pest images segmented out; (1b) Randomly select 100 single-class pest images that have undergone size normalization preprocessing for similarity statistics.
3. The adaptive feature fusion detection method for similar pest images under the lamp according to claim 1, wherein: Step (2) specifically includes the following steps: (2a) Extract the similarity texture of pests: In the RGB color space, use the Sobel operator to calculate for each of the three color channels, and return two corresponding vectors in the horizontal and vertical directions in each channel: vector a (Rx, Gx, Bx) and vector b (Ry, Gy, By), and calculate through formulas (1) to (4): a·b = R x ·R y +G x ·G y +B x ·B y (3) θ is the cosine angle between vectors a and b; Perform the measurement of the similarity color of pests: Quantize the results obtained from the R, G, and B channels into images of 64 colors respectively, perform texture detection in the image C(x, y) with 4 different primitives respectively to obtain the texture primitive image T(x, y), and finally describe the texture features with a multi-primitive histogram according to T(x, y). The definitions of the multi-primitive histogram are shown in formulas (5) and (6): H(T(P1)) = N{θ(P1) = v1 ∩ θ(P2) = v2 ||P1 - P2|| = D} (5) H(T(P1)) = N{θ(P1) = w1 ∩ θ(P2) = w2 ||P1 - P2|| = D} (6) Among them, P1 = (x1, y1) and P2 = (x2, y2) represent two adjacent pixel points with a distance of D in the original image. Their corresponding pixel points in the primitive image are T(P1) = w1 and T(p2) = w2 respectively. In the texture direction matrix θ(x, y), the directions of points P1 and P2 are θ(P1) = v1 and θ(P2) = v2 respectively; N represents the number of times v1 and v2 appear together, represents the number of times w1 and w2 appear together. v1 and v2 are four test primitives of image C(x, y), and w1 and w2 are the corresponding pixel points in the texture primitive image T(x, y); H(T(P1)) represents the number of times the same edge direction appears simultaneously under a certain color background; H(T(P1)) represents the number of times the same color appears under a certain edge direction; The calculation formula of the texture feature vector fv of the image is as follows: f(v) = H(T(P1)) ○ H(θ(P1)) where ○ represents concatenation, and the similarity between images I1 and I2 is defined as follows: S I (I1, I2) = ||f v (I1) - f v (I2)|| -1 (2b) Extract the similarity size of pests: Given an RGB image with the shape of H×W and the i-th pest bounding box Hi×Wi, the target relative size ORSi of this pest object is defined as: Calculate the ORS of the c-th category in the entire dataset in the following way: where M is the number of pest objects, and the function sgn(·) represents that the category of the i-th pest is the c-th category, and is defined as: Finally, obtain the ORS distribution map of pest species of all categories.
4. The adaptive feature fusion detection method for similar pest images under the lamp according to claim 1, wherein: Step (3) specifically includes the following steps: (3a) Set the input of the non-local similar pest detection network as the candidate region of the pest image and the output as the feature result of the pest image; (3b) Set the first layer of the non-local similar pest detection network to extract multi-scale local features of pest targets by the method of max pooling: According to the position and size of each pest feature region, extract the corresponding local features of pest targets in the similar pest feature map; (3c) Set the second layer of the non-local similar pest detection network to calculate the feature weights through a dual-channel attention module: According to the size of each pest feature region, calculate the corresponding feature weights of each pest feature region in the feature maps of each scale through the channel attention function: (3d) Set the third layer of the adaptive similar pest detection network to non-local feature fusion of similar pest targets: Use the weights calculated by the tg non-local feature function to fuse the similar pest features: Calculate the function output of each pest feature region, introduce the non-local feature fusion function, and construct a feature fusion that combines the features of the second layer of the extracted features with the features of the adjacent upper and lower layers. The formula of the non-local feature fusion function is as follows: Wherein, the input is x, the output is y, i and j respectively represent a certain spatial position of the input, and X i is a vector, the dimension of which is the same as the number of channels of X, f is a function for calculating the similarity relationship between any two points, and g is a mapping function that maps a point into a vector.
5. The adaptive feature fusion detection method for similar pest images under the lamp according to claim 1, wherein: The specific steps of step (4) are as follows: (4a) Add a skip convolution module to the pyramid feature network. Given a set of filter sets K with a shape of (C, C, kh, kw), the filter set K is composed of K1, K2, K3, K4, and K5, which are divided into two branches and perform different functions through K1, K2, K3, K4, and K5. In the skip convolution module, perform feature transformation on the original scale and the small size after downsampling. For the given X, it is divided into two parts, one part is X1, and the other part is X2; C is the number of channels, K1 to K5 are filters; kh is the height, and kw is the width; (4b) Use max pooling to reduce the scale: M1 = MaxPool r (X1) T1 = MaxPool r (M1) Among them, r is the downsampling rate and stride of the pooling process; MaxPool() is the max pooling function; (4c) T1 is used as the input of filters K2 and K3, and after the upsampling operation, the features are restored to the original scale, and the result is: X1' = Up(F2(T1)) = Up(T1 × K2) X1” = Up(F3(X1')) = Up(X1' × K3) Among them, Up() is the upsampling function, and X1’ is the result of two upsamplings of the product operation of filter K2 and T1; Among them, F2(T1)) = T1 × K2, F3(X1') = X1' × K3 is a simplified form of convolution. Then, the calibration operation is expressed as: Y1' = F4(X1) ⊙ Sigmoid(X1”) Among them, Y1’ is the dot product form of the value of the feature vector X1” passing through the activation function Sigmoid and X1 passing through the convolution function; F4(X1) = X1 × K4, Sigmoid(·) is an activation function, and the final result of calculating the skip convolution module is: Y1 = F5(Y1') F5(Y1')) = Y1' × K2 Among them, Y1 is the vector value of the feature vector Y1’ passing through the convolution function vector; (4d) Y2 = F(X2) = X2 × K1 Among them, Y2 is the vector value of the product of the feature vector X2 and filter K1; (4e) Sum Y1 and Y2 to obtain the final result Y, completing the construction of the skip convolution module pyramid feature network with the skip convolution module added; (4f) Output pest classification and reference detection boxes: Calculate the classification labels according to the Softmax regression model, sort them according to the confidence level, and sequentially select the values whose probability exceeds the confidence interval from largest to smallest; then use the loss function to train the non-local similarity pest detection network to output the reference detection boxes with classification labels.
6. The adaptive feature fusion detection method for similar pest images under the lamp according to claim 1, characterized in that: Step (5) specifically includes the following steps: (5a) Filter redundant boxes: Select the bounding box with the largest score for filtering; Select the boxes with an intersection over union greater than 0.6 as the reference detection boxes for regression, and use the non-maximum suppression method for screening. The specific steps include: (5a1) Sort the scores of all boxes, and select the highest score and its corresponding box: (5a2) Traverse the remaining boxes. If the overlapping area IOU between the box that meets the intersection over union greater than the specified condition and the current highest score box is greater than a certain threshold, delete the box; (5a3) Continue to select a box with the highest score from the unprocessed boxes, and repeat the above process until there are no detection boxes that meet the conditions; (5b) Perform center point optimization: (5b1) Define the center point as adding a single-layer branch parallel to the regression branch to predict the "centrality" of a position. Given the regression targets l, t, r, and b at a position, the center target is defined as: where l, t, r, and b are the distance values from the left, top, right, and bottom ends to the center point coordinates of the image respectively; (5b2) Use the binary cross-entropy BCE loss and the sigmoid activation function to participate in the training and jointly calculate the loss to ensure that the calculated detection boxes are close to the center position. The cross-entropy function used is: where M is the number of categories; y ic is the sign function, taking values of 0 or 1; if the label category of sample i is equal to C, the value is 1, otherwise it is 0; p ic is the predicted probability that the observed sample i belongs to category c.
Citation Information
Patent Citations
Agricultural pest image recognition method based on multi-feature deep learning technology
CN105488536A
Pest image detection method with similar size enhanced recognition
CN112733614A