An object detection method and system integrating global information of images
The multi-layer feature extraction network with dual pathways and self-attention optimization enhances pollen detection precision by addressing size variation issues in flower pollen images, improving detection accuracy.
Patent Information
- Application Number
- CN202111158075.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-28
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-09-28
AI Technical Summary
Current methods for pollen detection in images, such as using natural image detection techniques in flower pollen images, suffer from low precision due to the significant size variation of pollen grains, leading to grouped grains being detected as single grains, and existing automated detection methods are labor-intensive and time-consuming.
A method involving a multi-layer feature extraction network that fuses image features through dual pathways, utilizing self-attention modules for optimization, and a flower pollen detection module trained on sample features to enhance detection accuracy.
The method improves pollen detection precision by avoiding large anchor boxes and aggregating global context information, enhancing feature representation to accurately identify and classify pollen grains.
Smart Images

Figure CN113763381B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to an object detection method and system that fuse global image information. Background Art
[0002] Pollen is one of the most common allergens. When pollen comes into contact with a person's nose or eyes, it can trigger typical allergic reactions such as runny nose, nasal congestion, sneezing, tearing, and itchy eyes. In recent years, the number of people with pollen allergies has been increasing, making accurate pollen forecasts more necessary. Currently, the main method of pollen forecasting is to rely on staff to identify pollen grains one by one under a bright-field microscope and count the concentrations of various types of pollen. This method has a very long cycle and requires high professional knowledge of the staff, making it difficult to meet the real-time requirements of pollen forecasting.
[0003] Current automatic detection methods mainly apply detection methods that work well in natural images to pollen images. However, there are significant differences between pollen images and natural images. The sizes of objects in natural images vary widely, but the sizes of pollen grains are relatively concentrated. If overly large anchor boxes are used, clustered pollen grains will be detected as single pollen grains together, resulting in low detection accuracy. Summary of the Invention
[0004] The present invention provides an object detection method, system, electronic device, and storage medium that fuse global image information to solve the above technical problems and can effectively improve detection accuracy.
[0005] The present invention provides an object detection method that fuses global image information, including:
[0006] Inputting the acquired image to be detected into a trained multi-layer feature extraction network for feature extraction to obtain multi-layer feature results of the image to be detected; wherein, the multi-layer feature extraction network is trained based on sample images and corresponding sample feature results;
[0007] Performing dual-channel feature fusion on the multi-layer feature results to obtain multi-layer fusion feature results; wherein, the multi-layer fusion feature results include shallow fusion feature results, deep fusion feature results, and middle-layer fusion feature results corresponding to a preset detection branch layer;
[0008] Fusing the shallow fusion feature results with optimized features and the deep fusion feature results with optimized features into the middle-layer fusion feature results to obtain target fusion feature results;
[0009] Input the target feature fusion result into the trained pollen detection module to obtain the pollen detection result output by the pollen detection module; wherein, the pollen detection module is trained based on feature samples and corresponding sample detection results.
[0010] According to the object detection method for fusing global image information of the present invention, in the step of performing dual-channel feature fusion on the multi-layer feature results to obtain multi-layer fusion feature results, it specifically includes:
[0011] Perform feature fusion on the multi-layer feature results through the upsampling fusion path to obtain the shallow-layer fusion feature result, and perform feature fusion on the shallow-layer fusion feature result through the downsampling fusion path to obtain the middle-layer fusion feature result and the deep-layer fusion feature result.
[0012] According to the object detection method for fusing global image information of the present invention, the feature optimization methods for the shallow-layer fusion feature result and the deep-layer fusion feature result include:
[0013] Input the shallowest fusion feature in the shallow-layer fusion feature result into a preset parallel self-attention module for feature optimization to obtain the feature-optimized shallow-layer fusion feature result, and input the deepest fusion feature in the deep-layer fusion feature result into the parallel self-attention module for feature optimization to obtain the feature-optimized deep-layer fusion feature result; wherein, the parallel self-attention module includes a spatial self-attention unit and a channel self-attention unit.
[0014] According to the object detection method for fusing global image information of the present invention, the middle-layer fusion feature result includes a first middle-layer fusion feature result and a second middle-layer fusion feature result; the target fusion feature result includes a first fusion feature result and a second fusion feature result;
[0015] In the step of fusing the feature-optimized shallow-layer fusion feature result and the feature-optimized deep-layer fusion feature result into the middle-layer fusion feature result to obtain the target fusion feature result, it specifically includes:
[0016] Fuse the feature-optimized shallow-layer fusion feature result with the first middle-layer fusion feature result to obtain the first fusion feature result, and fuse the feature-optimized deep-layer fusion feature result with the second middle-layer fusion feature result to obtain the second fusion feature result.
[0017] According to the object detection method for fusing global image information of the present invention, the pollen detection module includes a position detection unit and a type detection unit; the pollen detection result includes a classification detection result and a position detection result;
[0018] In the step of inputting the target feature fusion result into the trained pollen detection module to obtain the pollen detection result output by the pollen detection module, it specifically includes:
[0019] Input the target feature fusion result into the position detection unit of the trained pollen detection module to obtain the position detection result; input the target feature fusion result into the type detection unit of the trained pollen detection module to obtain the classification detection result.
[0020] According to the object detection method for fusing global image information of the present invention, the upsampling fusion path performs upsampling operations using bilinear interpolation method or transposed convolution method; the downsampling path performs downsampling operations using convolution method or pooling method.
[0021] According to the object detection method for fusing global image information of the present invention, during the training process of the type detection unit, the focal loss function is used to calculate the loss value; during the training process of the position detection unit, the smoothL1 loss function is used to calculate the loss value.
[0022] The present invention also provides an object detection system for fusing global image information, including:
[0023] A feature extraction module, configured to input the acquired image to be detected into the trained multi-layer feature extraction network to extract the multi-layer feature results of the image to be detected; wherein, the multi-layer feature extraction network is trained based on the sample image and the corresponding sample feature results.
[0024] A first fusion module, configured to perform dual-path feature fusion on the multi-layer feature results to obtain multi-layer fusion feature results; wherein, the multi-layer fusion feature results include shallow fusion feature results, deep fusion feature results, and middle fusion feature results corresponding to the preset detection branch layer.
[0025] A second fusion module, configured to fuse the shallow fusion feature results with optimized features and the deep fusion feature results with optimized features into the middle fusion feature results to obtain the target fusion feature results.
[0026] A detection module, configured to input the target feature fusion result into the trained pollen detection module to obtain the pollen detection result output by the pollen detection module; wherein, the pollen detection module is trained based on the feature samples and the corresponding sample detection results.
[0027] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the object detection method for fusing global information of the fused image as described in any one of the above are implemented.
[0028] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the object detection method for fusing global information of the fused image as described in any one of the above are implemented.
[0029] The object detection method, system, electronic device, and storage medium for fusing global information of the fused image provided by the present invention perform dual-path feature fusion on the initially extracted multi-layer feature results, fuse the deep fused features and the shallow fused features into the features corresponding to the detection branch, and then use the fused features corresponding to the detection branch for pollen detection. This avoids using a detection branch with an overly large anchor box, prevents detecting multiple clustered pollen grains as single pollen grains, effectively improves the detection accuracy, and at the same time aggregates the global context information of the image through dual-path feature fusion, improving the expressiveness of the features, thereby contributing to improving the accuracy of pollen detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0031] Figure 1 It is a flowchart of the object detection method for fusing global information of the fused image provided by an embodiment of the present invention;
[0032] Figure 2 It is a technical route diagram of the object detection method for fusing global information of the fused image provided by an embodiment of the present invention;
[0033] Figure 3 It is a structural diagram of the object detection system for fusing global information of the fused image provided by an embodiment of the present invention;
[0034] Figure 4 It is a structural diagram of the electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.
[0036] As Figure 1 shown, an object detection method integrating global image information provided by an embodiment of the present invention includes the steps of:
[0037] S1. Input the acquired image to be measured into a trained multi-layer feature extraction network for feature extraction to obtain multi-layer feature results of the image to be measured; wherein, the multi-layer feature extraction network is trained based on a sample image and sample feature results corresponding to the sample image;
[0038] It should be noted that step S1 is to perform preliminary feature extraction on the image to be measured. In the embodiment of the present invention, multi-layer feature results are obtained through feature extraction by the multi-layer feature extraction network, and the multi-layer feature results include image features corresponding to multiple levels.
[0039] S2. Perform dual-channel feature fusion on the multi-layer feature results to obtain multi-layer fusion feature results; wherein, the multi-layer fusion feature results include shallow fusion feature results, deep fusion feature results, and middle fusion feature results corresponding to a preset detection branch layer.
[0040] In the embodiment of the present invention, further, step S2 specifically includes:
[0041] Perform feature fusion on the multi-layer feature results through an upsampling fusion path to obtain the shallow fusion feature results, and perform feature fusion on the shallow fusion feature results through a downsampling fusion path to obtain the middle fusion feature results and the deep fusion feature results.
[0042] In the embodiment of the present invention, further, the upsampling fusion path performs an upsampling operation using the bilinear interpolation method or the transposed convolution method; the downsampling path performs a downsampling operation using the convolution method or the pooling method.
[0043] It should be noted that step S2 is to perform dual-channel feature fusion on the image features of each layer. Specifically, first perform fusion through the upsampling fusion path to obtain the shallow fusion features corresponding to each layer, and then perform further fusion on these shallow fusion features through the downsampling fusion path to obtain the deep fusion features corresponding to each layer. Preferably, the middle two layers are used as the detection branch layer in the embodiment of the present invention. Therefore, the middle fusion feature results in the embodiment of the present invention refer to the fusion features corresponding to these two detection branch layers.
[0044] S3. Integrate the shallow fusion feature result with optimized features and the deep fusion feature result with optimized features into the middle-level fusion feature result to obtain the target fusion feature result.
[0045] In an embodiment of the present invention, further, the methods for optimizing the features of the shallow fusion feature result and the deep fusion feature result include:
[0046] Input the shallowest fusion feature in the shallow fusion feature result into a preset parallel self-attention module for feature optimization to obtain the shallow fusion feature result with optimized features, and input the deepest fusion feature in the deep fusion feature result into the parallel self-attention module for feature optimization to obtain the deep fusion feature result with optimized features; wherein, the parallel self-attention module includes a spatial self-attention unit and a channel self-attention unit.
[0047] In an embodiment of the present invention, the middle-level fusion feature result includes a first middle-level fusion feature result and a second middle-level fusion feature result; the target fusion feature result includes a first fusion feature result and a second fusion feature result; further, step S3 specifically includes:
[0048] Integrate the shallow fusion feature result with optimized features and the first middle-level fusion feature result to obtain the first fusion feature result, and integrate the deep fusion feature result with optimized features and the second middle-level fusion feature result to obtain the second fusion feature result.
[0049] In an embodiment of the present invention, after obtaining the shallow fusion feature result and the deep fusion feature result in step S2, it is necessary to use a parallel spatial and channel self-attention module to optimize the shallowest fusion feature and the deepest fusion feature, and then further integrate the optimized shallowest fusion feature and the deepest fusion feature into the middle-level fusion feature result to obtain the final target fusion feature result.
[0050] S4. Input the target feature fusion result into the trained pollen detection module to obtain the pollen detection result output by the pollen detection module; wherein, the pollen detection module is trained based on feature samples and the corresponding sample detection results.
[0051] In an embodiment of the present invention, further, the pollen detection module includes a position detection unit and a type detection unit; the pollen detection result includes a classification detection result and a position detection result; further, in the training process of the type detection unit, the focal loss function is used to calculate the loss value; in the training process of the position detection unit, the smooth L1 loss function is used to calculate the loss value.
[0052] Further, step S4 specifically includes:
[0053] Input the target feature fusion result into the position detection unit of the trained pollen detection module to obtain the position detection result; input the target feature fusion result into the type detection unit of the trained pollen detection module to obtain the classification detection result.
[0054] It should be noted that step S4 is to input the final target fusion feature result into the trained pollen detection module for pollen detection, including position detection and type detection.
[0055] Based on the object detection method for fusing image global information provided in the above embodiments, the following is an example of the specific implementation steps of the solution of the present invention:
[0056] Please refer to Figure 2 , the embodiment of the present invention is for the pollen detection process of bright-field microscope images, and the specific steps include:
[0057] 1. Image input, input the processed microscope image into the feature extraction network;
[0058] 2. Use the feature extraction network (convolutional neural network) to extract image features. The network is divided into different layers according to different downsampling multiples, and image features of different layers are extracted correspondingly for each layer;
[0059] 3. Fuse the multi-layer image features. First, use two fusion paths from deep to shallow and from shallow to deep to fuse the deep features and shallow features. Then, after optimizing the fusion features of the deepest layer and the shallowest layer using parallel spatial and channel self-attention modules, fuse them into the image features of the detection branch;
[0060] 4. The feature maps of the two detection branches are respectively input into the pollen prediction part (pollen detection module) to perform pollen classification prediction (type detection) and pollen grain position prediction (position detection).
[0061] The specific algorithm is as follows:
[0062] (1) Image input
[0063] The pollen images obtained by the bright-field microscope are very large and need to be cropped into images of size 512×512 first. It should be noted that during the training process, only two data augmentation methods, horizontal flipping and random cropping, are used, and color transformation is not used because the pollen grains in the images are dyed pink, which is an important feature for pollen recognition.
[0064] (2) Feature extraction
[0065] In the embodiment of the present invention, the feature extraction network selected is the ResNet-50 network. According to the increasing downsampling ratio of the network, the feature maps are set as P3, P4, P5, and P6, corresponding to downsampling ratios of 8, 16, 32, and 64 times respectively. Among them, P3, P4, and P5 respectively correspond to the feature maps of the C3, C4, and C5 parts of the ResNet-50 network, while P6 is obtained by performing convolution calculation on P5 through a convolutional kernel with a size of 3×3 and a stride of 2. At this time, the four feature maps P3 to P6 generated are input into the feature fusion module for feature fusion. It should be noted that the sizes of pollen grains are relatively concentrated and the differences are small. Using a detection branch with an overly large anchor box may cause multiple clustered pollen grains to be regarded as a single pollen grain, thus affecting the calculation of pollen quantity. Therefore, we only retain the detection branches corresponding to the P4 and P5 levels, and the corresponding anchor box sizes are 64 2 pixels and 128 2 pixels. The embodiment of the present invention uses three aspect ratios {1:2, 1:1, 2:1} and three scaling ratios {0.6, 1.3, 1.7}, so the anchor box range is [38, 218] pixels, which can cover the pollen size range [45, 190] pixels in the image.
[0066] (3) Multi-layer feature fusion
[0067] The purpose of using multi-layer feature fusion in the embodiment of the present invention is to fuse the shallow features containing more structural texture information and the deep features containing more abstract semantic information in the convolutional neural network to obtain more expressive features. These optimized feature maps are more conducive to the localization and classification of pollen.
[0068] Dual-path feature fusion: The embodiment of the present invention uses a fusion path from deep to shallow (upsampling fusion path) and a fusion path from shallow to deep (downsampling fusion path).
[0069] First, in the fusion path from deep to shallow, the P6 feature map is first upsampled so that P6 has the same size as P5. Then, the upsampled and deformed P6 and P5 are subjected to an element-wise addition operation to obtain a new feature map P5 m1 . Similarly, P4 m1 and P3 m1, as shown in Formulas 1, 2, and 3. The upsample() function here is the method of upsampling. The bilinear interpolation method is selected in the embodiments of the present invention, and it can also be replaced with other upsampling operation methods such as deconvolution.
[0070] P5 m1 = P5 + upsample(P6) Formula 1
[0071] P4 m1 = P4 + upsample(P5 m1 ) Formula 2
[0072] P3 m1 = P3 + upsample(P4 m1 ) Formula 3
[0073] In the fusion path from shallow to deep, P3 m1 will first undergo downsampling operation to be deformed into a feature map of the same size as P4 m1 , and then perform element-wise addition operation with P4 m1 to obtain a new feature map P4 m2 . Similarly, P5 m2 and P6 m2 can be obtained, as shown in Formulas 4, 5, and 6. The downsample() function here uses a convolution operation with a size of 3×3 and a stride of 2. It can also be replaced with downsampling operations such as pooling.
[0074] P4 m2 = P4 m1 + downsample(P3 m1 ) Formula 4
[0075] P5 m2 = P5 m1 + downsample(P4 m2 ) Formula 5
[0076] P6 m2 = P6 + downsample(P5 m2 ) Formula 6
[0077] At this time, both P3 m1 and P6 m2 have fused the spatial information of shallow features and the abstract semantic information of deep features. Since we only retain the detection branches corresponding to the P4 and P5 levels. Therefore, both P3 m1 and P6 m2 are used as supplementary features and fused into P4 m2 and P5 m2In the feature map. However, to improve the expressiveness of the features, we use a parallel spatial self-attention module and channel self-attention module to optimize P3 m1 and P6 m2 for feature optimization.
[0078] Feature optimization of the self-attention module: P3 m1 and P6 m2 will be optimized by the parallel spatial self-attention module and channel self-attention module respectively. Among them, the optimized features P3 out and P6 out are shown in Equation 7 and Equation 8.
[0079] Equation 7:
[0080]
[0081] The first term of Equation 7 here represents the calculation process of the spatial self-attention module. Among them and represent the i-th and j-th pixel points of the P3 m1 feature map respectively, N represents the total number of pixel points in the feature map, and exp() represents the exponential function with the natural constant e as the base used in the softmax function. The second term of Equation 7 represents the calculation process of the channel self-attention module. and represent the b-th and g-th channels of the P3 m1 feature map respectively, C represents the total number of channels in the feature map. α1 and β1 represent the weights of the spatial and channel self-attention modules for calculating features respectively. The feature maps calculated by the two modules will be weighted into the original feature map P3 m1 . To ensure the stability of the features, α1 and β1 are initialized to 0 and increase during the training process to aggregate more global information.
[0082] Equation 8:
[0083]
[0084] Similarly, the first term of Equation 8 represents the calculation process of the spatial self-attention module. Among them and represent the i-th and j-th pixel points of the P6 m2 feature map respectively, N represents the total number of pixel points in the P6 m2 feature map, and exp() represents the exponential function with the natural constant e as the base used in the softmax function. The second term of Equation 8 represents the calculation process of the channel self-attention module. and represent the b-th and g-th channels of the P6 m2The b-th and g-th channels of the feature map, where C represents the total number of channels of the feature map. α1 and β1 represent the weights of the features calculated by the spatial and channel self-attention modules respectively. The feature maps calculated by the two modules will be weighted to the original feature map P6 m2 . To ensure the stability of the features, α1 and β1 are initialized to 0 and increase during the training process to aggregate more global information.
[0085] Feature fusion: The optimized feature P3 out will be fused with P4 m2 , and P6 out will be fused with P5 m2 , as shown in Formula 9 and Formula 10.
[0086] P4 out = P4 m2 + downsample(P3 out ) Formula 9
[0087] P5 out = P5 m2 + upsample(P6 out ) Formula 10
[0088] Obtain P4 out and P5 out and input them into the corresponding detection branches for class prediction and location prediction.
[0089] (4) Pollen prediction
[0090] Classification prediction: Predict whether the current detection box contains pollen and the specific category of the pollen. The classification branch includes 4 convolutional layers of 3×3×256, and finally a prediction layer of 1×1×KA. Here, K represents the number of pollen categories to be predicted, and A represents the number of anchor boxes.
[0091] During the training process, to improve the classification accuracy of pollen grains that are difficult to classify. This patent uses the focal loss function to calculate the classification loss value. As shown in Formula 11,
[0092] FL(p t ) = -αt(1 - p t ) γ log(p t ) Formula 11
[0093] At this time, p t represents the classification probability of each category, γ is a value greater than 0, αt is a decimal between [0,1], both γ and αt are fixed values and do not participate in training. The optimal values of γ and αt affect each other, so they need to be adjusted in combination when evaluating the accuracy. In the embodiment of the present invention, γ = 2 and αt = 0.3 are used. When pt When it is larger, (1 - p t ) becomes smaller, and the classification loss of the target objects that are easy to distinguish will be suppressed, making the model pay more attention to the pollen that is difficult to distinguish. Thus, the pollen detection accuracy is effectively improved.
[0094] Location prediction: Predict the coordinate offset value of the anchor box. The location prediction branch includes 4 convolutional layers of 3×3×256, and finally a prediction layer of 1×1×4A. Here, A represents the number of anchor boxes, and the 4 prediction values are (p x , p y , p w , p h ), p x and p y represent the coordinates of the center point of the predicted bounding box, and p w and p h represent the width and height.
[0095] During the training process, the smooth L1 loss function is used for location prediction. If the true bounding box is (gt x , gt y , gt w , gt h ), then the error bias between the predicted value and the true value is shown in Equation 12:
[0096] bias = |p x - gt x | + |p y - gt y | + |p w - gt w | + |p h - gt h Equation 12
[0097] Then the smooth L1 loss function is as shown in Equation 13:
[0098]
[0099] Compared with the prior art, the embodiments of the present invention design two detection branches and anchor box ratios for the pollen particle measurement size range. Avoiding the detection branch with an overly large anchor box can prevent multiple pollen particles clustered together from being detected as single pollen grains, resulting in calculation errors in pollen density. At the same time, two feature fusion paths are used in the feature fusion part, and feature optimization is carried out through the spatial and channel self-attention modules, thereby aggregating more global context information, improving the expression ability of the features, and contributing to improving the pollen detection accuracy.
[0100] The object detection system that fuses global image information provided by the present invention will be described below. The object detection system that fuses global image information described below can be correspondingly referred to the object detection method that fuses global image information described above.
[0101] Please refer to Figure 3 , an embodiment of the present invention provides an object detection system that fuses global image information, including:
[0102] A feature extraction module 1, configured to input the acquired image to be detected into a trained multi-layer feature extraction network for feature extraction to obtain multi-layer feature results of the image to be detected; wherein, the multi-layer feature extraction network is trained based on a sample image and sample feature results corresponding to the sample image;
[0103] A first fusion module 2, configured to perform dual-channel feature fusion on the multi-layer feature results to obtain multi-layer fusion feature results; wherein, the multi-layer fusion feature results include shallow fusion feature results, deep fusion feature results, and middle fusion feature results corresponding to a preset detection branch layer;
[0104] A second fusion module 3, configured to fuse the shallow fusion feature results with optimized features and the deep fusion feature results with optimized features into the middle fusion feature results to obtain target fusion feature results;
[0105] A detection module 4, configured to input the target feature fusion results into a trained pollen detection module to obtain pollen detection results output by the pollen detection module; wherein, the pollen detection module is trained based on feature samples and sample detection results corresponding to the feature samples.
[0106] In an embodiment of the present invention, further, the first fusion module 2 is specifically configured to:
[0107] Perform feature fusion on the multi-layer feature results through an upsampling fusion path to obtain the shallow fusion feature results, and perform feature fusion on the shallow fusion feature results through a downsampling fusion path to obtain the middle fusion feature results and the deep fusion feature results.
[0108] In an embodiment of the present invention, further, the feature optimization methods for the shallow fusion feature results and the deep fusion feature results include:
[0109] Input the shallowest fusion feature in the shallow fusion feature result into a preset parallel self-attention module for feature optimization to obtain the shallow fusion feature result after feature optimization, and input the deepest fusion feature in the deep fusion feature result into the parallel self-attention module for feature optimization to obtain the deep fusion feature result after feature optimization; wherein, the parallel self-attention module includes a spatial self-attention unit and a channel self-attention unit.
[0110] In an embodiment of the present invention, further, the middle fusion feature result includes a first middle fusion feature result and a second middle fusion feature result; the target fusion feature result includes a first fusion feature result and a second fusion feature result;
[0111] The second fusion module 3 is specifically used for:
[0112] Fuse the shallow fusion feature result after feature optimization with the first middle fusion feature result to obtain the first fusion feature result, and fuse the deep fusion feature result after feature optimization with the second middle fusion feature result to obtain the second fusion feature result.
[0113] In an embodiment of the present invention, further, the pollen detection module includes a position detection unit and a type detection unit; the pollen detection result includes a classification detection result and a position detection result;
[0114] The detection module 4 is specifically used for:
[0115] Input the target feature fusion result into the position detection unit of the trained pollen detection module to obtain the position detection result; input the target feature fusion result into the type detection unit of the trained pollen detection module to obtain the classification detection result.
[0116] In an embodiment of the present invention, further, the upsampling fusion path performs upsampling operations using the bilinear interpolation method or the transposed convolution method; the downsampling path performs downsampling operations using the convolution method or the pooling method.
[0117] In an embodiment of the present invention, further, the training process of the type detection unit calculates the loss value using the focal loss function; the training process of the position detection unit calculates the loss value using the smooth L1 loss function.
[0118] The working principle of the object detection system for fusing image global information in the embodiments of this case is corresponding to the object detection method for fusing image global information in the above embodiments, and will not be elaborated here one by one.
[0119] Figure 4Illustrates a schematic diagram of the physical structure of an electronic device, as follows Figure 4 As shown, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communication bus 440. Among them, the processor 410, the communications interface 420, and the memory 430 complete mutual communication through the communication bus 440. The processor 410 may call the logical instructions in the memory 430 to execute the object detection method for fusing global image information. The method includes: inputting the acquired image to be measured into a trained multi-layer feature extraction network for feature extraction to obtain the multi-layer feature results of the image to be measured; wherein, the multi-layer feature extraction network is trained based on a sample image and the corresponding sample feature results of the sample image;
[0120] Performing dual-channel feature fusion on the multi-layer feature results to obtain multi-layer fusion feature results; wherein, the multi-layer fusion feature results include shallow fusion feature results, deep fusion feature results, and middle fusion feature results corresponding to a preset detection branch layer;
[0121] Fusing the shallow fusion feature results with optimized features and the deep fusion feature results with optimized features into the middle fusion feature results to obtain target fusion feature results;
[0122] Inputting the target feature fusion results into a trained pollen detection module to obtain the pollen detection results output by the pollen detection module; wherein, the pollen detection module is trained based on feature samples and the corresponding sample detection results of the feature samples.
[0123] In addition, when the logical instructions in the above-mentioned memory 430 are implemented in the form of software functional units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.
[0124] On the other hand, the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the object detection method for fusing global image information provided in the above embodiments. The method includes: inputting the acquired image to be detected into a trained multi-layer feature extraction network for feature extraction to obtain multi-layer feature results of the image to be detected; wherein, the multi-layer feature extraction network is trained based on a sample image and sample feature results corresponding to the sample image;
[0125] Performing dual-channel feature fusion on the multi-layer feature results to obtain multi-layer fusion feature results; wherein, the multi-layer fusion feature results include shallow fusion feature results, deep fusion feature results, and middle fusion feature results corresponding to a preset detection branch layer;
[0126] Fusing the shallow fusion feature results with optimized features and the deep fusion feature results with optimized features into the middle fusion feature results to obtain target fusion feature results;
[0127] Inputting the target feature fusion results into a trained pollen detection module to obtain pollen detection results output by the pollen detection module; wherein, the pollen detection module is trained based on feature samples and sample detection results corresponding to the feature samples.
[0128] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the object detection method for fusing global image information provided in the above embodiments. The method includes: inputting the acquired image to be detected into a trained multi-layer feature extraction network for feature extraction to obtain multi-layer feature results of the image to be detected; wherein, the multi-layer feature extraction network is trained based on a sample image and sample feature results corresponding to the sample image;
[0129] Performing dual-channel feature fusion on the multi-layer feature results to obtain multi-layer fusion feature results; wherein, the multi-layer fusion feature results include shallow fusion feature results, deep fusion feature results, and middle fusion feature results corresponding to a preset detection branch layer;
[0130] Fusing the shallow fusion feature results with optimized features and the deep fusion feature results with optimized features into the middle fusion feature results to obtain target fusion feature results;
[0131] Input the target feature fusion result into the trained pollen detection module to obtain the pollen detection result output by the pollen detection module; wherein, the pollen detection module is trained based on feature samples and corresponding sample detection results.
[0132] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative effort.
[0133] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0134] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or equivalently replace some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An object detection method that integrates global image information, characterized in that Including: Input the acquired image to be measured into the trained multi-layer feature extraction network for feature extraction to obtain the multi-layer feature results of the image to be measured; wherein, the multi-layer feature extraction network is trained based on the sample image and the corresponding sample feature results, the multi-layer feature extraction network is a ResNet-50 network, and the multi-layer feature results include feature map P3, feature map P4, feature map P5 and feature map P6. Among them, the downsampling ratios of the feature map P3, the feature map P4, the feature map P5 and the feature map P6 increase sequentially; Perform dual-channel feature fusion on the multi-layer feature results to obtain multi-layer fusion feature results; wherein, the multi-layer fusion feature results include shallow fusion feature results, deep fusion feature results and middle fusion feature results corresponding to the preset detection branch layer; Fuse the shallow fusion feature results with optimized features and the deep fusion feature results with optimized features into the middle fusion feature results to obtain the target feature fusion results; Input the target feature fusion results into the trained pollen detection module to obtain the pollen detection results output by the pollen detection module; wherein, the pollen detection module is trained based on the feature samples and the corresponding sample detection results; Among them, the shallow fusion feature result includes feature map P5 m1 , feature map P4 m1 and feature map P3 m1 . The deep fusion feature result includes feature map P6 m2 . The middle fusion feature result corresponding to the preset detection branch layer includes feature map P4 m2 and feature map P5 m2 ; In the step of performing dual-channel feature fusion on the multi-layer feature results to obtain multi-layer fusion feature results, it specifically includes: The feature map P6 is upsampled and then element-wise added to the feature map P5 to obtain the feature map P5 m1 ; Upsample the feature map P5 m1 and perform element-wise addition with the feature map P4 to obtain the feature map P4 m1 ; Upsample the feature map P4 m1 and add it element-wise to the feature map P3 to obtain the feature map P3 m1 ; Downsample the feature map P3 m1 and add it element-wise to the feature map P4 m1 to obtain the feature map P4 m2 ; Downsample the feature map P4 m2 and add it element-wise to the feature map P5 m1 to obtain the feature map P5 m2 ; Downsample the P5 m2 and add it element-wise to the feature map P6 to obtain the feature map P6 m2 ; Among them, the feature optimization methods in the shallow fusion feature results and the deep fusion feature results include: Input the feature map P3 in the shallow fusion feature result m1 into a preset parallel self-attention module for feature optimization to obtain the feature map P3 out , input the feature map P6 in the deep fusion feature result m2 into the parallel self-attention module for feature optimization to obtain the feature map P6 out ; wherein, the parallel self-attention module includes a spatially parallel self-attention module and a channel self-attention module; The step of fusing the shallow fusion feature results with optimized features and the deep fusion feature results with optimized features into the middle fusion feature results to obtain the target feature fusion results includes: Feature map P3 out is feature-fused with the said feature map P4 m2 to obtain feature map P4 out ; Fuse the feature map P6 out with the said P5 m2 to perform feature fusion and obtain the feature map P5 out .
2. The object detection method for fusing global information of the fused image according to claim 1, wherein, The pollen detection module includes a position detection unit and a type detection unit; the pollen detection results include classification detection results and position detection results; In the step of inputting the target feature fusion results into the trained pollen detection module to obtain the pollen detection results output by the pollen detection module, it specifically includes: Input the target feature fusion results into the position detection unit of the trained pollen detection module to obtain the position detection results; input the target feature fusion results into the type detection unit of the trained pollen detection module to obtain the classification detection results.
3. The object detection method for fusing global information of a fused image according to claim 1, wherein, Use bilinear interpolation method or deconvolution method for upsampling operation; use convolution method or pooling method for downsampling operation.
4. The object detection method for fusing global information of a fused image according to claim 2, wherein In the training process of the type detection unit, the focal loss function is used to calculate the loss value; in the training process of the position detection unit, the smooth L1 loss function is used to calculate the loss value.
5. An object detection system that integrates global information of an image, characterized in that, Including: A feature extraction module for inputting the acquired image to be measured into a trained multi-layer feature extraction network to extract features and obtain multi-layer feature results of the image to be measured; wherein, the multi-layer feature extraction network is trained based on sample images and corresponding sample feature results, the multi-layer feature extraction network is a ResNet-50 network, and the multi-layer feature results include feature map P3, feature map P4, feature map P5 and feature map P6, wherein the downsampling ratios of the feature map P3, the feature map P4, the feature map P5 and the feature map P6 increase sequentially; A first fusion module for performing dual-channel feature fusion on the multi-layer feature results to obtain multi-layer fusion feature results; wherein, the multi-layer fusion feature results include shallow fusion feature results, deep fusion feature results and middle fusion feature results corresponding to a preset detection branch layer; A second fusion module for fusing the feature-optimized shallow fusion feature results and the feature-optimized deep fusion feature results into the middle fusion feature results to obtain a target feature fusion result; A detection module for inputting the target feature fusion result into a trained pollen detection module to obtain a pollen detection result output by the pollen detection module; wherein, the pollen detection module is trained based on feature samples and corresponding sample detection results; Among them, the shallow fusion feature result includes the feature map P5 m1 , the feature map P4 m1 and the feature map P3 m1 . The deep fusion feature result includes the feature map P6 m2 . The middle fusion feature result corresponding to the preset detection branch layer includes the feature map P4 m2 and the feature map P5 m2 ; In the step of performing dual-channel feature fusion on the multi-layer feature results to obtain multi-layer fusion feature results, it specifically includes: The feature map P6 is upsampled and then added to the feature map P5 element by element to obtain the feature map P5 m1 ; Upsample the feature map P5 m1 and add it element-wise to the feature map P4 to obtain the feature map P4 m1 ; After upsampling the feature map P4 m1 and performing element-wise addition with the feature map P3, the feature map P3 m1 is obtained; Downsample the feature map P3 m1 and add it element-wise to the feature map P4 m1 to obtain the feature map P4 m2 ; Downsample the feature map P4 m2 and add it element-wise to the feature map P5 m1 to obtain the feature map P5 m2 ; Downsample the P5 m2 and add it element-wise to the feature map P6 to obtain the feature map P6 m2 ; Wherein, the feature optimization methods of the shallow fusion feature results and the deep fusion feature results include: Input the feature map P3 in the shallow fusion feature result m1 into a preset parallel self-attention module for feature optimization to obtain the feature map P3 out , and input the feature map P6 in the deep fusion feature result m2 into the parallel self-attention module for feature optimization to obtain the feature map P6 out ; wherein, the parallel self-attention module includes a parallel spatial self-attention module and a channel self-attention module; Fusing the feature-optimized shallow fusion feature results and the feature-optimized deep fusion feature results into the middle fusion feature results to obtain a target feature fusion result, including: Fuse the feature map P3 out with the feature map P4 m2 to obtain the feature map P4 out ; Fuse the feature map P6 out with the said P5 m2 to obtain the feature map P5 out .
6. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the object detection method for fusing global image information according to any one of claims 1 to 4.
7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the object detection method for fusing global image information according to any one of claims 1 to 4.
Citation Information
Patent Citations
Pollen detection method based on expansion convolution pyramid and multi-scale pyramid
CN112581450A