Object Detection Method, Device and Electronic Device Incorporating Attention Mechanism
Through the object detection method that integrates attention mechanism, the coordinate channel and spatial attention module in the primary visual perception cortex module and the object detection model are used to solve the detection error problem of the object detection model in interfering scenarios, improving the accuracy of small object detection and robustness in complex scenarios.
Patent Information
- Application Number
- CN202210880449.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-25
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-07-25
AI Technical Summary
The existing deep learning-based object detection model is prone to error detection in interference scenarios, especially in the case of seasonal variations in background and object occlusion, and the detection accuracy of small targets is low.
The object detection method with a fusion attention mechanism is adopted, and the preset imitation primary visual perception cortex module and object detection model are combined with the coordinate channel attention and spatial attention module to extract the image features and filter out interference information to improve the accuracy of small object detection.
Effectively filtering out image background interference improves the accuracy and robustness of object detection, especially in complex interference scenarios.
Smart Images

Figure CN115311468B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of object detection, and more specifically, to an object detection method, device and electronic device integrating an attention mechanism. Background Art
[0002] Currently, object detection models based on deep learning adopt a feature extraction method mainly based on a deep convolutional network, which can mine the deep features of image objects and apply them to object detection, showing excellent performance on various public data sets. However, such modules show vulnerability in interference scenarios, and are prone to false detection in cases where there are seasonal differences in the background, object occlusion, etc. The model based on the attention mechanism can focus the extraction of effective features on the target object of interest and filter out relevant interferences. The typical Convolutional Block Attention Module (CBAM) realizes channel attention and spatial attention. Since the channel attention of this module averages the elements of this channel, this method will ignore the small target features on the channel, affecting the detection accuracy of small targets. Summary of the Invention
[0003] An object of the present invention is to provide a new technical solution for object detection.
[0004] According to a first aspect of the present invention, there is provided an object detection method integrating an attention mechanism, the method comprising:
[0005] Obtain an input image;
[0006] Extract features from the input image through a preset primary visual perception cortex-like module to obtain a first feature map; the preset primary visual perception cortex-like model includes a VOneBlock layer, a Conv layer, and a feature fusion layer;
[0007] Perform object detection on the first feature map through a preset object detection model to obtain three object feature maps of different sizes; the object detection model includes a fusion attention module, and the fusion attention module is used to extract coordinate channel attention features and spatial attention features;
[0008] Perform object classification and coordinate positioning on the three object feature maps of different sizes to obtain an object detection result.
[0009] Optionally, the fusion attention module includes a coordinate channel attention module and a spatial attention module;
[0010] The coordinate-channel attention module is used to extract the coordinate-channel attention features of the original feature map, and obtain the feature map processed by the coordinate-channel attention; the original feature map is the feature map input to the fusion attention module;
[0011] The spatial attention module is used to extract the spatial attention features of the feature map processed by the coordinate-channel attention.
[0012] Optionally, the coordinate-channel attention module is used to extract the coordinate-channel attention features of the original feature map, and obtain the feature map processed by the channel attention, including:
[0013] Pool the X direction and Y direction of each channel of the original feature map respectively to obtain the X-direction feature map and the Y-direction feature map;
[0014] Concatenate the X-direction feature map and the Y-direction feature map to obtain a two-dimensional channel weight feature map;
[0015] Perform two-dimensional convolution, normalization, and non-linear transformation on the two-dimensional channel weight feature map in sequence to obtain an optimized feature map;
[0016] Obtain the X-direction channel attention feature map and the Y-direction channel attention feature map according to the optimized feature map;
[0017] Obtain the feature map processed by the coordinate-channel attention according to the X-direction channel attention feature map, the Y-direction channel attention feature map, and the original feature map.
[0018] Optionally, the spatial attention module is used to extract the spatial attention features of the feature map processed by the coordinate-channel attention, including:
[0019] The spatial attention module pools the feature map output by the coordinate-channel attention module in the channel direction to obtain a spatial attention weight feature map;
[0020] Obtain the spatial attention features according to the spatial attention weight feature map and the feature map output by the coordinate-channel attention module.
[0021] Optionally, the preset object detection model includes a backbone network and a head network, and the fusion attention module is provided in both the backbone network and the head network. The first feature map is subjected to object detection through the preset object detection model to obtain three object feature maps of different sizes, including:
[0022] The backbone network performs multiple size compressions and feature extractions on the first feature map to obtain multiple backbone feature maps of different sizes;
[0023] Input the multiple backbone feature maps of different sizes into the head network to obtain the three target feature maps of different sizes.
[0024] Optionally, the backbone network includes a first Conv layer, a first C3 layer, a second Conv layer, a second C3 layer, a third Conv layer, a third C3 layer, a fourth Conv layer, and a fourth C3 layer. The fusion attention module includes a first CCASA attention layer arranged after the first C3 layer, a second CCASA attention layer arranged after the second C3 layer, and a CCASA attention layer arranged after the third C3 layer;
[0025] The multiple backbone feature maps of different sizes are obtained by performing multiple size compressions and feature extractions on the first feature map through the backbone network, including:
[0026] Perform feature map size compression on the first feature map through the first Conv layer, perform feature extraction on the feature map output by the first Conv layer through the first C3 layer, perform extraction of interesting features on the feature map output by the first C3 layer through the first CCASA attention layer, perform feature map size compression on the feature map output by the first CCASA attention layer through the second Conv layer, and perform feature extraction on the feature map output by the second Conv layer through the second C3 layer to obtain a first backbone feature map;
[0027] Perform extraction of interesting features on the first backbone feature map through the second CCASA attention layer, perform feature map size compression on the feature map output by the second CCASA attention layer through the third Conv layer, and perform feature extraction on the feature map output by the third Conv layer through the third C3 layer to obtain a second backbone feature map;
[0028] Perform extraction of interesting features on the second backbone feature map through the third CCASA attention layer, perform feature map size compression on the feature map output by the third CCASA attention layer through the fourth Conv layer, and perform feature extraction on the feature map output by the fourth Conv layer through the fourth C3 layer to obtain a third backbone feature map.
[0029] Optionally, the head network includes two cascaded FPN modules, two feature aggregation modules, an SPPF layer, and a fourth CCASA attention layer. The FPN module includes a Conv layer, an Upsample layer, a Concat layer, and a C3 layer connected in sequence;
[0030] The input of the first feature map and the multiple backbone feature maps of different sizes into the head network to obtain the three target feature maps of different sizes includes:
[0031] Perform spatial information fusion on the third backbone feature map through the SPPF layer, and perform extraction of features of interest on the feature map output by the SPPF layer through the fourth CCASA attention layer to obtain a fourth feature map;
[0032] The first FPN module receives the fourth feature map and the second backbone feature map as inputs, the second FPN module receives the feature map output by the first FPN module and the first backbone feature map as inputs, and the second FPN module outputs a first target feature map;
[0033] Input the first target feature map and the feature map output by the Conv layer of the second FPN module into a first feature aggregation module, and the first feature aggregation module performs size compression and channel aggregation on the input multiple feature maps to obtain a second target feature map;
[0034] Input the second target feature map and the feature map output by the Conv layer of the first FPN module into a second feature aggregation module, and the second feature aggregation module performs size compression and channel aggregation on the input multiple feature maps to obtain a third target feature map.
[0035] According to a second aspect of the present invention, there is provided an object detection device integrating an attention mechanism, and the device includes:
[0036] An image acquisition module for acquiring an input image;
[0037] A preprocessing module for performing feature extraction on the input image through a preset imitation primary visual perception cortex module to obtain a first feature map; the preset imitation primary visual perception cortex model includes a VOneBlock layer, a Conv layer, and a feature fusion layer;
[0038] An object detection module for performing object detection on the first feature map through a preset object detection model to obtain three object feature maps of different sizes; the object detection model includes a fusion attention module for extracting coordinate channel attention features and spatial attention features; performing object classification and coordinate positioning on the three object feature maps of different sizes to obtain an object detection result.
[0039] Optionally, the object detection module includes:
[0040] A backbone network module for performing multiple size compressions and feature extractions on the first feature map to obtain multiple backbone feature maps of different sizes;
[0041] A head network module for receiving the multiple backbone feature maps of different sizes to obtain the three object feature maps of different sizes.
[0042] According to a third aspect of the present invention, there is provided an electronic device including a processor and a memory, where the memory stores a program that can run on the processor, and when the program is executed by the processor, it implements the object detection method with a fused attention mechanism as described in the first aspect of the present invention.
[0043] According to an embodiment of the present disclosure, by adding a fused attention module to the object detection module in the present invention and using the fused attention module to extract channel attention and spatial attention, it is possible to effectively filter out interference information such as image backgrounds, improve the accuracy of small object detection, and further enhance the detection accuracy and robustness of the object detection module in complex interference scenarios.
[0044] Other features and advantages of the present invention will become clear from the following detailed description of the exemplary embodiments of the present invention with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments of the invention and, together with the description, serve to explain the principles of the invention.
[0046] Figure 1 It is a flowchart of an object detection method with a fused attention mechanism in an embodiment of the present invention.
[0047] Figure 2 It is a schematic diagram of a fused attention module in an embodiment of the present invention.
[0048] Figure 3 It is a schematic diagram of a module imitating the primary visual cortex in an embodiment of the present invention.
[0049] Figure 4 It is a schematic diagram of an object detection model in an embodiment of the present invention.
[0050] Figure 5 It is a schematic diagram of an object detection device with a fused attention mechanism in an embodiment of the present invention.
[0051] Figure 6 It is a schematic diagram of an electronic device in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0052] Now, various exemplary embodiments of the present invention will be described in detail with reference to the accompanying drawings. It should be noted that: unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions, and numerical values set forth in these embodiments do not limit the scope of the present invention.
[0053] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the present invention, its application, or its use.
[0054] Techniques, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but in appropriate cases, such techniques, methods, and devices should be regarded as part of the specification.
[0055] In all the examples shown and discussed here, any specific values should be construed as merely exemplary, rather than as limitations. Thus, other examples of the exemplary embodiments may have different values.
[0056] It should be noted that like reference numerals and letters refer to like items in the following figures, and thus, once an item is defined in one figure, further discussion thereof is not required in subsequent figures.
[0057] In computer vision, a method that can focus attention on important regions of an image while discarding irrelevant ones is called an attention mechanism. In the human visual cortex, using the attention mechanism can analyze complex scene information more quickly and efficiently. The attention mechanism can be regarded as a dynamic selection process for important information in the image input, and this process is achieved by adaptive weights for features. The attention mechanism has greatly improved the performance levels of many computer vision tasks, such as playing an important role in tasks like classification, object detection, semantic segmentation, face recognition, action recognition, small sample detection, medical image processing, image generation, pose estimation, super-resolution, 3D vision, and multi-modal.
[0058] Currently, there is no clear definition of the concept of small targets in the industry. Small targets can be defined from two aspects: absolute scale and relative scale. For example, when defining small targets from the perspective of absolute scale, when the area of the target region is less than 32 * 32 pixel values, the target region can be considered a small target. When defining small targets from the perspective of relative scale, a small target is when the length and width of the target size account for 0.1 of the original image size.
[0059] As Figure 1 shown, the embodiments of the present invention introduce an object detection method integrating an attention mechanism, and the method includes steps S1 - S4.
[0060] S1: Obtain an input image.
[0061] S2: Extract features from the input image through a preset imitation primary visual perception cortex module to obtain a first feature map; the preset imitation primary visual perception cortex model includes a VOneBlock layer, a Conv layer, and a feature fusion layer.
[0062] S3: Perform object detection on the first feature map through a preset object detection model to obtain three object feature maps of different sizes; the object detection model includes a fusion attention module, and the fusion attention module is used to extract coordinate channel attention features and spatial attention features.
[0063] S4: Perform object classification and coordinate localization on the three object feature maps of different sizes to obtain an object detection result.
[0064] The input image is an image on which object detection needs to be performed. It can be an image captured by a camera, an image frame extracted from a video, or a partial region image cropped from a complete image.
[0065] After obtaining the input image, preprocess the input image. During the preprocessing process, perform feature extraction on the input image through a preset primary visual cortex-like model to obtain a first feature map. In the primary visual cortex-like model, there are a VOneBlock layer, a Conv (convolution) layer, and a feature fusion layer. Among them, the VOneBlock layer is a neural network layer constructed according to the primary visual cortex of primates, with a Gabor filter as the core component, and simulates the information processing mechanism of the human visual cortex to perform bionic visual feature extraction on the input image. After performing feature extraction on the input image through the primary visual cortex-like model, a first feature map closer to the features after human brain visual processing can be obtained. The primary visual cortex-like model reflects the mapping relationship between the input image and the first feature map.
[0066] After obtaining the first feature map, perform object detection on the first feature map through a preset object detection model to obtain three object feature maps of different sizes. There is a fusion attention module in the object detection module, and the fusion attention module is used to extract coordinate channel attention and spatial attention, and key object features are extracted from the image features to the deep semantic features, which can filter out interference information such as the image background.
[0067] After obtaining the three object feature maps of different sizes, perform object classification and coordinate localization on the three object feature maps of different sizes to obtain an object detection result. The object detection result contains the position information and classification information of the object to be detected in the input image, and the object detection model reflects the mapping relationship between the first feature map and the object detection result.
[0068] In the present invention, by adding a fusion attention module to the object detection module and using the fusion attention module to extract channel attention and spatial attention, it is possible to effectively filter out interference information such as the image background, improve the accuracy of small object detection, and further improve the detection accuracy and robustness of the object detection module in complex interference scenarios.
[0069] The above-mentioned step S2 includes S201 - S203.
[0070] S201: Extract features and compress the size of the input image through the VOneBlock layer to obtain a second feature map.
[0071] S202: Compress the size of the input image through the Conv layer to obtain a third feature map; the size of the third feature map is the same as that of the second feature map.
[0072] S203: Fuse the eigenvalues of the second feature map and the third feature map through the feature fusion layer to obtain the first feature map.
[0073] As Figure 3 shown, the primary visual perception cortex - like model in the present invention includes a VOneBlock layer, a Conv layer, and a feature fusion layer, where the VOneBlock layer and the Conv layer are in parallel. The input image serves as the input to the VOneBlock layer and the Conv layer. After the VOneBlock layer extracts features and compresses the size of the input image, a second feature map is obtained. After the Conv layer compresses the size of the input image, a third feature map is obtained. The size of the second feature map is the same as that of the third feature map. The second feature map and the third feature map serve as the input to the feature fusion layer, and the feature fusion layer fuses the eigenvalues of the second feature map and the third feature map, and finally outputs the first feature map.
[0074] For example, the size of the input image is 640*640. After the VOneBlock layer extracts features and compresses the size of the input image, the size of the obtained second feature map is 320*320, and the size of the second feature map is 1 / 2 of the size of the input image. After the Conv layer compresses the size of the input image, the size of the obtained third feature map is 320*320, and the size of the third feature map is also 1 / 2 of the size of the input image. The size of the third feature map is the same as that of the second feature map. Finally, the feature fusion layer fuses the eigenvalues of the second feature map and the third feature map, and the size of the obtained first feature map is also 320*320.
[0075] In the primary visual perception cortex - like model of the present invention, only the VOneBlock layer is in parallel with a Conv layer, which not only makes the model more concise, but also enables the primary visual perception cortex - like model to be more flexible in adjusting the size compression ratio of the feature map, improving the flexibility of the model.
[0076] As Figure 2As shown, in an embodiment of the present invention, the fusion attention module includes a coordinate channel attention module and a spatial attention module; the coordinate channel attention module is used to extract the coordinate channel attention of the original feature map to obtain a feature map processed by the coordinate channel attention, and the original feature map is the feature map input to the fusion attention module; the spatial attention module is used to extract the spatial attention features of the feature map output by the coordinate channel attention module.
[0077] The coordinate channel attention module performs pooling on the X direction and Y direction of each channel of the original feature map respectively to obtain an X-direction feature map and a Y-direction feature map; the X-direction feature map and the Y-direction feature map are concatenated to obtain a two-dimensional channel weight feature map; the two-dimensional channel weight feature map is successively subjected to two-dimensional convolution, normalization and nonlinear transformation to obtain an optimized feature map; an X-direction channel attention feature map and a Y-direction channel attention feature map are obtained according to the optimized feature map; according to the X-direction channel attention feature map, the Y-direction channel attention feature map and the original feature map, the feature map processed by the coordinate channel attention is obtained.
[0078] For example, the original feature map includes three channels. The coordinate channel attention module performs pooling on the X direction and Y direction of the three channels of the original feature map respectively, and concatenates them to form a two-dimensional channel weight feature map. Each channel of the original feature map corresponds to a two-dimensional channel weight feature map. Then, the two-dimensional weight feature map is successively subjected to two-dimensional convolution, normalization and nonlinear transformation. Finally, the two-dimensional weight feature map is divided again according to the X direction and Y direction to obtain an X-direction channel attention feature map and a Y-direction channel attention feature map. The above X-direction channel attention feature map and Y-direction channel attention feature map are respectively subjected to dot product with the original feature map to obtain the feature map processed by the coordinate channel attention.
[0079] The spatial attention module performs pooling on the feature map output by the coordinate channel attention module in the channel direction to obtain a spatial attention weight feature map; the spatial attention features are obtained according to the spatial attention weight feature map and the feature map output by the coordinate channel attention module.
[0080] As Figure 4 shown, in an embodiment of the present invention, the preset target detection model includes a backbone network 101 and a head network 102, and the fusion attention module is provided in both the backbone network 101 and the head network 103, and the step S3 includes S301-S302.
[0081] S301: The backbone network performs multiple size compressions and feature extractions on the first feature map to obtain multiple backbone feature maps of different sizes.
[0082] S302: Input the multiple backbone feature maps of different sizes into the head network to obtain the three target feature maps of different sizes.
[0083] The object detection model in the present invention is an improved model based on the Yolov5 model in the prior art. The object detection model includes a backbone network 101 and a head network 102, and a fusion attention module is added to both the backbone network 101 and the head network 102. Among them, the backbone network is used to perform multiple size compressions and feature extractions on the first feature map to obtain multiple backbone feature maps of different sizes. Then the head network receives the multiple backbone feature maps of different sizes obtained through the backbone network, and the head network will output five target feature maps of different sizes. As Figure 4 shown, the object detection model also includes a Detect (detection head) layer, and the three target feature maps of different sizes are subjected to object classification and coordinate positioning through the Detect layer to output the object detection result.
[0084] In an embodiment of the present invention, the backbone network includes a first Conv layer, a first C3 layer, a second Conv layer, a second C3 layer, a third Conv layer, a third C3 layer, a fourth Conv layer, and a fourth C3 layer. The fusion attention module includes a first CCASA (Coordinate channel attention meeting spatial attention) attention layer arranged after the first C3 layer, a second CCASA attention layer arranged after the second C3 layer, and a CCASA attention layer arranged after the third C3 layer. One CCASA attention layer represents a fusion attention module and is used to implement the functions of the above fusion attention module. The step S301 includes S3011 - S3013.
[0085] S3011: Compress the size of the first feature map through the first Conv layer, extract features from the feature map output by the first Conv layer through the first C3 layer, extract interesting features from the feature map output by the first C3 layer through the first CCASA attention layer, compress the size of the feature map output by the first CCASA attention layer through the second Conv layer, and extract features from the feature map output by the second Conv layer through the second C3 layer to obtain a first backbone feature map.
[0086] S3012: Extract the features of interest from the first backbone feature map through the second CCASA attention layer, compress the size of the feature map output by the second CCASA attention layer through the third Conv layer, and extract features from the feature map output by the third Conv layer through the third C3 layer to obtain the second backbone feature map;
[0087] S3013: Extract the features of interest from the second backbone feature map through the third CCASA attention layer, compress the size of the feature map output by the third CCASA attention layer through the fourth Conv layer, and extract features from the feature map output by the fourth Conv layer through the fourth C3 layer to obtain the third backbone feature map.
[0088] After obtaining the input image, the size of the first feature map obtained through the primary visual perception cortex model is 1 / 2 of the size of the input image. The first feature map is input into the backbone network, and the first Conv layer in the backbone network obtains the first feature map, compresses the size of the first feature map, and the size of the feature map output by the first Conv layer is 1 / 4 of the size of the input image. The feature map output by the first Conv layer is input into the first C3 layer for feature extraction. The first CCASA attention layer extracts the features of interest from the feature map output by the first C3 layer, the second Conv layer compresses the size of the feature map output by the first CCASA attention layer, and the feature map output by the second Conv layer is input into the second C3 layer for feature extraction. The second C3 layer outputs the first backbone feature map.
[0089] The second CCASA attention layer extracts the features of interest from the first backbone feature map, the third Conv layer compresses the size of the feature map output by the second CCASA attention layer, and the feature map output by the third Conv layer is input into the third C3 layer for feature extraction. The third C3 layer outputs the second backbone feature map.
[0090] The third CCASA attention layer extracts the features of interest from the second backbone feature map, the fourth Conv layer compresses the size of the feature map output by the third CCASA attention layer, and the feature map output by the fourth Conv layer is input into the fourth C3 layer for feature extraction. The fourth C3 layer outputs the third backbone feature map.
[0091] For example, if the size of the input image is 640*640, the primary visual perception cortex model will compress the size of the input image, and the output first feature map is 1 / 2 of the input image, with the size of the first feature map being 320*320. The first feature map is subjected to four size compressions through the backbone network, and feature extraction is performed after each size compression. After two size compressions and feature extraction processes through the first Conv layer, the second Conv layer, the first C3 layer, and the second C3 layer, the first backbone feature map with a size of 80*80 is obtained, and the size of the first backbone feature map is 1 / 8 of the size of the input image. After the third size compression and feature extraction process through the third Conv layer and the third C3 layer, the second backbone feature map with a size of 40*40 is obtained, and the size of the second backbone feature map is 1 / 16 of the size of the input image. After the fourth size compression and feature extraction process through the fourth Conv layer and the fourth C3 layer, the fourth backbone feature map with a size of 20*20 is obtained, and the size of the fourth backbone feature map is 1 / 32 of the size of the input image.
[0092] As Figure 4 shown, in an embodiment of the present invention, the head network includes two serially connected FPN modules, two feature aggregation modules, an SPPF layer, and a fourth CCASA attention layer. The FPN module includes a Conv layer, an Upsample (upsampling) layer, a Concat (feature aggregation) layer, and a C3 layer connected in sequence. The step S302 includes S3021-S3024.
[0093] S3021: Perform spatial information fusion on the third backbone feature map through the SPPF layer, and perform extraction of features of interest on the feature map output by the SPPF layer through the fourth CCASA attention layer to obtain a fourth feature map.
[0094] S3022: The first FPN module receives the fourth feature map and the second backbone feature map as inputs. The second FPN module receives the feature map output by the first FPN module and the first backbone feature map as inputs, and the second FPN module outputs a first target feature map.
[0095] S3023: Input the first target feature map and the feature map output by the Conv layer of the second FPN module into the first feature aggregation module. The first feature aggregation module performs size compression and channel aggregation on the input multiple feature maps to obtain a second target feature map.
[0096] S3024: Input the second target feature map and the feature map output by the Conv layer of the first FPN module into the second feature aggregation module. The second feature aggregation module performs size compression and channel aggregation on the input multiple feature maps to obtain a third target feature map.
[0097] The Conv layer, Upsample layer, Concat layer, and C3 layer in the FPN module are consistent with the existing Yolov5 model. The Conv layer is a basic convolutional unit that sequentially performs two-dimensional convolution, regularization, and activation operations on the input. The C3 layer consists of several Bottleneck modules. The Bottleneck is a classic residual structure where the input passes through two convolutional layers and then undergoes an Add operation with the original value to complete the transfer of residual features without increasing the output depth.
[0098] After the first feature map is input into the backbone network, the backbone network outputs three backbone feature maps of different sizes, namely the first backbone feature map, the second backbone feature map, and the third backbone feature map. The three target feature maps of different sizes are the first target feature map, the second target feature map, and the third target feature map. The two feature aggregation modules are the first feature aggregation module and the second feature aggregation module, and each aggregation module contains a Conv layer, a Concat layer, and a C3 layer. As Figure 4 shown, the first feature aggregation module includes the seventh Conv layer, the third Concat layer, and the seventh C3 layer, and the second feature aggregation module includes the eighth Conv layer, the fourth Concat layer, and the eighth C3 layer.
[0099] In the present invention, the two FPN modules are the first FPN module and the second FPN module. As Figure 4 shown, the first FPN module includes the fifth Conv layer, the first Upsample layer, the first Concat layer, and the fifth C3 layer, and the second FPN module includes the sixth Conv layer, the second Upsample layer, the second Concat layer, and the sixth C3 layer.
[0100] The output of the first FPN module is used as the input of the second FPN module, and the feature map output by the first FPN module is input into the sixth Conv layer of the second FPN module; the second FPN module also receives the first backbone feature map as input and inputs the first backbone feature map into the second Concat layer of the second FPN module. The first FPN module receives the fourth feature map and the second backbone feature map as input, inputs the fourth feature map into the fifth Conv layer of the first FPN module, and inputs the second backbone feature map into the second Concat layer of the first FPN module.
[0101] After obtaining the first target feature map, the first target feature map and the feature map output by the sixth Conv layer of the second FPN module are input into the first feature aggregation module. The seventh Conv layer in the first feature aggregation module compresses the size of the first target feature map. The third Concat layer in the first feature aggregation module performs channel aggregation on the feature map output by the seventh Conv layer in the first feature aggregation module and the feature map output by the sixth Conv layer of the first FPN module. The seventh C3 layer in the first feature aggregation module extracts features from the feature map output by the Concat layer in the first feature aggregation module. The feature map output by the seventh C3 layer in the first feature aggregation module is the second target feature map.
[0102] After obtaining the second target feature map, the second target feature map and the feature map output by the fifth Conv layer of the first FPN module are input into the second feature aggregation module. The eighth Conv layer in the second feature aggregation module compresses the size of the second target feature map. The fourth Concat layer in the second feature aggregation module performs channel aggregation on the feature map output by the Conv layer in the second feature aggregation module and the feature map output by the fifth Conv layer of the first FPN module. The eighth C3 layer in the second feature aggregation module extracts features from the feature map output by the Concat layer in the second feature aggregation module. The feature map output by the eighth C3 layer in the second feature aggregation module is the third target feature map.
[0103] As Figure 5 shown, an object detection device 200 with a fusion attention mechanism is introduced in an embodiment of the present invention. The object detection device 200 with the fusion attention mechanism is used to implement the object detection method with the fusion attention mechanism described in any embodiment of the present invention. The object detection device 200 with the fusion attention mechanism includes:
[0104] An image acquisition module 201 for acquiring an input image;
[0105] A preprocessing module 202 for extracting features from the input image through a preset primary visual cortex-like model to obtain a first feature map. The preset primary visual cortex-like model includes a VOneBlock layer, a Conv layer, and a feature fusion layer;
[0106] An object detection module 203 for performing object detection on the first feature map through a preset object detection model to obtain three target feature maps of different sizes. The object detection model includes a fusion attention module for extracting coordinate channel attention and spatial attention. Object classification and coordinate positioning are performed on the three target feature maps of different sizes to obtain an object detection result.
[0107] In the present invention, by adding a fusion attention module to the target detection module and using the fusion attention module to extract channel attention and spatial attention, it is possible to effectively filter out interference information such as the image background, improve the accuracy of small target detection, and further enhance the detection accuracy and robustness of the target detection module in complex interference scenarios.
[0108] In an embodiment of the present invention, the target detection module 203 includes a backbone network module and a head network module. The backbone network module is used to perform multiple size compressions and feature extractions on the first feature map to obtain multiple backbone feature maps of different sizes. The head network module is used to receive the multiple backbone feature maps of different sizes to obtain the three target feature maps of different sizes.
[0109] As Figure 6 shown, an embodiment of the present invention introduces an electronic device 300, including a processor 301 and a memory 302. The memory 302 stores a program that can run on the processor 301. When the program is executed by the processor 301, it implements the target detection method of the fusion attention mechanism as described in any embodiment of the present invention.
[0110] The present invention can be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present invention.
[0111] The computer-readable storage medium can be a tangible device that can hold and store instructions used by an instruction execution device. The computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as an instantaneous signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0112] The computer-readable program instructions described herein can be downloaded to various computing / processing devices from a computer-readable storage medium or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0113] The computer program instructions for carrying out operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present invention.
[0114] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0115] These computer-readable program instructions may be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture including instructions for implementing various aspects of the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0116] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0117] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, which comprises one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions. It is well known to those skilled in the art that implementation by hardware, implementation by software, and implementation by a combination of software and hardware are equivalent.
[0118] The embodiments of the present invention have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the technical improvement of the technology in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein. The scope of the present invention is defined by the appended claims.
Claims
1. A target detection method integrating an attention mechanism, characterized in that, The method includes: Obtaining an input image; Performing feature extraction on the input image through a preset primary visual perception cortex - like module to obtain a first feature map; the preset primary visual perception cortex - like model includes a VOneBlock layer, a Conv layer, and a feature fusion layer; Performing object detection on the first feature map through a preset object detection model to obtain three object feature maps of different sizes; the object detection model contains a fusion attention module, and the fusion attention module is used to extract coordinate - channel attention features and spatial attention features; Performing object classification and coordinate localization on the three object feature maps of different sizes to obtain an object detection result; Among them, the fusion attention module includes a coordinate - channel attention module and a spatial attention module; The coordinate - channel attention module is used to extract the coordinate - channel attention features of the original feature map to obtain a feature map processed by coordinate - channel attention; the original feature map is the feature map input to the fusion attention module; The spatial attention module is used to extract spatial attention features from the feature map processed by coordinate - channel attention; Among them, the coordinate - channel attention module is used to extract the coordinate - channel attention features of the original feature map to obtain a feature map processed by channel attention, including: Performing pooling on the X - direction and Y - direction of each channel of the original feature map respectively to obtain an X - direction feature map and a Y - direction feature map; Concatenating the X - direction feature map and the Y - direction feature map to obtain a two - dimensional channel weight feature map; Performing two - dimensional convolution, normalization, and non - linear transformation on the two - dimensional channel weight feature map in sequence to obtain an optimized feature map; Obtaining an X - direction channel attention feature map and a Y - direction channel attention feature map according to the optimized feature map; Obtaining the feature map processed by coordinate - channel attention according to the X - direction channel attention feature map, the Y - direction channel attention feature map, and the original feature map; Among them, the spatial attention module is used to extract spatial attention features from the feature map processed by coordinate - channel attention, including: The spatial attention module performs pooling on the feature map output by the coordinate - channel attention module in the channel direction to obtain a spatial attention weight feature map; Obtaining spatial attention features according to the spatial attention weight feature map and the feature map output by the coordinate - channel attention module.
2. The method according to claim 1, wherein The preset object detection model includes a backbone network and a head network, and the fusion attention module is provided in both the backbone network and the head network. Performing object detection on the first feature map through the preset object detection model to obtain three object feature maps of different sizes includes: Performing multiple size compressions and feature extractions on the first feature map through the backbone network to obtain multiple backbone feature maps of different sizes; Inputting the multiple backbone feature maps of different sizes into the head network to obtain the three object feature maps of different sizes.
3. The method according to claim 2, characterized in that, The backbone network includes a first Conv layer, a first C3 layer, a second Conv layer, a second C3 layer, a third Conv layer, a third C3 layer, a fourth Conv layer, and a fourth C3 layer. The fusion attention module includes a first CCASA attention layer arranged after the first C3 layer, a second CCASA attention layer arranged after the second C3 layer, and a CCASA attention layer arranged after the third C3 layer; The first feature map is subjected to multiple size compressions and feature extractions through the backbone network, resulting in multiple backbone feature maps of different sizes, including: The first Conv layer compresses the size of the first feature map, the first C3 layer extracts features from the feature map output by the first Conv layer, the first CCASA attention layer extracts interesting features from the feature map output by the first C3 layer, the second Conv layer compresses the size of the feature map output by the first CCASA attention layer, and the second C3 layer extracts features from the feature map output by the second Conv layer to obtain a first backbone feature map; The second CCASA attention layer extracts interesting features from the first backbone feature map, the third Conv layer compresses the size of the feature map output by the second CCASA attention layer, and the third C3 layer extracts features from the feature map output by the third Conv layer to obtain a second backbone feature map; The third CCASA attention layer extracts interesting features from the second backbone feature map, the fourth Conv layer compresses the size of the feature map output by the third CCASA attention layer, and the fourth C3 layer extracts features from the feature map output by the fourth Conv layer to obtain a third backbone feature map.
4. The method according to claim 3, wherein The head network includes two cascaded FPN modules, two feature aggregation modules, an SPPF layer, and a fourth CCASA attention layer. The FPN module includes a Conv layer, an Upsample layer, a Concat layer, and a C3 layer connected in sequence; The first feature map and the multiple backbone feature maps of different sizes are input into the head network to obtain three target feature maps of different sizes, including: The SPPF layer performs spatial information fusion on the third backbone feature map, and the fourth CCASA attention layer extracts interesting features from the feature map output by the SPPF layer to obtain a fourth feature map; The first FPN module receives the fourth feature map and the second backbone feature map as inputs, and the second FPN module receives the feature map output by the first FPN module and the first backbone feature map as inputs. The second FPN module outputs a first target feature map; Input the first target feature map and the feature map output by the Conv layer of the second FPN module into the first feature aggregation module. The first feature aggregation module performs size compression and channel aggregation on the input multiple feature maps to obtain a second target feature map; Input the second target feature map and the feature map output by the Conv layer of the first FPN module into the second feature aggregation module. The second feature aggregation module performs size compression and channel aggregation on the input multiple feature maps to obtain a third target feature map.
5. An object detection device integrating an attention mechanism, characterized in that, The device includes: An image acquisition module for acquiring an input image; A preprocessing module for extracting features from the input image through a preset primary visual perception cortex-like module to obtain a first feature map; the preset primary visual perception cortex-like model includes a VOneBlock layer, a Conv layer, and a feature fusion layer; A target detection module for performing target detection on the first feature map through a preset target detection model to obtain three target feature maps of different sizes; the target detection model includes a fusion attention module, and the fusion attention module is used to extract coordinate channel attention features and spatial attention features; perform target classification and coordinate localization on the three target feature maps of different sizes to obtain a target detection result; Among them, the fusion attention module includes a coordinate channel attention module and a spatial attention module; The coordinate channel attention module is used to extract the coordinate channel attention features of the original feature map to obtain a feature map processed by coordinate channel attention; the original feature map is the feature map input into the fusion attention module; The spatial attention module is used to extract spatial attention features from the feature map processed by coordinate channel attention; Among them, the coordinate channel attention module is used to extract the coordinate channel attention features of the original feature map to obtain a feature map processed by channel attention, including: Perform pooling on the X direction and Y direction of each channel of the original feature map respectively to obtain an X direction feature map and a Y direction feature map; Concatenate the X direction feature map and the Y direction feature map to obtain a two-dimensional channel weight feature map; Perform two-dimensional convolution, normalization, and non-linear transformation on the two-dimensional channel weight feature map in sequence to obtain an optimized feature map; Obtain an X direction channel attention feature map and a Y direction channel attention feature map according to the optimized feature map; Obtain the feature map processed by coordinate channel attention according to the X direction channel attention feature map, the Y direction channel attention feature map, and the original feature map; Among them, the spatial attention module is used to extract spatial attention features from the feature map processed by coordinate channel attention, including: The spatial attention module performs pooling on the feature map output by the coordinate channel attention module in the channel direction to obtain a spatial attention weight feature map; Obtain spatial attention features according to the spatial attention weight feature map and the feature map output by the coordinate channel attention module.
6. The device according to claim 5, wherein The target detection module includes: The backbone network module is used to perform multiple size compressions and feature extractions on the first feature map to obtain multiple backbone feature maps of different sizes; The head network module is used to receive the multiple backbone feature maps of different sizes to obtain the three target feature maps of different sizes.
7. An electronic device, characterized in that, It includes a processor and a memory. The memory stores a program that can run on the processor. When the program is executed by the processor, it implements the object detection method with the fusion attention mechanism as described in any one of claims 1-4.
Citation Information
Patent Citations
Target detection method based on information enhancement
CN111612017A
Image target detection method and device, equipment and storage medium
CN113936256A