An Object Detection Method Based on Feature Fusion and Attention Mechanism

By introducing the bidirectional feature fusion module IBFPN and the improved attention module ECBAM in the infrared object detection algorithm, the problems of missed detection and false detection in infrared object detection are solved, and higher detection accuracy and efficiency are achieved.

CN115424104BActive Publication Date: 2025-07-29XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210998016.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-19
Publication Date
2025-07-29
Estimated Expiration
2042-08-19

AI Technical Summary

Technical Problem

Existing infrared object detection algorithms are prone to missed and missed detection in complex backgrounds, and there is a lack of effective information exchange and supplementation between feature maps at different levels.

Method used

Using an improved method based on feature fusion and attention mechanism, the feature expression ability is enhanced through the bidirectional feature fusion module IBFPN and the improved attention module ECBAM, and the feature expression ability is realized, and the upper and lower layer information is integrated and the weight assignment of feature maps is achieved.

Benefits of technology

It improves the accuracy and efficiency of infrared target detection, reduces the error detection rate, and improves the accuracy of detection of targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115424104B_ABST
    Figure CN115424104B_ABST
Patent Text Reader

Abstract

An object detection method based on feature fusion and attention mechanism. The infrared image is input into the MobileNet network for layer-by-layer convolutional calculation to obtain feature maps of different scales. A bidirectional feature fusion module IBFPN is established, and the feature pyramid image for detection is input into the IBFPN to perform mutual fusion of upper and lower layer information. After passing through the IBFPN, each fused feature layer is input into the attention module ECBAM, and different weights are assigned to different features through the ECBAM. The feature maps of different scales processed by the IBFPN and ECBAM are sent to the detection module for detection, the category of each candidate box and the corresponding bounding box are obtained, and the prediction result is obtained. The prediction result is subjected to non-maximum suppression to delete redundant target boxes and obtain the final detection result. After comprehensively comparing various detection algorithms, the experimental results show that the present invention can effectively improve the detection accuracy of infrared targets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of infrared target detection, and particularly relates to a target detection method based on feature fusion and attention mechanism. Background Art

[0002] As an important branch in the fields of image processing and computer vision, target detection technology has always been the focus of research and attention. Target detection, also known as target extraction, essentially classifies and precisely locates multiple targets in a given image. Different from visible light imaging, infrared imaging is a passive imaging technology that forms images by detecting the infrared thermal radiation emitted by objects themselves, and has the advantages of long detection range, strong penetration, and all-day operation. Therefore, the infrared target detection technology that combines target detection with infrared imaging systems can effectively make up for the deficiencies of visible light detection and further meet the requirements of different industries for detection technology.

[0003] In recent years, deep learning algorithms have shown excellent performance in the field of image processing. Different from traditional manually designed features, they automatically extract features through convolutional neural networks and have good robustness and portability. Therefore, it has important research value to construct infrared target detection algorithms based on deep learning. Among them, the SSD detection algorithm uses the backbone network and additional convolutional layers to generate 6 groups of feature maps with different scales to predict the category and coordinates of candidate boxes. The depths of the convolutional layers are different, and the image information contained in their feature maps is also different. Among them, the shallow feature maps contain information such as the texture and position of the image, but the extracted features are not comprehensive; the deep feature maps contain richer semantic information, but a lot of details are lost during the convolution process. SSD only inputs the information of each branch into the detection module separately, and there is no mutual connection and complement between the features of each layer. In addition, due to the unclear target information and sparse feature structure in infrared images, complex background environments are also likely to cause phenomena such as missed detection and false detection of targets. Summary of the Invention

[0004] To overcome the above-mentioned shortcomings of the prior art, the purpose of the present invention is to provide a target detection method based on feature fusion and attention mechanism. By improving the traditional SSD detection algorithm, the expression ability of the detection branch is strengthened, the network expression ability is improved, and the influence brought by irrelevant information is suppressed, so as to solve the problems of missed detection and false detection caused by complex backgrounds and other reasons.

[0005] To achieve the above object, the technical solution adopted by the present invention is:

[0006] A target detection method based on feature fusion and attention mechanism, comprising the following steps:

[0007] Step 1: Input the infrared image into the MobileNet network for layer-by-layer convolution calculation to obtain feature maps of different scales. Among them, the feature maps of six layers, namely DWS11, DWS13, DWS14_2, DWS15_2, DWS16_2, and DWS17_2, are the feature pyramid images for detection.

[0008] Step 2: Establish a bidirectional feature fusion module IBFPN, and input the feature pyramid images for detection into the IBFPN to perform mutual fusion of upper and lower layer information.

[0009] Step 3: After passing through the IBFPN, input each fused feature layer into the attention module ECBAM to assign different weights to different features.

[0010] Step 4: Send DWS11, DWS13, DWS14_2, DWS15_2, DWS16_2, and DWS17_2 processed by the IBFPN and ECBAM into the detection module for detection, obtain the category and corresponding bounding box of each candidate box, and get the prediction result.

[0011] Step 5: Perform non-maximum suppression on the prediction result, delete the redundant target boxes, and obtain the final detection result.

[0012] In one embodiment, in Step 2, for each feature map for detection, compared with the traditional SSD algorithm that directly detects the feature map, in the present invention, each feature map is first input into the bidirectional feature fusion module IBFPN before detection to perform mutual fusion of upper and lower layer information. Among them, the IBFPN is based on the traditional bidirectional pyramid network, constructs a residual feature enhancement module RFA to strengthen the top layer features, introduces a bottom-up fusion path, and adds a connection from the starting input to the output for the same level.

[0013] In one embodiment, the steps for mutual fusion of upper and lower layer information in Step 2 are as follows:

[0014] Step 2.1: In the forward propagation process, the pyramid feature levels generated by the traditional SSD object detection model are {C1, C2, C3, C4, C5, C6} in sequence, corresponding to DWS11, DWS13, DWS14_2, DWS15_2, DWS16_2, and DWS17_2 in Step 1 respectively.

[0015] Step 2.2: Input the feature layer of C6 into the RFA residual feature enhancement module to improve its feature representation, and obtain the output feature layer R6 with multi-scale context information, that is, RFA(C6).

[0016] Step 2.3, add and fuse R6 with C6 after dimensionality reduction by 1×1 convolution to obtain the top-level feature C6_td;

[0017] Step 2.4, perform dimensionality reduction on C1 to C5 respectively through 1×1 convolution to obtain features C5_in to C1_in;

[0018] Step 2.5, fuse C1_in to C6_in from top to bottom using the following formula:

[0019] C6_td = Conv(C6) + RFA(C6)

[0020] C5_td = Conv(C5) + Resize(C6_td)

[0021]

[0022] C1_td = Conv(C1) + Resize(C2_td)

[0023] Among them, Conv is the 1×1 convolution operation for feature channel dimensionality reduction, Conv(C6) to Conv(C1) are C6_in to C1_in, RFA(*) is to perform residual feature enhancement on the feature map, Resize is the upsampling operation taken to match the resolutions of different layers, and "+" represents adding elements at corresponding positions; C1_td to C6_td are the top-level features obtained after addition and fusion;

[0024] Step 2.6, from bottom to top, fuse low-level features to high-level features, so that each layer not only has strong semantic information of the high level but also strong detail localization information of the low level, and add a connection from the starting input to the output to fuse more features without increasing the number of parameters, obtaining the fused feature layers C1_out to C6_out, and the fusion formula is as follows:

[0025] C1_out = C1_td = C1_in + Resize(C2_td)

[0026] C2_out = C2_td + Resize(C1_out) + C2_in

[0027]

[0028] C6_out = C6_td + Resize(C5_out) + C6_in

[0029] C1_out to C6_out are the feature layers of multi-scale feature fusion output after C1 to C6 pass through IBFPN.

[0030] In one embodiment, the residual feature enhancement module RFA combines the idea of residual feature enhancement in AugFPN in the traditional Res residual structure and incorporates an Adaptive Spatial Fusion (ASF) module. Whether in FPN or the bidirectional feature pyramid, the features of the highest layer are propagated by top-down upsampling and gradually fused with the features of lower layers. During this process, the features of lower layers are enhanced by the semantic information from higher layers, naturally making the fused features contain different context information; however, due to resizing, the channel dimension of the highest layer features is compressed, which will lead to the loss of some information. In response to this situation, the RFA designed in the present invention increases context information through residual features to reduce the loss of the highest layer features and improve the performance of the pyramid.

[0031] Exemplarily, the upsampling operation is bilinear interpolation.

[0032] In one embodiment, the attention module ECBAM is obtained by improving the original Convolutional Block Attention Module (CBAM). It mainly includes two modules: channel attention and spatial attention. Among them, the channel attention module focuses on the importance of different feature channels, including the average pooling layer AvgPool and the max pooling layer MaxPool, the one-dimensional convolutional module, and the Sigmoid mapping module. The spatial attention module focuses on the importance of features in different spaces, emphasizing "where" the information part is. This module is complementary to the channel attention and includes the average pooling layer AvgPool and the max pooling layer MaxPool, as well as the stacking module and the Sigmoid mapping module.

[0033] In one embodiment, the steps for assigning different weights to different features in step 3 are as follows:

[0034] Step 3.1, input each fused feature layer in step 2 into ECBAM respectively, and set the input fused feature layer (the output C1_out~C6_out of step 2) as where H and W are the height and width of each fused feature layer respectively, and C is the number of channels of the fused feature layer; for First, compress them through average pooling and max pooling respectively to obtain two feature maps and

[0035] Step 3.2, ECBAM considers the interaction between each channel and its adjacent k channels, and performs operations on I avg 、I max respectively using one-dimensional convolutions of size k;

[0036] Step 3.3: Add the elements of the two parts of the result processed in Step 3.2, and map the channel features to the range (0, 1) through the Sigmoid activation function to obtain the weight coefficient M for different channels. c (I), and its calculation method is as follows:

[0037]

[0038] where σ is the Sigmoid activation function, is a one-dimensional convolution of size k, and the size of k is adaptively determined by the following formula:

[0039]

[0040] where C represents the number of channels of the fused feature layer and | | odd denotes taking the odd number closest to the result, γ = 2, b = 1;

[0041] Step 3.4: Multiply M c by the fused feature layer to obtain the channel attention feature map I'.

[0042] Step 3.5: Take the average pooling and max pooling of all channels of the same feature point of I' respectively with I' as the new input feature map and input it into the spatial attention module to obtain and

[0043] Step 3.6: Stack and and then perform a standard convolution;

[0044] Step 3.7: Use the Sigmoid activation function to obtain the spatial weight M s (I′) of I', and the calculation process is as follows:

[0045] M s (I′) = σ(f 7×7 ([avgPool(I′); maxPool(I′)])

[0046] = σ(f 7×7 ([I a ′ vg ; I′ max )

[0047] where f 7×7 is a convolution kernel of size 7×7.

[0048] Step 3.8: Multiply the obtained spatial weight M s(I′) is multiplied by the channel attention feature map to obtain the final feature map that can be sent for detection. At this time, the feature map is the DWS11, DWS13, DWS14_2, DWS15_2, DWS16_2, and DWS17_2 feature maps processed by IBFPN and ECBAM.

[0049] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0050] The SSD detection algorithm based on feature fusion and attention mechanism designed by the present invention surpasses other algorithms in terms of accuracy. Among them, compared with the original SSD_VGG16 algorithm, the detection accuracy and algorithm efficiency of the algorithm of the present invention are respectively improved by 3.04% and 0.9 FPS; compared with the SSD_MobileNet algorithm, the detection accuracy of the algorithm of the present invention is improved by 5.15%; the accuracy of the DSSD algorithm is close to that of the algorithm of the present invention, but its detection speed is lower than that of the algorithm of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 It is a schematic diagram of the feature pyramid network FPN and the local structure.

[0052] Figure 2 It is a schematic diagram of the residual feature enhancement module structure.

[0053] Figure 3 It is a schematic diagram of the bidirectional feature pyramid network module IBFPN structure.

[0054] Figure 4 It is a schematic diagram of the convolutional attention module ECBAM structure.

[0055] Figure 5 It is a schematic diagram of the channel attention module structure.

[0056] Figure 6 It is a schematic diagram of the spatial attention module structure.

[0057] Figure 7 It is the detection result of vehicles and pedestrians actually by the present invention.

[0058] Figure 8 It is the detection result of the present invention under the condition of mutual occlusion in the vehicle environment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0059] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0060] The present invention mainly studies the targets in the infrared scene. To demonstrate the effectiveness of the method of the present invention, the infrared dataset FLIR is used for the analysis and verification of the algorithm. All experiments use the Ubuntu16.04 operating system and are equipped with an NVIDIA-GTX1050Ti graphics card. The algorithms are all written in python3.6.6, using the Cuda9.0.136 parallel development architecture, and the algorithm model is built, trained and tested through the Pytorch framework. The overall training process of the detection model is divided into two stages: the freezing stage and the thawing stage. When training the parameters in the freezing stage, the backbone network of the model is frozen. At this time, the feature extraction network will not change. The number of each batch is batch_size = 16, the learning rate is 0.0005, and a total of 50 epochs are trained; in the thawing training stage, the overall model is trained, and all parameters will change. At this time, the number of each batch is batch_size = 8, the learning rate is 0.0001, and a total of 100 epochs are trained. The basic process is as follows:

[0061] Step 1: Input the infrared image into the traditional lightweight network MobileNet for layer-by-layer convolution calculation to obtain feature maps of different scales. Among them, the 6-layer feature maps of DWS11, DWS13, DWS14_2, DWS15_2, DWS16_2, and DWS17_2 are the feature pyramid images for detection;

[0062] Step 2: SSD adopts a detection method at the pyramid feature level. The image is input into the feature extraction network and forward propagated to obtain feature maps of different scales, and then predictions are made for each feature map separately. This method naturally generates features of different scales through convolution without additional computational cost, but only one layer of feature is utilized for each prediction, and there is no information exchange and complementation between different levels of features. Specifically, in the SSD detection algorithm, the depths of the convolutional layers are different, and the image information contained in their feature maps is also different. Among them, the shallow feature map contains information such as the texture and position of the image, but the extracted features are not comprehensive; the deep feature map contains richer semantic information, but many details are lost during the convolution process. The traditional SSD detection algorithm only inputs the information of each branch into the detection module separately, and there is no connection and complementation between the features of each layer. Adding an FPN structure to the SSD algorithm can strengthen the shallow features layer by layer, but the position information contained in the shallow features is not effectively transmitted to the deep features.

[0063] Therefore, the present invention establishes a bidirectional feature fusion module IBFPN that can effectively transmit shallow information on the basis of the original pyramid feature level. The feature pyramid image for detection is input into the IBFPN for mutual fusion of the upper and lower layer information, thereby improving the detection performance of the target.

[0064] Step 2.1, The Feature Pyramid Network (FPN) is the most typical multi-scale fusion module. By continuously upsampling the top-level features and fusing them with the features of other layers, enhanced feature maps can be obtained. These feature maps then reconstitute a new feature pyramid hierarchy, which can be input into the detection module to predict the object category and location. As Figure 1 shown, the FPN structure can be mainly divided into two parts. One is the bottom-up path on the left in the figure, and the other is the top-down path on the right and the lateral connections at the same level in the middle.

[0065] The bottom-up path is actually the process of forward convolution operation of the basic network. In this process, the feature maps will pass through a series of convolutional kernels. After passing through the convolutional kernel with a stride of 2, the size of the feature map is reduced to half of the original. After passing through the convolutional kernel with a stride of 1, its size remains unchanged. Therefore, these convolutional layers with the same size and adjacent to each other are regarded as the same stage. Since the semantic information extracted by the network becomes richer as the convolution deepens, the last layer of feature maps in each stage is used as the representative of the features at that scale. The representative feature maps of different scales constitute the pyramid hierarchy, as Figure 1 shown by C1, C2, and C3 in

[0066] Contrary to the bottom-up process, the top-down is a process of gradually fusing deep features with shallow features, as Figure 1 shown by C4, C5, and C6 in Figure 1 . Specifically: First, the higher-level features with stronger features are upsampled to enlarge their size to be the same as that of the next layer of feature maps. Then, for the feature layer to be fused, 1×1 convolution is used to adjust its number of channels to be the same as that of the higher-level features Figure 1 , and then the above two parts of features are added element by element through lateral connection to obtain the fused feature map. At this point, the feature fusion operation between adjacent layers is completed. Repeating the above operation for other feature layers to be fused can obtain a new feature pyramid hierarchy.

[0067] Step 2.2, In the feature pyramid, the features of the highest layer are propagated by top-down upsampling and gradually fused with the features of lower layers. During this process, the features of lower layers are enhanced by the semantic information from higher layers, naturally making the fused features contain different context information. However, due to size adjustment, the channel dimension of the highest layer features is compressed, which will lead to the loss of some information. When designing the feature fusion module in the present invention, the idea of residual feature enhancement in AugFPN is combined in the traditional Res residual structure. The introduced Residual Feature Augmentation (RFA) module uses the residual branch to add different context information to the original feature layer to reduce the loss of the highest layer features and improve the performance of the pyramid. The specific structure is as Figure 2As shown in the figure, first, an adaptive pooling operation is performed on the highest-level features of scale Z to generate multiple context features with different scales (a1×Z, a2×Z, …, an×Z). Then, 1×1 convolutions are performed on each context feature separately to reduce the channel dimension to a fixed size. Finally, they are upsampled to the same scale Z by bilinear interpolation for subsequent fusion.

[0068] Step 2.3. Considering the aliasing effect brought by interpolation, the simple addition operation cannot be performed on each part of the context features. Therefore, an Adaptive Spatial Fusion (ASF) module is added to the RFA residual feature enhancement module. Its specific structure is as Figure 2 shown on the right. These context features can be better combined through the ASF module. Specifically, this module takes the upsampled features of each part as inputs, generates a spatial weight for each feature through operations such as concatenation, convolution, and activation, and aggregates the context features into a new feature layer R using these weights. This feature layer has multi-scale context information. After the ASF generates the aggregated feature layer R, the sum is used to add R to the highest-level features to strengthen the scale information, and then the enhanced highest-level features are used for subsequent scale transformation and feature fusion.

[0069] Step 2.4. Obviously, the original SSD algorithm does not complement each other between the detection branches at this time. Based on the above idea, the present invention designs a new bi-directional feature fusion module, namely the Improved Bi-directional Feature Pyramid Network (IBFPN). This fusion module is improved on the basis of the traditional bi-directional feature pyramid network. The RFA is introduced to strengthen the top-level features, and a bottom-up fusion path is introduced. At the same time, for the same level, a connection from the starting input to the output is added. For each feature map used for detection, as Figure 3 shown. The RFA module introduced above can input the features of the C6 layer into this module to improve its feature representation and obtain R6 with multi-scale context information. The bottom-up path on the right part means further fusing the low-level features into the high-level ones, so that each layer not only has strong semantic information of the high level but also strong detail localization information of the low level. In addition, the structure designed by the present invention also adds a new connection from the starting input to the output for the same level, as Figure 3 shown by the red connection line in Figure 3 . This can fuse more features without increasing the number of parameters. Compared with the traditional SSD algorithm that directly detects the feature maps, the present invention first inputs each feature map into the bi-directional feature fusion module IBFPN before detection to perform the mutual fusion of the upper and lower layer information.

[0070] Step 2.6. Specifically, assume that the pyramid feature levels generated by the traditional SSD object detection model in the forward propagation process are {C1, C2, C3, C4, C5, C6} in sequence, corresponding to DWS11, DWS13, DWS14_2, DWS15_2, DWS16_2, and DWS17_2 in Step 1 respectively.

[0071] Step 2.7. Input the feature of layer C6 into RFA to improve its feature representation, and then obtain R6 with multi-scale context information through the ASF module, that is, RFA(C6).

[0072] Step 2.8. Add and fuse R6 with C6 after dimensionality reduction by 1×1 convolution to obtain the top-layer feature C6_td. Respectively perform dimensionality reduction on C1 to C5 through 1×1 convolution to obtain features C5_in to C1_in;

[0073] Step 2.9. Fuse C1_in to C6_in from top to bottom using the following formula.

[0074] C6_td = Conv(C6) + RFA(C6)

[0075] C5_td = Conv(C5) + Resize(C6_td)

[0076]

[0077] C1_td = Conv(C1) + Resize(C2_td)

[0078] Among them, Conv is a 1×1 convolution operation for feature channel dimensionality reduction, Conv(C6) to Conv(C1) are C6_in to C1_in respectively, RFA(*) is for residual feature enhancement of the feature map, Resize is an upsampling operation taken to match the resolutions of different layers, which is bilinear interpolation here, and "+" represents element-wise addition at corresponding positions. C1_td to C6_td are the top-layer features obtained after addition and fusion.

[0079] Step 2.10. The present invention newly aggregates a bottom-up fusion path, which means fusing low-level features into high-level features further from bottom to top, so that each layer not only has strong semantic information of the high level but also has strong detail localization information of the low level. The present invention also newly adds a connection from the starting input to the output for the same level, as shown by the red connection line in the appendix Figure 3 shown, which can fuse more features without increasing the number of parameters to obtain the fused feature layers C1_out to C6_out, and the fusion formula is as follows.

[0080] C1_out = C1_td = C1_in + Resize(C2_td)

[0081] C2_out = C2_td + Resize(C1_out) + C2_in

[0082]

[0083] C6_out = C6_td + Resize(C5_out) + C6_in

[0084] C1_out to C6_out are the feature layers of the multi-scale feature fusion output after C1 to C6 pass through IBFPN.

[0085] Step 3. Since the target information in the infrared image is not obvious and the feature structure is sparse, the complex background environment is also likely to cause phenomena such as target missed detection and false detection. To address this problem and avoid the influence brought by irrelevant information in the original SSD detection algorithm, the present invention improves on the basis of CBAM to obtain a new hybrid attention module ECBAM. After passing through IBFPN, each fused feature layer is continuously input into ECBAM, and different weights are assigned to different features through ECBAM, so that the model focuses more on the target part and pays more attention to the target area of interest, thereby improving the network expression ability and suppressing the influence brought by irrelevant information.

[0086] Step 3.1. The original Convolutional Block Attention Module (CBAM) mainly includes a channel attention module and a spatial attention module. For the input feature map, the channel attention and spatial attention are processed in sequence to achieve the purpose of assigning different weights to features with different degrees of importance. However, during the channel attention operation, CBAM uses a fully connected process with parameter sharing for the results of max pooling and average pooling, which will cause the model to not be able to well balance the mapping of the two features at the same time; on the other hand, the fully connected layer considers the global feature correlation, so it will lead to too high model complexity and computational cost. To address this problem, the present invention proposes an improved convolutional attention module according to the idea of local cross-channel interaction.

[0087] Step 3.2. ECBAM is obtained by improving on the basis of CBAM. Its specific structure is as shown in the appendix Figure 4 and mainly includes a channel attention module and a spatial attention module. The channel attention module focuses on the importance of different feature channels and emphasizes "what" information to pay attention to. The spatial attention module focuses on the importance of features in different spaces and emphasizes "where" the information part is. This module is complementary to the channel attention. Assume that the input feature map is Pass through the channel attention module in sequence and the spatial attention module for operations, and then the final processing result can be obtained.

[0088] The specific structure of the channel attention module is as shown in the appendix Figure 5 and includes an average pooling layer AvgPool and a max pooling layer MaxPool, a one-dimensional convolutional module, and a Sigmoid mapping module. In order to highlight the relationship between channels, it is necessary to compress the length and width of the input image. Average pooling and max pooling can characterize the importance of different features from different aspects. The one-dimensional convolutional module considers the interaction of each channel and its adjacent k channels. While bringing performance gains, this module reduces the number of parameters to a constant order of magnitude and can well balance the mapping of two different features in the pooling layer. The Sigmoid mapping module maps each feature to the range (0, 1).

[0089] The specific structure of the spatial attention module is as shown in the appendix Figure 6 and includes an average pooling layer AvgPool and a max pooling layer MaxPool, a stacking module, and a Sigmoid mapping module. First, average pooling and max pooling are performed in the channel dimension, and then the feature maps generated by them are concatenated. Then, on the concatenated feature map, convolutional operations are used to generate the final spatial attention feature map.

[0090] For the step of assigning different weights to different features, each fused feature layer in step 2 is respectively input into ECBAM. The input fused feature layers (the outputs C1_out to C6_out of step 2) are set as where H and W are the height and width of each fused feature layer respectively, and C is the number of channels of the fused feature layer. Each fused feature layer passes through the channel attention module and the spatial attention module for operations, and then the final processing result can be obtained. The calculations of each part above can be expressed by the following formula.

[0091]

[0092]

[0093] where represents element-wise multiplication, I′ is the result after channel attention processing, and I″ is the final result.

[0094] Step 3.3, for the features in step 2 First, they are respectively compressed by average pooling and max pooling to obtain two feature maps and

[0095] Step 3.4. In the original CBAM, the parameter - shared fully - connected processing in CBAM causes the model to not be able to well balance the mapping of the two features simultaneously, and also leads to too high model complexity and computational cost. Different from the shared fully - connected operation in CBAM, ECBAM considers the interaction between each channel and its adjacent k channels, and performs operations on I avg and I max respectively using one - dimensional convolutions of size k. This operation not only brings performance gains but also reduces the number of parameters to an order of magnitude of constants.

[0096] Step 3.5. Add the results of the two parts processed in Step 3.4 element - by - element, and pass them through the Sigmoid activation function to map the features of each channel into the range (0, 1), then the weight coefficients M c (I) can be obtained. The calculation formula is as follows:

[0097]

[0098] where σ is the Sigmoid activation function, is the one - dimensional convolution of size k, and the size of k can be adaptively determined by the following formula.

[0099]

[0100] where C represents the number of channels of the fused feature layer , || odd represents taking the odd number closest to the result, γ = 2, b = 1.

[0101] Step 3.6. Finally, multiply M c (I) by the input feature, that is, the fused feature layer , then the channel - attention feature map I' can be obtained.

[0102] As Figure 5 shown, using I' as the new input feature map and inputting it into the spatial - attention module, perform average pooling and max - pooling on all channels of the same feature point of I' respectively, to obtain and

[0103] Step 3.7. Stack and and then perform standard convolution.

[0104] Step 3.8. Use the Sigmoid activation function to obtain the spatial weight M s (I′) of I'. The calculation process can be expressed as the following formula.

[0105] M s (I′) = σ(f 7×7([avgPool(I′); maxPool(I′)])

[0106] = σ(f 7×7 ([I avg ; I′ max )

[0107] where σ is the Sigmoid activation function, and f 7×7 is a convolutional kernel of size 7×7.

[0108] Multiply the obtained spatial weight M s (I′) with the channel attention feature map to obtain the final feature map that can be sent for detection.

[0109] Step 4: Send the DWS11, DWS13, DWS14_2, DWS15_2, DWS16_2, and DWS17_2 feature maps processed by IBFPN and ECBAM into the detection module for detection, obtain the category and corresponding bounding box of each candidate box, and get the prediction results.

[0110] Step 5: Perform non-maximum suppression on the prediction results, delete the redundant target boxes, and obtain the final detection results.

[0111] Step 6: Conduct experimental verification for each module and perform verification and analysis on the overall detection algorithm.

[0112] Step 7: Verification of the IBFPN module. The experiment is based on the MobileNet feature extraction network. Comparative experiments are designed by performing feature fusion using the original feature pyramid hierarchy, the traditional FPN structure, the bidirectional FPN structure, and the IBFPN designed in the present invention. The specific detection results are shown in Table 1.

[0113] Table 1 Experimental comparison results of feature fusion modules

[0114]

[0115] As can be seen from Table 1, compared with not using the feature fusion module, the traditional FPN has a certain improvement in the detection accuracy of small targets such as "person" and "bicycle", but the overall effect is not obvious, and the mAP (mean AP) increases by 0.63%; the bidirectional FPN not only improves the detection effect of small targets but also improves the detection accuracy of the large target "car", and its mAP value increases from the original 49.61% to 50.83%; the IBFPN module designed in the present invention significantly enhances the detection effect of both large and small targets, and the overall mAP value increases by 2.03%.

[0116] Step 8: Verification of the ECBAM module

[0117] To verify the effectiveness of the attention module ECBAM designed in the present invention for the SSD detection algorithm, experiments were conducted by adding different attention modules to each feature branch. Specifically: directly inputting into the detection module without additional processing for the branch, adding the channel attention module SE after the branch, adding the hybrid attention module CBAM, and adding the ECBAM module with the improved design of the present invention. The specific detection results are shown in Table 2.

[0118] Table 2 Experimental comparison results of attention modules

[0119] Attention Module Not Used SENet CBAM ECBAM mAP (%) 49.61% 50.17% 50.62% 51.04%

[0120] As can be seen from Table 2, the target detection effect has been improved to a certain extent after adding the attention module. Among them, the detection effect of the hybrid attention module is better than that of the single channel attention module. The mAP value of CBAM is 0.45% higher than that of SENet; ECBAM changes the shared fully connected layer in CBAM to a one-dimensional convolution, reducing the complexity and enhancing the accuracy, and its mAP value has increased by 0.42%; compared with not using the attention module, the mAP value of ECBAM has increased by 1.43%. This shows that the use of the attention module can effectively suppress the influence brought by the irrelevant background, thereby reducing the false detection rate and increasing the detection accuracy of the target.

[0121] Step 9, overall comparative analysis of the detection algorithm

[0122] To illustrate the effectiveness of the algorithm of the present invention as a whole, the SSD_BIFPN_ECBAM algorithm designed in the present invention was compared and analyzed with various detection algorithms on the FLIR dataset, including the SSD algorithm using different feature extraction networks, the SSD_MobileNet algorithm, and the DSSD algorithm with different feature fusions. The detection results are shown in Table 3.

[0123] Table 3 Experimental comparison results of detection algorithms

[0124] network Backbone mAP (%) FPS SSD VGG16 49.98% 16.7 SSD MobileNet 47.87% 20.5 DSSD ResNet 52.84% 15.7 SSD_BIFPN_ECBAM MobileNet 53.02% 17.6

[0125] As can be seen from Table 3, the object detection method based on feature fusion and attention mechanism designed in the present invention surpasses other algorithms in terms of accuracy. Among them, compared with the original SSD_VGG16 algorithm, the detection accuracy and speed of the algorithm of the present invention have increased by 3.04% and 0.9 FPS respectively; compared with the SSD_MobileNet algorithm, the detection accuracy of the algorithm of the present invention has increased by 5.15%; the accuracy of the DSSD algorithm is close to that of the algorithm of the present invention, but its detection speed is lower than that of the algorithm of the present invention.

[0126] As Figure 7 and Figure 8 shown. Among them, Figure 7It includes a car and a pedestrian, and it can be seen that the algorithm has successfully detected the targets in the picture. Only a small part of the front of a passing car is shown on the right side of the picture, but the algorithm still successfully detects it; Figure 8 The difficulty of the scenario lies in the mutual occlusion of cars or occlusion by the surrounding environment, but it can be seen that the algorithm still successfully detects all targets in the above two complex scenarios, with good detection results.

Claims

1. A target detection method based on feature fusion and attention mechanism, characterized in that It includes the following steps: Step 1: Input the infrared image into the MobileNet network for layer-by-layer convolutional calculation to obtain feature maps of different scales. Among them, the feature maps of 6 layers, namely DWS11, DWS13, DWS14_2, DWS15_2, DWS16_2, and DWS17_2, are the feature pyramid images for detection; Step 2: Establish a bidirectional feature fusion module IBFPN, and input the feature pyramid images for detection into the IBFPN to perform mutual fusion of upper and lower layer information; the IBFPN is based on the bidirectional pyramid network, constructs a residual feature enhancement module RFA to enhance the top layer features, introduces a bottom-up fusion path, and adds a connection from the starting input to the output for the same level; The RFA introduces the idea of residual feature enhancement in AugFPN into the Res residual structure and integrates the adaptive spatial fusion module. The RFA increases the context information through residual features to reduce the loss of the top layer features and improve the performance of the pyramid; Step 3: After passing through the IBFPN, input each fused feature layer into the attention module ECBAM to assign different weights to different features; Step 4: Send DWS11, DWS13, DWS14_2, DWS15_2, DWS16_2, and DWS17_2 processed by the IBFPN and ECBAM into the detection module for detection, obtain the category and corresponding bounding box of each candidate box, and get the prediction result; Step 5: Perform non-maximum suppression on the prediction result, delete the redundant target boxes, and obtain the final detection result.

2. The object detection method based on feature fusion and attention mechanism according to claim 1, wherein In step 2, the steps for mutual fusion of upper and lower layer information are as follows: Step 2.1: In the forward propagation process, the pyramid feature levels generated by the traditional SSD object detection model are {C1, C2, C3, C4, C5, C6} in sequence, corresponding to DWS11, DWS13, DWS14_2, DWS15_2, DWS16_2, and DWS17_2 in step 1 respectively; Step 2.2: Input the C6 layer feature into the RFA to improve its feature representation, obtain the output feature layer R6 with multi-scale context information, that is, RFA(C6); Step 2.3: Add and fuse R6 with C6 after dimension reduction by 1×1 convolution to obtain the top layer feature C6_td; Step 2.4: Respectively reduce the dimensions of C1~C5 by 1×1 convolution to obtain features C5_in~C1_in; Step 2.5: Use the following formula to perform top-down fusion of C1_in~C6_in: where Conv is the 1×1 convolution operation for feature channel dimension reduction, Conv(C6)~Conv(C1) are C6_in~C1_in, RFA(*) is to perform residual feature enhancement on the feature map, Resize is the upsampling operation taken to match the resolutions of different layers, and "+" represents element addition at the corresponding positions; C1_td~C6_td are the top layer features obtained after addition and fusion; Step 2.6, from bottom to top, fuse the low-level features into the high-level features, so that each layer has both strong semantic information of the high level and strong detail localization information of the low level, and add a connection from the starting input to the output to fuse more features without increasing the number of parameters, obtaining the fused feature layers C1_out to C6_out. The fusion formula is as follows: C1_out to C6_out are the feature layers of multi-scale feature fusion output after C1 to C6 pass through IBFPN.

3. The object detection method based on feature fusion and attention mechanism according to claim 2, wherein The upsampling operation is bilinear interpolation.

4. The object detection method based on feature fusion and attention mechanism according to claim 1 or 2 or 3, characterized in that, The attention module ECBAM includes a channel attention module and a spatial attention module. The channel attention module focuses on the importance of different feature channels, including an average pooling layer AvgPool, a max pooling layer MaxPool, a one-dimensional convolution module, and a Sigmoid mapping module. The spatial attention module focuses on the importance of features in different spaces, emphasizing "where" is the information part. The spatial attention module is complementary to the channel attention module and includes an average pooling layer AvgPool, a max pooling layer MaxPool, a stacking module, and a Sigmoid mapping module.

5. The object detection method based on feature fusion and attention mechanism according to claim 4, characterized in that For Step 3, the steps of assigning different weights to different features are as follows: Step 3.1, input each of the fused feature layers in Step 2 into ECBAM respectively, and set the input fused feature layer as where H and W are the height and width of each fused feature layer respectively, and C is the number of channels of the fused feature layer; for first, compress them respectively through average pooling and max pooling to obtain two feature maps and Step 3.2, ECBAM considers the interaction between each channel and its adjacent k channels, and performs operations on I avg and I max respectively using one-dimensional convolutions of size k; Step 3.3: Add the elements of the two parts of the result processed in Step 3.2, and map the feature of each channel to the range of (0, 1) through the Sigmoid activation function to obtain the weight coefficient M of different channels c (I), and its calculation method is as follows: where σ is the Sigmoid activation function, is a one-dimensional convolution of size k, and the size of k is adaptively determined by the following formula: Among them, || odd represents taking the odd number closest to the result, γ = 2, b = 1; Step 3.4, multiply M c (I) with the fused feature layer to obtain the channel attention feature map I', Step 3.5, take I' as the new input feature map and input it into the spatial attention module. Perform average pooling and max pooling on all channels of the same feature point of I' respectively to obtain and Step 3.6, stack and and then perform standard convolution; Step 3.7, use the Sigmoid activation function to obtain the spatial weight M of I', and the calculation process is as follows: s (I′) M s (I′) = σ(f 7×7 ([avgPool(I′); maxPool(I′)]) = σ(f 7×7 ([I′ avg ; I′ max ) where f 7×7 is a convolutional kernel of size 7×7; Step 3.8, multiply the obtained spatial weight M s (I′) by the channel attention feature map to obtain the final feature map that can be sent for detection. At this time, the feature map is the DWS11, DWS13, DWS14_2, DWS15_2, DWS16_2, and DWS17_2 feature maps processed by IBFPN and ECBAM.