A target detection method based on attention dual-mode feature fusion

CN118155035BActive Publication Date: 2026-10-09NORTHWEST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410333908.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-22
Publication Date
2026-10-09
Estimated Expiration
2044-03-22

AI Technical Summary

Technical Problem

首先,基于可见光RGB图像的方法对于光照变化和背景干扰比较敏感,导致在复杂环境下的目标检测准确性下降

Benefits of technology

[0058] This invention first introduces a dual-stream feature extraction network, considering features from both the visible light and infrared modes, enabling effective fusion of features from different modes. This overcomes the limitations of traditional single-mode feature extraction and improves the robustness of target detection. Secondly, by introducing coordinate attention and channel attention, this invention achieves refined enhancement of visible light and infrared mode features, effectively highlighting important information, suppressing redundancy and noise, and further improving detection accuracy. Finally, the introduction of a dual-mode fusion module enables information complementarity between different modes, fully utilizing the complementarity between them and effectively improving the overall performance of target detection. Compared with existing methods, this invention achieves significant advantages in accuracy, efficiency, and robustness. Therefore, this invention is of great significance for promoting the development of target detection technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118155035B_ABST
    Figure CN118155035B_ABST
Patent Text Reader

Abstract

The application discloses a target detection method based on attention dual-mode feature fusion, comprising the following steps: S1, acquiring a visible light and infrared paired data set; S2, constructing a dual-flow feature extraction backbone network for extracting visible light features and infrared features, respectively obtaining visible light convolutional features F RGB or infrared convolutional features F IR ; S3, fusing the visible light convolutional features F RGB or the infrared convolutional features F IR by an ADFM dual-mode fusion module to obtain fused features F fused ; S4, inputting the fused features F fused into a detection neck to further realize multi-scale feature fusion and obtain multi-scale features; and S5, inputting the multi-scale features obtained from the detection neck into a detection head to output detection frame positions, detection object categories and confidence information. The application can help researchers to quickly and accurately detect targets in images, especially under the interference of a complex background; and the scheme can fully utilize the advantages of various modalities to improve the accuracy and robustness of target detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of target detection technology, and in particular relates to a target detection method based on attention-based dual-mode feature fusion. Background Technology

[0002] Object detection has a history of more than 20 years. During this long development process, many pioneering methods have emerged, which can be roughly divided into methods based on traditional machine learning and methods based on deep learning.

[0003] Traditional machine learning methods were primarily used in the early stages, mostly based on hand-constructed features. Due to limitations in complex feature representations and limited computational resources, detection accuracy was difficult to improve. With the rise of convolutional neural networks, using deep convolutional neural networks to learn high-level feature representations of images has gradually become mainstream. As a large number of researchers transitioned from traditional methods to deep learning methods, object detection began to develop at an unprecedented pace.

[0004] However, both in the past and present, many object detection algorithms have used visible light RGB images as the sole input and processing object. In the real world, however, dynamic environmental changes such as illumination variations, occlusion, and weather conditions interfere with the input image, posing new challenges to the accuracy of models and algorithms. Relying solely on visible light RGB images to complete object detection tasks in specific scenarios is difficult. Specifically, this manifests in the following ways: First, methods based on visible light RGB images are sensitive to illumination changes and background interference, leading to decreased object detection accuracy in complex environments. Second, visible light RGB images cannot accurately capture the shape, texture, and detail information of objects in certain situations, limiting the accuracy and robustness of object detection. Furthermore, some objects are difficult to distinguish from the background in visible light RGB images, making them more susceptible to misclassification or missed detection.

[0005] To overcome these problems, introducing other modalities to achieve complementary advantages is an effective approach. In recent years, multimodal data has been widely used in many practical applications. Combining the internal information of multimodal data can effectively transmit complementary features and avoid the omission of certain information from a single modality. For example, introducing infrared images as an additional modality can compensate for the limitations of RGB images. Infrared images can capture the heat distribution and infrared radiation characteristics of a target, are unaffected by changes in lighting and background interference, and can provide more reliable target detection results at night or in low-light environments. Fusion of image features from two modalities to enhance a single modality is of great significance in some practical scenarios. At the same time, fusion methods can also provide valuable insights and references for problems in other fields, showing broad application prospects. Summary of the Invention

[0006] The purpose of this invention is to provide a target detection method based on attention-based dual-mode feature fusion, which leverages the complementary advantages of visible and infrared mode fusion to further improve the detection performance, accuracy, and robustness of the target detector.

[0007] To solve the above problems, the technical solution adopted by the present invention is as follows:

[0008] A target detection method based on attention-based dual-mode feature fusion is performed according to the following steps:

[0009] S1: Obtain the paired dataset of visible light and infrared light;

[0010] S2: Construct a dual-stream feature extraction backbone network to extract visible light and infrared features, respectively, to obtain visible light convolutional features F. RGB Or infrared convolution feature F IR ;

[0011] S3: Visible light convolutional features F are processed through the ADFM dual-mode fusion module. RGB Or infrared convolution feature F IR The fusion is performed to obtain the fusion feature F. fused ;

[0012] S4: Fuse features F fused Input the neck detection and further achieve multi-scale feature fusion to obtain multi-scale features;

[0013] S5: Input the multi-scale features obtained from the neck detection into the detection head, and output the detection box position, detection object category and confidence information.

[0014] Optionally, in S2, constructing the dual-stream feature extraction backbone network includes:

[0015] The dual-stream feature extraction backbone network consists of a first-layer CBS module, a third-layer CBS module + C3 module, and a fifth-layer SPPF module. The input visible light RGB image or infrared IR image is first downsampled by the first-layer CBS module, then further processed by a stacked combination of three CBS modules + C3 modules for feature extraction. Finally, the SPPF module fuses multi-scale features to obtain the visible light convolutional features F. RGB Or infrared convolution feature F IR .

[0016] Optionally, in S3, the visible light convolutional features F are processed by the ADFM dual-mode fusion module. RGB Or infrared convolution feature F IR The integration process specifically includes:

[0017] For the obtained visible light convolution features F RGB A CA coordinate attention module is introduced to enhance the representation of visible light features, as shown in the formula:

[0018]

[0019] Where i represents the vertical direction and j represents the horizontal direction. and Indicates the weights in two spatial directions;

[0020] The obtained infrared convolution features F IR The input is then fed into the SE channel attention module to perform nonlinear modeling on the features, resulting in the extracted features F. I ′ R The formula is:

[0021] F′ IR =s c F IRc ;

[0022] Where s c F represents the weight value of the c-th channel. IRc The feature representing the c-th channel;

[0023] After the above operations, enhanced visible light convolutional features F′ were obtained. RGB and enhanced infrared convolution features F′ IR By adjusting the number of feature channels using 1×1 convolutions, attention feature maps for different modalities in the spatial domain are obtained, as expressed by the following formula:

[0024] m RGB =f1(F′ RGB );

[0025] m IR =f2(F′ IR );

[0026] Where m RGB This represents the attention feature map of the visible light RGB modes in the spatial domain, m IR f1 and f2 represent the attention feature map of the infrared (IR) mode in the spatial domain, and f1 and f2 represent the 1×1 convolutional blocks of the RGB and IR modes, respectively.

[0027] Features F′ extracted from different modalities through an attention mechanism IR and F′ RGB By performing element-wise matrix multiplication with the corresponding spatial domain feature map, the internal spatial information between different modes can be obtained, specifically as follows:

[0028]

[0029]

[0030] Among them, F in1 For the internal spatial information of the visible light RGB mode, F in2 For the internal spatial information of the infrared (IR) mode, Represents element-wise matrix multiplication;

[0031] The internal spatial information of different modalities is added to the original convolutional features and then input into a 1×1 convolution to obtain the complete feature F. full1 and F full2 The specific formula is as follows:

[0032] F full1 =f3(F in1 +F RGB );

[0033] F full2 =f4(F in2 +F IR );

[0034] Where f3 and f4 represent 1×1 convolutional blocks;

[0035] Finally, the preceding features are concatenated and fused using SE attention blocks to obtain the final fused feature F. fused The formula is expressed as:

[0036] F fused =SE(Concat(F) full1 ,F full2 ));

[0037] SE(.) is the same as the SE channel attention module mentioned earlier; Concat(.) represents the concatenation operation along the channel axis.

[0038] Optionally, S5 specifically includes:

[0039] The multi-scale features obtained in S4 are input into three detection heads, and each detection head outputs the detection box position, the detection object category, and confidence information.

[0040] The loss value is calculated based on the prediction results and the true label values. The loss functions used are bounding box regression loss and... Classification loss and confidence loss

[0041] Optionally, the bounding box regression loss Classification loss and confidence loss Specifically as follows:

[0042] (1) The formula for the bounding box regression loss is as follows:

[0043]

[0044] Among them, S 2 N represents the number of image grids in the prediction process and the number of predicted bounding boxes in each grid; Representing the true value, the predicted bounding box, and the contained value, respectively. and The smallest closed frame; coefficient This indicates whether the j-th prediction box in the i-th grid is a positive sample.

[0045] (2) The classification loss formula is as follows:

[0046]

[0047] Where p(c) represents the probability that the true sample belongs to class c; This represents the probability that the network predicts a sample to be of class c; the coefficient. The meaning is the same as before. Maintain consistency with the central government;

[0048] (3) The confidence loss formula is as follows:

[0049]

[0050] Where the coefficient With the previous Conversely, c represents whether the j-th predicted box in the i-th grid is a negative sample; i and This represents the confidence level of the true value and the confidence level of the network prediction.

[0051] Optionally, the total loss function can be defined as follows: The calculation formula is as follows:

[0052]

[0053] Optionally, in step S4, the fused features F at different scales obtained in S3 are... fused Further feature extraction is performed by inputting the species from each layer of the neck area for detection.

[0054] In the detection of the neck, multi-scale feature fusion was achieved through top-down and bottom-up feature extraction and horizontal feature splicing.

[0055] Optionally, in S1, the acquired visible light and infrared paired dataset is:

[0056] The LLVIP and DroneVehicle datasets were used, and the datasets were divided into training and test sets according to a set ratio.

[0057] Compared with the prior art, the present invention has the following beneficial effects:

[0058] This invention first introduces a dual-stream feature extraction network, considering features from both the visible light and infrared modes, enabling effective fusion of features from different modes. This overcomes the limitations of traditional single-mode feature extraction and improves the robustness of target detection. Secondly, by introducing coordinate attention and channel attention, this invention achieves refined enhancement of visible light and infrared mode features, effectively highlighting important information, suppressing redundancy and noise, and further improving detection accuracy. Finally, the introduction of a dual-mode fusion module enables information complementarity between different modes, fully utilizing the complementarity between them and effectively improving the overall performance of target detection. Compared with existing methods, this invention achieves significant advantages in accuracy, efficiency, and robustness. Therefore, this invention is of great significance for promoting the development of target detection technology. Attached Figure Description

[0059] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some example drawings of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0060] Figure 1 This is a schematic diagram of the target detection method based on attention-based dual-mode feature fusion of the present invention;

[0061] Figure 2 This is a schematic diagram of the network structure of the target detection method based on attention dual-mode feature fusion of the present invention;

[0062] Figure 3 This is a schematic diagram of the attention dual-mode feature fusion module structure of the present invention;

[0063] Figure 4 This is a schematic diagram of the detection results of the present invention on the LLVIP dataset;

[0064] Figure 5 This is a schematic diagram of the detection results of the present invention on the DroneVehicle dataset. Detailed Implementation

[0065] The core of this invention is to provide a target detection method based on attention-based dual-modal feature fusion. This method enhances a single modality by fusing image features from two modalities, which is significant in some practical scenarios. Furthermore, the fusion method can also provide valuable insights and references for problems in other fields, showing broad application prospects.

[0066] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings, but this description is not intended to limit the present invention.

[0067] Includes the following steps:

[0068] S1: Relevant visible-infrared paired datasets were acquired: the LLVIP dataset and the DroneVehicle dataset. The LLVIP dataset, released in 2021, is a visible-infrared paired pedestrian dataset for low-light vision. This dataset collected 15,488 pairs of visible-infrared images from 26 different locations. Each of these 15,488 image pairs contains pedestrians, most of which were taken in low-light environments. The DroneVehicle dataset is a large visible-infrared dataset based on drone-based imagery, mainly containing 953,087 object instances from 56,878 images recorded from different scenes, half of which are visible-infrared RGB images and half are infrared images. The objects in this dataset are different types of vehicles, broadly divided into five categories: cars, trucks, buses, vans, and freight cars. These datasets were then proportionally divided into training and test sets.

[0069] S2: Based on the YOLOv5 single-stage object detection algorithm, a dual-stream feature extraction backbone network was constructed to extract features from visible light and infrared images separately, while providing feature inputs from different modalities for the subsequent feature fusion module. The constructed dual-stream feature extraction backbone network is used to extract features from visible light and infrared images separately, while providing feature inputs from different modalities for the subsequent feature fusion module. Combined with... Figure 2Specifically, the process includes: the input visible light RGB image first undergoes a CBS module for initial downsampling, then passes through a stacked module of three CBS and C3 layers for further feature extraction, and finally, the SPPF module fuses multi-scale features, stitching together feature representations from different scales of the same feature map. The infrared (IR) image processing is similar and will not be elaborated further. Simultaneously, in layers 3, 4, and 5, intermediate-level feature fusion is performed on the extracted features for use in subsequent neck and head detection. The backbone network of each branch primarily obtains five layers of visible light and infrared convolutional feature maps at different scales through stacked convolutional layers, preparing for subsequent fusion modules.

[0070] S3: Coordinate Attention (CA) and Squeeze-and-Excitation (SE) channel attention modules are introduced to enhance the extracted visible light and infrared convolutional features. For visible light convolutional features, a CA coordinate attention module is added to the fusion module. By considering the relationships between coordinates, the model can better capture long-range dependencies and spatial structures in visible light images, improving its modeling ability for visible light features. For complex object edges, structures, and temperature distributions in infrared images, nonlinear modeling is performed. By introducing an SE channel attention module, the expressive power of infrared features is improved.

[0071] S4: To better utilize the enhanced visible light and infrared features, an attention-based dual-mode fusion module (ADFM) was constructed. The visible light and infrared features extracted from the two-stream network are enhanced separately by the attention module, and then a residual network structure is used to fuse the features of each modality, improving the network's ability to perceive information from different modalities.

[0072] S5: The features output by the attention dual-mode feature fusion module ADFM are input again into the visible light feature extraction branch and the infrared feature extraction branch to achieve information complementarity. Three fusion modules are used in the network to realize multi-scale feature fusion, so that the network can adapt to multi-scale target detection tasks.

[0073] S6: By inputting the obtained multi-scale features into the detection head, the final output includes the detection box position, the detected object category, and confidence information. Researchers can classify and identify targets more quickly, improving work efficiency. The network uses three types of loss functions during training: bounding box regression loss... Classification loss and confidence loss Specifically as follows:

[0074] (1) Bounding box regression loss

[0075] Bounding box regression loss is used to measure the error between the predicted box and the actual box, enabling the detector to more accurately detect the size of the target. The specific formula is as follows:

[0076]

[0077] Among them, GIoU (Generalized Intersection over Union) loss It is used to predict bounding box regression loss. GIoU loss is a better choice than IoU loss; S 2 N represents the number of image grids in the prediction process and the number of predicted bounding boxes in each grid; These represent the true value, the predicted bounding box, and the contained value, respectively. and The smallest closed frame; coefficient This indicates whether the j-th prediction box in the i-th grid is a positive sample.

[0078] (2) Classification loss

[0079] To measure the difference between the class predicted by the model and the actual label, a binary cross-entropy loss function is used for each label in the network to reduce computational complexity and improve model performance. The specific formula is as follows:

[0080]

[0081] Where p(c) represents the probability that the true sample belongs to class c; This represents the probability that the network predicts a sample to be of class c; the coefficient. The meaning is the same as before. Maintain consistency.

[0082] (3) Confidence loss

[0083] The formula used to measure the difference between the model's confidence in the predicted bounding box and the actual label is as follows:

[0084]

[0085] Where the coefficient With the previous Conversely, c represents whether the j-th predicted box in the i-th grid is a negative sample; i and This represents the confidence level of the true value and the confidence level of the network prediction.

[0086] The network's final total loss function can be defined as follows: The calculation formula is as follows:

[0087]

[0088] For a more detailed process, please refer to the appendix. Figure 1 , Figure 1 The following is a logical diagram of the target detection method based on attention-based dual-mode feature fusion provided in an embodiment of the present invention, which is performed according to the following steps:

[0089] S1: Obtain the visible light and infrared paired dataset. Specifically, this step includes the following sub-steps:

[0090] S11: Process the dataset labels according to the YOLO data format;

[0091] S12: Store the dataset in the specified folder according to the ratio of 7 training sets: 3 test sets.

[0092] S2: Construct a dual-stream feature extraction backbone network to extract visible light and infrared features respectively. The detailed process is as follows:

[0093] S21: Based on the YOLOv5 feature extraction network, a two-stream feature extraction backbone network was constructed, the specific structure of which is as follows: Figure 2 As shown, the dual-stream feature extraction network is mainly built on the YOLOv5 algorithm backbone network, used to extract features from the visible light mode and the infrared mode respectively, and continuously shrinking the feature map. The main structure of the network consists of CBS modules, C3 modules, and SPPF modules. For example, the dual-stream feature extraction backbone network consists of a first-layer CBS module, three layers of CBS modules + C3 modules (second-layer CBS modules + C3 modules, third-layer CBS modules + C3 modules, and fourth-layer CBS modules + C3 modules), and a fifth-layer SPPF module. The input visible light RGB image first passes through a CBS module for first-layer downsampling, then passes through a stacked combination of three CBS and C3 modules for further feature extraction, and finally passes through the SPPF module to fuse multi-scale features, stitching together the feature representations of the same feature map at different scales. The infrared IR image follows a similar process and will not be described in detail.

[0094] S22: Extract the visible light features F from layers 3, 4, and 5. RGB Infrared signature F IR Then, input it into the ADFM module for fusion.

[0095] S3: The extracted features are fused using the ADFM dual-mode fusion module. Figure 3 The specific process is as follows:

[0096] S32: For the visible light convolution feature F obtained in S21RGB By introducing a CA coordinate attention module, enhanced representation of visible light features is achieved. The formula is:

[0097]

[0098] Where i represents the vertical direction and j represents the horizontal direction. and This represents the weights in two spatial directions. Finally, after normalization using the Sigmoid activation function, the features are reassigned weights and multiplied pixel-by-pixel with the original features to obtain the enhanced feature F. R ′ GB .

[0099] S32: For the infrared convolutional feature F obtained in S21 IR The feature is then input into the SE channel attention module to perform non-linear modeling on the feature, resulting in the extracted feature F′. IR The formula is:

[0100] F′ IR =SE(F IR ) = s c F IRc ;

[0101] Where s c F represents the weight value of the c-th channel. IRc This represents the feature of the c-th channel.

[0102] S33: Enhanced features F′ were obtained in S31 and S32 respectively. RGB and F′ IR By adjusting the number of feature channels using 1×1 convolutions, attention feature maps for different modalities in the spatial domain are obtained, as expressed by the following formula:

[0103] m RGB =f1(F′ RGB );

[0104] m IR =f2(F′ IR );

[0105] Where m RGB This represents the attention feature map of the visible light RGB modes in the spatial domain, m IR f1 and f2 represent the attention feature map of the infrared (IR) mode in the spatial domain, and f1 and f2 represent the 1×1 convolutional blocks of the RGB and IR modes, respectively.

[0106] S34: Features F′ extracted from different modalities via an attention mechanism IR and F′ RGBBy performing element-wise matrix multiplication with the corresponding spatial domain feature map, the internal spatial information between different modes can be obtained, specifically as follows:

[0107]

[0108]

[0109] Among them, F in1 For the internal spatial information of the visible light RGB mode, F in2 For the internal spatial information of the infrared (IR) mode, This represents element-wise matrix multiplication.

[0110] S35: To further integrate internal view information and spatial texture information, the internal spatial information of different modalities is added to the original convolutional features and then input into a 1×1 convolution to obtain the complete feature F. full1 and F full2 The specific formula is as follows:

[0111] F full1 =f3(F in1 +F RGB );

[0112] F full2 =f4(F in2 +F IR );

[0113] Where f3 and f4 represent 1×1 convolutional blocks.

[0114] S36: Finally, the features from the previous steps are concatenated and fused using the SE attention block to obtain the final fused feature F. fused The formula is expressed as:

[0115] F fused =SE(Concat(F) full1 ,F full2 ));

[0116] SE(.) is the same as the SE channel attention module mentioned earlier; Concat(.) represents the concatenation operation along the channel axis.

[0117] S4: Input the fused features into the neck detection area to further achieve multi-scale feature fusion. The specific process is as follows:

[0118] S41: The fusion features F obtained from S3 at three different scales fused Further feature extraction is performed by inputting the species from each layer of the neck area for detection.

[0119] S42: In the detection of the neck, multi-scale feature fusion is achieved through top-down and bottom-up feature extraction, as well as horizontal feature splicing.

[0120] S5: Input the multi-scale features obtained from neck detection into the detection head, and output the detection box position, detected object category, and confidence information. The detailed process is described below:

[0121] S51: The multi-scale features obtained in S4 are input into three detector heads to predict large, medium and small targets respectively, and output the detection box position, detection object category and confidence information. Each detector head will output three pieces of information, and finally the result with the highest confidence is taken.

[0122] S52: Calculate the loss value based on the prediction results and the true label values ​​to prepare for subsequent weight updates. The loss functions used are bounding box regression loss and... Classification loss and confidence loss Specifically as follows:

[0123] (1) Bounding box regression loss

[0124] Bounding box regression loss is used to measure the error between the predicted box and the actual box, enabling the detector to more accurately detect the size of the target. The specific formula is as follows:

[0125]

[0126] Among them, GIoU (Generalized Intersection over Union) loss It is used to predict bounding box regression loss. GIoU loss is a better choice than IoU loss; S 2 N represents the number of image grids in the prediction process and the number of predicted bounding boxes in each grid; These represent the true value, the predicted bounding box, and the contained value, respectively. and The smallest closed frame; coefficient This indicates whether the j-th prediction box in the i-th grid is a positive sample.

[0127] (2) Classification loss

[0128] To measure the difference between the class predicted by the model and the actual label, a binary cross-entropy loss function is used for each label in the network to reduce computational complexity and improve model performance. The specific formula is as follows:

[0129]

[0130] Where p(c) represents the probability that the true sample belongs to class c; This represents the probability that the network predicts a sample to be of class c; the coefficient. The meaning is the same as before. Maintain consistency.

[0131] (3) Confidence loss

[0132] The formula used to measure the difference between the model's confidence in the predicted bounding box and the actual label is as follows:

[0133]

[0134] Where the coefficient With the previous Conversely, c represents whether the j-th predicted box in the i-th grid is a negative sample; i and This represents the confidence level of the true value and the confidence level of the network prediction.

[0135] The network's final total loss function can be defined as follows: The calculation formula is as follows:

[0136]

[0137] S63: Set the initial learning rate for pre-training to 1e-2, the total number of iterations to 200, and the batch size to 4. After training, save the weights after the last training iteration and the weights of the model that performed best on the validation set during training.

[0138] S64: Figure 4 This is a schematic diagram of the detection results of the present invention on the LLVIP dataset, wherein... Figure 4 (a) is the GT (Ground Truth) result. Figure 4 (b) shows the prediction results of the benchmark algorithm YOLOv5. Figure 4 (c) shows the prediction results of this invention. The red boxes in the figure represent missed targets. As can be seen from the figure, in partially occluded scenarios, the baseline algorithm exhibits missed detections, while the detector of the method of this invention successfully detects the targets.

[0139] S65: Figure 5 This is a schematic diagram of the detection results of this invention on the LLVIP dataset. The meaning of each sub-graph is as follows: Figure 4 The diagram maintains consistency, with yellow boxes indicating false positives. As shown in the image, in dense scenes, the baseline algorithm misdetects the "freight car" as a regular "car" and misses the "freight car" in the lower right corner, while the method proposed in this invention correctly detects all objects in the diagram.

[0140] In summary, this invention proposes a target detection method based on attention-based dual-mode feature fusion, which achieves complementary advantages of visible light and infrared modal features, further improving the detection performance of the target detector and has significant implications in some practical scenarios.

[0141] The above provides a detailed description of the target detection method based on attention-based dual-mode feature fusion provided by this invention. Specific examples are used to illustrate the principles and implementation methods of this invention; however, the descriptions of the above embodiments are only intended to help understand the method and core ideas of this invention. It should be noted that those skilled in the art can make several improvements and equivalent substitutions to the technical solutions of this invention without departing from the principles of this invention, and these improvements and substitutions should also fall within the protection scope of the claims of this invention.

Claims

1. A target detection method based on attention-based dual-mode feature fusion, characterized in that, Follow these steps: S1: Obtain the paired dataset of visible light and infrared light; S2: Construct a dual-stream feature extraction backbone network to extract visible light and infrared features, respectively, to obtain visible light convolutional features. and infrared convolution features The dual-stream feature extraction backbone network consists of the following modules: a CBS module, a stacked module composed of a first-layer CBS module + C3 module, a second-layer CBS module + C3 module, and a third-layer CBS module + C3 module, and a fifth-layer SPPF module. S3: The second-layer CBS module + C3 module, the third-layer CBS module + C3 module, and the fifth-layer SPPF module of the dual-stream feature extraction backbone network respectively process visible light convolutional features through the ADFM dual-mode fusion module. and infrared convolution features The fusion process yields three fusion features. ; In S3, the visible light convolution features are processed by the ADFM dual-mode fusion module. and infrared convolution features The integration process specifically includes: For the obtained visible light convolution features A CA coordinate attention module is introduced to enhance the representation of visible light features, as shown in the formula: ; in Vertical direction In the horizontal direction, and Indicates the weights in two spatial directions; The obtained infrared convolution features The features are then input into the SE channel attention module to perform non-linear modeling, resulting in the extracted features. The formula is: ; in Representing the The weight values ​​of each channel, Representing the Characteristics of each channel; After the above operations, enhanced visible light convolutional features were obtained. and enhanced infrared convolution features By adjusting the number of feature channels using 1×1 convolutions, attention feature maps for different modalities in the spatial domain are obtained, as expressed by the following formula: ; ; in This represents the attention feature map of the visible light RGB modes in the spatial domain. This represents the attention feature map of the infrared (IR) mode in the spatial domain. and A 1×1 convolutional block representing RGB and IR modes; Features extracted from different modalities through an attention mechanism and By performing element-wise matrix multiplication with the corresponding spatial domain feature map, the internal spatial information between different modes can be obtained, specifically as follows: ; ; in, This refers to the internal spatial information of the visible light RGB mode. For the internal spatial information of the infrared (IR) mode, Represents element-wise matrix multiplication; The internal spatial information of different modalities is added to the original convolutional features and then input into a 1×1 convolution to obtain the complete features. and The specific formula is as follows: ; ; in and This represents a 1×1 convolutional block; Finally, the preceding features are concatenated and fused using SE attention blocks to obtain the final fused feature. The formula is expressed as: ; in Consistent with the SE channel attention module mentioned earlier; This indicates a splicing operation along the channel axis; S4: Integrate three features Input the neck detection and further achieve multi-scale feature fusion to obtain multi-scale features; S5: Input the multi-scale features obtained from the neck detection into the detection head, and output the detection box position, detection object category and confidence information.

2. The target detection method based on attention-based dual-mode feature fusion according to claim 1, characterized in that, In S2, constructing the dual-stream feature extraction backbone network includes: The dual-stream feature extraction backbone network consists of a first-layer CBS module, a third-layer CBS module + C3 module, and a fifth-layer SPPF module. The input visible light RGB image or infrared IR image is first downsampled by the first-layer CBS module, then further processed by a stacked combination of three CBS modules + C3 modules for feature extraction. Finally, the SPPF module fuses multi-scale features to obtain visible light convolutional features. and infrared convolution features .

3. The target detection method based on attention-based dual-mode feature fusion according to claim 1 or 2, characterized in that, The S5 specifically includes: The multi-scale features obtained in S4 are input into three detection heads, and each detection head outputs the detection box position, the detection object category, and confidence information. The loss value is calculated based on the prediction results and the true label values. The loss functions used are bounding box regression loss and... Classification loss and confidence loss .

4. The target detection method based on attention-based dual-mode feature fusion according to claim 3, characterized in that, The bounding box regression loss Classification loss and confidence loss The details are as follows: (1) The formula for the bounding box regression loss is as follows: ; in, and This indicates the number of image grids and the number of predicted bounding boxes in each grid during the prediction process; , , Representing the true value, the predicted bounding box, and the contained value, respectively. and The smallest closed frame; coefficient Representing the The first grid Is each predicted bounding box a positive sample? (2) The classification loss formula is as follows: ; in, Indicates the real sample is The probability of a class; Indicates that the network prediction sample is Probability of class; coefficient The meaning is the same as before. Maintain consistency with the central government; (3) The confidence loss formula is as follows: ; Where the coefficient With the previous The definition is the opposite, representing the first The first grid Is each predicted bounding box a negative sample? and This represents the confidence level of the true value and the confidence level of the network prediction.

5. The target detection method based on attention-based dual-mode feature fusion according to claim 3 or 4, characterized in that, The total loss function is defined as The calculation formula is as follows: 。 6. The target detection method based on attention-based dual-mode feature fusion according to claim 1 or 2, characterized in that, In step S4, the fused features of different scales obtained in S3 are... Further feature extraction is performed on each layer of the neck detection area; Multi-scale feature fusion was achieved in the detection of the neck by top-down and bottom-up feature extraction and horizontal feature splicing.

7. The target detection method based on attention-based dual-mode feature fusion according to claim 1 or 2, characterized in that, In S1, the acquired visible light and infrared paired dataset is as follows: The LLVIP and DroneVehicle datasets were used, and the datasets were divided into training and test sets according to a set ratio.

Citation Information

Patent Citations

  • Semantic segmentation method for RGB-D bimodal feature fusion

    CN114693929A

  • Feature-guided multi-modal fusion RGB-D saliency target detection based on coordinate attention filtering

    CN116246058A