Object Detection Method, Device and Medium Based on Cross-Attention Multi-Scale Fusion

Through the cross-attention multi-scale fusion detection network, the background perception module and differential and homogeneous attention fusion technology are used to solve the problem of low accuracy in multi-spectral object detection, achieving higher detection accuracy and robustness.

CN119478345BActive Publication Date: 2025-07-22NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411487094.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-23
Publication Date
2025-07-22
Estimated Expiration
2044-10-23

AI Technical Summary

Technical Problem

The existing object detection methods lack effective fusion mechanisms and prior knowledge guidance in multispectral fusion, resulting in low detection accuracy, especially in complex environments, and it is difficult to deal with sensor failures and dynamic perturbations.

Method used

The cross-attention multi-scale fusion detection network is used to extract image features through the backbone network, the background perception module performs lighting and contrast prediction, and the cross-attention fusion module performs differential and homogeneous attention fusion to generate a multimodal fusion feature map, and finally the target position and category are determined by the detection head.

Benefits of technology

It improves the accuracy of object detection, enhances the inherent and discernible characteristics within and between modes, and improves the robustness and detection accuracy in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119478345B_ABST
    Figure CN119478345B_ABST
Patent Text Reader

Abstract

The present application discloses an object detection method, device and medium based on cross-attention multi-scale fusion, which relates to the field of object detection. The method includes: obtaining a visible light image and an infrared image of an object to be detected, and performing object detection using a trained cross-attention multi-scale fusion detection network; wherein, the cross-attention multi-scale fusion detection network includes a backbone network, a background perception module, a cross-attention fusion module and a detection head; the backbone network is used to extract features of different scales; the background perception module is used for illumination and contrast prediction; the cross-attention fusion module is used to perform cross-fusion on the features extracted by the backbone network using differential attention and homogeneous attention according to the illumination and contrast predicted by the background perception module to obtain a multi-modal fusion feature map; the detection head is used to determine the position and category of the object to be detected according to the multi-modal fusion feature map. The present application improves the accuracy of object detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of object detection, and particularly to an object detection method, device and medium based on cross-attention multi-scale fusion. Background Art

[0002] Due to its wide applications in fields such as autonomous driving, search and rescue, and patrol surveillance, object detection has always been a core technology in computer vision tasks. However, limited by the sensing performance and adaptive ability of sensors, single-modal-based work often struggles to handle sensor failures and dynamic disturbances in the environment and lacks robustness.

[0003] To alleviate the limitations of single-modal imaging and achieve all-weather monitoring, more and more researchers have started to focus on multi-spectral object detection. By fusing visual information from different modalities, better detection accuracy can be achieved. However, most existing methods adopt simple fusion mechanisms, unable to learn complementary information between modalities and lacking the guidance of prior knowledge, thus resulting in relatively low accuracy of object detection results. Summary of the Invention

[0004] The purpose of the present application is to provide an object detection method, device and medium based on cross-attention multi-scale fusion, which can improve the accuracy of object detection.

[0005] To achieve the above purpose, the present application provides the following solutions:

[0006] In the first aspect, the present application provides an object detection method based on cross-attention multi-scale fusion, including:

[0007] Obtain a visible light image and an infrared image of the object to be detected;

[0008] According to the visible light image and the infrared image, use the trained cross-attention multi-scale fusion detection network to perform object detection to determine the position and category of the object to be detected;

[0009] Among them, the cross-attention multi-scale fusion detection network includes a backbone network, a background perception module, a cross-attention fusion module and a detection head; the backbone network is used to extract features of different scales of the visible light image and the infrared image; the background perception module is used to predict the illumination and contrast of the visible light image and the infrared image; the cross-attention fusion module is used to perform cross-fusion on the features extracted by the backbone network according to the illumination and contrast predicted by the background perception module, using differential attention and homogeneous attention, to obtain a multi-modal fusion feature map; the detection head is used to determine the position and category of the object to be detected according to the multi-modal fusion feature map.

[0010] In a second aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the above-mentioned object detection method based on cross-attention multi-scale fusion.

[0011] In a third aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the above-mentioned object detection method based on cross-attention multi-scale fusion.

[0012] According to the specific embodiments provided by the present application, the present application has the following technical effects:

[0013] The present application provides an object detection method, device, and medium based on cross-attention multi-scale fusion. An adaptive fusion of visible light and infrared features is achieved by using a cross-attention multi-scale fusion detection network. The background perception module is used to predict illumination and contrast. The feature map extracted by the backbone network is input into the cross-attention fusion module composed of differential attention and homogeneous attention. Homogeneous attention focuses on the shared features within the modality, avoiding the introduction of a large number of redundant features. Differential attention enhances the inter-modal difference features to promote the fusion of complementary information, so as to enhance the inherent and discriminative features within and between modalities. The enhanced features are cross-fused using the illumination and contrast predicted by the background perception module. The position and category of the object to be detected are determined based on the multi-modal fusion feature map, improving the accuracy of object detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0015] Figure 1 Schematic diagram of a visible light image in a low-light scene;

[0016] Figure 2 For Figure 1 the detection result graph;

[0017] Figure 3 Schematic diagram of an infrared image in a low-light scene;

[0018] Figure 4 For Figure 3 the detection result graph;

[0019] Figure 5 Schematic diagram of a visible light image in a normal lighting scene;

[0020] Figure 6 For Figure 5 the detection result diagram;

[0021] Figure 7 is a schematic diagram of an infrared image under normal illumination;

[0022] Figure 8 For Figure 7 the detection result diagram;

[0023] Figure 9 is a schematic flowchart of a target detection method based on cross-attention multi-scale fusion provided by an embodiment of the present application;

[0024] Figure 10 is a schematic structural diagram of a cross-attention multi-scale fusion detection network provided by an embodiment of the present application;

[0025] Figure 11 is a schematic diagram of a background perception module provided by an embodiment of the present application;

[0026] Figure 12 is a schematic diagram of the CSP structure;

[0027] Figure 13 is a schematic structural diagram of differential attention provided by an embodiment of the present application;

[0028] Figure 14 is a schematic structural diagram of homogeneous attention provided by an embodiment of the present application;

[0029] Figure 15 is a schematic diagram of the width distribution and scale change of three datasets;

[0030] Figure 16 is a schematic diagram of the height distribution and scale change of three datasets;

[0031] Figure 17 is a schematic diagram of a detection result of the VEDAI dataset;

[0032] Figure 18 is a schematic diagram of another detection result of the VEDAI dataset;

[0033] Figure 19 is a schematic diagram of the detection result of the FLIR dataset in the daytime scene;

[0034] Figure 20 is a schematic diagram of the detection result of the FLIR dataset in the night scene;

[0035] Figure 21 is a schematic diagram of a detection result of the LLVIP dataset;

[0036] Figure 22 Another schematic diagram of the detection results for the LLVIP dataset. Detailed implementation manners

[0037] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part rather than all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0038] Visible light sensors can provide rich color and texture information, but their images are vulnerable to the influence of isospectral foreign objects and different spectra of the same object, especially under complex conditions such as illumination changes, rain, and fog, where the contrast of some targets is poor. In contrast, infrared sensors are more sensitive to temperature and radiation and perform well under the above-mentioned harsh conditions. However, infrared images have low resolution and unclear edges, making it difficult to capture detailed visual features. As Figures 1 to 4 shown, in low-light scenes, the visible light spectrum has poor hierarchical properties, and it is impossible to distinguish the foreground and background in both the image and the feature space. In contrast, the targets in the infrared image can be clearly distinguished from the background. The features of the targets and the features of the background are also separated into two groups. Figures 5 to 8 The situation is just the opposite. Even though deep learning technology has made great progress, detection techniques using only a single sensor will still encounter difficulties in an increasingly complex environment.

[0039] To alleviate the limitations of single-modal imaging and achieve all-weather monitoring, more and more researchers have begun to focus on multi-spectral object detection. By fusing visual information from different spectra, higher detection accuracy can be achieved. The solutions for multi-modal object detection are currently divided into manual methods and deep learning methods. On the one hand, manual methods are implemented through the traditional "feature extraction + classifier" approach. However, the feature extraction ability of manually designed multi-modal operators is limited, and it is difficult to obtain reliable detection results. On the other hand, due to its strong representation ability, the feature fusion method based on Convolutional Neural Networks (CNN) has been widely used in multi-modal object detection, and most of them are based on two-stream convolutional neural networks. However, most of the fusion strategies involved in two-stream convolutional neural networks use simple element addition, multiplication, and concatenation. Although they have higher performance than single-modal detection, they do not fully consider cross-modal fusion and interaction, and cannot utilize the complementary information between different modalities, which means that the network has poor adaptability when dealing with highly similar objects. Worse still, these simple fusion mechanisms lack long-term dependence, which may exacerbate the imbalance of the network, resulting in unsatisfactory detection results. In addition, considering the differences in visible light and infrared imaging methods and their sensitivity to light, some works have proposed illumination-aware networks to help the network learn the weight ratios of different modalities. The main idea of these methods is to predict the illumination-aware weights of the input images through a predefined gate function, and then the two feature extraction branches are fused by weighted summation according to the ratio to obtain the final detection result. However, it is not enough to use only illumination information as prior knowledge to calculate the modality weights, because this method cannot measure other relevant factors and cannot achieve satisfactory correction for dawn, dusk, strong light, or dim scenes. Therefore, the goal of this application is to improve the accuracy of object detection results by designing an effective fusion mechanism and learning reliable prior knowledge to fully utilize the inherent and complementary information within and between modalities.

[0040] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0041] In an exemplary embodiment, as Figure 9 shown, a method for object detection based on cross-attention multi-scale fusion is provided. This method is executed by a computer device, and specifically can be executed alone by a computer device such as a terminal or a server, or can be jointly executed by a terminal and a server. In the embodiment of the present application, the method includes the following steps 101 to step 102.

[0042] Step 101, obtain a visible light image and an infrared image of the object to be detected.

[0043] Step 102: Based on the visible light image and the infrared image, use the trained cross-attention multi-scale fusion detection network to perform target detection to determine the position and category of the target to be detected.

[0044] Among them, as Figure 10 shown, the cross-attention multi-scale fusion detection network includes a backbone network, a background perception module, a cross-attention fusion module, and a detection head.

[0045] The backbone network is used to extract features of different scales of the visible light image and the infrared image.

[0046] The background perception module is used to perform illumination and contrast prediction on the visible light image and the infrared image. The background perception module uses illumination conditions and target contrast as prior knowledge to guide the fusion of the visible light modality and the infrared modality.

[0047] The cross-attention fusion module is used to perform cross-fusion on the features extracted by the backbone network by using differential attention and homogeneous attention according to the illumination and contrast predicted by the background perception module to obtain a multi-modal fusion feature map. The cross-attention fusion module consists of two parts: differential attention and homogeneous attention. Differential attention enhances the differences between the visible light modality and the infrared modality, and homogeneous attention suppresses a large amount of redundant features within the modality, and then the features of the two parts are weighted and fused to obtain a multi-modal fusion feature map.

[0048] The detection head is used to determine the position and category of the target to be detected according to the multi-modal fusion feature map. The detection head generates detection results through feature weighting.

[0049] This application uses a background-guided cross-attention multi-scale fusion detection network to achieve adaptive fusion of visible light and infrared features. First, use the background perception module to calculate the illumination and contrast weights. Subsequently, input the feature map extracted by the backbone network into the cross-attention fusion module composed of differential attention and homogeneous attention. Homogeneous attention focuses on the shared features within the modality to avoid introducing a large amount of redundant features; differential attention enhances the differential features between modalities to promote the fusion of complementary information to enhance the inherent and discriminative features within and between modalities. Then, sum the enhanced features according to the background perception weights to obtain a merged multi-modal fusion feature map. Finally, input the multi-modal fusion feature map into the detection head to generate detection results.

[0050] In an exemplary embodiment, step 102 includes the following steps 201 to step 204.

[0051] Step 201: Use the background perception module to predict the illumination and contrast of the visible light image and the infrared image, and determine the visible light illumination weight, infrared illumination weight, visible light contrast weight, and infrared contrast weight.

[0052] The factors affecting the reliability of the cross-attention multi-scale fusion detection network are complex and diverse. Although illumination information is effective for cross-modal fusion, there are still the following limitations in using only illumination information:

[0053] 1) Methods based on illumination perception perform better in fixed scenarios such as day and night, but it is difficult to distinguish the target and the surrounding area in dusk, dawn, or backgrounds with high similarity.

[0054] 2) Illumination information is usually used to guide the fusion of features between modalities, but it is insufficient for enhancing features within modalities.

[0055] Therefore, this application uses background perception weighting instead of illumination perception weighting. As Figure 11 shown, the background perception module calculates the illumination and target contrast as prediction weights, thereby more efficiently guiding the adaptive fusion of features within and between modalities. Figure 11 The blue boxes in [[ ]] represent convolutional layers, which are composed of convolutional modules, activation layers, and pooling layers. The connected gray rectangular boxes represent fully connected layers, and the last gray box represents mathematical operations; the orange rectangular box represents a convolutional layer.

[0056] In a specific example, Step 201 includes the following Steps 301 to 303.

[0057] Step 301: Predict the illumination intensity of the visible light image to determine the night illumination intensity and the day illumination intensity.

[0058] The background perception module plays a role in understanding different illumination conditions and background environments during the feature extraction and enhancement stages, guiding the complementary learning and adaptive fusion of the two modalities. Due to the different spectral bands, visible light images are more dependent on external light sources, so there is a large difference from day to night. On the contrary, infrared images belong to passive imaging, and they more reflect the radiation difference between the target and the background. Generally speaking, in the daytime scene, the visible light image has higher imaging quality and contains more color and texture information. In the night scene, the infrared image is clearer and the outline of the target can be obtained.

[0059] Therefore, first design a classification network to predict the illumination information to guide the fusion of information between modalities. Its output result is t d , representing the day illumination intensity. In natural scenes, the illumination conditions during the day are usually better than those at night, which means that the assigned weight is higher. The larger t d , the better the illumination condition.

[0060] Step 302: According to the nighttime illumination intensity and the daytime illumination intensity, use a gate function to determine the visible light illumination weight and the infrared light illumination weight.

[0061] The perception weights contributed by different modality images can be represented by illumination probability values. However, due to the binary nature of the light source, most of the calculated probability values will be close to 0 or 1. If the network directly multiplies these weights with the feature maps of the branches, then the modality with a lower probability value will be greatly suppressed, thus losing the meaning of fusion. To optimize the weights, this application designs a gate function to readjust the weights of the two modalities.

[0062] Specifically, the gate function is:

[0063]

[0064] where W v is the visible light illumination weight, W i is the infrared light illumination weight, W v and W i are the weights for guiding the fusion of the visible light and infrared modalities, W v +W i =1, t d is the daytime illumination intensity, t n is the nighttime illumination intensity, t d +t n =1, during the day t n tends to 0, and at night t n tends to 1.

[0065] When the probability during the day is large, the weight for the visible light modality will be greater than 1 / 2. And since the weight of the infrared modality is no longer close to 0 at this time, it can more fully promote the fusion of the two modalities, and vice versa. This application adjusts the image to a size of 128×128 and inputs it into the background perception module. The visible light image is input into the illumination prediction network, which includes a convolutional layer and a fully connected layer. After the convolutional layer, an activation layer and a 2×2 adaptive pooling are added to compress and extract the illumination features. Subsequently, it is input into the fully connected layer for calculation, and the output of the illumination prediction network is converted into the required weights. Among them, the convolutional layer consists of a convolutional module, an activation layer, and a pooling layer.

[0066] Step 303: According to the gray values of the pixels in the visible light image and the gray values of the pixels in the infrared image, perform contrast prediction on the visible light image and the infrared image respectively to determine the visible light contrast weight and the infrared light contrast weight.

[0067] Using only illumination information to guide modality fusion is insufficient. For example, the similarity of colors and textures between the target and the background in visible light images and the similarity of temperatures between the target and the background in infrared images will both affect the detection accuracy. Therefore, this application introduces contrast as prior knowledge on the basis of the illumination weight to guide the fusion of intra-modal features.

[0068] Generally speaking, the higher the similarity between the target and the background, the more difficult it is to distinguish. Contrast is measured by the gray-scale difference between the target and the surrounding background. Therefore, first, the target regions in the visible light image and the infrared image are determined by sliding the image frame (when the gray-scale difference between a certain image frame and other image frames is large and exceeds the set threshold of 2, this region is the target region), and then the target region is divided into 3×3 grids.

[0069] Define a variable m to represent the pixel mean of the remaining grid regions except the target. The specific calculation process is as follows:

[0070]

[0071] where m j is the gray-scale mean of the pixels in the j-th grid, j represents the serial number of the grid, j = 1 to 8, N j is the total number of pixel points in the j-th grid, k represents the serial number of the pixel point in each grid, is the gray-scale value of the k-th pixel point in the j-th grid.

[0072] Introduce a variable L to represent the maximum gray-scale value of the pixel points in the target region, and use the following formula to describe the contrast difference.

[0073]

[0074] where c j is the contrast of the j-th grid, L j is the maximum gray-scale value of the pixel points in the j-th grid.

[0075] By calculating the contrast, the target and the background can be better distinguished, and the target region in the image can be enhanced and the background region can be suppressed in the subsequent cross-attention fusion module.

[0076] Use the formula C = max(c i ) to determine the contrast weight, that is, obtain the visible light contrast weight C v and the infrared contrast weight C i .

[0077] Step 202, use the backbone network to extract features of different scales from the visible light image and the infrared image respectively, and obtain visible light feature maps of different scales and infrared feature maps of different scales.

[0078] In a specific example, to optimize the feature extraction ability of the network and improve the robustness of the network in complex scenarios, the backbone network uses a dual-branch network composed of Darknet53 and CSP modules to extract features of different scales of the visible light image and the infrared image, where the CSP module is as Figure 12 shown.

[0079] Figure 10 In, the blue rectangular box represents the visible light image branch, and the orange rectangular box represents the infrared image branch. Specifically, the square in the backbone network represents a convolutional module and an improved CSP module. The convolutional module consists of Conv2d (kernel is 3, stride is 2, padding is 1), batch normalization, and activation function. The improved CSP module consists of convolution, segmentation, convolution, and splicing operations. The input image features are initially 640*640*3. After passing through the backbone network and inputting into the cross-fusion module, they become 80*80*256, 40*40*512, and 40*40*512 respectively. Then, they are input into the subsequent network, and the detection results are generated through the iteration of the convolutional network.

[0080] Step 203, using the cross-attention fusion module, according to the visible light illumination weight, the infrared light illumination weight, the visible light contrast weight, and the infrared contrast weight, and using differential attention and homogeneous attention, perform cross-fusion on the visible light feature maps of different scales and the infrared feature maps of different scales to obtain fused features.

[0081] In an exemplary embodiment, as Figure 13 and Figure 14 shown, step 203 includes the following steps 401 to step 406.

[0082] Step 401, for any scale of visible light feature map and infrared feature map, calculate the differential feature and homogeneous feature of the visible light feature map and the infrared feature map.

[0083] Due to the differences in imaging methods, there are both difference information and complementary information in images of different modalities. How to fuse the two modalities of information is the key to multi-spectral object detection. However, most of the existing methods based on dual-branch networks use simple fusion schemes, such as addition, multiplication, and splicing between pixels of different modalities, and cannot fully utilize the inherent features between and within modalities. In addition, rough combination and connection will also increase the difficulty of network learning and fitting, resulting in a decline in detection performance. Inspired by the differential amplifier circuit, where the differential-mode signal and the common-mode signal are amplified and suppressed respectively, the background-aware guided cross-attention fusion module proposed in this application consists of two parts: differential attention and homogeneous attention. Given the visible light feature map M V and the infrared feature map MI , the differential features between classes and the homogeneous features within classes can be expressed as:

[0084]

[0085] where M D is the differential feature, and M C is the homogeneous feature.

[0086] The differential feature between classes can be regarded as the difference between two modalities, and the complementary specific features between modalities are enhanced by subtraction. Conversely, the homogeneous component can be regarded as the sum of two modalities, and the consistency features within modalities are enhanced by addition. Based on this, two new hybrid modalities are defined in this application to output a feature map with richer information for final fusion.

[0087] The inspiration for the differential attention between classes comes from the differential signal in the differential circuit. The differential attention calculates the difference between the visible light and infrared modalities to extract the differential features between modalities. Subsequently, the differential features are enhanced through channel attention.

[0088] Step 402: Use the global maximum average pooling operation and the global maximum pooling operation to globally enhance the differential features to obtain the channel attention map.

[0089] The differential features are encoded into the global vector through the global maximum average pooling and the global maximum pooling to extract the texture information of the target and integrate the global spatial information. The global vector can represent the difference in the channel characteristics between modalities. Specifically, the differential features obtain the first global vector g1 through the global maximum average pooling and the second global vector g2 through the global maximum pooling. Subsequently, through the shared convolution operation, the first feature map s1 and the second feature map s2 are obtained. Then, the first feature map and the second feature map are added together to obtain the channel attention map.

[0090] Step 403: Multiply the channel attention map with the visible light feature map and the infrared feature map respectively, and perform weighted fusion using the visible light illumination weight and the infrared light illumination weight to obtain the differential attention feature map.

[0091] Specifically, multiplying the channel attention map with the input features of the two modalities obtains the feature map of the key region, learns the importance of the visible light and infrared modalities, and performs weighted fusion on the enhanced feature map according to the illumination weight to obtain the differential attention feature map:

[0092]

[0093] where M DA is the differential attention feature map, representing the features after being processed by the inter-class differential module, and W vis the visible light illumination weight, W i is the infrared light illumination weight, M V is the visible light feature map, M I is the infrared feature map, V DA is the channel attention map.

[0094] In this application, through difference, compression, excitation, and weighted fusion, the importance of different modalities across channels is adaptively learned, and its generalization ability is also enhanced. It should be noted that the cross-attention fusion module draws on the residual network to enhance the stability of the network. The feature map after difference is input into a specific modality through a skip connection. Through the residual mapping, the inter-class difference attention can supplement the complementary information between modalities while avoiding directly affecting the features of a single modality.

[0095] Step 404, perform global average pooling and fully connected operations on the homogeneous features in sequence to obtain the visible light modality attention map and the infrared modality attention map

[0096] In the multi-spectral object detection task, the similarity between the foreground and the background will also affect the detection performance. In addition to the complementary features between modalities, the inherent features within a modality are also crucial for the extraction of target discriminative features. Therefore, in this application, under the guidance of the contrast weight, the inherent features within a modality are learned.

[0097] Step 405, multiply the visible light modality attention map by the visible light feature map, multiply the infrared modality attention map by the infrared feature map, and perform weighted fusion using the visible light contrast weight and the infrared contrast weight to obtain the homogeneous attention feature map.

[0098] Specifically, the following formula is used to determine the homogeneous attention feature map:

[0099]

[0100] where, M CA is the homogeneous attention feature map, C v is the visible light contrast weight, C i is the infrared contrast weight, is the visible light modality attention map, is the infrared modality attention map.

[0101] In this application, adaptive channel screening is achieved through summation, weight sharing, compression, and normalization. In steps 404 and 405, the sharing of fully connected layer parameters reduces the dimension of features and improves computational efficiency; through channel screening and the guidance of prior knowledge, the weights of the model are redistributed to the feature channels of visible light and infrared modalities, avoiding the introduction of a large number of redundant features; the use of skip connections enables the network to reuse shallow features and improves the representation ability for complex features.

[0102] Step 406: Fuse the differential attention feature maps of different scales and the homogeneous attention feature maps of different scales to obtain a multi-modal fusion feature map.

[0103] Generally speaking, a two-branch network can process features in a serial or parallel manner. However, a large number of studies have shown that simple serial and parallel methods will exacerbate the imbalance of the network and thus affect the performance of the network. In addition, the change of the target scale, especially some small-sized targets, will also lead to a decline in performance. Therefore, this application designs a multi-scale fusion strategy to achieve the fusion detection of cross-modal visible light and infrared images through the cross of feature maps and the fusion of three different-scale feature maps of small, medium, and large. In the feature extraction stage of the network, feature maps of different modalities and different scales (small, medium, and large) are extracted and fed into the cross-attention fusion module, and then feature weighting is performed to achieve the fusion of multi-scale features. The subsequent fusion steps are shown as follows:

[0104]

[0105] where F FUSE is the multi-modal fusion feature map, P is the number of scales, here P = 3, representing three different scales of small, medium, and large respectively, is the homogeneous attention feature map of the p-th scale, is the differential attention feature map of the p-th scale.

[0106] Specifically, the feature maps of the two modal branch networks are input into the cross-attention fusion module in a horizontally connected manner. To maintain the weight of global features, after each result, a convolutional layer is used to reduce its dimension to 1 / 2 of the original, and then the feature map is bilinearly interpolated to restore to the input length and width. Finally, the feature maps of different scales are concatenated as the global feature of multi-scale fusion. It should be noted that the features of visible light and infrared modalities have all been processed and refined by the cross-attention fusion module to avoid the loss of key information in the process of feature extraction and attention map generation.

[0107] Step 204: According to the multi-modal fusion feature map, use a detection head to determine the position and category of the target to be detected.

[0108] In an exemplary embodiment, a background-aware loss is introduced to enable the cross-attention multi-scale fusion detection network to adaptively capture inter-modal and intra-modal information according to background conditions. It should be noted that, as prior knowledge, contrast does not require an independent network for training, so no supervision is needed and it does not need to be part of the loss function. The background-aware loss includes a detection loss and an illumination condition loss. The detection loss is used to calibrate the multi-spectral detection results, and the illumination condition loss is used to supervise the inter-modal difference weights.

[0109] The accurate definition of the background-aware loss is shown as follows:

[0110] Loss = L d + L l ;

[0111] where Loss is the background-aware loss value, L d is the detection loss value, and L l is the illumination condition loss.

[0112] The detection loss includes a classification loss, a bounding box regression loss, and a confidence loss.

[0113] The classification loss uses the Binary Cross Entropy Loss (BCE). For the targets in the image, through the strong association of states, the labels can better guide the network to learn the class capabilities. The definition of the classification loss is: L cls = -[y a log(p a ) + (1 - y a )log(1 - p a )].

[0114] The bounding box regression loss uses the Generalized Intersection over Union (GIoU). GIoU adds a penalty term on the basis of the original IoU loss to alleviate the gradient problem that occurs when the detection boxes do not overlap. Its definition is:

[0115] The confidence loss draws on Focal Loss, mainly to solve the problem of serious imbalance in the proportions of different targets in the dataset, and is applicable to complex scenarios such as few samples and large target scales. Its definition is: L conf = -α a (1 - p a ) τ log(p a ).

[0116] where a represents the sample number, p a represents the probability predicted by the model, and ya is a binary variable with values of 0 or 1. A represents the area of the ground truth box, B represents the area of the predicted box, D represents the area of the smallest box containing both, and α a mainly addresses the imbalance between positive and negative samples, and τ mainly addresses the imbalance between easy and difficult samples. Here, α a = 0.25 and τ = 2.

[0117] The fusion of inter-modal complementarity and intra-modal inherent information depends to a large extent on the guidance of the background perception module, especially the perception of lighting conditions. The lighting conditions reflect the intensity of the lighting conditions in the image and can be regarded as a classifier that calculates the probabilities of belonging to day and night. Therefore, cross-entropy loss is used to constrain its training process, and its meaning is: L l = -zlogσ(x) - (1 - z)log(1 - σ(x)).

[0118] Among them, z is the lighting condition label of the input image, x represents the probability that the image belongs to day, and σ is the softmax function, which mainly normalizes the lighting condition probability to [0, 1]. To more fully characterize the intensity of the lighting conditions, the application sets the value space of z to 0, 0.5, 1.0, representing the night, low light, and day scenes respectively.

[0119] The object detection method based on cross-attention multi-scale fusion provided by this application can be inherited on multi-spectral detectors. In addition, this application conducts a large number of comparative experiments on three multi-spectral detection datasets (LLVIP, FLIR, and VEDAI), and the results show that the multi-spectral detector provided by this application achieves higher detection accuracy than the current state-of-the-art multi-spectral detectors.

[0120] First, the datasets and evaluation metrics are introduced below; second, the experimental settings and algorithm deployment are introduced; subsequently, the object detection method based on cross-attention multi-scale fusion provided by this application is compared with other advanced methods in terms of mAP and mAP 50 and other metrics in a large number of comparative experiments on 3 datasets; finally, ablation experiments are conducted for research.

[0121] (1) Datasets.

[0122] This application compares the publicly available dual-light fusion detection datasets and focuses on analyzing 3 datasets with remote sensing, road, and low illumination as the background, FLIR, LLVIP, and VEDAI. As Figure 15 (a) part of shows the width distribution and scale change of the VEDAI dataset, and as Figure 15 (b) part of shows the width distribution and scale change of the FLIR dataset, and as Figure 15Part (c) shows the distribution and scale variation of the width of the LLVIP dataset, as Figure 16 Part (a) shows the distribution and scale variation of the height of the VEDAI dataset, as Figure 16 Part (b) shows the distribution and scale variation of the height of the FLIR dataset, as Figure 16 Part (c) shows the distribution and scale variation of the height of the LLVIP dataset, indicating that the datasets used in this application cover instances of different scales and have good representativeness.

[0123] The VEDAI dataset is mainly used for optical remote sensing image target detection. The dataset was taken in various complex backgrounds such as cities, roads, fields, and forests, covering 9 different types of small and medium-sized vehicles, with a total of 1246 pairs of images and 3640 instances. This application is carried out on images with a resolution of 512 * 512. Its characteristic is that the target sizes vary greatly, which can verify the algorithm's fusion detection ability for targets of different scales.

[0124] FLIR mainly has 3 types of targets: pedestrians, bicycles, and cars, and its background is mainly on typical streets and roads. The FLIR aligned dataset manually deletes unaligned images based on the original dataset, avoiding the problem of network training not converging. This application is all carried out on the FLIR aligned dataset. For convenience, FLIR refers to the aligned version in the following. The FLIR aligned dataset contains 5142 pairs of bimodal images and is divided into training, validation, and test datasets according to a ratio of 6:2:2. Its sampling frame rate is mostly 2 frames per second, and it is mainly used for the autonomous driving scenario.

[0125] The LLVIP dataset aims to solve the problem of multi-spectral pedestrian detection in low-light scenarios. The ratio of night and day scenes in the dataset reaches 12:1. The dataset contains 15488 image pairs, of which 9292 pairs of images are used for training, and 3098 pairs of images are used for validation and testing respectively. Compared with other datasets, the visible light and infrared modality images in the LLVIP dataset are strictly aligned in time and space, which can enable the network to focus more on improving the fusion detection performance without considering the errors caused by calibration.

[0126] (2) Experimental results.

[0127] To verify the effectiveness of the cross-attention multi-scale fusion detection network provided by this application, it was compared with baseline detectors and state-of-the-art multi-spectral detection networks, including methods such as GAFF and CFT. The baseline detectors include the one-stage detection algorithms YOLOv9 and Faster R-CNN and the two-stream network that only uses pixel addition for fusion.

[0128] 1) VEDAI Dataset: In the remote sensing scenario, the method of this application was compared with other related works, and the experimental results are shown in Table 1. It can be observed that on the VEDAI dataset, the method of this application achieved the best detection performance compared with other algorithms. For single-modal detection, the method of this application improved the mAP by 15.2% and 18.2% respectively compared with the one-stage YOLOv9 and Faster R-CNN. For multi-modal detection, the method of this application was 1.3% and 0.7% higher than the best method respectively in terms of mAP 50 and mAP metrics.

[0129] Table 1 Comparison of Experimental Results on VEDAI Dataset

[0130]

[0131] In addition to the quantitative comparison, this application also conducted a qualitative analysis on the VEDAI dataset. Figure 17 and Figure 18 show the original input images, the detection results of the baseline and this application. The ground truth is labeled in the input infrared image. As Figure 17 (a) part of Figure 18 shows the original input visible light image, as Figure 17 (b) part of Figure 18 shows the original input infrared image, as Figure 17 (c) part of Figure 18 shows the schematic diagram of the detection results of the baseline, as Figure 17 (d) part of Figure 18 and (d) part of

[0132] show the schematic diagram of the detection results of this application. In the remote sensing scenario, the target size is small, the arrangement is dense, and the texture feature is weak. Using only a simple two-stream network is likely to cause misdetection or missed detection (the area where the triangle is located). However, the method of this application fully integrates the complementary and inherent features of the two modalities, which can significantly reduce the occurrence of misdetection and missed detection phenomena. Therefore, this application has better detection performance in the remote sensing scenario. 50 index, this application is 2.2% higher than the typical two-stream network LRAF-Net. In addition, compared with the single-modal detection algorithm, this application is 8.6% and 4.0% higher respectively in terms of mAP 50 and mAP metrics. This shows that the fusion algorithm of this application is effective and greatly improves the performance of single-modal detection.

[0133] Comparison of experimental results on the FLIR dataset

[0134]

[0135] Similarly, the present application conducts a qualitative analysis on the FLIR dataset and selects two scenarios, day and night. The detection results are as follows Figure 19 and Figure 20 shown. The ground truth of the target is marked by the white rectangular box in the infrared image. In the road scene, the targets are occluded from each other, resulting in a decline in the algorithm performance. As shown in the (a) part of Figure 19 is the visible light image input for the day scene. As shown in the (b) part of Figure 19 is the infrared image input for the day scene. As shown in the (c) part of Figure 19 is the schematic diagram of the detection result of the baseline for the day scene. As shown in the (d) part of Figure 19 is the schematic diagram of the detection result of the present application for the day scene. As shown in the (a) part of Figure 20 is the visible light image input for the night scene. As shown in the (b) part of Figure 20 is the infrared image input for the night scene. As shown in the (c) part of Figure 20 is the schematic diagram of the detection result of the baseline for the night scene. As shown in the (d) part of Figure 20 is the schematic diagram of the detection result of the present application for the night scene. In the detection results of the baseline, the pedestrians overlapping with the utility pole in the day scene are not detected, and the vehicle occluded by pedestrians in the night scene is detected twice. While in the detection results of the present application, the above targets are all detected. The reason is that the method of the present application takes into account the global and local information of the target, effectively enhancing the features of the target under occlusion.

[0136] 3) LLVIP dataset: In the low-light scene, the comparison results between the present application and other algorithms are shown in Table 3. It can be seen that compared with other algorithms, the present application has the state-of-the-art performance on the LLVIP dataset, with mAP 50 reaching 97.9% and mAP reaching 69.2%. Especially for the mAP index, compared with YOLOv9 and LRAF-Net, it is 7.3% and 2.9% higher respectively. In addition, from the detection results of single modality, the mAP of the infrared image is significantly higher than that of the visible light image, indicating that the infrared image has better feature representation ability in the low-light situation.

[0137] Table 3 Comparison of experimental results on the LLVIP dataset

[0138]

[0139]

[0140] Figure 21 and Figure 22 shows some experimental results of the baseline and this application. It can be seen that in low-light scenarios, the infrared images have better detection performance and can intuitively show the difference between the foreground and the background. As Figure 21 part (a) of Figure 22 and part (a) of Figure 21 show the original input visible light images, and as Figure 22 part (b) of Figure 21 and part (b) of Figure 22 show the original input infrared images. As Figure 21 part (c) of Figure 22 and part (c) of Figure 21 show the schematic diagrams of the detection results of the baseline. As Figure 22 part (d) of

[0141] show the schematic diagrams of the detection results of this application. When the targets overlap or are occluded, especially in

[0142] In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0143] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0144] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0145] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include Read-Only Memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0146] The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0147] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0148] In this article, specific examples are used to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A target detection method based on cross-attention multi-scale fusion, characterized in that, The object detection method based on cross-attention multi-scale fusion includes: Obtain the visible light image and the infrared image of the object to be detected; According to the visible light image and the infrared image, use the trained cross-attention multi-scale fusion detection network to perform object detection to determine the position and category of the object to be detected; Among them, the cross-attention multi-scale fusion detection network includes a backbone network, a background perception module, a cross-attention fusion module, and a detection head; the backbone network is used to extract features of different scales of the visible light image and the infrared image; the background perception module is used to perform illumination and contrast prediction on the visible light image and the infrared image; the cross-attention fusion module is used to perform cross-fusion on the features extracted by the backbone network using differential attention and homogeneous attention according to the illumination and contrast predicted by the background perception module to obtain a multi-modal fusion feature map; the detection head is used to determine the position and category of the object to be detected according to the multi-modal fusion feature map; According to the visible light image and the infrared image, use the trained cross-attention multi-scale fusion detection network to perform object detection to determine the position and category of the object to be detected, specifically including: Use the background perception module to perform illumination and contrast prediction on the visible light image and the infrared image to determine the visible light illumination weight, the infrared light illumination weight, the visible light contrast weight, and the infrared contrast weight; Use the backbone network to extract features of different scales of the visible light image and the infrared image respectively to obtain visible light feature maps of different scales and infrared feature maps of different scales; Use the cross-attention fusion module to perform cross-fusion on the visible light feature maps of different scales and the infrared feature maps of different scales using differential attention and homogeneous attention according to the visible light illumination weight, the infrared light illumination weight, the visible light contrast weight, and the infrared contrast weight to obtain a multi-modal fusion feature map, specifically including: For any scale of visible light feature map and infrared feature map, calculate the differential feature and homogeneous feature between the visible light feature map and the infrared feature map; Use the global maximum average pooling operation and the global maximum pooling operation to globally enhance the differential feature to obtain a channel attention map; Multiply the channel attention map with the visible light feature map and the infrared feature map respectively, and perform weighted fusion using the visible light illumination weight and the infrared illumination weight to obtain a differential attention feature map: where M DA is the differential attention feature map, W v is the visible light illumination weight, W i is the infrared illumination weight, M V is the visible light feature map, M I is the infrared feature map, V DA is the channel attention map; Perform global average pooling and fully connected operations on the homogeneous feature in sequence to obtain a visible light modality attention map and an infrared modality attention map; Multiply the visible light modality attention map with the visible light feature map, multiply the infrared modality attention map with the infrared feature map, and perform weighted fusion using the visible light contrast weight and the infrared contrast weight to obtain a homogeneous attention feature map: where M CA is the homogeneous attention feature map, C v is the visible light contrast weight, C i is the infrared contrast weight, is the visible light modality attention map, is the infrared modality attention map; Fuse the differential attention feature maps of different scales and the homogeneous attention feature maps of different scales to obtain a multi-modal fusion feature map; According to the multi-modal fusion feature map, use the detection head to determine the position and category of the object to be detected.

2. The object detection method based on cross-attention multi-scale fusion according to claim 1, characterized in that, Use the background perception module to perform illumination and contrast prediction on the visible light image and the infrared image to determine the visible light illumination weight, the infrared light illumination weight, the visible light contrast weight, and the infrared contrast weight, specifically including: Perform illumination intensity prediction on the visible light image to determine the night illumination intensity and the day illumination intensity; According to the night illumination intensity and the day illumination intensity, use the gate function to determine the visible light illumination weight and the infrared light illumination weight; According to the gray values of each pixel in the visible light image and the gray values of each pixel in the infrared image, the contrast of the visible light image and the infrared image is predicted respectively to determine the visible light contrast weight and the infrared contrast weight.

3. The object detection method based on cross-attention multi-scale fusion according to claim 2, characterized in that The gate function is: Among them, W v is the visible light illumination weight, and W i is the infrared light illumination weight, t d is the daytime illumination intensity, and t n is the nighttime illumination intensity.

4. The object detection method based on cross-attention multi-scale fusion according to claim 1, characterized in that, The backbone network uses a dual-branch network composed of Darknet53 and CSP structure to extract features of different scales of the visible light image and the infrared image.

5. The object detection method based on cross-attention multi-scale fusion according to claim 1, characterized in that The loss function during the training of the cross-attention multi-scale fusion detection network includes detection loss and illumination condition loss; the detection loss includes classification loss, bounding box regression loss, and confidence loss.

6. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the cross-attention multi-scale fusion-based object detection method according to any one of claims 1-5.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the cross-attention multi-scale fusion-based object detection method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Target detection method and system using illumination guidance and attention mechanism

    CN115131640A

  • Iris recognition control opening and closing method based on intelligent lock

    CN118644920A