Multi-source target detection method based on multi-source feature cross fusion and decomposition combination

By designing a detection model based on multi-source feature cross-fusion and decomposition combination, the problems of insufficient feature complementary fusion and modal imbalance in RGB-infrared multi-source object detection are solved, and high-precision and time-efficient target detection are achieved, which is suitable for multi-source object detection from the perspective of drones.

CN120298706APending Publication Date: 2025-07-11NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510293385.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing RGB-infrared multi-source object detection methods have problems such as insufficient complementary fusion of multi-source features, modal imbalance, and difficulty in balancing detection accuracy and speed in complex backgrounds and high-time application scenarios.

Method used

A detection model based on cross-fusion and decomposition combination of multi-source features is designed, and a dual-branch backbone network is used for feature extraction, decomposition and combination and cross-fusion network are used for feature decomposition and recombination, a dual-branch neck network is used for feature aggregation, and a target detection is carried out through a dual-branch detection network to achieve complementary fusion and independent detection of multi-source features.

Benefits of technology

It improves the accuracy and speed of multi-source target detection, and can achieve high-precision and time-efficient target detection in complex backgrounds, and is suitable for deployment of edge equipment such as drones.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298706A_ABST
    Figure CN120298706A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence and deep learning, in particular to a multi-source target detection method based on multi-source feature cross fusion and decomposition combination, and the method comprises the steps: carrying out the down-sampling and multi-scale feature extraction of an RGB-infrared image through a double-branch backbone network, and obtaining RGB features and infrared features of different scales; performing decomposition and recombination on the large-scale RGB features and the infrared features and the medium-scale RGB features and the infrared features through a decomposition combination and cross fusion network, and performing multi-source cross attention fusion on the small-scale RGB features and the infrared features to obtain fusion features of different scales; performing multi-path feature aggregation on the fusion features of different scales through a double-branch neck network to obtain aggregation features of different scales; and target detection is carried out on the aggregation features of different scales through a double-branch detection network, and an RGB detection result and an infrared detection result are obtained. According to the method, the real-time performance of detection is ensured while high-precision detection is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the technical fields of artificial intelligence and deep learning, and particularly to a multi-source object detection method based on multi-source feature cross-fusion and decomposition combination. Background Art

[0002] As one of the most basic tasks in computer vision, object detection aims to classify and locate specific objects in images. Thanks to the support of deep learning technology, a large number of object detection algorithms based on deep learning have been widely applied in fields such as intelligent transportation, autonomous driving, industrial inspection, and emergency rescue. Since RGB cameras are sensitive to light and weather, their performance is limited under conditions such as fog, object occlusion, camouflage, and night. Infrared cameras, on the other hand, image by capturing the thermal radiation information of objects, and their working wavelength can penetrate the night and smoke, being almost unaffected by changes in lighting and weather conditions, and can well complement the deficiencies of RGB cameras. Therefore, more and more research teams have begun to focus on the research of RGB-infrared multi-source collaborative object detection, improving the accuracy and robustness of object detection through the complementary fusion of two-modal features.

[0003] In RGB-infrared multi-source object detection methods, in order to fully achieve the complementary fusion of multi-source features, most of the currently proposed methods use means such as channel splicing, feature weighting, channel attention mechanism, and spatial attention mechanism for multi-source feature fusion. For example, Zhou et al. designed a differential modality perception fusion module to promote the in-depth fusion of complementary information between modalities through differential feature extraction and channel-level feature fusion. Althoupety et al. proposed a DAFF (Dual Attentive Feature Fusion) method for multi-spectral pedestrian detection, achieving complementary fusion of multi-source features by simultaneously using the global attention mechanism of spatial position and feature channels on RGB and infrared feature maps. Shen et al. proposed a dual cross-attention Transformer feature fusion framework, realizing cross-modal feature complementary fusion while performing global feature modeling. The introduction of the cross-attention mechanism enhances the discriminability of object features and improves object detection performance.

[0004] However, the currently proposed multi-source object detection methods still have the following deficiencies when facing complex backgrounds such as the drone perspective and high-timeliness application scenarios.

[0005] First, the currently proposed multi-source feature complementary fusion methods are relatively single, without fully considering the differences between the RGB and infrared modalities, as well as the differences between multi-source high-level features and low-level features, and the complementary fusion is not sufficient.

[0006] Second, due to the differences in the imaging mechanisms of RGB and infrared cameras and the differences in the shooting shutter times, weak registration phenomena are likely to occur in RGB and infrared images, which requires the multi-source target detection algorithm to be able to adapt to the problem of modal imbalance.

[0007] Third, in UAV application scenarios such as traffic monitoring and emergency rescue, high-precision and high-timeliness target detection is required, which requires the RGB-infrared multi-source target detection algorithm to achieve a balance between detection accuracy and detection speed. Summary of the Invention

[0008] In view of this, the embodiments of the present application propose a multi-source target detection method based on multi-source feature cross-fusion and decomposition combination, design a lightweight multi-source cross-attention fusion module and a feature decomposition combination module, starting from the high-level semantic cross-attention mechanism and the difference in the basic image frequency features, so as to improve the multi-modal feature complementary fusion ability of the detection model. At the same time, a dual-branch network is also designed to independently predict RGB-modal targets and infrared-modal targets to solve the target detection problem in the modal imbalance scenario, ensuring the real-time performance of detection while achieving high-precision detection.

[0009] To achieve the above object, the embodiments of the present application propose a multi-source target detection method based on multi-source feature cross-fusion and decomposition combination, which is implemented based on a pre-trained detection model composed of a dual-branch backbone network, a decomposition combination and cross-fusion network, a dual-branch neck network, and a dual-branch detection network. The method includes the following steps: performing downsampling and multi-scale feature extraction on the RGB-infrared image pair through the dual-branch backbone network to obtain RGB features and infrared features of different scales; through the decomposition combination and cross-fusion network, decomposing and recombining the large-scale RGB features and infrared features, and the medium-scale RGB features and infrared features respectively, and performing multi-source cross-attention fusion on the small-scale RGB features and infrared features to obtain fusion features of different scales; performing multi-path feature aggregation of RGB modality and infrared modality on the fusion features of different scales through the dual-branch neck network to obtain aggregated features of different scales; performing target detection of RGB modality and infrared modality on the aggregated features of different scales through the dual-branch detection network, and after performing non-maximum suppression processing, obtaining RGB detection results and infrared detection results.

[0010] To achieve the above object, the embodiments of the present application also propose an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a multi-source target detection method based on multi-source feature cross-fusion and decomposition combination as described above.

[0011] To achieve the above object, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, which, when executed by a processor, can implement a multi-source object detection method based on multi-source feature cross-fusion and decomposition combination as described above.

[0012] A multi-source object detection method based on multi-source feature cross-fusion and decomposition combination proposed in the present application designs, constructs, trains, and uses a detection model composed of a dual-branch backbone network, a decomposition combination and cross-fusion network, a dual-branch neck network, and a dual-branch detection network to achieve multi-source object detection. The dual-branch backbone network is used to downsample and extract multi-scale features from the RGB-infrared image pair, obtaining RGB features and infrared features at different scales, thereby expanding the number of features. The decomposition combination and cross-fusion network is used to decompose and recombine the large-scale RGB features and infrared features, and the medium-scale RGB features and infrared features respectively, and perform multi-source cross-attention fusion on the small-scale RGB features and infrared features, obtaining fusion features at different scales. The design of decomposition and recombination, as well as multi-source cross-attention fusion, realizes the complementary fusion of multi-source features, effectively improving the utilization degree of the detection model for multi-modal features. The dual-branch neck network is used to perform multi-path feature aggregation in the RGB modality and the infrared modality on the fusion features at different scales, obtaining aggregated features at different scales. The dual-branch detection network is then used to perform object detection in the RGB modality and the infrared modality on the aggregated features at different scales, and after performing non-maximum suppression processing, obtaining RGB detection results and infrared detection results. Using the dual-branch neck network and the dual-branch detection network to independently aggregate and detect objects in two modalities, while achieving complementary feature fusion, retains a certain degree of feature independence, thus being able to well solve the problem of object detection in weakly registered and other modality imbalance scenarios. The entire detection model adopts a lightweight basic network structure and is modularly integrated into one, and such a detection model can achieve multi-source object detection with high precision and high timeliness.

[0013] Optionally, the dual-branch backbone network is composed of an RGB feature extraction branch and an infrared feature extraction branch;

[0014] The RGB feature extraction branch is composed of five sequentially connected visible light convolution modules. The input of the first visible light convolution module is the RGB image in the RGB-infrared image pair, and the input of the subsequent visible light convolution module is the output of the previous visible light convolution module. All five visible light convolution modules are used to downsample and extract features from their own inputs, obtaining RGB features with resolutions of 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the RGB image respectively. The five RGB features are sequentially denoted as and

[0015] The infrared feature extraction branch consists of five sequentially connected thermal infrared convolutional modules. The input of the first thermal infrared convolutional module is the infrared image in the RGB-infrared image pair, and the input of the subsequent thermal infrared convolutional module is the output of the previous thermal infrared convolutional module. All five thermal infrared convolutional modules are used to downsample and extract features from their own inputs, obtaining infrared features with resolutions of 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the infrared image respectively. The five infrared features are sequentially denoted as and

[0016] Among them, both the visible light convolutional module and the thermal infrared convolutional module are composed of a basic convolutional unit and a cross-stage local convolutional unit connected in series. The basic convolutional unit is composed of a 3×3 convolutional layer, a batch normalization layer, and a SiLU activation function layer connected in series. The cross-stage local convolutional unit first uses a basic convolutional unit to increase the dimension of the input, and then splits it into two feature maps F p1 and F p2 along the channel. F p1 is used for residual connection, and F p2 then sequentially extracts features through multiple basic convolutional units. Finally, the two parts of the feature maps are concatenated along the channel and downsampled using a basic convolutional unit for output.

[0017] Optionally, the decomposition combination and cross-fusion network consists of two multi-source feature decomposition and combination modules, a multi-source cross-attention fusion module, and a spatial pyramid pooling module. The two multi-source feature decomposition and combination modules are the first multi-source feature decomposition and combination module and the second multi-source feature decomposition and combination module respectively;

[0018] As large-scale RGB features and infrared features, they are input into the first multi-source feature decomposition and combination module of the decomposition combination and cross-fusion network for the decomposition and recombination of basic features and detailed features, generating the first RGB combined feature and the first infrared combined feature

[0019] As medium-scale RGB features and infrared features, they are input into the second multi-source feature decomposition and combination module of the decomposition combination and cross-fusion network for the decomposition and recombination of basic features and detailed features, generating the second RGB combined feature and the second infrared combined feature

[0020] As small-scale RGB features and infrared features, they are input into the multi-source cross-attention fusion module of the decomposition-combination and cross-fusion network. The multi-source cross-attention fusion module uses the cross-attention mechanism to guide the cross-modal feature fusion between high-level semantic features and generates the first fusion feature.

[0021] They are input into the spatial pyramid pooling module of the decomposition-combination and cross-fusion network. The spatial pyramid pooling module performs spatial pyramid pooling processing and generates the second fusion feature.

[0022] Optionally, the multi-source feature decomposition-combination module consists of two basic feature extractors, two detail feature extractors, and two output units. The two basic feature extractors are the RGB basic feature extractor and the infrared basic feature extractor respectively. The two detail feature extractors are the RGB detail feature extractor and the infrared detail feature extractor respectively. The two output units are the RGB output unit and the infrared output unit respectively.

[0023] After feature extraction by the RGB basic feature extractor and the RGB detail feature extractor respectively, the RGB basic feature and the RGB detail feature After feature extraction by the infrared basic feature extractor and the infrared detail feature extractor respectively, the infrared basic feature and the infrared detail feature

[0024] The RGB output unit of the first multi-source feature decomposition-combination module will and perform an addition operation to obtain

[0025] The infrared output unit of the first multi-source feature decomposition-combination module will and perform an addition operation to obtain

[0026] The RGB output unit of the second multi-source feature decomposition-combination module will and perform an addition operation to obtain

[0027] The infrared output unit of the second multi-source feature decomposition-combination module will and perform an addition operation to obtain

[0028] Optionally, the basic feature extractor consists of a first LN normalization layer, a multi-head transposed attention module, a second LN normalization layer, and a depth convolutional gated feed-forward network connected in sequence. Let the input feature of the basic feature extractor be , First, it passes through the first LN normalization layer and the multi-head transposed attention module for local and global cross-channel feature enhancement, and after residual connection, it obtains , Then, it passes through the second LN normalization layer and the depth convolutional gated feed-forward network for gated feature enhancement, and after residual connection, it obtains the basic feature ; The calculation process of is expressed by the formula: ; Among them, represents the LN normalization layer, represents the multi-head transposed attention module, represents the depth convolutional gated feed-forward network; The detailed feature extractor adopts an affine coupling reversible neural network structure and consists of a local branch and a long-range branch. Let the input feature of the detailed feature extractor be , First, it is split into a local feature and a long-range feature in the channel dimension to serve as two reversible nodes. In the long-range branch, is mapped through the mapping function , and then added to to obtain the long-range node feature . In the local branch, are respectively mapped through the mapping function and the mapping function , and then perform an enhanced affine transformation with to obtain the local node feature . Finally, and are concatenated in the channel dimension to obtain the detailed feature ; The calculation process of is expressed by the formula: ; ; Among them, represents channel concatenation, Denotes the Hadamard product.

[0039] Optionally, the multi-source cross-attention fusion module includes two stages: a cross-modal feature interaction stage and a multi-source feature enhancement stage. The cross-modal feature interaction stage includes an RGB interaction branch and an infrared interaction branch, and each branch contains a multi-head attention module;

[0040] In the RGB interaction branch, the multi-head attention module takes as the query vector, takes as the key vector and the value of the key vector, calculates the cross-attention between modalities, and then performs a residual connection with to obtain the RGB interaction feature

[0041] In the infrared interaction branch, the multi-head attention module takes as the query vector, takes as the key vector and the value of the key vector, calculates the cross-attention between modalities, and then performs a residual connection with to obtain the infrared interaction feature

[0042] And The calculation process is expressed by the formula:

[0043]

[0044] where MHA(·) represents the multi-head attention module;

[0045] In the multi-source feature enhancement stage, first and are concatenated along the channel dimension, and after being processed by the LN normalization layer, the coarse fusion feature F att is obtained. Subsequently, feature enhancement is performed through the multi-head attention module and the feed-forward neural network layer, and a residual connection is established. Finally,

[0046] The calculation process is expressed by the formula:

[0047]

[0048] Among them, Concat(·) represents channel concatenation, LN(·) represents the LN normalization layer, MHA(·) represents the multi-head attention module, and FN(·) represents the feed-forward neural network layer.

[0049] Optionally, the dual-branch neck network consists of an RGB neck branch and an infrared neck branch, and the network structures of both the RGB neck branch and the infrared neck branch are top-down and bottom-up multi-path aggregation networks;

[0050] In the RGB neck branch, the top-down path consists of a common PAN module, a first visible light PAN module, and a second visible light PAN module connected in sequence, and the bottom-up path consists of a third visible light PAN module and a fourth visible light PAN module connected in sequence. The output of the common PAN module is also used as the input of the fourth visible light PAN module, and the output of the first visible light PAN module is also used as the input of the third visible light PAN module. and are respectively input into the common PAN module, the first visible light PAN module, and the second visible light PAN module, and the second visible light PAN module, the third visible light PAN module, and the fourth visible light PAN module respectively output three different scales of RGB aggregation features and

[0051] In the infrared neck branch, the top-down path consists of a common PAN module, a first thermal infrared PAN module, and a second thermal infrared PAN module connected in sequence, and the bottom-up path consists of a third thermal infrared PAN module and a fourth thermal infrared PAN module connected in sequence. The output of the common PAN module is also used as the input of the fourth thermal infrared PAN module, and the output of the first thermal infrared PAN module is also used as the input of the third thermal infrared PAN module. and are respectively input into the common PAN module, the first thermal infrared PAN module, and the second thermal infrared PAN module, and the second thermal infrared PAN module, the third thermal infrared PAN module, and the fourth thermal infrared PAN module respectively output three different scales of infrared aggregation features and

[0052] Optionally, when training the detection model, the total loss function used includes the classification loss the regression loss the distribution focusing loss and the multi-source feature decomposition loss is expressed by the formula as:

[0053]

[0054] h ∈ {rgb, ir};

[0055] where λ cls , λ box , λ dfl , λ decomp are the weights of respectively. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] To more clearly illustrate the embodiments of the present application or the technical solutions in the related art, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or the related art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0057] Figure 1 is a flowchart of a multi-source object detection method based on multi-source feature cross-fusion and decomposition combination provided in an embodiment of the present application;

[0058] Figure 2 is a schematic structural diagram of a detection model provided in an embodiment of the present application;

[0059] Figure 3 is a schematic structural diagram of a multi-source feature decomposition and combination module provided in an embodiment of the present application;

[0060] Figure 4 is a schematic structural diagram of a basic feature extractor provided in an embodiment of the present application;

[0061] Figure 5 is a schematic structural diagram of a detail feature extractor provided in an embodiment of the present application;

[0062] Figure 6 is a schematic structural diagram of a multi-source cross-attention fusion module provided in an embodiment of the present application;

[0063] Figure 7 is a graph of the comparative experiment results on the DroneVehicle dataset provided in an embodiment of the present application;

[0064] Figure 8 is a schematic structural diagram of an electronic device provided in another embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0065] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the embodiments of this application will be described in detail below with reference to the accompanying drawings. However, those of ordinary skill in the art can understand that in the embodiments of this application, many technical details are provided to help readers better understand this application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in this application can still be implemented. The division of the following embodiments is for convenience of description and should not constitute any limitation on the specific implementation of this application. The various embodiments can be combined and cross-referenced with each other on the premise of no contradiction.

[0066] An embodiment of this application proposes a multi-source object detection method based on multi-source feature cross-fusion and decomposition combination, which is implemented based on a pre-trained detection model composed of a dual-branch backbone network, a decomposition combination and cross-fusion network, a dual-branch neck network, and a dual-branch detection network. The implementation details of a multi-source object detection method based on multi-source feature cross-fusion and decomposition combination proposed in this embodiment will be specifically described below. The following content is only implementation details provided for convenience of understanding and is not necessary for implementing this solution.

[0067] The specific process of a multi-source object detection method based on multi-source feature cross-fusion and decomposition combination proposed in this embodiment can be as Figure 1 shown and includes:

[0068] Step 101, perform downsampling and multi-scale feature extraction on the RGB-infrared image pair through the dual-branch backbone network to obtain RGB features and infrared features of different scales.

[0069] In a specific implementation, the specific structure of the detection model is as Figure 2 shown. The RGB-infrared image pair includes an RGB image and an infrared image with the same content. The RGB image and the infrared image are fed into two branches of the dual-branch backbone network for downsampling and multi-scale feature extraction respectively, so as to obtain RGB features and infrared features of different scales.

[0070] In an example, the specific structure of the dual-branch backbone network is as Figure 2 shown. The dual-branch backbone network consists of an RGB feature extraction branch and an infrared feature extraction branch.

[0071] The RGB feature extraction branch consists of five sequentially connected visible light convolution modules ( Figure 2 in to ). The first visible light convolution module ( Figure 2 in ) The input of is the RGB image in the obtained RGB-infrared image pair, and the input of the subsequent visible light convolution module is the output of the previous visible light convolution module. All five visible light convolution modules are used to downsample and extract features from their own inputs, so as to obtain RGB features with resolutions of 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the RGB image respectively. The five RGB features are sequentially denoted as and

[0072] Correspondingly, the infrared feature extraction branch consists of five sequentially connected thermal infrared convolution modules ( Figure 2 in to ). The input of the first thermal infrared convolution module ( Figure 2 in ) is the infrared image in the obtained RGB-infrared image pair, and the input of the subsequent thermal infrared convolution module is the output of the previous thermal infrared convolution module. All five thermal infrared convolution modules are used to downsample and extract features from their own inputs, so as to obtain infrared features with resolutions of 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the infrared image respectively. The five infrared features are sequentially denoted as and

[0073] Among them, both the visible light convolution module and the thermal infrared convolution module are composed of a basic convolution unit and a cross-stage local convolution unit connected in series. The basic convolution unit is composed of a 3×3 convolution layer, a batch normalization layer, and a SiLU activation function layer connected in series. The cross-stage local convolution unit first uses a basic convolution unit to increase the dimension of the input, and then splits it into two feature maps F p1 and F p2 by channels. F p1 is used for residual connection, and F p2 then sequentially extracts features through multiple basic convolution units. Finally, the two parts of the feature maps are concatenated by channels and reduced in dimension using a basic convolution unit for output.

[0074] Step 102, through the decomposition combination and cross-fusion network, decompose and recombine the large-scale RGB features and infrared features, and the medium-scale RGB features and infrared features respectively, and perform multi-source cross-attention fusion on the small-scale RGB features and infrared features to obtain fusion features of different scales.

[0075] In a specific implementation, the dual-branch backbone network feeds the extracted RGB features and infrared features at different scales into the decomposition, combination, and cross-fusion network. The decomposition, combination, and cross-fusion network then decomposes and recombines the large-scale RGB features and infrared features, as well as the medium-scale RGB features and infrared features, respectively, and performs multi-source cross-attention fusion on the small-scale RGB features and infrared features, finally obtaining the fused features at different scales.

[0076] In one example, the specific structure of the decomposition, combination, and cross-fusion network is as Figure 2 shown. The decomposition, combination, and cross-fusion network consists of two multi-source feature decomposition and combination modules, a multi-source cross-attention fusion module, and a spatial pyramid pooling module. The two multi-source feature decomposition and combination modules are respectively denoted as the first multi-source feature decomposition and combination module and the second multi-source feature decomposition and combination module.

[0077] As the ultra-large-scale RGB features and infrared features, they do not participate in the processing of the decomposition, combination, and cross-fusion network.

[0078] As the large-scale RGB features and infrared features, they are input into the first multi-source feature decomposition and combination module of the decomposition, combination, and cross-fusion network for the decomposition and recombination of basic features and detailed features, generating the first RGB combined feature and the first infrared combined feature

[0079] As the medium-scale RGB features and infrared features, they are input into the second multi-source feature decomposition and combination module of the decomposition, combination, and cross-fusion network for the decomposition and recombination of basic features and detailed features, generating the second RGB combined feature and the second infrared combined feature

[0080] As the small-scale RGB features and infrared features, they are input into the multi-source cross-attention fusion module of the decomposition, combination, and cross-fusion network. The multi-source cross-attention fusion module uses the cross-attention mechanism to guide the cross-modal feature fusion between high-level semantic features, generating the first fused feature and inputting it into the spatial pyramid pooling module. The spatial pyramid pooling module performs spatial pyramid pooling processing to generate the second fused feature

[0081] So far, the decomposition, combination, and cross-fusion network outputs and as the fused features at different scales.

[0082] It should be noted that the theoretical basis of the multi-source feature decomposition and combination module stems from the frequency feature differences of images, that is, the RGB image and the infrared image are taken from the same scene. However, due to the differences in the imaging mechanisms of the sensors, the two modal images show a certain degree of correlation in low-frequency information, both containing common basic features such as the background, large-scale targets, and layouts, while being independent of each other in high-frequency information, representing the detailed features of their respective modalities, such as the color, texture, and details in the RGB image, as well as the thermal radiation information in the infrared image. The purpose of this design is to decompose the RGB-infrared multi-source features into low-frequency basic features and high-frequency detailed features, and introduce the low-frequency basic features of one modality into another modality to improve the richness and accuracy of the features of the other modality.

[0083] In one example, the structure of the multi-source feature decomposition and combination module is as Figure 3 shown, specifically consisting of two basic feature extractors (BFEs), two detailed feature extractors (DFEs), and two output units ( Figure 3 in ). The two basic feature extractors are the RGB basic feature extractor and the infrared basic feature extractor respectively, the two detailed feature extractors are the RGB detailed feature extractor and the infrared detailed feature extractor respectively, and the two output units are the RGB output unit and the infrared output unit respectively.

[0084] After passing through the RGB basic feature extractor and the RGB detailed feature extractor for feature extraction respectively, the RGB basic feature and the RGB detailed feature After passing through the infrared basic feature extractor and the infrared detailed feature extractor for feature extraction respectively, the infrared basic feature and the infrared detailed feature where i = 3, 4.

[0085] The RGB output unit of the first multi-source feature decomposition and combination module will and perform an addition operation to obtain

[0086] The infrared output unit of the first multi-source feature decomposition and combination module will and perform an addition operation to obtain

[0087] The RGB output unit of the second multi-source feature decomposition and combination module will and perform an addition operation to obtain

[0088] The infrared output unit of the second multi-source feature decomposition and combination module and are added together to obtain

[0089] It can be understood that this method of introducing the basic features of another modality in the current modality enriches the low-frequency basic features such as global and structural features, realizes cross-modal feature fusion, and at the same time takes into account the differences between the detailed features of the two modalities and the fact that the high-frequency detailed features may contain invalid information such as noise, thus avoiding problems such as feature conflicts and noise interference that may be caused by introducing the detailed features of another modality.

[0090] In one example, the specific structure of the basic feature extractor is as Figure 4 shown. The basic feature extractor consists of a first LN normalization layer, a multi-head transposed attention module, a second LN normalization layer, and a depth convolutional gated feed-forward network connected in sequence. For the convenience of description, let the input feature of the basic feature extractor be F t-1 , where H, W, and C are the height, width, and number of channels of F t-1 respectively. F t-1 first passes through the first LN normalization layer and the multi-head transposed attention module for local and global cross-channel feature enhancement, and after residual connection, F t is obtained. F t passes through the second LN normalization layer and the depth convolutional gated feed-forward network for gated feature enhancement, and after another residual connection, the basic feature

[0091] The calculation process can be expressed by the formula as:

[0092]

[0093] F t = MDTA[LN(F t-1 )]+F t-1 ;

[0094] where LN(·) represents the LN normalization layer, MDTA(·) represents the multi-head transposed attention module, and GDFN(·) represents the depth convolutional gated feed-forward network.

[0095] Such as Figure 4As shown, the multi-head transposed attention module first uses a 1×1 convolution to increase the dimensionality of the feature channels, and then uses a 3×3 depthwise separable convolution to extract feature information. After channel splitting, Q, K, and V feature vectors are generated. After reshaping the sizes of Q and K, matrix multiplication is performed to generate the transposed attention feature map A across channels. After applying A to V, a 1×1 convolution is used for mapping output. The depth convolutional gated feed-forward network first uses a 1×1 convolution to increase the dimensionality of the feature channels, and then uses a 3×3 depthwise separable convolution to enrich the local features and split them into two groups of parallel paths of features. After one group of features is activated using the GELU non-linear function, element-wise multiplication is performed with the other group of features to form a gating mechanism. Finally, a 1×1 convolution is used to restore the channel dimension.

[0096] In one example, the specific structure of the basic feature extractor is as Figure 5 shown. The detailed feature extractor adopts an affine coupling invertible neural network structure and consists of a local branch and a long-range branch. Similarly, for ease of description, assume the input feature of the detailed feature extractor is F t-1 , F t-1 is first split into local feature F g-1 and long-range feature F l-1 in the channel dimension as two invertible nodes. In the long-range branch, after F g-1 is mapped through the mapping function M1, it is added to F l-1 to obtain the long-range node feature F l . In the local branch, after F l is respectively mapped through the mapping function M2 and the mapping function M3, an enhanced affine transformation is performed with F g-1 to obtain the local node feature F g . Finally, F g and F l are concatenated in the channel dimension to obtain the detailed feature

[0097] The calculation process can be expressed by the formula as:

[0098]

[0099] F g = exp(F g-1 ) ⊙ [M2(F l )] + M3(F l );

[0100] F l = F l-1 + M1(F g-1 );

[0101] Among them, Concat(·) represents channel concatenation, and ⊙ represents the Hadamard product.

[0102] The structures of the mapping functions M1, M2, and M3 are as Figure 5 shown, taking the input feature F in as an example. F in successively passes through modules such as 1×1 convolution, ReLU activation function, 3×3 depthwise separable convolution, squeeze-and-excitation attention (SE), etc. for feature extraction, and finally outputs F out ,

[0103] F out The calculation process of can be expressed by the formula as:

[0104] F out = RELU{Conv 1×1 [SE(RELU<DConv 3×3 {RELU[Conv 1×1 (F in )]}>)]};

[0105] Among them, Conv 1×1 (·) represents 1×1 convolution, RELU(·) represents the ReLU activation function, DConv 3×3 (·) represents 3×3 depthwise separable convolution, and SE(·) represents the squeeze-and-excitation attention mechanism.

[0106] In one example, the specific structure of the multi-source cross-attention fusion module can be as Figure 6 shown. The multi-source cross-attention fusion module includes two stages: the cross-modal feature interaction stage and the multi-source feature enhancement stage. The cross-modal feature interaction stage includes an RGB interaction branch and an infrared interaction branch, and each branch contains a multi-head attention module ( Figure 6 the multi-head attention module 1 and multi-head attention module 2 in ).

[0107] In the RGB interaction branch, the multi-head attention module takes as the query vector, takes as the key vector and the value of the key vector, calculates the cross-modal attention, aiming to use the infrared feature to guide the RGB feature to capture the inter-modal correlation, and then performs a residual connection with to obtain the RGB interaction feature

[0108] In the infrared interaction branch, the multi-head attention module takes as the query vector, takes As the key vectors and the values of the key vectors, cross-modal attention is calculated with the aim of using RGB features to guide infrared features to capture the cross-modal correlation, and then perform a row residual connection to obtain infrared interaction features

[0109] and The calculation process is expressed by the formula as:

[0110]

[0111] where, MHA(·) represents the multi-head attention module.

[0112] In the multi-source feature enhancement stage, first and are concatenated in channels, and after being processed by the LN normalization layer, the rough fusion feature F att is obtained. Subsequently, feature enhancement is performed through the multi-head attention module and the feed-forward neural network layer, and a residual connection is established. Finally,

[0113] The calculation process is expressed by the formula as:

[0114]

[0115] where, Concat(·) represents channel concatenation, LN(·) represents the LN normalization layer, MHA(·) represents the multi-head attention module, and FN(·) represents the feed-forward neural network layer.

[0116] Step 103, through the dual-branch neck network, multi-path feature aggregation of the RGB modality and the infrared modality is respectively performed on the fusion features of different scales to obtain aggregation features of different scales.

[0117] In a specific implementation, the decomposition-combination and cross-fusion network sends the fusion features of different scales into the dual-branch neck network, and the dual-branch neck network respectively performs multi-path feature aggregation of the RGB modality and the infrared modality on the fusion features of different scales to obtain aggregation features of different scales.

[0118] In an example, the specific structure of the dual-branch neck network is as Figure 2 shown. The dual-branch neck network consists of an RGB neck branch and an infrared neck branch. The network structures of both the RGB neck branch and the infrared neck branch are top-down and bottom-up multi-path aggregation networks (Path Aggregation Network, PAN).

[0119] In the RGB neck branch, the top-down path is composed of an ordinary PAN module, a first visible light PAN module, and a second visible light PAN module connected in sequence. The bottom-up path is composed of a third visible light PAN module and a fourth visible light PAN module connected in sequence. The output of the ordinary PAN module is also used as the input of the fourth visible light PAN module, and the output of the first visible light PAN module is also used as the input of the third visible light PAN module.

[0120] and are respectively input into the ordinary PAN module, the first visible light PAN module, and the second visible light PAN module. The second visible light PAN module, the third visible light PAN module, and the fourth visible light PAN module respectively output three different scales of RGB aggregation features and

[0121] In the infrared neck branch, the top-down path is composed of an ordinary PAN module, a first thermal infrared PAN module, and a second thermal infrared PAN module connected in sequence. The bottom-up path is composed of a third thermal infrared PAN module and a fourth thermal infrared PAN module connected in sequence. The output of the ordinary PAN module is also used as the input of the fourth thermal infrared PAN module, and the output of the first thermal infrared PAN module is also used as the input of the third thermal infrared PAN module.

[0122] and are respectively input into the ordinary PAN module, the first thermal infrared PAN module, and the second thermal infrared PAN module. The second thermal infrared PAN module, the third thermal infrared PAN module, and the fourth thermal infrared PAN module respectively output three different scales of infrared aggregation features and

[0123] Step 104, perform object detection in RGB mode and infrared mode on the aggregation features of different scales through a dual-branch detection network, and after performing non-maximum suppression processing, obtain the RGB detection result and the infrared detection result.

[0124] In a specific implementation, the dual-branch neck network sends the aggregation features of different scales into the dual-branch detection network. The dual-branch detection network performs object detection in RGB mode and infrared mode on the aggregation features of different scales, and after performing non-maximum suppression processing, obtains the RGB detection result and the infrared detection result.

[0125] In an example, the specific structure of the dual-branch detection network can be as Figure 2 shown. The dual-branch detection network is composed of an RGB detection branch and an infrared detection branch, and each branch is provided with three detection modules.

[0126] In the RGB detection branch, three detection modules respectively perform and target detection in the RGB modality, obtaining RGB detection results at three scales and

[0127] In the infrared detection branch, three detection modules respectively perform and target detection in the infrared modality, obtaining infrared detection results at three scales and

[0128] In the training stage, independent annotations for RGB and infrared are used to supervise the RGB and infrared branch networks respectively. In the inference stage, the two branches respectively predict RGB and infrared images, avoiding the ambiguity brought by using a single detection result to represent the target distribution in the two modalities. For example, under dark night and no-light conditions, the RGB modality fails and the image presents a pure black style. It is unreasonable to use the infrared detection result to indicate the position of the target in the RGB image, which is misleading for practical applications.

[0129] In one example, when training the detection model, the total loss function includes the classification loss the regression loss the distribution focus loss and the multi-source feature decomposition loss which is expressed by the formula as:

[0130]

[0131] h ∈ {rgn, ir};

[0132] where λ cls , λ box , λ dfl , λ decomp are respectively weights.

[0133] In one example, λ cls , λ box , λ dfl , λ decomp are respectively set to 0.5, 7.5, 1.5, 20.

[0134] In one example, is the binary cross-entropy loss, is the CIoU loss, is the DFL loss, The purpose is to increase the correlation between multi-source basic features, reduce the correlation between multi-source detailed features, and promote the decomposition of multi-source features. The calculation process is expressed by the formula as:

[0135]

[0136] Where, is the Pearson correlation coefficient operator, ∈ = 1.01, which is used to ensure that the calculation result is positive.

[0137] A multi-source object detection method based on multi-source feature cross-fusion and decomposition combination proposed in this embodiment, compared with the existing multi-source object detection methods, designs, constructs, trains and uses a detection model composed of a dual-branch backbone network, a decomposition combination and cross-fusion network, a dual-branch neck network and a dual-branch detection network to achieve multi-source object detection. The dual-branch backbone network is used to downsample and extract multi-scale features from the RGB-infrared image pair, obtaining RGB features and infrared features at different scales, thus effectively expanding the number of features. The decomposition combination and cross-fusion network is used to decompose and recombine the large-scale RGB features and infrared features, and the medium-scale RGB features and infrared features respectively, and perform multi-source cross-attention fusion on the small-scale RGB features and infrared features, obtaining fusion features at different scales. The design of decomposition and recombination, as well as multi-source cross-attention fusion, realizes the complementary fusion of multi-source features, effectively improving the utilization degree of the detection model for multi-modal features. The dual-branch neck network is used to perform multi-path feature aggregation of RGB modality and infrared modality on the fusion features at different scales, obtaining aggregated features at different scales. The dual-branch detection network is then used to perform object detection of RGB modality and infrared modality on the aggregated features at different scales respectively, and after suppressing with non-maximum value, obtaining RGB detection results and infrared detection results. Using the dual-branch neck network and the dual-branch detection network to independently aggregate and detect objects in two modalities, while realizing feature complementary fusion, retains a certain degree of feature independence, so that the problem of object detection in weakly registered and other modality imbalance scenarios can be well solved. The entire detection model adopts a lightweight basic network structure and is modularly integrated into one, and such a detection model can achieve multi-source object detection with high precision and high timeliness.

[0138] The step division of the above various methods is only for clear description. When implemented, they can be combined into one step, or some steps can be split and decomposed into multiple steps. As long as they include the same logical relationship, they are all within the protection scope of this application; adding insignificant modifications or introducing insignificant designs to the algorithm or process, but not changing the core design of its algorithm and process are all within the protection scope of this application.

[0139] In one embodiment, the detection model proposed in this application is modified based on the lightweight single-source one-stage YOLOv8 network, which can be called UAV-DMDet. To verify the effectiveness of UAV-DMDet, we used the publicly available RGB-infrared multi-source object detection dataset DroneVehicle for the training and testing of the network framework and compared it with other mainstream methods. The DroneVehicle training set contains 17,990 pairs of RGB-infrared images, and the test set contains 8,980 pairs of RGB-infrared images. Its detection targets include 5 types of aerial detection targets: cars, vans, buses, trucks, and vans.

[0140] Based on this, we compared UAV-DMDet with mainstream RGB-infrared multi-source object detection algorithms (detection models), and the evaluation metrics involved multiple aspects such as accuracy, speed, and the number of parameters. The mean average precision (mAP@0.5 and mAP@0.5:0.95) was used to evaluate the algorithm accuracy, and the higher the performance, the better. The frames per second (FPS) was used to evaluate the algorithm inference time, and more than 30 FPS could be defined as real-time. Parms and GFLOPs were used to evaluate the number of model parameters and complexity. The smaller the value, the lighter the model, the less computing resources required, and it is more suitable for deployment and application on edge devices such as drones. The relevant experimental results are as Figure 7 shown. Among the current mainstream multi-source object detection methods, the 2 IMDet method showed relatively competitive results, with mAP@0.5 reaching 77.3% and mAP@0.5:0.95 reaching 46.2%. Compared with the 2 IMDet method, the UAV-DMDet method proposed in this application improved by 2.66% and 12.36% respectively in terms of the mAP@0.5 and mAP@0.5:0.95 accuracy metrics, and achieved the best accuracy in all sub-categories, demonstrating strong performance advantages. At the same time, the number of parameters of the UAV-DMDet model is less than 2 1 / 5 of the IMDet method, and the computational complexity is about 3 / 5, which is more suitable for deployment and application on platforms such as drones. In addition, the UAV-DMDet method achieved a detection speed of 41.15 frames per second on a single NVIDIA GeForce RTX3090 graphics card, realizing real-time object detection.

[0141] Another embodiment of this application proposes an electronic device, and its specific structure can be as Figure 8As shown, it includes: at least one processor 201; and a memory 202 communicatively connected to the at least one processor 201; wherein, the memory 202 stores instructions executable by the at least one processor 201, and the instructions are executed by the at least one processor 201 to enable the at least one processor 201 to execute a multi-source target detection method based on multi-source feature cross-fusion and decomposition combination as described in the above method embodiments.

[0142] Among them, the memory and the processor are connected in a bus manner. The bus can include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and the memory together. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, so they will not be further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be an element or multiple elements, such as multiple receivers and transmitters, and provides a unit for communicating with various other devices on the transmission medium. The data processed by the processor is transmitted over the wireless medium through the antenna. Further, the antenna also receives data and transmits the data to the processor.

[0143] The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. The memory can be used to store the data used by the processor when executing operations.

[0144] Another embodiment of the present application proposes a computer-readable storage medium storing a computer program, which when executed by a processor, can implement a multi-source target detection method based on multi-source feature cross-fusion and decomposition combination as described in the above method embodiments.

[0145] That is, those skilled in the art can understand that all or part of the steps in implementing the above embodiment methods can be completed by instructing relevant hardware through a program. This program is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The above storage medium includes: various media such as USB flash drives, mobile hard disks, ROM (Read-Only Memory), RAM (Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0146] Those of ordinary skill in the art can understand that the above embodiments are specific embodiments for implementing the present application. In actual applications, various changes can be made to them in form and details without departing from the spirit and scope of the present application.

Claims

1. A multi-source object detection method based on multi-source feature cross-fusion and decomposition combination, which is implemented based on a pre-trained detection model composed of a dual-branch backbone network, a decomposition combination and cross-fusion network, a dual-branch neck network, and a dual-branch detection network, characterized in that The method includes: Downsampling and multi-scale feature extraction are performed on the RGB-infrared image pair through a dual-branch backbone network to obtain RGB features and infrared features at different scales; Through a decomposition-combination and cross-fusion network, the large-scale RGB features and infrared features, and the medium-scale RGB features and infrared features are respectively decomposed and recombined, and the multi-source cross-attention fusion is performed on the small-scale RGB features and infrared features to obtain fusion features at different scales; Through a dual-branch neck network, multi-path feature aggregation in the RGB modality and the infrared modality is respectively performed on the fusion features at different scales to obtain aggregated features at different scales; Through a dual-branch detection network, object detection in the RGB modality and the infrared modality is respectively performed on the aggregated features at different scales, and after non-maximum suppression processing, RGB detection results and infrared detection results are obtained.

2. The multi-source target detection method based on multi-source feature cross-fusion and decomposition combination according to claim 1, wherein, The dual-branch backbone network consists of an RGB feature extraction branch and an infrared feature extraction branch; The RGB feature extraction branch consists of five sequentially connected visible light convolution modules. The input of the first visible light convolution module is the RGB image in the RGB-infrared image pair, and the input of the subsequent visible light convolution module is the output of the previous visible light convolution module. All five visible light convolution modules are used to downsample and extract features from their own inputs, obtaining RGB features with resolutions of 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the RGB image respectively. The five RGB features are sequentially denoted as and The infrared feature extraction branch consists of five sequentially connected thermal infrared convolutional modules. The input of the first thermal infrared convolutional module is the infrared image in the RGB-infrared image pair, and the input of the subsequent thermal infrared convolutional module is the output of the previous thermal infrared convolutional module. All five thermal infrared convolutional modules are used to downsample and extract features from their own inputs, obtaining infrared features with resolutions of 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the infrared image respectively. The five infrared features are sequentially denoted as and Among them, both the visible light convolution module and the thermal infrared convolution module are composed of a basic convolution unit and a cross-stage local convolution unit connected in series. The basic convolution unit is composed of a 3×3 convolution layer, a batch normalization layer, and a SiLU activation function layer connected in series. The cross-stage local convolution unit first uses a basic convolution unit to increase the dimension of the input, and then splits it into two feature maps F p1 and F p2 , F p1 for residual connection, and F p2 then sequentially performs feature extraction through multiple basic convolution units. Finally, the two parts of the feature maps are concatenated by channels and the basic convolution unit is used for dimensionality reduction and output.

3. A multi-source object detection method based on multi-source feature cross-fusion and decomposition combination according to claim 2, characterized in that The decomposition-combination and cross-fusion network consists of two multi-source feature decomposition-combination modules, a multi-source cross-attention fusion module, and a spatial pyramid pooling module. The two multi-source feature decomposition-combination modules are respectively the first multi-source feature decomposition-combination module and the second multi-source feature decomposition-combination module; As large-scale RGB features and infrared features, they are input into the first multi-source feature decomposition and combination module of the decomposition, combination and cross-fusion network for the decomposition and recombination of basic features and detailed features, generating the first RGB combined feature and the first infrared combined feature As mesoscale RGB features and infrared features, they are input into the second multi-source feature decomposition and combination module of the decomposition, combination and cross-fusion network to decompose and recombine basic features and detailed features, generating second RGB combined features and second infrared combined features As small-scale RGB features and infrared features, they are input into the multi-source cross-attention fusion module of the decomposition-combination and cross-fusion network. The multi-source cross-attention fusion module uses the cross-attention mechanism to guide the cross-modal feature fusion between high-level semantic features and generate the first fusion feature. It is input into the spatial pyramid pooling module of the decomposition, combination and cross-fusion network, and spatial pyramid pooling processing is performed by the spatial pyramid pooling module to generate a second fusion feature 4. A multi-source target detection method based on multi-source feature cross-fusion and decomposition and combination according to claim 3, characterized in that, The multi-source feature decomposition-combination module consists of two basic feature extractors, two detail feature extractors, and two output units. The two basic feature extractors are respectively the RGB basic feature extractor and the infrared basic feature extractor. The two detail feature extractors are respectively the RGB detail feature extractor and the infrared detail feature extractor. The two output units are respectively the RGB output unit and the infrared output unit; Feature extraction is performed through an RGB basic feature extractor and an RGB detailed feature extractor respectively to obtain RGB basic features and RGB detailed features Feature extraction is performed through an infrared basic feature extractor and an infrared detailed feature extractor respectively to obtain infrared basic features and infrared detailed features i = 3, 4; The RGB output unit of the first multi-source feature decomposition and combination module adds and to obtain The infrared output unit of the first multi-source feature decomposition and combination module will and be added to obtain The RGB output unit of the second multi-source feature decomposition and combination module adds and to obtain The infrared output unit of the second multi-source feature decomposition and combination module will and perform an addition process to obtain 5. A multi-source target detection method based on multi-source feature cross-fusion and decomposition combination according to claim 4, characterized in that, The basic feature extractor consists of a first LN normalization layer, a multi-head transposed attention module, a second LN normalization layer, and a depth convolutional gated feed-forward network connected in sequence. Let the input feature of the basic feature extractor be F t-1 , F t-1 First, it passes through the first LN normalization layer and the multi-head transposed attention module for local and global cross-channel feature enhancement, and obtains F through a residual connection t , F t Then it passes through the second LN normalization layer and the depth convolutional gated feed-forward network for gated feature enhancement, and obtains the basic feature through a residual connection The calculation process is expressed by the formula as follows: F t = MDTA[LN(F t-1 )] + F t-1 ; Among them, LN(·) represents the LN normalization layer, MDTA(·) represents the multi-head transposed attention module, and GDFN(·) represents the depth convolutional gated feed-forward network; The detailed feature extractor adopts an affine coupling reversible neural network structure and consists of a local branch and a long-range branch. Let the input feature of the detailed feature extractor be F t-1 , F t-1 First, it is split into local feature F g-1 and long-range feature F l-1 in the channel dimension as two reversible nodes. In the long-range branch, F g-1 is mapped through the mapping function M1 and then added to F l-1 to obtain the long-range node feature F l . In the local branch, F l is respectively mapped through the mapping function M2 and the mapping function M3 and then undergoes an enhanced affine transformation with F g-1 to obtain the local node feature F g . Finally, F g and F l are concatenated in the channel dimension to obtain the detailed feature The calculation process is expressed by the formula as follows: F g = exp(F g-1 ) ⊙ [M2(F l )] + M3(F l ); F l = F l-1 + M1(F g-1 ); Among them, Concat(·) represents channel concatenation, and ⊙ represents the Hadamard product.

6. The multi-source target detection method based on multi-source feature cross-fusion and decomposition combination according to claim 3, wherein The multi-source cross-attention fusion module includes two stages: a cross-modal feature interaction stage and a multi-source feature enhancement stage. The cross-modal feature interaction stage includes an RGB interaction branch and an infrared interaction branch, and each of the two branches contains a multi-head attention module; In the RGB interaction branch, the multi-head attention module takes as the query vector, takes as the key vector and the value of the key vector, calculates the cross-modal attention, and then performs a residual connection with to obtain the RGB interaction feature In the infrared interaction branch, the multi-head attention module takes as the query vector, takes as the key vector and the value of the key vector, calculates the cross-modal attention, and then performs a residual connection with to obtain the infrared interaction feature and The calculation process is expressed by the formula as follows: Among them, MHA(·) represents the multi-head attention module; In the multi-source feature enhancement stage, first, and are concatenated in channels. After being processed by the LN normalization layer, the rough fusion feature F att is obtained. Subsequently, feature enhancement is performed through the multi-head attention module and the feed-forward neural network layer, and a residual connection is established. Finally, The calculation process is expressed by the formula as follows: Among them, Concat(·) represents channel concatenation, LN(·) represents the LN normalization layer, MHA(·) represents the multi-head attention module, and FN(·) represents the feed-forward neural network layer.

7. A multi-source target detection method based on multi-source feature cross-fusion and decomposition combination according to claim 3, characterized in that, The dual-branch neck network consists of an RGB neck branch and an infrared neck branch. The network structures of the RGB neck branch and the infrared neck branch are both top-down and bottom-up multi-path aggregation networks; In the RGB neck branch, the top-down path is composed of a sequentially connected ordinary PAN module, a first visible light PAN module, and a second visible light PAN module. The bottom-up path is composed of a sequentially connected third visible light PAN module and a fourth visible light PAN module. The output of the ordinary PAN module also serves as the input of the fourth visible light PAN module, and the output of the first visible light PAN module also serves as the input of the third visible light PAN module. and are respectively input into the ordinary PAN module, the first visible light PAN module, and the second visible light PAN module. The second visible light PAN module, the third visible light PAN module, and the fourth visible light PAN module respectively output three different scales of RGB aggregation features. and In the infrared neck branch, the top-down path is composed of an ordinary PAN module, a first thermal infrared PAN module, and a second thermal infrared PAN module connected in sequence. The bottom-up path is composed of a third thermal infrared PAN module and a fourth thermal infrared PAN module connected in sequence. The output of the ordinary PAN module is also used as the input of the fourth thermal infrared PAN module, and the output of the first thermal infrared PAN module is also used as the input of the third thermal infrared PAN module. and are respectively input into the ordinary PAN module, the first thermal infrared PAN module, and the second thermal infrared PAN module. The second thermal infrared PAN module, the third thermal infrared PAN module, and the fourth thermal infrared PAN module respectively output three different scales of infrared aggregation features and 8. A multi-source target detection method based on multi-source feature cross-fusion and decomposition and combination according to claim 4, characterized in that The total loss function used when training the detection model includes the classification loss the regression loss the distribution focusing loss and the multi-source feature decomposition loss which is expressed by the formula as: h ∈ {rgb, ir}; Among them, λ cls , λ box , λ dfl , λ decomp are respectively weights.

9. An electronic device, characterized in that, Including: At least one processor; And a memory communicatively connected to the at least one processor; Among them, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a multi-source object detection method based on multi-source feature cross-fusion and decomposition-combination as described in any one of claims 1 to 8.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it can implement a multi-source object detection method based on multi-source feature cross-fusion and decomposition and combination as described in any one of claims 1 to 8.

Citation Information

Cited By

  • Sleep apnea detection method and system, electronic equipment and storage medium

    CN120477748A

  • Method and device for detecting power transmission line fault through line fault detection model

    CN120876400A

  • Power plant type detection method based on deep learning

    CN121353809A

  • A power plant type detection method based on deep learning

    CN121353809B