Multi-source Object Detection Method Based on Pixel-level Dynamic Modal Selection and Fusion

Through the combination of a dual-branch convolutional neural network and a pixel-level dynamic modal weight selection network, the problem of insufficient detection accuracy and speed of traditional methods in complex scenarios is solved, and efficient multimodal feature fusion and object detection are achieved.

CN120107568BActive Publication Date: 2025-07-04NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510587275.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-07-04
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

The existing visible light object detection algorithms have insufficient detection accuracy under insufficient light, haze weather and occlusion. The traditional multimodal feature fusion method has a large amount of calculation and a large amount of parameters, which is difficult to meet the needs of complex scenarios, and the fusion particle size is large, so fine-grained fusion cannot be achieved.

Method used

A dual-branch convolutional neural network encoder is used to perform preliminary feature extraction for visible light and infrared images, a pixel-level dynamic modal weight selection network is used for linear projection and position encoding, and a modal selection weight is generated in combination with the CoT attention mechanism, feature extraction and fusion are performed through the deep network, and object detection is finally realized in the detection head decoder.

Benefits of technology

It greatly improves the accuracy and speed of object detection in complex scenarios, reduces the amount of calculation and parameter, and realizes fine-grained multimodal feature fusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107568B_ABST
    Figure CN120107568B_ABST
Patent Text Reader

Abstract

The present application relates to the field of artificial intelligence technology, and discloses a multi-source object detection method based on pixel-level dynamic modality selection and fusion, including: using a dual-branch convolutional neural network encoder to perform preliminary feature extraction on visible light images and infrared images to obtain visible light preliminary features and infrared preliminary features; using a pixel-level dynamic modality weight selection network to perform linear projection, position encoding addition, and modality weight selection based on the CoT attention mechanism on the visible light preliminary features and infrared preliminary features respectively to generate modality selection weights, and finally performing pixel-level dynamic feature selection based on the modality selection weights to obtain visible light reselected features and infrared reselected features; using a deep network to perform deep feature extraction on the visible light reselected features and infrared reselected features, and then adding them element by element according to the resolution to obtain fusion features; inputting the fusion features into a detection head decoder for detection to obtain object detection results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of artificial intelligence technology, and particularly to a multi-source target detection method based on pixel-level dynamic modality selection and fusion. Background Art

[0002] Target detection is one of the basic tasks of computer vision technology. Its purpose is to detect objects of interest in a given image and give corresponding position and category information. Common visible light target detection algorithms are difficult to accurately detect targets in the scene under conditions of insufficient light (such as at night, on cloudy days), haze weather, and certain occlusion situations. Based on this, some researchers have attempted to combine visible light images and infrared images to improve the performance of target detection under complex conditions such as low light by combining the characteristics of rich color and texture detail information of visible light images and the characteristics of infrared images being unaffected by lighting conditions.

[0003] In existing research, technicians usually use methods based on channel splicing or methods based on attention mechanisms to handle the fusion problem of multi-modal feature maps. Through such a single convolutional neural network or a single attention mechanism, certain effects can be achieved. For example, Fang et al. designed a fusion module to fuse visible light features and infrared features into a unified representation to make full use of the visible light features and infrared features extracted by two parallel backbone networks.

[0004] However, in practical applications, methods based on channel splicing and methods based on attention mechanisms still have poor performance in many complex scenarios, and mainly have the following problems.

[0005] First, the scenes of visible light images and infrared images are complex and variable. For different scenes, it is more scientific to use different feature selection strategies. A single feature fusion method is difficult to meet the changing needs.

[0006] Second, most fusion algorithms are carried out before the multi-scale feature pyramid module. The fusion module needs to be reused at multiple scales, with a large amount of computational parameters, and there may be inconsistencies in the fusion at different scales.

[0007] Third, traditional algorithms often use channel attention and spatial attention to perform feature fusion between the visible light modality and the infrared modality. The granularity of this fusion is often large and cannot perform fine-grained fusion on the targets in the scene. Summary of the Invention

[0008] To solve the above technical problems, embodiments of the present application propose a multi-source target detection method based on pixel-level dynamic modality selection and fusion, aiming to extract deep features of visible light modality and infrared modality through multi-source collaborative learning, and perform pixel-level dynamic fusion, so as to make full use of the useful information in different modality images to improve the accuracy and speed of target detection in complex scenarios.

[0009] To achieve the above object, embodiments of the present application propose a multi-source target detection method based on pixel-level dynamic modality selection and fusion, including the following steps: using a dual-branch convolutional neural network encoder to compress and perform preliminary feature extraction on paired visible light images and infrared images to obtain preliminary visible light features and preliminary infrared features; using a pixel-level dynamic modality weight selection network, first linearly project the preliminary visible light features and preliminary infrared features respectively to obtain decoupled visible light features and decoupled infrared features, then add position encoding to the decoupled visible light features and decoupled infrared features, and perform modality weight selection based on the CoT (Contextual Transformer) attention mechanism to generate modality selection weights, and finally multiply the modality selection weights element-wise with the decoupled visible light features and decoupled infrared features respectively to obtain reselected visible light features and reselected infrared features; using a deep network to perform deep feature extraction on the reselected visible light features and reselected infrared features to sequentially obtain visible light deep features and infrared deep features with resolutions of 1 / 8, 1 / 16, and 1 / 32 of the original resolution, add the visible light deep features and infrared deep features with the same resolution element-wise to obtain three fusion features; input the three fusion features into a detection head decoder for detection to obtain target detection results.

[0010] To achieve the above object, embodiments of the present application also propose an electronic device, where the electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, instructions executable by the at least one processor are stored in the memory, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute a multi-source target detection method based on pixel-level dynamic modality selection and fusion as described above.

[0011] To achieve the above object, embodiments of the present application also propose a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it can implement a multi-source target detection method based on pixel-level dynamic modality selection and fusion as described above.

[0012] An multi-source object detection method based on pixel-level dynamic modality selection fusion proposed by an embodiment of the present application uses a dual-branch convolutional neural network encoder, a pixel-level dynamic modality weight selection network, a deep network, and a detection head decoder to jointly complete the multi-source object detection task. The dual-branch convolutional neural network encoder can compress and perform preliminary feature extraction on the input image, thereby retaining useful semantic information and removing redundant information. These semantic information represent global features and local features in the scene, preparing for subsequent processing. The pixel-level dynamic modality weight selection network can dynamically select important regions in visible light images and infrared images according to different scenes, and better combine multi-modal features through pixel-level dynamic feature selection. After sending the reselected visible light features and reselected infrared features into the deep network, only deep feature extraction and simple addition are required to obtain multi-scale fusion features, and then the output of the object detection result is realized through the detection head decoder, greatly reducing the number of parameters of traditional fusion methods, and thus significantly improving the accuracy and speed of object detection in complex scenes.

[0013] In some alternative embodiments, the two branches of the dual-branch convolutional neural network encoder are respectively a visible light preliminary extraction branch and an infrared preliminary extraction branch. The structures of the two branches are the same and are both composed of a series-connected convolutional neural network layer, a BN batch normalization layer, and a ReLU activation function layer;

[0014] The visible light preliminary extraction branch compresses and performs preliminary feature extraction on the visible light image to obtain visible light preliminary features with a resolution of 1 / 2 of the visible light image ;

[0015] The infrared preliminary extraction branch compresses and performs preliminary feature extraction on the infrared image to obtain infrared preliminary features with a resolution of 1 / 2 of the infrared image .

[0016] In some alternative embodiments, the pixel-level dynamic modality weight selection network is composed of a visible light linear projection unit, an infrared linear projection unit, a visible light position encoding unit, an infrared position encoding unit, and a modality weight selection unit;

[0017] The visible light linear projection unit uses 1×1 convolution to perform linear projection to achieve decoupling operation and obtain visible light decoupled features , and the infrared linear projection unit uses 1×1 convolution to perform linear projection to achieve decoupling operation and obtain infrared decoupled features ;

[0018] The visible light position encoding unit and the infrared position encoding unit generate two sets of identical position encoding vectors. The position encoding vectors are grids formed by taking pixel blocks as basic units, denoted as , , the visible light position encoding unit will and perform channel splicing to generate the visible light decoupled feature after adding the position encoding , the visible light position encoding unit will and perform channel splicing to generate the infrared decoupled feature after adding the position encoding ;

[0019] The modal weight selection unit first performs projection mapping and normalization processing on and to obtain the visible light projection normalized feature and the infrared projection normalized feature , then based on the CoT attention mechanism, respectively process and to generate the visible light weight and the infrared weight , then use the Sigmoid function to normalize and and then add them, and use the Softmax function to generate the modal selection weight , finally based on perform element-wise multiplication with and respectively according to the modality to obtain the visible light reselected feature and the infrared reselected feature .

[0020] In some alternative embodiments, the generation of the position encoding vector is achieved through the following steps:

[0021] Generate a one-dimensional vector with a fixed interval , The th element in is represented by the formula:

[0022] ;

[0023] ;

[0024] where is the preset window size, is The th element in;

[0025] Use the torch.mershgrid function to create a two-dimensional grid based on , ​It consists of two parts, namely the horizontal coordinate and the vertical coordinate , The size of is Expressed by the formula as:

[0026] ;

[0027] Among them, represents the torch.meshgrid function;

[0028] Stack into a tensor with a shape of , ,

[0029] Expand the dimension of , add a dimension in the zero-th dimension to obtain a tensor with a shape of

[0030] ; Repeat to match the batch size of the input data, and finally obtain a position encoding vector with a shape of

[0031] , Expressed by the formula as:

[0031] ;

[0032] Among them, is the batch size during training, is , the height of is , the width of

[0033] and perform projection mapping and normalization processing to obtain the visible light projection normalized feature and the infrared projection normalized feature

[0034] , Expressed by the formula as:

[0034] ;

[0035] ;

[0036] Among them, represents a 1×1 convolution, Presentation layer normalization processing;

[0037] Use the Sigmoid function to and After normalization, add them together, and use the Softmax function to generate the modality selection weights , which is expressed by the formula as:

[0038] ;

[0039] Among them, Represents the Sigmoid function, Represents the Softmax function, The shape of is , Is the batch size during training, Is the preset sequence length;

[0040] Based on Multiply element by element with and respectively according to the modality to obtain the visible light reselected feature and the infrared reselected feature , including:

[0041] For and Cut them pixel by pixel according to to obtain the visible light sequence and the infrared sequence , and The size of corresponds to the shape of ;

[0042] Regard as the weight corresponding to the visible light modality, regard as the weight corresponding to the infrared modality, multiply element by element with to obtain , multiply element by element with to obtain .

[0043] In some alternative embodiments, based on the CoT attention mechanism, process and respectively to generate the visible light weight and the infrared weight , including:

[0044] For with the shape of , , first, a vector of the same shape is obtained through the Key_Embed layer , and then is sent into the Value_Embed layer to obtain ; ;

[0045] is and are concatenated in the channel dimension to obtain , and then the attention weights are generated through the Attention_Embed layer, the mean function layer, and the Softmax function layer;

[0046] is multiplied by to obtain the weighted value feature , and then is reshaped back to the original size and added to to obtain the final weight ;

[0047] Among them, the Key_Embed layer, the Value_Embed layer, and the Attention_Embed layer are 3×3 convolutional layers, 1×1 convolutional layers, and 1×1 convolutional layers respectively, is , of the number of channels.

[0048] In some alternative embodiments, the deep network consists of two deep branches, namely the visible light deep extraction branch and the infrared deep extraction branch. The structures of the two deep branches are the same and are both composed of multiple ConvNormLayers and multiple ResBlocks;

[0049] The visible light deep extraction branch performs deep feature extraction on at different scales, and successively obtains visible light deep features , and with resolutions of 1 / 8, 1 / 16, and 1 / 32 of the original resolution;

[0050] The infrared deep extraction branch performs deep feature extraction on at different scales, and successively obtains infrared deep features , and with resolutions of 1 / 8, 1 / 16, and 1 / 32 of the original resolution;

[0051] is multiplied by , multiplied by , and as well as With , element-by-element addition is performed respectively to obtain three fused features in sequence 、 and .

[0052] In some alternative embodiments, the dual-branch convolutional neural network encoder, the pixel-level dynamic modality weight selection network, the deep network, and the detection head decoder together constitute a multi-source object detection model. When iteratively training the multi-source object detection model, the total loss function used is expressed by the formula:

[0053] ;

[0054] wherein is the L1 loss, is the generalized intersection over union loss, is the classification loss, and are used to calculate the difference between the position of the predicted object and the position of the true object, is used to calculate the difference between the category of the predicted object and the category of the true object, represents the total loss function. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] To more clearly illustrate the embodiments of the present application or the technical solutions in the related art, the following will briefly introduce the drawings required for the description of the embodiments of the present application or the related art. The following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. The drawings described herein are only used to explain the present application and are not used to limit the present application.

[0056] Figure 1 is a flowchart of a multi-source object detection method based on pixel-level dynamic modality selection fusion provided in an embodiment of the present application;

[0057] Figure 2 is a schematic structural diagram of a multi-source object detection model provided in an embodiment of the present application;

[0058] Figure 3 is a schematic structural diagram of a preliminary extraction branch provided in an embodiment of the present application;

[0059] Figure 4 is a schematic structural diagram of a pixel-level dynamic modality weight selection network provided in an embodiment of the present application;

[0060] Figure 5It is a schematic structural diagram of a modal weight selection unit provided in an embodiment of the present application;

[0061] Figure 6 It is a schematic structural diagram of a CoT attention mechanism provided in an embodiment of the present application;

[0062] Figure 7 It is a schematic diagram of pixel-level segmentation provided in an embodiment of the present application;

[0063] Figure 8 It is a schematic structural diagram of a deep network provided in an embodiment of the present application;

[0064] Figure 9 It is a schematic diagram of the performance comparison between a traditional multi-source target detection algorithm and the method proposed in the present application in an embodiment of the present application;

[0065] Figure 10 It is a schematic structural diagram of an electronic device provided in another embodiment of the present application. Detailed implementation manners

[0066] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the embodiments of the present application will be elaborated in detail below with reference to the accompanying drawings. Those of ordinary skill in the art can understand that in the embodiments of the present application, many technical details are provided for readers to better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions required to be protected by the present application can still be implemented. The following division of the embodiments is only for convenience of description and should not constitute any limitation on the specific implementation manner of the present application. The embodiments can be combined and cited with each other on the premise of not conflicting.

[0067] To solve the problem that the performance of traditional multi-source target detection methods is poor in complex scenarios, an embodiment of the present application proposes a multi-source target detection method based on pixel-level dynamic modal selection and fusion, which is applied to a server. The implementation details of a multi-source target detection method based on pixel-level dynamic modal selection and fusion proposed in this embodiment will be specifically described below. The following content is only the implementation details provided for convenience of understanding and is not necessary for implementing this solution.

[0068] The specific process of a multi-source target detection method based on pixel-level dynamic modal selection and fusion proposed in this embodiment can be as Figure 1 shown, including:

[0069] Step 101, using a dual-branch convolutional neural network encoder to compress and perform preliminary feature extraction on paired visible light images and infrared images to obtain visible light preliminary features and infrared preliminary features.

[0070] In a specific implementation, the dual-branch convolutional neural network encoder, the pixel-level dynamic modality weight selection network, the deep network, and the detection head decoder together form a multi-source object detection model as shown in Figure 2 . The input of the multi-source object detection model is a pair of visible light images and infrared images, and the output is the object detection result. The pair of visible light images and infrared images are input into the dual-branch convolutional neural network encoder for compression and preliminary feature extraction, so as to obtain the preliminary visible light features and the preliminary infrared features. The preliminary visible light features and the preliminary infrared features obtained through compression and preliminary feature extraction carry more semantic information, and these semantic information represent the global features and local features in the detection scene, preparing for subsequent detection operations.

[0071] In one example, the two branches of the dual-branch convolutional neural network encoder are the preliminary visible light extraction branch and the preliminary infrared extraction branch respectively. The structures of the two preliminary extraction branches are the same, both are ConvNormLayer, that is, they are both composed of a series of convolutional neural network layers, BN batch normalization layers, and ReLU activation function layers. The specific structure of the preliminary extraction branch can be as shown in Figure 3 . The preliminary visible light extraction branch compresses and preliminarily extracts features from the visible light image to obtain the preliminary visible light features with a resolution of 1 / 2 of the visible light image , and the preliminary infrared extraction branch compresses and preliminarily extracts features from the infrared image to obtain the preliminary infrared features with a resolution of 1 / 2 of the infrared image .

[0072] Step 102: Use the pixel-level dynamic modality weight selection network to first perform linear projection on the preliminary visible light features and the preliminary infrared features respectively to obtain the decoupled visible light features and the decoupled infrared features, then perform position encoding addition and modality weight selection based on the CoT attention mechanism on the decoupled visible light features and the decoupled infrared features to generate modality selection weights, and finally multiply the modality selection weights element-wise with the decoupled visible light features and the decoupled infrared features according to the modality to obtain the reselected visible light features and the reselected infrared features.

[0073] In a specific implementation, after the preliminary feature extraction is completed, the visible-light preliminary features and the infrared preliminary features will be fed into the pixel-level dynamic modality weight selection network for the most crucial step, namely, pixel-level dynamic feature selection. The pixel-level dynamic modality weight selection network first performs linear projections on the visible-light preliminary features and the infrared preliminary features respectively to obtain the visible-light decoupled features and the infrared decoupled features, then adds position encodings to the visible-light decoupled features and the infrared decoupled features, and performs modality weight selection based on the CoT attention mechanism to generate modality selection weights. Finally, the modality selection weights are multiplied element-wise with the visible-light decoupled features and the infrared decoupled features according to the modalities to obtain the visible-light reselected features and the infrared reselected features.

[0074] In one example, the specific structure of the pixel-level dynamic modality weight selection network is as Figure 4 shown, and it consists of a visible-light linear projection unit, an infrared linear projection unit, a visible-light position encoding unit, an infrared position encoding unit, and a modality weight selection unit.

[0075] The visible-light linear projection unit and the infrared linear projection unit are essentially 1×1 convolutional layers. The visible-light linear projection unit uses a 1×1 convolution to perform a linear projection to achieve a decoupling operation and obtain the visible-light decoupled features , and the infrared linear projection unit uses a 1×1 convolution to perform a linear projection to achieve a decoupling operation and obtain the infrared decoupled features . Such a decoupling operation can reduce the influence of weight selection on the original features.

[0076] After the linear projection is completed, and will be added position encoding vectors respectively to enhance the position representation of the features. The visible-light position encoding unit and the infrared position encoding unit will generate two sets of identical position encoding vectors. The position encoding vectors are grids formed with pixel blocks as the basic units, denoted as , . The visible-light position encoding unit will concatenate with through channels to generate the visible-light decoupled features after adding position encoding , and the visible-light position encoding unit will concatenate with

[0077] and can be represented by the formula:

[0078] ;

[0079] ;

[0080] Among them, represents the channel splicing operation.

[0081] The modal weight selection unit first projects and normalizes and to obtain the visible light projection normalized feature and the infrared projection normalized feature . Subsequently, based on the CoT attention mechanism, it processes and respectively to generate the visible light weight and the infrared weight . Then, the Sigmoid function is used to normalize and add and , and the Softmax function is used to generate the modal selection weight . Finally, based on , it multiplies element-wise with and respectively according to the modality to obtain the visible light reselected feature and the infrared reselected feature .

[0082] In one example, when generating the position encoding vector, first, a one-dimensional vector with a fixed interval needs to be generated. The -th element in

[0083] is expressed by the formula:

[0084] ;

[0085] Among them, is the preset window size, is The -th element in .

[0086] Subsequently, using the torch.meshgrid function, based on a two-dimensional grid is created. consists of two parts, namely the abscissa and the ordinate . The size of is expressed by the formula:

[0087] ;

[0088] Among them, represents the torch.meshgrid function.

[0089] Next, stack into a tensor with a shape of , , , represents the stacking operation.

[0090] After that, expand the dimension of , add a dimension to the zeroth dimension of , and obtain a tensor with a shape of , , , represents the dimension addition.

[0091] Finally, repeat to match the batch size of the input data, and finally obtain a position encoding vector with a shape of , , which is expressed by the formula as:

[0092] ;

[0093] Among them, is the batch size during training, is , 's height, is , 's width, represents the repetition operation.

[0094] In an example, the graph structure of the modal weight selection unit can be as shown in Figure 5 . The modal weight selection unit first performs projection mapping on and through a 1×1 convolutional layer, and then performs normalization processing through a LayerNorm layer to obtain the visible light projection normalized feature and the infrared projection normalized feature . and are expressed by the formula as:

[0095] ;

[0096] ;

[0097] Among them, represents a 1×1 convolution, represents layer normalization processing, is a commonly used normalization function. For the input vector of a certain layer in a neural network, layer normalization independently normalizes all dimensions of each sample, and then normalizes the input vector using the calculated mean and variance to obtain the normalized value. Similar to batch normalization, layer normalization also includes learnable scale parameters and translation parameters, which are used to scale and shift the normalized value respectively.

[0098] In one example, the structure of the CoT attention mechanism can be as Figure 6 shown.

[0099] For with a shape of , , first, a vector with the same shape of is obtained through the Key_Embed layer, and then is sent into the Value_Embed layer to obtain .

[0100] Next, and are concatenated in the channel dimension to obtain , and then the attention weights are generated through the Attention_Embed layer, the mean function layer, and the Softmax function layer.

[0101] Finally, is multiplied by to obtain the weighted value feature , and then is reshaped back to the original size and added to to obtain the final weight .

[0102] It should be noted that the Key_Embed layer, the Value_Embed layer, and the Attention_Embed layer are 3×3 convolution layer, 1×1 convolution layer, and 1×1 convolution layer respectively, is , the number of channels of.

[0103] In one example, the Sigmoid function is used to normalize and and then add them, and the Softmax function is used to generate the modality selection weight , which is expressed by the formula as:

[0104] ;

[0105] Among them, represents the Sigmoid function, represents the Softmax function, has a shape of , is the batch size during training, is the preset sequence length, 2 represents the weights corresponding to two modalities, the value is between 0 and 1, and the sum is 1. The larger the value, the higher the importance of the corresponding modality on this pixel-level sub-graph.

[0106] In one example, during the process of separately multiplying element-wise with and to obtain the visible light reselected feature and the infrared reselected feature , first and are sliced at the pixel level according to to obtain the visible light sequence and the infrared sequence , and are the same size as the shape of . During this slicing process, and are respectively divided into windows of size . The slicing process is as shown in Figure 7 .

[0107] Next, is regarded as the weight corresponding to the visible light modality, and is regarded as the weight corresponding to the infrared modality. Multiply element-wise with to obtain . Subsequently, multiply element-wise with to obtain .

[0108] and are expressed by the formula as:

[0109] ;

[0110] ;

[0111] Among them, represents element-wise multiplication.

[0112] Step 103: Use a deep network to perform deep feature extraction on the visible light reselected features and the infrared reselected features, successively obtaining visible light deep features and infrared deep features with resolutions of 1 / 8, 1 / 16, and 1 / 32 of the original resolution. Add the visible light deep features and the infrared deep features with the same resolution element by element to obtain three fusion features.

[0113] In a specific implementation, after completing pixel-level dynamic feature selection, the visible light reselected features and the infrared reselected features will enter a deep network for deep feature extraction, successively obtaining visible light deep features and infrared deep features with resolutions of 1 / 8, 1 / 16, and 1 / 32 of the original resolution. Add the visible light deep features and the infrared deep features with the same resolution element by element to obtain three fusion features.

[0114] In an example, the specific structure of the deep network can be as Figure 8 shown, consisting of two deep branches, namely the visible light deep extraction branch and the infrared deep extraction branch. Similar to the dual-branch convolutional neural network encoder, the structures of the two deep branches of the deep network are the same, both consisting of multiple ConvNormLayer (convolution layer and normalization layer) and multiple ResBlock (residual block).

[0115] The visible light deep extraction branch performs deep feature extraction at different scales on to successively obtain visible light deep features with resolutions of 1 / 8, 1 / 16, and 1 / 32 of the original resolution , and .

[0116] The infrared deep extraction branch performs deep feature extraction at different scales on to successively obtain infrared deep features with resolutions of 1 / 8, 1 / 16, and 1 / 32 of the original resolution , and .

[0117] Finally, add and , and , as well as and element by element respectively to successively obtain three fusion features , and .

[0118] , and It can be expressed by the formula:

[0119] ;

[0120] ;

[0121] ;

[0122] wherein, represents element-wise addition.

[0123] Step 104, input the three fused features into the detection head decoder for detection to obtain the target detection result.

[0124] In a specific implementation, the three fused features , and will enter the detection head decoder, and the detection head decoder will perform detection based on , and to obtain the target detection result.

[0125] In an example, the dual-branch convolutional neural network encoder, the pixel-level dynamic modality weight selection network, the deep network, and the detection head decoder together constitute a multi-source target detection model. When performing iterative training on the multi-source target detection model, the total loss function used is expressed by the formula:

[0126] ;

[0127] wherein, is the L1 loss, is the generalized intersection over union loss, is the classification loss, and are used to calculate the difference between the position of the predicted target and the position of the true target, is used to calculate the difference between the category of the predicted target and the category of the true target, represents the total loss function. The server performs backpropagation on the multi-source target detection model based on the total loss value calculated by the total loss function, and repeats the iteration until the model converges, so as to obtain the finally trained multi-source target detection model, and deploy it in the application scenarios where needed to complete the multi-source target detection task.

[0128] A multi-source object detection method based on pixel-level dynamic modality selection and fusion proposed in this embodiment uses a dual-branch convolutional neural network encoder, a pixel-level dynamic modality weight selection network, a deep network, and a detection head decoder to jointly complete the multi-source object detection task. The dual-branch convolutional neural network encoder can compress the input image and perform preliminary feature extraction, retain useful semantic information, and remove redundant information. These semantic information represent the global and local features in the scene, preparing for subsequent processing. The pixel-level dynamic modality weight selection network can dynamically select important regions in visible light images and infrared images according to different scenes, and better combine multi-modal features through pixel-level dynamic feature selection. After sending the reselected visible light features and reselected infrared features into the deep network, only deep feature extraction and simple addition are required to obtain multi-scale fusion features, and then the object detection results are output through the detection head decoder, greatly reducing the number of parameters of traditional fusion methods, and thus significantly improving the accuracy and speed of object detection in complex scenes.

[0129] The step division of the above various methods is only for clear description. When implemented, they can be combined into one step, or some steps can be split into multiple steps. As long as the same logical relationship is included, they are all within the protection scope of this application. Minor modifications added to the algorithm or process, or minor designs introduced, but without changing the core design of the algorithm and process, are all within the protection scope of this application.

[0130] In one embodiment, to verify the effectiveness of the multi-source object detection method based on pixel-level dynamic modality selection and fusion proposed in this application, we used the publicly available dataset FLIR to train and test the multi-source object detection model (abbreviated as Ours), and compared it with other methods. The FLIR dataset is a commonly used dataset in visible light and infrared multi-modal object detection. This dataset contains 4,129 pairs of training images and 1,013 pairs of test images.

[0131] The performance comparison between the traditional multi-source object detection model (algorithm) and Ours is as Figure 9As shown. Among many traditional models, including convolutional-based models and Transformer-based models, algorithms such as CFT and ICAFusion use the Transformer self-attention mechanism to interact and fuse visible light and infrared modalities, achieving certain results. However, compared to Ours, this method has a larger computational cost and more parameters. Methods such as GAFF that use simple convolutional operations for inter-modal and intra-modal fusion improve performance to a certain extent while minimizing the number of parameters, but the performance is still not good. The evaluation metric for multi-source object detection is mAP (Mean Average Precision), which is used to measure the average detection performance of the model across multiple classes. This is a comprehensive metric that can be understood as the average accuracy of the model across all classes. Depending on the different intersection-over-union ratios it determines, mAP50 and mAP50-95 are usually selected as performance metrics. map50 is the average of the AP values calculated at an intersection-over-union threshold of 0.50, and the mAP50-95 metric calculates the average of the AP values in the range of intersection-over-union thresholds from 0.50 to 0.95. In the convolutional-based multi-modal collaborative semantic segmentation model, Ours shows highly competitive results, with 79.5% mAP50 and 44% mAP50-95.

[0132] Another embodiment of the present application proposes an electronic device, the structure of which is as Figure 10 shown, including: at least one processor 201; and a memory 202 communicatively connected to the at least one processor 201; wherein, the memory 202 stores instructions executable by the at least one processor 201, and the instructions are executed by the at least one processor 201 to enable the at least one processor 201 to execute a multi-source object detection method based on pixel-level dynamic modality selection and fusion as described in the above method embodiment.

[0133] Among them, the memory and the processor are connected by a bus. The bus can include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and the memory together. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and will not be further described in this application. The bus interface provides an interface between the bus and the transceiver. The transceiver can be a single component or multiple components, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. The data processed by the processor is transmitted over a wireless medium through the antenna. Further, the antenna also receives data and transmits the data to the processor.

[0134] The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. The memory can be used to store the data used by the processor during operation.

[0135] Another embodiment of the present application provides a computer-readable storage medium storing a computer program, which when executed by a processor, can implement a multi-source target detection method based on pixel-level dynamic mode selection and fusion as described in the above method embodiment.

[0136] That is, those skilled in the art can understand that all or part of the steps in the above method embodiment can be completed by instructing relevant hardware through a program. The program is stored in a storage medium, including several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of a multi-source target detection method based on pixel-level dynamic mode selection and fusion as described in the method embodiment of the present application. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs and other various media that can store program codes.

[0137] Those of ordinary skill in the art can understand that the above embodiments are specific embodiments for implementing the present application, and in practical applications, various changes can be made in form and details without departing from the spirit and scope of the present application.

Claims

1. A multi-source target detection method based on pixel-level dynamic mode selection and fusion, characterized in that Including: Using a dual-branch convolutional neural network encoder to compress and perform preliminary feature extraction on paired visible light images and infrared images, obtaining visible light preliminary features and infrared preliminary features; Using a pixel-level dynamic modality weight selection network, first linearly project the visible light preliminary features and infrared preliminary features respectively to obtain visible light decoupled features and infrared decoupled features, then perform position encoding addition on the visible light decoupled features and infrared decoupled features, as well as modality weight selection based on the CoT attention mechanism to generate modality selection weights, and finally multiply the modality selection weights element-wise with the visible light decoupled features and infrared decoupled features respectively according to the modality to obtain visible light reselected features and infrared reselected features; Using a deep network to perform deep feature extraction on the visible light reselected features and infrared reselected features, successively obtaining visible light deep features and infrared deep features with resolutions of 1 / 8, 1 / 16, and 1 / 32 of the original resolution, and adding the visible light deep features and infrared deep features with the same resolution element-wise to obtain three fusion features; Inputting the three fusion features into a detection head decoder for detection to obtain target detection results; The two branches of the dual-branch convolutional neural network encoder are respectively a visible light preliminary extraction branch and an infrared preliminary extraction branch, and the structures of the two preliminary extraction branches are the same, both consisting of a series-connected convolutional neural network layer, a BN batch normalization layer, and a ReLU activation function layer; The visible light preliminary extraction branch compresses the visible light image and performs preliminary feature extraction to obtain the visible light preliminary features with a resolution of 1 / 2 of the visible light image ; The infrared preliminary extraction branch compresses and preliminarily extracts features from the infrared image, obtaining infrared preliminary features with a resolution of 1 / 2 of the infrared image. ; The pixel-level dynamic modality weight selection network consists of a visible light linear projection unit, an infrared linear projection unit, a visible light position encoding unit, an infrared position encoding unit, and a modality weight selection unit; The deep network consists of two deep branches, namely a visible light deep extraction branch and an infrared deep extraction branch, and the structures of the two deep branches are the same, both consisting of multiple ConvNormLayers and multiple ResBlocks; The dual-branch convolutional neural network encoder, the pixel-level dynamic modality weight selection network, the deep network, and the detection head decoder together constitute a multi-source target detection model. When performing iterative training on the multi-source target detection model, the total loss function used is expressed by the formula: ; Among them, is the L1 loss, is the generalized intersection over union loss, is the classification loss, and is used to calculate the difference between the predicted target position and the true target position, is used to calculate the difference between the predicted target category and the true target category, represents the total loss function.

2. The multi-source target detection method based on pixel-level dynamic mode selection and fusion according to claim 1, wherein The visible light linear projection unit uses 1×1 convolution to perform linear projection to achieve decoupling operation, and obtain visible light decoupled features , and the infrared linear projection unit uses 1×1 convolution to perform linear projection to achieve decoupling operation, and obtain infrared decoupled features ; The visible light position encoding unit and the infrared position encoding unit generate two sets of identical position encoding vectors. The position encoding vectors are grids formed with pixel blocks as the basic units, denoted as , . The visible light position encoding unit will and perform channel concatenation to generate the visible light decoupled feature after adding the position encoding . The visible light position encoding unit will and perform channel concatenation to generate the infrared decoupled feature after adding the position encoding ; The modal weight selection unit first projects and normalizes and to obtain the visible light projection normalized feature and the infrared projection normalized feature . Subsequently, based on the CoT attention mechanism, it processes and respectively to generate the visible light weight and the infrared weight . Then, it uses the Sigmoid function to normalize and and then adds them together, and uses the Softmax function to generate the modal selection weight . Finally, based on , it multiplies element by element with and respectively according to the modality to obtain the visible light reselected feature and the infrared reselected feature .

3. A multi-source target detection method based on pixel-level dynamic mode selection and fusion according to claim 2, characterized in that, The generation of the position encoding vector is achieved through the following steps: Generate a one-dimensional vector with a fixed interval , The th element in is expressed by the formula as follows: ; ; Among them, is the preset window size, is the th element in; Using the torch.meshgrid function, based on Create a two-dimensional grid , Consisting of two parts, namely the abscissa and the ordinate , The size of is , Expressed by the formula as: ; Among them, represents the torch.meshgrid function; Stack into a tensor with a shape of ; , , denotes the stacking operation; Expand the dimension, add a dimension in the zero dimension to obtain a tensor with a shape of ; ; Repeat to match the batch size of the input data, and finally obtain a positional encoding vector with the shape of , which is expressed by the formula as:​ ; Among them, is the batch size during training, is , the height of is , the width of represents a repeated operation.

4. A multi-source target detection method based on pixel-level dynamic mode selection and fusion according to claim 2, characterized in that, Pair and are subjected to projection mapping and normalization to obtain visible light projection normalized features and infrared projection normalized features , which are expressed by the formula as: ; ; Among them, represents a 1×1 convolution, represents layer normalization processing; Use the Sigmoid function to and perform normalization and then add them together, and use the Softmax function to generate the modality selection weights , which is expressed by the formula as: ; Among them, represents the Sigmoid function, represents the Softmax function, has a shape of , is the batch size during training, is the preset sequence length; Based on Multiply element by element with and respectively according to the modality to obtain the visible light re-selection feature and the infrared re-selection feature , including: Pair And According to Perform pixel-level segmentation to obtain the visible light sequence And the infrared sequence , And The size of corresponds to the shape of ; Regard as the weight corresponding to the visible light modality, and regard as the weight corresponding to the infrared modality. Multiply element-wise with to obtain . Multiply element-wise with to obtain .

5. A multi-source target detection method based on pixel-level dynamic modality selection and fusion according to claim 4, characterized in that, Based on the CoT attention mechanism, and Processing to generate visible light weights and infrared weight ,include: For the shape of of , , first obtain a vector with the same shape through the Key_Embed layer, and then send to the Value_Embed layer to obtain ; ; Concatenate and along the channel dimension to obtain , and then generate attention weights through the Attention_Embed layer, mean function layer, and Softmax function layer ; Multiply by to obtain the weighted value feature . Then reshape back to its original size and add it to to obtain the final weight ; Among them, the Key_Embed layer, Value_Embed layer, and Attention_Embed layer are 3×3 convolutional layer, 1×1 convolutional layer, and 1×1 convolutional layer respectively, is , the number of channels of.

6. The multi-source target detection method based on pixel-level dynamic mode selection and fusion according to claim 4, wherein Visible light deep extraction branch pair Performs deep feature extraction at different scales, and sequentially obtains visible light deep features with resolutions of 1 / 8, 1 / 16, and 1 / 32 of the original resolution , and ; Infrared deep extraction branch pair Performs deep feature extraction at different scales, and successively obtains infrared deep features with resolutions of 1 / 8, 1 / 16, and 1 / 32 of the original resolution , and ; Add with , with , and with , respectively perform element-by-element addition to obtain three fused features , and .

7. An electronic device, characterized in that, Including: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute a multi-source target detection method according to any one of claims 1 to 6.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it can implement a multi-source target detection method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Infrared and visible light image fusion system and method

    CN114187214A

  • Transform-based dual-band image semantic segmentation method

    CN116469100A