Multi-source target detection method based on pixel-level dynamic mode selection fusion

By introducing pixel-level dynamic modal selection fusion technology into the multi-source object detection method, using dual-branch convolutional neural network and dynamic modal weight selection network, the problem of poor performance of existing methods in complex scenarios is solved, and high-precision and high-speed object detection is achieved.

CN120107568AActive Publication Date: 2025-06-06NORTHWESTERN POLYTECHNICAL UNIV

Patent Information

Application Number
CN202510587275.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-06-06
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

The existing multimodal object detection methods have poor performance in complex scenarios, especially in conditions such as insufficient light and haze, making it difficult to achieve high-precision and high-speed object detection.

Method used

A multi-source object detection method based on pixel-level dynamic modal selection fusion is proposed. A dual-branch convolutional neural network encoder and pixel-level dynamic modal weight selection network is used to perform deep feature extraction of visible light and infrared images and pixel-level dynamic fusion, and generate multi-scale fusion features to improve object detection performance.

Benefits of technology

Through the pixel-level dynamic modal selection fusion method, the target detection accuracy and speed in complex scenarios are significantly improved, and the calculation amount and parameter amount in traditional methods are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107568A_ABST
    Figure CN120107568A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and discloses a multi-source target detection method based on pixel-level dynamic modal selection fusion, comprising the following steps: using a double-branch convolutional neural network encoder to perform preliminary feature extraction on a visible light image and an infrared image to obtain a visible light preliminary feature and an infrared preliminary feature; using a pixel-level dynamic modal weight selection network to respectively carry out linear projection, position code addition and modal weight selection based on a CoT attention mechanism on the visible light preliminary feature and the infrared preliminary feature, generating a modal selection weight, and finally carrying out pixel-level dynamic feature selection based on the modal selection weight. Obtaining visible light reselection features and infrared reselection features; performing deep feature extraction on the visible light reselection features and the infrared reselection features by using a deep network, and adding elements one by one according to the resolution to obtain fusion features; and inputting the fusion feature into a detection head decoder for detection to obtain a target detection result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and in particular to a multi-source target detection method based on pixel-level dynamic modality selection fusion. Background Art

[0002] Object detection is one of the basic tasks of computer vision technology, which aims to detect objects of interest in a given image and give the corresponding location and category information. Commonly used visible light object detection algorithms are difficult to accurately detect objects in scenes with insufficient lighting (such as at night, on cloudy days), foggy weather, and certain occlusions. Based on this, some researchers have tried to combine visible light images with infrared images, and improve the performance of object detection under complex conditions such as weak light by combining the rich color and texture detail information of visible light images with the characteristics of infrared images that are not affected by lighting conditions.

[0003] In existing research, technicians usually use channel splicing-based methods or attention-based methods to deal with the fusion problem of multimodal feature maps. Through this single convolutional neural network or a single attention mechanism, certain effects can be achieved. For example, Fang et al. fused visible light features and infrared features into a unified representation through a carefully designed fusion module to make full use of the visible light features and infrared features extracted by two parallel backbone networks.

[0004] However, in practical applications, the channel splicing-based methods and attention mechanism-based methods still perform poorly in many complex scenarios, mainly due to the following problems.

[0005] First, the scenes of visible light images and infrared images are complex and changeable. It is more scientific to use different feature selection strategies for different scenes. A single feature fusion method is difficult to meet the changing needs.

[0006] Second, most fusion algorithms are performed before the multi-scale feature pyramid module. The fusion module needs to be reused at multiple scales, which requires a large amount of calculation and parameters, and there may be inconsistencies in the fusion of different scales.

[0007] Third, traditional algorithms often use channel attention and spatial attention to perform feature fusion between visible light modalities and infrared modalities. The granularity of this fusion is often large and it is impossible to perform fine-grained fusion of targets in the scene. Summary of the invention

[0008] In order to solve the above technical problems, an embodiment of the present application proposes a multi-source target detection method based on pixel-level dynamic modality selection fusion, which aims to extract deep features of visible light modality and infrared modality through multi-source collaborative learning, and perform pixel-level dynamic fusion, so as to make full use of useful information in different modality images to improve the accuracy and speed of target detection in complex scenes.

[0009] In order to achieve the above-mentioned purpose, the embodiment of the present application proposes a multi-source target detection method based on pixel-level dynamic modality selection fusion, comprising the following steps: using a dual-branch convolutional neural network encoder to compress and preliminarily extract features of paired visible light images and infrared images to obtain visible light preliminary features and infrared preliminary features; using a pixel-level dynamic modality weight selection network to first linearly project the visible light preliminary features and the infrared preliminary features respectively to obtain visible light decoupling features and infrared decoupling features, and then position encoding and adding the visible light decoupling features and the infrared decoupling features, and based on CoT (Contextual Transformer (context transformer) attention mechanism selects the modality weights to generate modality selection weights, and finally multiplies the modality selection weights with the visible light decoupling features and the infrared decoupling features element by element according to the modality to obtain the visible light reselection features and the infrared reselection features; a deep network is used to perform deep feature extraction on the visible light reselection features and the infrared reselection features, and visible light deep features and infrared deep features with resolutions of 1 / 8, 1 / 16 and 1 / 32 of the original resolution are obtained respectively, and the visible light deep features and infrared deep features with the same resolution are added element by element to obtain three fused features; the three fused features are input into the detection head decoder for detection to obtain the target detection result.

[0010] In order to achieve the above-mentioned purpose, an embodiment of the present application also proposes an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a multi-source target detection method based on pixel-level dynamic modality selection fusion as described above.

[0011] In order to achieve the above-mentioned purpose, an embodiment of the present application also proposes a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it can implement a multi-source target detection method based on pixel-level dynamic modality selection fusion as described above.

[0012] The embodiment of the present application proposes a multi-source target detection method based on pixel-level dynamic modality selection fusion, which uses a dual-branch convolutional neural network encoder, a pixel-level dynamic modality weight selection network, a deep network and a detection head decoder to jointly complete the multi-source target detection task. The dual-branch convolutional neural network encoder can compress and preliminarily extract features of the input image, thereby retaining useful semantic information and removing redundant information. These semantic information represents the global and local features in the scene, and is ready for subsequent processing. The pixel-level dynamic modality weight selection network can dynamically select important areas in visible light images and infrared images according to different scenes, and better combine multi-modal features through pixel-level dynamic feature selection. After the visible light reselection features and infrared reselection features are sent to the deep network, only deep feature extraction and simple addition are required to obtain multi-scale fusion features, and then the output of the target detection results is realized through the detection head decoder, which greatly reduces the number of parameters of the traditional fusion method, thereby greatly improving the accuracy and speed of target detection in complex scenes.

[0013] In some optional embodiments, the two branches of the dual-branch convolutional neural network encoder are respectively a visible light preliminary extraction branch and an infrared preliminary extraction branch, and the two branches have the same structure, both consisting of a series of convolutional neural network layers, BN batch normalization layers, and ReLU activation function layers; The visible light preliminary extraction branch compresses the visible light image and performs preliminary feature extraction to obtain the preliminary visible light features with a resolution of 1 / 2 of the visible light image. ; The infrared preliminary extraction branch compresses the infrared image and extracts preliminary features to obtain infrared preliminary features with a resolution of 1 / 2 of the infrared image. .

[0014] In some optional embodiments, the pixel-level dynamic modal weight selection network is composed of a visible light linear projection unit, an infrared linear projection unit, a visible light position encoding unit, an infrared position encoding unit, and a modal weight selection unit; The visible light linear projection unit uses a 1×1 convolution to Linear projection is performed to achieve decoupling operation and obtain visible light decoupling features , the infrared linear projection unit uses a 1×1 convolution Perform linear projection to achieve decoupling operation and obtain infrared decoupling features ; The visible light position coding unit and the infrared position coding unit generate two sets of identical position coding vectors. The position coding vector is a grid composed of pixel blocks as the basic unit, denoted as , , the visible light position encoding unit will and Perform channel stitching to generate visible light decoupling features after adding position coding , the visible light position encoding unit will and Perform channel stitching to generate infrared decoupling features after adding position coding ; The modal weight selection unit first selects and Perform projection mapping and normalization processing to obtain visible light projection normalization features and infrared projection normalized features , and then based on the CoT attention mechanism, and Processing to generate visible light weights and infrared weight , and then use the Sigmoid function to and After normalization, add them together and use the Softmax function to generate the modality selection weights , and finally based on According to the mode, and Perform element-by-element multiplication to obtain the visible light reselection feature and infrared reselection characteristics .

[0015] In some optional embodiments, the generation of the position encoding vector is achieved by the following steps: Generate a fixed-spaced one-dimensional vector , The The elements are expressed by the formula: ; ; in, is the default window size, for The elements; Using torch.mershgrid function, based on Create a 2D grid , It consists of two parts: the horizontal axis and the vertical coordinate , The size is , It is expressed by the formula: ; in, Represents the torch.mershgrid function; Will Stacked into a shape Tensor , , Indicates stacking operation; expansion Add a dimension to the zeroth dimension to get a shape of Tensor ; repeat To match the batch size of the input data, the final shape is The position encoding vector , It is expressed by the formula: ; in, is the batch size during training, for , Height, for , The width of Indicates a repeated operation.

[0016] In some optional embodiments, and Perform projection mapping and normalization processing to obtain visible light projection normalization features and infrared projection normalized features , expressed by the formula: ; ; in, represents 1×1 convolution, Representation layer normalization processing; Use the Sigmoid function to and After normalization, add them together and use the Softmax function to generate the modality selection weights , expressed by the formula: ; in, represents the Sigmoid function, represents the Softmax function, The shape is , is the batch size during training, is the preset sequence length; based on According to the mode, and Perform element-by-element multiplication to obtain the visible light reselection feature and infrared reselection characteristics ,include: right and according to Perform pixel-level segmentation to obtain visible light sequence and infrared sequences , and The size and The shape corresponds to; Will is regarded as the weight corresponding to the visible light mode, is regarded as the weight corresponding to the infrared mode, and Multiplying element by element, we get ,Will and Multiplying element by element, we get .

[0017] In some optional embodiments, based on the CoT attention mechanism, and Processing to generate visible light weights and infrared weight ,include: For the shape of , First, we use the Key_Embed layer to get the same shape as Vector , and then Send it to the Value_Embed layer to get ; Will and Concatenate in the channel dimension to get , and then generate attention weights through the Attention_Embed layer, mean function layer and Softmax function layer ; Will and Multiply them together to get the weighted feature , and then Reshape back to original size and Add together to get the final weight ; Among them, the Key_Embed layer, Value_Embed layer, and Attention_Embed layer are 3×3 convolutional layer, 1×1 convolutional layer, and 1×1 convolutional layer respectively. for , The number of channels.

[0018] In some optional embodiments, the deep network is composed of two deep branches, namely a visible light deep extraction branch and an infrared deep extraction branch, and the two deep branches have the same structure, both consisting of multiple ConvNormLayers and multiple ResBlocks; Visible light deep extraction branch pair Perform deep feature extraction at different scales to obtain visible light deep features with resolutions of 1 / 8, 1 / 16, and 1 / 32 of the original resolution. , and ; Infrared deep extraction branch pair Perform deep feature extraction at different scales to obtain infrared deep features with resolutions of 1 / 8, 1 / 16, and 1 / 32 of the original resolution. , and ; Will and , and ,as well as and , respectively, and add them element by element to obtain three fusion features in turn , and .

[0019] In some optional embodiments, a dual-branch convolutional neural network encoder, a pixel-level dynamic modal weight selection network, a deep network, and a detection head decoder together constitute a multi-source target detection model. When iteratively training the multi-source target detection model, the total loss function used is expressed by the formula: ; in, is the L1 loss, is the generalized intersection-combination loss, is the classification loss, and Used to calculate the difference between the predicted target position and the actual target position. Used to calculate the difference between the predicted target category and the true target category. Represents the total loss function. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the related technologies, the drawings required for use in the embodiments of the present application or the related technical descriptions will be briefly introduced below. The following drawings are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. The drawings described here are only used to explain the present application and are not used to limit the present application.

[0021] Figure 1 is a flowchart of a multi-source target detection method based on pixel-level dynamic modality selection fusion provided in one embodiment of the present application; Figure 2 is a schematic diagram of the structure of a multi-source target detection model provided in an embodiment of the present application; Figure 3 is a schematic diagram of the structure of a preliminary extraction branch provided in one embodiment of the present application; Figure 4 is a schematic diagram of the structure of a pixel-level dynamic modality weight selection network provided in one embodiment of the present application; Figure 5 is a structural schematic diagram of a modal weight selection unit provided in one embodiment of the present application; Figure 6 is a schematic diagram of the structure of the CoT attention mechanism provided in one embodiment of the present application; Figure 7 is a schematic diagram of pixel-level segmentation provided in one embodiment of the present application; Figure 8 is a schematic diagram of the structure of a deep network provided in one embodiment of the present application; Fig. 9 is a schematic diagram of a performance comparison between a traditional multi-source target detection algorithm provided in an embodiment of the present application and the method proposed in the present application; Fig.10 It is a structural schematic diagram of an electronic device provided by another embodiment of the present application. DETAILED DESCRIPTION

[0022] In order to make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the embodiments of the present application will be described in detail below in conjunction with the accompanying drawings. It will be appreciated by those skilled in the art that in the embodiments of the present application, many technical details are proposed in order to enable the reader to better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical scheme claimed in the present application can also be implemented. The division of the following embodiments is only for the convenience of description, and the specific implementation of the present application should not constitute any limitation, and the various embodiments can be combined and referenced to each other under the premise of no contradiction.

[0023] In order to solve the problem of poor performance of traditional multi-source target detection methods in complex scenarios, an embodiment of the present application proposes a multi-source target detection method based on pixel-level dynamic modal selection fusion, which is applied to a server. The implementation details of the multi-source target detection method based on pixel-level dynamic modal selection fusion proposed in this embodiment are described in detail below. The following content is only the implementation details provided for the convenience of understanding and is not necessary for the implementation of this solution.

[0024] The specific process of the multi-source target detection method based on pixel-level dynamic modality selection fusion proposed in this embodiment can be as follows: Figure 1 As shown, including: Step 101, using a dual-branch convolutional neural network encoder to compress and perform preliminary feature extraction on a pair of visible light images and infrared images to obtain preliminary visible light features and preliminary infrared features.

[0025] In the specific implementation, the dual-branch convolutional neural network encoder, the pixel-level dynamic modal weight selection network, the deep network and the detection head decoder together constitute the following Figure 2 The multi-source target detection model shown in FIG. 1 is inputted by a pair of visible light images and infrared images, and outputted by a target detection result. The pair of visible light images and infrared images are inputted into a dual-branch convolutional neural network encoder for compression and preliminary feature extraction, thereby obtaining preliminary visible light features and preliminary infrared features. The preliminary visible light features and preliminary infrared features obtained after compression and preliminary feature extraction carry more semantic information, which represents the global and local features in the detection scene, and is ready for subsequent detection operations.

[0026] In one example, the two branches of the dual-branch convolutional neural network encoder are the visible light preliminary extraction branch and the infrared preliminary extraction branch. The two preliminary extraction branches have the same structure, both of which are ConvNormLayer, that is, they are composed of a series of convolutional neural network layers, BN batch normalization layers, and ReLU activation function layers. The specific structure of the preliminary extraction branch can be as follows: Figure 3The visible light preliminary extraction branch compresses the visible light image and performs preliminary feature extraction to obtain a preliminary visible light feature with a resolution of 1 / 2 of the visible light image. The infrared preliminary extraction branch compresses the infrared image and extracts preliminary features to obtain infrared preliminary features with a resolution of 1 / 2 of the infrared image. .

[0027] Step 102, using a pixel-level dynamic modal weight selection network, first linearly project the visible light preliminary features and the infrared preliminary features respectively to obtain visible light decoupled features and infrared decoupled features, then perform position encoding and addition on the visible light decoupled features and the infrared decoupled features, and select modal weights based on the CoT attention mechanism to generate modal selection weights, and finally multiply the modal selection weights element-by-element with the visible light decoupled features and the infrared decoupled features according to the modality to obtain visible light reselected features and infrared reselected features.

[0028] In the specific implementation, after completing the preliminary feature extraction, the preliminary visible light features and the preliminary infrared features will be sent to the pixel-level dynamic modal weight selection network for the most critical step, namely, pixel-level dynamic feature selection. The pixel-level dynamic modal weight selection network first performs linear projection on the preliminary visible light features and the preliminary infrared features to obtain the visible light decoupling features and the infrared decoupling features, then performs position encoding and addition on the visible light decoupling features and the infrared decoupling features, and performs modal weight selection based on the CoT attention mechanism to generate modal selection weights. Finally, the modal selection weights are element-wise multiplied with the visible light decoupling features and the infrared decoupling features according to the modality to obtain the visible light reselection features and the infrared reselection features.

[0029] In an example, the specific structure of the pixel-level dynamic modality weight selection network is as follows: Figure 4 As shown, it consists of a visible light linear projection unit, an infrared linear projection unit, a visible light position encoding unit, an infrared position encoding unit and a modal weight selection unit.

[0030] The visible light linear projection unit and the infrared linear projection unit are essentially 1×1 convolutional layers. The visible light linear projection unit uses 1×1 convolution to Linear projection is performed to achieve decoupling operation and obtain visible light decoupling features , the infrared linear projection unit uses a 1×1 convolution Perform linear projection to achieve decoupling operation and obtain infrared decoupling features Such decoupling operations can reduce the impact of weight selection on the original features.

[0031] After completing the linear projection, and Position encoding vectors will be added to enhance the position representation of features. The visible light position encoding unit and the infrared position encoding unit will generate two sets of identical position encoding vectors. The position encoding vector is a grid composed of pixel blocks as the basic unit, denoted as , , the visible light position encoding unit will and Channel stitching to generate visible light decoupled features after position encoding is added , the visible light position encoding unit will and Channel stitching to generate infrared decoupled features after position encoding is added .

[0032] and It can be expressed by the formula: ; ; in, Represents a channel concatenation operation.

[0033] The modal weight selection unit first selects and Perform projection mapping and normalization processing to obtain visible light projection normalization features and infrared projection normalized features , and then based on the CoT attention mechanism, and Processing to generate visible light weights and infrared weight , and then use the Sigmoid function to and After normalization, add them together and use the Softmax function to generate the modality selection weights , and finally based on According to the mode, and Perform element-by-element multiplication to obtain the visible light reselection feature and infrared reselection characteristics .

[0034] In one example, when generating a position encoding vector, it is first necessary to generate a one-dimensional vector with a fixed interval , The The elements are expressed by the formula: ; ; in, is the default window size, for The Elements, general settings .

[0035] Then, using the torch.mershgrid function, based on Create a 2D grid , It consists of two parts: the horizontal axis and the vertical coordinate , The size is , It is expressed by the formula: ; in, Represents the torch.mershgrid function.

[0036] Next, Stacked into a shape Tensor , , Indicates a stacking operation.

[0037] Then, expansion The dimension, in Add a dimension to the zeroth dimension of Tensor , , Indicates adding a dimension.

[0038] Finally, repeat To match the batch size of the input data, the final shape is The position encoding vector , It is expressed by the formula: ; in, is the batch size during training, for , Height, for , The width of Indicates a repeated operation.

[0039] In one example, the graph structure of the modality weight selection unit can be as follows Figure 5 As shown in the figure, the modality weight selection unit first passes through a 1×1 convolutional layer to and Perform projection mapping, and then perform normalization through a LayerNorm layer to obtain the visible light projection normalization feature and infrared projection normalized features . and It is expressed by the formula: ; ; in, represents 1×1 convolution, Representation layer normalization processing, It is a commonly used normalization function. For the input vector of a layer in a neural network, layer normalization independently normalizes all dimensions of each sample. Then, the input vector is normalized using the calculated mean and variance to obtain the normalized value. Similar to batch normalization, layer normalization also includes learnable scale parameters and translation parameters, which are used to scale and move the normalized values, respectively.

[0040] In an example, the structure of the CoT attention mechanism can be as follows Figure 6 shown.

[0041] For the shape of , First, we use the Key_Embed layer to get the same shape as Vector , and then Send it to the Value_Embed layer to get .

[0042] Next, and Concatenate in the channel dimension to get , and then generate attention weights through the Attention_Embed layer, mean function layer and Softmax function layer .

[0043] Finally, and Multiply them together to get the weighted feature , and then Reshape back to original size and Add together to get the final weight .

[0044] It should be noted that the Key_Embed layer, Value_Embed layer, and Attention_Embed layer are 3×3 convolutional layers, 1×1 convolutional layers, and 1×1 convolutional layers, respectively. for , The number of channels.

[0045] In an example, using the Sigmoid function and After normalization, add them together and use the Softmax function to generate the modality selection weights , expressed by the formula: ; in, represents the Sigmoid function, represents the Softmax function, The shape is , is the batch size during training, is the preset sequence length, 2 represents the weights corresponding to the two modes, the value is between 0 and 1, and the sum is 1. The larger the value, the higher the importance of the corresponding mode on this pixel-level sub-image.

[0046] In one example, based on According to the mode, and Perform element-by-element multiplication to obtain the visible light reselection feature and infrared reselection characteristics In the process, first and according to Perform pixel-level segmentation to obtain visible light sequence and infrared sequences , and The size and The shape corresponds to that of and Divide into indivual The size of the window is as follows: Figure 7 shown.

[0047] Next, is regarded as the weight corresponding to the visible light mode, and is regarded as the weight corresponding to the infrared mode, and Performing element-wise multiplication, we get , and then and Performing element-wise multiplication, we get .

[0048] and It is expressed by the formula: ; ; in, Represents element-wise multiplication.

[0049] Step 103, use a deep network to perform deep feature extraction on the visible light reselection features and the infrared reselection features, and obtain visible light deep features and infrared deep features with resolutions of 1 / 8, 1 / 16 and 1 / 32 of the original resolution respectively, and add the visible light deep features and infrared deep features with the same resolution element by element to obtain three fused features.

[0050] In the specific implementation, after completing the pixel-level dynamic feature selection, the visible light reselected features and the infrared reselected features will enter the deep network for deep feature extraction, and the visible light deep features and infrared deep features with resolutions of 1 / 8, 1 / 16 and 1 / 32 of the original resolution will be obtained respectively. The visible light deep features and infrared deep features with the same resolution will be added element by element to obtain three fused features.

[0051] In an example, the specific structure of the deep network can be as follows Figure 8 As shown in the figure, it consists of two deep branches, namely the visible light deep extraction branch and the infrared deep extraction branch. Similar to the dual-branch convolutional neural network encoder, the structures of the two deep branches of the deep network are also the same, both consisting of multiple ConvNormLayer (convolutional layer and normalization layer) and multiple ResBlock (residual block).

[0052] Visible light deep extraction branch pair Perform deep feature extraction at different scales to obtain visible light deep features with resolutions of 1 / 8, 1 / 16, and 1 / 32 of the original resolution. , and .

[0053] Infrared deep extraction branch pair Perform deep feature extraction at different scales to obtain infrared deep features with resolutions of 1 / 8, 1 / 16, and 1 / 32 of the original resolution. , and .

[0054] Finally, and , and ,as well as and , respectively, and add them element by element to obtain three fusion features in turn , and .

[0055] , and It can be expressed by the formula: ; ; ; in, Represents element-by-element addition.

[0056] Step 104, input the three fused features into the detection head decoder for detection to obtain the target detection result.

[0057] In the specific implementation, three fusion features , and It will enter the detection head decoder, which will be based on the detection head decoder , and Detection is performed to obtain the target detection result.

[0058] In an example, a dual-branch convolutional neural network encoder, a pixel-level dynamic modal weight selection network, a deep network, and a detection head decoder together constitute a multi-source target detection model. When iteratively training the multi-source target detection model, the total loss function used is expressed by the formula: ; in, is the L1 loss, is the generalized intersection-combination loss, is the classification loss, and Used to calculate the difference between the predicted target position and the actual target position. Used to calculate the difference between the predicted target category and the true target category. Represents the total loss function. The server performs backpropagation on the multi-source target detection model based on the total loss value calculated by the total loss function, and repeats the iteration until the model converges, thereby obtaining the final trained multi-source target detection model, which is deployed in the application scenario in need to complete the multi-source target detection task.

[0059] The present embodiment proposes a multi-source target detection method based on pixel-level dynamic modality selection fusion, which uses a dual-branch convolutional neural network encoder, a pixel-level dynamic modality weight selection network, a deep network and a detection head decoder to jointly complete the multi-source target detection task. The dual-branch convolutional neural network encoder can compress and preliminarily extract features of the input image, retain useful semantic information, and remove redundant information. These semantic information represents the global and local features in the scene, and prepares for subsequent processing. The pixel-level dynamic modality weight selection network can dynamically select important areas in visible light images and infrared images according to different scenes, and better combine multi-modal features through pixel-level dynamic feature selection. After the visible light reselection features and infrared reselection features are sent to the deep network, only deep feature extraction and simple addition are required to obtain multi-scale fusion features, and then the output of the target detection results is realized through the detection head decoder, which greatly reduces the number of parameters of the traditional fusion method, thereby greatly improving the accuracy and speed of target detection in complex scenes.

[0060] The steps of the above methods are divided only for the purpose of clear description. They can be combined into one step or some steps can be split and decomposed into multiple steps during implementation. As long as they include the same logical relationship, they are all within the scope of protection of this application. Insignificant modifications added to the algorithm or process or insignificant designs introduced, but do not change the core design of the algorithm and process, are all within the scope of protection of this application.

[0061] In one embodiment, in order to verify the effectiveness of the multi-source target detection method based on pixel-level dynamic modality selection fusion proposed in this application, we used the public dataset FLIR to train and test the multi-source target detection model (abbreviated as Ours), and compared it with other methods. The FLIR dataset is a commonly used dataset in visible light infrared multi-modal target detection, which contains 4129 pairs of training images and 1013 pairs of test images.

[0062] The performance comparison between the traditional multi-source target detection model (algorithm) and Ours is as follows Fig. 9As shown. Among many traditional models, including convolution-based models and Transformer-based models, algorithms such as CFT and ICAFusion use the Transformer self-attention mechanism to interactively fuse visible light and infrared modalities, achieving certain results, but compared with Ours, this method has a large amount of calculation and a large number of parameters. Methods such as GAFF that use simple convolution operations to perform inter-modal and intra-modal fusion have improved performance to a certain extent while minimizing the number of parameters, but the performance is still poor. The evaluation index of multi-source target detection is mAP (Mean Average Precision), which is used to measure the average detection performance of the model on multiple categories. This is a comprehensive index, which can be understood as the average accuracy of the model on all categories. According to the different intersection-over-union ratios determined, mAP50 and mAP50-95 are usually selected as performance indicators. Map50 is the average value of the AP value calculated when the intersection-over-union ratio threshold is 0.50, and the mAP50-95 index calculates the average value of the AP value in the intersection-over-union ratio threshold range from 0.50 to 0.95. Ours shows highly competitive results in convolution-based multimodal cooperative semantic segmentation models, with 79.5% mAP50 and 44% mAP50-95.

[0063] Another embodiment of the present application provides an electronic device, the structure of which is as follows: Fig.10 As shown, it includes: at least one processor 201; and a memory 202 that is communicatively connected to the at least one processor 201; wherein the memory 202 stores instructions that can be executed by the at least one processor 201, and the instructions are executed by the at least one processor 201 so that the at least one processor 201 can execute a multi-source target detection method based on pixel-level dynamic modality selection fusion as described in the above method embodiment.

[0064] Among them, the memory and the processor are connected in a bus manner, and the bus may include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and memories together. The bus can also connect various other circuits such as peripherals, voltage regulators, and power management circuits, which are well known in the art and will not be further described in this application. The bus interface provides an interface between the bus and the transceiver. The transceiver can be one element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices on a transmission medium. The data processed by the processor is transmitted on a wireless medium through an antenna, and further, the antenna also receives data and transmits the data to the processor.

[0065] The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory can be used to store data used by the processor when performing operations.

[0066] Another embodiment of the present application proposes a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it can implement a multi-source target detection method based on pixel-level dynamic modality selection fusion as described in the above method embodiment.

[0067] That is, those skilled in the art can understand that all or part of the steps in the above method embodiment can be completed by instructing the relevant hardware through a program, and the program is stored in a storage medium, including a number of instructions to enable a device (which can be a single-chip microcomputer, chip, etc.) or a processor (processor) to execute all or part of the steps of a multi-source target detection method based on pixel-level dynamic mode selection fusion described in the method embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as: ROM), random access memory (Random Access Memory, referred to as: RAM), disk or optical disk and other media that can store program codes.

[0068] Those skilled in the art will appreciate that the above embodiments are specific embodiments for implementing the present application, and in actual applications, various changes may be made thereto in form and detail without departing from the spirit and scope of the present application.

Claims

1. A multi-source target detection method based on pixel-level dynamic modality selection fusion, characterized in that: include: A dual-branch convolutional neural network encoder is used to compress and perform preliminary feature extraction on paired visible light images and infrared images to obtain preliminary visible light features and preliminary infrared features. Using a pixel-level dynamic modal weight selection network, first linearly project the visible light preliminary features and infrared preliminary features to obtain visible light decoupling features and infrared decoupling features, then perform position encoding and add the visible light decoupling features and infrared decoupling features, and select the modal weight based on the CoT attention mechanism to generate the modal selection weight. Finally, the modal selection weight is element-wise multiplied with the visible light decoupling features and infrared decoupling features according to the modality to obtain the visible light reselection features and infrared reselection features. A deep network is used to perform deep feature extraction on visible light reselection features and infrared reselection features, and visible light deep features and infrared deep features with resolutions of 1 / 8, 1 / 16, and 1 / 32 of the original resolution are obtained respectively. The visible light deep features and infrared deep features with the same resolution are added element by element to obtain three fusion features; The three fused features are input into the detection head decoder for detection to obtain the target detection result.

2. According to claim 1, a multi-source target detection method based on pixel-level dynamic mode selection fusion is characterized in that: The two branches of the dual-branch convolutional neural network encoder are the visible light preliminary extraction branch and the infrared preliminary extraction branch. The two preliminary extraction branches have the same structure, which consists of a series of convolutional neural network layers, BN batch normalization layers, and ReLU activation function layers. The visible light preliminary extraction branch compresses the visible light image and performs preliminary feature extraction to obtain the visible light preliminary features with a resolution of 1 / 2 of the visible light image. ; The infrared preliminary extraction branch compresses the infrared image and extracts preliminary features to obtain infrared preliminary features with a resolution of 1 / 2 of the infrared image. .

3. The multi-source target detection method based on pixel-level dynamic modality selection fusion according to claim 2 is characterized in that: The pixel-level dynamic modal weight selection network consists of a visible light linear projection unit, an infrared linear projection unit, a visible light position encoding unit, an infrared position encoding unit, and a modal weight selection unit; The visible light linear projection unit uses a 1×1 convolution to Linear projection is performed to achieve decoupling operation and obtain visible light decoupling features , the infrared linear projection unit uses a 1×1 convolution Perform linear projection to achieve decoupling operation and obtain infrared decoupling features ; The visible light position coding unit and the infrared position coding unit generate two sets of identical position coding vectors. The position coding vector is a grid composed of pixel blocks as the basic unit, denoted as , , the visible light position encoding unit will and Perform channel stitching to generate visible light decoupling features after adding position coding , the visible light position encoding unit will and Perform channel stitching to generate infrared decoupling features after adding position coding ; The modal weight selection unit first selects and Perform projection mapping and normalization processing to obtain visible light projection normalization features and infrared projection normalized features , and then based on the CoT attention mechanism and Processing to generate visible light weights and infrared weight , and then use the Sigmoid function to and After normalization, add them together and use the Softmax function to generate the modality selection weights , and finally based on According to the mode, and Perform element-by-element multiplication to obtain the visible light reselection feature and infrared reselection characteristics .

4. The multi-source target detection method based on pixel-level dynamic modality selection fusion according to claim 3 is characterized in that: The generation of the position encoding vector is achieved through the following steps: Generate a fixed-spaced one-dimensional vector , The The elements are expressed by the formula: ; ; in, is the default window size, for The elements; Using torch.mershgrid function, based on Create a 2D grid , It consists of two parts: the horizontal axis and the vertical coordinate , The size is , It is expressed by the formula: ; in, Represents the torch.mershgrid function; Will Stacked into a shape Tensor , , Indicates stacking operation; expansion Add a dimension to the zeroth dimension to get a shape of Tensor ; repeat To match the batch size of the input data, the final shape is The position encoding vector , It is expressed by the formula: ; in, is the batch size during training, for , Height, for , The width of Indicates a repeated operation.

5. The multi-source target detection method based on pixel-level dynamic modality selection fusion according to claim 3 is characterized in that: right and Perform projection mapping and normalization processing to obtain visible light projection normalization features and infrared projection normalized features , expressed by the formula: ; ; in, represents 1×1 convolution, Representation layer normalization processing; Use the Sigmoid function to and After normalization, add them together and use the Softmax function to generate the modality selection weights , expressed by the formula: ; in, represents the Sigmoid function, represents the Softmax function, The shape is , is the batch size during training, is the preset sequence length; based on According to the mode, and Perform element-by-element multiplication to obtain the visible light reselection feature and infrared reselection characteristics ,include: right and according to Perform pixel-level segmentation to obtain visible light sequence and infrared sequences , and The size and The shape corresponds to; Will is regarded as the weight corresponding to the visible light mode, is regarded as the weight corresponding to the infrared mode, and Multiplying element by element, we get ,Will and Multiplying element by element, we get .

6. The multi-source target detection method based on pixel-level dynamic modality selection fusion according to claim 5 is characterized in that: Based on the CoT attention mechanism, and Processing to generate visible light weights and infrared weight ,include: For the shape of , First, we use the Key_Embed layer to get the same shape as Vector , and then Send it to the Value_Embed layer to get ; Will and Concatenate in the channel dimension to get , and then generate attention weights through the Attention_Embed layer, mean function layer and Softmax function layer ; Will and Multiply them together to get the weighted feature , and then Reshape back to original size and Add together to get the final weight ; Among them, the Key_Embed layer, Value_Embed layer, and Attention_Embed layer are 3×3 convolutional layer, 1×1 convolutional layer, and 1×1 convolutional layer respectively. for , The number of channels.

7. The multi-source target detection method based on pixel-level dynamic modality selection fusion according to claim 5 is characterized in that: The deep network consists of two deep branches, namely the visible light deep extraction branch and the infrared deep extraction branch. The two deep branches have the same structure, both consisting of multiple ConvNormLayers and multiple ResBlocks. Visible light deep extraction branch pair Perform deep feature extraction at different scales to obtain visible light deep features with resolutions of 1 / 8, 1 / 16, and 1 / 32 of the original resolution. , and ; Infrared deep extraction branch pair Perform deep feature extraction at different scales to obtain infrared deep features with resolutions of 1 / 8, 1 / 16, and 1 / 32 of the original resolution. , and ; Will and , and ,as well as and , respectively, and add them element by element to obtain three fusion features in turn , and .

8. A multi-source target detection method based on pixel-level dynamic modality selection fusion according to any one of claims 1 to 7, characterized in that: The dual-branch convolutional neural network encoder, pixel-level dynamic modal weight selection network, deep network and detection head decoder together constitute the multi-source target detection model. When iteratively training the multi-source target detection model, the total loss function used is expressed by the formula: ; in, is the L1 loss, is the generalized intersection-combination loss, is the classification loss, and Used to calculate the difference between the predicted target position and the actual target position. Used to calculate the difference between the predicted target category and the true target category. Represents the total loss function.

9. An electronic device, characterized in that: include: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a multi-source target detection method based on pixel-level dynamic modality selection fusion as described in any one of claims 1 to 8.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it can implement a multi-source target detection method based on pixel-level dynamic modality selection fusion as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Infrared and visible light image fusion system and method

    CN114187214A

  • Transform-based dual-band image semantic segmentation method

    CN116469100A

  • Target tracking method and system under visible light and infrared images

    CN116758117A

  • Small target detection method based on attention mechanism

    CN117292117A

  • RGBT target tracking method based on target perception enhancement fusion structure

    CN117474957A

Cited By

  • Road area environment safety evaluation method and system in highway engineering construction period

    CN120355239A