Visible light-infrared dual-mode fusion target detection system and method and medium

Through the dual-modal fusion target detection system, the deep feature weighted summation and multi-scale feature fusion of visible light and infrared images are utilized to solve the problems of insufficient detection accuracy and robustness in existing technologies, and realize adaptive high-precision target detection.

CN120707829APending Publication Date: 2025-09-26YUNNAN POWER GRID CO LTD ELECTRIC POWER RES INST
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510826191.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing visible light-infrared dual-modal fusion methods have limitations in cross-modal feature utilization and scene adaptability, resulting in limited improvement in detection performance, especially insufficient detection accuracy and robustness in low-visibility scenarios.

Method used

A dual-modal fusion target detection system is adopted. The deep features of visible light and infrared images are weighted and summed with the feature weight map to generate a first fused feature image, and multi-scale feature fusion is used for target detection. At the same time, the model selection module selects the optimal sub-model according to the probability distribution of the fused features to achieve adaptive detection.

Benefits of technology

It improves the accuracy and robustness of target detection, enhances the system's adaptability to environmental changes, and improves detection performance and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707829A_ABST
    Figure CN120707829A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a visible light-infrared bimodal fusion target detection system and method and a medium, and the system comprises the steps: a bimodal fusion target detection model carries out the weighted summation of deep features of visible light and infrared images and a feature weight map, and obtains a first fusion feature image for target detection; the visible light or infrared light single-mode target detection model carries out two times of down-sampling on the image and then carries out standard convolution to obtain an initial feature image, after four times of convolution output splicing, multi-scale feature fusion is carried out, and a second fusion feature image is obtained for detection; the model selection module selects an optimal sub-model from candidate models (bimodal fusion, visible light or infrared single-modal model) for detection according to the probability distribution of the six-channel spliced image. According to the system, through feature level and decision level fusion, the target detection accuracy and stability in a complex scene are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of target detection technology, and in particular to a visible light-infrared dual-modal fusion target detection system, method and medium. Background Art

[0002] In recent years, dual-modal fusion detection strategies based on visible and infrared imagery have attracted widespread attention in fields such as security surveillance, intelligent transportation, and military reconnaissance due to their rich information and strong robustness. Visible light images, due to their high resolution and rich texture details, perform well in everyday lighting conditions, especially in well-lit environments, providing clear target detection information. However, their performance degrades significantly in low-visibility scenarios such as nighttime, haze, or backlighting. In contrast, infrared images, by capturing the thermal radiation difference between the target and the background, can provide clear images in both day and night environments and have the significant advantage of penetrating smoke and hot targets. However, infrared images have weak texture information, blurred edges, and are susceptible to thermal interference, limiting their detection capabilities when used alone. The complementary spatial detail and thermal signatures of the two imaging modalities give them great potential for target detection in complex scenes.

[0003] Despite this, current bimodal fusion methods still have limitations in terms of cross-modal feature utilization and scene adaptability, which restricts further improvements in detection performance. Fusion strategies typically rely on fixed structures or manually designed modules and lack the ability to dynamically adjust fusion strategies based on environmental changes. Therefore, how to effectively fuse visible light and infrared imagery to improve the accuracy and robustness of target detection has become an urgent problem to be solved. Although existing feature-level and decision-level fusion methods have made breakthroughs in improving target detection performance, they still face problems such as difficulty in dynamically adapting feature selection weights, poor cross-modal semantic consistency, and insufficient adaptability to complex scenes. These issues make it difficult for existing bimodal fusion methods to meet the requirements of high accuracy and strong robustness in practical applications. Therefore, it is particularly important to develop a new bimodal fusion target detection technology that can effectively fuse visible light and infrared image information while also having strong environmental adaptability and detection robustness. Summary of the Invention

[0004] Based on this, it is necessary to propose a visible light-infrared dual-modal fusion target detection system, method and medium to address the above problems.

[0005] A visible light-infrared dual-modal fusion target detection system, the system comprising:

[0006] The dual-modal fusion target detection model is used to perform weighted summation on the deep features of the three-channel visible light image and infrared image with the feature weight map to obtain a first fused feature image, and perform target detection on the first fused feature image.

[0007] A single-modal target detection model for visible light or infrared light is used to perform two downsampling operations on the visible light image or infrared image, then perform a standard convolution operation to obtain an initial feature image, perform four convolution operations on the initial feature image, stitch the output images of each convolution operation to obtain a stitched image, perform multi-scale feature fusion on the stitched image to obtain a second fused feature image, and perform target detection on the second fused feature image.

[0008] A model selection module is used to select candidate models based on the probability distribution of the six-channel stitched image formed by stitching the three-channel visible light image and the infrared image, obtain the optimal sub-model among the candidate models, and perform target detection using the optimal sub-model. The candidate models include a dual-modal fusion target detection model and a single-modal target detection model of visible light or infrared light.

[0009] The dual-modal fusion target detection model specifically includes:

[0010] The dual-branch feature extraction module is used to extract shallow features of the three-channel visible light image and infrared image.

[0011] A feature weight generation network is used to stitch the three-channel visible light image and infrared image to obtain a six-channel stitched image, and extract feature weights of the six-channel stitched image.

[0012] The fusion module is used to perform weighted summation on the deep features of the three-channel visible light image and infrared image respectively according to the feature weights to obtain a first fused feature image.

[0013] The target detection module is used to perform target detection on the first fused feature image.

[0014] The dual-branch feature extraction module specifically includes:

[0015] The first convolutional neural network is used to extract shallow features of the three-channel visible light image and infrared image.

[0016] a second convolutional neural network, configured to extract shallow features of the three-channel visible light image and infrared image, respectively, to obtain deep features of the three-channel visible light image and infrared image;

[0017] Among them, the three-channel visible light image and infrared image input are respectively input into the sub-models of the candidate model, and the output of the sub-model with the best detection accuracy is used as the sample label, and the model selection module is trained according to the sample label.

[0018] A visible light-infrared dual-modal fusion target detection method, the method comprising:

[0019] Acquire three-channel visible light images and infrared images.

[0020] A weighted sum is performed on the deep features of the three-channel visible light image and infrared image and the feature weight map to obtain a first fused feature image, and target detection is performed on the first fused feature image.

[0021] After performing two downsampling operations on the visible light image or the infrared light image, a standard convolution operation is performed to obtain an initial feature image. After performing four convolution operations on the initial feature image, the output images of each convolution operation are spliced ​​to obtain a spliced ​​image. Multi-scale feature fusion is performed on the spliced ​​image to obtain a second fused feature image. Target detection is performed on the second fused feature image.

[0022] Candidate models are selected according to the probability distribution of the fused feature image, and the optimal sub-model in the candidate model is obtained. Target detection is performed using the optimal sub-model. The candidate models include a dual-modal fusion target detection model and a single-modal target detection model of visible light or infrared light. The fused feature image includes a first fused feature image and a second fused feature image.

[0023] The step of performing weighted summation on the deep features of the three-channel visible light image and infrared image and the corresponding feature weight maps to obtain a first fused feature image, and performing target detection on the first fused feature image specifically includes:

[0024] Shallow features of the three-channel visible light image and infrared image are extracted.

[0025] Feature extraction is performed on shallow features of the three-channel visible light image and infrared image respectively to obtain deep features of the three-channel visible light image and infrared image.

[0026] The three-channel visible light image and infrared image are stitched together to obtain a six-channel stitched image, and feature weights of the six-channel stitched image are extracted.

[0027] The deep features of the three-channel visible light image and infrared image are weighted and summed using the feature weights to obtain a first fused feature image.

[0028] The step of performing weighted summation on the deep features of the three-channel visible light image and infrared image using the feature weights to obtain a first fused feature image specifically includes:

[0029] according to Perform weighted summation on the deep features of the three-channel visible light image and infrared image to obtain a first fusion feature image, where F vis is the deep feature of the visible light image, F ir is the deep feature of infrared image, F fused is the first fusion feature image, and w is the feature weight.

[0030] The performing multi-scale feature fusion on the stitched image to obtain a second fused feature image, and performing target detection on the second fused feature image specifically includes:

[0031] A convolution operation is performed on the spliced ​​image to obtain a multi-level feature image, wherein the multi-level feature image includes a shallow feature image, a middle feature image, and a deep feature image.

[0032] An upsampling operation is performed on the deep feature image, and after splicing the upsampled deep feature image with the middle feature image, a 3×3 convolution operation is performed to output an enhanced feature map.

[0033] An upsampling operation is performed on the middle-layer feature image, and after splicing the upsampled middle-layer feature image with the shallow-layer feature image, a 3×3 convolution operation is performed to output a second fused feature image of the first scale.

[0034] A downsampling operation is performed on the second fused feature image of the first scale, and the downsampled second fused feature image of the first scale is concatenated with the first enhanced feature map, and a 3×3 convolution operation is performed again to output the second fused feature image of the second scale.

[0035] A downsampling operation is performed on the second fused feature image of the second scale, and the downsampled second fused feature image of the second scale is spliced ​​with the deep feature image, and a 3×3 convolution operation is performed again to output the second fused feature image of the third scale.

[0036] Perform target detection on the second fused feature images of different scales to generate final detection results.

[0037] Among them, the candidate models are selected according to the probability distribution of the six-channel spliced ​​image formed by splicing the three-channel visible light image and the infrared image, and the optimal sub-model in the candidate models is obtained. Target detection is performed using the optimal sub-model. The candidate models include a dual-modal fusion target detection model and a single-modal target detection model of visible light or infrared light, specifically including:

[0038] The three-channel visible light image and infrared image are stitched together to obtain a six-channel stitched image.

[0039] The six-channel stitched image is mapped into a probability distribution.

[0040] A three-category probability distribution is determined according to the probability distribution.

[0041] The candidate models are selected according to the three-category probability distribution, the optimal sub-model in the candidate models is obtained, and target detection is performed using the optimal sub-model. The candidate models include a dual-modal fusion target detection model and a single-modal target detection model of visible light or infrared light.

[0042] The step of mapping the fused feature image into a probability distribution specifically includes:

[0043] Compressing the fused feature image into a one-dimensional feature image through an average pooling operation;

[0044] Performing a linear transformation on the one-dimensional feature image, mapping the average pooled one-dimensional feature image to a target dimension, and obtaining an output vector;

[0045] The probability distribution of the output vector is calculated by the softmax function.

[0046] A computer-readable storage medium stores a computer program, which, when executed by a processor, causes the processor to perform the steps of the above method.

[0047] The embodiments of the present invention have the following beneficial effects:

[0048] The visible light and infrared dual-modal fusion target detection system provided by this invention primarily consists of three key modules: a dual-modal fusion target detection model, a single-modal target detection model for visible light or infrared light, and a model selection module. The dual-modal fusion target detection model performs a weighted summation of deep features from visible light and infrared images with a feature weight map to generate a first fused feature image. This image is used for target detection, fusing information from both lighting conditions to improve detection accuracy and robustness.

[0049] The single-modality target detection model for visible light or infrared light first performs two downsampling operations on the single-modality visible light or infrared image, performs standard convolution to generate an initial feature image, and then, after four convolution operations, splices the output images to form a spliced ​​image. Based on the principle of multi-scale feature fusion, a second fused feature image is generated for target detection.

[0050] The model selection module selects the optimal sub-model and the optimal detection model according to the probability distribution of the fusion features. By learning the detection effects of the training set images under different models, it can automatically determine which model should be selected in the current scenario to obtain the optimal detection results, and realize the adaptive selection of the optimal solution for detection performance, effectively enhancing the system's adaptability to environmental changes, while ensuring the detection accuracy, and improving the robustness and application breadth of the overall system. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0052] in:

[0053] Figure 1 A schematic structural diagram of an embodiment of a visible light-infrared dual-modal fusion target detection system provided by the present invention;

[0054] Figure 2 A schematic structural diagram of another embodiment of a visible light-infrared dual-modal fusion target detection system provided by the present invention;

[0055] Figure 3 A schematic structural diagram of an embodiment of a dual-branch feature extraction module provided by the present invention;

[0056] Figure 4 A schematic structural diagram of an embodiment of a fusion module based on a feature weight generation network provided by the present invention;

[0057] Figure 5 A flow chart of an embodiment of a visible light-infrared dual-modal fusion target detection method provided by the present invention;

[0058] Figure 6 This is a schematic structural diagram of an embodiment of the medium provided by the present invention. DETAILED DESCRIPTION

[0059] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0060] like Figure 1 As shown, Figure 1 This is a schematic diagram of the structure of an embodiment of a visible light-infrared dual-modal fusion target detection system provided by the present invention. A visible light-infrared dual-modal fusion target detection system 10 includes:

[0061] The dual-modal fusion target detection model 11 is used to perform weighted summation on the deep features of the three-channel visible light image and infrared image with the feature weight map to obtain a first fused feature image and perform target detection on the first fused feature image.

[0062] For example, the dual-modal fusion target detection model based on the feature-level fusion module mainly consists of a dual-branch feature extraction module, a fusion module based on a feature weight generation network, and a target detection module. The dual-branch feature extraction module includes a first convolutional neural network and a second convolutional neural network. The first convolutional neural network extracts shallow features of the three-channel visible light image and infrared image, and the second convolutional neural network extracts features of the shallow features of the three-channel visible light image and infrared image respectively, obtains deep features of the three-channel visible light image and infrared image, and completes the extraction of deep features of the visible light image and infrared image.

[0063] At the same time, the feature weight generation network in the fusion module based on the feature weight generation network is used to splice the three-channel visible light image and the infrared image to obtain a six-channel spliced ​​image, and the feature weights of the six-channel spliced ​​image are extracted.

[0064] Furthermore, the fusion module in the fusion module based on the feature weight generation network realizes weighted fusion of the dual-path features according to the feature weight map output by the feature weight generation network to obtain a first fused feature image.

[0065] Finally, the target detection module is used to perform target detection on the first fused feature image.

[0066] A single-modal target detection model 12 for visible light or infrared light is used to perform two downsampling operations on a visible light image or an infrared image, then perform a standard convolution operation to obtain an initial feature image, perform four convolution operations on the initial feature image, stitch the output images of each convolution operation to obtain a stitched image, perform multi-scale feature fusion on the stitched image, obtain a second fused feature image, and perform target detection on the second fused feature image.

[0067] For example, a single-modality target detection model for visible light or infrared light performs target detection based solely on visible light or infrared single-modality images. The detection model framework can be flexibly designed and can also be designed based on currently advanced target detection model frameworks. In the present invention, the single-modality target detection model for visible light or infrared light adopts the Yolov7 model.

[0068] Specifically, assuming that the input image (visible light image or infrared light image) is Yolov7 first extracts features through a backbone network consisting of multiple ELAN (Efficient Layer Aggregation Network) modules. The core concepts behind this network are path convolution, feature fusion, and residual connections. In the backbone network, the visible or infrared image is downsampled twice, followed by a standard convolution operation to obtain the initial feature image F1. The ELAN module is then used for calculations, employing a four-way convolutional architecture. This four-way convolutional architecture performs four convolution operations on the initial feature image, as shown in the following equation:

[0069] A1=Conv 1×1 (F1);

[0070] A2=Conv 3×3 (A1);

[0071] A3=Conv 3×3 (A2);

[0072] A4=Conv 1×1 (A3);

[0073] Among them, A1 is the output image of the first convolution operation, Conv1×1 is a 1×1 convolution operation, A2 is the output image of the second convolution operation, Conv3×3 is a 3×3 convolution operation, A3 is the output image of the third convolution operation, and A4 is the output image of the fourth convolution operation.

[0074] Furthermore, stitch them together to obtain a stitched image:

[0075] F ELAN =concat(F1,A2,A3,A4);

[0076] Among them, FELAN is the stitching image, and concat is the stitching operation.

[0077] Furthermore, convolution operation is performed on the spliced ​​image to obtain multi-level feature images:

[0078] F out =Conv 1×1 (F ELAN );

[0079] Among them, FELAN is the spliced ​​image and Fout is the multi-level feature image.

[0080] After the backbone network, multi-scale feature fusion is performed, which mainly includes top-down fusion (FPN) and bottom-up enhanced semantics (PAN). The multi-level feature map includes shallow feature images, middle feature images, and deep feature images. Assume that the shallow feature images, middle feature images, and deep feature images are The outputs of the ELAN modules from different layers are arranged in order from shallow to deep. The calculation method of FPN is:

[0081]

[0082] Among them, U1 is the first intermediate feature, is the deep feature image, F4 is the first enhanced feature map, is the middle feature image, U2 is the second middle feature, is the shallow feature image, and F5 is the second fused feature image of the first scale.

[0083] Furthermore, the fine-grained features are passed downward through downsampling, as shown in the following formula:

[0084] D4=Conv 3×3 (Concat(Downsample(F5),F4));

[0085]

[0086] Among them, F4 is the first enhanced feature map, F5 is the second fused feature image of the first scale, D4 ​​is the second fused feature image of the second scale, and D5 is the second fused feature image of the third scale. is a deep feature image.

[0087] Finally, the feature fusion images of three scales are obtained. Then, the target detection head performs target detection on the second fused feature images of different scales to generate the final detection results.

[0088] The model selection module 13 is used to select candidate models based on the probability distribution of the six-channel spliced ​​image formed by splicing the three-channel visible light image and the infrared image, obtain the optimal sub-model among the candidate models, and perform target detection using the optimal sub-model. The candidate models include a dual-modal fusion target detection model and a single-modal target detection model of visible light or infrared light. The fused feature image includes a first fused feature image and a second fused feature image.

[0089] Exemplarily, a three-channel visible light image and an infrared image are spliced ​​to obtain a six-channel spliced ​​image; the six-channel spliced ​​image is mapped to a probability distribution, specifically, the fused feature image is compressed into a one-dimensional feature image through an average pooling operation; a linear transformation is performed on the one-dimensional feature image, and the one-dimensional feature image after average pooling is mapped to a target dimension to obtain an output vector; and the probability distribution of the output vector is calculated through a softmax function.

[0090] Furthermore, a three-class probability distribution is determined based on the probability distribution. Based on the three-class probability distribution, candidate models are selected to obtain the optimal sub-model among the candidate models. Specifically, if the three-class probability distribution result is Class 1, the dual-modal fusion target detection model is selected; if Class 2, the visible light single-modal target detection model is selected; and if Class 3, the infrared light single-modal target detection model is selected.

[0091] Then, target detection is performed through the optimal sub-model. The candidate models include a dual-modal fusion target detection model and a single-modal target detection model of visible light or infrared light.

[0092] Furthermore, the three-channel visible light image and infrared image are respectively input into the sub-model of the candidate model to obtain the detection accuracy score, as shown in the following formula:

[0093] S1=score(M1([X vis ,X ir ]));

[0094] S2=score(M2(X vis ));

[0095] S3=score(M3(X ir ));

[0096] Among them, score(·) is the comprehensive performance indicator of the sub-model. In this experiment, it is mAP50. S1 is the score of the dual-modal fusion target detection model, S2 is the score of the single-modal target detection model for visible light, and S3 is the score of the single-modal target detection model for infrared light.

[0097] The output of the sub-model with the best detection accuracy is used as the sample label, and the model selection module is trained based on the sample label.

[0098] Furthermore, we design the loss function

[0099] Using standard cross entropy loss:

[0100] Let the output of the model selection module be:

[0101] p i =softmax(f MSN (Xi ))∈R 3 ;

[0102] The true label is:

[0103] y i ∈{0,1,2};

[0104] Then the loss function is:

[0105]

[0106] in, Indicates that in the i-th sample, the model selection module selects the true label y i The predicted probability of .

[0107] Furthermore, back propagation and gradient descent are used to update the parameters of MSN:

[0108]

[0109] Among them, θ is all the parameters of the model selection module, and η is the learning rate.

[0110] As can be seen from the above description, the visible light and infrared dual-modal fusion target detection system provided by the present invention mainly includes three key modules: a dual-modal fusion target detection model, a single-modal target detection model for visible light or infrared light, and a model selection module. The dual-modal fusion target detection model generates a first fused feature image by weighted summing the deep features of the visible light and infrared images with a feature weight map. This image is used for target detection, fusing information from both lighting conditions to improve detection accuracy and robustness.

[0111] The single-modality target detection model for visible light or infrared light first performs two downsampling operations on the single-modality visible light or infrared image, performs standard convolution to generate an initial feature image, and then, after four convolution operations, splices the output images to form a spliced ​​image. Based on the principle of multi-scale feature fusion, a second fused feature image is generated for target detection.

[0112] The model selection module selects the optimal sub-model and the optimal detection model according to the probability distribution of the fusion features. By learning the detection effects of the training set images under different models, it can automatically determine which model should be selected in the current scenario to obtain the optimal detection results, and realize the adaptive selection of the optimal solution for detection performance, effectively enhancing the system's adaptability to environmental changes, while ensuring the detection accuracy, and improving the robustness and application breadth of the overall system.

[0113] like Figure 2 As shown, Figure 2This is a schematic diagram of another embodiment of a visible light-infrared dual-modal fusion target detection system provided by the present invention. The dual-modal fusion target detection model specifically includes:

[0114] A dual-branch feature extraction module is used to extract shallow features of three-channel visible light images and infrared images.

[0115] For example, in combination Figure 3 , Figure 3 This is a schematic diagram of the structure of an embodiment of the dual-branch feature extraction module provided by the present invention. The first convolutional neural network of the dual-branch feature extraction module is used to extract shallow features of the three-channel visible light image and infrared image;

[0116] First, the visible light and infrared images are sequentially passed through several layers of a shared network to obtain shallow feature maps. This shared network includes the first convolutional neural network. This network extracts common features between the two modal images. Furthermore, by sharing network parameters, it effectively reduces computational effort, avoids unnecessary redundant calculations, and improves overall computational efficiency. Furthermore, this shared network parameterization improves the model's generalization capabilities.

[0117] Furthermore, the output of the shared network passes through two independent branch modules to obtain deep feature maps corresponding to the visible light image and infrared image. Since the network parameters of the visible light branch and the infrared branch are independent of each other, each focuses on mining the unique features of its corresponding modality, providing more valuable feature data for subsequent feature-level fusion. Specifically, the visible light branch network and the infrared branch network include a second convolutional neural network, which extracts the shallow features of the three-channel visible light image and infrared image respectively through the second convolutional neural network to obtain the deep features of the three-channel visible light image and infrared image.

[0118] The feature weight generation network 22 is used to stitch the three-channel visible light image and the infrared image to obtain a six-channel stitched image, and extract the feature weights of the six-channel stitched image.

[0119] The fusion module 23 is used to perform weighted summation on the deep features of the three-channel visible light image and infrared image respectively through feature weights to obtain a first fused feature image.

[0120] For example, referring to Figure 4 , Figure 4 This is a schematic diagram of the structure of an embodiment of a fusion module based on a feature weight generation network provided by the present invention. The fusion module based on a feature weight generation network includes a feature weight generation network and a fusion module. The feature weight generation network generates a feature weight map, and then performs a weighted summation of the two feature maps output by the dual-branch feature extraction network based on the feature weight map to achieve feature fusion.

[0121] Specifically, the feature weighted generation network (FWGN) is based on the three-channel visible light image data X vis and X ir The input is a six-channel mosaic of three-channel infrared image data. The output is a feature weight W with the same size as the feature map output by the dual-branch feature extraction module. Each element in the feature weight W ranges from [0, 1]. This network uses a convolutional neural network as its basic architecture. Hyperparameters such as network depth and the number of convolution kernels per layer can be flexibly designed based on actual conditions. It is important to note that since each element in W ranges from [0, 1], the final layer should use a sigmoid activation function.

[0122] The first fused feature image F fused Calculated by weighted summation:

[0123]

[0124] Among them, F vis is the deep feature of the visible light image, F ir is the deep feature of infrared image, F fused is the first fusion feature image, and w is the feature weight.

[0125] From the above formula, we can see that when the value of w is large, it means that in the fusion feature map F fused In the generation process of W, the contribution of visible light modal features is greater; on the contrary, if W is smaller, the influence of infrared modal features on the fusion features is more obvious. The FWGN network outputs different values ​​of w based on the input visible light-infrared image content, which can automatically adjust the weight distribution of visible light and infrared modalities according to the scene content of the input image, realize dynamic balance of the contribution of dual modal features, and thus make the fused feature F fused This approach maximizes the advantages of both modal features to improve target detection performance. By guiding the network to learn the actual contribution of different modalities to the detection task, it enables effective extraction and fusion of target features in a variety of scenarios, including strong and weak light, occlusion, and complex backgrounds, significantly improving the accuracy and robustness of target detection.

[0126] The target detection module 24 is configured to perform target detection on the first fused feature image.

[0127] like Figure 5 As shown, Figure 5 A flow chart of an embodiment of a visible light-infrared dual-modal fusion target detection method provided by the present invention. A visible light-infrared dual-modal fusion target detection method, the method comprising:

[0128] S101: Acquire three-channel visible light images and infrared images.

[0129] S102: Perform weighted summation on the deep features of the three-channel visible light image and infrared image and the feature weight map to obtain a first fused feature image, and perform target detection on the first fused feature image.

[0130] Exemplarily, shallow features of the three-channel visible light image and infrared image are extracted; feature extraction is performed on the shallow features of the three-channel visible light image and infrared image respectively to obtain deep features of the three-channel visible light image and infrared image; the three-channel visible light image and infrared image are spliced ​​to obtain a six-channel spliced ​​image, and feature weights of the six-channel spliced ​​image are extracted; the deep features of the three-channel visible light image and infrared image are weighted and summed using the feature weights to obtain a first fused feature image. Specifically, the deep features of the three-channel visible light image and infrared image are weighted and summed according to the following formula to obtain the first fused feature image:

[0131]

[0132] Among them, F vis is the deep feature of the visible light image, F ir is the deep feature of infrared image, F fused is the first fusion feature image, and w is the feature weight.

[0133] S103: After performing two downsampling operations on the visible light image or the infrared light image, a standard convolution operation is performed to obtain an initial feature image, the initial feature image is convolved four times, and the output images of each convolution operation are spliced ​​to obtain a spliced ​​image, multi-scale feature fusion is performed on the spliced ​​image to obtain a second fused feature image, and target detection is performed on the second fused feature image.

[0134] Exemplarily, a convolution operation is performed on the spliced ​​image to obtain a multi-level feature image, which includes a shallow feature image, a middle feature image, and a deep feature image. An upsampling operation is performed on the deep feature image, and the upsampled deep feature image is spliced ​​with the middle feature image, and a 3×3 convolution operation is performed to output an enhanced feature map, as shown in the following formula:

[0135]

[0136] Among them, U1 is the first intermediate feature, is the deep feature image, F4 is the first enhanced feature map, is the middle-level feature image.

[0137] Furthermore, the middle-layer feature image is upsampled, and the upsampled middle-layer feature image is concatenated with the shallow-layer feature image, and then a 3×3 convolution operation is performed to output the second fused feature image of the first scale, as shown in the following formula:

[0138]

[0139] Among them, U1 is the first intermediate feature, U2 is the second intermediate feature, is the shallow feature image, and F5 is the second fused feature image of the first scale.

[0140] Furthermore, a downsampling operation is performed on the second fused feature image of the first scale, and the downsampled second fused feature image of the first scale is spliced ​​with the first enhanced feature map, and a 3×3 convolution operation is performed again to output the second fused feature image of the second scale, as shown in the following formula:

[0141] D4=Conv 3×3 (Concat(Downsample(F5),F4));

[0142] Among them, F4 is the first enhanced feature map, F5 is the second fused feature image of the first scale, and D4 is the second fused feature image of the second scale.

[0143] The second fused feature image of the second scale is downsampled, and the downsampled second fused feature image of the second scale is concatenated with the deep feature image. Then, a 3×3 convolution operation is performed again to output the second fused feature image of the third scale, as shown in the following formula:

[0144]

[0145] Among them, D4 is the second fused feature image of the second scale, D5 is the second fused feature image of the third scale, is a deep feature image.

[0146] Furthermore, target detection is performed on the second fused feature images of different scales to generate a final detection result.

[0147] S104: Select candidate models based on the probability distribution of a six-channel spliced ​​image formed by splicing three-channel visible light images and infrared images, obtain the optimal sub-model among the candidate models, and perform target detection using the optimal sub-model. The candidate models include a dual-modal fusion target detection model and a single-modal target detection model of visible light or infrared light.

[0148] Exemplarily, a three-channel visible light image and an infrared image are spliced ​​to obtain a six-channel spliced ​​image; the six-channel spliced ​​image is mapped to a probability distribution, specifically, the fused feature image is compressed into a one-dimensional feature image through an average pooling operation; a linear transformation is performed on the one-dimensional feature image, and the one-dimensional feature image after average pooling is mapped to a target dimension to obtain an output vector; and the probability distribution of the output vector is calculated through a softmax function.

[0149] Furthermore, a three-class probability distribution is determined based on the probability distribution. Based on the three-class probability distribution, candidate models are selected to obtain the optimal sub-model among the candidate models. Specifically, if the three-class probability distribution result is Class 1, the dual-modal fusion target detection model is selected; if Class 2, the visible light single-modal target detection model is selected; and if Class 3, the infrared light single-modal target detection model is selected.

[0150] Then, target detection is performed through the optimal sub-model. The candidate models include a dual-modal fusion target detection model and a single-modal target detection model of visible light or infrared light.

[0151] like Figure 6 As shown, Figure 6 The structure diagram of an embodiment of the medium provided by the present invention. The medium 10 stores at least one computer program 11, which is executed by a processor to implement the following Figure 5 In one embodiment, the medium 10 can be a memory chip, a hard disk, a mobile hard disk, a USB flash drive, an optical disk, or other readable and writable storage tools, or a server.

[0152] Additionally, the processes depicted in the accompanying figures do not necessarily have to be performed in the particular order shown, or sequential order, to achieve desired results. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.

[0153] The various embodiments in this specification are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from the other embodiments. In particular, the device, apparatus, and non-volatile computer-readable storage medium embodiments are described briefly because they are generally similar to the method embodiments. For relevant portions, refer to the description of the method embodiments.

[0154] The apparatus, device, non-volatile computer-readable storage medium and method provided in the embodiments of this specification correspond to each other. Therefore, the apparatus, device, and non-volatile computer storage medium also have similar beneficial technical effects as the corresponding method. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the corresponding apparatus, device, and non-volatile computer storage medium will not be repeated here.

[0155] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0156] For the convenience of description, when describing the above device, various units are divided into functions and described separately. Of course, when implementing this specification, the functions of each unit can be implemented in the same one or more software and / or hardware. It should be understood by those skilled in the art that this specification embodiment can be provided as a method, system, or computer program product. Therefore, this specification embodiment can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, this specification embodiment can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0157] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0158] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0159] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0160] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0161] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0162] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0163] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0164] This specification may be described in the general context of computer-executable instructions, such as program modules, executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communications network. In a distributed computing environment, program modules may be located in both local and remote computer storage media, including storage devices.

[0165] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are described briefly because they are generally similar to the method embodiments. For relevant parts, refer to the description of the method embodiments.

[0166] The above disclosure is merely a preferred embodiment of the present invention and certainly cannot be used to limit the scope of the present invention. Therefore, equivalent changes made according to the claims of the present invention are still within the scope of the present invention.

Claims

1. A visible light-infrared dual-modal fusion target detection system, characterized in that: The system comprises: A dual-modal fusion target detection model is used to perform weighted summation on the deep features of the three-channel visible light image and infrared image with the feature weight map to obtain a first fused feature image, and perform target detection on the first fused feature image; A single-modality target detection model for visible light or infrared light is configured to perform two downsampling operations on a visible light image or infrared image, then perform a standard convolution operation to obtain an initial feature image, perform four convolution operations on the initial feature image, stitch the output images of each convolution operation to obtain a stitched image, perform multi-scale feature fusion on the stitched image to obtain a second fused feature image, and perform target detection on the second fused feature image; A model selection module is used to select candidate models based on the probability distribution of the six-channel stitched image formed by stitching the three-channel visible light image and the infrared image, obtain the optimal sub-model among the candidate models, and perform target detection using the optimal sub-model. The candidate models include a dual-modal fusion target detection model and a single-modal target detection model of visible light or infrared light.

2. The visible light-infrared dual-modal fusion target detection system according to claim 1, characterized in that: The dual-modal fusion target detection model specifically includes: A dual-branch feature extraction module, used to extract shallow features of the three-channel visible light image and infrared image; a feature weight generation network for stitching the three-channel visible light image and the infrared image to obtain a six-channel stitched image, and extracting feature weights of the six-channel stitched image; a fusion module, configured to perform weighted summation on the deep features of the three-channel visible light image and infrared image respectively using the feature weights to obtain a first fused feature image; The target detection module is used to perform target detection on the first fused feature image.

3. The visible light-infrared dual-modal fusion target detection system according to claim 2, characterized in that: The dual-branch feature extraction module specifically includes: A first convolutional neural network is used to extract shallow features of the three-channel visible light image and infrared image; The second convolutional neural network is used to extract shallow features of the three-channel visible light image and infrared image respectively, and obtain deep features of the three-channel visible light image and infrared image.

4. The visible light-infrared dual-modal fusion target detection system according to claim 1, characterized in that: The three-channel visible light image and infrared image input are respectively input into the sub-models of the candidate model, and the output of the sub-model with the best detection accuracy is used as the sample label, and the model selection module is trained according to the sample label.

5. A visible light-infrared dual-modal fusion target detection method, characterized in that: The method comprises: Acquire three-channel visible light images and infrared images; Performing weighted summation on the deep features of the three-channel visible light image and infrared image and the feature weight map to obtain a first fused feature image, and performing target detection on the first fused feature image; After performing two downsampling operations on the visible light image or the infrared light image, a standard convolution operation is performed to obtain an initial feature image, the initial feature image is convolved four times, and the output images of each convolution operation are spliced ​​to obtain a spliced ​​image, multi-scale feature fusion is performed on the spliced ​​image to obtain a second fused feature image, and target detection is performed on the second fused feature image; Candidate models are selected according to the probability distribution of the fused feature image, and the optimal sub-model in the candidate model is obtained. Target detection is performed using the optimal sub-model. The candidate models include a dual-modal fusion target detection model and a single-modal target detection model of visible light or infrared light. The fused feature image includes a first fused feature image and a second fused feature image.

6. The visible light-infrared dual-modal fusion target detection method according to claim 5, characterized in that: The step of performing weighted summation on the deep features of the three-channel visible light image and infrared image and the corresponding feature weight maps to obtain a first fused feature image, and performing target detection on the first fused feature image specifically includes: Extracting shallow features of the three-channel visible light image and infrared image; Extracting shallow features of the three-channel visible light image and infrared image to obtain deep features of the three-channel visible light image and infrared image; Stitching the three-channel visible light image and infrared image to obtain a six-channel stitched image, and extracting feature weights of the six-channel stitched image; The deep features of the three-channel visible light image and infrared image are weighted and summed using the feature weights to obtain a first fused feature image.

7. The visible light-infrared dual-modal fusion target detection method according to claim 6, characterized in that: The step of performing weighted summation on the deep features of the three-channel visible light image and infrared image using the feature weights to obtain a first fused feature image specifically includes: according to Perform weighted summation on the deep features of the three-channel visible light image and infrared image to obtain a first fusion feature image, where F vis is the deep feature of the visible light image, F ir is the deep feature of infrared image, F fused is the first fusion feature image, and w is the feature weight.

8. The visible light-infrared dual-modal fusion target detection method according to claim 5, characterized in that: The performing multi-scale feature fusion on the stitched image to obtain a second fused feature image, and performing target detection on the second fused feature image specifically includes: Performing a convolution operation on the spliced ​​image to obtain a multi-level feature image, wherein the multi-level feature image includes a shallow feature image, a middle feature image, and a deep feature image; Performing an upsampling operation on the deep feature image, concatenating the upsampled deep feature image with the middle feature image, performing a 3×3 convolution operation, and outputting an enhanced feature map; performing an upsampling operation on the middle-layer feature image, concatenating the upsampled middle-layer feature image with the shallow-layer feature image, and performing a 3×3 convolution operation to output a second fused feature image of the first scale; Downsampling the second fused feature image of the first scale, concatenating the downsampled second fused feature image of the first scale with the first enhanced feature map, and performing a 3×3 convolution operation again to output the second fused feature image of the second scale; Downsampling the second fused feature image of the second scale, concatenating the downsampled second fused feature image of the second scale with the deep feature image, and performing a 3×3 convolution operation again to output the second fused feature image of the third scale; Perform target detection on the second fused feature images of different scales to generate final detection results.

9. The visible light-infrared dual-modal fusion target detection system according to claim 5, characterized in that: The candidate models are selected based on the probability distribution of the six-channel spliced ​​image formed by splicing the three-channel visible light image and the infrared image, and the optimal sub-model in the candidate models is obtained. Target detection is performed using the optimal sub-model. The candidate models include a dual-modal fusion target detection model and a single-modal target detection model of visible light or infrared light, specifically including: Stitching the three-channel visible light image and infrared image to obtain a six-channel stitched image; Mapping the six-channel stitched image into a probability distribution; Determine a three-category probability distribution based on the probability distribution; The candidate models are selected according to the three-category probability distribution, the optimal sub-model in the candidate models is obtained, and target detection is performed using the optimal sub-model. The candidate models include a dual-modal fusion target detection model and a single-modal target detection model of visible light or infrared light.

10. The visible light-infrared dual-modal fusion target detection system according to claim 9, characterized in that: Mapping the fused feature image into a probability distribution specifically includes: Compressing the fused feature image into a one-dimensional feature image through an average pooling operation; Performing a linear transformation on the one-dimensional feature image, mapping the average pooled one-dimensional feature image to a target dimension, and obtaining an output vector; The probability distribution of the output vector is calculated by the softmax function.

11. A computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor is caused to perform the steps of the method according to any one of claims 5 to 10.

Citation Information

Cited By

  • Conditional perception-based infrared-visible light fusion target detection method and device

    CN120997491A

  • Infrared-visible light fusion target detection method and device based on conditional perception

    CN120997491B

  • Electric power well environment detection method, system and equipment based on image processing and medium

    CN121482461A