An infrared image and a visible light image fusion method, device and storage medium
Patent Information
- Application Number
- CN202410979162.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-22
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2044-07-22
AI Technical Summary
[0004]尽管在早期研究中,许多学者提出了诸多有效方法用于图像融合,但其大都主要关注于可见光与红外图像互补特征信息充分融合的过程,而忽略了针对复杂环境下源图像质量下降的问题
[0053] Compared with the prior art, the infrared image and visible light image fusion method, apparatus and storage medium provided in the embodiments of the present invention achieve the following beneficial effects:
Smart Images

Figure CN118657671B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method, apparatus, and storage medium for fusing infrared and visible light images, belonging to the field of image processing technology. Background Technology
[0002] Image fusion technology has made significant progress in the field of computer vision, especially infrared and visible light image fusion. The goal of image fusion is to effectively extract and fuse complementary feature information from source images to generate a fused image with richer information and better scene representation, thereby improving the user's visual perception experience and enhancing the performance of high-level semantic tasks such as image segmentation, object detection and tracking, and scene understanding.
[0003] However, image fusion methods in dense fog environments face numerous challenges. In foggy conditions, water vapor in the air scatters light propagation, interfering with the imaging process and degrading the quality of visible light images. Specifically, in foggy weather, most of the detailed texture information in visible light images collected by sensors is obscured, resulting in the loss of most complementary feature information compared to fog-free visible light images, thus significantly reducing the quality of the fused image. However, because infrared wavelengths have a higher ability to penetrate atmospheric particles than visible light, infrared information can capture significant target information that is obscured or distorted by fog in visible images.
[0004] Although many scholars proposed effective methods for image fusion in early research, most of them mainly focused on the process of fully fusing complementary feature information from visible and infrared images, while neglecting the problem of source image quality degradation in complex environments. This resulted in most existing infrared and visible image fusion methods producing images with blurred texture details and numerous artifacts in complex environments, severely affecting the quality of the fused image. Summary of the Invention
[0005] The purpose of this invention is to overcome the shortcomings of existing high-resolution image classification models, such as numerous parameters, large computational load, and difficulty in practical application. This invention provides a method, apparatus, and storage medium for fusing infrared and visible light images, which can retain more semantic information and improve the quality of infrared and visible light image fusion even in extreme environments such as heavy fog. To achieve the above objective, this invention employs the following technical solution:
[0006] In a first aspect, the present invention provides a method for fusing infrared images and visible light images, comprising:
[0007] Acquire real-time images of infrared and visible light images;
[0008] The acquired real-time images are input into a pre-trained bimodal feature extraction model for feature extraction, resulting in feature maps of the infrared image and the visible light image.
[0009] The real-time image and feature map of the visible light image are input into a pre-trained semantic information model for semantic information fusion to obtain a feature map of the visible light image containing semantic information.
[0010] The feature map of the visible light image containing semantic information and the feature map of the infrared image are fused to obtain a fused image.
[0011] In conjunction with the first aspect, optionally, the pre-trained bimodal feature extraction model includes a feature extractor and a weight generator in sequence;
[0012] The feature extractor is used to: extract features from the infrared image and the visible light image based on the real-time images of the infrared image and the visible light image, and generate a dual-modal feature map;
[0013] The weight generator is used to: obtain a weight score matrix based on the bimodal feature map, and output feature maps of the infrared image and the visible light image based on the bimodal feature map and the weight score matrix.
[0014] In conjunction with the first aspect, optionally, the pre-trained semantic information model includes a semantic information generator and a semantic information fusion unit in sequence;
[0015] The semantic generator is used to: perform semantic mapping of the visible light image based on the real-time image of the visible light image to obtain semantic guidance features;
[0016] The semantic information fusion unit is used to: perform further feature extraction based on the feature map of the visible light image to obtain a further feature map of the visible light image, and fuse the further feature map of the visible light image with the semantic guidance feature to obtain a feature map of the visible light image containing semantic information.
[0017] In conjunction with the first aspect, optionally, the feature map of the visible light image containing semantic information and the feature map of the infrared image are fused, and the fusion formula is expressed by the following equation:
[0018] (1)
[0019] In equation (1), F To merge images, f () performs semantic information fusion on a pre-trained semantic information processing model. I v This is a real-time image of a visible light image. I rThis is a real-time image of an infrared image. W This is the weighted score matrix. For dot product operation, This is for addition operations.
[0020] In conjunction with the first aspect, optionally, the pre-trained bimodal feature extraction model is obtained by training a pre-constructed bimodal feature extraction model based on sample image data using visible light contrast constraint loss and infrared joint constraint loss. The training process includes:
[0021] The infrared image and the foggy visible light image in the sample image data are input into a pre-constructed dual-modal feature extraction model for feature extraction, resulting in the infrared feature map and the visible light feature map of the sample image.
[0022] Based on the visible light images with fog in the sample image data and the visible light feature maps of the sample images, semantic information is fused using a pre-trained semantic information model to obtain the visible light feature maps of the sample images containing semantic information.
[0023] The visible light feature map and infrared feature map of the sample image containing semantic information are fused to obtain a fused image of the sample image;
[0024] A mask is generated based on the infrared image in the sample image data;
[0025] The fog-free visible light image, the infrared image, and the fused image of the sample image data are processed using a mask to obtain the fog-free visible light image, the infrared image, and the fused image of the sample image after mask processing.
[0026] Based on the sample image data, a fused image consisting of a foggy visible light image, a fog-free visible light image after masking, a masked infrared image, and a masked sample image, a pre-constructed bimodal feature extraction model is trained using a preset bimodal global perception loss to obtain a pre-trained bimodal feature extraction model; wherein, the preset bimodal global perception loss includes visible light contrast constraint and infrared joint constraint loss.
[0027] In conjunction with the first aspect, optionally, the mask is generated based on the infrared image in the sample image data, using the following formula:
[0028] (2)
[0029] In equation (2), For the mask in the first i Line 1 j Column elements, For infrared images in the first i Line 1j Column elements, The mean of all elements in the infrared image. t For predefined thresholds.
[0030] In conjunction with the first aspect, optionally, a pre-built bimodal feature extraction model can be trained using visible light contrast constraints, including:
[0031] The sample image data includes foggy visible light images as negative samples, masked fog-free visible light images as positive samples, and the fused image of the masked sample images as anchor samples.
[0032] The negative samples, positive samples, and anchored samples are input into a pre-trained contrastive learning feature extraction module with fixed parameters to obtain the hidden features of each sample.
[0033] The contrast constraint loss between the hidden features of each sample is calculated using the following formula:
[0034] (3)
[0035] In equation (3), L c To compare the constraint loss, GT i The first positive sample i One hidden feature, F i For the anchored sample i One hidden feature, I Vi The first negative sample i One hidden feature, p i The output of the pre-built bimodal feature extraction model is the first... i The weight coefficients of each hidden feature;
[0036] In response to the contrast constraint loss being less than a preset first threshold, the anchor sample is moved closer to the positive sample and further away from the negative sample, and the parameters of the pre-constructed bimodal feature extraction model are updated.
[0037] In conjunction with the first aspect, optionally, the pre-built bimodal feature extraction model can be trained using infrared joint constraint loss, including:
[0038] The total infrared joint constraint loss between the masked infrared image and the fused image of the masked sample image is calculated using the following formula:
[0039] (4)
[0040] In equation (4), L rFor the total loss of infrared joint constraint, As the first hyperparameter, The content loss is calculated using the following formula:
[0041] (5)
[0042] In equation (5), MSE () is used to calculate the mean square error loss. This is the fused image of the sample images after masking. As a mask, This is the fused image of the sample images after masking. The infrared image is from the sample image data. The infrared image after masking;
[0043] In equation (4), This is the second hyperparameter. The multi-scale structural similarity loss is calculated using the following formula:
[0044] (6)
[0045] In equation (6), This is the infrared image after masking. This is the fused image of the sample images after masking. for The mean, for The mean, for standard deviation for standard deviation The covariance of the fused image of the masked sample images and the masked infrared image; As the first relative importance, As the second most important, It is the first constant. It is the second constant. For scale quantity;
[0046] In response to the total loss due to infrared joint constraints being less than a preset second threshold, the parameters of the pre-constructed dual-modal feature extraction model are updated.
[0047] Secondly, this application provides an infrared image and visible light image fusion device, comprising:
[0048] Acquisition module: Used to acquire real-time images of infrared and visible light images;
[0049] Feature extraction module: This module is used to input the acquired real-time images into a pre-trained bimodal feature extraction model for feature extraction, resulting in feature maps of the infrared image and the visible light image.
[0050] Semantic information fusion module: This module is used to input real-time images and feature maps of visible light images into a pre-trained semantic information model to perform semantic information fusion, thereby obtaining a feature map of the visible light image containing semantic information.
[0051] Image fusion module: used to fuse the feature map of the visible light image containing semantic information and the feature map of the infrared image to obtain a fused image.
[0052] Thirdly, the present invention provides a computer-readable storage medium having a computer program / instructions stored thereon, which, when executed by a processor, implements the steps of the low-resolution image classification method described in the first aspect.
[0053] Compared with the prior art, the infrared image and visible light image fusion method, apparatus and storage medium provided in the embodiments of the present invention achieve the following beneficial effects:
[0054] This invention inputs the acquired real-time image into a pre-trained bimodal feature extraction model for feature extraction, obtaining feature maps of the infrared image and the visible light image; this invention can extract bimodal information from the infrared image and the visible light image to generate a weighted score matrix, and can combine the contribution of each pixel of the infrared image and the visible light image to the fused image according to the weights.
[0055] This invention inputs real-time images and feature maps of visible light images into a pre-trained semantic information model for semantic information fusion, resulting in a feature map of the visible light image containing semantic information; this invention can further mine the texture structure information preserved in the visible light image;
[0056] This invention fuses the feature map of the visible light image containing semantic information and the feature map of the infrared image to obtain a fused image; this invention can effectively combine the content of dual-modal images to generate an ideal fused image containing rich details and significant information; this invention can retain more semantic information and improve the image fusion quality even in the face of extreme environmental conditions such as heavy fog. Attached Figure Description
[0057] Figure 1 This is a schematic flowchart of an infrared image and visible light image fusion method provided in Embodiment 1 of the present invention;
[0058] Figure 2 This is a schematic diagram of the overall framework of an infrared image and visible light image fusion method provided in Embodiment 1 of the present invention;
[0059] Figure 3 This is a schematic diagram of the network structure of a dual-modal feature extraction model in an infrared image and visible light image fusion method provided in Embodiment 1 of the present invention;
[0060] Figure 4 This is a schematic diagram of the network structure of the semantic information model in an infrared image and visible light image fusion method provided in Embodiment 1 of the present invention. Detailed Implementation
[0061] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and should not be used to limit the scope of protection of the present invention.
[0062] Example 1:
[0063] like Figure 1 As shown, the present invention provides a method for fusing infrared images and visible light images, comprising:
[0064] Acquire real-time images of infrared and visible light images;
[0065] The acquired real-time images are input into a pre-trained bimodal feature extraction model for feature extraction, resulting in feature maps of the infrared image and the visible light image.
[0066] The real-time image and feature map of the visible light image are input into a pre-trained semantic information model for semantic information fusion to obtain a feature map of the visible light image containing semantic information.
[0067] The feature map of the visible light image containing semantic information and the feature map of the infrared image are fused to obtain a fused image.
[0068] The specific steps include:
[0069] Step 1: Acquire real-time images of infrared and visible light.
[0070] In this embodiment, the real-time image of the visible light image is a real-time image of the visible light image with fog.
[0071] Step 2: Input the acquired real-time image into the pre-trained dual-modal feature extraction model for feature extraction to obtain the feature map of the infrared image and the feature map of the visible light image.
[0072] like Figure 2 As shown, the pre-trained bimodal feature extraction model consists of a feature extractor and a weight generator.
[0073] The feature extractor is used to extract features from the infrared image and the visible light image based on the real-time images of the infrared image and the visible light image, and generate a bimodal feature map.
[0074] The weight generator is used to: obtain a weight score matrix from the bimodal feature map generated by the feature extractor, and output feature maps of the infrared image and the visible light image based on the bimodal feature map generated by the feature extractor and the weight score matrix.
[0075] like Figure 3 As shown, the dual-modal feature extraction model sequentially includes one component... Composed of a large kernel (9×9) convolution block (Convolution Block 1), and 2 convolutional blocks composed of... It consists of small kernel (3×3) convolution blocks, 6 residual blocks, 2 small kernel (3×3) convolution blocks, and a sigmoid activation function.
[0076] In this embodiment, the feature maps of the infrared image and the visible light image are obtained as follows:
[0077] Step 2.1: Stitch together the real-time infrared image and the real-time visible light image affected by heavy fog as a dual-channel image.
[0078] Step 2.2: Input the stitched image into a shared fully convolutional network to extract features from the infrared image and the visible light image, generating a dual-modal feature map.
[0079] Unlike most current fusion methods that manually pad the input data, this embodiment uses a shared fully convolutional network to pad each layer without changing the image scale.
[0080] In order to obtain more information from the input visible light and infrared images, this embodiment first feeds the stitched image into a large kernel convolution block (Convolution Block 1), and then feeds it into two small kernel convolution blocks (Convolution Block 2) for further refined feature extraction.
[0081] Step 2.3: Input the bimodal feature map into a shared fully convolutional network to obtain the weight score matrix.
[0082] To avoid gradient vanishing and learn a strong weight representation, six residual blocks are input. Considering that the gradient of the visible image can well preserve texture details, even under the influence of haze, a large amount of environmental texture details are still preserved at the near imaging point. This embodiment combines the gradient image with the output of each residual block.
[0083] The output of the residual block is input into two small kernel convolution blocks (Convolution Block 2), and then the Sigmoid activation function is used to transform the output to the range [0,1] to obtain the weight score matrix.
[0084] The weighted score matrix shows the contribution of each pixel in the infrared and visible light images to the fused image.
[0085] Step 2.4: Output the feature maps of the infrared image and the visible light image based on the bimodal feature map and weight score matrix generated by the feature extractor.
[0086] Step 3: Input the real-time image and feature map of the visible light image into the pre-trained semantic information model to perform semantic information fusion and obtain the feature map of the visible light image containing semantic information.
[0087] The pre-trained semantic information model consists of a semantic information generator and a semantic information fusion unit.
[0088] The semantic generator is used to perform semantic mapping of visible light images based on real-time images of visible light images, and obtain semantic guidance features.
[0089] The semantic information fusion unit is used to: perform further feature extraction based on the feature map of the visible light image to obtain a further feature map of the visible light image, and fuse the further feature map of the visible light image with the semantic guidance features to obtain a feature map of the visible light image containing semantic information.
[0090] like Figure 4 As shown, the semantic generator first uses a pre-trained DeepLabv3+ network to estimate semantic predictions from real-time images of the input foggy visible light image, and then employs depthwise separable convolutions to extract semantic guidance features.
[0091] like Figure 4As shown, the semantic information fusion processor first uses an efficient residual dense block (RDB) to further extract features from the feature map of the visible light image, obtaining a further feature map of the visible light image to extract more texture detail information. Finally, the extracted semantic guided features are fused with the further feature map of the visible light image extracted by the RDB to obtain a feature map of the visible light image containing semantic information, thereby completing a more refined processing of the input visible light image.
[0092] RDB is an enhanced residual network that combines residual learning and dense connections in a unified module, allowing for the extraction of multi-scale and richer feature representations.
[0093] The semantic information module is designed to consider that objects within the same semantic category often have similar structures and colors, providing strong clues for recovering object structures and details in visible light images that are obscured or distorted by fog. It utilizes the semantic correlations within objects to constrain a reasonable solution space, achieving better fusion results while preserving both large-scale structures and small-scale details from visible light images, allowing for further extraction of information within these images. The semantic information module treats high-level semantic information as important prior knowledge, guiding the proposed network to better extract detailed information preserved in visible light images, thus achieving more efficient fusion of infrared and visible light images.
[0094] Guided by advanced semantic information, this embodiment can make full use of the remaining texture structure information in visible light images affected by fog.
[0095] Step 4: Fuse the feature map of the visible light image containing semantic information and the feature map of the infrared image to obtain a fused image.
[0096] The fusion formula is expressed by the following equation:
[0097] (1)
[0098] In equation (1), F To merge images, f () performs semantic information fusion on a pre-trained semantic information processing model. I v This is a real-time image of a visible light image. I r This is a real-time image of an infrared image. W This is the weighted score matrix. For dot product operation, This is for addition operations.
[0099] In this embodiment, the content of dual-modal images is effectively combined to generate an ideal fused image containing rich details and significant information, which greatly reduces the impact of foggy environment on the fusion effect. At the same time, the fusion formula constrains the pixel intensity of the fused image to be consistent with that of the source image.
[0100] Furthermore, such as Figure 2 As shown in the figure, this embodiment provides the training steps of the pre-trained bimodal feature extraction model in step 2. The pre-trained bimodal feature extraction model is trained based on sample image data using visible light contrast constraint loss and infrared joint constraint loss.
[0101] The specific training steps include:
[0102] Step 1: Input the infrared image and the foggy visible light image from the sample image data into the pre-constructed dual-modal feature extraction model to extract features, and obtain the infrared feature map and the visible light feature map of the sample image.
[0103] Step 2: Based on the visible light image with fog in the sample image data and the visible light feature map of the sample image, perform semantic information fusion using a pre-trained semantic information model to obtain the visible light feature map of the sample image containing semantic information.
[0104] Step 3: Fuse the visible light feature map and the infrared feature map of the sample image containing semantic information to obtain the fused image of the sample image.
[0105] Step 4: Generate a mask based on the infrared image in the sample image data.
[0106] Considering that the salient information in the infrared images of the sample image data consists of content with relatively large pixel values, this embodiment designs an infrared image mask to learn more valuable salient information in the infrared images.
[0107] However, different infrared images have different pixel values. In order to capture valuable information in infrared images, contrast pixels are used as a measure, and a mask is generated using the following formula:
[0108] (2)
[0109] In equation (2), For the mask in the first i Line 1 j Column elements, For infrared images in the first i Line 1 j Column elements, The mean of all elements in the infrared image. t For predefined thresholds.
[0110] In this embodiment, the predefined threshold t Set to 50. Meanwhile, to make the mask image more clearly visible, this embodiment uses a guided filter to blur the edges, and the boundaries in the improved mask better conform to human visual perception.
[0111] Step 5: Use a mask to process the haze-free visible light image, the infrared image, and the fused image of the sample image data to obtain the masked haze-free visible light image, the masked infrared image, and the masked fused image of the sample image.
[0112] Step 6: Based on the fused image of the sample image data, including the foggy visible light image, the fog-free visible light image after masking, the masked infrared image, and the masked sample image, the pre-built bimodal feature extraction model is trained using a preset bimodal global perception loss to obtain the pre-trained bimodal feature extraction model.
[0113] The preset dual-modal global perception loss includes visible light contrast constraint and infrared joint constraint loss.
[0114] Step 6.1: Train the pre-built bimodal feature extraction model using visible light contrast constraints.
[0115] Considering that the information contained in visible light images affected by heavy fog is not ideal, and a large amount of texture structure and target information is lost, contrastive learning is applied to train a pre-built bimodal feature extraction model.
[0116] The goal of contrastive learning is to bring anchor samples together with positive samples while pushing anchor samples away from negative samples.
[0117] Step 6.1.1: Use the hazy visible light image in the sample image data as the negative sample, the hazy-free visible light image after masking as the positive sample, and the fused image of the masked sample images as the anchor sample.
[0118] Step 6.1.2: Input the negative samples, positive samples, and anchored samples into the pre-trained contrastive learning feature extraction module with fixed parameters to obtain the hidden features of each sample.
[0119] In this embodiment, the pre-trained contrastive learning feature extraction module with fixed parameters is: Network. Specifically, it adopts... The hidden features of negative samples, positive samples, and anchor samples are extracted from layers 1, 3, 5, 9, and 13 of the network.
[0120] Step 6.1.3: Calculate the contrast constraint loss between the hidden features of each sample using the following formula:
[0121] (3)
[0122] In equation (3), L c To compare the constraint loss, GT i The first positive sample i One hidden feature, F i For the anchored sample i One hidden feature, I Vi The first negative sample i One hidden feature, p i The output of the pre-built bimodal feature extraction model is the first... i The weight coefficients of each hidden feature.
[0123] In this embodiment, the weight coefficient of the i-th hidden feature output by the pre-constructed bimodal feature extraction model Set to respectively .
[0124] Step 6.1.4: In response to the contrast constraint loss being less than the preset first threshold, the anchor sample is moved closer to the positive sample and further away from the negative sample, and the parameters of the pre-constructed bimodal feature extraction model are updated.
[0125] Introducing contrast constraint loss enables the bimodal feature extraction model trained in this embodiment to better generate ideal fused images by utilizing information from both foggy and fog-free images.
[0126] Step 6.2: Train the pre-built bimodal feature extraction model using infrared joint constraint loss.
[0127] Step 6.2.1: Calculate the total infrared joint constraint loss between the fused image of the masked infrared image and the masked sample image, using the following formula:
[0128] (4)
[0129] In equation (4), L r For the total loss of infrared joint constraint, This is the first hyperparameter.
[0130] In this embodiment, the first hyperparameter The value was set to 0.5 based on the task context and experimental results.
[0131] In equation (4), The content loss is calculated using the following formula:
[0132] (5)
[0133] In equation (5), MSE () is used to calculate the mean square error loss. This is the fused image of the sample images after masking. As a mask, This is the fused image of the sample images after masking. The infrared image is from the sample image data. This is the infrared image after masking.
[0134] Content loss aims to learn the difference information between the infrared image and the fused image, and generate a fused image containing salient targets and rich information.
[0135] It can extract significant information from infrared images and make full use of the valuable information present in infrared images.
[0136] In equation (4), This is the second hyperparameter. In this embodiment, the second hyperparameter... The value is set to 1 based on the task context and experimental results.
[0137] In equation (4), The multi-scale structural similarity loss is calculated using the following formula:
[0138] (6)
[0139] In equation (6), This is the infrared image after masking. This is the fused image of the sample images after masking. for The mean, for The mean, for standard deviation for standard deviation The covariance of the fused image of the masked sample images and the masked infrared image; As the first relative importance, As the second most important, It is the first constant. It is the second constant. For the scale quantity.
[0140] In this embodiment, a first constant is set. Second constant This is to prevent division by zero. (Scale quantity) Set to 5, first relative importance Second relative importance All are set to 1.
[0141] Step 6.2.2: In response to the total loss of infrared joint constraints being less than the preset second threshold, update the parameters of the pre-constructed dual-modal feature extraction model.
[0142] It should be noted that steps 6.1 and 6.2 are performed simultaneously, and the step numbers are only used to distinguish the two steps as different training steps.
[0143] In summary, this embodiment can effectively combine bimodal image content to generate an ideal fused image containing rich details and significant information; guided by advanced semantic information, this embodiment can make full use of the remaining texture structure information in visible light images affected by fog; and can retain more semantic information and improve image fusion quality even in the face of extreme environments such as heavy fog.
[0144] Example 2:
[0145] Based on the same inventive concept as Embodiment 1, this embodiment provides an infrared image and visible light image fusion device, comprising:
[0146] Acquisition module: Used to acquire real-time images of infrared and visible light images;
[0147] Feature extraction module: This module is used to input the acquired real-time images into a pre-trained bimodal feature extraction model for feature extraction, resulting in feature maps of the infrared image and the visible light image.
[0148] Semantic information fusion module: This module is used to input real-time images and feature maps of visible light images into a pre-trained semantic information model to perform semantic information fusion, thereby obtaining a feature map of the visible light image containing semantic information.
[0149] Image fusion module: used to fuse the feature map of the visible light image containing semantic information and the feature map of the infrared image to obtain a fused image.
[0150] Example 3:
[0151] Based on the same inventive concept as other embodiments, this embodiment provides a computer-readable storage medium having a computer program / instructions stored thereon, which, when executed by a processor, implements the steps of the low-resolution image classification method described in Embodiment 1.
[0152] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0153] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0154] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0155] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0156] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.
Claims
1. A method for fusing infrared and visible light images, characterized in that, include: Acquire real-time images of infrared and visible light images; The acquired real-time images are input into a pre-trained bimodal feature extraction model for feature extraction, resulting in feature maps of the infrared image and the visible light image. The real-time image and feature map of the visible light image are input into a pre-trained semantic information model for semantic information fusion to obtain a feature map of the visible light image containing semantic information. The feature map of the visible light image containing semantic information and the feature map of the infrared image are fused to obtain a fused image; The pre-trained bimodal feature extraction model is obtained by training a pre-constructed bimodal feature extraction model based on sample image data using visible light contrast constraint loss and infrared joint constraint loss. The training process includes: The infrared image and the foggy visible light image in the sample image data are input into a pre-constructed dual-modal feature extraction model for feature extraction, resulting in the infrared feature map and the visible light feature map of the sample image. Based on the visible light images with fog in the sample image data and the visible light feature maps of the sample images, semantic information is fused using a pre-trained semantic information model to obtain the visible light feature maps of the sample images containing semantic information. The visible light feature map and infrared feature map of the sample image containing semantic information are fused to obtain the fused image of the sample image; A mask is generated based on the infrared image in the sample image data; The fog-free visible light image, the infrared image, and the fused image of the sample image data are processed using a mask to obtain the fog-free visible light image, the infrared image, and the fused image of the sample image after mask processing. Based on the sample image data, a hazy visible light image, a hazy-free visible light image after masking, an infrared image after masking, and a fused image of the sample image after masking, a pre-constructed bimodal feature extraction model is trained using a preset bimodal global perception loss to obtain a pre-trained bimodal feature extraction model; wherein, the preset bimodal global perception loss includes visible light contrast constraint and infrared joint constraint loss. The training of a pre-built bimodal feature extraction model using visible light contrast constraints includes: The sample image data includes foggy visible light images as negative samples, masked fog-free visible light images as positive samples, and the fused image of the masked sample images as anchor samples. The negative samples, positive samples, and anchored samples are input into a pre-trained contrastive learning feature extraction module with fixed parameters to obtain the hidden features of each sample. The contrast constraint loss between the hidden features of each sample is calculated using the following formula: (3), In formula (3), L c GT i is the i-th hidden feature of the positive sample, F i is the i-th hidden feature of the anchor sample, I Vi is the i-th hidden feature of the negative sample, p i is the weight coefficient of the i-th hidden feature output by the pre-constructed dual-modal feature extraction model; In response to the contrast constraint loss being less than a preset first threshold, the anchor sample is moved closer to the positive sample and further away from the negative sample, and the parameters of the pre-built bimodal feature extraction model are updated. The training of a pre-built bimodal feature extraction model using infrared joint constraint loss includes: The total infrared joint constraint loss between the masked infrared image and the fused image of the masked sample image is calculated using the following formula: (4) In equation (4), L r For the total loss of infrared joint constraint, As the first hyperparameter, The content loss is calculated using the following formula: (5) In equation (5), MSE() is the loss for calculating the mean square error. The fused image is from the sample image data. As a mask, This is the fused image after masking. The infrared image is from the sample image data. The infrared image after masking; In equation (4), This is the second hyperparameter. The multi-scale structural similarity loss is calculated using the following formula: (6) In equation (6), This is the fused image of the sample images after masking. The infrared image after masking; for The mean, for The mean, for standard deviation for standard deviation The covariance of the fused image of the masked sample image and the masked infrared image; As the first relative importance, As the second most important, It is the first constant. It is the second constant. For scale quantity; In response to the total loss due to infrared joint constraints being less than a preset second threshold, the parameters of the pre-constructed dual-modal feature extraction model are updated.
2. The infrared image and visible light image fusion method according to claim 1, characterized in that, The pre-trained bimodal feature extraction model includes a feature extractor and a weight generator. The feature extractor is used to: extract features from the infrared image and the visible light image based on the real-time images of the infrared image and the visible light image, and generate a dual-modal feature map; The weight generator is used to: obtain a weight score matrix based on the bimodal feature map, and output feature maps of the infrared image and the visible light image based on the bimodal feature map and the weight score matrix.
3. The method for fusing infrared and visible light images according to claim 1, characterized in that, The pre-trained semantic information model includes a semantic information generator and a semantic information fusion unit. The semantic information generator is used to: perform semantic mapping of the visible light image based on the real-time image of the visible light image to obtain semantic guidance features; The semantic information fusion unit is used to: perform further feature extraction based on the feature map of the visible light image to obtain a further feature map of the visible light image, and fuse the further feature map of the visible light image with the semantic guidance feature to obtain a feature map of the visible light image containing semantic information.
4. The method for fusing infrared and visible light images according to claim 1, characterized in that, The feature map of the visible light image containing semantic information and the feature map of the infrared image are fused together, and the fusion formula is expressed by the following equation: (1) In equation (1), F represents the fused image, f() represents the semantic information fusion performed by the pre-trained semantic information model, and I v For real-time images of visible light, I r Here is a real-time infrared image, and W is the weighted score matrix. For dot product operation, This is for addition operations.
5. The method for fusing infrared and visible light images according to claim 1, characterized in that, The mask is generated based on the infrared image in the sample image data using the following formula: (2) In equation (2), Let i be the element in the i-th row and j-th column of the mask. Let be the element in the i-th row and j-th column of the infrared image. is the mean of all elements in the infrared image, and t is a predefined threshold.
6. An infrared image and visible light image fusion apparatus, used to perform the infrared image and visible light image fusion method according to any one of claims 1-5, characterized in that, include: Acquisition module: Used to acquire real-time images of infrared and visible light images; Feature extraction module: This module is used to input the acquired real-time images into a pre-trained bimodal feature extraction model for feature extraction, resulting in feature maps of the infrared image and the visible light image. Semantic information fusion module: This module is used to input real-time images and feature maps of visible light images into a pre-trained semantic information model to perform semantic information fusion, thereby obtaining a feature map of the visible light image containing semantic information. Image fusion module: used to fuse the feature map of the visible light image containing semantic information and the feature map of the infrared image to obtain a fused image.
7. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instruction is executed by the processor, it implements the steps of the infrared image and visible light image fusion method according to any one of claims 1-5.