Image fusion model training method, image fusion method, electronic equipment and medium
By calculating and adjusting the network model for segmentation information loss calculations on infrared images and visible light images, the problem of low effectiveness of fusion images in subsequent segmentation tasks is solved, and a fusion image that is more suitable for segmentation tasks is achieved.
Patent Information
- Application Number
- CN202510157952.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art After the infrared image is fused with visible light image, the fused image has low effectiveness in subsequent segmentation tasks, resulting in poor accuracy of segmentation results.
By obtaining infrared images and visible light images taken on the same scene, input them into the preset network model for image fusion, and segmenting the original image and fusion images respectively, calculate the loss value to adjust the network model parameters, and obtain an image fusion model that is more suitable for the segmentation task.
The fusion image is achieved that participates more effectively in the subsequent segmentation task, which improves the performance of the fusion image in the segmentation task and ensures the effectiveness of the fusion image.
Smart Images

Figure CN120147150A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of model training, and particularly relates to an image fusion model training method, an image fusion method, an electronic device, and a medium. Background Art
[0002] The images captured by visible light sensors have rich texture details but are easily affected by the environment; the images captured by infrared sensors can highlight thermal targets under harsh environmental conditions but are prone to losing details. The technology of fusing infrared (IR) images and visible light (VIS) images is one of the most important branches of multi-source image fusion technology. This technology effectively fuses the thermal radiation information in infrared images with complementary information such as textures and edges in visible light images, and has received extensive attention and research, and has been widely applied in fields such as aerial remote sensing.
[0003] The infrared image and visible light image fusion task needs to start from an infrared image and a visible light image to obtain a fused image. Currently, the relatively common and unified approach is to use a network model to perform image fusion on the infrared image and the visible light image to obtain a fused image, and then calculate the pixel intensity loss value and the gradient loss value of the fused image with the infrared image and the visible light image respectively. The loss value is used for backpropagation training of the network model, so that the fused image obtained by the network model can retain the significant features of the two source images. It can be seen that in the prior art, the network model is trained only based on the pixel intensity loss value and the gradient loss value between images, and it focuses more on the improvement of the current image quality evaluation index.
[0004] However, the fused image is generally used for segmentation tasks in subsequent processes. If only focusing on the improvement of the current image quality evaluation index, it may lead to low effectiveness of the fused image in subsequent segmentation tasks, resulting in poor accuracy of the segmentation results obtained when the subsequent fused image is used for segmentation tasks. Summary of the Invention
[0005] The embodiments of this application provide an image fusion model training method, an image fusion method, an electronic device, and a medium, which can obtain a fused image that can more effectively participate in subsequent segmentation tasks.
[0006] In a first aspect, the embodiments of this application provide an image fusion model training method, which includes:
[0007] Obtain training samples, where the training samples include a first infrared image and a first visible light image taken of the same scene;
[0008] Input the training samples into a preset network model to obtain a first fused image corresponding to the training samples;
[0009] Segment the target objects in the first infrared image, the first visible light image, and the first fusion image respectively to obtain first segmentation information, second segmentation information, and third segmentation information;
[0010] Calculate the loss based on the first segmentation information, the second segmentation information, and the third segmentation information to determine the target loss value;
[0011] Adjust the parameters of the preset network model based on the target loss value corresponding to the training sample to obtain an image fusion model, which is used to perform image fusion on the infrared image and the visible light image to obtain a fusion image.
[0012] Combined with the first aspect, in some embodiments, the above-mentioned step of inputting the training sample into the preset network model to obtain the first fusion image corresponding to the training sample includes: inputting the training sample into the preset network model, so that the preset network model performs the following processing: adding first Gaussian noise to the first infrared image to obtain a second infrared image; adding second Gaussian noise to the first visible light image to obtain a second visible light image; performing image fusion processing on the second infrared image and the second visible light image to obtain a second fusion image; obtaining the first fusion image based on the first infrared image, the first visible light image, and the second fusion image.
[0013] Combined with the first aspect, in some embodiments, the above-mentioned step of performing image fusion processing on the second infrared image and the second visible light image to obtain a second fusion image includes: determining a target prediction noise according to the first infrared image, the second infrared image, the first visible light image, and the second visible light image, where the target prediction noise is the noise that needs to be removed when performing image fusion processing on the second infrared image and the second visible light image; using the DDIM algorithm to obtain the second fusion image based on the target prediction noise.
[0014] In combination with the first aspect, in some embodiments, the determination of the target prediction noise based on the first infrared image, the second infrared image, the first visible light image, and the second visible light image includes: predicting a first prediction noise based on the first infrared image and the second infrared image, where the first prediction noise is the noise that needs to be removed from the second infrared image when performing image fusion processing on the second infrared image and the second visible light image; predicting a second prediction noise based on the first visible light image and the second visible light image, where the second prediction noise is the noise that needs to be removed from the second visible light image when performing image fusion processing on the second infrared image and the second visible light image; performing a weighted summation calculation on the first prediction noise and the second prediction noise to determine the target prediction noise.
[0015] In combination with the first aspect, in some embodiments, the obtaining of the first fusion image based on the first infrared image, the first visible light image, and the second fusion image includes: respectively performing feature extraction on the first infrared image, the first visible light image, and the second fusion image to obtain corresponding first image features, second image features, and third image features; performing fusion processing on the first image features, the second image features, and the third image features to obtain target image features; generating a third fusion image based on the target image features; and embedding the first segmentation information and the second segmentation information into the third fusion image to obtain the first fusion image.
[0016] In combination with the first aspect, in some embodiments, the first segmentation information includes the target values corresponding to each pixel in the first infrared image, the second segmentation information includes the target values corresponding to each pixel in the first visible light image, the third segmentation information includes the target values corresponding to each pixel in the first fusion image, and the target value is used to represent whether the corresponding pixel is the area where the target object is located; the determination of the target loss value based on the first segmentation information, the second segmentation information, and the third segmentation information includes: using a preset first loss function to calculate the loss between the target values corresponding to each pixel in the first infrared image and the target values corresponding to each pixel in the first fusion image to obtain a first loss value, and using the first loss function to calculate the loss between the target values corresponding to each pixel in the first visible light image and the target values corresponding to each pixel in the first fusion image to obtain a second loss value; performing a summation calculation on the first loss value and the second loss value to determine the target loss value.
[0017] In combination with the first aspect, in some embodiments, before adjusting the parameters of the preset network model based on the target loss value corresponding to the training sample to obtain an image fusion model, the method further includes: calculating a loss between the noise distribution function corresponding to the first predicted noise and the Gaussian distribution function corresponding to the first Gaussian noise by using a preset second loss function to obtain a third loss value; calculating a loss between the noise distribution function corresponding to the second predicted noise and the Gaussian distribution function corresponding to the second Gaussian noise by using the second loss function to obtain a fourth loss value; the first segmentation information includes target values corresponding to each pixel in the first infrared image, the second segmentation information includes target values corresponding to each pixel in the first visible light image, the third segmentation information includes target values corresponding to each pixel in the first fused image, and the target value is used to represent whether the corresponding pixel is the area where the target object is located;
[0018] In combination with the first aspect, in some embodiments, calculating the target loss value based on the first segmentation information, the second segmentation information, and the third segmentation information includes: calculating a loss between the target value corresponding to each pixel in the first infrared image and the target value corresponding to each pixel in the first fused image by using a preset first loss function to obtain a first loss value; calculating a loss between the target value corresponding to each pixel in the first visible light image and the target value corresponding to each pixel in the first fused image by using the first loss function to obtain a second loss value; performing a summation calculation on the first loss value, the second loss value, the third loss value, and the fourth loss value to determine the target loss value.
[0019] In combination with the first aspect, in some embodiments, before performing the summation calculation on the first loss value, the second loss value, the third loss value, and the fourth loss value to determine the target loss value, the method further includes: calculating a loss between the first fused image and the first infrared image by using a preset third loss function to obtain a fifth loss value; calculating a loss between the first fused image and the first visible light image by using the third loss function to obtain a sixth loss value.
[0020] In combination with the first aspect, in some embodiments, performing the summation calculation on the first loss value, the second loss value, the third loss value, and the fourth loss value to determine the target loss value includes: performing a summation calculation on the first loss value, the second loss value, the third loss value, the fourth loss value, the fifth loss value, and the sixth loss value to determine the target loss value.
[0021] In a second aspect, an embodiment of the present application provides an image fusion method, and the method includes:
[0022] Obtain an infrared image and a visible light image captured for the same scene;
[0023] Input the infrared image and the visible light image into an image fusion model to obtain a target fusion image, where the image fusion model is obtained according to the image fusion model training method described in the first aspect.
[0024] In a third aspect, an embodiment of the present application provides an image fusion model training device, and the device includes:
[0025] A sample acquisition module, configured to acquire training samples, where the training samples include a first infrared image and a first visible light image captured for the same scene;
[0026] A model output module, configured to input the training samples into a preset network model to obtain a first fusion image corresponding to the training samples;
[0027] A segmentation module, configured to respectively segment target objects in the first infrared image, the first visible light image, and the first fusion image to obtain first segmentation information, second segmentation information, and third segmentation information;
[0028] A calculation module, configured to perform loss calculation based on the first segmentation information, the second segmentation information, and the third segmentation information to determine a target loss value;
[0029] A training module, configured to adjust parameters of the preset network model based on the target loss value corresponding to the training samples to obtain an image fusion model, where the image fusion model is used to perform image fusion on an infrared image and a visible light image to obtain a fusion image.
[0030] In a fourth aspect, an embodiment of the present application provides an electronic device, and the electronic device includes: a processor and a memory storing computer program instructions;
[0031] When the processor executes the computer program instructions, the image fusion model training method described in the first aspect is implemented.
[0032] In a fifth aspect, an embodiment of the present application provides a computer storage medium, and computer program instructions are stored on the computer-readable storage medium. When the computer program instructions are executed by a processor, the image fusion model training method described in the first aspect is implemented.
[0033] In a sixth aspect, an embodiment of the present application provides a computer program product. When instructions in the computer program product are executed by a processor of an electronic device, the electronic device is caused to execute the image fusion model training method described in the first aspect.
[0034] The image fusion model training method, device, equipment, computer storage medium and computer program product according to the embodiments of the present application can obtain a fused image corresponding to a first infrared image and a first visible light image, respectively segment target objects in the first infrared image, the first visible light image and the fused image to obtain first segmentation information, second segmentation information and third segmentation information, calculate a loss based on the first segmentation information, the second segmentation information and the third segmentation information to determine a target loss value, adjust the parameters of a preset network model based on the target loss value to obtain an image fusion model, and the image fusion model is used to perform image fusion on an infrared image and a visible light image to obtain a fused image. It can be seen that the present application trains a preset network model based on the target loss value determined by calculating the loss of the segmentation information of the image. Thus, the obtained image fusion model synthesizes the segmentation information of the image, making the fused image obtained by the image fusion model more suitable for subsequent segmentation tasks, and a more practical fused image can be obtained through the image fusion model. The present application can fully combine the image fusion task with the segmentation task after the fused image is completed, and can ensure the effectiveness of the fused image when used in subsequent segmentation tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments of the present application. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0036] Figure 1 is a flowchart of the image fusion model training method provided by the embodiments of the present application;
[0037] Figure 2 is a processing flowchart of the diffusion model provided by the embodiments of the present application;
[0038] Figure 3 is a processing flowchart of the fused image modification module provided by the embodiments of the present application;
[0039] Figure 4 is a flowchart of a loss calculation provided by the embodiments of the present application;
[0040] Figure 5 is a flowchart of an image fusion method provided by the embodiments of the present application;
[0041] Figure 6 is a structural block diagram of an image fusion model training device provided by the embodiments of the present application;
[0042] Figure 7 is a structural block diagram of an image fusion device provided by the embodiments of the present application;
[0043] Figure 8 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0044] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than limiting the present application. For those skilled in the art, the present application can be implemented without some of these specific details. The following description of the embodiments is only intended to provide a better understanding of the present application by showing examples of the present application.
[0045] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, elements defined by the statement "including..." do not exclude the existence of additional identical elements in the process, method, article or device including the said elements.
[0046] Currently, the fusion effect of infrared images and visible light images is usually measured by various image evaluation metrics. For example, mutual information (MI), cross entropy (EN), peak signal-to-noise ratio (PSNR), etc. Mutual information is used to measure the amount of features transferred from the source image to the fused image; cross entropy is used to measure the amount of information transferred from the source image to the fused image; peak signal-to-noise ratio characterizes the ratio of peak power to noise power in the fused image, and can reflect the distortion situation in the fusion process from the pixel level.
[0047] Although the above image evaluation metrics can reflect the effect of image fusion to a certain extent, they cannot directly illustrate the performance of the fused image in other downstream tasks (such as image segmentation, object detection, etc.). Therefore, evaluating the algorithm solution in this way will inevitably lead to a situation where the evaluation metric of the fused image is improved but the performance of the downstream task cannot be further improved. This has limited the further popularization of the infrared image and visible light image fusion technology to a certain extent.
[0048] To solve the problems of the prior art, the embodiments of the present application provide an image fusion model training method, an image fusion method, an image fusion model training device, an electronic device, a computer storage medium, and a computer program product. First, the image fusion model training method provided by the embodiments of the present application will be introduced below.
[0049] Referring to Figure 1 , Figure 1 shows a schematic flowchart of an image fusion model training method provided by an embodiment of the present application. As Figure 1 shown, the image fusion model training method includes the following steps S101 - step S105:
[0050] Step S101, obtain training samples.
[0051] Step S102, input the training samples into a preset network model to obtain a first fused image corresponding to the training samples.
[0052] Step S103, respectively segment the target objects in the first infrared image, the first visible light image, and the first fused image to obtain first segmentation information, second segmentation information, and third segmentation information.
[0053] Step S104, calculate the loss based on the first segmentation information, the second segmentation information, and the third segmentation information to determine the target loss value.
[0054] Step S105, adjust the parameters of the preset network model based on the target loss value corresponding to the training samples to obtain an image fusion model.
[0055] The above image fusion model is used to perform image fusion on infrared images and visible light images to obtain a fused image.
[0056] The above training samples include a first infrared image and a first visible light image taken of the same scene.
[0057] In step S101 of some embodiments, the training samples can be historical infrared images and visible light images.
[0058] The above preset network model can be a network model composed of a diffusion model and a Restormer model.
[0059] In step S102 of some embodiments, when the training samples are input into the preset network model, the preset network model adds noise to the training samples to obtain the noisy training samples, and then performs denoising fusion processing on the noisy training samples to obtain the first denoised fused image. That is, the noisy training samples are gradually denoised and fused to obtain the first fused image.
[0060] The above first segmentation information may include target values corresponding to each pixel in the first infrared image, the above second segmentation information may include target values corresponding to each pixel in the first visible light image, and the above third segmentation information may include target values corresponding to each pixel in the first fused image, where the target value is used to characterize whether the corresponding pixel is the area where the target object is located.
[0061] In step S103 of some embodiments, the first segmentation information, the second segmentation information, and the third segmentation information may be a first segmentation mask, a second segmentation mask, and a third segmentation mask respectively. Among them, the segmentation mask is usually a binary image or a multi-valued image, and the target value may be the corresponding pixel value in the segmentation mask. In one implementation, a SAM image predictor can be used to segment the target object in the first infrared image, the first visible light image, and the first fused image to obtain the first segmentation mask, the second segmentation mask, and the third segmentation mask.
[0062] For a segmentation mask that is a binary image, there are only two pixel values, 0 or 1. 0 represents the mask area, and 1 represents the non-mask area. In an image segmentation task, each target object can be represented by a unique segmentation mask, where the area with a pixel value of 1 represents the part belonging to the target object, and the area with a pixel value of 0 represents the part that does not belong to the target object.
[0063] In one implementation, after processing the first infrared image, the first visible light image, and the first fused image through the SAM image predictor, the first initial float-type mask, the second initial float-type mask, and the third initial float-type mask are correspondingly output. The sigmoid operation is used to process the first initial float-type mask, the second initial float-type mask, and the third initial float-type mask to obtain the first segmentation mask, the second segmentation mask, and the third segmentation mask.
[0064] In step S104 of some embodiments, a preset loss function can be used to calculate the loss of the first segmentation information, the second segmentation information, and the third segmentation information to determine the target loss value.
[0065] The above image fusion model is used to perform image fusion on an infrared image and a visible light image to obtain a fused image.
[0066] In step S105 of some embodiments, the parameters of the preset network model are adjusted according to the target loss value to obtain an image fusion model with the loss function meeting the requirements.
[0067] In the embodiments of the present application, based on the target loss value determined by calculating the loss of the segmentation information of the image, a preset network model is trained. The image fusion model obtained in this way incorporates the segmentation information of the image, making the fused image obtained through the image fusion model more suitable for subsequent segmentation tasks. Through the image fusion model, a more practical fused image can be obtained. The present application can fully combine the image fusion task with the segmentation task after the fused image is completed, and can ensure the effectiveness of the fused image when used in subsequent segmentation tasks.
[0068] In some embodiments, in step S102 above, the training samples are input into the preset network model so that the preset network model can perform but is not limited to performing the following steps:
[0069] Add first Gaussian noise to the first infrared image to obtain a second infrared image, and add second Gaussian noise to the first visible light image to obtain a second visible light image.
[0070] Among them, for adding noise to the image, the formula of the forward noise addition process of the diffusion model can be used. For example, for the time step t, using the formula of the forward noise addition process of the diffusion model, add the corresponding first Gaussian noise and second Gaussian noise to the first infrared image and the second visible light source image respectively. The noise addition formula is as follows:
[0071]
[0072] Perform image fusion processing on the second infrared image and the second visible light image to obtain a second fused image.
[0073] Since the prediction network model has a certain randomness as a generative model, in order to further improve the image quality, the second fused image can also be adjusted and optimized.
[0074] Based on the first infrared image, the first visible light image, and the second fused image, obtain a first fused image.
[0075] In this embodiment, by gradually adding noise to source images such as the first infrared image and the first visible light image, the second fused image and the first fused image for training the preset network model are obtained, so that the preset network model can learn the features of source images such as the first infrared image and the first visible light image, which is beneficial to improving the image fusion ability of the preset network model.
[0076] In some embodiments, the step of performing image fusion processing on the second infrared image and the second visible light image to obtain a second fused image may specifically further include the following steps:
[0077] Determine the target prediction noise according to the first infrared image, the second infrared image, the first visible light image, and the second visible light image.
[0078] The above target prediction noise is the noise that needs to be removed when performing image fusion processing on the second infrared image and the second visible light image.
[0079] The DDIM algorithm is adopted to obtain the second fused image based on the target prediction noise.
[0080] Among them, in the training stage, a time step t (t ∈ [1, 1000]) is randomly sampled for each training iteration. That is to say, each iteration of training is the denoising of the t-th single step. Using the image generation sampling principle of DDIM (Denoising Diffusion Implicit Models), the noise-free fused image at the t-th step can be obtained. The formula for DDIM sampling is as follows:
[0081]
[0082]
[0083] This formula represents denoising the noisy image x at the k-th step k to obtain the image x at the s-th step s , x 0 is the second fused image obtained by removing the target prediction noise from x k , and ε is the standard Gaussian distribution. For the acquisition of the above noisy image, k = 1000 and s = t. Among them, α t and β t are a series of preset parameters. The noisy image is the noise image corresponding to the target prediction noise.
[0084] In this embodiment, the target prediction noise is determined through the first infrared image, the second infrared image, the first visible light image, and the second visible light image, and the second infrared image and the second visible light image are denoised and fused based on the target prediction noise to obtain the second fused image, which can restore an image containing the features of the source image from a Gaussian noise image through step-by-step denoising, improving the effectiveness of obtaining the second fused image.
[0085] In some embodiments, the above determination of the target prediction noise may include but is not limited to the following steps:
[0086] Based on the first infrared image and the second infrared image, the first prediction noise is predicted.
[0087] The above first prediction noise is the noise that needs to be removed from the second infrared image when performing image fusion processing on the second infrared image and the second visible light image.
[0088] Based on the first visible light image and the second visible light image, the second predicted noise is obtained through prediction.
[0089] The second predicted noise is the noise that needs to be removed from the second visible light image when performing image fusion processing on the second infrared image and the second visible light image.
[0090] In one implementation, the above prediction network model includes a first noise prediction module and a second noise prediction module. Among them, the first noise prediction module is used to predict the first predicted noise, and the second noise prediction module is used to predict the second predicted noise.
[0091] Specifically, the first infrared image and the second infrared image are concatenated in the channel dimension to obtain a first concatenated image, and the first concatenated image is used as the input of the first noise prediction module to obtain the first predicted noise output by the first noise prediction module. The first visible light image and the second visible light image are concatenated in the channel dimension to obtain a second concatenated image, and the second concatenated image is used as the input of the second noise prediction module to obtain the second predicted noise output by the second noise prediction module.
[0092] Perform a weighted sum calculation on the first predicted noise and the second predicted noise to determine the target predicted noise. Among them, the weights corresponding to the first predicted noise and the second predicted noise can each be 0.5.
[0093] In this embodiment, the two noise prediction modules can respectively learn the features of the infrared image and the visible light image, so that in the subsequent process of generating a fused image based on the target predicted noise, each step of denoising actually includes two parts: denoising for the infrared image and denoising for the visible light image. In other words, the noise removed for the infrared image is the gradient from the noisy image to the infrared image, and the noise removed for the visible light image is the gradient from the noisy image to the visible light image. Then, the gradient of generating an image towards the fused image can be considered as the weighted sum of the above two parts of gradients. Therefore, the process of generating a fused image can start from a Gaussian noise, and the noise removed in each step is the weighted sum of the infrared image noise and the visible light image noise. The finally generated fused image will simultaneously contain the features of the infrared image and the visible light image, which can improve the effectiveness of obtaining the first fused image.
[0094] In some embodiments, the steps of obtaining the first fused image may include but are not limited to the following steps:
[0095] Feature extraction is respectively performed on the first infrared image, the first visible light image, and the second fused image to obtain corresponding first image features, second image features, and third image features.
[0096] The first image features, the second image features, and the third image features are subjected to fusion processing to obtain target image features.
[0097] Generate a third fused image based on the target image features.
[0098] Embed the first segmentation information and the second segmentation information into the third fused image to obtain the first fused image.
[0099] In one embodiment, the step of optimizing the second fused image to obtain the first fused image can be performed by the Restormer model.
[0100] Specifically, the Restormer model performs image fusion processing on the first infrared image, the first visible light image, and the second fused image to obtain the third fused image output by the Restormer model in the prediction network model. Input the first segmentation information and the second segmentation information into the prediction network model, and the first fused image is output by the Restormer model in the prediction network model.
[0101] Therefore, the Restormer model performs the following processing: respectively extract the features of the first infrared image, the first visible light image, and the second fused image to obtain the corresponding first image features, second image features, and third image features, and fuse the first image features, second image features, and third image features to obtain the target image features. Finally, generate a third fused image based on the target image features. Embed the first segmentation information and the second segmentation information into the third fused image to obtain the first fused image.
[0102] In one embodiment, the time step can also be input into the Restormer model, and the Restormer model outputs the first fused image embedded with the time step information.
[0103] In this embodiment, the present invention uses the Restormer model and combines it with the SAM image predictor to perform segmentation task-guided image fusion. The Restormer model is used to extract the features of the stacked source images and fused images respectively, and complete dimensionality reduction and fusion through convolution, and then perform the embedding of the time step and segmentation information, so that the obtained first fused image is more suitable for training the network model to obtain a fused image network model that can obtain a fused image more beneficial to downstream tasks.
[0104] In some embodiments, the first segmentation information includes the target values corresponding to each pixel in the first infrared image, the second segmentation information includes the target values corresponding to each pixel in the first visible light image, and the third segmentation information includes the target values corresponding to each pixel in the first fused image. The target value is used to represent whether the corresponding pixel is the area where the target object is located. The above step S104 may include but is not limited to the following steps:
[0105] Calculate the loss between the target values corresponding to each pixel in the first infrared image and the target values corresponding to each pixel in the first fusion image using a preset first loss function to obtain a first loss value.
[0106] Calculate the loss between the target values corresponding to each pixel in the first visible light image and the target values corresponding to each pixel in the first fusion image using the first loss function to obtain a second loss value.
[0107] The above first loss function can be the L1 - Loss function, and the L1 - Loss function can be as follows:
[0108]
[0109] Where, can be the first loss value. If is the first loss value, then m is the number of pixels in the first infrared image, i represents the pixel serial number, and y (i) represents the target value corresponding to the i - th pixel in the first fusion image, represents the target value corresponding to the i - th pixel in the first infrared image.
[0110] If is the second loss value, then m is the number of pixels in the first visible light image, i represents the pixel serial number, and y (i) represents the target value corresponding to the i - th pixel in the first fusion image, represents the target value corresponding to the i - th pixel in the first visible light image.
[0111] Sum the first loss value and the second loss value to determine the target loss value.
[0112] The summation calculation formula is as follows:
[0113]
[0114] Where, Loss T represents the target loss value, a and b are the coefficients corresponding to the first loss value and the second loss value respectively, are the first loss value and the second loss value respectively.
[0115] In one implementation, a and b can both be 1, that is, the first loss value and the second loss value are added to obtain the target loss value.
[0116] In this embodiment, by comparing the target values in the two images, the difference in segmentation between the two can be effectively compared, and the target loss value can be accurately calculated.
[0117] In some embodiments, before the above step S105, it may further include but is not limited to the following steps:
[0118] Use a preset second loss function to calculate the loss between the noise distribution function corresponding to the first predicted noise and the Gaussian distribution function corresponding to the first Gaussian noise, obtaining a third loss value.
[0119] Use the second loss function to calculate the loss between the noise distribution function corresponding to the second predicted noise and the Gaussian distribution function corresponding to the second Gaussian noise, obtaining a fourth loss value.
[0120] Among them, the noise distribution function corresponding to the first predicted noise is the probability density function corresponding to the first predicted noise, the Gaussian distribution function corresponding to the first Gaussian noise is the probability density function that follows a Gaussian distribution corresponding to the first Gaussian noise, the noise distribution function corresponding to the second predicted noise is the probability density function corresponding to the second predicted noise, and the Gaussian distribution function corresponding to the second Gaussian noise is the probability density function that follows a Gaussian distribution corresponding to the second Gaussian noise.
[0121] The above second loss function can be the same as the first loss function, and the second loss function can be the L1 - Loss function.
[0122] If the second loss function is the L1 - Loss function, in the L1 - Loss function, if is the third loss value, then m is the number of independent variables, i represents the serial number of the independent variable, y (i) represents the function value corresponding to the i - th independent variable in the noise distribution function corresponding to the first predicted noise, represents the function value corresponding to the i - th independent variable in the Gaussian distribution function corresponding to the first Gaussian noise; if is the fourth loss value, then m is the number of independent variables, i represents the serial number of the independent variable, y (i) represents the function value corresponding to the i - th independent variable in the noise distribution function corresponding to the first predicted noise, represents the function value corresponding to the i - th independent variable in the Gaussian distribution function corresponding to the second Gaussian noise.
[0123] Then the above step S104 may include but is not limited to the following steps:
[0124] Use a preset first loss function to calculate the loss between the target values corresponding to each pixel in the first infrared image and the target values corresponding to each pixel in the first fusion image, obtaining a first loss value.
[0125] Use the first loss function to calculate the loss between the target values corresponding to each pixel in the first visible - light image and the target values corresponding to each pixel in the first fusion image, obtaining a second loss value.
[0126] The calculation methods of the first loss value and the second loss value can be referred to as above, and will not be elaborated here one by one.
[0127] Sum up the first loss value, the second loss value, the third loss value and the fourth loss value to determine the target loss value.
[0128] The summation calculation formula can be as follows:
[0129]
[0130] Among them, Loss T represents the target loss value, a, b, c, and d are the coefficients corresponding to the first loss value, the second loss value, the third loss value, and the fourth loss value respectively, are the first loss value, the second loss value, the third loss value, and the fourth loss value respectively.
[0131] In one embodiment, a, b, c, and d can all be 1, that is, the first loss value, the second loss value, the third loss value, and the fourth loss value are added together to obtain the target loss value.
[0132] In this embodiment, by adding the third loss value and the fourth loss value to the target loss value, the subsequent adjustment of the prediction ability in the model can be realized, and the effectiveness of model training can be improved.
[0133] In some embodiments, before summing up the first loss value, the second loss value, the third loss value, and the fourth loss value to determine the target loss value, it may include but is not limited to the following steps:
[0134] Use a preset third loss function to calculate the loss between the first fused image and the first infrared image to obtain a fifth loss value.
[0135] Use the third loss function to calculate the loss between the first fused image and the first visible light image to obtain a sixth loss value.
[0136] The above third loss function can be the same as the first loss function or the second loss function. The third loss function can be the L1-Loss function, or can be a conventional pixel intensity loss, gradient loss, and structural similarity loss calculation function.
[0137] If the third loss function is the L1-Loss function, in the L1-Loss function, if it is the fifth loss value, then m is the number of pixels in the first infrared image, i represents the pixel serial number, and y (i) represents the pixel value corresponding to the i-th pixel in the first infrared image, represents the pixel value corresponding to the i-th pixel in the first fused image; if is the sixth loss value, where m is the number of pixels in the first visible light image, i represents the pixel serial number, and y (i) represents the pixel value corresponding to the i-th pixel in the first visible light image, represents the pixel value corresponding to the i-th pixel in the first fused image.
[0138] Sum up the first loss value, the second loss value, the third loss value, the fourth loss value, the fifth loss value, and the sixth loss value to determine the target loss value.
[0139] The summation calculation formula is as follows:
[0140]
[0141] where Loss T represents the target loss value, and a, b, c, d, e, f are the coefficients corresponding to the first loss value, the second loss value, the third loss value, the fourth loss value, the fifth loss value, and the sixth loss value respectively, are the first loss value, the second loss value, the third loss value, the fourth loss value, the fifth loss value, and the sixth loss value respectively.
[0142] In one implementation, a, b, c, d, e, f can all be 1, that is, the first loss value, the second loss value, the third loss value, the fourth loss value, the fifth loss value, and the sixth loss value are added together to obtain the target loss value.
[0143] In this embodiment, by adding the fifth loss value and the sixth loss value to the target loss value, the subsequent adjustment of the feature fusion ability in the model can be realized, and the effectiveness of model fusion can be improved.
[0144] In some embodiments, there can be multiple training samples. For each training sample in the multiple training samples, execute step S101-step S105 to train the preset network model.
[0145] To better understand the above image fusion model training method, the embodiment of the present invention provides a complete process example of the image fusion model training method, which is as follows.
[0146] The preset network model is a network structure composed of a diffusion model and a Restormer model. In this embodiment, the structure of the diffusion model is improved so that the diffusion model includes two noise prediction modules, namely the first noise prediction module and the second noise prediction module.
[0147] Obtain multiple training samples. The training samples include the first infrared image and the first visible light image taken of the same scene, and use the multiple training samples to train the preset network model.
[0148] Specifically, the following processing is performed on each of the multiple training samples:
[0149] Refer to Figure 2 , Figure 2 which is the processing flow chart of the diffusion model shown in the embodiments of the present application.
[0150] Input the first infrared image and the first visible light image in the training sample into a preset network model.
[0151] The diffusion model performs the following processing:
[0152] Add the first Gaussian noise to the first infrared image to obtain a second infrared image, and add the second Gaussian noise to the first visible light image to obtain a second visible light image. Wherein, the first Gaussian noise and the second Gaussian noise can be the same Gaussian noise or different Gaussian noises.
[0153] Among them, adding noise to the image can use the formula of the forward noise addition process of the diffusion model. For example, for the time step t, use the formula of the forward noise addition process of the diffusion model to add the corresponding first Gaussian noise and second Gaussian noise to the first infrared image and the second visible light source image respectively. The noise addition formula is as follows:
[0154]
[0155] Then, splice the first infrared image and the second infrared image after noise addition to obtain a first spliced image, send the first spliced image into the corresponding first noise prediction module to obtain a first predicted noise, splice the first visible light image and the second visible light image after noise addition to obtain a second spliced image, and send the second spliced image into the corresponding second noise prediction module to obtain a second predicted noise.
[0156] Perform a weighted sum calculation on the first predicted noise and the second predicted noise to determine the target predicted noise. Among them, the weights corresponding to the first predicted noise and the second predicted noise can each be 0.5.
[0157] Adopt the DDIM algorithm to obtain a second fused image based on the target predicted noise.
[0158] In the training stage, a time step t (t ∈ [1, 1000]) is randomly sampled in each training iteration. That is to say, each iteration of training is to denoise the t-th step single step. The t-th step may be an intermediate process. At this time, it is necessary to first obtain the noisy image of this intermediate process. This noisy image can be denoised from the pure noise image to the t-th step by the image generation sampling principle of DDIM (Denoising Diffusion Implicit Models) to obtain the denoised fused image of the t-th step.
[0159] Among them, the formula for DDIM sampling is as follows:
[0160]
[0161] This formula represents denoising the noisy image \(x_k\) at the \(k\)-th step k to obtain the denoised image \(x_s\) at the \(s\)-th step s , where \(x_{s'}\) 0 is the second fused image after removing the target prediction noise from \(x_{s}\) k , and \(\epsilon\) follows a standard Gaussian distribution. The noisy image is the noise image corresponding to the target prediction noise.
[0162] For the acquisition of the above-mentioned noisy image, \(k = 1000\) and \(s=t\). Among them, \(\alpha\) t and \(\beta\) t are a series of pre-set parameters.
[0163] The above is the processing process performed by the diffusion model.
[0164] Next, the processing process of the fused image modification module is described. Referring to Figure 3 , Figure 3 shows the processing flow chart of the fused image modification module provided by the embodiment of the present application. The fused image modification module includes the SAM (Segment Anything Model) and the Restormer model.
[0165] Use the SAM (Segment Anything Model) to segment the target objects in the first infrared image and the first visible light image to obtain the first segmentation mask and the second segmentation mask. That is, input the first infrared image and the first visible light image into the SAM to obtain the first segmentation mask and the second segmentation mask output by the SAM.
[0166] Among them, the image predictor in the SAM (Segment Anything Model) is used for the segmentation task. The SAM needs to provide points or rectangular boxes as segmentation hints. To enable the solution of the present invention to adapt to various data sets without specifically preparing hint inputs for segmentation, the LOD technology is used to implement an automated method for generating segmentation hint rectangular boxes, which supports artificially specifying the LOD (Level of Detail), and generates rectangular boxes according to levels for hinting the SAM segmentation.
[0167] The method for generating the segmentation hint rectangular box is specifically as follows:
[0168] It can be hierarchically divided into LOD0, LOD1, LOD2, etc., denoted as LODi, where i = 0, 1, 2... LOD0 means the rectangular box is the size of the entire input image selected by the box. On the basis of LOD0, LOD1 divides the input image into blocks according to a preset number. For example, if the preset number is 2 i+1 , then the preset number is 2 * 2, generating rectangular boxes that frame four equal parts. Adding the large rectangular box of LOD0 gives a total of 5 rectangular box prompts; and so on for LOD2, which divides each sub-image into a new layer of four equal parts. Therefore, there are a total of 1 + 2 * 2 + 2 * 2 * 2 * 2 = 21 rectangular boxes. The algorithm is default set to LOD2, and the input image can be the first infrared image or the first visible light image.
[0169] The mask output by the original SAM is a Boolean - type tensor. It compares the initial float - type mask with a set mask threshold. Pixels exceeding the threshold are True, and vice versa. The pixels finally marked as True are the pixels occupied by the segmented target object. However, the Boolean - type tensor is not convenient for calculating the loss to back - propagate and train the network. Therefore, the present invention makes certain modifications to SAM so that it directly obtains the original float - type mask tensor, that is, it removes the operation of comparing the final part of the SAM - output mask with the threshold to generate a Boolean tensor, and directly outputs the obtained float mask. Then, a sigmoid operation is performed on it to control its value between 0 and 1. At this time, each pixel value in the mask represents the probability that the pixel belongs to the part of the image to be segmented, which is convenient for subsequent Loss calculation.
[0170] Then, the first infrared image, the first visible light image, the second fusion image, the first segmentation mask, the second segmentation mask, and the time step are input into the Restormer model, and the Restormer model outputs the first fusion image. In this embodiment, the Restormer model used is the Restormer model with good performance in the field of high - resolution image reconstruction, and the encoding of the time step is the same as the time step when adding noise to the first infrared image and the first visible light image.
[0171] Specifically, the Restormer model makes the following processing:
[0172] Mainly, the TransformerBlock of the Restormer model is used to extract features from the first infrared image, the first visible light image, and the second fused image respectively, obtaining the corresponding first image feature, second image feature, and third image feature. And through convolution, dimensionality reduction and fusion processing of the first image feature, second image feature, and third image feature are completed to obtain the target image feature, and the third fused image is generated based on the target image feature. The time step and the first segmentation mask and the second segmentation mask are embedded into the third fused image to obtain the first fused image, enabling the model to truly learn the segmentation information when entering the subsequent stage with time step information and segmentation information.
[0173] The above steps are preset to be performed four times in the algorithm, and the final output part is a connection of a TransformerBlock and Conv2d, obtaining the final modified fused image that comprehensively combines the information of the first infrared image and the first visible light image and the segmentation information for guidance, that is, the first fused image.
[0174] In this embodiment, the input first infrared image, first visible light image, and second fused image are all convolved to the preset number of channels (the algorithm is preset to 36 channels). As mentioned above, the SAM image predictor algorithm supports different LOD levels, and the number, size, and area of the masks are different under different settings. Therefore, it is necessary to encode the input original masks. The Encoding module here stacks and convolves the masks at each level respectively, and the number of channels of its convolutional output is the LOD level number + 1. For example, LOD0 contains 1 mask, with 1 input channel and 0 + 1 = 1 output channel; LOD1 contains 4 masks, with 4 input channels and 1 + 1 = 2 output channels. The algorithm defaults to LOD2, so the input mask has 1 + 4 + 16 = 21 channels and the output has 1 + 2 + 3 = 6 channels. The subsequent MaskEmbedding mainly performs channel normalization, then performs a Swish operation to add non-linear features, and finally convolves it to the channels suitable for the intermediate tensor for directly adding to complete the embedding of the segmentation mask information.
[0175] The first fused image is input into SAM to obtain the third segmentation mask output by SAM.
[0176] Refer to Figure 4 , Figure 4 which is a schematic flowchart of a loss calculation provided by an embodiment of the present invention.
[0177] The L1-Loss function is used to calculate the loss between the noise distribution function corresponding to the first predicted noise and the Gaussian distribution function corresponding to the first Gaussian noise, obtaining a third loss value denoted as Loss3. The L1-Loss function is also used to calculate the loss between the noise distribution function corresponding to the second predicted noise and the Gaussian distribution function corresponding to the second Gaussian noise, obtaining a fourth loss value denoted as Loss4.
[0178] Conventional pixel intensity loss, gradient loss, and structural similarity loss are calculated for the first fused image and the first infrared image, obtaining a fifth loss value denoted as Loss5. For example, by separately calculating the conventional pixel intensity loss, gradient loss, and structural similarity loss for the first fused image and the first infrared image to obtain the corresponding loss values, and denoting all the loss values obtained from calculating the conventional pixel intensity loss, gradient loss, and structural similarity loss for the first fused image and the first infrared image as the fifth loss value. Conventional pixel intensity loss, gradient loss, and structural similarity loss are calculated for the first fused image and the first visible light image, obtaining a sixth loss value denoted as Loss6. For example, by separately calculating the conventional pixel intensity loss, gradient loss, and structural similarity loss for the first fused image and the first visible light image to obtain the corresponding loss values, and denoting all the loss values obtained from calculating the conventional pixel intensity loss, gradient loss, and structural similarity loss for the first fused image and the first visible light image as the sixth loss value.
[0179] The L1-Loss function is used to calculate the loss between the first segmentation mask and the third segmentation mask, obtaining a first loss value denoted as Loss1. The L1-Loss function is used to calculate the loss between the second segmentation mask and the third segmentation mask, obtaining a second loss value denoted as Loss2.
[0180] All the losses for the model training have been calculated. Adding Loss1, Loss2, Loss3, Loss4, Loss5, and Loss6 gives the final target loss value Loss. In each iteration, backpropagation is performed to update the model parameters to complete the training. Among them, Loss3 and Loss4 can be used to train the noise prediction module, while Loss5, Loss6, Loss1, and Loss2 are used to guide the training of the Restormer model.
[0181] In the embodiments of this application, the above method provides a network model based on a diffusion model, which is an infrared image and visible light image fusion algorithm without the prior condition of the fused image, avoiding the degradation of the fusion quality caused by the difficult-to-guarantee quality of the prior estimation of the fused image features, and making the training and use of the model more convenient.
[0182] Secondly, the structure of the diffusion model is expanded and improved to enable it to better learn the respective features of infrared images and visible light images, thereby better generating fused images. Finally, combined with the SAM large vision model, a method for fusing infrared images and visible light images guided by image segmentation tasks is provided. By leveraging the advantages of SAM as a large model, the small model for the image fusion task is trained, aiming to improve the effect with only a small increase in the parameters of the image fusion model.
[0183] The image segmentation result is also added as part of the loss function to the training of the model, so that the final fused image obtained by the trained network model can be more closely integrated with downstream tasks such as semantic segmentation and object detection and achieve better performance in them.
[0184] Based on the above image fusion model, an embodiment of the present application provides a schematic flowchart of an image fusion method. Referring to Figure 5 , the image fusion method specifically includes the following steps S501 - step S502:
[0185] Step S501, obtain an infrared image and a visible light image taken of the same scene.
[0186] Step S502, input the infrared image and the visible light image into the image fusion model to obtain a target fused image.
[0187] The image fusion model is obtained by the above image fusion model training method.
[0188] The target fused image obtained through the image fusion model can be more closely integrated with downstream tasks such as semantic segmentation and object detection.
[0189] To better implement the above image fusion model training method, an embodiment of the present invention provides an image fusion model training device. Referring to Figure 6 , Figure 6 is a structural block diagram of an image fusion model training device provided by an embodiment of the present application. As shown in Figure 6 , the device 60 includes:
[0190] A sample acquisition module 601, configured to acquire training samples, where the training samples include a first infrared image and a first visible light image taken of the same scene.
[0191] A model output module 602, configured to input the training samples into a preset network model to obtain a first fused image corresponding to the training samples.
[0192] A segmentation module 603 is configured to segment target objects in the first infrared image, the first visible light image, and the first fused image respectively, to obtain first segmentation information, second segmentation information, and third segmentation information.
[0193] A calculation module 604 is configured to calculate a loss based on the first segmentation information, the second segmentation information, and the third segmentation information, and determine a target loss value.
[0194] A training module 605 is configured to adjust parameters of a preset network model based on the target loss value corresponding to a training sample, to obtain an image fusion model, where the image fusion model is configured to perform image fusion on an infrared image and a visible light image to obtain a fused image.
[0195] In one implementation, the model output module 602 is specifically configured to: add first Gaussian noise to the first infrared image to obtain a second infrared image; add second Gaussian noise to the first visible light image to obtain a second visible light image; perform image fusion processing on the second infrared image and the second visible light image to obtain a second fused image; and obtain a first fused image based on the first infrared image, the first visible light image, and the second fused image.
[0196] In one implementation, the model output module 602 is specifically configured to: determine a target prediction noise based on the first infrared image, the second infrared image, the first visible light image, and the second visible light image, where the target prediction noise is the noise that needs to be removed when performing image fusion processing on the second infrared image and the second visible light image; and use the DDIM algorithm to obtain a second fused image based on the target prediction noise.
[0197] In one implementation, the model output module 602 is specifically configured to: predict a first prediction noise based on the first infrared image and the second infrared image, where the first prediction noise is the noise that needs to be removed from the second infrared image when performing image fusion processing on the second infrared image and the second visible light image; predict a second prediction noise based on the first visible light image and the second visible light image, where the second prediction noise is the noise that needs to be removed from the second visible light image when performing image fusion processing on the second infrared image and the second visible light image; and perform a weighted summation calculation on the first prediction noise and the second prediction noise to determine the target prediction noise.
[0198] In one implementation, the model output module 602 is specifically configured to: extract features of the first infrared image, the first visible light image, and the second fused image respectively to obtain corresponding first image features, second image features, and third image features; perform fusion processing on the first image features, the second image features, and the third image features to obtain target image features; generate a third fused image based on the target image features; and embed the first segmentation information and the second segmentation information into the third fused image to obtain a first fused image.
[0199] The first segmentation information includes the target values corresponding to each pixel in the first infrared image, the second segmentation information includes the target values corresponding to each pixel in the first visible light image, and the third segmentation information includes the target values corresponding to each pixel in the first fused image. The target value is used to characterize whether the corresponding pixel is the area where the target object is located.
[0200] In one implementation, the calculation module 604 is specifically configured to: calculate the loss between the target values corresponding to each pixel in the first infrared image and the target values corresponding to each pixel in the first fused image by using a preset first loss function to obtain a first loss value, and calculate the loss between the target values corresponding to each pixel in the first visible light image and the target values corresponding to each pixel in the first fused image by using the first loss function to obtain a second loss value; perform a summation calculation on the first loss value and the second loss value to determine the target loss value.
[0201] In one implementation, the calculation module 604 is specifically configured to: calculate the loss between the noise distribution function corresponding to the first predicted noise and the Gaussian distribution function corresponding to the first Gaussian noise by using a preset second loss function to obtain a third loss value; calculate the loss between the noise distribution function corresponding to the second predicted noise and the Gaussian distribution function corresponding to the second Gaussian noise by using the second loss function to obtain a fourth loss value. Perform a summation calculation on the first loss value, the second loss value, the third loss value, and the fourth loss value to determine the target loss value.
[0202] In one implementation, the calculation module 604 is specifically configured to: calculate the loss between the first fused image and the first infrared image by using a preset third loss function to obtain a fifth loss value; calculate the loss between the first fused image and the first visible light image by using the third loss function to obtain a sixth loss value; perform a summation calculation on the first loss value, the second loss value, the third loss value, the fourth loss value, the fifth loss value, and the sixth loss value to determine the target loss value.
[0203] The steps that the above device 60 can implement are similar to those of the above image fusion model training method, and will not be elaborated here one by one.
[0204] Based on the above device 60, an image fusion model is realized that can output a fused image that is more closely combined with downstream tasks such as semantic segmentation and object detection.
[0205] To better implement the above image fusion method, an embodiment of the present invention provides an image fusion device. Refer to Figure 7 , Figure 7 For the structural block diagram of an image fusion device provided by an embodiment of the present application, as Figure 7 shown, the device 70 includes:
[0206] An image acquisition module 701 is configured to acquire an infrared image and a visible light image captured of the same scene.
[0207] An image fusion module 702 is configured to input the infrared image and the visible light image into an image fusion model to obtain a target fusion image, where the image fusion model is obtained according to the above-mentioned image fusion model training method.
[0208] Based on the above device 70, the final fusion image obtained through the trained image fusion model can be more closely combined with downstream tasks such as semantic segmentation and object detection, and achieve better performance therein.
[0209] Figure 8 The schematic hardware structure diagram of an electronic device provided by an embodiment of the present application is shown.
[0210] The electronic device may include a processor 801 and a memory 802 storing computer program instructions.
[0211] Specifically, the above-mentioned processor 801 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0212] The memory 802 may include a mass storage for data or instructions. By way of example and not limitation, the memory 802 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disc, a magneto-optical disc, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. In a suitable case, the memory 802 may include a removable or non-removable (or fixed) medium. In a suitable case, the memory 802 may be internal or external to the integrated gateway disaster recovery device. In a specific embodiment, the memory 802 is a non-volatile solid state memory.
[0213] In some embodiments, the memory 802 may include a read only memory (ROM), a random access memory (RAM), a magnetic disk storage media device, an optical storage media device, a flash memory device, an electrical, optical, or other physical / tangible memory storage device. Thus, in general, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described in reference to the method according to one aspect of the present disclosure.
[0214] The processor 801 reads and executes the computer program instructions stored in the memory 802 to implement any one of the image fusion model training methods in the above embodiments.
[0215] In one example, the electronic device may further include a communication interface 803 and a bus 810. Among them, as Figure 8 shown, the processor 801, the memory 802, and the communication interface 803 are connected through the bus 810 to complete communication with each other.
[0216] The communication interface 803 is mainly used to implement communication between various modules, devices, units, and / or devices in the embodiments of the present application.
[0217] The bus 810 includes hardware, software, or both, and couples the components of the online data flow billing device to each other. By way of example and not limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses or a combination of two or more of these. In a suitable case, the bus 810 may include one or more buses. Although the embodiments of the present application describe and illustrate specific buses, the present application contemplates any suitable bus or interconnect.
[0218] The electronic device can execute the image fusion model training method in the embodiments of the present application, thereby implementing the image fusion model training method and device described in combination with Figure 1 and Figure 6 described.
[0219] In addition, in combination with the image fusion model training method in the above embodiments, the embodiments of the present application can be implemented by providing a computer storage medium. Computer program instructions are stored on the computer storage medium; when the computer program instructions are executed by a processor, any one of the image fusion model training methods in the above embodiments is implemented.
[0220] It should be clear that the present application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present application.
[0221] The functional blocks shown in the above-described structural block diagrams can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, and so on. When implemented in software, the elements of the present application are programs or code segments for performing the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted via a data signal carried in a carrier wave over a transmission medium or a communication link. A "machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical discs, hard disks, fiber optic media, radio frequency (RF) links, and so on. The code segment can be downloaded via a computer network such as the Internet, an intranet, and so on.
[0222] It should also be noted that the exemplary embodiments mentioned in the present application describe some methods or systems based on a series of steps or devices. However, the present application is not limited to the order of the above steps, that is, the steps can be executed in the order mentioned in the embodiments, can be different from the order in the embodiments, or several steps can be executed simultaneously.
[0223] Aspects of the present disclosure have been described above with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block in the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine such that the instructions executed by the processor of the computer or other programmable data processing device enable the implementation of the functions / actions specified in one or more blocks of the flowcharts and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and the combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by dedicated hardware for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0224] The above are only specific embodiments of the present application. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application.
Claims
1. A method for training an image fusion model, characterized in that: The method comprises: Acquire a training sample, where the training sample includes a first infrared image and a first visible light image obtained by photographing the same scene; Inputting the training sample into a preset network model to obtain a first fused image corresponding to the training sample; Segmenting the target object in the first infrared image, the first visible light image, and the first fused image respectively to obtain first segmentation information, second segmentation information, and third segmentation information; Perform loss calculation based on the first segmentation information, the second segmentation information, and the third segmentation information to determine a target loss value; The parameters of the preset network model are adjusted based on the target loss value corresponding to the training sample to obtain an image fusion model, and the image fusion model is used to fuse the infrared image and the visible light image to obtain a fused image.
2. The method according to claim 1, characterized in that The step of inputting the training sample into a preset network model to obtain a first fused image corresponding to the training sample includes: The training sample is input into a preset network model so that the preset network model performs the following processing: adding a first Gaussian noise to the first infrared image to obtain a second infrared image; adding a second Gaussian noise to the first visible light image to obtain a second visible light image; performing image fusion processing on the second infrared image and the second visible light image to obtain a second fused image; The first fused image is obtained based on the first infrared image, the first visible light image, and the second fused image.
3. The method according to claim 2, characterized in that The performing image fusion processing on the second infrared image and the second visible light image to obtain a second fused image includes: determining a target prediction noise according to the first infrared image, the second infrared image, the first visible light image, and the second visible light image, wherein the target prediction noise is noise that needs to be removed when performing image fusion processing on the second infrared image and the second visible light image; A DDIM algorithm is used to obtain a second fused image based on the target prediction noise.
4. The method according to claim 3, characterized in that The determining target prediction noise according to the first infrared image, the second infrared image, the first visible light image, and the second visible light image includes: Predicting a first predicted noise according to the first infrared image and the second infrared image, where the first predicted noise is noise that needs to be removed from the second infrared image when performing image fusion processing on the second infrared image and the second visible light image; Predicting a second predicted noise according to the first visible light image and the second visible light image, where the second predicted noise is noise that needs to be removed from the second visible light image when performing image fusion processing on the second infrared image and the second visible light image; A weighted sum calculation is performed on the first prediction noise and the second prediction noise to determine the target prediction noise.
5. The method according to claim 2, characterized in that: The obtaining the first fused image based on the first infrared image, the first visible light image and the second fused image includes: Extracting features from the first infrared image, the first visible light image, and the second fused image respectively to obtain corresponding first image features, second image features, and third image features; The first image feature, the second image feature and the third image feature are fused to obtain a target image feature; generating a third fused image based on the target image feature; The first segmentation information and the second segmentation information are embedded into the third fused image to obtain the first fused image.
6. The method according to claim 1, characterized in that The first segmentation information includes a target value corresponding to each pixel in the first infrared image, the second segmentation information includes a target value corresponding to each pixel in the first visible light image, and the third segmentation information includes a target value corresponding to each pixel in the first fused image, and the target value is used to indicate whether the corresponding pixel is an area where a target object is located; The performing loss calculation based on the first segmentation information, the second segmentation information, and the third segmentation information to determine a target loss value includes: Using a preset first loss function to perform loss calculation on a target value corresponding to each pixel in the first infrared image and a target value corresponding to each pixel in the first fused image, to obtain a first loss value; Using the first loss function to perform loss calculation on a target value corresponding to each pixel in the first visible light image and a target value corresponding to each pixel in the first fused image, to obtain a second loss value; The first loss value and the second loss value are summed to determine the target loss value.
7. The method according to claim 4, characterized in that The first segmentation information includes a target value corresponding to each pixel in the first infrared image, the second segmentation information includes a target value corresponding to each pixel in the first visible light image, and the third segmentation information includes a target value corresponding to each pixel in the first fused image, and the target value is used to indicate whether the corresponding pixel is an area where a target object is located; Before adjusting the parameters of the preset network model based on the target loss value corresponding to the training sample to obtain the image fusion model, the method further includes: Using a preset second loss function to perform loss calculation on a noise distribution function corresponding to the first predicted noise and a Gaussian distribution function corresponding to the first Gaussian noise, to obtain a third loss value; Using the second loss function to perform loss calculation on a noise distribution function corresponding to the second predicted noise and a Gaussian distribution function corresponding to the second Gaussian noise, to obtain a fourth loss value; The performing loss calculation based on the first segmentation information, the second segmentation information, and the third segmentation information to determine a target loss value includes: Using a preset first loss function to perform loss calculation on a target value corresponding to each pixel in the first infrared image and a target value corresponding to each pixel in the first fused image, to obtain a first loss value; Using the first loss function to perform loss calculation on a target value corresponding to each pixel in the first visible light image and a target value corresponding to each pixel in the first fused image, to obtain a second loss value; The first loss value, the second loss value, the third loss value, and the fourth loss value are summed to determine the target loss value.
8. The method according to claim 7, characterized in that Before calculating the sum of the first loss value, the second loss value, the third loss value, and the fourth loss value to determine the target loss value, the method further includes: Using a preset third loss function to perform loss calculation on the first fused image and the first infrared image to obtain a fifth loss value; Using the third loss function to perform loss calculation on the first fused image and the first visible light image to obtain a sixth loss value; The summing up the first loss value, the second loss value, the third loss value, and the fourth loss value to determine the target loss value includes: The first loss value, the second loss value, the third loss value, the fourth loss value, the fifth loss value and the sixth loss value are summed up to determine the target loss value.
9. An image fusion method, characterized in that: The method comprises: Acquire an infrared image and a visible light image obtained by photographing the same scene; The infrared image and the visible light image are input into an image fusion model to obtain a target fused image, and the image fusion model is obtained according to the image fusion model training method according to claims 1-8.
10. An electronic device, characterized in that: The device comprises: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, the image fusion model training method as described in any one of claims 1-8 is implemented.
Citation Information
Cited By
General camera near-infrared image generation method
CN121810828A
A general camera near-infrared image generation method
CN121810828B