An image fusion network with the function of removing scene degradation
By designing an image fusion network with desceived degradation function, and using a scene degradation discriminator and a fusion image generator to process the images, the problems of poor robustness of the image fusion method and lack of degradation information repair in the prior art are solved, and high-quality image fusion effect is achieved.
Patent Information
- Application Number
- CN202411570508.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-06
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-11-06
AI Technical Summary
The existing image fusion method is poor in dealing with factors such as light changes and weather, has high computational complexity, and lacks repair of degraded information in visible light images, resulting in poor quality of the fusion result.
An image fusion network with desceived degradation function was designed, and the scene degradation discriminator and fusion image generator were used to model the scene degradation information through the image encoder and text encoder of the text image big model CLIP, and the degradation information in the visible light image was repaired using BANet and DFNet.
The robustness of the image fusion method in harsh environments is improved, the effective repair of degraded information in visible light images is achieved, and the quality of the generated fusion image is significantly improved.
Smart Images

Figure CN119444596B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image fusion, and specifically relates to an image fusion network with a function of removing scene degradation. Background Art
[0002] The infrared and visible light image fusion technology is an advanced technology that combines infrared images and visible light images. Its purpose is to make full use of the complementary information of both to improve the visual effect and recognition ability of images, and it is widely used in military reconnaissance, security monitoring, remote sensing detection and other fields. Infrared images can capture the infrared radiation emitted by objects and reflect the thermal distribution of objects, and are suitable for target detection in night or low visibility environments. Visible light images, on the other hand, can provide rich color and texture information, which is beneficial to the recognition of targets and the understanding of scenes. The image fusion technology effectively combines the thermal information of infrared images and the visual information of visible light images through a certain algorithm to generate a fusion image that contains both infrared characteristics and visible light details.
[0003] Currently, common image fusion methods are mainly divided into two categories: Pixel-level fusion methods: This method directly operates on image pixels, such as weighted average, multi-scale decomposition, etc. Feature-based fusion methods: This method first extracts image features and then performs feature fusion, such as fusion methods based on deep learning.
[0004] Although the existing image fusion methods have achieved certain results, there are still the following problems:
[0005] Poor robustness: The existing fusion methods are sensitive to factors such as light changes and weather, which easily lead to a decline in the fusion effect.
[0006] High computational complexity: The fusion methods based on deep learning require a large amount of training data and have a high computational complexity.
[0007] Lack of repair of degraded information: The existing fusion methods do not repair the degraded information in visible light images, resulting in poor quality of the fusion results. Summary of the Invention
[0008] The purpose of the present invention is to provide an image fusion network with a function of removing scene degradation in order to solve the above-mentioned problems.
[0009] The technical solution adopted by the present invention is as follows: An image fusion network with a function of removing scene degradation, the image fusion network includes: The network consists of a scene degradation discriminator and a fusion image generator;
[0010] The scene degradation discriminator is composed of an image encoder and a text encoder of the large model CLIP for text images;
[0011] The image encoder is cascaded with a learnable small neural network for fine-tuning the image encoding vectors generated by the image encoder; the input of the text encoder is scene information, including different text Prompts such as Fog, Low Light, Overexposure, and normal. It is trained with a contrastive learning loss to complete the modeling of scene degradation information.
[0012] The fusion image generator consists of two independent feature extraction branches and a fusion image reconstruction branch. Among them, the visible light branch degrades the visible light features with the parameters generated by five consecutive Degradation Correction Unit (DCU) brightness adjustment networks BANet and a defogging network DFNet.
[0013] The visible light branch of the fusion image generator degrades the visible light features with the parameters generated by five consecutive Degradation Correction Unit (DCU) brightness adjustment networks BANet and a defogging network DFNet. The infrared branch consists of five consecutive convolutional, Batch Norm, and ReLU layers for feature extraction of infrared images. After the visible light image features and infrared image features after degradation correction are concatenated on the channel, they are input into a fusion image reconstruction network consisting of five consecutive convolutional, Batch Norm, and ReLU layers to obtain the final fusion image.
[0014] In a preferred embodiment, the training of the image fusion network consists of three stages. In the first stage of training, the small neural network cascaded with the image encoder is trained to fine-tune the image encoding of the input visible light to match the corresponding degraded Prompt. In the second stage of training, the fusion image generator is trained to generate a preliminary fusion result. In the third stage, the scene degradation repair networks BANet and DFNet based on prior knowledge are fine-tuned to further repair the degradation information in the visible light image.
[0015] In a preferred embodiment, the Degradation Correction Unit (DCU) brightness adjustment network uses a basic physical model to repair the degradation of visible light image features. Denote the visible light features input to the DCU as
[0016] , and first perform preliminary feature extraction on it through a convolutional, Batch Norm, and ReLU layer:
[0017]
[0018] Among them represents the visible light features after preliminary feature extraction. Considering that directly adjusting the brightness will destroy the fog information in the original image, thus increasing the difficulty of defogging.
[0019] In a preferred embodiment, the defogging network uses the fog degradation in the visible light features of the atmospheric scattering model for removal. The atmospheric scattering model represents the observed fog image as the sum of the direct component of the object's reflected light and the scattering component caused by the atmospheric light, formulated as:
[0020]
[0021] Among them represents the observed fog image, is the unaffected scene information, A is the atmospheric light intensity, and t(x) is the transmittance. Therefore, the process of recovering clean features from the fog-contaminated features can be expressed as:
[0022]
[0023] Among them represents the visible light features after defogging. A and t(x) are unknowns, which are generated by the DFNet learning in the present invention. Considering the instability of the degradation information in the input visible light image and avoiding the unnecessary negative impact of the defogging model on the already clean visible light features, the present invention adds gating to the defogging process to adaptively control the degree of defogging:
[0024]
[0025] Among them, w is the learnable weight and b is the learnable bias. is the scene encoding vector obtained by the scene degradation discriminator. represents the Sigmoid activation function. Through the gating method, the network will adaptively decide whether to perform the degradation removal process according to the degradation information of the current scene.
[0026] In a preferred embodiment, the input of the degradation correction network is the visible light image and the scene encoding vector obtained by the degradation scene discriminator. For the DFNet, the input visible light image is first processed by four consecutive convolutional and activation layers for feature extraction:
[0027]
[0028] Then, the scene encoding vector is embedded into the visible light features through the cross-attention mechanism:
[0029]
[0030] Among them, CAM() represents the cross-attention mechanism, and this process can be expressed as:
[0031]
[0032] Among them represents the transpose of. Finally, the visible light features after being embedded with the scene encoding vector are input into two different branches for generating the transmittance and the atmospheric light intensity:
[0033]
[0034] For BANet, four consecutive convolutional layers and activation layers are also adopted for feature processing. Then, the cross-attention mechanism is used to embed the scene encoding vector to obtain . Then, is decomposed along the channels to obtain the parameters of n brightness adjustment curves. For the best efficiency, n in the present invention is set to 8:
[0035] .
[0036] In a preferred embodiment, during the training of the first stage, the degenerate scene discriminator is trained with the help of the contrast learning theory. Define the input image as , the degenerate Prompt set as . The calculation of the similarity between the current image input and the Prompt can be formulated as:
[0037]
[0038] Among them represents the calculation of the cosine similarity, represents the image encoder, represents the text encoder, represents the current input Prompt. Then, the calculation of the loss can be expressed as:
[0039]
[0040] Among them represents the loss used in the first stage of training. y is the binarized label, which takes the value of 1 when the input image matches the text Prompt; otherwise, it takes the value of 0.
[0041] In a preferred embodiment, during the training of the second stage, the RIDCP[] algorithm and the CLAHE[] algorithm are first used to manually repair the degenerate information therein to obtain the repaired visible light image as the training label, denoted as . The loss of the second stage can be formulated as:
[0042]
[0043] wherein represents the first norm, and max() represents taking the maximum element-wise. represents edge extraction using the Sobel operator. represents the fused image obtained by the fused image generation network. represents the infrared image.
[0044] In a preferred embodiment, during the training of the third stage, a pre-trained scene degradation discrimination network is used to fine-tune the scene degradation repair network to force it to generate higher-quality scene correction parameters. During the training of the third stage, the parameters other than BANet and DFNet are frozen. Specifically, the loss of the third stage can be defined as:
[0045]
[0046] wherein represents the Prompt with the content "Norm". In this way, the distance between the generated fused image and the non-degraded Prompt is continuously reduced, and the distance between the generated fused image and the degraded Prompt is continuously increased, thereby supplementing and refining the insufficient degradation repair knowledge learned from the labels, and further improving the degradation removal ability of the fusion network.
[0047] In a preferred embodiment, the brightness adjustment curve of the Degradation Correction Unit (DCU) brightness adjustment network can be defined as:
[0048]
[0049] where x represents the pixel value size in the original image, y represents the pixel value size of the image after brightness correction, and n represents the order of the brightness adjustment curve. The larger the value, the more accurate the generated curve. Therefore, BANet is used to generate scene brightness adjustment parameters to further correct the defogged visible light features:
[0050] .
[0051] In summary, due to the adoption of the above technical solutions, the beneficial effects of the present invention are:
[0052] In the present invention, the network can repair the degradation information in the visible light image during the execution of the infrared-visible light image fusion task, improving the robustness of the fusion method in harsh environments. A scene discriminator based on the text-image large model is proposed, and the relationship between the scene degradation Prompts and the corresponding images is established through the contrastive learning theory. After training, the scene discriminator can automatically generate the degradation type of the input image. Based on prior knowledge, BANet and DFNet are designed to repair uncontrolled light degradation and smoke degradation. BANet and DFNet have a simple and interpretable structure, and can use the text-image large model as a supervision to adaptively generate repair parameters to accurately restore the degradation information. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 It is a schematic diagram of the overall network structure of the present invention;
[0054] Figure 2 It is a visualization effect diagram of the brightness adjustment curve in the present invention, where (a) represents the correction of the visible light image in a dark environment by histogram equalization, and (b) represents the reverse process. The abscissa of the curve represents the pixel value of the image before correction, and the ordinate is the pixel value of the image after correction. The blue curve represents the original mapping result, and the red curve represents the mapping result after smoothing;
[0055] Figure 3 It is a network structure diagram of DFNet and BANet in the present invention, where Conv represents a convolutional layer with a convolutional kernel of 3, ReLU and Sigmoid represent activation functions, AvgPool represents global average pooling, FC represents a fully connected layer, and CAM represents a cross-attention mechanism. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0056] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0057] Refer to Figures 1-3 ,
[0058] An image fusion network with the function of removing scene degradation consists of a scene degradation discriminator and a fused image generator. The scene degradation discriminator consists of the image encoder and the text encoder of the large text-image model CLIP, where the image encoder is cascaded with a learnable small neural network for fine-tuning the image encoding vector generated by the image encoder. The input of the text encoder is scene information, including several different text Prompts such as Fog, Low Light, Overexposure, and normal. It is trained by contrastive learning loss to complete the modeling of scene degradation information. The fused image generator consists of two independent feature extraction branches and a fused image reconstruction branch. Among them, the visible light branch uses the parameters generated by five consecutive Degradation Correction Unit (DCU) brightness adjustment networks BANet and defogging network DFNet to remove degradation from the visible light features. The infrared branch consists of five consecutive convolutional, Batch Norm, and ReLU layers for feature extraction of infrared images. After the visible light image features and infrared image features after degradation correction are concatenated on the channel, they are input into a fused image reconstruction network consisting of five consecutive convolutional, Batch Norm, and ReLU layers to obtain the final fused image.
[0059] The training of the proposed network consists of three stages. In the first stage of training, the small neural network cascaded with the image encoder is trained to fine-tune the input visible light image encoding to match the corresponding degradation Prompt. In the second stage of training, the fused image generator is trained to generate a preliminary fused result. In the third stage, the scene degradation repair networks BANet and DFNet based on prior knowledge are fine-tuned to further repair the degradation information in the visible light image.
[0060] The Degradation Correction Unit uses a basic physical model to repair the degradation of visible light image features. Denote the visible light feature input to the DCU as , and first perform preliminary feature extraction on it through a convolutional, Batch Norm, and ReLU layer:
[0061]
[0062] where Represents the visible light features after preliminary feature extraction. Considering that directly adjusting the brightness will destroy the fog information in the original image, thereby increasing the difficulty of defogging. Therefore, the present invention first uses the atmospheric scattering model [] to remove the fog degradation in the visible light features. The atmospheric scattering model represents the observed fog image as the sum of the direct component of the object's reflected light and the scattering component caused by the atmospheric light, and is formulated as:
[0063]
[0064] where represents the observed fog image, is the unaffected scene information, A is the atmospheric light intensity, and t(x) is the transmittance. Therefore, the process of recovering clean features from fog-contaminated features can be expressed as:
[0065]
[0066] where represents the visible light features after defogging. A and t(x) are unknowns, which are generated by the present invention using DFNet. Considering the instability of the degradation information in the input visible light image and avoiding the unnecessary negative impact of the defogging model on the already clean visible light features, the present invention adds gating to the defogging process to adaptively control the degree of defogging:
[0067]
[0068] where w is the learnable weight and b is the learnable bias, is the scene encoding vector obtained by the scene degradation discriminator, represents the Sigmoid activation function. Through the gating method, the network will adaptively decide whether to perform the degradation removal process according to the degradation information of the current scene.
[0069] After correcting the fog degradation in the visible light features, the present invention further corrects the scene brightness degradation. By observing the process of image brightness adjustment, we find that the relationship between the image after brightness adjustment and the original image can be mapped by a brightness correction, Figure 2 which is specifically presented in []. To ensure that the value range of the image does not change during the brightness adjustment process, the brightness correction curve must pass through the points (0,0) and (1,1). Since the change in brightness is slow during the brightness adjustment process and drastic brightness changes will bring potential noise information, the brightness correction curve must be continuously differentiable. Therefore, the brightness adjustment curve can be defined as:
[0070]
[0071] Where x represents the pixel value size in the original image, and y represents the pixel value size of the image after brightness correction. n represents the order of the brightness adjustment curve. The larger the value, the more accurate the resulting curve. Therefore, the present invention uses BANet to generate scene brightness adjustment parameters and further corrects the dehazed visible light features:
[0072]
[0073] Considering that the scene brightness of some images is at a good level, a gating method is still adopted to regulate the brightness correction process.
[0074] Generally speaking, in the brightness adjustment unit, the present invention does not directly operate on the degraded visible light features, but indirectly adjusts them using prior knowledge. The indirect adjustment method (1) avoids the destruction of the original visible light information, and (2) provides the possibility for fine-tuning of the degradation correction. In the training of the third stage, all parameters except BANet and DFNet will be frozen, so as to fine-tune BANet and DFNet to generate degradation correction parameters with continuously improved accuracy.
[0075] The overall structure of the scene degradation repair network includes the overall network structures of DFNet for generating fog correction parameters and BANet for generating brightness correction parameters. Considering that the degradation information of the scene promotes the generation of correction parameters, the input of the degradation correction network is the visible light image and the scene encoding vector obtained by the degradation scene discriminator . For DFNet, the input visible light image is first processed by four consecutive convolutional and activation layers for feature extraction:
[0076]
[0077] Then, the scene encoding vector is embedded into the visible light features through the cross-attention mechanism:
[0078]
[0079] Where CAM() represents the cross-attention mechanism, and this process can be expressed as:
[0080]
[0081] Where represents the transpose of. Finally, the visible light features embedded with the scene encoding vector are input into two different branches to generate the transmittance and atmospheric light intensity:
[0082]
[0083] For BANet, four consecutive convolutional layers are also adopted, and the activation layer is used for feature processing. Then, the cross-attention mechanism is used to embed the scene encoding vector to obtain . Then, is decomposed along the channels to obtain the parameters of n brightness adjustment curves. For the best efficiency, n is set to 8 in the present invention:
[0084]
[0085] The training of the network consists of three stages. In the first stage, the degraded scene discriminator is trained to match the scene encoding vector of the image and the text encoding vector of the corresponding degraded Prompt. In the second stage, the training of the fused image generation network is carried out, including the training of BANet and DFNet. In the third stage, using the text-image large model as supervision, the parameters of BANet and DFNet are fine-tuned to further improve the degradation removal ability of the network, specifically including:
[0086] Stage 1: Training of the degraded scene discriminator
[0087] In the present invention, the degraded scene discriminator is trained with the help of the contrastive learning theory. Define the input image as , and the set of degraded Prompts is . The calculation of the similarity between the current image input and the Prompt can be formulated as:
[0088]
[0089] where represents the calculation of the cosine similarity, represents the image encoder, represents the text encoder, represents the current input Prompt. Then, the calculation of the loss can be expressed as:
[0090]
[0091] where represents the loss used in the first stage of training, y is the binarized label, and when the input image and the text Prompt match, y takes the value of 1; otherwise, it takes the value of 0.
[0092] Stage 2: Training of the fused image generation network
[0093] In the second - stage training, the parameters of the scene degradation discriminator are frozen. For the input visible - light images containing degradation, the present invention first uses the RIDCP[] algorithm and the CLAHE[] algorithm to manually repair the degradation information therein, obtaining the repaired visible - light image as the training label, denoted as . The loss in the second stage can be formulated as:
[0094]
[0095] where represents the L1 - norm, max() represents taking the maximum element - by - element, represents edge extraction using the Sobel operator, represents the fused image obtained by the fused - image generation network, represents the infrared image.
[0096] Stage 3: Fine - tuning of the scene degradation repair network
[0097] In the third - stage training, the present invention uses the pre - trained scene degradation discriminative network to fine - tune the scene degradation repair network to force it to generate higher - quality scene correction parameters. In the third - stage training, the parameters other than BANet and DFNet are frozen. Specifically, the loss in the third stage can be defined as:
[0098]
[0099] where represents the Prompt with the content "Norm". In this way, the distance between the generated fused image and the non - degraded Prompt is continuously reduced, and the distance between the generated fused image and the degraded Prompt is continuously increased, thereby supplementing and refining the insufficient degradation repair knowledge learned from the labels, and further enhancing the degradation - removal ability of the fusion network.
[0100] In the present invention, the network can repair the degradation information in the visible - light image during the process of performing the infrared - visible - light image fusion task, improving the robustness of the fusion method in harsh environments. A scene discriminator based on the text - image large model is proposed, and the relationship between scene degradation Prompts and corresponding images is established through the contrast - learning theory. The trained scene discriminator can automatically generate the degradation type of the input image. Based on prior knowledge, BANet and DFNet are designed to repair uncontrolled illumination degradation and smoke degradation. BANet and DFNet have a simple and interpretable structure, and can use the text - image large model as supervision to adaptively generate repair parameters to accurately restore the degradation information.
[0101] It should be noted that in the present invention, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0102] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An image fusion network with scene degradation removal function, characterized by: The image fusion network includes: the network consists of a scene degradation discriminator and a fusion image generator; The scene degradation discriminator is composed of an image encoder and a text encoder of a large text image model CLIP; The image encoder is cascaded with a small learnable neural network to fine-tune the image encoding vector generated by the image encoder; the input of the text encoder is scene information, including fog, low light, overexposure and normal text prompts; it is trained by contrastive learning loss to complete the modeling of scene degradation information The fused image generator consists of two independent feature extraction branches and a fused image reconstruction branch; the visible light branch uses parameters generated by five consecutive Degradation Correction Unit (DCU) brightness adjustment networks BANet and defogging networks DFNet to degrade visible light features; The visible light branch of the fused image generator uses parameters generated by five consecutive Degradation Correction Unit (DCU) brightness adjustment networks BANet and defogging networks DFNet to degrade visible light features; the infrared branch consists of five consecutive convolution, Batch Norm and ReLU layers to extract features from infrared images; the visible light image features and infrared image features after degradation correction are spliced on the channels and input into the fused image reconstruction network composed of five consecutive convolution, Batch Norm and ReLU layers to obtain the final fused image; The input of the degradation correction network is the visible light image I vis and the scene encoding vector F obtained by the degraded scene discriminator scene ; For DFNet, the input visible light image I vis First, four consecutive convolution and activation layers are used for feature processing: Next, the scene encoding vector F scene Embed visible light features through cross-attention mechanism: Where CAM() represents the cross attention mechanism, the process can be expressed as: CAM(x,y)=softmax(y·x T )·x where x T represents the transpose of x; finally, it is embedded in the scene encoding vector F scene The resulting visible light features are fed into two different branches to generate transmittance and atmospheric light intensity: For BANet, four consecutive convolutional layers are also used, the activation layer performs feature processing, and then the scene encoding vector F is transformed into scene Embed it and get F BA ; Then, F BA It is decomposed along the channel to obtain the parameters of n brightness adjustment curves. In order to achieve the best efficiency, n is set to 8: [r1,r2,...,r n ]←F BA 。 2. The image fusion network with scene degradation removal function as claimed in claim 1, characterized in that: The training of the image fusion network consists of three stages. In the first stage of training, a small neural network cascaded with the image encoder is trained to fine-tune the input visible light image encoding to match the corresponding degraded prompt. In the second stage of training, the fused image generator is trained to generate preliminary fusion results. In the third stage, the scene degradation restoration networks BANet and DFNet based on prior knowledge are fine-tuned to further repair the degradation information in the visible light image.
3. The image fusion network with scene degradation removal function as claimed in claim 1, characterized in that: The Degradation Correction Unit (DCU) brightness adjustment network uses the basic physical model to repair the degradation of visible light image features; the visible light feature of the input DCU is denoted as F vis First, a convolution, Batch Norm and ReLU layer are used to perform preliminary feature extraction: F′ vis =CBR(F vis ) where F′ vis represents the visible light features after preliminary feature extraction; Considering that directly adjusting the brightness will destroy the fog information in the original image, it will increase the difficulty of defogging.
4. The image fusion network with scene degradation removal function as claimed in claim 3, characterized in that: The Degradation Correction Unit (DCU) brightness adjustment network removes fog degradation in the visible light characteristics of the atmospheric scattering model; the atmospheric scattering model represents the observed fog image as the sum of the direct component of the object reflected light and the scattered component caused by the atmospheric light, which is formulated as: I(x)=J(x)t(x)+A(1-t(x)) Where I(x) represents the observed fog image, J(x) is the unaffected scene information, A is the atmospheric light intensity, and t(x) is the transmittance; therefore, the process of recovering clean features from features contaminated by fog can be expressed as: in It represents the visible light features after defogging. A and t(x) are unknown quantities and are generated by DFNet learning. Considering the instability of the degraded information in the input visible light image, the defogging model is prevented from having unnecessary negative impact on the already clean visible light features. Gating is added to the defogging process to adaptively control the degree of defogging: Where w is the learnable weight, b is the learnable bias, and F scene is the scene encoding vector obtained by the scene degradation discriminator, and σ(·) represents the Sigmoid activation function. Through gating, the network will adaptively decide whether to perform the de-degradation process based on the degradation information of the current scene.
5. The image fusion network with scene degradation removal function as claimed in claim 2, characterized in that: In the first stage of training, the degraded scene discriminator is trained with the help of contrastive learning theory, and the input image is defined as I vis , the degenerate Prompt set is P i ∈{Fog, Low light, Overexposure, Norm}, the calculation of the similarity between the current image input and Prompt can be formulated as: Where cos(·,·) represents the calculation of cosine similarity, Φ I (·) represents the image encoder, Φ T (·) represents the text encoder, P represents the current input prompt; then, the loss calculation can be expressed as: Where L1 represents the loss used in the first stage of training, y is the binary label, and when the input image and the text Prompt match, y takes the value of 1; otherwise, it takes the value of 0.
6. The image fusion network with scene degradation removal function as claimed in claim 2, characterized in that: In the second stage of training, the degradation information is first manually repaired using the RIDCP[] algorithm and the CLAHE[] algorithm, and the repaired visible light image is obtained as the training label, denoted as The loss of the second stage can be formulated as: Among them, ||·||1 represents a norm, and max() represents taking the largest value element by element. Indicates edge extraction using the Sobel operator, I fuse represents the fused image obtained by the fusion image generation network, I inf Represents an infrared image.
7. The image fusion network with scene degradation removal function as claimed in claim 2, characterized in that: In the third stage of training, the pre-trained scene degradation identification network is used to fine-tune the scene degradation restoration network to force it to generate higher quality scene correction parameters; In the third stage of training, the parameters except BANet and DFNet are frozen; specifically, the loss of the third stage can be defined as: Where P Norm Indicates the Prompt with the content "Norm". In this way, the distance between the generated fused image and the non-degraded Prompt continues to decrease, and the distance between the generated fused image and the degraded Prompt continues to increase, thereby supplementing and refining the insufficient degradation restoration knowledge learned from the label, and further improving the de-degradation ability of the fusion network.
8. The image fusion network with scene degradation removal function as claimed in claim 1, characterized in that: The brightness adjustment curve of the Degradation Correction Unit (DCU) brightness adjustment network can be defined as: Among them, x represents the pixel value in the original image, y represents the pixel value of the image after brightness correction; n represents the order of the brightness adjustment curve, and the larger the value, the more accurate the curve generated. Therefore, BANet is used to generate scene brightness adjustment parameters and further correct the visible light features after defogging:
Citation Information
Patent Citations
End-to-end image defogging method based on multi-modal fusion
CN117575925A
Feature decomposition-based infrared image and visible light image fusion method
CN118134780A