Isovariant consistency image fusion method and system based on task driving

Through the task-driven isovariant consistent image fusion method, the problems of structure restoration, detail retention and semantic enhancement in infrared and visible image fusion are solved, and high-quality fusion images are generated, which improves the performance of downstream tasks.

CN120339092AActive Publication Date: 2025-07-18HEBEI UNIV OF TECH

Patent Information

Application Number
CN202510773419.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-07-18
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

The existing infrared and visible image fusion methods have shortcomings in structural restoration, detail retention, semantic enhancement and adaptability to downstream tasks, and it is difficult to effectively take into account texture details and global semantic information while maintaining geometric transformation robustness.

Method used

The task-driven isovariant consistency image fusion method is adopted, and the structural information, detailed information and semantic information in the fusion process are modeled layer by dividing and governance ideas, and multiple loss constraints are constructed, and the isovariant consistency constraints and semantic segmentation models are introduced to improve the structural restoration and detail retention capabilities of the fusion image, and enhance the support ability for downstream tasks.

Benefits of technology

The fusion image is achieved in the coordinated improvement of geometric transformation robustness, detail retention and semantic enhancement. The generated fusion image performs excellently in downstream tasks such as object detection and semantic segmentation, and has higher application potential.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339092A_ABST
    Figure CN120339092A_ABST
Patent Text Reader

Abstract

The invention discloses an isovariant consistency image fusion method and system based on task driving, and the method comprises the steps: carrying out the feature decomposition of an input infrared image and a visible image, obtaining the basic features and detail features, and carrying out the geometric transformation operation meeting the isovariant consistency constraint in the subsequent fusion processing, so as to guarantee the transformation stability of the features. And performing semantic constraint on the fusion model by utilizing supervision information of a semantic segmentation task of the semantic segmentation model. According to the method, the equal-variant consistency constraint and the semantic loss are collaboratively optimized through the loss function part, the robustness of geometric transformation can be kept, meanwhile, the reservation of local details and the enhancement of global semantic information are effectively considered, a high-quality fused image is generated, and the performance of a downstream advanced visual task is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of infrared and visible light fusion, and in particular to an equivariant consistency image fusion method and system based on task-driven, which is used for the fusion of infrared and visible light images, aiming to improve the image quality and semantic information by fusing images of two modalities to support advanced vision tasks. Background Art

[0002] The infrared and visible light image fusion technology is committed to integrating the information captured by different sensors to generate a single image containing more comprehensive scene information, which is crucial for improving the situation awareness ability in complex environments (such as low light, smoke). Through fusion, the ability of infrared images to detect targets under adverse conditions can be utilized, and at the same time, the advantages of visible light images providing rich texture and color details can be combined. Currently, using deep neural network models to process image fusion tasks has become a research hotspot.

[0003] For example, Chinese Patent with application number 202411048544.5 discloses a visible light imaging and infrared imaging fusion method, device, electronic device and storage medium. This method directly extracts semantic information from visible light imaging, obtains task prompt words through a large language model with preset task information, then predicts an infrared image with the task prompt words as input, and finally fuses the current frame visible light image and the predicted infrared image to obtain a target fusion image. However, this task-guided method only focuses on global semantic information and lacks fine-grained modeling of local structures and details, resulting in that while the fusion image improves the performance of downstream tasks, it often sacrifices some details and structural information of the original image. In addition, it is difficult to fully explore and utilize multi-level and multi-scale semantic features by relying on task prompt words, and the adaptability of the fusion result to complex scenes is limited.

[0004] Therefore, the existing infrared and visible light image fusion methods still have obvious deficiencies in aspects such as structural restoration, detail retention, semantic enhancement, and adaptability to downstream tasks. How to more effectively combine the transformation consistency decomposition and the task guidance mechanism, and achieve active perception and enhancement of multi-level information is a technical problem that urgently needs to be solved in this field. Summary of the Invention

[0005] In view of the above deficiencies, the object of the present invention is to provide a task-driven equivariant consistency image fusion method and system. By introducing the divide-and-conquer idea, the method hierarchically models the structural information (basic features), detail information (detail features), and semantic information in the fusion process, and constructs multiple loss constraints. The feature decomposition loss and structural similarity loss effectively improve the structural restoration ability and detail retention ability of the fused image, and the semantic loss improves the support ability for downstream tasks, thus overcoming the defect that the existing loss function design is single and it is difficult to take into account multi-level information expression, and solving the technical problem that the existing fusion technology is difficult to effectively take into account texture details and global semantic information while maintaining the robustness of geometric transformation.

[0006] To achieve the above object, the technical solution of the present invention is implemented as follows: In a first aspect, the present invention provides a task-driven equivariant consistency image fusion method, and the method includes the following steps: Obtain an infrared image and a visible light image; Construct a fusion model, where the fusion model includes an encoder, a decoder, and a geometric transformation operation; Perform feature decomposition on the infrared image and the visible light image in the encoder to obtain their respective basic features and detail features; Add the basic features of the infrared image and the basic features of the visible light image element-wise to obtain a unified basic feature; add the detail features of the infrared image and the detail features of the visible light image element-wise to obtain a unified detail feature; Apply the geometric transformation operation to the unified basic feature and the unified detail feature to obtain a transformed basic feature and a transformed detail feature respectively; At the same time, input the unified basic feature and the unified detail feature into the decoder to obtain a fused image; Apply the geometric transformation operation to the fused image to obtain a transformed fused image, and re-input the transformed fused image into the encoder and decoder to obtain a reconstructed fused image; Among them, the geometric transformation operation satisfies the equivariant consistency constraint; During the training process, the fused image is simultaneously processed by a semantic segmentation model to obtain a predicted segmentation map, and the semantic loss between the predicted segmentation map and the corresponding true segmentation label map is calculated; the semantic loss participates in the total loss calculation and backpropagation of the entire model, guiding the update of the parameters of the fusion model, and the total loss is obtained by weighting the fusion loss and the semantic loss; Use the trained fusion model for the fusion of infrared images and visible light images.

[0007] Furthermore, the fusion loss includes a structural similarity loss, a feature decomposition loss, an intensity loss, and a gradient loss.

[0008] Furthermore, the structural similarity loss is the weighted sum of the similarity between the transformed base features and the reconstructed base features, the similarity between the transformed detail features and the reconstructed detail features, and the similarity between the fused images before and after reconstruction; The feature decomposition loss is obtained from the correlation coefficients between the base features of the infrared image and the visible light image, the correlation coefficients between the detail features, the correlation coefficients between the transformed detail features and the reconstructed detail features, and the correlation coefficients between the transformed base features and the reconstructed base features; The intensity loss is used to calculate the difference in pixel intensity between the transformed fused image and the reconstructed fused image; The gradient loss is used to calculate the loss in gradient between the transformed fused image and the reconstructed fused image.

[0009] Furthermore, the semantic loss is the cross-entropy loss.

[0010] Furthermore, the encoder includes a shared feature encoder, a base feature encoder, and a detail feature encoder. After the infrared image and the visible light image are respectively input into the shared feature encoder and then processed by the base feature encoder and the detail feature encoder, their respective base features and detail features are obtained; The shared feature encoder is implemented using the Restormer module, the base feature encoder is implemented using the Transformer module, and the detail feature encoder is implemented using the convolutional neural network module.

[0011] Furthermore, the geometric transformation operations include translation operations, rotation operations, and flipping operations; it is set to perform a translation by a random distance, a rotation of 10 degrees, a left mirror flip, and a right mirror flip once each; The result of concatenating the transformed base features and the transformed detail features is equivalent to the result of concatenating the unified base features and the unified detail features and then applying the same transformation; and the transformed fused image and the reconstructed fused image are consistent.

[0012] In a second aspect, the present invention provides a task-driven equivariant consistency image fusion system, the system comprising: An image acquisition module: used to acquire paired infrared images and visible light images; A fusion model: used for the fusion of infrared images and visible light images; wherein the fusion model includes an encoder, an equivariant consistency constraint module, and a decoder; the encoder is used to perform feature decomposition on the input infrared image and visible light image to obtain their respective base features and detail features; the equivariant consistency constraint module is configured to impose an equivariant consistency constraint based on geometric transformation operations; The semantic guidance module includes a semantic segmentation model related to a preset downstream visual task, configured to generate a semantic constraint signal according to the supervision information of the downstream visual task; The decoder receives the constraints from the equivariant consistency constraint module and the semantic guidance module, and is used to fuse the basic features and detailed features input into the decoder, and decode and generate a fused image.

[0013] The present invention also protects a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method can be implemented.

[0014] Compared with the prior art, the beneficial effects of the present invention are: (1) The present invention first organically combines equivariant consistency constraints and downstream segmentation semantic guidance and introduces them into the fusion model, and uses the feature decomposition loss to constrain the detailed features and basic features after equivariance, as well as the detailed features and basic features after reconstruction, realizing the coordinated improvement of the fused image in terms of geometric transformation robustness, detail retention, and semantic enhancement. It overcomes the deficiency of the prior art that in practical applications, it often only focuses on global consistency and lacks effective fusion of local high-frequency details and semantic information, resulting in deficiencies in the detail performance and semantic expression of the fused image.

[0015] (2) The present invention introduces equivariant consistency constraints, which not only improves the robustness of the fused image to geometric transformations such as rotation, translation, and flipping, but also realizes the retention of local detail information. The fused image with equivariant consistency constraints can retain more detail information of the image.

[0016] (3) The present invention performs excellently in downstream tasks and has higher application potential. By introducing the semantic segmentation task and the corresponding semantic loss term, the fusion model can actively enhance the semantic feature expression beneficial to subsequent visual tasks. Compared with the existing methods that only rely on pixel-level or perceptual-level losses, the generated fused image shows better performance in downstream tasks such as object detection and semantic segmentation. The specific effects are as Figure 7 with Figure 8 shown.

[0017] (4) The present invention maximally retains the transformed basic features and detailed features, as well as their reconstructed basic features and detailed features through the feature decomposition loss, thereby realizing the balance of the overall performance of the model. In addition, by introducing a reasonable semantic loss term and regulating its weight in the total loss, the optimal compromise between the detail clarity and semantic richness of the fused image is achieved. Experiments show that when the semantic loss weight is set to 0.2, all evaluation indicators reach the optimal state. Brief Description of the Drawings

[0018] The accompanying drawings that form a part of the present invention are used to provide a further understanding of the present invention. The illustrative embodiments and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention.

[0019] Figure 1 It is a schematic diagram of the network structure of an embodiment of the task-driven equivariant consistency image fusion method of the present invention; Figure 2 It is a visualization comparison diagram of the basic features and detailed features in the embodiments of the present invention with and without equivariant consistency constraints; Figure 3 It is a comparison diagram of the experimental effects of the method of the present invention and the existing method on the M3FD daytime dataset; Figure 4 It is a comparison diagram of the experimental effects of the method of the present invention and the existing method on the M3FD night dataset; Figure 5 It is a comparison diagram of the experimental effects of the method of the present invention and the existing method on the TNO dataset; Figure 6 It is a comparison diagram of the experimental effects of the method of the present invention and the existing method on the DroneVehicle dataset; Figure 7 It is a performance comparison diagram of the method of Embodiment 1 of the present invention and the existing method when applied to the downstream target detection task; Figure 8 It is a performance comparison diagram of the method of Embodiment 1 of the present invention and the existing method when applied to the downstream semantic segmentation task. Detailed implementation manners

[0020] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.

[0021] The task-driven equivariant consistency image fusion method of the present invention uses the supervision information of semantic segmentation to guide the fusion process, so that the generated fused image contains richer and more accurate semantic information. The aim is to generate a high-quality fused image that not only retains local texture details but also is rich in global semantic information and is robust to geometric transformations, so as to improve the performance of downstream high-level vision tasks.

[0022] For the equivariant image fusion paradigm to solve the problems that the existing fusion methods are sensitive to geometric transformations (such as translation, rotation, and flipping) and are prone to losing spatial details, the present invention introduces the equivariant consistency principle. Its core idea is to ensure that the fusion process F has equivariance for the transformations in the transformation group, that is, it satisfies the property, where I represents the input image. This ensures that applying the transformation to the input image and then performing fusion is equivalent to first performing fusion and then applying the same transformation to the result.

[0023] Equivariant Consistency Constraint: The present invention constructs a structural similarity loss function to constrain the consistency between the transformed fused image, the transformed detailed features, the transformed basic features and the reconstructed fused image, the reconstructed detailed features, and the reconstructed basic features, thereby introducing a strong prior mechanism to enable the model to maintain good robustness in the face of geometric transformations such as rotation, translation, and flipping.

[0024] Apply a preset geometric transformation operation to the unified basic features and the unified detailed features to obtain the transformed basic features and detailed features. The result after splicing the transformed basic features and the transformed detailed features is equivalent to the result after splicing the unified basic features and the unified detailed features and then applying the same transformation. It is defined as follows:

[0025] Among them, represents the translation, rotation, and flipping operations applied to the image; represents the unified detailed features, represents the unified basic features, represents the fusion process.

[0026] Semantic Segmentation Guidance: Aiming at the deficiency of the existing image fusion model in semantic information extraction, which in turn affects its applicability in downstream visual tasks, the fusion model of the present invention uses a semantic segmentation model for guidance to guide the fusion process to focus more on the retention and enhancement of semantic features. Adopt the joint training method of the semantic segmentation model and the fusion model, and achieve effective coordination between the two through shared features and multi-task optimization to determine the optimal coordination strategy between the fusion model and the semantic segmentation model. Introduce a semantic loss term and adjust its weight to achieve a good balance between image clarity and semantic expression ability, thereby significantly improving the application performance of the fused image in downstream tasks such as object detection and semantic segmentation. After training is completed, the semantic segmentation model is only used as an auxiliary module during training to guide the fusion model to learn more accurate semantic features and is not retained during the inference stage, thus ensuring the inference efficiency of the model.

[0027] Both the encoder and decoder structures in the present invention can be implemented using conventional techniques in the art.

[0028] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will describe the technical solutions of the present invention in detail with reference to specific embodiments.

[0029] Embodiment 1: This embodiment is a task-driven equivariant consistency image fusion method, including the following steps: Obtain an infrared image and a visible light image; Construct a fusion model, which includes an encoder, a decoder, and a geometric transformation operation; for the overall framework of the fusion model, refer to Figure 1 , The encoder is used to receive paired infrared images and visible light images, and output their respective basic features and detailed features, which includes three main parts: Shared feature encoder: In this embodiment, the shared feature encoder is implemented using the Restormer module. Its main function is to extract and encode features from the input infrared and visible light images, and can effectively capture global and local features in the images.

[0030] Basic feature encoder: Use the Transformer module to further process the shallow features and extract low-frequency basic features. Denote the basic features extracted from the infrared image as , and denote the basic features extracted from the visible light image as . The Transformer module is good at capturing global structural information, which helps to ensure semantic consistency.

[0031] Detail feature encoder: Use the convolutional neural network module to process the deep features and extract high-frequency detailed features. Denote the detailed features extracted from the infrared image as , and denote the detailed features extracted from the visible light image as .

[0032] Add the basic features of the infrared image and the basic features of the visible light image element by element to obtain the unified basic feature . Add the detailed features of the infrared image and the detailed features of the visible light image element by element to obtain the unified detailed feature .

[0033] Apply the geometric transformation operation to the unified basic feature and the unified detailed feature respectively, that is, perform translation, rotation, and flipping once each to obtain the transformed basic feature and the transformed detailed feature .

[0034] At the same time, input the unified basic feature and the unified detailed feature into the decoder to obtain the fused image ; Apply the same geometric transformation operation to the fused image to obtain the transformed fused image .

[0035] Input the transformed fused image back into the encoder to obtain the reconstructed basic features and the reconstructed detail features . Then input the reconstructed basic features and the reconstructed detail features into the decoder to generate the reconstructed fused image .

[0036] Meanwhile, input the fused image into the semantic segmentation model to obtain the predicted segmentation map , which carries segmentation labels. In this embodiment, the semantic segmentation model uses the BiSeNet semantic segmentation network. The main feature of BiSeNet is that through the parallel spatial path and context path structure, it can take into account both the spatial details and global semantic information of the segmentation result, so as to achieve accurate segmentation of different semantic regions in the image. The specific formula for the semantic segmentation model to process can be expressed as:

[0037] Calculate the semantic loss between the predicted segmentation map and the corresponding ground truth segmentation label map . The ground truth segmentation label comes from the annotation in the training dataset.

[0038] Incorporate the semantic loss into the overall loss calculation and backpropagation of the entire model, so as to guide the update of the parameters of the fusion model, making the fused image generated by the fusion model not only have good visual effects, but also be beneficial to subsequent semantic segmentation tasks.

[0039] The total loss function of this embodiment is composed of the fusion loss and the semantic loss weighted, aiming to optimize both image fidelity and semantic consistency at the same time.

[0040] The fusion loss includes structural similarity loss, feature decomposition loss, intensity and gradient loss: Structural similarity loss : Based on the SSIM metric, it is used to quantify and ensure the retention degree of the information of the transformed and reconstructed images. It mainly calculates the similarity among the following three:

[0041] Among them, represents the visualization image of the transformed basic features, represents the visualization image of the reconstructed basic features, represents the visualization image of the transformed detail features Represents the visualized image of the detailed features after reconstruction, Represents the fused image after transformation, Represents the fused image after reconstruction; Represents the similarity between the basic features before and after transformation; Represents the similarity between the detailed features before and after transformation; Represents the similarity between the fused images before and after reconstruction; Is a weight parameter used to control the balance between the fused image loss and the decomposed image loss, which is set to 0.1 in this embodiment.

[0042] Feature decomposition loss : Used to ensure that the basic features and detailed features are effectively separated and maintain consistency. It calculates the correlation between the original decomposed features (the decomposed features include basic features and detailed features), as well as between the transformed and reconstructed decomposed features, to ensure that the basic features and detailed features can still be retained after undergoing equivariant consistency and semantic tasks, expressed as:

[0043] Among them, Is the correlation coefficient operator, Represents the detailed features of the infrared image, Represents the detailed features of the visible light image, Represents the basic features of the infrared image, Represents the basic features of the visible light image, Represents the basic features after transformation, Represents the detailed features after transformation, Represents the basic features after reconstruction, Represents the detailed features after reconstruction; Is a constant. To prevent the denominator from being 0, it is set here To be 0.01.

[0044] Intensity and gradient loss : Intensity loss Calculates the difference in pixel intensity between the fused image after transformation and the fused image after reconstruction to maintain the overall brightness and contrast information; Gradient loss Calculates the loss in gradient between the fused image after transformation and the fused image after reconstruction, which is used to retain edge and texture details; Then the intensity and gradient loss is expressed as:

[0045] Among them, λ is a parameter used to control the proportion of the gradient loss.

[0046] Semantic loss : The cross - entropy loss function is adopted to calculate the difference between the predicted segmentation map output by the semantic segmentation model and the corresponding ground - truth segmentation label map. The semantic loss is used to improve the semantic expression ability of the fused image and its adaptability to downstream tasks. The specific formula is:

[0047] where represents the cross - entropy function.

[0048] Therefore, the total loss can be expressed as: Total loss:

[0049] where is the fusion loss, is the semantic loss, is a hyper - parameter used to balance the importance between fusion fidelity and semantic information enhancement. In this embodiment is set to 0.2.

[0050] To verify the effectiveness and superiority of the method in the embodiments of the present invention, on the publicly available M3FD and RoadScene infrared and visible - light image datasets, using commonly used objective evaluation indexes in this field such as entropy (EN), standard deviation (SD), spatial frequency (SF), sum of differential correlation (SCD), visual information fidelity (VIF), and gradient - based fusion performance index (Qabf), the performance of the method of the present invention is compared with a variety of existing representative fusion methods (such as SwinFusion, SDNet, LRRNet, SeAFusion, EMMA, DDFM, CDDFuse). The objective evaluation results shown in Table 1 and Table 2 clearly indicate that the method in this embodiment shows significant advantages in multiple key indexes: it achieves the highest values among all the comparison methods in each index, which means that the fused images generated by the method of the present invention have richer information, higher contrast, clearer texture details, and can more effectively retain the complementary information of the source images. Generally speaking, these quantitative data strongly prove that compared with the prior art, the embodiments of the present invention have significant superiority and beneficial effects in generating high - quality, information - rich, detail - clear and well - structured infrared and visible - light fused images.

[0051]

[0052]

[0053] Figure 2 This is the visualization comparison diagram of the basic features and detail features in the embodiments of the present invention with and without the equivariant consistency constraint. Figure 2In the first row, there is the original infrared image and visible light image of the front part of a car under high exposure conditions. In the second row, there is a visualization of the unified basic features and unified detailed features directly processed by the encoder and stitched together through geometric transformation operations without setting the equivariant consistency constraint. In the third row, there is a visualization of the transformed basic features and transformed detailed features with the equivariant consistency constraint set. It can be seen from Figure 2 that the equivariant consistency constraint effectively improves the model's ability to retain structural and texture information. The feature decomposition loss function imposes a correlation constraint on the transformed detailed features and the reconstructed detailed features, further enhancing the model's performance in retaining detailed information.

[0054] It can be seen from Figures 3 to 6 that, compared with other models, the fused image of the present invention contains more texture details and structural information, the edge contours are clearer, and the target features are more prominent. In addition, the method of the present invention effectively retains the thermal radiation information in the infrared image and the rich color details in the visible light image, realizing the complementary fusion of information. The overall visual effect is more natural and the detail level is richer. Figure 7 It shows that the monitoring downstream task is the object detection task, and the model used for the object detection task is the YOLOv5 model. After using the fused model trained in the embodiment of the present invention and other existing fused models to perform image fusion on the infrared image and the visible light image, and then using the YOLOv5 model to detect vehicles and personnel, the detection accuracy of the method of the present invention is significantly improved compared with other models, indicating that the method of the present invention has an advantage in enhancing the expression of target features, can more accurately identify and locate targets, and improves the overall performance and reliability of detection. Figure 8 It shows that the monitoring downstream task is the semantic segmentation task, and the model used for the semantic segmentation task is the BiSNet semantic segmentation network. After using the fused model trained in the embodiment of the present invention and other existing fused models to perform image fusion on the infrared image and the visible light image, and then using the BiSNet semantic segmentation network to perform semantic segmentation, the object structure in the fused image obtained by the method of the present invention is clear and the boundaries are distinct, and the number of segmented targets is significantly more than that of other comparison methods, reflecting the advantage of the present method in retaining and expressing semantic information.

[0055] To verify the influence of the weight of the semantic loss term on the overall performance of the semantic segmentation model, in this embodiment, under the condition that other parameter settings remain unchanged, the weights of this loss term are respectively set to 0.1, 0.2, and 0.3, and a comparative experiment is conducted to explore the optimal weight configuration. The experimental results are shown in Table 3. When the weight is set to 0.2, optimal or relatively optimal performance is achieved in multiple evaluation indicators such as EN, SD, SF, MI, VIF, and Qabf, indicating that this setting achieves a good balance between image clarity and semantic expression ability, and verifies the effectiveness of the semantic loss in improving the comprehensive quality of the fused image.

[0056]

[0057] The present invention can effectively fuse infrared and visible light images, generate fused images with excellent visual quality, retain fine details, be robust to geometric transformations, and be rich in semantic information, thereby significantly improving the performance of downstream advanced vision tasks such as object detection and semantic segmentation.

[0058] Embodiment 2: This embodiment is a task-driven equivariant consistency image fusion system, including: An image acquisition module: used to acquire paired infrared images and visible light images; A fusion model: used for fusing infrared images and visible light images; the fusion model includes an encoder, an equivariant consistency constraint module, and a decoder; the encoder is used to perform feature decomposition on the input infrared image and visible light image to obtain their respective basic features and detail features; the equivariant consistency constraint module is configured to impose an equivariant consistency constraint based on geometric transformation operations; A semantic guidance module, including a semantic segmentation model related to a preset downstream vision task, configured to generate a semantic constraint signal according to the supervision information of the downstream vision task; The decoder receives the constraints from the equivariant consistency constraint module and the semantic guidance module, and is used to fuse the basic features and detail features input into the decoder, and decode to generate a fused image.

[0059] In this embodiment, the geometric transformation operations include a combination of translation operations, rotation operations, and flipping operations, which are set to perform a translation of a random distance, a rotation of 10 degrees, a left mirror flip, and a right mirror flip once each.

[0060] The present invention imposes an equivariant consistency constraint based on geometric transformation operations, which can ensure the robustness of the fusion process to geometric transformations; the semantic guidance module can impose semantic constraints on the fusion process, and perform fusion reconstruction based on the equivariant consistency constraint and the semantic constraint to obtain the final fused image.

[0061] Embodiment 3: In the system of this embodiment, the encoder includes a shared feature encoder, a basic feature encoder, and a detailed feature encoder; the basic feature encoder is used to extract basic information, and the detailed feature encoder is used to extract detailed information. The decoder includes a splicing operation and a Restormer module. It should be noted that the Restormer modules used in the encoder and decoder have similar structures, but their functions and parameter settings are different, which are determined through training to adapt to the encoding and decoding tasks respectively. The two inputs of the decoder are processed through the splicing operation and then enter the Restormer module for decoding.

[0062] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

[0063] Matters not described in the present invention apply to the prior art.

Claims

1. A task-driven equivariant consistency image fusion method, characterized in that, The method includes the following steps: Obtain an infrared image and a visible light image; Construct a fusion model, which includes an encoder, a decoder, and a geometric transformation operation; Perform feature decomposition on the infrared image and the visible light image in the encoder to obtain their respective basic features and detail features; Add the basic features of the infrared image and the basic features of the visible light image element by element to obtain a unified basic feature; add the detail features of the infrared image and the detail features of the visible light image element by element to obtain a unified detail feature; Apply the geometric transformation operation to the unified basic feature and the unified detail feature to obtain a transformed basic feature and a transformed detail feature respectively; Input the unified basic feature and the unified detail feature into the decoder simultaneously to obtain a fused image; Apply the geometric transformation operation to the fused image to obtain a transformed fused image, and re - input the transformed fused image into the encoder and decoder to obtain a reconstructed fused image; Among them, the geometric transformation operation satisfies the equivariant consistency constraint; During the training process, the fused image is simultaneously processed by a semantic segmentation model to obtain a predicted segmentation map, and the semantic loss between the predicted segmentation map and the corresponding true segmentation label map is calculated; the semantic loss participates in the calculation of the total loss of the entire model and backpropagation, guiding the update of the parameters of the fusion model, and the total loss is obtained by weighting the fusion loss and the semantic loss; Use the trained fusion model for the fusion of infrared images and visible light images.

2. The task-driven equivariant consistency image fusion method according to claim 1, wherein The fusion loss includes a structural similarity loss, a feature decomposition loss, an intensity loss, and a gradient loss.

3. The task-driven equivariant consistency image fusion method according to claim 2, wherein, The structural similarity loss is the weighted sum of the similarity between the transformed basic feature and the reconstructed basic feature, the similarity between the transformed detail feature and the reconstructed detail feature, and the similarity between the fused images before and after reconstruction; The feature decomposition loss is obtained through the correlation coefficients between the basic features of the infrared image and the visible light image, the correlation coefficients between the detail features, the correlation coefficients between the transformed detail feature and the reconstructed detail feature, and the correlation coefficients between the transformed basic feature and the reconstructed basic feature; The intensity loss is used to calculate the difference in pixel intensity between the transformed fused image and the reconstructed fused image; The gradient loss is used to calculate the loss in gradient between the transformed fused image and the reconstructed fused image.

4. The task-driven equivariant consistency image fusion method according to claim 1, wherein The semantic loss is a cross - entropy loss.

5. The task-driven equivariant consistency image fusion method according to claim 1, characterized in that The encoder includes a shared feature encoder, a basic feature encoder, and a detail feature encoder. The infrared image and the visible light image are respectively input into the shared feature encoder and then processed by the basic feature encoder and the detail feature encoder to obtain their respective basic features and detail features; The shared feature encoder is implemented using a Restormer module, the basic feature encoder is implemented using a Transformer module, and the detail feature encoder is implemented using a convolutional neural network module.

6. The task-driven equivariant consistency image fusion method according to claim 1, wherein The geometric transformation operation includes a translation operation, a rotation operation, and a flipping operation; it is set to perform a translation by a random distance, a rotation of 10 degrees, a left - hand mirror flip, and a right - hand mirror flip once each.

7. A task-driven equivariant consistency image fusion system, characterized in that, The system includes: Image acquisition module: used to acquire paired infrared images and visible light images; Fusion model: used for the fusion of infrared images and visible light images; the fusion model includes an encoder, an equivariant consistency constraint module, and a decoder; the encoder is used to perform feature decomposition on the input infrared image and visible light image to obtain their respective basic features and detailed features; the equivariant consistency constraint module is configured to impose an equivariant consistency constraint based on geometric transformation operations; Semantic guidance module, including a semantic segmentation model related to a preset downstream visual task, configured to generate a semantic constraint signal according to the supervision information of the downstream visual task; The decoder receives the constraints from the equivariant consistency constraint module and the semantic guidance module, and is used to fuse the basic features and detailed features input into the decoder and decode to generate a fused image.

Citation Information

Patent Citations

  • Visible light imaging and infrared imaging fusion method and device, electronic equipment and storage medium

    CN119027768A

  • High-performance infrared-visible fusion detection method

    CN108364272A

  • Infrared and visible light image fusion system based on distillation-fusion-semantic joint driving

    CN117274759A

  • Feature decomposition-based infrared image and visible light image fusion method

    CN118134780A

  • Infrared and visible light image perception enhancement fusion method and system based on deep learning

    CN119722492A

Cited By

  • Image editing method and device, equipment and storage medium

    CN121033227A

  • Semantic segmentation method and device, electronic equipment and storage medium

    CN121170299A

  • Semantic segmentation method and device, electronic equipment and storage medium

    CN121170299B

  • Semantic segmentation method and device, electronic equipment and storage medium

    CN121170300A