An infrared and visible light image fusion method based on fine-grained semantic space alignment and double-branch feature calibration

By constructing a fine-grained semantic description space and a bi-branch feature calibration method, the problem of insufficient cross-modal alignment accuracy in infrared and visible light image fusion is solved, achieving efficient fusion of infrared and visible light images and improving the semantic expression and target detection capabilities of the images.

CN122510107APending Publication Date: 2026-08-04WUHAN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
WUHAN INST OF TECH
Filing Date
2026-05-15
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing infrared and visible light image fusion methods have insufficient cross-modal alignment accuracy, making it difficult to accurately distinguish between salient targets and background areas. This results in infrared highlight areas masking visible light textures or visible light noise interfering with the representation of infrared targets during the fusion process. Furthermore, their semantic expression capabilities are limited, making it difficult to describe the detailed attributes of targets.

Method used

By constructing a fine-grained semantic description space, semantic retrieval and aggregation are performed to generate enhanced text representations. Spatial awareness calibration and distribution bias calibration are carried out in the dual-branch feature calibration. Spatial weights are constructed by combining the semantic response map, and infrared and visible light features are differentially modulated and fused. Multiple constraints are introduced to optimize the fusion process.

Benefits of technology

It significantly improves the semantic accuracy and consistency of image fusion, preserves key structural information, enhances target representation capabilities, and improves the stability and detail clarity of the fusion results, making it suitable for target detection and recognition in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122510107A_ABST
    Figure CN122510107A_ABST
Patent Text Reader

Abstract

The application discloses an infrared and visible light image fusion method based on fine-grained semantic space alignment and double-branch feature calibration, relates to the technical fields of computer vision, image processing and multi-modal information fusion, and comprises the following steps: generating diversified semantic descriptions for target categories and clustering, constructing a fine-grained semantic description space containing appearance features and thermal radiation features, and obtaining enhanced text representation through semantic retrieval and aggregation.The application constructs fine-grained semantic descriptions and enhances text expression, guides cross-modal alignment, improves detail preservation and consistency from spatial relationship and distribution relationship through double-branch feature calibration, generates spatial weights based on semantic response, realizes adaptive fusion of infrared and visible light features, and improves fusion quality and stability at the pixel, structure and semantic levels in combination with multi-constraint collaborative optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision, image processing and multimodal information fusion technology, specifically to an infrared and visible light image fusion method based on fine-grained semantic space alignment and bi-branch feature calibration. Background Technology

[0002] Infrared and visible light image fusion is an important research direction in the field of multi-source sensor information fusion, with wide application value in scenarios such as military reconnaissance, target detection, intelligent security, and autonomous driving. Infrared sensors image by sensing the thermal radiation of objects, effectively highlighting heat source targets, such as pedestrians and vehicles, in complex environments such as nighttime, low light, and fog. However, their imaging results often suffer from low resolution, lack of texture information, and insufficient background contrast. Visible light sensors, on the other hand, rely on ambient light reflection for imaging, providing rich texture, structure, and color information, but their performance degrades significantly in insufficient lighting or when there is occlusion. Therefore, effectively combining salient target information from infrared images with detailed information from visible light images through image fusion technology is of great significance for improving the overall expressiveness and environmental adaptability of images.

[0003] With the development of deep learning technology, image fusion methods based on convolutional neural networks (CNN), generative adversarial networks (GAN), and Transformer architectures have gradually become mainstream. However, existing technologies still have the following shortcomings: First, most fusion methods are still at the pixel-level or feature-level processing level, lacking high-level semantic understanding of image content, making it difficult to accurately distinguish salient targets from background areas. This results in infrared highlight areas easily obscuring visible light textures or visible light noise interfering with the expression of infrared targets during the fusion process. Second, while introducing pre-trained visual language models (such as CLIP) can enhance semantic information, their mechanism based on global image-text pairing training is prone to producing an overly uniform attention distribution, leading to the loss of fine-grained structure and edge information in dense prediction tasks. In addition, existing methods generally rely on fixed template text prompts (such as aphotoof[class]), which have limited semantic expressive power and are difficult to describe the detailed attributes of targets (such as material, shape, and thermal radiation features), resulting in insufficient cross-modal alignment accuracy, which in turn limits the performance improvement of fused images in scene understanding and target localization.

[0004] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0005] The purpose of this invention is to provide an infrared and visible light image fusion method based on fine-grained semantic space alignment and bi-branch feature calibration to solve the problems mentioned in the background art.

[0006] To achieve the above objectives, the present invention provides the following technical solution: an infrared and visible light image fusion method based on fine-grained semantic space alignment and bi-branch feature calibration, comprising the following steps: Diverse semantic descriptions are generated for the target category and clustered to construct a fine-grained semantic description space that includes appearance features and thermal radiation features. Enhanced text representations are obtained through semantic retrieval and aggregation. Infrared and visible light images are input into a visual pre-trained model to extract features. The infrared and visible light features are calibrated separately under an isomorphic parallel structure. The calibration process includes spatial perception calibration based on autocorrelation operation and distribution bias calibration based on learnable parameters to obtain calibrated infrared and visible light features. Based on the similarity relationship calculated by calibrated infrared features, calibrated visible light features and enhanced text representation, a fine-grained semantic response map representing the degree of semantic association at different spatial locations is generated. Spatial weights are constructed based on fine-grained semantic response maps. The calibrated infrared and visible light features are differentially modulated and fused to obtain fused features and reconstruct a fused image. A joint optimization constraint is constructed around the fused image, and collaborative optimization is carried out through guided preservation constraint, feature similarity constraint, perceptual constraint, cross-modal consistency constraint, and semantic feature consistency constraint.

[0007] Preferably, the process of constructing a fine-grained semantic description space includes: designing prompts for the target category to generate a multi-dimensional description containing appearance, shape, material, and thermal radiation attributes; and inputting each description into a text encoder to obtain a set of semantic vectors. ,in This represents a text encoding function. Indicates the first Description, Indicates the number of descriptions for each category. Indicates the number of categories; clustering the semantic vector set yields the implicit attribute vector set. ; Based on category-based cue vectors The cosine similarity between the implicit attribute vector and the relevant attribute set is selected, and the enhanced text representation is generated through weighted aggregation. .

[0008] Preferably, the spatial perception calibration process in the dual-branch feature calibration includes: extracting the first branch feature from the visual encoding process. The intermediate features of the layer are mapped to query representation, key representation, and value representation; autocorrelation relationships are calculated in their respective spaces to generate a spatially aware attention representation. ,in The weighting coefficients are used to weight the original features using spatially perceptual attention representation, resulting in spatially calibrated infrared and visible light features.

[0009] Preferably, the distribution bias calibration process in dual-branch feature calibration includes: extracting multi-layer intermediate features from the visual encoder. The fusion features are obtained through channel mapping and splicing. ; Calculate the spatial correlation matrix of the fused features ; Filter the correlation matrix to generate distribution offset terms The calibrated features are then superimposed onto the spatial attention representation.

[0010] Preferably, the process of generating the fine-grained semantic response map includes: calculating the cosine similarity between the calibrated infrared features and visible light features and the enhanced text representation, respectively, to obtain the initial response result; and then applying the minimum-maximum normalization function... The response results are normalized to obtain the infrared semantic response map. With visible light semantic response map ,in .

[0011] Preferably, the adaptive modulation fusion process includes: calculating spatial weights based on the semantic response graph. The infrared and visible light features are then fused element-wise using weighted methods to obtain the fused features. Then, the image is processed by a decoder to generate a fused image.

[0012] Preferably, the total loss function is Each loss term is used to constrain pixel representation, structural consistency, semantic representation, and cross-modal relationships, respectively.

[0013] Preferably, the guided retention loss is This is used to constrain the fused image to retain the main information of the corresponding modality in different semantic response regions.

[0014] Preferably, the feature similarity loss is This is used to constrain the fusion result to maintain structural consistency with the source image. (Dependent claims) Preferably, the perceived loss is: in This indicates that the pre-trained network is in the first... Features extracted from layers; The cross-modal consistency loss is: Used to constrain the consistency of responses from different modalities at the same semantic location; The semantic feature consistency constraint loss is: This is used to constrain the fused features to remain consistent with the text representation in the semantic space.

[0015] The technical effects and advantages provided by the present invention in the above technical solution are as follows: This invention constructs a fine-grained semantic description system by introducing a large language model, refining traditional coarse-grained category descriptions into multi-dimensional semantic expressions that include appearance, material characteristics, and thermal radiation attributes. Based on this, implicit attribute sets are formed through semantic clustering, and enhanced text representations are generated by combining attribute retrieval and feature aggregation, transforming semantic information from a single label into a multi-attribute combination expression. This approach significantly improves the information density of semantic priors, providing clearer reference points for cross-modal alignment, thereby enabling differentiated understanding of different target regions, enhancing semantic guidance during the fusion process, and improving the accuracy and consistency of image content expression.

[0016] This invention constructs a dual-branch feature calibration mechanism to jointly correct visual features at two levels: spatial relationship and distribution relationship. In the spatial dimension, internal correlation calculation strengthens the structural expression within a single modality, avoiding detail weakening caused by cross-modal information mixing. In the distribution dimension, multi-layer feature joint modeling achieves global relationship reconstruction, guiding features towards a more discriminative direction. This mechanism effectively alleviates feature homogenization while maintaining the original pre-training capabilities, ensuring more complete preservation of edge information and local details, while reducing expression bias between different modalities and improving the stability of the fusion result.

[0017] This invention proposes a spatial adaptive adjustment method based on semantic response, directly converting semantic alignment results into spatial weights and adjusting different modal features positionally. In regions with strong semantic association, infrared features are enhanced to highlight thermal target information; in regions with weak semantic association, visible light features are preferentially preserved to maintain texture details. By directly involving semantic information in weight allocation, the fusion process acquires region selection capabilities, enabling dynamic coordination of different modalities within a spatial range. This ensures both target prominence and overall visual quality, achieving a balance between structural representation and information integrity in the output image.

[0018] This invention employs an optimization approach that leverages multiple constraints to adjust the fusion result from various perspectives, including pixel representation, structural consistency, semantic association, and cross-modal coordination. Each constraint operates at a different level, enabling the fused image to gradually converge towards semantic consistency while preserving information from the source images. Through the synergistic effect of these multi-dimensional constraints, information loss and feature conflicts are effectively reduced, resulting in a comprehensive improvement in detail clarity, semantic accuracy, and overall stability of the fusion result. This provides a more reliable input foundation for subsequent recognition and detection tasks. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.

[0020] Figure 1 This is a flowchart illustrating the overall process of the infrared and visible light image fusion method based on fine-grained semantic space alignment and bi-branch feature calibration of this invention.

[0021] Figure 2 This is a diagram of the overall network structure of the present invention.

[0022] Figure 3 This is a diagram showing the overall structure of the text semantic clustering module of this invention.

[0023] Figure 4 This is a structural diagram of the visual feature calibration module of the present invention.

[0024] Figure 5 The present invention includes infrared images, visible light images, and fused images. Detailed Implementation

[0025] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that the description of this disclosure will be more complete and fully convey the concept of the exemplary embodiments to those skilled in the art.

[0026] This invention provides, for example Figures 1-5 This paper presents an infrared and visible light image fusion method based on fine-grained semantic space alignment and bi-branch feature calibration. To address the issue of insufficient semantic representation in multimodal image fusion tasks, a fine-grained semantic description system can be constructed and combined with text feature encoding and attribute aggregation mechanisms to systematically expand category semantics, thereby improving the ability to express key regions and detailed information during the fusion process.

[0027] In the semantic expansion stage, a multi-dimensional description set is first constructed around the various target objects involved in the fusion scenario. Traditional prompts, such as a photo of a person, only provide coarse category information, lacking descriptions of appearance details, thermal radiation distribution, and material properties, making it difficult to support cross-modal feature alignment. Therefore, by setting a unified description template, the language generation model is guided to output information descriptions containing multiple attributes. Specifically, it takes the form: a photo of [category], which has…, covering key dimensions such as appearance, color distribution, structural outline, thermal radiation characteristics, and material properties. For all categories included in the fusion dataset, a total of… is generated. Item description, in which This indicates the number of descriptions generated for each category. This represents the total number of categories. This set of descriptions constitutes a global semantic knowledge base, providing rich input for subsequent semantic encoding.

[0028] In the text feature encoding stage, all the above descriptions are input into the text encoder for vectorization representation. The encoding process is performed through a function. The implementation, in essence, maps natural language descriptions to a high-dimensional semantic space, thereby forming an embedded representation. The encoded result is represented as follows: in, Indicates the first A text description, This represents the corresponding semantic vector representation. This represents the set consisting of all the encoded descriptions. This set contains a large amount of redundant information in the semantic space, as well as rich fine-grained attribute features.

[0029] To improve the compactness and generalization ability of semantic representation, clustering is introduced into the feature space to process the set. Compression and reconstruction are performed. The K-means method is used to divide the semantic vector set into... There are 12 cluster centers, which constitute an implicit attribute space: in, Represents a collection of attributes. Indicates the first A vector of attributes Indicates the number of cluster centers. This represents the feature dimension output by the text encoder. Through this process, the original large-scale set of descriptions is compressed into a finite number of semantic prototypes, thereby achieving the structured organization of semantic expressions.

[0030] In the semantic enhancement stage, general suggestions are provided based on the input category. In attribute space Relevance retrieval is performed within the scope of the search. Cosine similarity is calculated to select the category most semantically relevant to the current category. Attribute vectors, construct attribute subsets Based on this, the original prompt is enhanced, and the enhanced semantic representation is calculated as follows: The meanings of each parameter are as follows: This represents the enhanced text semantic representation vector; The basic cue vector representing the input category; This represents the adjustment coefficient, used to control the fusion strength of attribute information, and its value is... ; This indicates the number of selected attributes, used to limit the scale of semantic information participating in the fusion; This represents the set of relevant attributes selected from the attribute space; Represents the first in the attribute set A vector of attributes; This indicates the similarity calculation result between the basic prompt and the attribute set; This represents the normalization function, used to rotate similarity into a weighted distribution so that each attribute has a different contribution during the fusion process.

[0031] This expression incorporates multiple relevant attributes into the base cue through a weighted summation, thereby constructing a more expressive semantic vector. The weights are dynamically determined by similarity, ensuring that attributes with higher relevance to the current category have a greater weight in the fusion process.

[0032] Overall, this method achieves a transformation from coarse-grained category information to fine-grained semantic representation through a layered processing flow of description expansion, semantic encoding, attribute extraction, and semantic enhancement. In image fusion tasks, this enhanced semantic representation can more accurately guide the selection and combination of different modal features, avoid infrared highlight areas from obscuring visible light textures, and retain key structural information, thereby improving the clarity and target representation ability of the fusion results.

[0033] To address the feature alignment problem in the fusion of infrared and visible light images, this paper proposes a method that uses unified visual coding and staged calibration to progressively correct the differences in spatial and distributional representations of cross-modal features, thereby obtaining feature representations that combine detail preservation with semantic consistency. The overall process consists of three consecutive parts: feature extraction, spatial awareness calibration, and distributional offset calibration, with close integration between each part.

[0034] In the feature extraction stage, infrared images With visible light images The input is processed using a unified visual coding system. Through a consistent coding path, the two types of images are made to correspond in feature dimension and spatial scale, providing a unified foundation for subsequent spatial relationship calculations. The coding process generates multiple layers of intermediate feature representations, which contain both shallow texture information and deep semantic information.

[0035] In the spatial perception calibration stage, the focus is on enhancing the representation of spatial relationships within a single modality to avoid interference with detailed information caused by direct cross-modal fusion. Intermediate features from several later layers are selected from the encoding results. And construct query representations respectively through linear mapping. , key representation Value representation The three are unified as In matrix form. Here Indicates the number of feature channels. This represents the spatial resolution. This process transforms the relationship between spatial locations into a problem of calculating the relationship between vectors.

[0036] Based on this, autocorrelation is calculated for each representation, and the expression is as follows: in, , representing the three types of spatial representations after mapping; This indicates a transpose operation, used to convert the spatial location dimension into a similarity calculation dimension; The inner product of any two spatial locations is used to characterize the degree of similarity; the denominator contains... Used for scaling to avoid excessively large values ​​affecting the stability of normalization; This indicates that the similarity results are normalized and transformed into a weighted distribution; the final output is... This represents the attention expression after relation modeling is completed within this space.

[0037] To integrate the advantages of different spatial representations, the three autocorrelation results are weighted and fused to obtain a unified spatial perception representation: in, This represents three spatial representations. This represents the corresponding weight coefficient, used to adjust the proportion of influence of different representations on the overall result, ensuring that the sum of the weights is one. This fusion process simultaneously considers local details, contextual relationships, and numerical distribution characteristics, making the spatial representation more complete.

[0038] After constructing the spatial relationships, the feature distribution is further refined to address the distribution discrepancy between the pre-trained features and the fusion task. First, a multi-layered intermediate feature set is extracted from the encoding process. The features of each layer are then scaled to ensure consistent spatial dimensions. This is followed by a mapping function. The channel dimensions of features from different layers are adjusted to a unified space, and multi-layer information is integrated by concatenation. Then, convolutional operations are used to fuse the information to form a comprehensive feature representation. in, This indicates the channel mapping process; This indicates splicing at the channel level; This indicates that the splicing result is reconstructed using point convolution; Indicates the channel dimension after fusion; This represents the multi-layered feature representation after fusion.

[0039] Based on this, the spatial correlation between features is calculated and centered to obtain the distribution offset representation: in, The cosine similarity matrix represents the number of spatial locations, and its elements are defined as follows: here, and They represent the first The and the first The eigenvectors corresponding to each spatial location; the numerator represents the inner product result; the denominator represents the vector norm product, used for normalization; the entire expression is used to characterize the directional similarity between two locations.

[0040] This represents the average cosine similarity across all pairs of spatial locations, used as a global reference. Used to amplify the difference; Used to control the degree of centralization; This represents the offset result after removing the average relationship, used to reflect the relative differences between positions.

[0041] To highlight the effectively relevant regions, the correlation matrix is ​​filtered: in, This represents the corresponding element in the matrix; it is retained when it is non-negative to indicate a positive correlation; when it is negative, it is assigned a value of negative infinity to suppress its influence in subsequent normalization. Finally, the distribution relationship and spatial relationship are superimposed to form the calibrated attention expression: in, This indicates that the filtered relevance matrix is ​​normalized to obtain the spatial weight distribution; This represents spatial perception. This represents the final attention expression result after integrating spatial and distributional relationships.

[0042] Through the above continuous processing, we obtain the feature representation after spatial relationship enhancement and distribution relationship correction. and This result preserves detailed structure at the spatial level and enhances discriminative power at the distribution level, thus providing a stable and consistent feature base for subsequent fusion.

[0043] In multimodal image fusion, relying solely on spatial alignment between visual features is insufficient to fully express the semantic differences of the target, especially when there are significant differences in imaging mechanisms between infrared and visible light images. Simple spatial calibration is inadequate for accurately locating semantic targets. Therefore, it is necessary to introduce textual semantics as guiding information to unify and align visual features with semantic expressions, thereby constructing a semantically aware response representation to guide the subsequent fusion process.

[0044] In the semantic alignment stage, the visual features that have undergone spatial and distribution calibration are first... 、 With enhanced text representation Projected onto a unified feature space. Through a learnable linear mapping method, the three types of features are made consistent in dimension, uniformly mapped to a dimension of... In the semantic space. The core purpose of this process is to eliminate the differences in feature representation scale between different modalities, so that visual information and text semantics can be directly compared.

[0045] After dimensional unification, the similarity between visual features and text semantics is calculated point-by-point along the spatial dimension. The similarity calculation for the relationship between infrared features and text semantics is as follows: The similarity between visible light features and text semantics is calculated as follows: The parameters mentioned above have the following meanings: Indicates the spatial location of infrared features The degree of matching between the location and the semantics of the text; This indicates the degree of matching between visible light features and text semantics at the same location; : Indicates the spatial location of infrared features The corresponding feature vector at that location; Indicates the spatial location of visible light features The corresponding eigenvector; symbol The first part represents the vector transpose operation, used to calculate the inner product; the denominator represents the norm product of the corresponding vectors, used for normalization to limit the result to a specific value. Within range; spatial index Represents the position coordinates in the feature map, where , , which correspond to the height and width directions respectively.

[0046] Through the above calculations, the semantic response distributions of the infrared and visible light modes in the spatial dimension can be obtained respectively. This response distribution reflects the degree of correlation between each pixel location and the target semantics, thus enabling explicit guidance of the target region in space.

[0047] However, due to differences in feature distribution among different image contents and modalities, directly using the original similarity matrix leads to inconsistent numerical ranges, thus affecting the subsequent fusion process. Therefore, it is necessary to normalize the similarity results to unify their distribution range and enhance the contrast between different regions. The normalization calculation is as follows: The parameters are explained below: This represents the minimum value in the infrared similarity matrix; This represents the maximum value in the infrared similarity matrix; This represents the minimum value in the visible light similarity matrix; This represents the maximum value in the visible light similarity matrix; : This is a minimal constant term used to prevent the denominator from being zero, thus ensuring numerical stability; the normalization result maps matrix elements to The interval allows for comparability of response strengths between different locations.

[0048] After normalization and These represent the semantic response maps of infrared and visible light images in the spatial dimension. A higher value indicates a stronger match between the location and the current semantic description; a lower value indicates a weaker correlation. This method can highlight the target region spatially while suppressing background interference.

[0049] Overall, this process establishes a direct link between visual and semantic information through continuous processing of feature alignment, similarity calculation, and numerical normalization. This allows the fusion process to no longer rely solely on pixels or feature intensity, but to make spatial selections based on semantic guidance. This enables the more accurate preservation of key target information in complex scenes and enhances the expressive power of the fusion results.

[0050] After constructing the semantic response, the visual features have obtained a response distribution consistent with the semantic description in the spatial dimension. However, at this point, the infrared and visible light features are still in an independent expression state. In order to achieve adaptive collaboration between the two types of features in different regions, the semantic response needs to be further transformed into spatial weights, and the features of different modalities should be differentially adjusted accordingly to form a fused feature expression.

[0051] During the weight generation stage, the semantic response results are transformed into a normalized spatial weight distribution. This is specifically expressed as follows: The meanings of the parameters in the above formula are as follows: Indicates spatial location The weighting ratio of infrared features; This indicates the weighting proportion of visible light features at the same location; This represents the similarity response value between infrared features and semantics; This represents the similarity response value between visible light features and semantics; The exponential function is used to non-linearly amplify the similarity, so that larger response values ​​occupy a higher proportion in the normalization process; the denominator represents the joint normalization of the two types of responses, so that the sum of the two weights is 1, thereby forming a complementary relationship in the same spatial location.

[0052] Through the above calculations, each location in the space corresponds to a set of weight values, and satisfies the following conditions: and When the infrared semantic response is strong at a certain location, the exponential amplification effect will cause... When the response is close to 1, the infrared signature is enhanced; when the visible light response is higher, more visible light information is retained at the corresponding location.

[0053] In the feature fusion stage, the two types of features are weighted point-by-point using the aforementioned spatial weights to form a unified fused representation: The parameters have the following meanings: This represents the feature result after fusion; Indicates the infrared characteristics after prior calibration; Indicates the visible light characteristics after the same processing; symbol The first operation represents element-wise multiplication, which means multiplying the weight value with the corresponding feature at each spatial location; the second operation represents superimposing the weighted results of the two modalities to obtain the final fused feature.

[0054] This fusion method achieves differentiated adjustment in the spatial dimension, enabling different regions to automatically select dominant features based on semantic response. For example, in high-temperature target regions, the infrared response is higher, resulting in increased weight and highlighting of the target region; in textured regions, the visible light response is stronger, thus preserving more detailed information. The entire process is completed at the pixel level, enabling fine-grained spatial control.

[0055] After obtaining the fused features, they are input into the decoding stage for image reconstruction. The decoding process maps the fused features back to the image space and restores the spatial resolution. The decoding structure consists of multiple levels of upsampling units, each responsible for progressively increasing the spatial size while smoothing and reconstructing the features. Each stage expands the feature map size through upsampling operations and reorganizes the channel information using convolution operations, thereby maintaining feature continuity while restoring resolution.

[0056] In the final output stage, the channel dimension is adjusted to 3 through convolution mapping to obtain the fused image representation: in, This represents the output fused image; 3 Indicates the number of image channels; This indicates the spatial resolution of the image, which is consistent with the input image.

[0057] The overall process involves three consecutive stages: semantically guided weight generation, spatial weighted fusion, and decoding reconstruction. This effectively integrates semantic information into visual features, achieving adaptive adjustment to different modalities. During the fusion process, semantic responses no longer serve merely as auxiliary information but directly participate in weight allocation, thus having a real impact on feature selection within the spatial range. This approach can simultaneously preserve infrared target information and visible light detail information in complex scenes, achieving a balance between visual representation and target expression in the final output.

[0058] During the training of a multimodal fusion network, multiple constraints are needed to guide the fusion result to achieve a balance in semantic consistency, structural preservation, and cross-modal alignment. The training strategy keeps the parameters of the pre-trained visual and text encoding parts fixed, only adjusting the weight parameters during the dual-branch calibration process. Channel mapping function Fusion characteristics The corresponding convolution parameters, adaptive fusion process, and decoding part are updated so that the features are gradually adapted to the fusion task without destroying the original semantic expressive power.

[0059] The overall loss consists of five parts, each constraining the fusion result from a different perspective.

[0060] The first part is the guided retention loss, which is expressed as follows: in, Indicates guiding constraints; This indicates the fused output image; Indicates infrared image input; Indicates a visible light image input; and This represents the semantic response graph, corresponding to the weight distribution of spatial locations; This represents the first norm, used to calculate the sum of absolute errors.

[0061] This method compares the fusion result with the weighted original image to make the fusion result closer to the content of the corresponding source image in areas with strong semantic response.

[0062] The second part is the feature similarity loss, which is expressed as follows: in, This indicates that the structure maintains constraints; This represents a structural similarity metric function, used to measure the similarity between two images in terms of brightness, contrast, and structure. Indicates the structural similarity between the fused image and the infrared image; Represents the structural similarity between the fused image and the visible light image; coefficient This indicates that the similarity between the two parts is averaged. This term is used to constrain the fusion result to retain the structural information in the original image.

[0063] The third part is the perceived loss, which is expressed as follows: in, This indicates high-level semantic consistency constraints; This represents the selected set of feature layers; Indicates the first Feature mappings extracted at each level; 、 、 These represent the number of channels, height, and width of the feature in this layer, respectively. This indicates an element-wise maximum value operation, used to select the stronger response between two modes; This represents the square of the L2 norm, used to calculate the mean square error.

[0064] This approach constrains high-level features to make the fusion result semantically closer to a more information-rich modal expression.

[0065] The fourth part is the cross-modal consistency loss, which is expressed as follows: in, This indicates a response consistency constraint; 、 Indicates the image space size; 、 Semantic response diagrams for infrared and visible light, respectively; This indicates the calculation of the squared error. This term is used to constrain the two modes to have a consistent response distribution at the same nominal location, thereby reducing spatial offset.

[0066] The fifth part is the semantic feature consistency constraint, which is expressed as follows: in, Indicates semantic alignment constraints; 、 Indicates the calibrated visual features; 、 Represents the corresponding response graph; symbols This indicates element-wise multiplication, applying the response map to the feature; The cosine similarity function measures the consistency of the directions of two vectors. This expression ensures consistency between visual representation and semantic description by comparing the relationship between weighted features and textual semantics. Finally, the total loss function is expressed as: in, Indicates the overall optimization goal; This represents the weight coefficient corresponding to each loss term, used to adjust the degree of influence of different constraints during the training process.

[0067] Through the combined effect of the above multiple constraints, the fusion results can be comprehensively adjusted at the pixel layer, structural layer, semantic layer, and cross-modal level, so that the output image can achieve a coordinated state in terms of detail representation, target expression, and overall consistency.

[0068] This invention constructs a fine-grained semantic description system by introducing a large language model, refining traditional coarse-grained category descriptions into multi-dimensional semantic expressions that include appearance, material characteristics, and thermal radiation attributes. Based on this, implicit attribute sets are formed through semantic clustering, and enhanced text representations are generated by combining attribute retrieval and feature aggregation, transforming semantic information from a single label into a multi-attribute combination expression. This approach significantly improves the information density of semantic priors, providing clearer reference points for cross-modal alignment, thereby enabling differentiated understanding of different target regions, enhancing semantic guidance during the fusion process, and improving the accuracy and consistency of image content expression.

[0069] This invention constructs a dual-branch feature calibration mechanism to jointly correct visual features at two levels: spatial relationship and distribution relationship. In the spatial dimension, internal correlation calculation strengthens the structural expression within a single modality, avoiding detail weakening caused by cross-modal information mixing. In the distribution dimension, multi-layer feature joint modeling achieves global relationship reconstruction, guiding features towards a more discriminative direction. This mechanism effectively alleviates feature homogenization while maintaining the original pre-training capabilities, ensuring more complete preservation of edge information and local details, while reducing expression bias between different modalities and improving the stability of the fusion result.

[0070] This invention proposes a spatial adaptive adjustment method based on semantic response, directly converting semantic alignment results into spatial weights and adjusting different modal features positionally. In regions with strong semantic association, infrared features are enhanced to highlight thermal target information; in regions with weak semantic association, visible light features are preferentially preserved to maintain texture details. By directly involving semantic information in weight allocation, the fusion process acquires region selection capabilities, enabling dynamic coordination of different modalities within a spatial range. This ensures both target prominence and overall visual quality, achieving a balance between structural representation and information integrity in the output image.

[0071] This invention employs an optimization approach that leverages multiple constraints to adjust the fusion result from various perspectives, including pixel representation, structural consistency, semantic association, and cross-modal coordination. Each constraint operates at a different level, enabling the fused image to gradually converge towards semantic consistency while preserving information from the source images. Through the synergistic effect of these multi-dimensional constraints, information loss and feature conflicts are effectively reduced, resulting in a comprehensive improvement in detail clarity, semantic accuracy, and overall stability of the fusion result. This provides a more reliable input foundation for subsequent recognition and detection tasks.

[0072] The foregoing has only described certain exemplary embodiments of the present invention by way of illustration. Undoubtedly, those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the foregoing drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.

Claims

1. A method for fusing infrared and visible light images based on fine-grained semantic space alignment and bi-branch feature calibration, characterized in that, Includes the following steps: Diverse semantic descriptions are generated for the target category and clustered to construct a fine-grained semantic description space that includes appearance features and thermal radiation features. Enhanced text representations are obtained through semantic retrieval and aggregation. Infrared and visible light images are input into a visual pre-trained model to extract features. The infrared and visible light features are calibrated separately under an isomorphic parallel structure. The calibration process includes spatial perception calibration based on autocorrelation operation and distribution bias calibration based on learnable parameters to obtain calibrated infrared and visible light features. Based on the similarity relationship calculated by calibrated infrared features, calibrated visible light features and enhanced text representation, a fine-grained semantic response map representing the degree of semantic association at different spatial locations is generated. Spatial weights are constructed based on fine-grained semantic response maps. The calibrated infrared and visible light features are differentially modulated and fused to obtain fused features and reconstruct a fused image. A joint optimization constraint is constructed around the fused image, and collaborative optimization is carried out through guided preservation constraint, feature similarity constraint, perceptual constraint, cross-modal consistency constraint, and semantic feature consistency constraint.

2. The infrared and visible light image fusion method based on fine-grained semantic space alignment and bi-branch feature calibration according to claim 1, characterized in that, The process of constructing a fine-grained semantic description space includes: designing prompts for the target category to generate multi-dimensional descriptions containing appearance, shape, material, and thermal radiation attributes; and inputting each description into a text encoder to obtain a set of semantic vectors. ,in This represents a text encoding function. Indicates the first Description, Indicates the number of descriptions for each category. Indicates the number of categories; clustering the semantic vector set yields the implicit attribute vector set. ; Based on category-based cue vectors The cosine similarity between the implicit attribute vector and the relevant attribute set is selected, and the enhanced text representation is generated through weighted aggregation. .

3. The infrared and visible light image fusion method based on fine-grained semantic space alignment and bi-branch feature calibration according to claim 1, characterized in that, The spatial perception calibration process in bi-branch feature calibration includes: extracting the first branch feature from the visual encoding process. The intermediate features of the layer are mapped to query representation, key representation, and value representation; autocorrelation relationships are calculated in their respective spaces to generate a spatially aware attention representation. ,in The weighting coefficients are used to weight the original features using spatially perceptual attention representation, resulting in spatially calibrated infrared and visible light features.

4. The infrared and visible light image fusion method based on fine-grained semantic space alignment and bi-branch feature calibration according to claim 1, characterized in that, The distribution bias calibration process in bi-branch feature calibration includes: extracting multi-layer intermediate features from the visual encoder. The fusion features are obtained through channel mapping and splicing. ; Calculate the spatial correlation matrix of the fused features ; Filter the correlation matrix to generate distribution offset terms The calibrated features are then superimposed onto the spatial attention representation.

5. The infrared and visible light image fusion method based on fine-grained semantic space alignment and bi-branch feature calibration according to claim 1, characterized in that, The process of generating fine-grained semantic response maps includes: calculating the cosine similarity between the calibrated infrared and visible light features and the enhanced text representation, respectively, to obtain the initial response results; and then applying the min-max normalization function. The response results are normalized to obtain the infrared semantic response map. With visible light semantic response map ,in .

6. The infrared and visible light image fusion method based on fine-grained semantic space alignment and bi-branch feature calibration according to claim 1, characterized in that, The adaptive modulation fusion process includes: calculating spatial weights based on the semantic response graph. The infrared and visible light features are then fused element-wise using weighted methods to obtain the fused features. Then, the image is processed by a decoder to generate a fused image.

7. The infrared and visible light image fusion method based on fine-grained semantic space alignment and bi-branch feature calibration according to claim 1, characterized in that, The total loss function is Each loss term is used to constrain pixel representation, structural consistency, semantic representation, and cross-modal relationships, respectively.

8. The infrared and visible light image fusion method based on fine-grained semantic space alignment and bi-branch feature calibration according to claim 7, characterized in that, Guided retention loss is This is used to constrain the fused image to retain the main information of the corresponding modality in different semantic response regions.

9. The infrared and visible light image fusion method based on fine-grained semantic space alignment and bi-branch feature calibration according to claim 7, characterized in that, Feature similarity loss is This is used to constrain the fusion result to maintain structural consistency with the source image.

10. The infrared and visible light image fusion method based on fine-grained semantic space alignment and bi-branch feature calibration according to claim 7, characterized in that, The perceived loss is: in This indicates that the pre-trained network is in the first... Features extracted from layers; The cross-modal consistency loss is: Used to constrain the consistency of responses from different modalities at the same semantic location; The semantic feature consistency constraint loss is: This is used to constrain the fused features to remain consistent with the text representation in the semantic space.