An infrared-visible light image fusion processing method based on cross-modal compensation and structure self-calibration

CN122415353BActive Publication Date: 2026-09-22SHANDONG JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610856072.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-15
Publication Date
2026-09-22
Estimated Expiration
2046-06-15

AI Technical Summary

Technical Problem

普通模态内相关性聚合虽然能够建模同一特征图内部不同位置之间的关系,但如果缺乏局部一致性约束,红外热噪声、可见光阴影区域中的局部异常响应或局部异常纹理可能参与其他位置的特征更新,并在特征融合和解码重建过程中被传播或放大,影响最终融合图像的局部连续性和结构自然性

Benefits of technology

(1)本发明通过特征相关性支路、响应差异引导支路和位置约束支路共同生成跨模态补偿权重,使红外模态与可见光模态之间的补偿关系同时受到特征匹配程度、模态响应差异和空间位置合理性的约束;由此能够避免仅依据特征相似性进行跨模态聚合而引入重复信息的问题,提高跨模态互补信息选择的准确性和稳定性;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122415353B_ABST
    Figure CN122415353B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of image processing, in particular to an infrared-visible light image fusion processing method based on cross-modal compensation and structure self-calibration, taking the current modal original feature map as the compensated object, taking the original feature map of another modal as the compensation information source, generating cross-modal compensation weights through three branches, and weighting and aggregating the value features of another modal by using the weights to obtain the compensated features of the current modal; meanwhile, taking the single modal original feature map as the input, using the modal internal correlation and local consistency constraints together for self-calibration weight calculation in the local neighborhood, and weighting and aggregating the value features of the same modal neighborhood by using the weights to obtain the self-calibration features of the corresponding modal; finally, integrating the feature maps and reconstructing the fusion image by the decoder. The present application simultaneously considers the constraints of the internal correlation and local structure consistency of the same modal, and reduces the influence of the neighborhood position with too large difference from the current position on the feature update.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and more specifically to an infrared-visible image fusion processing method based on cross-modal compensation and structural self-calibration. Background Technology

[0002] Image fusion is an important research direction in the field of multi-source image processing, with infrared and visible light images being two common types of heterogeneous image modalities. Infrared images can generally present high thermal radiation response areas in complex environments such as nighttime, low light, and smoke, but their ability to express texture details, edge levels, and background structure is limited. Visible light images can provide rich texture, edge, and scene details, but they are prone to problems such as insufficient local intensity response, loss of detail, and decreased contrast under low light, shadow, or complex weather conditions. Therefore, infrared and visible light images are complementary, and their fusion helps to balance the thermal radiation response information in infrared images and the texture detail information in visible light images to a certain extent, thus providing a foundation for obtaining a more complete and structurally clear image representation.

[0003] In existing image fusion methods, traditional rule-driven and transform-domain methods rely heavily on manually designed fusion rules, making it difficult to adaptively select effective information based on the modal contributions of different image regions. This often leads to problems such as blurred edges, lost textures, or insufficient local contrast. In some learning-based fusion methods, the interaction between infrared and visible light features still mainly relies on simple stitching, element-wise addition, or global weighting, making it difficult to locally model cross-modal complementary regions. Consequently, infrared intensity information and visible light texture details cannot fully participate in the fusion and reconstruction.

[0004] Furthermore, some fusion methods based on local correlation weights primarily establish cross-modal associations based on feature similarity. However, in infrared-visible image fusion, similarity is not equivalent to complementarity. Relying solely on similarity may aggregate repetitive information, making it difficult to highlight complementary regions with strong infrared responses but weak visible responses, or those with rich visible textures but insufficient infrared details. Simultaneously, spatially distant regions with similar responses may be incorrectly associated, affecting the structural consistency of the fused image. On the other hand, relying solely on modal response differences may also be affected by noise, misalignment, or anomalous responses, making it difficult to reliably distinguish between effective complementary differences and ineffective anomalous differences. Therefore, cross-modal compensation relationships need to simultaneously consider feature correlation, modal response differences, and spatial relationships to improve the rationality of compensation information selection.

[0005] Besides cross-modal complementarity, the quality of fused images is also affected by the structural stability within a single modality. While ordinary intramodal correlation aggregation can model the relationships between different locations within the same feature map, without local consistency constraints, local anomalous responses or textures in infrared thermal noise, visible light shadow regions, or other areas may participate in feature updates at other locations and be propagated or amplified during feature fusion and decoding reconstruction, affecting the local continuity and structural naturalness of the final fused image. Therefore, intramodal feature aggregation needs to consider not only the correlation within the same modality but also the constraints of local structural consistency to reduce the impact of neighboring locations with responses that differ significantly from the current location on feature updates. Summary of the Invention

[0006] The purpose of this invention is to overcome the shortcomings of existing technologies and propose an infrared-visible image fusion processing method based on cross-modal compensation and structural self-calibration. This method considers the correlation within the same mode and the constraints of local structural consistency, thereby reducing the impact of neighboring locations with large differences in response from the current location on feature updates.

[0007] The technical solution adopted by this invention to solve its technical problem is: An infrared-visible image fusion processing method based on cross-modal compensation and structural self-calibration includes: Step S1: Using the infrared image and the visible light image as input, extract the original feature maps of the infrared image and the visible light image respectively through the dual-stream feature extraction unit; Step S2: The cross-modal compensation processing unit takes the original feature map of the current modality as the object to be compensated and the original feature map of another modality as the source of compensation information. It generates cross-modal compensation weights through feature correlation branches, response difference guidance branches, and position constraint branches. These weights are then used to weighted aggregate the value features of the other modality to obtain the compensated feature map of the current modality, i.e., the cross-modal compensation feature map. Specifically, the cross-modal compensation processing unit performs compensation processing using the infrared mode and the visible light mode as the current modality, respectively, to obtain infrared compensation feature maps and visible light compensation feature maps. The structural self-calibration processing unit takes the original feature map of a single mode as input, and uses the intramodal correlation and local consistency constraints in the local neighborhood to calculate the self-calibration weight. The weight is then used to perform weighted aggregation of the neighborhood value features of the same mode, thereby obtaining the self-calibration features of the corresponding mode, i.e., the structural self-calibration feature map. Step S3: Integrate the original feature map, cross-modal compensation feature map, and structural self-calibration feature map through the feature fusion unit to obtain the enhanced fused feature map; Step S4: Reconstruct the enhanced fusion feature map into a fused image using the decoder.

[0008] Furthermore, in step S2, the processing procedure of the cross-modal compensation processing unit is as follows: Let the current mode be m The other mode is n , m and n These correspond to the infrared mode or the visible light mode, respectively, and m≠n The original feature map of the current modality is denoted as... F m ∈ R C×H×W The original feature map of another modality is denoted as F n ∈ R C×H×W ,in, C , H and W These represent the number of channels, height, and width of the feature, respectively. Spatial position in the current mode i Spatial location in another mode j ,in i , j ∈ [1, N First, the original feature map of the current modality is analyzed. F m Perform query mapping to obtain query features Q m ; For the original feature map of another modality F n Perform key mapping and value mapping separately to obtain key features. K n Sum value characteristics V n ; Q m and K n Used for subsequent similarity calculations. V n Used as compensating candidate information in a weighted aggregation in another modality; query features come from the current modality, while key features and value features come from another modality, enabling the current modality to proactively obtain complementary information related to its spatial location from the other modality; Cross-modal compensation scoring calculation includes three parallel calculation branches: feature correlation branch, response difference guided branch, and location constraint branch. Each of the three branches calculates the correlation score. S corr ij Response difference score S diff ij and position constraint score S pos ijThe current modal position is constrained from three aspects: matchability, complementarity, and spatial rationality. i With another modality candidate position j The compensation relationship between them; The correlation scores obtained from the three branches S corr ij Response difference score S diff ij and position constraint score S pos ij Perform weighted summation to obtain the current modal position. i With another modality candidate position j Cross-modal compensation scoring: ; in, Indicates the change from another mode n To the current mode m The location at which compensation is provided is used for scoring; α The moderating coefficient representing the response difference score. c This represents the adjustment coefficient for the position constraint score. Through the combined scoring of these three factors, the compensation score simultaneously considers the feature correlation, response difference, and spatial relationship between the current mode and another mode, thereby reducing the risk of aggregating repetitive information solely based on similarity or mistakenly using anomalous responses from distant locations as compensation information. Subsequently, the cross-modal compensation scores are normalized along another modal candidate position dimension to obtain the compensation weights: ; in, Indicates the current modal position i For another modal candidate position j The compensation weights are normalized at another modality candidate position. j This is performed along the dimension, so that the candidate compensation weights corresponding to the same current modal position form a standardized weight distribution; After obtaining the compensation weights, use the compensation weights For another modal value feature V n ( j Perform weighted aggregation to obtain the current position. i Cross-modal compensation aggregation results: ; in, C m i This represents the value obtained by weighted aggregation of features from another modality, used to compensate for the current modal position. iThe compensation result; Finally, the compensation aggregation results will be... C m i After output mapping W m o (•) Convert to the original feature map of the current modality F m Matching feature representations and comparing them with the original feature map of the current modality. F m Perform residual connections at the corresponding positions to obtain the compensated features of the current mode: F m c ( i )= F m ( i ) + W m o ( C m i ); in, F m c ( i ) indicates the current mode is in position i Features after compensation at the location, F m c This represents the compensated feature map of the current mode. W m o ( C m i This is used to introduce compensation information from another modality that is relevant to the current location, has a different response, and is spatially reasonable. The cross-modal compensation processing unit can provide directional complementary enhancement features for subsequent feature fusion and image reconstruction.

[0009] Furthermore, in the feature correlation branch, the current modality position is... i Query features Q m ( i ) and another modality candidate position j Key features K n ( j Similarity calculations are performed to obtain a relevance score: ; in, dThe channel dimension represents the query feature and key feature; the relevance score is used to characterize the current modality position. i With another modality candidate position j The degree of matchability in the feature space makes compensation relationships preferentially established between cross-modal position pairs with certain feature correlations; In the response difference-guided branch, the original feature map of the current modality is retrieved. F m Middle position i response characteristics F m ( i ), and another modality's original feature map F n Candidate positions j response characteristics F n ( j The difference between the two is calculated, and the response difference score is obtained through difference mapping: S diff ij = G m,n d ( F m ( i ) - F n ( j )); in, G m,n d (•) represents the difference mapping function between the current mode and another mode. The response difference score is used to reflect the differences between the two modes in the position pair ( i, j The response differences on the surface make the compensation scoring not only dependent on feature similarity, but also able to focus on response difference regions with the possibility of cross-modal complementarity; In the position-constrained branch, based on the current modal position i and another modality candidate position j Spatial coordinates are used to calculate position constraint scores; let's assume... p i and p j Representing positions respectively i and location j In the two-dimensional spatial coordinates of the feature map, the position constraint score is expressed as: S pos ij = f pos ( i, j ); in,f pos ( i, j ) represents the position constraint calculation function.

[0010] Furthermore, in step S2, the processing procedure of the structural self-calibration processing unit is as follows: For a single mode to be self-calibrated, its original feature map is denoted as: F m ∈ R Cm×H×W ,in, m It can represent infrared modes or visible light modes. C m Indicates the first m The number of channels in the original feature map of each modality. H and W These represent the feature map height and width, respectively. The structural self-calibration processing unit is in... F m Internally, local neighborhood feature aggregation is performed to generate the first... m Self-calibration feature maps of each modality; First, regarding the first m Original feature maps of each modality F m By performing query mapping, key mapping, and value mapping respectively, self-calibrating query features are obtained. Q m s Self-calibration key features K m s and self-calibration value characteristics V m s : set up p Indicates the current spatial location. q Indicates the current position p The local neighborhood position; let N ( p ) indicates by position p The local spatial neighborhood centered on the center, in q ∈ N ( p Perform feature correlation calculation and local consistency constraint within a local range; The structural self-calibration scoring includes two parallel calculation branches: an intramodal correlation branch and a local consistency branch. The two branches calculate the intramodal local correlation score respectively. S m att ( p, q Local consistency weight M m ( p, q ); To incorporate local consistency constraints into the attention normalization process, the local consistency weights are converted into modulation terms, and intramodal local correlation scores are added to obtain the structural self-calibration score: ; in, or The weighting coefficients of the local consistency modulation term are represented. e A very small constant is set to avoid numerical instability in logarithmic operations; since the local consistency modulation term is added to the score before normalization, the subsequent normalization can still obtain a normal attention weight distribution, avoiding the problem of the weight distribution losing normalization constraints due to directly multiplying by the local consistency weight after Softmax; Then, at the current location p local neighborhood N ( p The structural self-calibration score is normalized to obtain the structural self-calibration weight: ; in, A m s ( p, q () indicates the current position p Neighborhood location q The structure self-calibration weight; this weight simultaneously considers feature correlation and local consistency within the same mode, so that the feature update of the current position prioritizes the reference of the neighboring position that has a high correlation with it and a relatively consistent local response. Utilizing structural self-calibration weights to characterize the self-calibration values ​​of the same mode V m s ( q Perform weighted aggregation to obtain the current position. p Intramodal self-calibration aggregation results: ; in, R m s ( p ) indicates that by the first m The self-calibrated aggregation result is obtained by weighted aggregation of neighborhood value features within each modality; Finally, the intramodal self-calibration aggregation results are... R m s ( p After output mapping W m so (•), and with the original features F m ( p Perform residual join to obtain the first...m The modality at position p Self-calibration characteristics at the location: F m s ( p ) = F m ( p ) + W m so ( R m s ( p )); in, F m s ( p ) indicates the first m The modality at position p Self-calibration characteristics at the location, W m so (•) denotes the self-calibration output mapping function. The self-calibration aggregation result after output mapping is used to introduce relevant and consistent structural information within the local neighborhood of the same mode.

[0011] Furthermore, in the intramodal correlation branch, the current position is... p Self-calibration query features Q m s ( p ) and neighborhood location q Self-calibration bond features K m s ( q Similarity calculations are performed to obtain intramodal local correlation scores: ; in, d This represents the channel dimension of the self-calibration query features and the self-calibration key features. The intra-modal local correlation score is used to characterize the current position. p With neighboring locations q The degree of correlation within the same modal feature space; In the locally consistent branch, based on the original feature map F m Current location p response characteristics F m ( p and neighborhood location q response characteristics F m ( qCalculate the local consistency weight: ; in, M m ( p, q ) indicates the first m Position in each modality p Its neighboring location q The local consistency weight between them, ||•||1 represents L 1-norm, β This represents the local difference sensitivity coefficient. The local consistency weight is used to reflect the neighborhood location. q With current location p The degree of similarity in the original modal characteristic response.

[0012] Furthermore, in step S3, the processing procedure of the feature fusion unit is as follows: Align feature maps from different sources by channel and spatial dimensions; The aligned original feature map, compensated feature map, and self-calibrated feature map are concatenated along the channel dimension to obtain the concatenated multi-source feature map. F cat ; To map multi-source features from different origins to a unified fusion space, a convolutional fusion mapping pair is used. F cat Feature integration is performed to obtain an enhanced fused feature map. F fuse : F fuse = f ( W f ( F cat )); in, W f (•) represents a convolutional fusion mapping, which can be composed of 1×1 convolution, 3×3 convolution or a combination thereof; f (•) represents a non-linear activation function.

[0013] The collaborative input of the aforementioned multi-source features forms an enhanced fusion feature map, providing a feature basis for subsequent decoder to reconstruct the fused image.

[0014] Furthermore, in step S4, the decoder reconstructs the enhanced fusion feature map into a fused image as follows: Enhance the fusion feature map F fuse As the initial input to the decoder, let it be denoted as R 0= F fuse; in the l In the multi-level decoding process, the features reconstructed from the previous layer are first... R l-1 Upsampling is performed to obtain reconstructed features with improved spatial resolution; To supplement low-level edge, texture, and local spatial structure information during the decoding stage, shallow structural features from the two-stream feature extraction stage are introduced. Let the two-stream feature extraction unit in step S1 be at the [missing information - likely a specific step or stage]. l The infrared shallow features output at the corresponding scale are S ir l The shallow visible light characteristics are S vis l To simultaneously utilize the thermal response structure in infrared images and the texture edge structure in visible light images, the two are shallowly fused and mapped to obtain the first... l Shallow structural features: S l =Φ shallow l ( Concat ( S ir l , S vis l )); Where, Φ shallow l (•) indicates shallow structure fusion mapping; S l Indicates the fused first l Shallow structural features; In the shallow structure feature S l Before adding it to the upsampled reconstructed features, ψ is first mapped by channel and scale alignment. l (•) for S l Process it to make it consistent with U l ( R l-1 They have the same spatial dimensions and number of channels; l The level decoding process is represented as: ; ; in, U l (•) indicates the first l Level upsampling operation; ψ l (•) represents the mapping function used for channel and scale alignment; This represents the intermediate reconstruction features after fusing shallow structural information; Cl (•) indicates the first l Level convolution reconstruction operation; R l Indicates the first l The reconstructed features of the output are determined by the number of levels. The output spatial size of each decoding unit gradually increases with the number of levels until the spatial resolution corresponding to the fused image of the target is restored. go through L After the levels are gradually restored, the final reconstructed features are obtained. R L Subsequently, a fused image is generated through output convolution and pixel range normalization mapping: I fuse = s ( W out ( R L )); in, W out (•) indicates the output convolution mapping; s (•) represents the pixel range normalization function, which can be used to constrain the output pixel values ​​to the range [0,1] by using the Sigmoid function or normalization mapping.

[0015] Through the above-mentioned step-by-step restoration structure, the decoder can introduce shallow spatial information from the dual-stream feature extraction stage while restoring spatial resolution, so that the complementary information provided by the cross-modal compensation processing unit, the intramodal stable structural information provided by the structural self-calibration processing unit, and the basic structural information in the original features are jointly transformed into the final fused image.

[0016] Furthermore, the decoder includes: Upsampling units are used to progressively restore the spatial resolution of the feature map; Convolutional reconstruction units are used to reconstruct local edges, textures, and region structures; Shallow structural connections are used to supplement the low-level spatial information retained during the feature extraction stage in the spatial resolution restoration process; The output mapping unit is used to generate the final fused image.

[0017] Technical effects of the present invention: Compared with existing technologies, the infrared-visible image fusion processing method based on cross-modal compensation and structural self-calibration of the present invention has the following advantages: (1) The present invention generates cross-modal compensation weights by using feature correlation branch, response difference guidance branch and position constraint branch together, so that the compensation relationship between infrared mode and visible light mode is simultaneously constrained by feature matching degree, modal response difference and spatial position rationality; thereby avoiding the problem of introducing duplicate information by cross-modal aggregation based solely on feature similarity, and improving the accuracy and stability of cross-modal complementary information selection; (2) In the cross-modal compensation process, the present invention introduces response difference guidance, so that the compensation weight not only focuses on the similar regions between the two modes, but also on complementary regions with strong infrared response but insufficient visible light response, or rich visible light texture but insufficient infrared detail. Therefore, it can enhance the comprehensive expression of infrared target intensity information, visible light texture detail information and scene background structure information of the fused image, and improve the problems of target not being prominent, texture missing or insufficient local contrast in the fused image; (3) The present invention restricts the spatial rationality of cross-modal candidate compensation positions by using position constraint branches, so that the compensation information comes first from areas with similar spatial positions or reasonable structural correspondences; thereby reducing the impact of slight image misalignment, abnormal response or distant similar areas on the compensation results, reducing the probability of spatially distant but similar response areas being incorrectly associated, and improving the structural consistency in the cross-modal feature aggregation process; (4) The present invention sets up a structural self-calibration processing unit, which simultaneously uses intramodal correlation and local consistency constraints to calculate self-calibration weights in the local neighborhood within a single modality, so that the feature update of the current position prioritizes the reference of the neighborhood position with higher correlation and more consistent local response; thus, it can reduce the propagation of infrared thermal noise, visible light shadows, local abnormal textures or isolated abnormal responses in the feature aggregation and decoding reconstruction process, and improve the local continuity and structural naturalness of the fused image; (5) The present invention integrates the original feature map, cross-modal compensation feature map and structural self-calibration feature map through the feature fusion unit, so that the basic spatial structure in the original feature, the complementary information in the cross-modal compensation feature and the stable structural information in the self-calibration feature can jointly participate in the reconstruction of the fused image; thereby improving the integrity of the fused feature expression and enabling the final fused image to achieve better results in terms of target saliency, texture detail preservation and local structural continuity. Attached Figure Description

[0018] Figure 1 This is a block diagram of the image fusion processing structure of the present invention; Figure 2 This is a block diagram of the cross-modal compensation processing unit of the present invention; Figure 3 This is a structural block diagram of the self-calibration processing unit of the present invention. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.

[0020] Example 1: like Figure 1 As shown, this embodiment relates to an infrared-visible image fusion processing method based on cross-modal compensation and structural self-calibration, comprising: Let the infrared image be I ir Visible light image is I vis The two images are different modalities captured in the same scene. The sizes of the two modalities are unified and then input into the dual-stream feature extraction unit. For image pairs with slight spatial deviations, weak registration can be performed by overlapping region cropping, center alignment, affine transformation, feature point registration, or a combination thereof, so that the main edges, contours, and background structures of the two images have a basic spatial correspondence at the feature map scale. The slight spatial deviation can be set as an offset within a preset ratio range of the feature map width or height, or as an offset within 1 to 5 feature map pixels. The dual-stream feature extraction unit extracts features from the infrared image and the visible light image respectively, obtaining the original infrared feature map and the original visible light feature map: F ir = E ir ( I ir ); F vis = E vis ( I vis ); in, E ir and E vis These represent the infrared feature extraction branch and the visible light feature extraction branch, respectively. In order to meet the needs of subsequent cross-modal compensation, feature fusion and image reconstruction, this invention retains the two-dimensional spatial structure of the original feature map instead of globally pooling it into a feature vector. After channel mapping and scale alignment, the original infrared feature map and the original visible light feature map are uniformly represented as follows: F ir ,F vis ∈R C×H’×W’ ; in, CIndicates the number of feature channels. H’ and W’ Indicates the spatial dimensions of the feature map; To reduce cross-modal error association caused by spatial offset, positional constraints are introduced in the subsequent cross-modal compensation processing unit so that the calculation of local association weights considers both feature response and spatial proximity. Overall image fusion comprises two parallel enhancement paths: a cross-modal compensation path and a structural self-calibration path. The cross-modal compensation path models the complementary relationship between infrared and visible light features, outputting an infrared compensated feature map. F ir c Visible light compensation feature map F vis c The structural self-calibration path is used to enhance the structural stability within a single mode and outputs an infrared self-calibration feature map. F ir S Visible light self-calibration feature map F vis S Subsequently, the original feature map, the compensated feature map, and the self-calibrated feature map are input into the feature fusion unit to obtain the enhanced fusion feature map. F fuse The decoder outputs the fused image: I fuse = D ( F fuse ); in, I fuse This represents the final fused image output. D (•) indicates the decoder.

[0021] 1. Cross-modal compensation processing unit The processing procedure of the cross-modal compensation processing unit described in this embodiment is as follows: Figure 2 As shown, the cross-modal compensation processing unit is used to establish a compensation relationship between the current modal feature and another modal feature, so that the current modality can obtain effective compensation information from another modality based on the feature matching relationship, modal response difference and spatial position constraints.

[0022] Specifically, let the current mode be... m The other mode is n ,in m and n These correspond to the infrared mode or the visible light mode, respectively, and m≠n The original feature map of the current modality is denoted as... F m ∈ R C×H×WThe original feature map of another modality is denoted as F n ∈ R C×H×W ,in, C , H and W These represent the number of channels, height, and width of the feature, respectively; to facilitate calculations across spatial locations, the spatial location of the feature map is unfolded as follows: N = H × W One position.

[0023] Spatial position in the current mode i Spatial location in another mode j ,in i , j ∈ [1, N First, the original feature map of the current modality is analyzed. F m Perform query mapping to obtain query features Q m ; For the original feature map of another modality F n Perform key mapping and value mapping separately to obtain key features. K n Sum value characteristics V n , is represented as: Q m = W m q ( F m ); K n = W n k ( F n ); V n = W n v ( F n ); in, W m q (•) indicates the query mapping for the current modality. W n k (•) indicates a key mapping for another mode. W n v (•) indicates a value mapping for another mode;Q m and K n Used for subsequent similarity calculations. V n Used as compensating candidate information in another modality that is weighted and aggregated; query features come from the current modality, key features and value features come from another modality, enabling the current modality to actively obtain complementary information related to its spatial location from another modality.

[0024] Cross-modal compensation scoring calculation includes three parallel calculation branches: feature correlation branch, response difference-guided branch, and position constraint branch. These three branches constrain the current modal position from the aspects of matchability, complementarity, and spatial rationality, respectively. i With another modality candidate position j The compensation relationship between them.

[0025] In the feature correlation branch, the current modality position is... i Query features Q m ( i ) and another modality candidate position j Key features K n ( j Similarity calculations are performed to obtain a relevance score: ; in, d The channel dimension represents the query feature and key feature; the relevance score is used to characterize the current modality position. i With another modality candidate position j The degree of matchability in the feature space makes compensation relationships preferentially established between cross-modal position pairs with certain feature correlations.

[0026] In the response difference-guided branch, the original feature map of the current modality is retrieved. F m Middle position i response characteristics F m ( i ), and another modality's original feature map F n Candidate positions j response characteristics F n ( j The difference between the two is calculated, and the response difference score is obtained through difference mapping: S diff ij = G m,nd ( F m ( i ) - F n ( j )); in, G m,n d (•) represents the difference mapping function between the current mode and another mode. In other implementations, the difference calculation may also employ absolute difference, concatenated difference, or a combination thereof. F m and F n When the channel dimensions are inconsistent, they can be adjusted to the same dimension through channel mapping before difference calculation; the response difference score is used to reflect the positional differences between the two modes ( i, j The response differences on the surface allow the compensation scoring to not only rely on feature similarity, but also to focus on response difference regions with the potential for cross-modal complementarity.

[0027] In the position-constrained branch, based on the current modal position i and another modality candidate position j Spatial coordinates are used to calculate position constraint scores; let's assume... p i and p j Representing positions respectively i and location j In the two-dimensional spatial coordinates of the feature map, the position constraint score can be expressed as: S pos ij = f pos ( i, j ); in, f pos ( i, j The ) represents the position constraint calculation function. As an optional method, the position constraint score can be determined based on the position... i With position j Spatial distance construction between them, for example: ; in, s pThis represents the location constraint scale parameter. Following the above format, spatially closer location pairs receive higher location constraint scores, while spatially farther location pairs receive lower scores, thereby reducing the probability of unrelated, distant locations being incorrectly associated in subsequent compensation scoring. In other implementations, location constraint scores can also be obtained through location encoding, distance lookup tables, or learnable location biases.

[0028] The current modal position is obtained by weighting the correlation score, response difference score, and position constraint score obtained from the above three branches. i With another modality candidate position j Cross-modal compensation scoring: ; in, Indicates the change from another mode n To the current mode m The location at which compensation is provided is used for scoring; α The moderating coefficient representing the response difference score. c This represents the adjustment coefficient for the position constraint score. Through the combined scoring of these three factors, the compensation score simultaneously considers the feature correlation, response difference, and spatial relationship between the current mode and another mode, thereby reducing the risk of aggregating repetitive information solely based on similarity or mistakenly using anomalous responses from distant locations as compensation information.

[0029] Subsequently, the cross-modal compensation scores are normalized along another modal candidate position dimension to obtain the compensation weights: ; in, Indicates the current modal position i For another modal candidate position j The compensation weights are normalized at another modality candidate position. j This is performed on the dimension of the algorithm, so that the candidate compensation weights corresponding to the same current modal position form a standardized weight distribution.

[0030] After obtaining the compensation weights, use the compensation weights For another modal value feature V n ( j Perform weighted aggregation to obtain the current position. i Cross-modal compensation aggregation results: ; in, C m i This represents the value obtained by weighted aggregation of features from another modality, used to compensate for the current modal position. i The compensation result.

[0031] Finally, the compensation aggregation results will be... C m i After output mapping W m o (•) Convert to the original feature map of the current modality F m Matching feature representations and comparing them with the original feature map of the current modality. F m Perform residual connections at the corresponding positions to obtain the compensated features of the current mode: F m c ( i )= F m ( i ) + W m o ( C m i ); in, F m c ( i ) indicates the current mode is in position i Features after compensation at the location, F m c This represents the compensated feature map of the current mode. W m o ( C m i It is used to introduce compensation information from another mode that is related to the current position, has a response difference, and is spatially reasonable.

[0032] Current mode m Visible light mode, another mode n When the mode is infrared, the compensated feature map is denoted as... F vis c This helps to introduce infrared thermal radiation response information into visible light characteristics; current mode m Infrared mode, another mode n When the mode is visible light, the compensated feature map is denoted as: F ir c This helps to introduce visible light texture, edge, and background details into infrared features. Therefore, the cross-modal compensation processing unit can provide directional complementary enhancement features for subsequent feature fusion and image reconstruction.

[0033] 2. Structural self-calibration processing unit The processing procedure of the structural self-calibration processing unit described in this embodiment is as follows: Figure 3 As shown, the structural self-calibration processing unit is used to perform intramodal local self-calibration on the original feature map of a single mode to enhance the local structural stability within the single mode. The cross-modal compensation processing unit is mainly used to enhance the utilization of complementary information between infrared and visible light images, while the structural self-calibration processing unit is used to constrain the feature aggregation process within a single mode, reducing the impact of isolated noise, local anomalous responses, and responses that differ too much from the neighborhood structure on feature updates.

[0034] The structural self-calibration processing unit takes the original feature map of a single mode as input, rather than the mixed features after cross-modal compensation.

[0035] Specifically, for a single mode to be self-calibrated, its original feature map is denoted as: F m ∈ R Cm×H×W ,in, m It can represent infrared modes or visible light modes. C m Indicates the first m The number of channels in the original feature map of each modality. H and W These represent the feature map height and width, respectively. The structural self-calibration processing unit is in... F m Internally, local neighborhood feature aggregation is performed to generate the first... m Self-calibration feature map of each modality.

[0036] First, regarding the first m Original feature maps of each modality F m By performing query mapping, key mapping, and value mapping respectively, self-calibrating query features are obtained. Q m s Self-calibration key features K m s and self-calibration value characteristics V m s : Q m s = W m sq ( F m ); K ms = W m sk ( F m ); V m s = W m sv ( F m ); in, W m sq (•) W m sk (•)and W m sv (•) represent the corresponding query map, key map, and value map, respectively, which can be implemented by 1×1 convolution, linear mapping, or a combination thereof. The query features, key features, and value features mentioned above all come from the original feature map of the same modality. F m Therefore, this process is an intramodal self-calibration rather than a cross-modal compensation.

[0037] set up p Indicates the current spatial location. q Indicates the current position p The local neighborhood location. Let N ( p ) indicates by position p A local spatial neighborhood centered on a location, wherein the local spatial neighborhood can be set based on location. p Centered r × r A window, a fixed-radius neighborhood, or a preset set of local candidate locations. The structural self-calibration processing unit... q ∈ N ( p The structural self-calibration scoring performs feature correlation calculations and local consistency constraints within a local range. The scoring includes two parallel calculation branches: the intramodal correlation branch and the local consistency branch.

[0038] In the intramodal correlation branch, the current position is... p Self-calibration query features Q m s ( p ) and neighborhood location q Self-calibration bond features K m s ( qSimilarity calculations are performed to obtain intramodal local correlation scores: ; in, d This represents the channel dimension of the self-calibration query features and the self-calibration key features. The intra-modal local correlation score is used to characterize the current position. p With neighboring locations q The degree of correlation within the same modal feature space.

[0039] In the locally consistent branch, based on the original feature map F m Current location p response characteristics F m ( p and neighborhood location q response characteristics F m ( q Calculate the local consistency weight: ; in, M m ( p, q ) indicates the first m Position in each modality p Its neighboring location q The local consistency weight between them, ||•||1 represents L 1-norm, β This represents the local difference sensitivity coefficient. The local consistency weight is used to reflect the neighborhood location. q With current location p The degree of similarity in the original modal characteristic response.

[0040] To incorporate local consistency constraints into the attention normalization process, the local consistency weights are converted into modulation terms, and intramodal local correlation scores are added to obtain the structural self-calibration score: ; in, or The weighting coefficients of the local consistency modulation term are represented. e A very small constant is set to avoid numerical instability in logarithmic operations. Since the local consistency modulation term is added to the score before normalization, the subsequent normalization can still obtain a normalized attention weight distribution, avoiding the problem of the weight distribution losing normalization constraints due to directly multiplying by the local consistency weight after Softmax.

[0041] Then, at the current location p local neighborhood N ( pThe structural self-calibration score is normalized to obtain the structural self-calibration weight: ; in, A m s ( p, q () indicates the current position p Neighborhood location q The structure self-calibration weights consider both feature correlation and local consistency within the same modality, ensuring that feature updates at the current position preferentially reference neighboring positions with high correlation and consistent local responses.

[0042] Using this weight to define the self-calibration value characteristics of the same mode V m s ( q Perform weighted aggregation to obtain the current position. p Intramodal self-calibration aggregation results: ; in, R m s ( p ) indicates that by the first m The self-calibrated aggregation result is obtained by weighted aggregation of neighborhood value features within each modality.

[0043] Finally, the intramodal self-calibration aggregation results are... R m s ( p After output mapping W m so (•), and with the original features F m ( p Perform residual join to obtain the first... m The modality at position p Self-calibration characteristics at the location: F m s ( p ) = F m ( p ) + W m so ( R m s ( p )); in, F ms ( p ) indicates the first m The modality at position p Self-calibration characteristics at the location, W m so (•) denotes the self-calibration output mapping function. The self-calibration aggregation result after output mapping is used to introduce relevant and consistent structural information within the local neighborhood of the same mode.

[0044] Through the above process, the structure self-calibration processing unit can perform local structural constraints within a single mode, making the feature update process not only dependent on intra-modal correlations but also subject to local consistency modulation. For the infrared mode, this unit helps reduce the impact of isolated thermal noise and local anomalous responses during the intra-modal aggregation process; for the visible light mode, this unit helps reduce the interference of local shadows, anomalous texture responses, and neighborhood structural discontinuities on feature updates.

[0045] Therefore, the structural self-calibration processing unit does not simply perform ordinary autocorrelation enhancement on single-modal features, but introduces local consistency constraints during the intramodal local aggregation process, making the feature updates of the same mode more consistent with the requirements of local spatial structural continuity. The resulting infrared and visible light self-calibration feature maps can provide more structurally stable intramodal features for subsequent feature fusion and decoding reconstruction.

[0046] 3. Feature fusion unit After dual-stream feature extraction, cross-modal compensation, and structural self-calibration, the processed structure yields three types of feature maps: original feature map, compensated feature map, and self-calibrated feature map. Among these, the original infrared feature map... F ir and raw feature map of visible light F vis Used to preserve the basic spatial structure of the input image; infrared compensated feature map F ir c This represents the compensated feature obtained by using the infrared mode as the current mode and providing compensation information from the visible light mode. It can be used to introduce texture, edge, and background details from the visible light image into the infrared feature; Visible light compensated feature map. F vis c This represents the compensated feature obtained by using the visible light mode as the current mode and providing compensation information from the infrared mode. It can be used to introduce the thermal radiation intensity response from the infrared image into the visible light features. (Infrared self-calibration feature map) F ir s Visible light self-calibration feature map F vis sThey are used to provide stable features within the corresponding single mode after local structural constraints.

[0047] Therefore, the original feature map, the compensated feature map, and the self-calibrated feature map correspond to basic structural information, cross-modal complementary information, and intramodal structural stability information, respectively. The feature fusion unit is used to map features from these different sources to a unified fusion space to form an enhanced fusion feature map for subsequent image reconstruction.

[0048] Before feature fusion, feature maps from different sources are first aligned in terms of channel and spatial dimensions. For feature maps with different numbers of channels, the channel dimension can be unified using 1×1 convolution; for feature maps with different spatial dimensions, the spatial resolution can be unified using upsampling or downsampling. Let the aligned feature maps be... , , , , , After alignment, each feature map has a consistent spatial size and meets the requirements for channel dimension splicing.

[0049] Subsequently, the aligned original feature map, compensated feature map, and self-calibrated feature map are concatenated along the channel dimension: ; in, F cat This represents the concatenated multi-source feature map. Concat (•) indicates a splicing operation on the channel dimension.

[0050] To map multi-source features from different origins to a unified fusion space, a convolutional fusion mapping pair is used. F cat Perform feature integration: F fuse = f ( W f ( F cat )); in, W f (•) represents a convolutional fusion mapping, which can be composed of 1×1 convolution, 3×3 convolution or a combination thereof; f (•) represents a nonlinear activation function; F fuse This represents the enhanced fusion feature map.

[0051] The feature fusion unit described in this embodiment does not rely solely on channel stitching to achieve its technical effect. Instead, it inputs the original structural information, cross-modal compensation information, and intra-modal self-calibration information into the fusion map, enabling the convolutional fusion map to reorganize and select information from different sources within a unified feature space. The original feature map preserves the basic spatial structure of the input image, the compensation feature map introduces complementary image information from another modality, and the self-calibration feature map provides a stable response after intra-modal local structural constraints. The collaborative input of these multi-source features forms an enhanced fusion feature map, providing a feature foundation for subsequent decoder image reconstruction.

[0052] 4. Decoder and Fusion Image Reconstruction Enhanced fused feature map output by the feature fusion unit F fuse Still residing in a high-dimensional feature space, the enhanced fusion feature map cannot be directly output as the fused image. Therefore, a stepwise restoration decoder is designed to gradually map the enhanced fusion feature map back to the image space. The decoder includes an upsampling unit, a convolutional reconstruction unit, shallow structural connections, and an output mapping unit. The upsampling unit is used to progressively restore the spatial resolution of the feature map; the convolutional reconstruction unit is used to reconstruct local edges, textures, and region structures; the shallow structural connections are used to supplement the low-level spatial information retained during the feature extraction stage during the spatial resolution restoration process; and the output mapping unit is used to generate the final fused image.

[0053] Enhance the fusion feature map F fuse As the initial input to the decoder, let it be denoted as R 0= F fuse In the first l In the multi-level decoding process, the features reconstructed from the previous layer are first... R l-1 Upsampling is performed to obtain reconstructed features with improved spatial resolution. The upsampling operation can be implemented using at least one of bilinear interpolation, nearest neighbor interpolation, deconvolution, or pixel rearrangement.

[0054] To supplement low-level edge, texture, and local spatial structure information during the decoding stage, shallow structural features from the two-stream feature extraction stage are introduced. Let the two-stream feature extraction unit in step S1 be at the [missing information - likely a specific step or stage]. l The infrared shallow features output at the corresponding scale are S ir l The shallow visible light characteristics are S vis l To simultaneously utilize the thermal response structure in infrared images and the texture edge structure in visible light images, the two are shallowly fused and mapped to obtain the first... l Shallow structural features: S l =Φ shallow l ( Concat ( S ir l , S vis l )); Where, Φ shallow l (•) represents a shallow structure fusion mapping, which can be implemented by 1×1 convolution, 3×3 convolution, or a combination thereof; S l Indicates the fused first l Shallow structural features.

[0055] In the shallow structure feature S l Before adding it to the upsampled reconstructed features, ψ is first mapped by channel and scale alignment. l (•) for S l Process it to make it consistent with U l ( R l-1 They have the same spatial dimensions and number of channels. l The level decoding process is represented as: ; ; in, U l (•) indicates the first l Level upsampling operation; ψ l (•) represents the mapping function used for channel and scale alignment; This represents the intermediate reconstruction features after fusing shallow structural information; C l (•) indicates the first l Level convolution reconstruction operation; R l Indicates the first l The reconstructed features of the output at each stage. The output spatial size of each decoding unit gradually increases with the number of stages until it is restored to the spatial resolution corresponding to the fused image of the target.

[0056] go through L After the levels are gradually restored, the final reconstructed features are obtained. R L Subsequently, a fused image is generated through output convolution and pixel range normalization mapping: I fuse = s ( Wout ( R L )); in, W out (•) indicates the output convolution mapping; s (•) represents the pixel range normalization function, which can be used to constrain the output pixel values ​​to the range [0,1] by using the Sigmoid function or normalization mapping.

[0057] For grayscale fusion tasks, the output image is the fused image. I fuse It can be set to a single-channel image. In an optional implementation, when the visible light image is a color image, the visible light image can be converted to a luminance-chrominance space, and the visible light luminance channel can be fused with the infrared image to obtain a fused luminance channel. The fused luminance channel can then be combined with the chrominance channel of the visible light image to generate a color fused image.

[0058] Through the above-mentioned step-by-step restoration structure, the decoder can introduce shallow spatial information from the dual-stream feature extraction stage while restoring spatial resolution, so that the complementary information provided by the cross-modal compensation processing unit, the intramodal stable structural information provided by the structural self-calibration processing unit, and the basic structural information in the original features are jointly transformed into the final fused image.

[0059] 5. Training and Reasoning Process During the training phase, pairs of infrared and visible light images are input. Let the input infrared image be... I ir Input visible light image is I vis For color visible light images, they can be converted to a luminance-chrominance space, and the luminance channel can be extracted. Y vis The training constraints are applied; for grayscale visible light images, they can be directly used as visible light brightness images. After the input image is size-uniformed, it sequentially passes through a dual-stream feature extraction unit, a cross-modal compensation processing unit, a structure self-calibration processing unit, a feature fusion unit, and a decoder to generate a fused image. I fuse .

[0060] For infrared-visible image pairs with slight spatial deviations, weak registration, cropping alignment, or scale unification can be performed during the input stage to ensure that the two modalities have a basic correspondence in terms of major edges, target contours, and background structures, thereby improving the reliability of pixel-level intensity constraints, gradient constraints, and structural constraints during training.

[0061] Since infrared-visible image fusion typically lacks a real fused image as a supervisory label, the fusion processing parameters can be optimized by reconstructing the target class from a self-supervised image without a real fused label. The training loss can be expressed as: L = l 1 L int +λ 2 L grad +λ 3 L ssim +λ 4 L tv ; in, L int Indicates strength retention loss, L grad This indicates a loss in edge texture preservation. L ssim This represents the loss of structural consistency. L tv This represents the smoothing constraint loss. l 1. l 2. l 3. l 4 represents the weight parameter of the corresponding loss term.

[0062] Strength retention loss L int This is used to constrain the fused image to retain the main intensity response information from the input image. Alternatively, it can be based on the fused image... I fuse With infrared images I ir Visible light brightness image Y vis Calculation of the strength difference between them: L int = || I fuse - max ( I ir , Y vis )||1; in, max ( I ir , Y visThe value represents the larger intensity response from the infrared image and the visible light brightness image at the corresponding pixel location. This constraint allows the fused image to retain infrared intensity information in areas of high thermal response, while retaining visible light brightness information in areas of strong visible light brightness response. In other embodiments, the intensity retention loss can also be calculated using a weighted combination of the infrared image and the visible light brightness image.

[0063] Edge texture preservation loss L grad This is used to constrain the fused image to retain edge and texture information from the input image. Alternatively, gradients can be calculated based on gradient operators for the fused image, infrared image, and visible light brightness image, making the gradient of the fused image approximate the larger response in the input image gradient. ; Here, ▽ represents the gradient calculation operator, which can be the Sobel operator, the difference operator, or other edge gradient operators. This loss is used to enhance the ability of the fused image to preserve the outline of infrared targets and the edges of visible light textures.

[0064] Structural consistency loss L ssim This is used to constrain the consistency of the fused image with the input image in terms of local structure. As an optional approach, the structural similarity between the fused image and the infrared image and the visible light brightness image can be calculated separately, and the following loss can be constructed: L ssim = (1 - SSIM ( I fuse , I ir )) + (1 - SSIM ( I fuse , Y vis )); in, SSIM (•) denotes the structural similarity metric function. This loss is used to make the fused image reference both the infrared and visible light brightness images in terms of local structure, reducing structural distortion and local discontinuities.

[0065] Smoothing Constraint Loss L tv This is used to suppress isolated noise and local discontinuities in the fused image. As an alternative, a total variational constraint can be employed: ; in,( x , y) represents the pixel location in the image. This loss is used to improve the smoothness and continuity of local regions in the fused image, but its weight can be set to be less than the weights of the intensity preservation loss and the edge texture preservation loss to avoid over-smoothing that would weaken edge and texture details.

[0066] During the inference phase, the image fusion processing parameters are fixed. After inputting a pair of infrared and visible light images of the same scene, the images undergo dual-stream feature extraction, cross-modal compensation, structural self-calibration, feature fusion, and decoding reconstruction sequentially, following the same forward processing path as in the training phase, to generate a fused image. I fuse During the inference process, there is no need to manually set pixel-level fusion rules or fixed fusion weights. The output fused image can be used for display, storage, transmission, image enhancement observation, or subsequent image processing.

[0067] The above-described specific embodiments are merely specific examples of the present invention. The patent protection scope of the present invention includes, but is not limited to, the above-described specific embodiments. Any appropriate changes or modifications made by a person skilled in the art that conform to the claims of the present invention should fall within the patent protection scope of the present invention.

Claims

1. An infrared-visible image fusion processing method based on cross-modal compensation and structural self-calibration, characterized in that, include: Step S1: Using the infrared image and the visible light image as input, extract the original feature maps of the infrared image and the visible light image respectively through the dual-stream feature extraction unit; Step S2: The cross-modal compensation processing unit takes the original feature map of the current modality as the object to be compensated and the original feature map of another modality as the source of compensation information. It generates cross-modal compensation weights through feature correlation branches, response difference guidance branches, and position constraint branches. These weights are then used to weight and aggregate the features of the other modality to obtain the compensated feature map of the current modality, i.e., the cross-modal compensation feature map. Specifically, the cross-modal compensation processing unit performs compensation processing using the infrared mode and the visible light mode as the current modality, respectively, to obtain infrared compensation feature maps and visible light compensation feature maps. The structural self-calibration processing unit takes the original feature map of a single mode as input, and uses the intramodal correlation and local consistency constraints in the local neighborhood to calculate the self-calibration weight. The self-calibration weight is then used to weight and aggregate the neighborhood value features of the same mode to obtain the self-calibration features of the corresponding mode, i.e., the structural self-calibration feature map. Step S3: Integrate the original feature map, cross-modal compensation feature map, and structural self-calibration feature map through the feature fusion unit to obtain the enhanced fused feature map; Step S4: Reconstruct the enhanced fusion feature map into a fused image using the decoder.

2. The infrared-visible image fusion processing method based on cross-modal compensation and structural self-calibration according to claim 1, characterized in that, In step S2, the processing procedure of the cross-modal compensation processing unit is as follows: Let the current mode be m The other mode is n , m and n These correspond to the infrared mode or the visible light mode, respectively, and m≠n The original feature map of the current modality is denoted as... F m ∈ R C×H×W The original feature map of another modality is denoted as F n ∈ R C×H×W ,in, C , H and W These represent the number of channels, height, and width of the feature, respectively. Spatial position in the current mode i Spatial location in another mode j ,in i , j ∈ [1, N First, the original feature map of the current modality is analyzed. F m Perform query mapping to obtain query features Q m ; For the original feature map of another modality F n Perform key mapping and value mapping separately to obtain key features. K n Sum value characteristics V n ; Cross-modal compensation scoring calculation includes three parallel calculation branches: feature correlation branch, response difference guided branch, and location constraint branch. Each of the three branches calculates the correlation score. S corr ij Response difference score S diff ij and position constraint score S pos ij ; The correlation scores obtained from the three branches S corr ij Response difference score S diff ij and position constraint score S pos ij Perform weighted summation to obtain the current modal position. i With another modality candidate position j Cross-modal compensation scoring: ; in, α The moderating coefficient representing the response difference score. γ The adjustment coefficient representing the position constraint score; Subsequently, the cross-modal compensation scores are normalized along another modal candidate position dimension to obtain the compensation weights: ; Among them, normalization at another modality candidate position j It is carried out in terms of dimensions; After obtaining the compensation weights, use the compensation weights For another modal value feature V n ( j Perform weighted aggregation to obtain the current position. i Cross-modal compensation aggregation results: ; Finally, the compensation aggregation result will be... C m i After output mapping W m o (•) Convert to the original feature map of the current modality F m Matching feature representations and comparing them with the original feature map of the current modality. F m Perform residual connections at the corresponding positions to obtain the compensated features of the current mode: F m c ( i )= F m ( i ) + W m o ( C m i ); in, F m c This represents the compensated feature map of the current mode. W m o ( C m i It is used to introduce compensation information from another mode that is related to the current position, has a response difference, and is spatially reasonable.

3. The infrared-visible image fusion processing method based on cross-modal compensation and structural self-calibration according to claim 2, characterized in that, in In the feature correlation branch, the current modality position is... i Query features Q m ( i ) and another modality candidate position j Key features K n ( j Similarity calculations are performed to obtain a relevance score: ; in, d The channel dimension represents the query feature and key feature; the relevance score is used to characterize the current modality position. i With another modality candidate position j The degree of matching in the feature space; In the response difference-guided branch, the original feature map of the current modality is retrieved. F m Middle position i response characteristics F m ( i ), and another modality's original feature map F n Candidate positions j response characteristics F n ( j The difference between the two is calculated, and the response difference score is obtained through difference mapping: S diff ij = G m,n d ( F m ( i ) - F n ( j )); in, G m,n d (•) represents the difference mapping function between the current mode and another mode. The response difference score is used to reflect the positional differences between the two modes. i, j Differences in response on ) In the position-constrained branch, based on the current modal position i and another modality candidate position j Spatial coordinates are used to calculate position constraint scores; let's assume... p i and p j Representing positions respectively i and location j In the two-dimensional spatial coordinates of the feature map, the position constraint score is expressed as: S pos ij = f pos ( i, j ); in, f pos ( i, j ) represents the position constraint calculation function.

4. The infrared-visible image fusion processing method based on cross-modal compensation and structural self-calibration according to claim 1, characterized in that, In step S2, the processing procedure of the structure self-calibration processing unit is as follows: For a single mode to be self-calibrated, its original feature map is denoted as: F m ∈ R Cm×H×W , m Indicates infrared mode or visible light mode. C m Indicates the first m The number of channels in the original feature map of each modality. H and W These represent the height and width of the feature map, respectively; in F m Internally, local neighborhood feature aggregation is performed to generate the first... m Self-calibration feature map of each modality; First, regarding the first m Original feature maps of each modality F m By performing query mapping, key mapping, and value mapping respectively, self-calibrating query features are obtained. Q m s Self-calibration key features K m s and self-calibration value characteristics V m s : set up p Indicates the current spatial location. q Indicates the current position p The local neighborhood location; make N ( p ) indicates by position p The local spatial neighborhood centered on the center, in q ∈ N ( p Perform feature correlation calculation and local consistency constraint within a local range; The structural self-calibration scoring includes two parallel calculation branches: an intramodal correlation branch and a local consistency branch. The two branches calculate the intramodal local correlation score respectively. S m att ( p, q Local consistency weight M m ( p, q ); The local consistency weights are converted into modulation terms, and intramodal local correlation scores are added to obtain the structural self-calibration score: ; in, η The weighting coefficients of the local consistency modulation term are represented. ε A very small constant set to avoid numerical instability in logarithmic operations; Then, at the current location p local neighborhood N ( p The structural self-calibration score is normalized to obtain the structural self-calibration weight: ; Utilizing structural self-calibration weights to characterize the self-calibration values ​​of the same mode V m s ( q Perform weighted aggregation to obtain the current position. p Intramodal self-calibration aggregation results: ; Finally, the intramodal self-calibration aggregation results are... R m s ( p After output mapping W m so (•), and with the original features F m ( p Perform residual join to obtain the first... m The modality at position p Self-calibration features at the location: F m s ( p ) = F m ( p ) + W m so ( R m s ( p )); in, W m so (•) represents the self-calibration output mapping function.

5. The infrared-visible image fusion processing method based on cross-modal compensation and structural self-calibration according to claim 4, characterized in that, In the intramodal correlation branch, the current position is... p Self-calibration query features Q m s ( p ) and neighborhood location q Self-calibration bond features K m s ( q Similarity calculations are performed to obtain intramodal local correlation scores: ; in, d The channel dimension represents the self-calibration query feature and the self-calibration key feature; In the locally consistent branch, based on the original feature map F m Current location p response characteristics F m ( p and neighborhood location q response characteristics F m ( q Calculate the local consistency weight: ; Where, ||•||1 represents L 1-norm, β This represents the local difference sensitivity coefficient.

6. The infrared-visible image fusion processing method based on cross-modal compensation and structural self-calibration according to claim 1, characterized in that, In step S3, the processing procedure of the feature fusion unit is as follows: Align feature maps from different sources by channel and spatial dimensions; The aligned original feature map, compensated feature map, and self-calibrated feature map are concatenated along the channel dimension to obtain the concatenated multi-source feature map. F cat ; Using convolutional fusion mapping pairs F cat Feature integration is performed to obtain an enhanced fused feature map. F fuse : F fuse = φ ( W f ( F cat )); in, W f (•) represents a convolutional fusion mapping; φ (•) represents a non-linear activation function.

7. The infrared-visible image fusion processing method based on cross-modal compensation and structural self-calibration according to claim 1, characterized in that, In step S4, the decoder reconstructs the enhanced fusion feature map into a fused image as follows: Enhance the fusion feature map F fuse As the initial input to the decoder, let it be denoted as R 0= F fuse ; In the l In the multi-level decoding process, the features reconstructed from the previous layer are first... R l-1 Upsampling is performed to obtain reconstructed features with improved spatial resolution; Let the dual-stream feature extraction unit in step S1 be in the... l The infrared shallow layer features output at the corresponding scale are S ir l The shallow visible light characteristics are S vis l The two are then fused and mapped to obtain the first... l Shallow structural features: S l = F shallow l ( Concat ( S ir l , S vis l )); Where, Φ shallow l (•) indicates shallow structure fusion mapping; In the shallow structure feature S l Before adding it to the upsampled reconstructed features, ψ is first mapped by channel and scale alignment. l (•) for S l Process it to make it consistent with U l ( R l-1 They have the same spatial dimensions and number of channels; l The level decoding process is represented as: ; ; in, U l (•) indicates the first l Level upsampling operation; ψ l (•) represents the mapping function used for channel and scale alignment; This represents the intermediate reconstruction features after fusing shallow structural information; C l (•) indicates the first l Level convolution reconstruction operation; R l Indicates the first l Reconstructed features of the level output; go through L After the levels are gradually restored, the final reconstructed features are obtained. R L Subsequently, a fused image is generated through output convolution and pixel range normalization mapping: I fuse = σ ( W out ( R L )); in, W out (•) indicates the output convolution mapping; σ (•) represents the pixel range normalization function.

8. The infrared-visible image fusion processing method based on cross-modal compensation and structural self-calibration according to any one of claims 1-7, characterized in that, The decoder includes: Upsampling units are used to progressively restore the spatial resolution of the feature map; Convolutional reconstruction units are used to reconstruct local edges, textures, and region structures; Shallow structural connections are used to supplement the low-level spatial information retained during the feature extraction stage in the spatial resolution restoration process; The output mapping unit is used to generate the final fused image.

Citation Information

Patent Citations

  • High-precision image processing method and system based on illumination adaptive compensation

    CN121032846A

  • Low-light image enhancement method and system based on multi-mode cooperation, storage medium and equipment

    CN121481869A