A multi-modal image fusion method based on structural causal model and texture compensation mechanism

CN122312399BActive Publication Date: 2026-08-21NAVAL AVIATION UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610758379.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-29
Publication Date
2026-08-21
Estimated Expiration
2046-05-29

AI Technical Summary

Technical Problem

[0005]本发明的目的在于,针对上述现有技术存在的未对成像物理机制建模,导致的源图像中跨模态干扰伪影无法有效剔除缺陷,以及局限于统计特征与竞争性融合规则,导致的传统融合策略存在信息吞噬的缺陷,提供设计一种基于结构因果模型与纹理补偿机制的多模态图像融合方法及系统,以解决上述技术问题

Benefits of technology

[0078]本发明的有益效果在于,能够有效剔除跨模态干扰伪影,提升融合图像的真实性与可靠性:通过采用结构因果模型,从物理成像机制层面建模并分离不同模态信号的潜在因果结构,不仅关注像素层面的统计特征,更通过因果推理区分真实信号与跨模态干扰(如红外图像中由可见光反射造成的虚假高亮区域),从而在融合前识别并抑制伪影,避免其被放大进入融合结果,降低虚警率,增强图像在后续识别、检测等任务中的可信度;能够避免信息吞噬,实现双模态信息的协同增强与互补融合:针对传统加权融合或竞争性规则导致的信息丢失问题,本发明提出纹理补偿机制,在保留目标区域显著特征(如红外热目标)的同时,通过纹理提取与自适应补偿,将另一模态(如可见光)中对应的关键细节纹理(如车牌、图案、边缘结构)有效注入融合结果中。克服了单一模态主导导致的纹理信息被压制的问题,实现了真实目标信息与细节纹理信息的协同保留与叠加,提高了信息利用率与融合图像的完整性;能够提升融合图像的视觉质量与任务适应性,并且具有良好的可解释性与扩展性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122312399B_ABST
    Figure CN122312399B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of multi-modal image fusion, and relates to a multi-modal image fusion method based on a structural causal model and a texture compensation mechanism, which comprises the following steps: step S1, a step of mixed factor extraction; step S2, a step of local causal effect coefficient estimation; step S3, a step of counterfactual intervention denoising; step S4, a step of thermal source probability weight nonlinear mapping; and step S5, a step of image reconstruction and fusion. The application models physical imaging by constructing a structural causal model, and actively separates and removes cross-modal artifacts by using causal intervention. In the fusion process, the suppressed detail information is recovered and enhanced by using a texture compensation mechanism, the multi-modal information is fully reserved and complementarily fused while the interference is suppressed, and the target prominence and visual quality of the fused image are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of multimodal image fusion technology, specifically relating to a multimodal image fusion method and system based on a structural causal model and texture compensation mechanism. Background Technology

[0002] Multimodal image fusion aims to integrate the advantages of different sensors. Existing fusion methods have evolved from pixel-level weighting and multi-scale transformation to sparse representation and deep learning, but bottlenecks still exist. First, cross-modal interference artifacts in the source image cannot be effectively removed. In the physical mechanism of multimodal imaging, signals from different modalities are coupled. For example, infrared images are often affected by visible light texture reflections, forming false bright areas and causing false alarms in the fused image. Existing methods only focus on pixel statistical features (numerical magnitude or gradient strength), lacking the ability to physically discriminate the signal source, and cannot distinguish between real signals and interference artifacts, resulting in the amplification of artifacts.

[0003] Secondly, traditional fusion strategies suffer from information attrition. Some weighted fusion methods use normalized weight maps; once a region is identified as a target and assigned a high weight, the corresponding information in the other image is suppressed, leading to the loss of key textures on the target surface (such as license plates and clothing patterns). Other methods employ competitive fusion rules, which cannot superimpose effective bimodal information on the same pixel, resulting in low information utilization.

[0004] In view of this, it is very necessary to provide a multimodal image fusion method and system based on structural causal model and texture compensation mechanism to solve the above-mentioned defects in the prior art. Summary of the Invention

[0005] The purpose of this invention is to address the shortcomings of existing technologies, such as the inability to effectively remove cross-modal interference artifacts in source images due to the lack of modeling of the imaging physical mechanism, and the information-consuming nature of traditional fusion strategies due to limitations in statistical features and competitive fusion rules. This invention provides a multimodal image fusion method and system based on a structural causal model and texture compensation mechanism to solve the aforementioned technical problems.

[0006] To achieve the above objectives, the present invention provides the following technical solution: A multimodal image fusion method based on a structural causal model and texture compensation mechanism includes the following steps: Step S1, the step of extracting confounding factors, in which: Obtain the first source image and the second source image; Guided filtering is applied to the second source image to obtain the base layer of the second source image; The second source image is subjected to a difference operation with its base layer to obtain a detail layer of the second source image, which is then used as a confounding factor. Step S2, the step of estimating the local causal effect coefficient, in which: A local linear causal model is established between the first source image and the confounding factor, and the interference coefficient of the local linear causal model is estimated. Step S3, the counterfactual intervention denoising step, in which: Counterfactual intervention is performed based on the interference coefficients of the local linear causal model to remove cross-modal interference from the first source image and obtain the denoised physical signal. Step S4, the nonlinear mapping step of heat source probability weights, in which: Generate a target probability weight map based on the denoised physical signal; Step S5, the image reconstruction and fusion step, in which: Based on the target probability weight map, a method combining weighted base layer and incremental compensation of detail layer is used to reconstruct and generate a fused image.

[0007] Preferably, step S1 specifically includes: Before establishing a local linear causal model, the interference sources causing cross-modal artifacts are precisely defined and extracted. High-frequency textures in the second source image (such as a visible light image) are defined as confounding factors, and a cross-guided filtering technique is used to separate these high-frequency textures from the second source image. This technique simultaneously ensures that the overall skeletal structure of the image is not destroyed. The first and second source images are geometrically registered and normalized to the double-precision floating-point domain. Preprocessing yields the first source image matrix and the second source image matrix.

[0008] A guided filter is applied to the second source image matrix to obtain the base layer of the second source image. To ensure that the edge structure of the base layer is consistent with that of the first source image, the first source image matrix is ​​used as the guiding image to apply guided filtering to the second source image matrix; this involves solving the following local linear model optimization problem:

[0009] in, The output image of the guide filter, For guiding purposes, Represents pixels, Represented in pixels A local window centered on the center. Indicates the position (or index) of the pixel within the current local window. and These are the linear coefficients within a local window.

[0010] Solving using the Ridge Regression cost function and The optimal value:

[0011] in, It is a regularization parameter used to prevent overfitting. Represents the second source image matrix At pixel position The value at that location, Let be the ridge regression cost function. The value is obtained by minimizing the ridge regression cost function. and Value:

[0012]

[0013] in, For guiding diagrams In local window The mean within, For guiding diagrams In local window within variance, For the second source image matrix In local window The mean within, For local windows The total number of pixels contained therein.

[0014] Obtain linear coefficients and After determining the value, the guided filter averages the output images of all local windows covering the same pixel to obtain the base layer of the second source image. .

[0015] Mathematical difference operations are used to extract the detail layer of the second source image, and the detail layer of the second source image is defined as the confounding factor. :

[0016] in, and These are the coordinate components of a pixel in the image. The confounding factor is a floating-point matrix containing positive and negative values. It accurately carries the high-frequency detail information in the second source image and is the direct cause variable for artifacts in the first source image.

[0017] Preferably, step S2 specifically includes: After extracting the confounding factors, a local linear causal model is constructed between the first source image matrix and the confounding factors. Based on the physical imaging laws, this local linear causal model decomposes the observed signal of the first source image matrix into real physical signals. and interference caused by texture The expression is:

[0018] in, It is a spatially variable interference coefficient, representing the interference at the pixel level. The magnitude of the interference intensity (causal effect) of the texture of the second source image on the intensity of the first source image. Additive noise, Indicated in pixels The actual physical signal observed from the first source image matrix. Indicated in pixels The interference term of the observed signal in the first source image matrix.

[0019] The interference coefficient is solved using statistical properties within a local window of the image. Set a local window Assuming the interference coefficient is constant, according to the least squares principle, the optimal interference coefficient estimate should make the actual physical signal... The objective function of the locally linear causal model, which minimizes the local variance, is:

[0020] in, The first source image matrix, This is the intercept of the local linear causal model.

[0021] right Find the partial derivative and set it to zero. Through mathematical derivation, The optimal solution is:

[0022] in, The local covariance between the first source image matrix and the confounding factor is represented by the following method: , The local variance of the confounding factor is calculated as follows: , Indicates in a local window Mean filtering operation within, The minimal regularization constant (e.g.) is set to prevent the denominator from being zero. ).

[0023] The physical meaning of the above formula for calculating the interference coefficient is: if the trend of the intensity change of the first source image in the current local window is positively correlated with the trend of the texture change of the second source image, then the calculated interference coefficient... A larger value indicates that the observed signal in the current local window is mostly composed of interference artifacts; conversely, if the intensity variation trend of the first source image is unrelated to the texture variation trend of the second source image in the current local window, then the calculated value... Approaching zero indicates that the observed signal in the current local window is an independent physical signal.

[0024] All The values ​​are arranged according to their corresponding pixel positions to form a causality coefficient graph of the same size as the source image, denoted as . Figure. To ensure... To improve the smoothness of the graph and reduce estimation noise, guided filtering can also be used to refine the calculated original graph. The image was post-processed and optimized.

[0025] Preferably, step S3 specifically includes: After obtaining the causal coefficient graph, counterfactual intervention is performed based on the Do operator theory in causal inference: Do operator The ideal state after physically removing environmental texture interference is mathematically simulated. Based on a local linear causal model, the estimated interference term is subtracted from the first source image to achieve counterfactual intervention:

[0026] in, This is a counterfactual residual image.

[0027] At the pixel location of image artifacts such as ground reflection and anatomical structure interference, the interference coefficient is... Large value and confounding factor Significantly, after counterfactual intervention, the interference items... This will be precisely subtracted, thus eliminating false highlight features or structural interference. (In the interference coefficient...) Pixel positions with values ​​approaching zero are affected by interference coefficients. The signal is very small, and the observed signal mainly originates from independent physical signal components. Through counterfactual intervention, the original observed signal strength was preserved.

[0028] Through counterfactual intervention, the image is transformed from an "observational state" containing confounding variables into a "counterfactual state" free from interference, thus obtaining a pure physical signal.

[0029] Preferably, step S4 specifically includes: A nonlinear mapping method is used to transform the denoised physical signal into a weight map to guide fusion: Counterfactual residual image Perform linear normalization to map its numerical range to The interval is used to obtain the normalized residual image. :

[0030] Target probability weight map generated using the Sigmoid activation function :

[0031] in, The slope of the activation function (e.g., a value of 15) determines the sharpness of the transition between the foreground and background in the target probability weight map; The threshold value for the activation function (e.g., 0.05) determines the signal strength at which a signal is considered a target (an object with specific physical properties that is of interest to the user or downstream tasks). The calculated... The range of values ​​is Indicates the pixel position The probability that the image content belongs to the target.

[0032] The target probability weight map not only eliminates regions that are identified as artifacts by the local linear causal model (these regions have extremely low weights), but also highlights the target.

[0033] Preferably, step S5 specifically includes: Step S51: Construct the base layer of the fused image Using the calculated target probability weight map Linear weighting is applied to the first and second source images to determine the overall brightness and contrast tone of the fused image:

[0034] in, This represents the base layer of the fused image.

[0035] The above formula ensures that... For target regions approaching 1, the fused image exhibits high-intensity features from the first source image; Background areas approaching 0 are merged to present the natural background features of the second source image.

[0036] Step S52: Construct the detail layer of the fused image In traditional methods, if the weight of the target region is too high, it can suppress information from the second source image in the fused image. To recover these lost details, this invention employs an incremental compensation method to reduce the confounding factor. (Second source image texture) is re-injected into the high-weight target region, expressed as:

[0037] in, To blend the detail layers of the image, This is the compensation gain coefficient.

[0038] The physical meaning of the above formula is that even within a defined target area, the surface texture is still considered valuable, and this is applied incrementally. The ratio is superimposed on the target brightness.

[0039] Step S53: Generate fused image The base and detail layers of an image are linearly superimposed and blended to generate the final blended image. :

[0040] To prevent pixel value overflow and optimize visual effects, the fusion result is numerically clamped (limited to...). Within the specified range, adaptive histogram equalization (CLAHE) is used for post-processing. CLAHE divides the image into several non-overlapping blocks and limits the height of the histogram of each block (i.e., limits the magnitude of contrast enhancement), thereby enhancing the local contrast of the image while avoiding excessive amplification of noise.

[0041] The resulting fused image removes cross-modal interference artifacts, retains the target's physical intensity information, and clearly presents the texture details of the target surface.

[0042] Furthermore, this invention also provides a multimodal image fusion system based on a structural causal model and a texture compensation mechanism, comprising: The confounding factor extraction module contains: Obtain the first source image and the second source image; Guided filtering is applied to the second source image to obtain the base layer of the second source image; The second source image is subjected to a difference operation with its base layer to obtain a detail layer of the second source image, which is then used as a confounding factor. The module for estimating the local causal effect coefficients contains: A local linear causal model is established between the first source image and the confounding factor, and the interference coefficient of the local linear causal model is estimated. The counterfactual intervention noise reduction module contains: Counterfactual intervention is performed based on the interference coefficients of the local linear causal model to remove cross-modal interference from the first source image and obtain the denoised physical signal. The heat source probability weight nonlinear mapping module contains: Generate a target probability weight map based on the denoised physical signal; The image reconstruction and fusion module contains: Based on the target probability weight map, a method combining weighted base layer and incremental compensation of detail layer is used to reconstruct and generate a fused image.

[0043] Preferably, the confounding factor extraction module specifically includes: Before establishing a local linear causal model, the interference sources causing cross-modal artifacts are precisely defined and extracted. High-frequency textures in the second source image (such as a visible light image) are defined as confounding factors, and a cross-guided filtering technique is used to separate these high-frequency textures from the second source image. This technique simultaneously ensures that the overall skeletal structure of the image is not destroyed. The first and second source images are geometrically registered and normalized to the double-precision floating-point domain. Preprocessing yields the first source image matrix and the second source image matrix.

[0044] A guided filter is applied to the second source image matrix to obtain the base layer of the second source image. To ensure that the edge structure of the base layer is consistent with that of the first source image, the first source image matrix is ​​used as the guiding image to apply guided filtering to the second source image matrix; this involves solving the following local linear model optimization problem:

[0045] in, The output image of the guide filter, For guiding purposes, Represents pixels, Represented in pixels A local window centered on the center. Indicates the position (or index) of the pixel within the current local window. and These are the linear coefficients within a local window.

[0046] Solving using the Ridge Regression cost function and The optimal value:

[0047] in, It is a regularization parameter used to prevent overfitting. Represents the second source image matrix At pixel position The value at that location, Let be the ridge regression cost function. The value is obtained by minimizing the ridge regression cost function. and Value:

[0048]

[0049] in, For guiding diagrams In local window The mean within, For guiding diagrams In local window within variance, For the second source image matrix In local window The mean within, For local windows The total number of pixels contained therein.

[0050] Obtain linear coefficients and After determining the value, the guided filter averages the output images of all local windows covering the same pixel to obtain the base layer of the second source image. .

[0051] Mathematical difference operations are used to extract the detail layer of the second source image, and the detail layer of the second source image is defined as the confounding factor. :

[0052] in, and These are the coordinate components of a pixel in the image. The confounding factor is a floating-point matrix containing positive and negative values. It accurately carries the high-frequency detail information in the second source image and is the direct cause variable for artifacts in the first source image.

[0053] Preferably, the local causal effect coefficient estimation module specifically includes: After extracting the confounding factors, a local linear causal model is constructed between the first source image matrix and the confounding factors. Based on the physical imaging laws, this local linear causal model decomposes the observed signal of the first source image matrix into real physical signals. and interference caused by texture The expression is:

[0054] in, It is a spatially variable interference coefficient, representing the interference at the pixel level. The magnitude of the interference intensity (causal effect) of the texture of the second source image on the intensity of the first source image. Additive noise, Indicated in pixels The actual physical signal observed from the first source image matrix. Indicated in pixels The interference term of the observed signal in the first source image matrix.

[0055] The interference coefficient is solved using statistical properties within a local window of the image. Set a local window Assuming the interference coefficient is constant, according to the least squares principle, the optimal interference coefficient estimate should make the actual physical signal... The objective function of the locally linear causal model, which minimizes the local variance, is:

[0056] in, The first source image matrix, This is the intercept of the local linear causal model.

[0057] right Find the partial derivative and set it to zero. Through mathematical derivation, The optimal solution is:

[0058] in, The local covariance between the first source image matrix and the confounding factor is represented by the following method: , The local variance of the confounding factor is calculated as follows: , Indicates in a local window Mean filtering operation within, The minimal regularization constant (e.g.) is set to prevent the denominator from being zero. ).

[0059] The physical meaning of the above formula for calculating the interference coefficient is: if the trend of the intensity change of the first source image in the current local window is positively correlated with the trend of the texture change of the second source image, then the calculated interference coefficient... A larger value indicates that the observed signal in the current local window is mostly composed of interference artifacts; conversely, if the intensity variation trend of the first source image is unrelated to the texture variation trend of the second source image in the current local window, then the calculated value... Approaching zero indicates that the observed signal in the current local window is an independent physical signal.

[0060] All The values ​​are arranged according to their corresponding pixel positions to form a causality coefficient graph of the same size as the source image, denoted as . Figure. To ensure... To improve the smoothness of the graph and reduce estimation noise, guided filtering can also be used to refine the calculated original graph. The image was post-processed and optimized.

[0061] Preferably, the counterfactual intervention noise reduction module specifically includes: After obtaining the causal coefficient graph, counterfactual intervention is performed based on the Do operator theory in causal inference: Do operator The ideal state after physically removing environmental texture interference is mathematically simulated. Based on a local linear causal model, the estimated interference term is subtracted from the first source image to achieve counterfactual intervention:

[0062] in, This is a counterfactual residual image.

[0063] At the pixel location of image artifacts such as ground reflection and anatomical structure interference, the interference coefficient is... Large value and confounding factor Significantly, after counterfactual intervention, the interference items... This will be precisely subtracted, thus eliminating false highlight features or structural interference. (In the interference coefficient...) Pixel positions with values ​​approaching zero are affected by interference coefficients. The signal is very small, and the observed signal mainly originates from independent physical signal components. Through counterfactual intervention, the original observed signal strength was preserved.

[0064] Through counterfactual intervention, the image is transformed from an "observational state" containing confounding variables into a "counterfactual state" free from interference, thus obtaining a pure physical signal.

[0065] Preferably, the heat source probability weight nonlinear mapping module specifically includes: A nonlinear mapping method is used to transform the denoised physical signal into a weight map to guide fusion: Counterfactual residual image Perform linear normalization to map its numerical range to The interval is used to obtain the normalized residual image. :

[0066] Target probability weight map generated using the Sigmoid activation function :

[0067] in, The slope of the activation function (e.g., a value of 15) determines the sharpness of the transition between the foreground and background in the target probability weight map; The threshold value for the activation function (e.g., 0.05) determines the signal strength at which a signal is considered a target (an object with specific physical properties that is of interest to the user or downstream tasks). The calculated... The range of values ​​is Indicates the pixel position The probability that the image content belongs to the target.

[0068] The target probability weight map not only eliminates regions that are identified as artifacts by the local linear causal model (these regions have extremely low weights), but also highlights the target.

[0069] Preferably, the image reconstruction and fusion module specifically includes: a sub-module for constructing the base layer of the fused image, a sub-module for constructing the detail layer of the fused image, and a sub-module for generating the fused image; The aforementioned sub-module for constructing the fused image base layer includes: Using the calculated target probability weight map Linear weighting is applied to the first and second source images to determine the overall brightness and contrast tone of the fused image:

[0070] in, This represents the base layer of the fused image.

[0071] The above formula ensures that... For target regions approaching 1, the fused image exhibits high-intensity features from the first source image; Background areas approaching 0 are merged to present the natural background features of the second source image.

[0072] The aforementioned submodule for constructing the fused image detail layer includes: In traditional methods, if the weight of the target region is too high, it can suppress information from the second source image in the fused image. To recover these lost details, this invention employs an incremental compensation method to reduce the confounding factor. (Second source image texture) is re-injected into the high-weight target region, expressed as:

[0073] in, To blend the detail layers of the image, This is the compensation gain coefficient.

[0074] The physical meaning of the above formula is that even within a defined target area, the surface texture is still considered valuable, and this is applied incrementally. The ratio is superimposed on the target brightness.

[0075] The aforementioned submodule for generating fused images includes: The base and detail layers of an image are linearly superimposed and blended to generate the final blended image. :

[0076] To prevent pixel value overflow and optimize visual effects, the fusion result is numerically clamped (limited to...). Within the specified range, adaptive histogram equalization (CLAHE) is used for post-processing. CLAHE divides the image into several non-overlapping blocks and limits the height of the histogram of each block (i.e., limits the magnitude of contrast enhancement), thereby enhancing the local contrast of the image while avoiding excessive amplification of noise.

[0077] The resulting fused image removes cross-modal interference artifacts, retains the target's physical intensity information, and clearly presents the texture details of the target surface.

[0078] The beneficial effects of this invention are that it can effectively eliminate cross-modal interference artifacts and improve the realism and reliability of fused images: by adopting a structural causal model, it models and separates the potential causal structure of different modal signals from the perspective of physical imaging mechanism. It not only focuses on the statistical features at the pixel level, but also distinguishes real signals from cross-modal interference (such as false bright areas caused by visible light reflection in infrared images) through causal reasoning. This allows for the identification and suppression of artifacts before fusion, preventing them from being amplified into the fusion result, reducing the false alarm rate, and enhancing the credibility of the image in subsequent recognition, detection, and other tasks. It can also avoid information swallowing and achieve synergistic enhancement and complementary fusion of dual-modal information: to address the information loss problem caused by traditional weighted fusion or competitive rules, this invention proposes a texture compensation mechanism. While retaining the significant features of the target area (such as infrared thermal targets), it effectively injects the corresponding key detail textures (such as license plates, patterns, and edge structures) from another modality (such as visible light) into the fusion result through texture extraction and adaptive compensation. It overcomes the problem of suppressed texture information caused by single-modality dominance, and achieves the collaborative preservation and superposition of real target information and detailed texture information, improving information utilization and the integrity of fused images; it can improve the visual quality and task adaptability of fused images, and has good interpretability and scalability.

[0079] Therefore, it is evident that the present invention has outstanding substantive features and significant progress compared with the prior art, and the beneficial effects of its implementation are also obvious. Attached Figure Description

[0080] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0081] Figure 1 This is a flowchart of a multimodal image fusion method based on a structural causal model and texture compensation mechanism provided by the present invention.

[0082] Figure 2 This is a block diagram of the principle of a multimodal image fusion system based on a structural causal model and texture compensation mechanism provided by the present invention.

[0083] Among them, 1-confounding factor extraction module, 2-local causal effect coefficient estimation module, 3-counterfactual intervention denoising module, 4-heat source probability weight nonlinear mapping module, and 5-image reconstruction and fusion module. Detailed Implementation

[0084] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. The following embodiments are explanations of the present invention, but the present invention is not limited to the following implementation methods.

[0085] Example 1: like Figure 1 As shown in the figure, this embodiment provides a multimodal image fusion method based on a structural causal model and a texture compensation mechanism, which includes the following steps: Step S1, the step of extracting confounding factors, in which: Infrared thermal imaging is used as the first source image, and visible light image is used as the second source image; Guided filtering is applied to the second source image to obtain the base layer of the second source image; The second source image is subjected to a difference operation with its base layer to obtain a detail layer of the second source image, which is then used as a confounding factor. Step S1 specifically includes: Before establishing a local linear causal model, the interference sources causing cross-modal artifacts are precisely defined and extracted. High-frequency textures in the second source image are defined as confounding factors, and a cross-guided filtering technique is used to separate these high-frequency textures from the second source image. This technique simultaneously ensures that the overall skeleton structure of the image is not destroyed. The first and second source images are geometrically registered and normalized to the double-precision floating-point domain. Preprocessing yields the first source image matrix and the second source image matrix.

[0086] A guided filter is applied to the second source image matrix to obtain the base layer of the second source image. To ensure that the edge structure of the base layer is consistent with that of the first source image, the first source image matrix is ​​used as the guiding image to apply guided filtering to the second source image matrix; this involves solving the following local linear model optimization problem:

[0087] in, The output image of the guide filter, For guiding purposes, Represents pixels, Represented in pixels A local window centered on the center. Indicates the position (or index) of the pixel within the current local window. and These are the linear coefficients within a local window.

[0088] Solving using the Ridge Regression cost function and The optimal value:

[0089] in, It is a regularization parameter used to prevent overfitting. Represents the second source image matrix At pixel position The value at that location, Let be the ridge regression cost function. The value is obtained by minimizing the ridge regression cost function. and Value:

[0090]

[0091] in, For guiding diagrams In local window The mean within, For guiding diagrams In local window within variance, For the second source image matrix In local window The mean within, For local windows The total number of pixels contained therein.

[0092] Obtain linear coefficients and After determining the value, the guided filter averages the output images of all local windows covering the same pixel to obtain the base layer of the second source image. .

[0093] Mathematical difference operations are used to extract the detail layer of the second source image, and the detail layer of the second source image is defined as the confounding factor. :

[0094] in, and These are the coordinate components of a pixel in the image. The confounding factor is a floating-point matrix containing positive and negative values. It accurately carries the high-frequency detail information in the second source image and is the direct cause variable for artifacts in the first source image.

[0095] Step S2, the step of estimating the local causal effect coefficient, in which: A local linear causal model is established between the first source image and the confounding factor, and the interference coefficient of the local linear causal model is estimated. Step S2 specifically includes: After extracting the confounding factors, a local linear causal model is constructed between the first source image matrix and the confounding factors. Based on the physical imaging laws, this local linear causal model decomposes the observed signal of the first source image matrix into real physical signals. and interference caused by texture The expression is:

[0096] in, It is a spatially variable interference coefficient, representing the interference at the pixel level. The magnitude of the interference intensity (causal effect) of the texture of the second source image on the intensity of the first source image. Additive noise, Indicated in pixels The actual physical signal observed from the first source image matrix. Indicated in pixels The interference term of the observed signal in the first source image matrix.

[0097] The interference coefficient is solved using statistical properties within a local window of the image. Set a local window Assuming the interference coefficient is constant, according to the least squares principle, the optimal interference coefficient estimate should make the actual physical signal... The objective function of the locally linear causal model, which minimizes the local variance, is:

[0098] in, The first source image matrix, This is the intercept of the local linear causal model.

[0099] right Find the partial derivative and set it to zero. Through mathematical derivation, The optimal solution is:

[0100] in, The local covariance between the first source image matrix and the confounding factor is represented by the following method: , The local variance of the confounding factor is calculated as follows: , Indicates in a local window Mean filtering operation within, The minimal regularization constant (e.g.) is set to prevent the denominator from being zero. ).

[0101] The physical meaning of the above formula for calculating the interference coefficient is: if the trend of the intensity change of the first source image in the current local window is positively correlated with the trend of the texture change of the second source image, then the calculated interference coefficient... A larger value indicates that the observed signal in the current local window is mostly composed of interference artifacts; conversely, if the intensity variation trend of the first source image is unrelated to the texture variation trend of the second source image in the current local window, then the calculated value... Approaching zero indicates that the observed signal in the current local window is an independent physical signal.

[0102] All The values ​​are arranged according to their corresponding pixel positions to form a causality coefficient graph of the same size as the source image, denoted as . Figure. To ensure... To improve the smoothness of the graph and reduce estimation noise, guided filtering can also be used to refine the calculated original graph. The image was post-processed and optimized.

[0103] Step S3, the counterfactual intervention denoising step, in which: Counterfactual intervention is performed based on the interference coefficients of the local linear causal model to remove cross-modal interference from the first source image and obtain the denoised physical signal. Step S3 specifically includes: After obtaining the causal coefficient graph, counterfactual intervention is performed based on the Do operator theory in causal inference: Do operator The ideal state after physically removing environmental texture interference is mathematically simulated. Based on a local linear causal model, the estimated interference term is subtracted from the first source image to achieve counterfactual intervention:

[0104] in, This is a counterfactual residual image.

[0105] At the pixel location of image artifacts such as ground reflection and anatomical structure interference, the interference coefficient is... Large value and confounding factor Significantly, after counterfactual intervention, the interference items... This will be precisely subtracted, thus eliminating false highlight features or structural interference. (In the interference coefficient...) Pixel positions with values ​​approaching zero are affected by interference coefficients. The signal is very small, and the observed signal mainly originates from independent physical signal components. Through counterfactual intervention, the original observed signal strength was preserved.

[0106] Through counterfactual intervention, the image is transformed from an "observational state" containing confounding variables into a "counterfactual state" free from interference, thus obtaining a pure physical signal.

[0107] Step S4, the nonlinear mapping step of heat source probability weights, in which: Generate a target probability weight map based on the denoised physical signal; Step S4 specifically includes: A nonlinear mapping method is used to transform the denoised physical signal into a weight map to guide fusion: Counterfactual residual image Perform linear normalization to map its numerical range to The interval is used to obtain the normalized residual image. :

[0108] Target probability weight map generated using the Sigmoid activation function :

[0109] in, The slope of the activation function (e.g., a value of 15) determines the sharpness of the transition between the foreground and background in the target probability weight map; The threshold value for the activation function (e.g., 0.05) determines the signal strength at which a signal is considered a target (an object with specific physical properties that is of interest to the user or downstream tasks). The calculated... The range of values ​​is Indicates the pixel position The probability that the image content belongs to the target.

[0110] The target probability weight map not only eliminates regions that are identified as artifacts by the local linear causal model (these regions have extremely low weights), but also highlights the target.

[0111] Step S5, the image reconstruction and fusion step, in which: Based on the target probability weight map, a method combining weighted base layer and incremental compensation of detail layer is used to reconstruct and generate a fused image.

[0112] Step S5 specifically includes: Step S51: Construct the base layer of the fused image Using the calculated target probability weight map Linear weighting is applied to the first and second source images to determine the overall brightness and contrast tone of the fused image:

[0113] in, This represents the base layer of the fused image.

[0114] The above formula ensures that... For target regions approaching 1, the fused image exhibits high-intensity features from the first source image; Background areas approaching 0 are merged to present the natural background features of the second source image.

[0115] Step S52: Construct the detail layer of the fused image In traditional methods, if the weight of the target region is too high, it can suppress information from the second source image in the fused image. To recover these lost details, this invention employs an incremental compensation method to reduce the confounding factor. (Second source image texture) is re-injected into the high-weight target region, expressed as:

[0116] in, To blend the detail layers of the image, In this embodiment, the value of the compensation gain coefficient is 0.6.

[0117] Step S53: Generate fused image The base and detail layers of an image are linearly superimposed and blended to generate the final blended image. :

[0118] To prevent pixel value overflow and optimize visual effects, the fusion result is numerically clamped (limited to...). Within the specified range, adaptive histogram equalization (CLAHE) is used for post-processing. CLAHE divides the image into several non-overlapping blocks and limits the height of the histogram of each block (i.e., limits the magnitude of contrast enhancement), thereby enhancing the local contrast of the image while avoiding excessive amplification of noise.

[0119] The resulting fused image removes cross-modal interference artifacts, retains the target's physical intensity information, and clearly presents the texture details of the target surface.

[0120] Example 2: like Figure 2 As shown in the figure, this embodiment provides a multimodal image fusion system based on a structural causal model and a texture compensation mechanism, comprising: Confounding factor extraction module 1, in which: Obtain the first source image and the second source image; Guided filtering is applied to the second source image to obtain the base layer of the second source image; The second source image is subjected to a difference operation with its base layer to obtain a detail layer of the second source image, which is then used as a confounding factor. The aforementioned confounding factor extraction module 1 specifically includes: Before establishing a local linear causal model, the interference sources causing cross-modal artifacts are precisely defined and extracted. High-frequency textures in the second source image (such as a visible light image) are defined as confounding factors, and a cross-guided filtering technique is used to separate these high-frequency textures from the second source image. This technique simultaneously ensures that the overall skeletal structure of the image is not destroyed. The first and second source images are geometrically registered and normalized to the double-precision floating-point domain. Preprocessing yields the first source image matrix and the second source image matrix.

[0121] A guided filter is applied to the second source image matrix to obtain the base layer of the second source image. To ensure that the edge structure of the base layer is consistent with that of the first source image, the first source image matrix is ​​used as the guiding image to apply guided filtering to the second source image matrix; this involves solving the following local linear model optimization problem:

[0122] in, The output image of the guide filter, For guiding purposes, Represents pixels, Represented in pixels A local window centered on the center. Indicates the position (or index) of the pixel within the current local window. and These are the linear coefficients within a local window.

[0123] Solving using the Ridge Regression cost function and The optimal value:

[0124] in, It is a regularization parameter used to prevent overfitting. Represents the second source image matrix At pixel position The value at that location, Let be the ridge regression cost function. The value is obtained by minimizing the ridge regression cost function. and Value:

[0125]

[0126] in, For guiding diagrams In local window The mean within, For guiding diagrams In local window within variance, For the second source image matrix In local window The mean within, For local windows The total number of pixels contained therein.

[0127] Obtain linear coefficients and After determining the value, the guided filter averages the output images of all local windows covering the same pixel to obtain the base layer of the second source image. .

[0128] Mathematical difference operations are used to extract the detail layer of the second source image, and the detail layer of the second source image is defined as the confounding factor. :

[0129] in, and These are the coordinate components of a pixel in the image. The confounding factor is a floating-point matrix containing positive and negative values. It accurately carries the high-frequency detail information in the second source image and is the direct cause variable for artifacts in the first source image.

[0130] Module 2 for estimating local causal effect coefficients, in which: A local linear causal model is established between the first source image and the confounding factor, and the interference coefficient of the local linear causal model is estimated. The local causal effect coefficient estimation module 2 specifically includes: After extracting the confounding factors, a local linear causal model is constructed between the first source image matrix and the confounding factors. Based on the physical imaging laws, this local linear causal model decomposes the observed signal of the first source image matrix into real physical signals. and interference caused by texture The expression is:

[0131] in, It is a spatially variable interference coefficient, representing the interference at the pixel level. The magnitude of the interference intensity (causal effect) of the texture of the second source image on the intensity of the first source image. Additive noise, Indicated in pixels The actual physical signal observed from the first source image matrix. Indicated in pixels The interference term of the observed signal in the first source image matrix.

[0132] The interference coefficient is solved using statistical properties within a local window of the image. Set a local window Assuming the interference coefficient is constant, according to the least squares principle, the optimal interference coefficient estimate should make the actual physical signal... The objective function of the locally linear causal model, which minimizes the local variance, is:

[0133] in, The first source image matrix, This is the intercept of the local linear causal model.

[0134] right Find the partial derivative and set it to zero. Through mathematical derivation, The optimal solution is:

[0135] in, The local covariance between the first source image matrix and the confounding factor is represented by the following method: , The local variance of the confounding factor is calculated as follows: , Indicates in a local window Mean filtering operation within, The minimal regularization constant (e.g.) is set to prevent the denominator from being zero. ).

[0136] The physical meaning of the above formula for calculating the interference coefficient is: if the trend of the intensity change of the first source image in the current local window is positively correlated with the trend of the texture change of the second source image, then the calculated interference coefficient... A larger value indicates that the observed signal in the current local window is mostly composed of interference artifacts; conversely, if the intensity variation trend of the first source image is unrelated to the texture variation trend of the second source image in the current local window, then the calculated value... Approaching zero indicates that the observed signal in the current local window is an independent physical signal.

[0137] All The values ​​are arranged according to their corresponding pixel positions to form a causality coefficient graph of the same size as the source image, denoted as . Figure. To ensure... To improve the smoothness of the graph and reduce estimation noise, guided filtering can also be used to refine the calculated original graph. The image was post-processed and optimized.

[0138] Counterfactual intervention denoising module 3, in which: Counterfactual intervention is performed based on the interference coefficients of the local linear causal model to remove cross-modal interference from the first source image and obtain the denoised physical signal. The counterfactual intervention denoising module 3 specifically includes: After obtaining the causal coefficient graph, counterfactual intervention is performed based on the Do operator theory in causal inference: Do operator The ideal state after physically removing environmental texture interference is mathematically simulated. Based on a local linear causal model, the estimated interference term is subtracted from the first source image to achieve counterfactual intervention:

[0139] in, This is a counterfactual residual image.

[0140] At the pixel location of image artifacts such as ground reflection and anatomical structure interference, the interference coefficient is... Large value and confounding factor Significantly, after counterfactual intervention, the interference items... This will be precisely subtracted, thus eliminating false highlight features or structural interference. (In the interference coefficient...) Pixel positions with values ​​approaching zero are affected by interference coefficients. The signal is very small, and the observed signal mainly originates from independent physical signal components. Through counterfactual intervention, the original observed signal strength was preserved.

[0141] Through counterfactual intervention, the image is transformed from an "observational state" containing confounding variables into a "counterfactual state" free from interference, thus obtaining a pure physical signal.

[0142] Heat source probability weight nonlinear mapping module 4, in which: Generate a target probability weight map based on the denoised physical signal; The heat source probability weight nonlinear mapping module 4 specifically includes: A nonlinear mapping method is used to transform the denoised physical signal into a weight map to guide fusion: Counterfactual residual image Perform linear normalization to map its numerical range to The interval is used to obtain the normalized residual image. :

[0143] Target probability weight map generated using the Sigmoid activation function :

[0144] in, The slope of the activation function (e.g., a value of 15) determines the sharpness of the transition between the foreground and background in the target probability weight map; The threshold value for the activation function (e.g., 0.05) determines the signal strength at which a signal is considered a target (an object with specific physical properties that is of interest to the user or downstream tasks). The calculated... The range of values ​​is Indicates the pixel position The probability that the image content belongs to the target.

[0145] The target probability weight map not only eliminates regions that are identified as artifacts by the local linear causal model (these regions have extremely low weights), but also highlights the target.

[0146] Image reconstruction and fusion module 5, in which: Based on the target probability weight map, a method combining weighted base layer and incremental compensation of detail layer is used to reconstruct and generate a fused image.

[0147] The image reconstruction and fusion module 5 specifically includes: a sub-module for constructing the base layer of the fused image, a sub-module for constructing the detail layer of the fused image, and a sub-module for generating the fused image; The aforementioned sub-module for constructing the fused image base layer includes: Using the calculated target probability weight map Linear weighting is applied to the first and second source images to determine the overall brightness and contrast tone of the fused image:

[0148] in, This represents the base layer of the fused image.

[0149] The above formula ensures that... For target regions approaching 1, the fused image exhibits high-intensity features from the first source image; Background areas approaching 0 are merged to present the natural background features of the second source image.

[0150] The aforementioned submodule for constructing the fused image detail layer includes: In traditional methods, if the weight of the target region is too high, it can suppress information from the second source image in the fused image. To recover these lost details, this invention employs an incremental compensation method to reduce the confounding factor. (Second source image texture) is re-injected into the high-weight target region, expressed as:

[0151] in, To blend the detail layers of the image, This is the compensation gain coefficient.

[0152] The physical meaning of the above formula is that even within a defined target area, the surface texture is still considered valuable, and this is applied incrementally. The ratio is superimposed on the target brightness.

[0153] The aforementioned submodule for generating fused images includes: The base and detail layers of an image are linearly superimposed and blended to generate the final blended image. :

[0154] To prevent pixel value overflow and optimize visual effects, the fusion result is numerically clamped (limited to...). Within the specified range, adaptive histogram equalization (CLAHE) is used for post-processing. CLAHE divides the image into several non-overlapping blocks and limits the height of the histogram of each block (i.e., limits the magnitude of contrast enhancement), thereby enhancing the local contrast of the image while avoiding excessive amplification of noise.

[0155] The resulting fused image removes cross-modal interference artifacts, retains the target's physical intensity information, and clearly presents the texture details of the target surface.

[0156] The above-disclosed embodiments are merely preferred embodiments of the present invention, but the present invention is not limited thereto. Any non-creative variations that can be conceived by those skilled in the art, as well as any improvements and modifications made without departing from the principles of the present invention, should fall within the protection scope of the present invention.

Claims

1. A multimodal image fusion method based on a structural causal model and a texture compensation mechanism, characterized in that, Includes the following steps: Step S1, the step of extracting confounding factors, in which: Obtain the first source image and the second source image; Guided filtering is applied to the second source image to obtain the base layer of the second source image; The second source image is subjected to a difference operation with its base layer to obtain the detail layer of the second source image, and this detail layer is used as the confounding factor. Step S2, the step of estimating the local causal effect coefficient, in which: A local linear causal model is established between the first source image and the confounding factor, and the interference coefficient of the local linear causal model is estimated. Step S3, the counterfactual intervention denoising step, in which: Counterfactual intervention is performed based on the interference coefficients of the local linear causal model to remove cross-modal interference from the first source image and obtain the denoised physical signal. Step S4, the nonlinear mapping step of heat source probability weights, in which: Generate a target probability weight map based on the denoised physical signal; Step S5, the image reconstruction and fusion step, in which: Based on the target probability weight map, a method combining weighted basic layer and incremental compensation of detail layer is used to reconstruct and generate a fused image; Step S2 specifically includes: Based on the physical imaging laws, a local linear causal model is constructed between the first source image matrix and the confounding factor, decomposing the observed signal of the first source image matrix into real physical signals. and interference caused by texture The expression is: in, It is a spatially variable interference coefficient, representing the interference at the pixel level. The magnitude of the interference intensity of the second source image texture on the intensity of the first source image. Additive noise, Indicated in pixels The actual physical signal observed from the first source image matrix. The confounding factor is represented as a floating-point matrix containing both positive and negative values. Indicated in pixels The interference term of the observed signal from the first source image matrix; The interference coefficient is solved by using the statistical properties within a local window of the image. Internal interference coefficient Assuming is a constant, according to the principle of least squares, the objective function of the local linear causal model is defined as: in, The first source image matrix, The intercept of the local linear causal model; right Find the partial derivative and set it to zero. Through mathematical derivation, The optimal solution is: in, The local covariance between the first source image matrix and the confounding factor is represented by the following method: , The local variance of the confounding factor is calculated as follows: , Indicates in a local window Mean filtering operation within, is the regularization constant.

2. The multimodal image fusion method based on structural causal model and texture compensation mechanism according to claim 1, characterized in that, Step S1 specifically includes: The first and second source images are geometrically registered and normalized to the double-precision floating-point domain. Preprocessing yields the first source image matrix and the second source image matrix; Using the first source image matrix as the guiding image, guided filtering is applied to the second source image matrix to solve the following local linear model optimization problem: in, The output image of the guide filter, For guiding purposes, Represents pixels, Represented in pixels A local window centered on the center. Indicates the position of the pixel within the current local window. and These are the linear coefficients within the local window; Solving using the ridge regression cost function and The optimal value: in, It is a regularization parameter. Represents the second source image matrix At pixel position The value at that location, The cost function for ridge regression; The result is obtained by minimizing the ridge regression cost function. and Value: in, For guiding diagrams In local window The mean within, For guiding diagrams In local window within variance, For the second source image matrix In local window The mean within, For local windows The total number of pixels contained; Obtain linear coefficients and After determining the value, the guided filter averages the output images of all local windows covering the same pixel to obtain the base layer of the second source image. .

3. The multimodal image fusion method based on structural causal model and texture compensation mechanism according to claim 2, characterized in that, Step S1 further includes: Mathematical difference operations are used to extract the detail layer of the second source image, and the detail layer of the second source image is defined as the confounding factor. : in, and These are the coordinate components of a pixel in the image. The confounding factor is represented as a floating-point matrix containing both positive and negative values.

4. The multimodal image fusion method based on structural causal model and texture compensation mechanism according to claim 1, characterized in that, Step S2 further includes: All The values ​​are arranged according to their corresponding pixel positions to form a causality coefficient graph of the same size as the source image, denoted as . Figure; Guided filtering is used to calculate the original... The image was post-processed and optimized.

5. The multimodal image fusion method based on structural causal model and texture compensation mechanism according to claim 1, characterized in that, Step S3 specifically includes: Based on the Do operator theory and the local linear causal model in causal inference, the estimated interference term is subtracted from the first source image to perform counterfactual intervention: in, This is a counterfactual residual image.

6. The multimodal image fusion method based on structural causal model and texture compensation mechanism according to claim 5, characterized in that, Step S4 specifically includes: Counterfactual residual image Perform linear normalization to map its numerical range to The interval is used to obtain the normalized residual image. : Target probability weight map generated using the Sigmoid activation function : in, The slope of the activation function. The threshold of the activation function; calculated The range of values ​​is .

7. The multimodal image fusion method based on structural causal model and texture compensation mechanism according to claim 1, characterized in that, Step S5 specifically includes: Step S51: Construct the base layer of the fused image Using the calculated target probability weight map Linear weighting is applied to the first source image and the second source image: in, This represents the base layer of the fused image. Represents the second source image matrix; Step S52: Construct the detail layer of the fused image Incremental compensation method is used to reduce contamination factors. Re-injected into the high-weight target region, the expression is: in, To blend the detail layers of the image, This is the compensation gain coefficient; Step S53: Generate fused image The base and detail layers of an image are linearly superimposed and blended to generate the final blended image. : Numerical clamping is performed on the generated fused image; Adaptive histogram equalization is used for post-processing.

8. The multimodal image fusion method based on structural causal model and texture compensation mechanism according to claim 7, characterized in that, The numerical clamping process performed on the generated fused image in step S53 is limited to the following range: .

9. A multimodal image fusion method based on a structural causal model and texture compensation mechanism according to claim 8, characterized in that, The adaptive histogram equalization in step S53 includes parameters such as block size and contrast limit threshold.

Citation Information

Patent Citations

  • Multi-modal image fusion model construction method

    CN121147033A

  • Water body color recognition regression method and system based on space-time causality and manifold learning

    CN121170610A