X-ray-visible light multi-modal image fusion method
The dual-stream feature extraction and dynamic attention-based fusion method enhances X-ray and visible light image fusion by aligning modal features and adjusting weights adaptively, addressing issues of gradient distortion and incomplete detail retention in existing methods.
Patent Information
- Application Number
- CN202510413060.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-15
AI Technical Summary
The prior art is difficult to effectively improve the visual consistency and feature completeness of the fusion of X-ray and visible light images, and there are problems of gradient distortion, artifact interference and cross-modal consistency.
Dual-flow feature extraction, cross-modal dynamic attention guidance module and adaptive normalization module are used to pre-process images through gradient enhancement convolution blocks and contrast sensitivity convolution blocks, and combined with dynamic weight fusion and gradient constraints, deep fusion of images is achieved.
It significantly improves the visual nature and detail integrity of X-ray-visible images, eliminates artifact interference, and maintains the gradient consistency and feature integrity across modalities.
Smart Images

Figure CN120318091A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and in particular relates to an X-ray-visible light multimodal image fusion method. Background Art
[0002] As an important technology in the field of computer vision, multimodal image fusion is widely used in industrial inspection, medical imaging diagnosis and other scenarios. X-ray images can penetrate the surface of objects to reveal internal structural features, such as cracks inside metal components and foreign matter inclusions, but are limited by the imaging principle and have problems such as blurred surface texture and low signal-to-noise ratio. Although visible light images can clearly present the microscopic details of the surface of objects, they lack the ability to characterize the internal state. How to achieve deep coupling of the advantages of both and improve the visual consistency and feature completeness of the fused image has become a technical problem that needs to be solved urgently in this field.
[0003] In the prior art, the following methods are often used to achieve the fusion of X-ray and visible light images.
[0004] Fusion methods based on multi-scale decomposition: For example, wavelet transform and Laplace pyramid decomposition are used to separate the frequency domain of different modal images, and high-frequency components and low-frequency components are fused according to preset rules. However, such methods do not perform differential preprocessing on the low-contrast characteristics of X-ray images and the high-frequency texture characteristics of visible light images, resulting in gradient distortion of the fused image in the area of sudden density changes, and insufficient retention of high-frequency details.
[0005] Regional fusion method based on fixed weights: For example, a single measure such as regional variance or spatial frequency is used to calculate static fusion weights, and pixel-level weighted fusion is performed on dual-modal images. However, due to the significant nonlinear correlation between the energy distribution of X-ray images and the directional gradient characteristics of visible light images, the fixed weight strategy is difficult to dynamically adapt to the energy intensity changes in local areas, resulting in contrast suppression in high-density areas and loss of edge continuity in weak texture areas.
[0006] Fusion method based on traditional attention mechanism: Enhance the feature response of a specific modality through spatial attention map or channel attention mechanism. However, the existing attention model does not build cross-modal directional gradient association, resulting in semantic misalignment between the internal density distribution of X-rays and the surface texture of visible light in the edge area, reducing the consistency of the feature space.
[0007] Post-processing methods based on static parameters: After fusion, fixed threshold histogram matching or mean-variance normalization is often used for brightness correction. Since the spatial heterogeneity of the bimodal grayscale distribution is not considered, such methods cannot eliminate local artifacts (such as coupling interference between X-ray noise patches and visible light highlight areas) and will destroy the consistent expression of cross-modal gradients.
[0008] Therefore, a method is needed that can effectively improve the accuracy of multi-modal feature fusion, eliminate artifact interference, and maintain cross-modal consistency. Summary of the Invention
[0009] In view of the above-mentioned defects of the prior art, the technical problem to be solved by the present invention is: in view of the above problems of the prior art, to provide an X-ray-visible light multi-modal image fusion method. This method respectively extracts and enhances features of visible light and X-ray images through dual-stream feature extraction, uses a cross-modal dynamic attention guidance module to achieve image alignment in multiple angular directions, and then uses a dynamic weight fusion and adaptive normalization module to achieve deep fusion of images, improving the visual naturalness and detail integrity of the internal and external features of X-ray-visible light images. Finally, the final image is obtained after inverse normalization and channel compression.
[0010] To achieve the above object, the present invention adopts the following technical solutions: an X-ray-visible light multi-modal image fusion method, characterized in that: the method includes the following steps:
[0011] S1 Respectively perform dual-stream feature extraction on the original visible light image and the X-ray image. The visible light branch uses a gradient-enhanced convolutional block (GECB) to strengthen the image texture features, and the X-ray branch uses a contrast sensitivity convolutional block (CSCB) to improve the image contrast;
[0012] S2 Input the images with optimized features into the cross-modal dynamic attention guidance module, align the two-modal features on the edge structure through direction gradient consistency constraints, adjust the modal contribution weights through dynamic energy-sensitive mapping, and synthesize the direction consistency and energy weights to generate the final dynamic attention map;
[0013] S3 Adaptively fuse the X-ray and visible light features according to the attention weight map to achieve local optimal modal contribution allocation;
[0014] S4 Perform adaptive normalization fusion on the image after dynamic weight fusion. First, perform dynamic learning of statistical parameters; then map the image after dynamic weight fusion to the target statistical distribution for feature distribution alignment; finally, add a gradient constraint loss for gradient consistency constraint to retain key edge information;
[0015] S5 Perform an inverse normalization operation on the fused image to restore the physical pixel value range, and then perform channel compression. The multi-channel feature map is adaptively folded into a single-channel output through a learnable 1×1 convolutional kernel to obtain the final fused image.
[0016] As a preferred technical solution of the present invention: the structures of the gradient-enhanced convolutional block and the contrast sensitivity convolutional block in S1 are respectively as shown in S1-1 to S1-2:
[0017] S1-1: The gradient-enhanced convolutional block consists of a multi-directional gradient convolutional layer, a directional feature fusion module, a residual connection, and a normalization module;
[0018] S1-2: The contrast-sensitive convolutional block consists of a dilated convolutional layer, a local contrast attention module, and a weighted feature output module.
[0019] As a preferred technical solution of the present invention: The steps of aligning the two-modal features on the edge structure through the direction gradient consistency constraint in S2 include:
[0020] Multi-directional edge detection, using convolution kernels of 0°, 45°, 90°, and 135° to extract edge responses from X-ray (X) and visible light (V) respectively, as shown in the following formula:
[0021]
[0022] In the formula: K θ is the direction convolution kernel;
[0023] Calculate the cosine similarity in each direction to measure the consistency of the bimodal edges. The cosine similarity calculation formula is as follows:
[0024]
[0025] In the formula: S θ ∈[0, 1], the closer to 1 indicates the direction is consistent;
[0026] Set a threshold T to generate a direction consistency mask and set a direction validity flag, as shown in the following formula:
[0027]
[0028] Superimpose the consistency results in the four directions to generate a global direction weight, as shown in the following formula:
[0029]
[0030] As a preferred technical solution of the present invention: The steps of adjusting the modal contribution weight through dynamic energy-sensitive mapping in S2 include:
[0031] For visible light, calculate its gradient magnitude as the texture intensity. The calculation formula is as follows:
[0032]
[0033] For X-ray, calculate its local standard deviation as the structure intensity. The calculation formula is as follows:
[0034]
[0035] Where: Ω is the local window, and μ Ω is the mean value within the window;
[0036] Combining the two energy generation regions to generate region weights, as shown in the following formula:
[0037]
[0038] Where: σ(X) / σ(V) is the global standard deviation of the image.
[0039] As a preferred technical solution of the present invention: The calculation formula for generating the final dynamic attention map in S2 is as follows:
[0040] A(x,y) = Sigmoid(λ dir ·M dir (x,y) + λ energy ·W e (x,y))
[0041] Subsequently, the final fusion output is performed, as shown in the following formula:
[0042] F fused (x,y) = A(x,y)·X(x,y) + (1 - A(x,y))·V(x,y)
[0043] As a preferred technical solution of the present invention: The steps of adaptively fusing X-ray and visible light features according to the attention weight map in S3 include:
[0044] Step 1: Expand the attention map A to the feature map dimension (add the channel dimension) and perform weight dimension expansion;
[0045] Step 2: Linearly combine the bimodal features according to the weight map and perform feature weighted fusion.
[0046] As a preferred technical solution of the present invention: The operation of anti-normalizing the fused image in S5 is specifically to map the pixel values to the physical quantity value range of the original image by restoring the normalization coefficient in the preprocessing stage, so as to ensure the compatibility of the output image with the signal response characteristics of the target detection device.
[0047] As a preferred technical solution of the present invention: The operation of compressing the channels of the image in S5 is specifically to use a trainable 1×1 convolutional kernel to perform cross-channel information aggregation on the multi-channel feature map, and autonomously adjust the contribution weights of each feature channel through the dynamic optimization of the convolutional kernel parameters, and finally adaptively fold the high-dimensional features into a high-quality single-channel grayscale image that meets the detection requirements.
[0048] The beneficial effects of the present invention are as follows: A method for X-ray-visible light multimodal image fusion provided by the present invention, through multi-directional gradient enhancement convolution of visible light images and contrast sensitivity optimization preprocessing of X-ray images, strengthens the continuity of surface textures and the details of the energy distribution of internal structures. Using a cross-modal dynamic attention guidance module, it dynamically models the direction consistency of edge gradients and the correlation of energy intensities, realizes pixel-level alignment of feature structures between different source modalities, and avoids edge blurring and transition distortion caused by traditional weight assignment. Subsequently, combined with a dynamic fusion strategy with local energy-direction dual constraints, it adaptively adjusts the feature contribution weights, enhances the complementarity of the two modalities, and eliminates cross-modal brightness differences and high-frequency artifacts through a global adaptive normalization and reverse reconstruction module. Through the above methods, the visual naturalness and detail integrity of the internal and external feature fusion of X-ray-visible light images are significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 It is a brief flowchart of a method for X-ray-visible light multimodal image fusion provided by the present invention.
[0050] Figure 2 It is a flowchart of an adaptive normalization fusion method provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0052] The present invention proposes a method for X-ray-visible light multimodal image fusion, including:
[0053] Perform two-stream feature extraction on the original visible light image and X-ray image. The visible light branch uses a gradient enhancement convolution block to strengthen the image texture features, and the X-ray branch uses a contrast sensitivity convolution block to improve the image contrast;
[0054] Input the image with optimized features into a cross-modal dynamic attention guidance module. Align the features of the two modalities on the edge structure through the direction gradient consistency constraint, and adjust the modal contribution weights through dynamic energy sensitive mapping. Combine the direction consistency and energy weights to generate a final dynamic attention map;
[0055] Adaptive fusion of X-ray and visible light features according to the attention weight map to achieve local optimal modal contribution allocation;
[0056] Perform adaptive normalization fusion on the image after dynamic weight fusion. First, perform dynamic learning of statistical parameters; then map the image after dynamic weight fusion to the target statistical distribution for feature distribution alignment; finally, add gradient constraint loss for gradient consistency constraint to retain key edge information;
[0057] Perform denormalization operation on the fused image to restore the physical pixel value range, and then perform channel compression. The multi-channel feature map is adaptively folded into a single-channel output through a learnable 1×1 convolutional kernel to obtain the final fused image.
[0058] Furthermore, perform two-stream feature extraction on the original visible light image and X-ray image, including:
[0059] Adopt a gradient enhancement convolutional block to strengthen the texture features of the visible light image, and adopt a contrast sensitivity convolutional block to enhance the contrast of the X-ray image.
[0060] Specifically, the gradient enhancement convolutional block is used to capture and strengthen the complex texture features of the visible light image. The process includes:
[0061] Step 1: Use a four-direction convolutional kernel to extract edge responses, as shown in the following formula:
[0062]
[0063] In the formula, V ∈ R H×W×3 represents the visible light image, and K θ (k) ∈ R 3×3 is the convolutional kernel in the k-th direction (k ∈ {0°, 45°, 90°, 135°}), and F dir (k) ∈ R H×W corresponds to the edge response map in each direction;
[0064] Step 2: Concatenate the four-direction responses and reduce the dimension, as shown in the following formula:
[0065]
[0066] In the formula, W fusion ∈ R 4×C is the 1×1 convolutional weight, and the dimension is reduced to the number of channels C;
[0067] Step 3: Residual connection and normalization to retain low-frequency information, as shown in the following formula:
[0068]
[0069] W res : R 3→C maps the input visible light V to the feature space and performs instance normalization after superimposing with
[0070] Specifically, the contrast-sensitive convolution block enhances the visual saliency of key regions by minimizing the contrast noise ratio. The process includes:
[0071] Step 1: Use dilated convolution to extract large-scale structures, expand the receptive field to cover a wider density variation, as shown in the following formula:
[0072]
[0073] In the formula, K dilated ∈R 3×3 is a convolution kernel with a dilation rate d = 2;
[0074] Step 2: Use local contrast attention to calculate the local region variance, and generate spatial weights through the Sigmoid function. The calculation formula is as follows:
[0075]
[0076] In the formula, Var(·) calculates the variance of the local k×k block (k = 3), which characterizes the contrast strength. Conv a is a single-channel convolution and Sigmoid activation to generate a contrast attention map;
[0077] Step 3: Weighted feature output, retain the original information through residual connection, and the attention mechanism strengthens the high-contrast regions, as shown in the following formula:
[0078]
[0079] Furthermore, the direction gradient consistency constraint in the cross-modal dynamic attention guidance module incorporates the multi-directional cosine similarity into the fusion weight to match the bimodal edges and align the features of the two modalities on the edge structure.
[0080] Specifically, the direction gradient consistency constraint extracts the edge responses of X-ray and visible light images through multi-directional edge detection, then calculates the direction similarity to measure the consistency of the bimodal edges, then performs direction validity marking to generate a direction consistency mask, and finally superimposes the four-direction consistency results to generate a global direction weight.
[0081] Furthermore, the dynamic energy-sensitive mapping in the cross-modal dynamic attention guidance module calculates the gradient magnitude as the gradient energy of the visible light, calculates the local standard deviation as the contrast energy of the X-ray image, and combines the two energies to generate a region weight.
[0082] Furthermore, by integrating the direction consistency and the energy weight, a final dynamic attention map is generated.
[0083] Specifically, adaptively fusing X-ray and visible light features according to the attention weight map, the process includes:
[0084] Step 1: Expand the attention map A to the feature map dimension (add the channel dimension) and perform weight dimension expansion;
[0085] Step 2: Linearly combine the bimodal features according to the weight map for feature weighted fusion.
[0086] Specifically, the statistical parameter learning in adaptive normalization fusion, the process includes:
[0087] Step 1: Extract the input feature statistics, and calculate the mean and standard deviation of the current batch using the feature map output by dynamic weight fusion. The calculation formulas are as follows:
[0088]
[0089] Step 2: Generate adaptive target mean and standard deviation according to the global statistical characteristics of the input X-ray (X) and visible light (V), and perform target statistical parameter prediction, as shown in the following formula:
[0090] μ target = α·μ X +(1 - α)·μ V
[0091] σ target = β·σ X +(1 - β)·σ V
[0092] In the formula: α, β ∈ [0, 1] are learnable parameters, and the initial values are both 0.5.
[0093] Specifically, the calculation formula for feature distribution alignment is as follows:
[0094]
[0095] Specifically, the image feature denormalization operation maps the features from the normalized distribution to the actual pixel value range. The process includes:
[0096] Step 1: Perform linear denormalization, as shown in the following formula:
[0097] F denorm = F final × s scale + s shift
[0098] In the formula: F final ∈ R H×W×C is the feature map after adaptive normalization, and s scale is the scaling coefficient used to restore the contrast, sshift is the offset coefficient used to adjust the brightness;
[0099] Step 2: Perform truncation constraints to ensure that the pixel values are within the legal range, as shown in the following formula:
[0100] I denorm = Clip(F denorm , 0, 255)
[0101] Specifically, the image channel compression operation weights the multi-channel feature map by a coefficient and converts the multi-channel feature map into a single-channel number output, as shown in the following formula:
[0102]
[0103] In the formula: w c is the weight coefficient.
[0104] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations according to the concept of the present invention without creative work. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field based on the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the prior art should be within the protection scope determined by the claims.
Claims
1. An X-ray - visible light multimodal image fusion method, characterized in that: It includes the following steps: S1: Perform two-stream feature extraction on the original visible light image and X-ray image respectively. The visible light branch uses a gradient enhancement convolutional block to strengthen the image texture features, and the X-ray branch uses a contrast sensitivity convolutional block to enhance the image contrast. S2: Input the image with optimized features into the cross-modal dynamic attention guidance module. Align the two-modal features on the edge structure through the direction gradient consistency constraint, adjust the modal contribution weights through the dynamic energy-sensitive mapping, and generate the final dynamic attention map by synthesizing the direction consistency and energy weights. S3: Adaptively fuse the X-ray and visible light features according to the attention weight map to achieve the optimal local modal contribution allocation. S4: Perform adaptive normalization fusion on the image after dynamic weight fusion. First, perform dynamic learning of statistical parameters. Then map the image after dynamic weight fusion to the target statistical distribution for feature distribution alignment. Finally, add a gradient constraint loss for gradient consistency constraint to retain key edge information. S5: Perform an inverse normalization operation on the fused image to restore the physical pixel value range, and then perform channel compression. Adaptively fold the multi-channel feature map into a single-channel output through a learnable 1×1 convolutional kernel to obtain the final fused image.
2. The X-ray - visible light multimodal image fusion method according to claim 1, characterized in that: The operation process of using the gradient enhancement convolutional block to strengthen the visible light image features is as follows: Step 1: Use gradient convolutional kernels in the directions of 0°, 45°, 90°, and 135° to capture multi-angle edges and perform multi-directional edge detection. Step 2: Concatenate the four-direction responses and reduce the dimension for direction feature fusion operation. Step 3: Retain the low-frequency information of the image and perform residual connection and normalization operations.
3. An X-ray - visible light multi-modal image fusion method according to claim 1, characterized in that: The operation process of using the contrast sensitivity convolutional block to enhance the X-ray image contrast is as follows: Step 1: Use dilated convolution to extract large-scale structures and expand the receptive field to cover a wider density change. Step 2: Use local contrast attention to calculate the local region variance and generate spatial weights through the Sigmoid function. Step 3: Output weighted features, retain the original information through residual connection, and strengthen the high-contrast regions through the attention mechanism.
4. The X-ray - visible light multimodal image fusion method according to claim 1, characterized in that: The dynamic attention guidance module includes: a direction gradient consistency constraint unit that calculates the correlation of the gradient direction histograms of visible light and X-ray as the structure alignment coefficient; a dynamic energy-sensitive mapping unit that establishes a weight allocation model based on the energy ratio of X-ray pixels in the infrared and visible light bands; a hybrid attention generation unit that performs non-linear product fusion on the direction alignment coefficient and the energy weight.
5. A method for X-ray-visible light multimodal image fusion according to claim 1, characterized in that: The operation process of the adaptive normalization fusion is as follows: Step 1: Obtain dynamic statistical parameters, and based on the feature map after dynamic weight fusion, calculate the mean and standard deviation of the current batch of data in real time. Step 2: Generate the target statistics. According to the global statistical features of the X-ray image and the visible light image, predict the target mean and target standard deviation through a linear interpolation module, and the interpolation weight is a trainable parameter and is initialized to an equal ratio value. Step 3: Map the fused features to the feature space defined by the target statistical parameters so that the two-modal features follow the unified distribution characteristics. Step 4: Gradient retention optimization, applying an edge direction consistency constraint during the distribution alignment process to maintain the geometric structure alignment of cross-modal features.