Dual-energy X-ray image fusion method and device, computer equipment and storage medium

The dual-energy X-ray images are fused by expanding the convolution of the receptive field and multi-attention mechanism, which solves the problem of insufficient image fusion in the prior art, and achieves a higher quality image fusion effect, meeting the needs of high-precision detection.

CN119991472AInactive Publication Date: 2025-05-13TECHIK INSTR SHANGHAI

Patent Information

Application Number
CN202510465452.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing dual-energy X-ray image fusion method is difficult to fully extract and utilize the complementary information between low-energy X-ray images and high-energy images, resulting in blurred details, loss of edges, incomplete global features of the fusion image, and unable to meet the needs of high-precision detection.

Method used

The convolution method of expanding the receptive field is used to convolve the low-energy and high-energy X-ray images, generate convolution feature maps and fuse them, and enhance the initial fusion feature maps through a multivariate attention mechanism. Finally, feature extraction of the enhanced features through an expanded convolution layer to generate the final fusion image.

Benefits of technology

By extracting wider context information and automatically focusing on key areas, the visual and semantic consistency and clarity of the fusion image are significantly improved, detail expressiveness and global consistency are enhanced, and high-precision detection needs are met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991472A_ABST
    Figure CN119991472A_ABST
Patent Text Reader

Abstract

The invention relates to a dual-energy X-ray image fusion method and device, computer equipment and a storage medium. The method comprises the following steps: acquiring an X-ray low-energy image and an X-ray high-energy image; respectively carrying out convolution on the X-ray low-energy image and the X-ray high-energy image by adopting a receptive field expanding convolution mode to generate convolution feature maps, and fusing the convolution feature maps to obtain a preliminary fusion feature map; performing enhancement processing on the preliminary fusion feature map through a multivariate attention mechanism to enhance complementary information expression represented by the fusion feature so as to obtain an enhanced feature; and performing feature extraction on the enhanced features through an expansion convolution layer to obtain convolution output features, and performing channel adjustment and format reduction processing on the convolution output features to generate a final fusion image. The method has the effect of improving the image fusion quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image processing, and in particular to a dual-energy X-ray image fusion method, device, computer equipment and storage medium. Background Art

[0002] At present, X-ray imaging technology is widely used in industrial inspection, medical imaging, security inspection and other fields, mainly for visual inspection of the internal structure of objects. In order to improve imaging accuracy, dual-energy X-ray imaging technology uses two X-ray sources with different energies and utilizes the differences in the absorption characteristics of materials to X-rays of different energies to effectively enhance the material information of the image and improve the image contrast and resolution.

[0003] Existing dual-energy X-ray image fusion methods usually use simple pixel-level weighted fusion or traditional convolutional neural network (CNN) for image synthesis. Such methods are difficult to fully extract and utilize the complementary information of low-energy X-ray images and high-energy images during the fusion process, and due to the lack of accurate enhancement of key area features, they easily lead to blurred details, lost edges, and incomplete global features of the fused image, which cannot meet the needs of high-precision detection.

[0004] The above-mentioned existing technical solutions have the following defects: the existing dual-energy X-ray image fusion method is insufficient in feature extraction and fusion accuracy, resulting in the inability to effectively take into account the edge details and global features of the fused image, so there is room for improvement. Summary of the invention

[0005] In order to improve the quality of image fusion, the present application provides a dual-energy X-ray image fusion method, apparatus, computer equipment and storage medium.

[0006] The above-mentioned invention objective of the present application is achieved through the following technical solutions: A dual-energy X-ray image fusion method, the method comprising: Acquire low-energy X-ray images and high-energy X-ray images; The X-ray low-energy image and the X-ray high-energy image are respectively convolved by a convolution method that expands the receptive field to generate a convolution feature map, and the convolution feature map is fused to obtain a preliminary fused feature map; The preliminary fused feature map is enhanced by a multi-attention mechanism to strengthen the complementary information expression represented by the fused feature, thereby obtaining an enhanced feature; The enhanced features are extracted through an expanded convolution layer to obtain convolution output features, and then channel adjustment and format restoration processing are performed on the convolution output features to generate a final fused image.

[0007] By adopting the above technical solution, by acquiring X-ray low-energy images and X-ray high-energy images, information sources of different energy levels can be provided for image fusion, thereby ensuring that the fused image can contain more details and information; by using the convolution method of expanding the receptive field to convolve the low-energy and high-energy images respectively, the wider contextual information in the feature map can be enhanced, thereby improving the feature expression and recognition capabilities of the image; by using the multi-attention mechanism to enhance the preliminary fused feature map, the key areas in the image can be automatically focused, thereby effectively improving the visual and semantic consistency and clarity of the fused image; by using the expanded convolution layer to extract the enhanced features, the feature information at different scales can be captured, thereby enhancing the detail expression and global consistency of the fused image.

[0008] In one example, the present application may be further configured as follows: the convolution method of expanding the receptive field is used to convolve the X-ray low-energy image and the X-ray high-energy image respectively to generate a convolution feature map, and the convolution feature map is fused, specifically including: Using a 7×7 convolution kernel to perform a convolution operation on the X-ray low-energy image and the X-ray high-energy image to generate a convolution feature map; By presetting the fusion formula The convolution feature maps are fused, where F1 is the preliminary fusion feature map, I L is the X-ray low-energy image, I H is the X-ray high energy image.

[0009] By adopting the above technical solution, by using a 7×7 convolution kernel to perform convolution operations on low-energy and high-energy images, the receptive field of the convolution layer can be improved, thereby extracting richer image context information and enhancing the expressiveness of the feature map; by fusing the convolution feature map through a preset fusion formula, it is possible to optimize according to the feature differences between the low-energy and high-energy images, thereby ensuring the consistency of the final fused image in details and structure, and improving the quality of image fusion.

[0010] In one example, the present application may be further configured as follows: the preliminary fused feature map is enhanced by a multi-attention mechanism to strengthen the complementary information expression represented by the fused feature, thereby obtaining enhanced features, specifically including: The feature channels and spatial regions of the preliminary fused feature map are weighted respectively by an efficient channel attention mechanism and a spatial attention mechanism to enhance key region features in the preliminary fused feature map; The preliminary fusion feature map and the enhanced preliminary fusion feature map are residually connected to obtain the enhanced feature.

[0011] By adopting the above technical scheme, the feature channels of the preliminary fused feature map are weighted by an efficient channel attention mechanism, which can automatically enhance the influence of important channels, thereby highlighting the key feature information in the image and improving the expressiveness of the image in details and key areas; by weighting the spatial area of ​​the image by the spatial attention mechanism, the important areas in the image can be highlighted and the influence of irrelevant areas can be reduced, thereby improving the visual effect and structural performance of the fused image; by fusing the preliminary fused feature map with the enhanced feature map through the residual connection, the integrity of the original information can be maintained, while effectively enhancing the expressiveness of the image and improving the quality and accuracy of the image.

[0012] In one example, the present application may be further configured as follows: extracting the enhanced features through the dilated convolution layer to obtain convolution output features, and then performing channel adjustment and format restoration processing on the convolution output features to generate a final fused image, specifically including: Using convolution kernels of different sizes to perform secondary convolution on the enhanced features to capture corresponding global features and obtain the convolution output features; By default output formula The convolution output features are restored to obtain the final fused image, where Out is the final fused image, and F out Output features for the convolution.

[0013] By adopting the above technical solution, by using convolution kernels of different sizes to perform secondary convolution on the enhanced features, it is possible to capture feature information at different scales, thereby improving the local and global expressiveness of the image; by restoring the convolution output features through a preset output formula, it is possible to effectively integrate the feature information and generate the final fused image, thereby ensuring the detail retention and global consistency of the image, and improving the image quality and visual effect.

[0014] In one example, the present application may be further configured as follows: the dual-energy X-ray image fusion method further includes: Based on the preset loss function , optimize the image fusion process, where L P is the perceptual loss, L G is the gradient loss, L S is the structural loss, σ and γ are weight coefficients.

[0015] By adopting the above technical solution, by using a combination of perceptual loss, gradient loss and structural loss to optimize the output image quality, the image can be effectively optimized in terms of perceptual quality, edge details and structural consistency through the combined effect of multiple loss functions, thereby improving the overall expressiveness and accuracy of the fused image.

[0016] In one example, the present application may be further configured as follows: the preset loss function specifically includes: By perceiving loss , constraining the high-level semantic features between the final fused image, the X-ray low-energy image and the X-ray high-energy image, wherein, , , They are high-level features extracted by pre-trained convolutional neural networks, I F is the final fused image, α is a balance parameter used to adjust the contribution of the loss term; Through gradient loss , maintain the edge and detail information of the image, where is the gradient representation of the image, to obtain a maximum gradient between the X-ray low-energy image and the X-ray high-energy image; Through structural loss , constraining the local structural information of the final fused image, the X-ray low-energy image and the X-ray high-energy image to improve the structural similarity of the fused image, wherein SSIM(·) is the structural similarity index, The part with the strongest structure is selected for comparison between the X-ray low-energy image and the X-ray high-energy image.

[0017] By adopting the above technical solution, the differences in the high-level feature space are calculated through perceptual loss, which can ensure the consistency of the image at the semantic level, thereby improving the perceptual quality of the image; the gradient difference of the image is calculated through gradient loss, which can effectively maintain the edge information of the image and avoid detail loss, thereby improving the clarity of the image; the structural consistency of the image is evaluated through structural loss, which can optimize the structural performance of the image and ensure that the details and texture of the image are retained, thereby improving the structural restoration ability and visual quality of the image.

[0018] The second object of the invention is achieved by the following technical solutions: A dual-energy X-ray image fusion device, comprising: An image acquisition module, used for acquiring X-ray low-energy images and X-ray high-energy images; A convolution operation module, used to respectively convolve the X-ray low-energy image and the X-ray high-energy image using a convolution method that expands the receptive field to generate a convolution feature map, and fuse the convolution feature map to obtain a preliminary fused feature map; A multi-attention mechanism module, used for enhancing the preliminary fused feature map through a multi-attention mechanism to strengthen the complementary information expression represented by the fused feature, thereby obtaining enhanced features; The dilated convolution module is used to extract the enhanced features through the dilated convolution layer to obtain convolution output features, and then perform channel adjustment and format restoration processing on the convolution output features to generate a final fused image.

[0019] By adopting the above technical solution, by acquiring X-ray low-energy images and X-ray high-energy images, information sources of different energy levels can be provided for image fusion, thereby ensuring that the fused image can contain more details and information; by using the convolution method of expanding the receptive field to convolve the low-energy and high-energy images respectively, the wider contextual information in the feature map can be enhanced, thereby improving the feature expression and recognition capabilities of the image; by using the multi-attention mechanism to enhance the preliminary fused feature map, the key areas in the image can be automatically focused, thereby effectively improving the visual and semantic consistency and clarity of the fused image; by using the expanded convolution layer to extract the enhanced features, the feature information at different scales can be captured, thereby enhancing the detail expression and global consistency of the fused image.

[0020] The third objective of the present application is achieved through the following technical solutions: A computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the dual-energy X-ray image fusion method when executing the computer program.

[0021] The fourth objective of the present application is achieved through the following technical solutions: A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the dual-energy X-ray image fusion method.

[0022] In summary, this application includes the following beneficial technical effects: 1. By acquiring X-ray low-energy images and X-ray high-energy images, it is possible to provide information sources of different energy levels for image fusion, thereby ensuring that the fused image can contain more details and information; by using the convolution method of expanding the receptive field to convolve the low-energy and high-energy images respectively, it is possible to enhance the wider contextual information in the feature map, thereby improving the feature expression and recognition capabilities of the image; by using the multi-attention mechanism to enhance the preliminary fusion feature map, it is possible to automatically focus on the key areas in the image, thereby effectively improving the visual and semantic consistency and clarity of the fused image; by using the dilated convolution layer to extract the enhanced features, it is possible to capture the feature information at different scales, thereby enhancing the detail expression and global consistency of the fused image; 2. By using a 7×7 convolution kernel to perform convolution operations on low-energy and high-energy images, the receptive field of the convolution layer can be improved, thereby extracting richer image context information and enhancing the expressiveness of the feature map; by fusing the convolution feature map through a preset fusion formula, it can be optimized according to the feature differences between low-energy and high-energy images, thereby ensuring the consistency of the final fused image in details and structure, and improving the quality of image fusion; 3. By weighting the feature channels of the preliminary fused feature map through an efficient channel attention mechanism, the influence of important channels can be automatically enhanced, thereby highlighting the key feature information in the image and improving the image's expressiveness in details and key areas; by weighting the spatial areas of the image through a spatial attention mechanism, the important areas in the image can be highlighted and the influence of irrelevant areas can be reduced, thereby improving the visual effect and structural performance of the fused image; by fusing the preliminary fused feature map with the enhanced feature map through a residual connection, the integrity of the original information can be maintained, while effectively enhancing the image's expressiveness, and improving the image's quality and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 is a flow chart of a dual-energy X-ray image fusion method in one embodiment of the present application; Figure 2 This is a principle block diagram of a dual-energy X-ray image fusion device in one embodiment of the present application; Figure 3 It is a schematic diagram of a device in an embodiment of the present application. DETAILED DESCRIPTION

[0024] The present application is further described in detail below in conjunction with the accompanying drawings.

[0025] In one embodiment, if Figure 1 As shown, the present application discloses a dual-energy X-ray image fusion method, which specifically includes the following steps: S10: Acquire a low-energy X-ray image and a high-energy X-ray image.

[0026] Specifically, an X-ray low-energy image and an X-ray high-energy image of the target object are synchronously acquired through an image acquisition device. The X-ray low-energy image and the X-ray high-energy image are respectively generated by a low-energy X-ray light source and a high-energy X-ray light source. The low-energy X-ray image contains structural information of low-density or light elements in the target object, and the high-energy X-ray image contains structural information of high-density or heavy elements in the target object. The acquisition process of the low-energy image and the high-energy image is completed through a multi-energy switching mode to ensure the consistency of the two images in spatial position, imaging angle and image scale, and to ensure that the pixel positions between the low-energy image and the high-energy image can correspond one by one, thereby avoiding spatial dislocation and information loss during the image fusion process.

[0027] S20: The X-ray low-energy image and the X-ray high-energy image are convolved respectively by using a convolution method with an expanded receptive field to generate a convolution feature map, and the convolution feature map is fused to obtain a preliminary fused feature map.

[0028] Specifically, by performing convolution processing on the X-ray low-energy image and the X-ray high-energy image respectively, the convolution kernel with a large receptive field is used to expand the spatial perception range and obtain more comprehensive image features. The convolution operation uses a combination of deep convolution and point convolution. First, deep convolution is performed using a large convolution kernel to capture the global features of the target area, and then point convolution is used to adjust the channel information to ensure the integrity and distinguishability of the feature space. The feature map after convolution is merged through a feature-level fusion strategy, and feature integration is performed using element-by-element addition or channel splicing to ensure the effective fusion of low-energy images and high-energy images at multi-scale information, and generate a preliminary fused feature map. The preliminary fused feature map can fully express the complementary information in the low-energy image and the high-energy image, providing a basis for subsequent feature enhancement and output image generation.

[0029] S30: The preliminary fused feature map is enhanced through a multi-attention mechanism to strengthen the complementary information expression represented by the fused features, thereby obtaining enhanced features.

[0030] Specifically, multi-dimensional feature enhancement is performed on the preliminary fused feature map. First, the global statistical information of each feature channel is calculated in the channel dimension through the channel attention mechanism, and different weights are assigned to each feature channel according to the channel importance to enhance the expression of key information. The spatial attention mechanism is further applied in the spatial dimension to extract the saliency information of the local area in the feature map, and the spatial position is weighted to highlight the spatial characteristics and edge details of the target area. Finally, the channel enhanced features are residually connected with the spatial enhanced features to retain the basic information in the original feature map, ensuring that the important information of the original image is not lost while enhancing the features, so as to form an enhanced feature map. The enhanced feature map has a stronger feature distinction ability and can improve the recognition and segmentation accuracy of complex targets.

[0031] S40: extracting the enhanced features through the dilated convolution layer to obtain convolution output features, and then performing channel adjustment and format restoration processing on the convolution output features to generate a final fused image.

[0032] Specifically, a dilated convolution layer is used to extract multi-scale features from the enhanced feature map. The dilated convolution introduces a hole expansion factor to expand the receptive field while avoiding the decrease in feature resolution. Convolution kernels with different expansion rates are used to perform multi-scale convolution operations on the enhanced features to capture local detail features and global background information respectively, generating a multi-layer feature map. The multi-layer feature map is then channel-adjusted and the adjusted feature map is deconvolved to restore the spatial resolution consistent with the input image, ensuring that the spatial details of the output image are consistent with the original image. The final fused image retains a variety of features in the X-ray low-energy image and the X-ray high-energy image, and can generate clearer and more complete image results in complex scenes.

[0033] In one embodiment, in step S20, the X-ray low-energy image and the X-ray high-energy image are convolved respectively by using a convolution method for expanding the receptive field to generate a convolution feature map, and the convolution feature map is fused, specifically including: S21: Use a 7×7 convolution kernel to perform convolution operations on the X-ray low-energy image and the X-ray high-energy image to generate a convolution feature map.

[0034] Specifically, a 7×7 convolution kernel is used to extract features from X-ray low-energy images and X-ray high-energy images respectively. In the convolution operation, the input X-ray low-energy image and X-ray high-energy image will be scanned in sequence by the sliding 7×7 convolution kernel. Each sliding of the convolution kernel will perform a dot product operation with the corresponding local area in the image, and the calculation result will be used as the feature value of the area. The convolution operation will extract multi-scale information from the image. Compared with the commonly used 3×3 convolution kernel, the 7×7 convolution kernel can cover a larger receptive field and capture a wider range of spatial feature information. It is particularly suitable for obtaining edge features and material differences across regions in X-ray low-energy images and X-ray high-energy images.

[0035] S22: By pre-setting the fusion formula The convolution feature maps are fused, where F1 is the initial fusion feature map, I L For X-ray low-energy images, I H It is a high-energy X-ray image.

[0036] Specifically, a preset fusion formula is used to fuse the convolutional feature maps generated from the X-ray low-energy image and the X-ray high-energy image, and the channel-level splicing method, i.e., Concat, is used to fuse the two convolutional feature maps in the channel dimension. The two types of energy information are mapped to a unified feature space to ensure that the fused feature map can completely retain the material information of the low-energy image and the edge and density information of the high-energy image.

[0037] In one embodiment, in step S30, the preliminary fused feature map is enhanced by a multi-attention mechanism to strengthen the complementary information expression represented by the fused features, thereby obtaining enhanced features, specifically including: S31: The feature channels and spatial regions of the preliminary fused feature map are weighted respectively through efficient channel attention mechanism and spatial attention mechanism to enhance the key area features in the preliminary fused feature map.

[0038] Specifically, the feature channels are weighted through the efficient channel attention mechanism ECA, including performing global average pooling on the preliminary fused feature map F1 to obtain the global statistics of each channel, and then performing a one-dimensional convolution operation on the global statistics of the channel to learn the dependency relationship of each channel and obtain the channel attention weight vector. The preliminary fused feature map F1 is channel-weighted according to the attention weight vector to calculate the enhanced feature map F ECA ; The spatial dimension is weighted through the spatial attention mechanism BMA, including performing maximum pooling and average pooling on the preliminary fusion feature map F1 at the same time to obtain a spatial feature map, and then performing a convolution operation on the spatial feature map to obtain a spatial attention weight map. The preliminary fusion feature map F1 is spatially weighted through the spatial attention weight map to calculate the spatially enhanced feature map F BAM ; Through RC, the preliminary fusion feature map is convolved twice to obtain the convolved feature map F RC ; Finally, the ECA and BAM enhanced feature maps are multiplied to obtain the feature F MC and F MS ,Right now .

[0039] Residual connection is performed on the preliminary fusion feature map and the enhanced preliminary fusion feature map to obtain enhanced features.

[0040] Specifically, after channel and space enhancement of the preliminary fusion feature map, a residual connection method is used , retaining the complementary relationship between the original information and the enhanced features, adding the initial fused feature map and the enhanced feature map element by element according to the channel dimension to obtain the enhanced features. Among them, the residual connection can effectively avoid the loss of information during the enhancement process, while enhancing the stability and anti-interference of the feature map. The enhanced features not only retain the basic information in the original fused features, but also strengthen the key area features extracted by the attention mechanism, and improve the adaptability and discrimination ability of the features to complex scenes.

[0041] In one embodiment, in step S40, feature extraction is performed on the enhanced features through the dilated convolution layer to obtain convolution output features, and then channel adjustment and format restoration processing are performed on the convolution output features to generate a final fused image, specifically including: S41: Use convolution kernels of different sizes to perform secondary convolution on the enhanced features to capture the corresponding global features and obtain the convolution output features.

[0042] Specifically, multi-scale dilated convolution is used for feature extraction to enhance features. Dilated convolution introduces a dilation factor to expand the receptive field of the convolution kernel without increasing the computational complexity. as well as , the enhanced feature map is convolved twice using 3×3, 5×5, 7×7 and 9×9 dilated convolution kernels respectively. The larger the size of the convolution kernel, the wider the receptive field, and the longer-distance feature dependencies can be captured. Using multiple convolution kernels can simultaneously extract local detail features and global context information, and generate multi-scale convolution output feature maps. The convolution output feature map has more comprehensive spatial information and rich semantic features, providing multi-level feature support for the generation of the final fused image.

[0043] S42: Through the preset output formula The convolution output features are restored to obtain the final fused image, where Out is the final fused image and F out is the convolution output feature.

[0044] Specifically, through a 1×1 convolution, The output channel of the convolution output feature is adjusted to restore the spatial resolution of the feature map to be consistent with the input X-ray image, and the final fused image Out is generated that conforms to the spatial layout of the original image. The final fused image can take into account the characteristics of both the X-ray low-energy image and the high-energy image, presenting higher image contrast and richer detail information.

[0045] In one embodiment, the dual-energy X-ray image fusion method further includes: S50: Based on the preset loss function , optimize the image fusion process, where L P is the perceptual loss, L G is the gradient loss, L S is the structural loss, σ and γ are weight coefficients.

[0046] Specifically, in the training stage of image fusion, a composite loss function consisting of perceptual loss, gradient loss and structural loss is defined to optimize the fusion effect. The perceptual loss ensures that the fused image can retain the semantic information of the original image by calculating the difference between the fused image and the low-energy image and the high-energy image in the deep feature space. The gradient loss retains the edges and details in the image by comparing the edge gradients of the fused image with the source image. The structural loss ensures the fidelity of the global structure and texture information of the image by calculating the structural similarity index between the fused image and the source image. The perceptual loss, gradient loss and structural loss are multiplied by the set weight parameters σ and γ respectively to control the contribution of each loss to the optimization process, and finally the quality of the fused image is improved by minimizing the total loss.

[0047] In one embodiment, in step S10, the preset loss function specifically includes: S51: Through Perceptual Loss , constrains the high-level semantic features between the final fused image, X-ray low-energy image and X-ray high-energy image, where , , They are high-level features extracted by pre-trained convolutional neural networks, I F is the final fused image, and α is a balance parameter used to adjust the contribution of the loss term.

[0048] Specifically, high-level features of X-ray low-energy images, X-ray high-energy images and the final fused image are extracted through pre-trained convolutional neural networks such as VGG16, and the Euclidean distance between these feature maps is calculated to measure the similarity of the feature maps in the high-level semantic feature space. The perceptual quality between the fused image and the source image is measured by comparing the high-level feature differences of the images. Finally, the network optimization is guided by the loss function to ensure that the final generated fused image is similar to the low-energy image and the high-energy image in high-level semantics, thereby improving the perceptual effect of the image.

[0049] S52: Through gradient loss , maintain the edge and detail information of the image, where is the gradient representation of the image, To obtain the maximum gradient between the X-ray low-energy image and the X-ray high-energy image.

[0050] Specifically, the edge information in the image is obtained by calculating the gradients of the X-ray low-energy image, the X-ray high-energy image and the final fused image, the gradient of the image is calculated using a convolution operation, the edge features of the image are extracted, and then the gradients between the X-ray low-energy image and the X-ray high-energy image are compared, and the maximum gradient value is selected as a reference to maintain the key details in the fused image. The gradient loss calculates the gradient difference between the fused image and the source image, and ultimately minimizes the loss, ensuring that the fused image avoids detail loss while retaining edge information.

[0051] S53: Through structural loss , constrain the local structural information of the final fused image, X-ray low-energy image and X-ray high-energy image to improve the structural similarity of the fused image, where SSIM(·) is the structural similarity index, To select the strongest structure part between the X-ray low-energy image and the X-ray high-energy image for comparison.

[0052] Specifically, the structural similarity index SSIM is used to measure the structural similarity of images. The brightness, contrast and structural information of the image are evaluated by SSIM, and the similarity of the strongest structural part between the X-ray low-energy image and the X-ray high-energy image is calculated. Then, the similarity is compared with the fused image to obtain the structural loss, and finally the structural loss is minimized to ensure the structural consistency of the fused image, thereby improving the image fusion quality.

[0053] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0054] In one embodiment, a dual-energy X-ray image fusion device is provided, and the dual-energy X-ray image fusion device corresponds to the dual-energy X-ray image fusion method in the above embodiment. Figure 2 As shown, the dual-energy X-ray image fusion device includes an image acquisition module, a convolution operation module, a multi-attention mechanism module, and a dilated convolution module. The functional modules are described in detail as follows: An image acquisition module, used for acquiring X-ray low-energy images and X-ray high-energy images; A convolution operation module is used to convolve the X-ray low-energy image and the X-ray high-energy image respectively by using a convolution method that expands the receptive field to generate a convolution feature map, and to fuse the convolution feature map to obtain a preliminary fused feature map; The multi-attention mechanism module is used to enhance the preliminary fusion feature map through the multi-attention mechanism to strengthen the complementary information expression of the fusion feature representation, thereby obtaining enhanced features; The dilated convolution module is used to extract the enhanced features through the dilated convolution layer to obtain the convolution output features, and then perform channel adjustment and format restoration processing on the convolution output features to generate the final fused image.

[0055] Optionally, the convolution operation module specifically includes: A convolution operation submodule is used to perform a convolution operation on the X-ray low-energy image and the X-ray high-energy image using a 7×7 convolution kernel to generate a convolution feature map; Feature fusion submodule, used to preset fusion formula The convolution feature maps are fused, where F1 is the initial fusion feature map, I L For X-ray low-energy images, I H It is a high-energy X-ray image.

[0056] Optionally, the multi-attention mechanism module specifically includes: The attention mechanism submodule is used to weight the feature channels and spatial regions of the preliminary fused feature map respectively through an efficient channel attention mechanism and a spatial attention mechanism, so as to enhance the key region features in the preliminary fused feature map; The residual connection submodule is used to perform residual connection on the preliminary fusion feature map and the enhanced preliminary fusion feature map to obtain enhanced features.

[0057] Optionally, the dilated convolution module specifically includes: The dilated convolution submodule is used to perform secondary convolution on the enhanced features using convolution kernels of different sizes to capture the corresponding global features and obtain the convolution output features; Format restoration submodule, used to output the formula by default The convolution output features are restored to obtain the final fused image, where Out is the final fused image and F out is the convolution output feature.

[0058] Optionally, the dual-energy X-ray image fusion method further includes: Training module, used based on preset loss function , optimize the image fusion process, where L P is the perceptual loss, L G is the gradient loss, L S is the structural loss, σ and γ are weight coefficients.

[0059] Optionally, the training modules include: The perceptual loss calculation submodule is used to calculate the perceptual loss , constrains the high-level semantic features between the final fused image, X-ray low-energy image and X-ray high-energy image, where , , They are high-level features extracted by pre-trained convolutional neural networks, I F is the final fused image, α is a balance parameter used to adjust the contribution of the loss term; Gradient loss calculation submodule, used to calculate the gradient loss , maintain the edge and detail information of the image, where is the gradient representation of the image, To obtain the maximum gradient between the X-ray low energy image and the X-ray high energy image; The structural loss calculation submodule is used to calculate the structural loss , constrain the local structural information of the final fused image, X-ray low-energy image and X-ray high-energy image to improve the structural similarity of the fused image, where SSIM(·) is the structural similarity index, To select the strongest structure part between the X-ray low-energy image and the X-ray high-energy image for comparison.

[0060] For the specific definition of the dual-energy X-ray image fusion device, please refer to the definition of the dual-energy X-ray image fusion method above, which will not be repeated here. Each module in the above-mentioned dual-energy X-ray image fusion device can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0061] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 3 As shown. The computer device includes a processor, a memory, a network interface and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a dual-energy X-ray image fusion method is implemented.

[0062] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the following steps when executing the computer program: Acquire low-energy X-ray images and high-energy X-ray images; The X-ray low-energy image and the X-ray high-energy image are convolved respectively by using a convolution method with an expanded receptive field to generate a convolution feature map, and the convolution feature map is fused to obtain a preliminary fused feature map; The preliminary fusion feature map is enhanced through the multi-attention mechanism to strengthen the complementary information expression of the fusion feature representation, thereby obtaining enhanced features; The enhanced features are extracted through the dilated convolution layer to obtain the convolution output features, and then the convolution output features are subjected to channel adjustment and format restoration processing to generate the final fused image.

[0063] In one embodiment, a computer readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented: Acquire low-energy X-ray images and high-energy X-ray images; The X-ray low-energy image and the X-ray high-energy image are convolved respectively by using a convolution method with an expanded receptive field to generate a convolution feature map, and the convolution feature map is fused to obtain a preliminary fused feature map; The preliminary fusion feature map is enhanced through the multi-attention mechanism to strengthen the complementary information expression of the fusion feature representation, thereby obtaining enhanced features; The enhanced features are extracted through the dilated convolution layer to obtain the convolution output features, and then the convolution output features are subjected to channel adjustment and format restoration processing to generate the final fused image.

[0064] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0065] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0066] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A dual-energy X-ray image fusion method, characterized in that: The method comprises: Acquire low-energy X-ray images and high-energy X-ray images; The X-ray low-energy image and the X-ray high-energy image are respectively convolved by a convolution method that expands the receptive field to generate a convolution feature map, and the convolution feature map is fused to obtain a preliminary fused feature map; The preliminary fused feature map is enhanced by a multi-attention mechanism to strengthen the complementary information expression represented by the fused features, thereby obtaining enhanced features; The enhanced features are extracted through an expanded convolution layer to obtain convolution output features, and then channel adjustment and format restoration processing are performed on the convolution output features to generate a final fused image.

2. The dual-energy X-ray image fusion method according to claim 1, characterized in that: The convolution method of expanding the receptive field is used to convolve the X-ray low-energy image and the X-ray high-energy image respectively to generate a convolution feature map, and the convolution feature map is fused, specifically including: Using a 7×7 convolution kernel to perform a convolution operation on the X-ray low-energy image and the X-ray high-energy image to generate a convolution feature map; By pre-setting the fusion formula The convolution feature maps are fused, where F1 is the preliminary fusion feature map, I L is the X-ray low-energy image, I H is the X-ray high energy image.

3. The dual-energy X-ray image fusion method according to claim 1, characterized in that: The method of enhancing the preliminary fused feature map by a multi-attention mechanism to strengthen the complementary information expression represented by the fused feature, thereby obtaining enhanced features, specifically includes: The feature channels and spatial regions of the preliminary fused feature map are weighted respectively by an efficient channel attention mechanism and a spatial attention mechanism to enhance key region features in the preliminary fused feature map; The preliminary fusion feature map and the enhanced preliminary fusion feature map are residually connected to obtain the enhanced feature.

4. The dual-energy X-ray image fusion method according to claim 1, characterized in that: The step of extracting the enhanced features through the dilated convolution layer to obtain convolution output features, and then performing channel adjustment and format restoration processing on the convolution output features to generate a final fused image specifically includes: Using convolution kernels of different sizes to perform secondary convolution on the enhanced features to capture corresponding global features and obtain the convolution output features; By default output formula The convolution output features are restored to obtain the final fused image, where Out is the final fused image, and F out Output features for the convolution.

5. The dual-energy X-ray image fusion method according to claim 1, characterized in that: The dual-energy X-ray image fusion method also includes: Based on the preset loss function , optimize the image fusion process, where L P is the perceptual loss, L G is the gradient loss, L S is the structural loss, σ and γ are weight coefficients.

6. The dual-energy X-ray image fusion method according to claim 5, characterized in that: The preset loss function specifically includes: By perceiving loss , constraining the high-level semantic features between the final fused image, the X-ray low-energy image and the X-ray high-energy image, wherein, , , They are high-level features extracted by pre-trained convolutional neural networks, I F is the final fused image, α is a balance parameter used to adjust the contribution of the loss term; Through gradient loss , maintain the edge and detail information of the image, where is the gradient representation of the image, to obtain a maximum gradient between the X-ray low-energy image and the X-ray high-energy image; Through structural loss , constraining the local structural information of the final fused image, the X-ray low-energy image and the X-ray high-energy image to improve the structural similarity of the fused image, wherein SSIM(·) is the structural similarity index, The part with the strongest structure is selected for comparison between the X-ray low-energy image and the X-ray high-energy image.

7. A dual-energy X-ray image fusion device, characterized in that: The device comprises: An image acquisition module, used for acquiring X-ray low-energy images and X-ray high-energy images; A convolution operation module, used to respectively convolve the X-ray low-energy image and the X-ray high-energy image using a convolution method that expands the receptive field to generate a convolution feature map, and fuse the convolution feature map to obtain a preliminary fused feature map; A multi-attention mechanism module, used for enhancing the preliminary fused feature map through a multi-attention mechanism to strengthen the complementary information expression represented by the fused feature, thereby obtaining enhanced features; The dilated convolution module is used to extract the enhanced features through the dilated convolution layer to obtain convolution output features, and then perform channel adjustment and format restoration processing on the convolution output features to generate a final fused image.

8. The dual-energy X-ray image fusion device according to claim 7, characterized in that: The convolution operation module specifically includes: A convolution operation submodule, used for performing a convolution operation on the X-ray low-energy image and the X-ray high-energy image using a 7×7 convolution kernel to generate a convolution feature map; Feature fusion submodule, used to preset fusion formula The convolution feature maps are fused, where F1 is the preliminary fusion feature map, I L is the X-ray low-energy image, I H is the X-ray high energy image.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the dual-energy X-ray image fusion method according to any one of claims 1 to 6 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the dual-energy X-ray image fusion method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Infrared visible light image fusion method based on depth guided filtering

    CN117437143A

  • Infrared and visible light image fusion method and system based on segmentation task driving and improved lightweight Transform, and storage medium

    CN118172260A

  • Object detection method, dual-energy detector, system and storage medium

    CN118396972A

  • Medical image fusion method and device, equipment and storage medium

    CN118505529A

  • Infrared and visible light image fusion method based on attention and Transform

    CN119130827A

Cited By

  • Lightweight end-to-end infrared visible light adaptive image fusion method

    CN120746864A

  • A lightweight end-to-end infrared-visible light adaptive image fusion method

    CN120746864B