A medical image fusion system based on a semantic guidance network
Through the medical image fusion system based on semantic guidance network, the problems of edge artifacts and time redundancy in multimodal medical image fusion are solved, the clarity and semantic information are improved, and the fusion process is optimized.
Patent Information
- Application Number
- CN202211070837.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-02
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-09-02
AI Technical Summary
The prior art has problems such as edge artifacts, semantic information ignorance, network scale redundancy and excessive fusion time in multimodal medical image fusion.
A medical image fusion system based on semantic guidance network is adopted, including edge enhancement module, region mask module and global refinement module, and the feature extraction and fusion process is optimized through gradient filters, region mask generators and loss functions.
It reduces edge artifacts, improves clarity and visual effects of fusion results, saves fusion time, and enhances the transmission of semantic information.
Smart Images

Figure CN115423731B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image fusion, and particularly to a medical image fusion system based on a semantic guidance network. Background Art
[0002] In the development of medical images, multimodal medical imaging has been widely applied to clinical diagnosis, medical research, and surgical navigation. Due to different imaging technologies, the information highlighted by different medical images varies. Generally, they can be roughly divided into two categories: structural medical images and functional medical images. For example, magnetic resonance (MRI) images, as a typical structural medical image, can highlight soft tissue information; computed tomography (CT) images can clearly provide structural contours and brain anatomical information with high resolution. However, structural medical images are not sensitive to functional information in human metabolism. In the field of functional medical images, positron emission tomography (PET) images play an important role, which can characterize metabolism, blood flow, and some tumor information in the brain tissue. In addition, single photon emission computed tomography (SPECT) images can highlight tissue damage and organ information. Nevertheless, functional medical images still have the drawback of low resolution and cannot accurately display structural information.
[0003] Deep learning-based and traditional fusion methods have been widely applied to existing multimodal image fusion methods. Their common goal is to extract actual features from different single-source images and generate a fused image through a designed fusion strategy or network model. In most traditional methods, fusion rules based on spatial domain and transform domain transformations are usually adopted, which generate a new fused image by transforming specific regions and then reconstructing them together. However, in these methods, it is inevitable to manually calculate complex fusion strategies, which reduces the efficiency of the fusion process. In addition, by using the same decomposition operation to process source images of different modalities, artifacts may appear in the fusion result.
[0004] In recent years, to improve the drawbacks of traditional methods, researchers have introduced deep learning-based methods to perform multimodal image fusion tasks. They use an end-to-end model to avoid the complexity of manually designed fusion rules. In addition, different modules in the architecture can correctly extract their unique features from multiple single-modal images. However, there are still certain limitations: (1) In multimodal fusion tasks, semantic information is often ignored, so halos may appear, reducing the quality of the generated result; (2) Some existing deep learning-based methods increase the network scale to improve the quality of the fused image, which leads to a large amount of redundant calculations and a long running time. (3) Due to the different prominent features in each source image, it is not easy to achieve the performance of certain texture details in the fused image. Summary of the Invention
[0005] The object of the present invention is to provide a medical image fusion system based on a semantic guidance network, which can reduce the appearance of edge artifacts and obtain a real fusion result, saving the fusion time.
[0006] To achieve the above object, the present application proposes a medical image fusion system based on a semantic guidance network, including:
[0007] An edge enhancement module that extracts a shallow feature map ε from the source image using two 3×3 convolutions with a dense connection pattern and one 1×1 convolution Conv ; learns the gradient information ε through a new gradient filter G ; then combines the shallow feature map ε Conv and the gradient information ε G according to element-wise addition to obtain the final edge-enhanced feature ε f ;
[0008] A region masking module that learns and optimizes the edge-enhanced feature ε through a region mask generator RMG and multiple region mask convolutions RMC f , to obtain a feature
[0009] A global refinement module that imports the feature into two 3×3 convolutions to obtain a globally masked optimized feature Then, the edge-enhanced feature ε from the MRI branch f is associated with the globally masked optimized feature ; finally, the feature ε f and are integrated through element-wise addition to generate global refinement information, and two 1×1 convolutions are used to eliminate the difference in the channel dimension; the refinement process of this module is quantified as:
[0010]
[0011] where ⊙ represents the concatenation operation.
[0012] Further, the new gradient filter deploys 3×3 convolutions with Sobel operators in the horizontal and vertical directions respectively to obtain the horizontal gradient information ε G , and the gradient information ε G is input into a 1×1 convolution to eliminate the difference in the channel dimension.
[0013] Further, in the region masking module, the edge-enhanced feature ε f is first fed into a 3×3 convolution with an LReLU layer and an average pooling layer. After being modified by another 3×3 convolution in the LReLU layer, the extracted feature map is input into a transposed convolution to obtain an upsampled feature
[0014] Furthermore, in order to achieve self - adjustment of the spatial mask, the region mask module estimates a One - hot distribution through the Gumbel softmax distribution, specifically as follows:
[0015]
[0016] where g and w represent the factors in the vertical and horizontal directions; G sp is the intermediate noise tensor in Gumbel softmax, and all elements follow the Gumbel distribution; when the hyperparameter θ approaches infinity, the feature map can execute a uniform distribution; when the hyperparameter θ approaches 0, a One - hot distribution will appear.
[0017] Furthermore, in order for the region mask module to mark the "redundant" regions through the channel mask M c the edge - enhanced feature ε f is randomly converted into a sampled feature on a Gaussian distribution before feeding the Gumbel softmax where M c is defined as:
[0018]
[0019] where c is the number of channels, and G c is the intermediate noise tensor.
[0020] Furthermore, in the training stage, the spatial and channel masks are used to mark the "important" and "redundant" regions in the region mask generator RMG respectively; in the testing stage, an Argmax layer is introduced to replace the Gumbel softmax to obtain the spatial mask and the channel mask, and the feature optimized by the mask is obtained
[0021] Even further, skip connections are used to connect the edge - enhancement module on the CT / PET / SPECT branches to the 3×3 convolution behind the global refinement module.
[0022] Even further, the medical image fusion system is trained through a loss function, and the loss function includes an edge loss function L E , a structural similarity loss function L SSIM and a semantic loss; thus the total loss function L total is defined as:
[0023] L total = L E + αL SSIM + βL S
[0024] where α and β are hyperparameters for balancing L total ;
[0025] In the training stage, the edge loss function is divided into two parts:
[0026] L E = L c + γL g
[0027] where, L c and L g represent content and gradient losses respectively; γ is a hyperparameter for controlling the magnitude of L g ; in the content loss L c , the L1 norm is used to measure the difference between the generated output image and the source image; the L c is defined as:
[0028]
[0029] where, H F and W F represent the height and width of I F ; max(*) and ||*||1 represent the maximum selection strategy and the L1 norm respectively; I F is the fused image, and I A , I B is the source image.
[0030] Furthermore, the gradient value in the pixel domain is measured by the gradient loss, specifically:
[0031]
[0032] where, represents the Sobel operator for calculating the gradient value.
[0033] As a further step, the structural similarity loss function L SSIM measures the structural difference through the structural similarity index SSIM, which includes three kinds of information: luminance, structure, and contrast; it is quantified as:
[0034] L SSIM = (1 - SSIM(I F , I A )) + (1 - SSIM(I F , I B ))
[0035] where, SSIM(I F , I * ) is defined as:
[0036]
[0037] Among them, I * represents the source image I A or I B ; μ and σ respectively represent the mean and standard deviation; C1, C2, and C3 are constants for maintaining stability.
[0038] As a further step, the semantic loss includes the main semantic loss L main and the secondary semantic loss L sub :
[0039] L S = L main + δL sub
[0040]
[0041]
[0042] Among them, the parameter δ can maintain the stability of L S ; I O represents the One-hot distribution generated by the extracted semantic features; S main and S sub respectively represent the main semantic information and the auxiliary semantic information.
[0043] The above technical solutions adopted by the present invention, compared with the prior art, have the following advantages: This method is used to perform medical image fusion tasks. Through the edge enhancement module and the corresponding edge loss function, the edge texture of the fusion result is clearer. The region mask module can divide different regions of the original image to focus on feature extraction, while the global refinement module can optimize the overall visual effect of the fusion result. In addition, more semantic information is transmitted through the semantic loss function to improve the quality of the fused image. This system can generate vivid fusion results in visual perception and also ensure the quantitative indicators. Brief Description of the Drawings
[0044] Figure 1 is the structural block diagram of the medical image fusion system based on the semantic guidance network;
[0045] Figure 2 is the structural diagram of the edge enhancement module;
[0046] Figure 3 is the structural diagram of the region mask module;
[0047] Figure 4 is the structural diagram of the global refinement module;
[0048] Figure 5 is the qualitative comparison diagram between this method and other fusion methods on the Harvard medical image dataset;
[0049] Figure 6 It is a quantitative comparison graph between this method and other fusion methods on the SICE dataset. Specific implementation manners
[0050] In order to make the objectives, technical solutions and advantages of this application clearer, the following further elaborates on this application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application, that is, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments.
[0051] Therefore, the following detailed description of the embodiments of this application provided in the drawings is not intended to limit the scope of this application claimed, but merely represents selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative efforts fall within the scope of protection of this application.
[0052] Embodiment 1
[0053] As Figure 1 shown, this application provides a medical image fusion system based on a semantic guidance network, specifically including:
[0054] An edge enhancement module, as Figure 2 shown, in the main layer structure, two 3×3 convolutions with a dense connection pattern and a 1×1 convolution are used to extract a shallow feature map ε from the source image Conv ; in the residual structure, a new gradient filter is used to learn gradient information ε G ; then the shallow feature map ε Conv and the gradient information ε G are merged according to element-wise addition to obtain the final edge enhancement feature ε f ;
[0055] Specifically, the new gradient filter deploys 3×3 convolutions with Sobel operators in the horizontal and vertical directions respectively to obtain the horizontal gradient gradient ε G ; the gradient information ε G is input into a 1×1 convolution to eliminate the difference in the channel dimension.
[0056] A region mask module, as Figure 3 shown, learns and optimizes the edge enhancement feature ε f through a region mask generator RMG and multiple region mask convolutions RMC to obtain the feature
[0057] Specifically, in the training stage, spatial and channel masks are used to label the "important" and "redundant" regions in the RMG. The edge-enhanced feature ε f is first fed into a 3×3 convolution with an LReLU layer and an average pooling layer; after the LReLU layer is modified by another 3×3 convolution, the extracted feature map is input into a transposed convolution to obtain an upsampled feature In this way, the learnable parameters of the transposed convolution can be updated through backpropagation, thereby effectively performing the upsampling operation. To achieve self-adjustment of the spatial mask, the Gumbel softmax distribution is used to estimate a One-hot distribution, and its formula is:
[0058]
[0059] where h and w represent the factors in the vertical and horizontal directions; G sp is the intermediate noise tensor in Gumbel softmax, and all elements follow the Gumbel distribution. When the hyperparameter θ approaches infinity, the feature map can perform a uniform distribution; when the hyperparameter θ approaches 0, a One-hot distribution will appear. In this embodiment, due to the range limitation, the hyperparameter θ is set to 0.5 for balance. To mark the "redundant" regions through the channel mask M c before feeding Gumbel softmax, ε f is randomly converted on a Gaussian distribution into where M c can be defined as:
[0060]
[0061] where c represents the number of channels.
[0062] In the testing stage, an Argmax layer is introduced to replace Gumbel softmax to obtain the spatial mask and the channel mask. The Argmax layer can return the corresponding index of the maximum value in the feature, and finally obtain the feature optimized by the mask
[0063] The global refinement module, as Figure 4 shown, imports the said feature into two 3×3 convolutions to obtain the globally mask-optimized feature Then, the edge-enhanced feature ε f from the MRI branch is associated with the globally mask-optimized feature to retain the information obtained previously and thus improve the performance of the fusion result; finally, the feature ε f is integrated with To generate global refinement information and use two 1×1 convolutions to eliminate differences in the channel dimension; the refinement process of this module is quantified as:
[0064]
[0065] Among them, ⊙ represents the concatenation operation.
[0066] In addition, in order to make the fused result colorful and retain more functional details, skip connections are used to connect the edge enhancement module on the CT / PET / SPECT branch to the 3×3 convolution behind the global refinement module.
[0067] To ensure the quality of the fused image by retaining more meaningful extracted information, the designed loss function consists of three parts, including the edge loss function L E , the structural similarity loss function L SSIM and the semantic loss L S ; therefore, the total loss function L total is defined as:
[0068] L total = L E + αL SSIM + βL S
[0069] Among them, α and β are hyperparameters that balance L total .
[0070] In the training stage, the edge loss function can guide the source image to generate a fused image that can highlight content awareness and gradient information. Therefore, the edge loss function is divided into two parts:
[0071] L E = L c + γL g
[0072] Among them, L c and L g represent content and gradient losses respectively. γ is a hyperparameter that controls the magnitude of L g . In the content loss L c , the L1 norm is used to measure the difference between the generated output image and the source image. L c is defined as:
[0073]
[0074] Among them, H F and W F represent I FThe height and width. max(*) and ||*||1 represent the maximum selection strategy and the L1 norm respectively. According to the content loss, pixel-level information can be transmitted into the system for image fusion.
[0075] In order to retain more edge textures while transmitting content information. Therefore, a gradient loss is proposed to measure the gradient value in the pixel domain, which can be:
[0076]
[0077] where represents the Sobel operator for calculating the gradient value. Since negative gradients are not available, the absolute value operation |*| is introduced to solve this problem.
[0078] L SSIM The structural difference can be measured by the structural similarity index (SSIM), which includes three kinds of information: luminance, structure, and contrast. It can be quantified as:
[0079] L SSIM =(1 - SSIM(I F , I A ))+(1 - SSIM(I F , I B )
[0080] where SSIM(I F , I * ) can be defined as:
[0081]
[0082] where I * represents the source image I A or I B . μ and σ represent the mean and standard deviation respectively. C1, C2, and C3 are constants to maintain stability.
[0083] In addition, a semantic loss is introduced to feedback the semantic information in the source image into the fusion result. The semantic loss is divided into the main semantic loss L main and the secondary semantic loss L sub :
[0084] L S = L mai n + δL sub
[0085]
[0086]
[0087] where δ can keep L S stable. IO represents the One-hot distribution generated by the extracted semantic features. S main and S sub respectively represent the main semantic information and the auxiliary semantic information.
[0088] In this embodiment, test image sequences are selected from the Harvard Medical Image Dataset and compared with nine state-of-the-art medical image fusion methods. For a full comparison, Figure 5 shows the overall effect and local feature details respectively; the first group makes a comparison of the MRI-CT results. A fusion result with obvious contrast and the coexistence of bone information and soft tissue information can be obtained well. The result is in an optimal state compared with other methods. The second and third groups are the comparison results between MRI-PET and MRI-SPECT respectively. It can be seen that when performing these two types of tasks, on the premise of ensuring edge information, no color distortion and edge artifacts are generated, which plays a crucial role for researchers and doctors in pathological diagnosis and condition analysis of patients.
[0089] In addition to subjective qualitative analysis, objective quantitative indicators such as SSIM, SD, MI, VIF, SCD, and Q ab / f are used to evaluate the performance of the fusion results. 21 pairs of MRI-CT images, 42 pairs of MRI-PET images, and 73 pairs of MRI-SPECT images are used as the test sets to complete different medical image fusion tasks, and the quantitative results are as Figure 6 shown. Obviously, the obtained results reach a relatively high level among the six evaluation indicators. From the distribution of the scatter plot, it can be seen that most of the indicators of the test results are higher than those of other methods. The higher the SSIM value, the better the visual effect can be provided and the more similar it is to the source image. SD can represent the degree of deviation between the image pixel values and the average pixel value. Generally, the larger the MI value, the better the performance of the fusion result, and the more information is interacted between the fused image and the source image. A higher VIF can indicate high information fidelity of the fused image. SCD and Q ab / f are used to measure the differences and edge information between the source image and the fused image in each element.
[0090] The foregoing description of the specific exemplary embodiments of the present invention is for the purposes of illustration and exemplification. These descriptions are not intended to limit the invention to the precise forms disclosed, and obviously, many changes and variations are possible in light of the above teachings. The purpose of selecting and describing the exemplary embodiments is to explain the specific principles of the invention and its practical applications, so that those skilled in the art can implement and utilize various different exemplary embodiments of the invention, as well as various different selections and changes. The scope of the present invention is intended to be defined by the claims and their equivalents.
Claims
1. A medical image fusion system based on a semantic guidance network, characterized in that, Including: Edge enhancement module, which extracts the shallow feature map ε from the source image using two 3×3 convolutions with a dense connection pattern and a 1×1 convolution Conv ; learns the gradient information ε through a new gradient filter G ; then combines the shallow feature map ε Conv and the gradient information ε G by element-wise addition to obtain the final edge-enhanced feature ε f ; Region masking module, which learns and optimizes the edge enhancement feature ε through a region mask generator RMG and multiple region mask convolutions RMC f , to obtain the feature Global refinement module, which imports the features into two 3×3 convolutions to obtain globally masked optimized features Then, the edge-enhanced feature ε from the MRI branch f is associated with the globally masked optimized feature ; finally, the features ε f and are integrated by element-wise addition to generate global refinement information, and two 1×1 convolutions are used to eliminate the difference in channel dimensions; the refinement process of this module is quantified as: Among them, ⊙ represents the connection operation; Training a medical image fusion system using a loss function, the loss function including an edge loss function LE, a structural similarity loss function L SS I M and a semantic loss LS; thus the total loss function L total is defined as: L total = L E + αL SSIM + βL S where α and β are hyperparameters that balance L total ; In the training stage, the edge loss function is divided into two parts: L E = L c + γL g Among them, L c and L g represent content and gradient loss respectively; γ is a hyperparameter that controls the magnitude of L g ; in the content loss L c , the L1 norm is used to measure the difference between the generated output image and the source image; the L c is defined as: Among them, H F and W F represent the F height and width of I; max(*) and ||*||1 represent the maximum selection strategy and the L1 norm respectively; I F is the fused image, and I A , I B are the source images. The gradient loss is used to measure the gradient value in the pixel domain, specifically: Among them, represents the Sobel operator for calculating the gradient value.
2. The medical image fusion system based on a semantic guidance network according to claim 1, wherein The new gradient filter deploys 3×3 convolutions with Sobel operators in the horizontal and vertical directions respectively to obtain horizontal gradient information ε G , and the gradient information ε G is input into a 1×1 convolution.
3. The medical image fusion system based on a semantic guidance network according to claim 1, wherein In the regional mask module, the edge enhancement feature ε f is first fed into a 3×3 convolution with an LReLU layer and an average pooling layer. After being modified by another 3×3 convolution in the LReLU layer, the extracted feature map is input into a transposed convolution to obtain an upsampled feature 4. The medical image fusion system based on a semantic guidance network according to claim 3, wherein In order to achieve spatial mask self-adjustment, the region mask module estimates a One-hot distribution through the Gumbel softmax distribution, specifically: where h and w represent the factors in the vertical and horizontal directions; G sp is the intermediate noise tensor in Gumbel softmax, and all elements follow the Gumbel distribution; when the hyperparameter θ approaches infinity, the feature map can perform a uniform distribution; when the hyperparameter θ approaches 0, a One-hot distribution will occur.
5. The medical image fusion system based on a semantic guidance network according to claim 4, characterized in that The area masking module marks the "redundant" area through the channel mask M c Before feeding the Gumbel softmax, randomly convert the edge-enhanced feature ε f into a sampled feature on the Gaussian distribution where M c is defined as: where c is the number of channels, and G c is the intermediate noise tensor.
6. The medical image fusion system based on a semantic guidance network according to claim 1, wherein, In the training stage, spatial and channel masks are used to label the "important" and "redundant" regions in the region mask generator RMG respectively; In the testing phase, the Argmax layer was introduced to replace the Gumbel softmax to obtain the spatial mask and channel mask, and the features optimized by the mask were obtained.
7. The medical image fusion system based on a semantic guidance network according to claim 1, wherein The edge enhancement modules on the CT, PET, and SPECT branches are connected to the 3×3 convolution behind the global refinement module using skip connections.
8. The medical image fusion system based on a semantic guidance network according to claim 1, wherein Structural similarity loss function L SSIM The structural difference is measured by the structural similarity index SSIM, which includes three kinds of information: brightness, structure, and contrast; and it is quantified as: L SSIM = (1 - SSIM(I F , I A )) + (1 - SSIM(I F , I B )) Among them, SSIM(I F ,I * ) is defined as: where I * represents the source image I A or I B ; μ and σ represent the mean and standard deviation respectively; C1, C2, and C3 are constants for maintaining stability.
9. The medical image fusion system based on a semantic guidance network according to claim 1, wherein The semantic loss includes a primary semantic loss L main and a secondary semantic loss L sub : L S = L main + δL sub Among them, the parameter δ can keep L S stable; I O represents the One-hot distribution generated by the extracted semantic features; S main and S sub represent the main semantic information and the auxiliary semantic information respectively.