A three-dimensional medical image fusion method based on deep learning

By using a deep learning-based approach, combining in-modal and extramodal preprocessing with a 3D generative adversarial network, high-quality 3D medical fusion images are generated, solving the problem that 2D fusion images cannot reflect 3D positions and improving the accuracy and efficiency of diagnosis.

CN119863377BActive Publication Date: 2025-10-24XIDIAN UNIV HANGZHOU RES INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411938497.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-10-24
Estimated Expiration
2044-12-26

AI Technical Summary

Technical Problem

Existing multimodal medical image fusion technology generates two-dimensional fused images that are difficult to accurately reflect the three-dimensional position of the human body, with limited information content, resulting in low interpretability and a high risk of errors in diagnosis.

Method used

A deep learning-based approach is employed, utilizing intra-modal preprocessing, inter-modal preprocessing, deep learning network models, and an improved 3D generative adversarial network to generate 3D medical fusion images, thereby improving image quality and matching accuracy. Furthermore, convolutional neural networks and transformer networks are combined to extract local and global information, resulting in high-quality 3D medical fusion images.

Benefits of technology

It accelerates the medical diagnosis process, improves medical decision-making capabilities, enhances the visualization of images and the accuracy and real-time nature of diagnosis, provides richer three-dimensional perspectives, and improves the accuracy and robustness of medical image analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119863377B_ABST
    Figure CN119863377B_ABST
Patent Text Reader

Abstract

The application relates to a three-dimensional medical image fusion method, device and equipment based on deep learning and a medium. The method comprises the following steps: acquiring two-dimensional medical source images corresponding to at least two modes, and performing intra-mode preprocessing on the two-dimensional medical source images to obtain a single-mode two-dimensional medical original image group corresponding to each mode; performing inter-mode preprocessing on the single-mode two-dimensional medical original image group to obtain a multi-mode two-dimensional medical original image group; inputting the multi-mode two-dimensional medical original image group into a deep learning network model to generate a two-dimensional medical fusion image group; and inputting the two-dimensional medical fusion image group into an improved three-dimensional generative adversarial network taking the two-dimensional medical original image group as a real image to generate a three-dimensional medical fusion image. The method can enrich the information content of a medical fusion picture display, improve the explainability of a medical fusion picture fusion process, and improve the efficiency and accuracy of medical diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of medical image data processing, and particularly relates to a three-dimensional medical image fusion method based on deep learning. BACKGROUND

[0002] Multi-modal medical fusion images are important auxiliary tools for doctors to diagnose and treat various diseases, and play an indispensable role in auxiliary diagnosis, treatment guidance and progress detection. In the past few decades, a large number of medical image fusion methods have been proposed in the prior art. Traditional fusion methods usually use specific transformation algorithms to decompose images, then use a designed fusion strategy for fusion, and finally use inverse transformation to obtain the final fusion image.

[0003] However, in the existing multi-modal medical image fusion technology, the generated fusion image is basically displayed in two dimensions. Since the human body is three-dimensional, it is difficult to accurately reflect the three-dimensional position coordinates of the patient's lesion in a two-dimensional image. Moreover, the two-dimensional fusion image generated by the existing multi-modal medical fusion image technology has less information content and lower interpretability, which can easily lead to diagnostic errors. SUMMARY

[0004] Therefore, it is necessary to provide a three-dimensional medical image fusion method based on deep learning, which can accelerate the medical diagnosis process and improve the medical decision-making ability.

[0005] The application provides a three-dimensional medical image fusion method based on deep learning, which comprises the following steps:

[0006] Obtaining at least two modal corresponding two-dimensional medical source images, and performing intra-modal preprocessing on the two-dimensional medical source images to obtain a single modal two-dimensional medical original image group corresponding to each modality, wherein the intra-modal preprocessing is used to improve the image quality of the two-dimensional medical source image corresponding to a single modality;

[0007] Performing inter-modal preprocessing on the single modal two-dimensional medical original image group to obtain a multi-modal two-dimensional medical original image group, wherein the inter-modal preprocessing is used to align the two-dimensional medical source images corresponding to different modalities;

[0008] Inputting the multi-modal two-dimensional medical original image group into a deep learning network model to generate a two-dimensional medical fusion image group, wherein the two-dimensional medical fusion image group is used to represent a two-dimensional medical fusion image in which the features of the two-dimensional medical source images of different modalities corresponding to the same spatial pose in the multi-modal two-dimensional medical original image group are fused, and the deep learning network model is obtained by combining a convolutional neural network and a transformer network;

[0009] The two-dimensional medical fusion image group is input into the improved three-dimensional generative adversarial network taking the two-dimensional medical original image group as a real image, and a three-dimensional medical fusion image is generated.

[0010] The three-dimensional medical image fusion method based on deep learning can accelerate the medical diagnosis process and improve the medical decision-making ability by fusing multi-modal medical images to obtain a multi-modal three-dimensional medical fusion image. The intra-modal preprocessing and inter-modal preprocessing can effectively improve the image quality and improve the spatial matching degree of the multi-modal two-dimensional medical original image group.

[0011] Further, the inter-modal preprocessing pairs each single-modal medical original image group to obtain a multi-modal two-dimensional medical original image group, which provides an efficient data structure for further image fusion and three-dimensional image generation. By combining the convolutional neural network (CNN) branch and the transformer network (Transformer) branch to extract local and global information of the image, the visual representation capability can be enhanced, and the clarity of the structure and texture of the fusion image can be improved.

[0012] Further, the improved three-dimensional generative adversarial network generates a two-dimensional medical fusion image group by taking a two-dimensional medical fusion image group as an input image and a two-dimensional medical original image group as a real image, thereby improving the quality and efficiency of three-dimensional medical image fusion. BRIEF DESCRIPTION OF DRAWINGS

[0013] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the drawings needed to be used in the embodiments or related art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0014] Figure 1 The schematic diagram of the application environment of the scheme provided by an embodiment of the present application;

[0015] Figure 2 The flowchart of a three-dimensional medical image fusion method based on deep learning provided by an embodiment of the present application;

[0016] Figure 3 The flowchart of a method for generating a two-dimensional medical fusion image group provided by an embodiment of the present application;

[0017] Figure 4 The structural schematic diagram of a deep learning network model provided by an embodiment of the present application;

[0018] Figure 5 The structural schematic diagram of another deep learning network model provided by an embodiment of the present application;

[0019] Figure 6 A structural diagram of an improved three-dimensional generative adversarial network is provided for an embodiment of the present application.

[0020] Figure 7 A flowchart of a method for generating a multi-modal two-dimensional medical original image group is provided for an embodiment of the present application.

[0021] Figure 8 A flowchart of another deep learning-based three-dimensional medical image fusion method is provided for an embodiment of the present application.

[0022] Figure 9 A structural diagram of a deep learning-based three-dimensional medical image fusion device is provided for an embodiment of the present application. DETAILED DESCRIPTION

[0023] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.

[0024] The deep learning-based three-dimensional medical image fusion method provided by the embodiments of the present application can be applied in an application environment as shown in Figure 1 . Among them, the computing platform 101 communicates with the medical image acquisition device 102 through a communication channel. The medical image acquisition device 102 can acquire data required by the computing platform 101 for processing. The computing platform 101 can be integrated on the medical image acquisition device 102 and / or the medical service platform carried by the medical image acquisition device 102, or can be placed on the cloud or other terminal medium. The computing platform 101 can generate medical fusion images for disease diagnosis and detection based on the data acquired by the medical image acquisition device 102.

[0025] Illustratively, the computing platform 101 can be, but is not limited to, a high-performance computing cluster, an edge computing platform, and a hybrid computing platform. The medical image acquisition device 102 can include, but is not limited to, an MMIF device based on electromagnetic energy imaging technology to form static images, an X-ray computed tomography (CT) device, a single photon emission computed tomography (SPECT) device, a positron emission tomography (PET) device, and a magnetic resonance imaging (MRI) device.

[0026] In an exemplary embodiment, as shown in Figure 2 , a deep learning-based three-dimensional medical image fusion method is provided. Taking the computing platform 101 in Figure 1 as an example for illustration, the method includes:

[0027] In step S201, two-dimensional medical source images corresponding to at least two modalities are acquired, and the two-dimensional medical source images are preprocessed within the modalities to obtain a single-modality two-dimensional medical original image group corresponding to each modality.

[0028] Specifically, the computing platform 101 can acquire two-dimensional medical source images corresponding to at least two modalities through a medical image acquisition device. The computing platform 101 can preprocess the acquired two-dimensional medical source images corresponding to at least two modalities within the modalities to obtain a single-modality two-dimensional medical original image group corresponding to each modality. The computing platform 101 can improve the image quality of the two-dimensional medical source images corresponding to a single modality in the at least two modalities through the intra-modality preprocessing.

[0029] Illustratively, the two-dimensional medical source images can include, but are not limited to, CT source images, SPECT source images, PET source images, and MRI source images. The intra-modality preprocessing can include, but is not limited to, denoising processing, contrast enhancement, and distortion correction.

[0030] In step S202, the single-modality two-dimensional medical original image group is preprocessed between modalities to obtain a multi-modality two-dimensional medical original image group.

[0031] Specifically, the computing platform 101 can preprocess the single-modality two-dimensional medical original image group obtained after the intra-modality preprocessing between modalities to obtain a multi-modality two-dimensional medical original image group. The computing platform 101 can spatially align the two-dimensional medical source images corresponding to different modalities through the inter-modality preprocessing.

[0032] Illustratively, the inter-modality preprocessing can include, but is not limited to, image resampling processing, image normalization processing, and image spatial registration processing.

[0033] In step S203, the multi-modality two-dimensional medical original image group is input into a deep learning network model to generate a two-dimensional medical fusion image group.

[0034] Specifically, the computing platform 101 can input the multi-modality two-dimensional medical original image group into a deep learning network model to generate a two-dimensional medical fusion image, and then construct a two-dimensional medical fusion image group based on the spatial pose corresponding to the two-dimensional medical fusion image. The two-dimensional medical fusion image group is used to represent a two-dimensional medical fusion image that fuses the features of the two-dimensional medical source images of different modalities corresponding to the same spatial pose in the multi-modality two-dimensional medical original image group. The deep learning network model is obtained by combining a convolutional neural network and a transformer network.

[0035] Exemplarily, the convolutional neural network and the transformer network in the deep learning network model are multi-stage structures, and the computing platform 101 can fuse the outputs of the convolutional neural network and the transformer network at the same stage as the output of the deep learning network model at the stage.

[0036] Further, the computing platform 101 can determine the number of layers of the multi-layer structure of the deep learning network model based on the down-sampling multiple of each stage in the convolutional neural network and the minimum resolution of the multi-modal two-dimensional medical original image group.

[0037] Step S204, inputting the two-dimensional medical fusion image group into the improved three-dimensional generative adversarial network taking the two-dimensional medical original image group as real images to generate a three-dimensional medical fusion image.

[0038] Exemplarily, the computing platform 101 can generate a three-dimensional medical fusion image in a supervised manner through the improved three-dimensional generative adversarial network taking the two-dimensional medical fusion image group as input images and the two-dimensional medical original image group as real images.

[0039] In the above deep learning-based three-dimensional medical image fusion method, the quality of the single-modal two-dimensional medical source image can be improved through intra-modal preprocessing, thereby ensuring the quality and consistency of the multi-modal two-dimensional medical original image group input into the deep learning model. The different modal images can be accurately aligned in space through inter-modal preprocessing, thereby laying a foundation for subsequent multi-modal data fusion. The features of the two-dimensional medical original images of different modalities can be fused by inputting the multi-modal two-dimensional medical original image group into the deep learning network model combining the convolutional neural network and the transformer network, thereby improving the accuracy and robustness of medical image analysis. The three-dimensional medical fusion image can be generated by inputting the two-dimensional medical fusion image group into the improved three-dimensional generative adversarial network, thereby enhancing the visualization effect of the medical fusion image, providing a comprehensive and accurate three-dimensional perspective for doctors, and improving the accuracy and real-time performance of medical diagnosis.

[0040] Further, the above deep learning-based three-dimensional medical image fusion method can provide a powerful tool for cross-modal medical image analysis research and promote the development of the field of medical image analysis.

[0041] In one exemplary embodiment, as shown in Figure 3 and Figure 4 inputting the multi-modal two-dimensional medical original image group into the deep learning network model to generate a two-dimensional medical fusion image group, comprising:

[0042] Step 301, extracting local feature maps of each image in the multi-modal two-dimensional medical original image group based on the convolutional neural network branch in the deep learning network model.

[0043] Specifically, the convolutional neural network branch in the deep learning network model can be divided into three stages. The first stage can be a convolutional neural network initial module including a convolutional layer with a convolution kernel size of 7 and a convolution step size of 2, a max-pooling layer, a residual convolutional module, and an average-pooling layer. The computing platform can implement 4 times down-sampling based on the convolutional neural network initial module, wherein the max-pooling layer and the average-pooling layer respectively implement 2 times down-sampling. The second stage and the third stage can be a convolutional neural network standard module including a residual convolutional module and an average-pooling layer. The computing platform can implement 2 times down-sampling based on the average-pooling layer in the convolutional neural network initial module.

[0044] Illustratively, when the minimum resolution of the multi-modal two-dimensional medical original image group exceeds 640x640, a trained third convolutional neural network standard module can be added as a fourth stage in the third stage of the convolutional neural network branch.

[0045] Step 302, extracting global feature maps of each image in the multi-modal two-dimensional medical original image group based on the transformer network branch in the deep learning network model.

[0046] Specifically, the number of stages of the transformer network branch in the deep learning network model can be the same as the number of stages of the convolutional neural network branch. The side length of the image block corresponding to the global feature map output by the transformer network branch in the same stage of the deep learning network model can be the same as the down-sampling multiple of the convolutional neural network branch. The transformer network branch in the deep learning network model is obtained by replacing the multi-head attention module in the original transformer network with a spatially reduced multi-head attention module.

[0047] Illustratively, the transformer network branch in the deep learning network model can be divided into three stages, each stage including a patch embedding module and an encoder. The transformer network branch can control the size of the image block corresponding to the output global feature map through the patch embedding module, and the encoder can include a normalization layer, a spatially reduced multi-head attention module, and a multi-layer perceptron.

[0048] Illustratively, the side length of the image block corresponding to the global feature map output by the first stage of the transformer network branch can be 4; the side length of the image block corresponding to the global feature map output by the second stage can be 8; and the side length of the image block corresponding to the global feature map output by the third stage can be 16.

[0049] Step 303, coupling the local feature maps and the global feature maps to generate coupled feature maps, and inputting the coupled feature maps into a cross-scale attention module to generate multi-scale feature maps.

[0050] Specifically, the computing platform can perform feature coupling on the local feature map output by the convolutional neural network branch at the same stage and the global feature map output by the transformer network branch at the same stage through a feature fusion module in the deep learning network model, and the computing platform can add the global feature map output by the transformer network branch at the same stage and the matched local feature map output by the convolutional neural network branch at the same stage to obtain a coupled feature map.

[0051] Further, the computing platform can input the coupled feature maps at each stage into the cross-scale attention module to generate multi-scale feature maps.

[0052] At step 304, a two-dimensional medical fusion image is generated based on the multi-scale feature maps of the images corresponding to different modalities that need to be fused in the multi-modal two-dimensional medical original image group, and a two-dimensional medical fusion image group is constructed.

[0053] Specifically, the computing platform can input the coupled feature map output at the first stage and the multi-scale feature maps output by each cross-scale attention module into a channel concatenation layer to obtain a concatenated feature map. Then, the concatenated feature maps of the images corresponding to different modalities that need to be fused in the multi-modal two-dimensional medical original image group are input into a feature decoding network after weighted summation, to generate a two-dimensional medical fusion image, and after spatial registration and fine-tuning optimization, a two-dimensional medical fusion image group is constructed.

[0054] In the above deep learning-based three-dimensional medical image fusion method, the convolutional neural network branch can effectively extract local key information such as details, edges, and other local information of the lesion site, and at the same time, the transformer network branch can obtain macro features such as tissue distribution and overall structure morphology reflected on the overall level of the image, thereby realizing comprehensive mining of medical image features and improving the explainability of medical image fusion.

[0055] Further, by coupling and cross-scale fusing the extracted local feature map and global feature map, features of different dimensions and different levels can be supplemented and fused with each other, richer and more comprehensive image information can be obtained, and multi-dimensional basis for accurately judging the lesion condition can be provided.

[0056] In an exemplary embodiment, please refer to Figure 5 Based on the multi-scale feature maps of the images corresponding to different modalities that need to be fused in the multi-modal two-dimensional medical original image group, a two-dimensional medical fusion image is generated, including:

[0057] The multi-scale feature maps of the images corresponding to different modalities that need to be fused in the multi-modal two-dimensional medical original image group are input into a feature map fusion module to generate a fused feature map, and the feature map fusion module includes a global attention mechanism submodule and a fusion submodule.

[0058] The fusion feature map is input into a feature decoding network to generate a two-dimensional medical fusion image.

[0059] The two-dimensional medical fusion image is subjected to spatial registration to construct a two-dimensional medical fusion image group, which is used to represent the spatial pose of the two-dimensional medical fusion image.

[0060] In the above three-dimensional medical image fusion method based on deep learning, the global attention mechanism submodule in the feature map fusion module can pay attention to the correlation and importance distribution between the features of different modal images from a global perspective. By weighting the multi-scale feature maps of the corresponding images of different modalities, the feature information that is more critical to fusion is highlighted, avoiding important features being ignored or weakened during fusion, making the fusion more in line with the needs of medical diagnosis and fully integrating the advantages of each modality.

[0061] Further, through the spatial registration operation, the spatial pose of the two-dimensional medical fusion image can be accurately determined to ensure the consistency of the spatial positions of the fusion images, thereby improving the accuracy of the three-dimensional medical fusion image.

[0062] In an exemplary embodiment, referring to Figure 6 The improved three-dimensional generative adversarial network is obtained by the following method:

[0063] The initial improved three-dimensional generative adversarial network is constructed based on the three-dimensional VEA-GAN model combined with the discriminator adopting the gradient penalty method of the Wasserstein distance loss function.

[0064] Specifically, the improved three-dimensional generative adversarial network includes an encoder 601, a generator 602, and a discriminator 603. Since the quality of the three-dimensional image generated by the traditional three-dimensional GAN model through random hidden variable input into the generator is low, the three-dimensional VEA-GAN model is constructed by combining the VEA network to obtain the improved three-dimensional generative adversarial network, which can improve the quality of the three-dimensional image generated by the improved three-dimensional generative adversarial network. The input data of the improved three-dimensional generative adversarial network can be a two-dimensional medical fusion image group, and the real data of the improved three-dimensional generative adversarial network can be a two-dimensional medical original image group.

[0065] The pseudo three-dimensional medical fusion image generated by the generator of the improved three-dimensional generative adversarial network is two-dimensionally split at the spatial position corresponding to the image in the single-modality medical original image group to obtain a pseudo two-dimensional medical fusion image group corresponding to each single-modality medical original image group.

[0066] The loss of the discriminator on the real image, i.e. each single modality medical original image group, and the loss on the false image, i.e. the pseudo two-dimensional medical fusion image group, are calculated by inputting each single modality medical original image group and the pseudo two-dimensional medical fusion image group into the discriminator of the initial improved three-dimensional generative adversarial network, and the loss is back propagated to update the parameters of the discriminator, and the improved three-dimensional generative adversarial network is obtained.

[0067] The Wasserstein distance loss function expression using the gradient penalty method is as follows:

[0068]

[0069] In the formula, X is an input image, X' is a generated image, P g is a distribution law followed by the generated image X', P r is a distribution law followed by the input image X, is a straight line connecting the sampling points sampled from the distribution law P r and the distribution law P g , and λ is a penalty factor, is a gradient penalty term.

[0070] In the above three-dimensional medical image fusion method based on deep learning, the initial improved three-dimensional generative adversarial network is constructed based on the three-dimensional VEA-GAN model combined with the discriminator of the Wasserstein distance loss function using the gradient penalty method, which can more accurately measure the distribution difference between the generated image and the real image, and can more smoothly guide the generator to generate images closer to the real distribution, so that the generated image quality is higher and more consistent with the actual situation, and the possibility of mode collapse of the generated image is reduced.

[0071] Further, by two-dimensionally splitting the pseudo three-dimensional medical fusion image generated by the generator at the spatial position corresponding to the image in the single modality medical original image group, the pseudo two-dimensional medical fusion image group corresponding to each single modality medical original image group is obtained, which can more meticulously explore the correlation and difference between the three-dimensional fusion image and the two-dimensional original single modality image.

[0072] In an exemplary embodiment, the two-dimensional medical source images are preprocessed within the modalities to obtain a single modality two-dimensional medical original image group corresponding to each modality, including:

[0073] The two-dimensional medical source images are filtered and optimized to identify and remove noise and artifacts in the two-dimensional medical source images, and the optimized two-dimensional medical source images are obtained.

[0074] The optimized two-dimensional medical source images are modality-intra spatially registered to generate two-dimensional medical original images, so that the projection coordinates of the two-dimensional medical original images in each modality in the vertical direction of the same spatial pose are consistent, and a single-modality two-dimensional medical original image group corresponding to each modality is assembled.

[0075] In an exemplary embodiment, referring to Figure 7 The single-modality two-dimensional medical original image group is modality-inter processed to obtain a multi-modality two-dimensional medical original image group, including:

[0076] Step 701, the two-dimensional medical original images in the single-modality medical original image group that are spatially and pose-overlapped are paired and then modality-inter image-registered to generate a first matched multi-modality medical original image group, and a spatially and pose-displaced medical original image is identified.

[0077] Step 702, based on a spline function, the two-dimensional medical original images in the other single-modality medical original image group closest to the spatially and pose-displaced medical original image are generated into a reconstructed medical original image at the spatial pose where the spatially and pose-displaced medical original image is located.

[0078] Step 703, the spatially and pose-displaced medical original image and the reconstructed medical original image are paired and then modality-inter image-registered to generate a second matched multi-modality medical original image group.

[0079] Step 704, the first matched multi-modality medical original image group and the second matched multi-modality medical original image group are combined to generate a multi-modality two-dimensional medical original image group.

[0080] In the above-described three-dimensional medical image fusion method based on deep learning, through image pairing, reconstructed medical original image generation and modality-inter image registration, the consistency and accuracy of the to-be-fused images in the spatial pose in the first matched multi-modality medical original image group and the second matched multi-modality medical original image group are guaranteed.

[0081] In an exemplary embodiment, the deep learning network model adopts a comprehensive loss function combining an image content loss function and an image similarity loss function;

[0082] The expression of the comprehensive loss function is as follows:

[0083] L z =ρ1×L con +ρ2×L cos

[0084]

[0085] In the formula, L z is the comprehensive loss function, L con is the image content loss function, and Lcos is an image similarity loss function, pi is a weight coefficient of an image content loss function, p2 is a weight coefficient of the image similarity loss function, H is a height of a two-dimensional medical original image, W is a width of the two-dimensional medical original image, N is a total number of fused modalities, I f is a two-dimensional medical fusion image, I k is a two-dimensional medical original image corresponding to the k-th modality in a multi-modality two-dimensional medical original image group, is a gradient operator, a k is a modality pixel weight coefficient, b k is a modality gradient weight coefficient, cos is a vector cosine operator, g k is a modality similarity weight coefficient.

[0086] In an exemplary embodiment, as Figure 8 shown, a deep learning-based three-dimensional medical image fusion method includes:

[0087] Step S801, acquiring two-dimensional medical source images corresponding to at least two modalities.

[0088] Step S802, filtering and optimizing the two-dimensional medical source images to identify and remove noise and artifacts in the two-dimensional medical source images, to obtain optimized two-dimensional medical source images.

[0089] Step S803, performing intra-modality spatial registration on the optimized two-dimensional medical source images to generate two-dimensional medical original images, so that the projection coordinates of the two-dimensional medical original images in the same spatial pose in the vertical direction are consistent, and a single-modality two-dimensional medical original image group corresponding to each modality is established.

[0090] Step S804, performing inter-modality image registration on the two-dimensional medical original images in the single-modality medical original image group that are paired after spatial pose overlap, to generate a first matched multi-modality medical original image group, and identifying spatially displaced medical original images.

[0091] Step S805, based on a spline function, generating a reconstructed medical original image at the spatial pose where the spatially displaced medical original image is located, for the two-dimensional medical original images in the other single-modality medical original image group closest to the spatially displaced medical original image.

[0092] Step S806, performing inter-modality image registration on the spatially displaced medical original image and the reconstructed medical original image after pairing, to generate a second matched multi-modality medical original image group.

[0093] Step S807, combining the first matched multi-modality medical original image group and the second matched multi-modality medical original image group to generate a multi-modality two-dimensional medical original image group.

[0094] Step S808, local feature maps of each image in the multi-modal two-dimensional medical original image group are extracted based on the convolutional neural network branch in the deep learning network model.

[0095] Step S809, global feature maps of each image in the multi-modal two-dimensional medical original image group are extracted based on the transformer network branch in the deep learning network model.

[0096] Step S810, the local feature maps and the global feature maps are coupled to generate coupled feature maps, and the coupled feature maps are input into the cross-scale attention module to generate multi-scale feature maps.

[0097] Step S811, the multi-scale feature maps of the images corresponding to different modalities in the multi-modal two-dimensional medical original image group that need to be fused are input into the feature map fusion module to generate fused feature maps.

[0098] Step S812, the fused feature maps are input into the feature decoding network to generate two-dimensional medical fusion images.

[0099] Step S813, the two-dimensional medical fusion images are constructed into a two-dimensional medical fusion image group after spatial registration.

[0100] Step S814, the two-dimensional medical fusion image group is input into the improved three-dimensional generative adversarial network taking the two-dimensional medical original image group as real images to generate three-dimensional medical fusion images.

[0101] In the above deep learning-based three-dimensional medical image fusion method, effective optimization, fusion and three-dimensional reconstruction of multi-modal two-dimensional medical images can be achieved, providing richer and more accurate image data for medical diagnosis and research, and improving the efficiency and accuracy of medical diagnosis.

[0102] It should be understood that although each step in the flowchart involved in each embodiment as described above is displayed in sequence according to the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other orders. Moreover, at least part of the steps in the flowchart involved in each embodiment as described above can include multiple steps or stages, which are not necessarily executed at the same time but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential but can be alternately or alternately executed with at least part of other steps or steps or stages in other steps.

[0103] Based on the same inventive concept, the embodiments of the present application also provide a deep learning-based three-dimensional medical image fusion device for implementing the deep learning-based three-dimensional medical image fusion method described above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme described in the above method, so the specific limitations in one or more deep learning-based three-dimensional medical image fusion device embodiments provided below can refer to the limitations of the deep learning-based three-dimensional medical image fusion method described above, which will not be repeated here.

[0104] In one exemplary embodiment, as shown in Figure 9 a deep learning-based three-dimensional medical image fusion device 900 is provided, comprising:

[0105] The original image acquisition module 901 can be used to acquire two-dimensional medical source images corresponding to at least two modalities, and to perform intra-modality preprocessing on the two-dimensional medical source images to obtain a single-modality two-dimensional medical original image group corresponding to each modality. Intra-modality preprocessing is used to improve the image quality of the two-dimensional medical source images corresponding to a single modality.

[0106] The multi-modality preprocessing module 902 can be used to perform inter-modality preprocessing on the single-modality two-dimensional medical original image group to obtain a multi-modality two-dimensional medical original image group. Inter-modality preprocessing is used to align two-dimensional medical source images corresponding to different modalities.

[0107] The two-dimensional image fusion module 903 can be used to input the multi-modality two-dimensional medical original image group into a deep learning network model to generate a two-dimensional medical fusion image group; the two-dimensional medical fusion image group is used to represent a two-dimensional medical fusion image that fuses the features of two-dimensional medical source images of different modalities corresponding to the same spatial pose in the multi-modality two-dimensional medical original image group; and the deep learning network model is obtained by combining a convolutional neural network and a transformer network.

[0108] The three-dimensional image fusion module 904 can be used to input the two-dimensional medical fusion image group into an improved three-dimensional generative adversarial network with the two-dimensional medical original image group as the real image to generate a three-dimensional medical fusion image.

[0109] In one embodiment, the two-dimensional image fusion module 903 can also be used to:

[0110] extract local feature maps of each image in the multi-modality two-dimensional medical original image group based on a convolutional neural network branch in the deep learning network model;

[0111] extract global feature maps of each image in the multi-modality two-dimensional medical original image group based on a transformer network branch in the deep learning network model, the transformer network branch being obtained by replacing a multi-head attention module in an original transformer network with a spatially reduced multi-head attention module.

[0112] The local feature map and the global feature map are coupled to generate a coupled feature map, and the coupled feature map is input into a cross-scale attention module to generate a multi-scale feature map;

[0113] Based on the multi-scale feature maps of the images corresponding to different modalities that need to be fused in the multi-modal two-dimensional medical original image group, a two-dimensional medical fusion image is generated, and a two-dimensional medical fusion image group is constructed.

[0114] In one embodiment, the two-dimensional image fusion module 903 can also be used for:

[0115] The multi-scale feature maps of the images corresponding to different modalities that need to be fused in the multi-modal two-dimensional medical original image group are input into a feature map fusion module to generate a fusion feature map, and the feature map fusion module includes a global attention mechanism submodule and a fusion submodule;

[0116] The fusion feature map is input into a feature decoding network to generate a two-dimensional medical fusion image;

[0117] The two-dimensional medical fusion image is used to construct a two-dimensional medical fusion image group after spatial registration, and the two-dimensional medical fusion image group is used to represent the spatial pose of the two-dimensional medical fusion image.

[0118] In one embodiment, the three-dimensional medical image fusion device 900 based on deep learning can be used for:

[0119] An initial improved three-dimensional generative adversarial network is constructed based on a three-dimensional VEA-GAN model and a discriminator using a gradient penalty method Wasserstein distance loss function;

[0120] The pseudo three-dimensional medical fusion image generated by the generator of the improved three-dimensional generative adversarial network is two-dimensionally split at the spatial position corresponding to the image in the single modality medical original image group to obtain a pseudo two-dimensional medical fusion image group corresponding to each single modality medical original image group;

[0121] Each single modality medical original image group and the pseudo two-dimensional medical fusion image group are input into the discriminator of the initial improved three-dimensional generative adversarial network to calculate the loss of the discriminator on the real image, i.e., each single modality medical original image group, and the loss on the fake image, i.e., the pseudo two-dimensional medical fusion image group, and the loss is back propagated to update the parameters of the discriminator to obtain an improved three-dimensional generative adversarial network;

[0122] The Wasserstein distance loss function using the gradient penalty method is expressed as follows:

[0123]

[0124] Where X is the input image, X' is the generated image, and P g To generate the distribution law of the image X', P r is the distribution law obeyed by the input image X, is the sampling self-distribution law P r and distribution law P g The straight line connecting the sampling points between them, λ is the penalty factor, is the gradient penalty term.

[0125] In one embodiment, the original image acquisition module 901 may also be used to:

[0126] Performing filtering optimization on the two-dimensional medical source image, identifying and removing noise and artifacts in the two-dimensional medical source image, and obtaining an optimized two-dimensional medical source image;

[0127] The optimized two-dimensional medical source images are subjected to intra-modal spatial registration to generate two-dimensional medical original images, so that the projection coordinates of the two-dimensional medical original images in each modality in the same spatial position are consistent in the vertical direction, and a single-modality two-dimensional medical original image group corresponding to each modality is formed.

[0128] In one embodiment, the multimodal preprocessing module 902 may also be used to:

[0129] After pairing the two-dimensional medical original images with overlapping spatial postures in the single-modality medical original image group, inter-modality image registration is performed to generate a first matching multi-modality medical original image group, and spatially heterotopic medical original images are identified;

[0130] For a two-dimensional medical original image in another single-modality medical original image group that is closest to the spatially heterotopic medical original image, a reconstructed medical original image is generated at the spatial position of the spatially heterotopic medical original image based on a spline function;

[0131] performing inter-modality image registration after pairing the spatially heterotopic original medical image and the reconstructed original medical image to generate a second matched multi-modality original medical image group;

[0132] The first matching multimodal medical original image group and the second matching multimodal medical original image group are combined to generate a multimodal two-dimensional medical original image group.

[0133] In one embodiment, a computer device is provided, comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the three-dimensional medical image fusion method based on deep learning as described above are implemented.

[0134] In one embodiment, a computer readable storage medium is provided, having stored thereon a computer program, which when executed by a processor implements the steps of any of the above method embodiments.

[0135] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts are described in the part of the method embodiments. The above described device embodiments are only illustrative, wherein the components described as separate components can or can not be physically separated, the components displayed as units can or can not be physical units, i.e. can be located in one place or can be distributed to multiple network units. Part or all of the modules can be selected to achieve the purpose of the present disclosure according to actual needs. Those skilled in the art can understand and implement it without creative labor.

[0136] The above described embodiments only express several implementation manners of the present application, the description is more specific and detailed, but it should not be understood as a limitation to the patent scope of the application. It should be pointed out that, for those skilled in the art, without departing from the concept of the present application, several modifications and improvements can be made, which are all within the protection scope of the present application.

Claims

1. A deep learning-based three-dimensional medical image fusion method, characterized by, The method comprises: acquiring two-dimensional medical source images corresponding to at least two modalities, and performing intra-modality preprocessing on the two-dimensional medical source images to obtain a single-modality two-dimensional medical original image group corresponding to each modality, wherein the intra-modality preprocessing is used to improve the image quality of the two-dimensional medical source images corresponding to a single modality; performing inter-modality preprocessing on the single-modality two-dimensional medical original image group to obtain a multi-modality two-dimensional medical original image group, wherein the inter-modality preprocessing is used to align the two-dimensional medical source images corresponding to different modalities; inputting the multi-modality two-dimensional medical original image group into a deep learning network model to generate a two-dimensional medical fusion image group, wherein the two-dimensional medical fusion image group is used to represent a two-dimensional medical fusion image that fuses the features of two-dimensional medical source images of different modalities corresponding to the same spatial pose in the multi-modality two-dimensional medical original image group, and the deep learning network model is obtained by combining a convolutional neural network and a transformer network; inputting the two-dimensional medical fusion image group into an improved three-dimensional generative adversarial network taking the two-dimensional medical original image group as real images to generate a three-dimensional medical fusion image; wherein the inputting the multi-modality two-dimensional medical original image group into the deep learning network model to generate a two-dimensional medical fusion image group comprises: extracting local feature maps of each image in the multi-modality two-dimensional medical original image group based on a convolutional neural network branch in the deep learning network model; extracting global feature maps of each image in the multi-modality two-dimensional medical original image group based on a transformer network branch in the deep learning network model, wherein the transformer network branch is obtained by replacing a multi-head attention module in an original transformer network with a spatially reduced multi-head attention module; performing feature coupling on the local feature maps and the global feature maps to generate coupled feature maps, and inputting the coupled feature maps into a cross-scale attention module to generate multi-scale feature maps; generating the two-dimensional medical fusion image based on the multi-scale feature maps of images of different modalities that need to be fused in the multi-modality two-dimensional medical original image group, and constructing the two-dimensional medical fusion image group; wherein the improved three-dimensional generative adversarial network is obtained by the following method: constructing an initial improved three-dimensional generative adversarial network based on a three-dimensional VEA-GAN model combined with a discriminator adopting a gradient penalty method Wasserstein distance loss function; performing two-dimensional splitting on pseudo three-dimensional medical fusion images generated by a generator of the improved three-dimensional generative adversarial network at spatial positions corresponding to images in a single-modality medical original image group to obtain a pseudo two-dimensional medical fusion image group corresponding to each single-modality medical original image group; inputting each single-modality medical original image group and the pseudo two-dimensional medical fusion image group into the discriminator of the initial improved three-dimensional generative adversarial network to calculate the loss of the discriminator on real images, i.e., the loss of the discriminator on the single-modality medical original image group, and the loss of the discriminator on fake images, i.e., the loss of the discriminator on the pseudo two-dimensional medical fusion image group, and reversely propagating the loss to update the parameters of the discriminator, thereby obtaining the improved three-dimensional generative adversarial network.

2. The method of claim 1, wherein, The multi-scale feature maps of the images corresponding to different modalities that need to be fused in the multi-modal two-dimensional medical original image group are input into a feature map fusion module to generate a fusion feature map, and the feature map fusion module includes a global attention mechanism submodule and a fusion submodule; The fusion feature map is input into a feature decoding network to generate the two-dimensional medical fusion image; The two-dimensional medical fusion image is subjected to spatial registration to construct the two-dimensional medical fusion image group, which is used to represent the spatial pose of the two-dimensional medical fusion image. The Wasserstein distance loss function using the gradient penalty method is expressed as follows:

3. The method of claim 1, wherein, The two-dimensional medical source images are subjected to modal intra-preprocessing to obtain a single-modal two-dimensional medical original image group corresponding to each modality, including: where X is the input image, X' is the generated image, P g is the distribution law that the generated image X' is subject to, P r is the distribution law that the input image X is subject to, is the sampling distribution from the distribution law P r and the distribution law P g is the straight line connecting the sampling points between the distribution law P is the gradient penalty term.

4. The method of claim 1, wherein, The two-dimensional medical source images are subjected to filtering optimization to identify and remove noise and artifacts in the two-dimensional medical source images, thereby obtaining the optimized two-dimensional medical source images; The optimized two-dimensional medical source images are subjected to modal intra-spatial registration to generate two-dimensional medical original images, so that the projection coordinates of the two-dimensional medical original images in the same spatial pose in the vertical direction are consistent, and a single-modal two-dimensional medical original image group corresponding to each modality is formed. The single-modal two-dimensional medical original image group is subjected to inter-modal preprocessing to obtain a multi-modal two-dimensional medical original image group, including:

5. The method of claim 4, wherein, The two-dimensional medical original images in the single-modal medical original image group that overlap in spatial pose are paired and subjected to inter-modal image registration to generate a first matched multi-modal medical original image group, and a spatially displaced medical original image is identified; Based on a spline function, the two-dimensional medical original images in the other single-modal medical original image group closest to the spatially displaced medical original image are used to generate a reconstructed medical original image at the spatial pose of the spatially displaced medical original image; The spatially displaced medical original image and the reconstructed medical original image are paired and subjected to inter-modal image registration to generate a second matched multi-modal medical original image group; The multi-modal two-dimensional medical original image group is generated by combining the first matched multi-modal medical original image group and the second matched multi-modal medical original image group. The deep learning network model uses a comprehensive loss function combining an image content loss function and an image similarity loss function; 6. The method according to any one of claims 1 to 5, characterized in that, The expression of the comprehensive loss function is as follows: The device includes various functional modules required for implementing the method of any one of claims 1 to 6. L z = p1 x L con + p2 x L cos In the formula, L z is a comprehensive loss function, L con is an image content loss function, L cos is an image similarity loss function, ρ1 is a weight coefficient of the image content loss function, ρ2 is a weight coefficient of the image similarity loss function, H is the height of a two-dimensional medical original image, W is the width of the two-dimensional medical original image, N is the total number of fused modes, I f is a two-dimensional medical fusion image, I k is a two-dimensional medical original image corresponding to the kth mode in a multi-modal two-dimensional medical original image group, is a gradient operator, α k is a mode pixel weight coefficient, β k is a mode gradient weight coefficient, cos is a vector cosine operator, γ k is a mode similarity weight coefficient. 7.A three-dimensional medical image fusion apparatus based on deep learning, characterized by The processor executes the computer program to implement the steps of the method of any one of claims 1 to 6.

8. A computer device comprising a memory and a processor, the memory storing a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 6.

9. A computer readable storage medium having stored thereon a computer program, characterized in that, ​

Citation Information

Patent Citations

  • Weak supervised lesion segmentation

    CN115004225A

  • Image style migration method, model and device, electronic equipment and storage medium

    CN116309014A