Medical image unified diffusion generation method oriented to random mode deficiency and related device

Through the unified diffusion generation method of medical images oriented towards random modal deletion, the visual cue diffusion generation model is used to generate missing modal medical images, which solves the problem of modal deletion in magnetic resonance imaging, and improves the image generation quality and the accuracy of clinical diagnosis.

CN120107393APending Publication Date: 2025-06-06XI AN JIAOTONG UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510269135.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively deal with the problem of random modes in magnetic resonance imaging in clinical practice, resulting in the impact of disease diagnosis and treatment planning.

Method used

A unified diffusion generation method for random modal deletion is adopted for medical images, and a pre-trained visual cue diffusion generation model, including visual cue extraction subnet and residual enhanced U-Net network, is used to generate missing modal medical images.

Benefits of technology

High-quality medical imaging generation in the absence of any modality is achieved, improving the accuracy of clinical diagnosis and the reliability of treatment plans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107393A_ABST
    Figure CN120107393A_ABST
Patent Text Reader

Abstract

The invention discloses a random mode missing-oriented medical image unified diffusion generation method and a related device. The generation method comprises the following steps: acquiring one or more mode missing medical images and corresponding mode acquisition codes; sending the medical image and the modal acquisition code into a pre-trained visual prompt diffusion generation model to obtain a missing modal medical image; the pre-trained visual cue diffusion generation model comprises a visual cue extraction sub-network and a residual error enhanced U-Net network, and the visual cue extraction sub-network is used for extracting visual cues from the medical image according to modal acquisition codes; and the residual-enhanced U-Net network is used for continuously iteratively denoising the collected random Gaussian noise based on visual cues and target modal coding until a missing modal medical image is generated. The method focuses on exploring and improving the self-adaptive generation capability of the medical image of the diffusion generation model, and adopts the visual prompt extraction sub-network and the residual error enhanced U-Net network to realize unified and accurate generation of the medical image under the condition of random mode deficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of medical image analysis, and specifically relates to a unified diffusion generation method for medical images with random modality missing and a related device, which is used to accurately generate the missing modality images based on the acquired modality images when several modality images are randomly missing in the multimodality medical images of a subject, thereby filling and constructing a complete multimodality medical image of the subject, which is used to assist clinicians in diagnosing the disease and subsequently applying medical image automatic analysis software for analysis and diagnosis. Background Art

[0002] Magnetic resonance imaging (MRI) is an important medical imaging modality that provides high-resolution, multimodal images and is essential for accurate diagnosis of diseases and treatment planning. However, in clinical practice, it is often impractical to obtain a complete set of MRI data for each patient due to various constraints such as limited scanning time, patient motion, and resource availability. Missing modalities in MRI scans pose a huge challenge to comprehensive medical analysis and downstream tasks. The absence of certain modalities can hinder accurate diagnosis of diseases, treatment planning, and subsequent monitoring, ultimately affecting the quality of patient care. In addition, different patients may be missing different modalities, which also exacerbates the severity and complexity of the problem. How to generate missing modality images based on the medical images of the patient's acquired modality to construct a complete MRI data set is of great clinical significance.

[0003] To address this problem, there are now some medical image generation methods that can generate missing modality images. Most of these generation methods focus on the mutual generation between two modality images, or the generation of target modality images based on a specific source modality. How to extend these methods to multimodal generation to handle the situation where any modality is missing remains a major challenge. A few generation methods can handle the generation of any missing modality images with a unified model, but they are all based on traditional deep learning technologies such as deep autoencoders and generative adversarial networks. They do not apply advanced deep diffusion generation model technology, and the quality of the generated medical images is poor, which limits their application in clinical scenarios. Summary of the invention

[0004] In order to further improve the quality of missing modality medical image generation under random modality missing conditions, the purpose of the present invention is to propose a unified diffusion generation method and related devices for medical images with random modality missing conditions. This method achieves higher missing modality image generation quality by processing the unified diffusion of medical images with different modality missing conditions.

[0005] To achieve the above purpose, the technical solution adopted by the present invention is as follows:

[0006] A unified diffusion generation method for medical images with random modality missing includes the following steps:

[0007] Acquire one or more missing modality medical images and corresponding modality acquisition codes;

[0008] The acquired medical images and modality acquisition codes are fed into a pre-trained visual cue diffusion generation model to obtain the missing modality medical images;

[0009] Among them, the pre-trained visual cue diffusion generation model includes a visual cue extraction subnetwork and a residual enhanced U-Net network. The visual cue extraction subnetwork is used to extract visual cues from medical images according to the modality acquisition coding; the residual enhanced U-Net network is used to iteratively denoise the acquired random Gaussian noise based on visual cues and target modality coding until a missing modality medical image is generated.

[0010] Furthermore, the visual cue extraction subnetwork includes a main network and a modality acquisition code embedding supernetwork, which respectively take the medical image and the modality acquisition code as input, wherein the modality acquisition code is a 0-1 code for indicating the corresponding modality acquisition or absence; the main network includes three 2D convolutional layers connected in sequence, four first repeating units and a first residual module; the modality acquisition code embedding supernetwork includes five fully connected layers connected in sequence;

[0011] 2D convolutional layer, used to extract features from medical images;

[0012] A first repeating unit, for extracting multi-scale visual cues based on features, comprising a second residual module and a downsampling layer;

[0013] A first residual module for extracting the lowest-scale visual cue from the output of the first repeating unit;

[0014] The fully connected layer is used to generate parameters according to the modal acquisition encoding and adjust the first residual module and the second residual module in the corresponding main network.

[0015] Furthermore, the residual-enhanced U-Net network includes a main network and a target modality encoding embedding super-network;

[0016] The target modality coding embedding hypernetwork includes seventeen fully connected layers connected in sequence, taking the target modality coding as input and generating parameters for adjusting the corresponding residual module in the main network, wherein the target modality coding is a one-hot coding, which is used to indicate the missing modality medical image to be generated.

[0017] Furthermore, the main network in the residual enhanced U-Net network takes the noisy image and visual cues as input to produce a denoised image; the main network includes an encoder, four skip connections and a decoder, and the overall structure is U-shaped and symmetrical; each skip connection contains two convolutional layers, which are connected from the residual module of the encoder to the splicing layer of the corresponding resolution of the decoder.

[0018] Furthermore, the encoder of the main network is used to extract multi-resolution noise image features, including a 2D convolutional layer, four second repeating units, a concatenation layer, and a third residual module;

[0019] 2D convolutional layer, used to extract initial features from noisy images;

[0020] A second repeating unit, for extracting multi-resolution noise image features according to the initial features, comprises a concatenation layer, a fourth residual module and a downsampling layer;

[0021] A splicing layer, used for splicing the noise image features output by the second repeating unit and the visual prompts of the corresponding resolution;

[0022] The third residual module is used to extract the lowest resolution noise image features based on the output of the concatenated layer.

[0023] Further, the decoder of the main network is used to generate a denoised image according to the multi-resolution noise image features, including a fifth residual module, a sixth residual module, four third repeating units, a seventh residual module, an eighth residual module, a 2-dimensional convolutional layer and a tanh activation function;

[0024] a fifth residual module, for extracting features based on the noise image features of the lowest resolution;

[0025] a sixth residual module, used for extracting denoised image features of the lowest resolution according to the output of the fifth residual module;

[0026] A third repeating unit, for extracting higher-resolution denoised image features according to the output of the sixth residual module, comprising a ninth residual module, an upsampling layer, and a splicing layer;

[0027] The seventh residual module is used to extract the denoised image features with the highest resolution according to the output of the third repeating unit,

[0028] The eighth residual module is used to extract more complete denoising image features according to the output of the seventh residual module.

[0029] A 2D convolutional layer for extracting an unnormalized denoised image based on the output of the eighth residual module;

[0030] The tanh activation function is used to map the output of the 2D convolutional layer to between -1 and 1 and output a denoised image.

[0031] Furthermore, the loss function of the pre-trained visual cue diffusion generation model is for:

[0032]

[0033] in,‖·‖ 2 represents the 2-norm, N is the number of modes, m i is the ith mode, For a clean image, is the noise image at time t, F vp (X obs ) is a visual cue, X obs For medical images, c i is the target modality encoding, and t is the time.

[0034] A unified diffusion generation system for medical images with random modality missing, including:

[0035] Acquire one or more missing modality medical images and corresponding modality acquisition codes;

[0036] The acquired medical images and modality acquisition codes are fed into a pre-trained visual cue diffusion generation model to obtain the missing modality medical images;

[0037] Among them, the pre-trained visual cue diffusion generation model includes a visual cue extraction subnetwork and a residual enhanced U-Net network. The visual cue extraction subnetwork is used to extract visual cues from medical images according to the modality acquisition coding; the residual enhanced U-Net network is used to iteratively denoise the acquired random Gaussian noise based on visual cues and target modality coding until a missing modality medical image is generated.

[0038] An electronic device comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the unified diffusion generation method for medical images with random modality missing is implemented.

[0039] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the unified diffusion generation method for medical images with random modality missing is implemented.

[0040] Compared with the prior art, the present invention has at least the following beneficial technical effects:

[0041] The present invention realizes the accurate generation of missing modality medical images in the case of any missing modality through a pre-trained visual cue diffusion generation model. Compared with the existing cross-modality medical image generation method, the method of the present invention trains a unified medical image generation model based on the diffusion generation model, which can handle the problem of medical image generation in the case of any missing modality. Compared with the existing unified generation method of missing modality medical images, the method of the present invention is based on a more advanced diffusion generation model, adopts a visual cue extraction subnetwork and a residual enhanced U-Net network, and realizes the unified generation of medical images in the case of any missing modality. It can generate higher quality medical images, and the sampling generation process is more intuitive and visible. The present invention can be mainly used for the generation of missing modalities of medical images, that is, for the situation where any number of modalities are missing in multimodal medical images, the missing modality images can be accurately generated based on the existing modality images, which has important application value in assisting clinicians in accurate diagnosis, treatment planning and disease monitoring.

[0042] Furthermore, the visual cue extraction subnetwork of the present invention realizes the adaptive extraction of uniform visual cues for medical images in any missing modality by designing modality acquisition coding and modality acquisition coding embedded supernetwork, which only requires training one network. The extracted visual cues will be used as conditions for the visual cue back diffusion process to generate any missing modality image.

[0043] Furthermore, the residual enhanced U-Net network of the present invention serves as a core module in the back-diffusion process of visual cues. By designing the target modality coding and the target modality coding embedded in the hypernetwork, it is achieved that only one network needs to be trained. By inputting different target modality codings, the network can be controlled to generate missing modality medical images corresponding to the target modality coding. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 It is a specific implementation flow chart of the present invention;

[0045] Figure 2 This is a framework diagram of the visual cue diffusion generation model;

[0046] Figure 3 This is the structure diagram of the visual cue extraction subnetwork of the visual cue diffusion generation model;

[0047] Figure 4 It is the residual enhanced U-Net network structure diagram of the visual cue diffusion generation model;

[0048] Figure 5 is the missing modality image generation result of the present invention on different acquired medical images; wherein (a) is an acquired medical image, (b) is a real missing modality medical image, and (c) is a missing modality medical image generated by the present invention;

[0049] Figure 6 is a flowchart of the medical image generation method of the present invention;

[0050] Figure 7 Schematic diagram of a medical image generation system of the present invention. DETAILED DESCRIPTION

[0051] In order to facilitate the understanding of the present invention, the present invention will be described more fully below with reference to the relevant drawings. Preferred embodiments of the present invention are given in the drawings. However, the present invention can be implemented in a variety of different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present invention more thoroughly and comprehensively understood. It should be noted that the terms "including" and "having" in the specification and claims of the present invention and the above-mentioned drawings and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0052] like Figure 6 As shown, the present invention provides a unified diffusion generation method for medical images with random modality loss, comprising the steps of:

[0053] S1, obtaining one or more missing modality medical images and corresponding modality acquisition codes;

[0054] S2, sending the acquired medical image and modality acquisition code into a pre-trained visual cue diffusion generation model to obtain the missing modality medical image;

[0055] Among them, the pre-trained visual cue diffusion generation model includes a visual cue extraction subnetwork and a residual enhanced U-Net network. The visual cue extraction subnetwork is used to extract visual cues from medical images according to the modality acquisition coding; the residual enhanced U-Net network is used to iteratively denoise the acquired random Gaussian noise based on visual cues and target modality coding until a missing modality medical image is generated.

[0056] See also Figure 1, the present invention realizes the generation of medical images in any modality missing situation through the visual cue diffusion generation model: construct a visual cue extraction subnetwork to adaptively extract visual cues from the collected medical images; continuously add noise to the clean medical images through the forward diffusion process, reconstruct the reverse diffusion process, design the residual enhanced U-Net network, learn to denoise the noisy image based on visual cues, and predict the clean image; in the model training stage, define the denoising objective function, and use the AdamW algorithm to learn the optimal parameters of the network in the reverse diffusion process based on the training data set; after the model is trained, apply the learned reverse diffusion process to the collected medical images (one or more modalities are randomly missing), iteratively sample from pure Gaussian noise until a clean image is restored, which is the generated missing modality medical image. Specifically includes the following steps:

[0057] 1. Visual Cue Extraction Subnetwork Construction

[0058] The present invention constructs a visual cue extraction subnetwork F vp , used to obtain medical images X obs (where one or more modalities are randomly missing). Assume that a complete multimodal medical image has N modalities. Acquired medical images X obs Medical images of 1 to N-1 modalities may be randomly missing. For different modality missing situations, the extraction process will use the modality acquisition code v a Perform adaptive adjustment. Modal acquisition coding v a is a 0-1 code of length N, where 0 indicates the corresponding modality is missing and 1 indicates the corresponding modality has been acquired. The extracted visual cues will be used as a uniform condition for generating clean medical images and then fed into the diffusion process of visual cues, see Figure 2 .

[0059] See also Figure 3 The visual cue extraction subnetwork of the present invention consists of the main network ( Figure 3 above) and modality acquisition encoding embedded hypernetwork ( Figure 3 The main network consists of two parts, which take medical images and modality acquisition codes as inputs respectively. The modality acquisition code is a 0-1 code, which is used to indicate the corresponding modality acquisition or missing. The modality acquisition code is embedded in the super network to dynamically adjust the parameters of the main network. Specifically, the main network includes three 2D convolutional layers connected in sequence, four first repeating units and a first residual module.

[0060] Among them, the 2D convolutional layer is used to extract features from medical images;

[0061] A first repeating unit, for extracting multi-scale visual cues based on features, comprising a second residual module and a downsampling layer;

[0062] A first residual module for extracting the lowest-scale visual cue from the output of the first repeating unit;

[0063] The fully connected layer is used to generate parameters according to the modal acquisition encoding and adjust the first residual module and the second residual module in the corresponding main network.

[0064] Modality acquisition coding embedding hypernetwork with modality acquisition coding v a As input, it includes five fully connected layers connected in sequence, and the corresponding layers in the main network are adjusted by the filter scaling strategy as described below.

[0065] Filter scaling strategy: For the residual module of the main network, there are d filters The convolutional layer, κ i is the filter, i is the filter number, and the corresponding fully connected layer of the modal acquisition code embedded in the hypernetwork will produce d scaling factors '

[0066] α i Then, each filter κ of the convolutional layer i will be modified to the modified positive filter κ i =α i k i , where α i Scaling the filter κ by element-wise multiplication i This strategy allows the main network to extract visual cues based on the missing modality (i.e., modality acquisition encoding v a )Adaptive adjustment.

[0067] 2. The Diffusion Process of Visual Cues

[0068] See also Figure 2 The present invention constructs a visual cue diffusion process for generating a multimodal medical image based on the extracted unified visual cue. Specifically, the constructed visual cue diffusion process includes a forward diffusion process and a backward diffusion process. The forward diffusion process gradually adds Gaussian noise to the multimodal medical image through a Markov chain. The backward diffusion process of the visual cue gradually denoises the noisy image based on the visual cue to generate a clean multimodal medical image. These forward and backward diffusion processes operate on each modality of the medical image respectively. The following is a schematic diagram of the forward and backward diffusion processes of the modality m. i Take the following as an example to introduce these processes. i The corresponding modal code c i is a one-hot encoding of length N, where the 1 element position indicates that the corresponding mode is m i .

[0069] 1) Forward diffusion process construction

[0070] The present invention has no special requirements or designs for the forward diffusion process, and only requires a Markov process to realize continuous noise addition to the clean medical image until pure Gaussian noise is obtained. Here, a feasible construction method of the forward diffusion process is described.

[0071] Forward diffusion through a Markov chain For mode m i Clean image of Continuously add Gaussian noise to produce a series of increasing noise images in Approximate pure Gaussian noise. Given the noise image at time t-1 The forward diffusion process can be defined as:

[0072]

[0073] Among them, α t is the variance scaling parameter for adding noise at time t, is the noise image at time t, is the noise image at time t-1, is the noise image at time 1, is the noise image at time 2, is the noise image at time T, is a Gaussian distribution, α t is the variance scaling parameter, and I is the identity matrix.

[0074] 2) Construction of the reverse diffusion process of visual cues

[0075] The visual cue back diffusion process of the present invention iteratively denoises the noisy image to generate a modal m i Clean image of This process can be done by approximating the posterior distribution Specifically, based on formula (1), the posterior distribution It can be deduced as

[0076]

[0077] in,

[0078]

[0079] γ t is the cumulative variance scaling parameter at time t, defined as However, since the clean image Unable to obtain, the posterior distribution It cannot be directly applied in the back diffusion process.

[0080] Among them, μt is the mean function of the posterior distribution, σ t is the standard deviation, γ t-1 is the cumulative variance scaling parameter at time t-1.

[0081] To address this problem, the visual cue back diffusion process of the present invention does not generate medical images unconditionally, but generates missing modality images based on the acquired medical images, which is a conditional generation process. In this task, the visual cue back diffusion process attempts to learn a conditional distribution in the training phase.

[0082]

[0083] To approximate the posterior distribution Among them, F vp (X obs ) represents the collected medical image X obs The visual cues extracted from θ is the conditional distribution mean function. Further, the present invention does not directly design a network to learn the mean function Instead, a residual enhanced U-Net network f is designed θ As a denoiser to directly reconstruct clean medical images Specifically, the residual enhanced U-Net network f θ The noisy image Unified visual cuesF vp (X obs ), modal code c i and time point t as input, the goal is to output a clean medical image Based on the denoising network f θ , the present invention uses the network to learn the mean function Modeled as follows

[0084]

[0085] According to the conditional distribution in formula (3) The present invention can generate mode m by reverse diffusion sampling at multiple time points. i This process gradually removes the pure Gaussian noise and obtains the desired clean medical image. Using the learned conditional distribution To ensure that the generated medical images With the collected medical image X obs The consistency between them.

[0086] Next, the core denoising network in the back-diffusion process of visual cues, namely the residual enhanced U-Net network designed by the present invention, is introduced in detail. θ. Residual enhanced U-Net network f θ For encoding the extracted visual cues and target modality c i Under the condition of Restoring noise-free medical images See also Figure 4 , the residual enhanced U-Net network f proposed in this invention θ A main network ( Figure 4 top) and a corresponding target modality encoding embedding hypernetwork ( Figure 4 The target modality encoding embedding hypernetwork consists of seventeen fully connected layers connected in sequence, with the target modality encoding c i As input, each fully connected layer corresponds to a residual module in the main network, and its parameters are adjusted by the filter scaling strategy in step one. The target modality is encoded as a one-hot encoding, which is used to indicate the missing modality medical image to be generated. The main network takes the noisy image and the visual cue as input to produce a denoised image. The basic architecture of the main network follows the U-Net network in the existing denoising diffusion model, and makes some modifications to it to improve performance. Specifically, the present invention introduces two additional convolutional layers in the skip connection of each resolution to transform the extracted features of the input noisy image, replaces the nearest neighbor upsampling layer with a bilinear upsampling layer, and adds a tanh activation function at the end to map the range of the final output image to between -1 and 1. In addition, the visual cues of different resolutions extracted by the visual cue extraction subnetwork are spliced ​​with the deep features of the corresponding resolutions in the residual enhanced U-Net network main network, and then sent to the subsequent residual module, see. Figure 4 .

[0087] Specifically, the residual enhanced U-Net network f θ The main network in the middle includes an encoder, four jump connections and a decoder, and the overall structure is U-shaped and symmetrical. The encoder is used to extract multi-resolution noise image features. The encoder includes a 2D convolution layer, four groups of splicing layers, residual modules and downsampling layers, and one group of splicing layers and residual modules. The outputs of the first four residual modules are fed into the jump connection, and the output of the last residual module is fed into the decoder. The decoder is used to generate denoised images based on multi-resolution noise image features. Each jump connection contains two convolution layers, which are connected from the residual module of the encoder to the splicing layer of the corresponding resolution of the decoder. The decoder includes two residual modules, four groups of residual modules, upsampling layers and splicing layers, two residual modules, a convolution module and a tanh activation function. Specifically, the decoder includes a fifth residual module, a sixth residual module, four third repeating units, a seventh residual module, an eighth residual module, a 2D convolution layer and a tanh activation function;

[0088] a fifth residual module, for extracting features based on the noise image features of the lowest resolution;

[0089] a sixth residual module, used for extracting denoised image features of the lowest resolution according to the output of the fifth residual module;

[0090] A third repeating unit, for extracting higher-resolution denoised image features according to the output of the sixth residual module, comprising a ninth residual module, an upsampling layer, and a splicing layer;

[0091] The seventh residual module is used to extract the denoised image features with the highest resolution according to the output of the third repeating unit,

[0092] The eighth residual module is used to extract more complete denoising image features according to the output of the seventh residual module.

[0093] A 2D convolutional layer for extracting an unnormalized denoised image based on the output of the eighth residual module;

[0094] The tanh activation function is used to map the output of the 2D convolutional layer to between -1 and 1 and output a denoised image.

[0095] 3. Training Objectives of Visual Cue Diffusion Generative Model

[0096] In the training phase, the present invention uses an early stopping strategy to train a U-Net segmenter to generate segmentation labels of the training set as pseudo labels when training the domain generalization diffusion segmentation model. In order to make the learned back diffusion process accurate and domain generalized, the present invention uses the following three training loss functions.

[0097] In the training phase, in order to learn the visual cue back-diffusion process applicable to any modality missing situation, the present invention trains a visual cue extraction subnetwork F vp and a residual enhanced U-Net network f θ , for each mode m at each time step t i Restoring noise-free clean medical images The final training goal of the visual cue diffusion generation model of the present invention is the denoising objective function for

[0098]

[0099] in,‖·‖ 2 represents the 2-norm.

[0100] The present invention will complete multimodal medical imaging data As training data, randomly sample some modalities to construct the collected images X each time during training. obs, the back-diffusion process of the visual cue diffusion generation model is trained using the AdamW algorithm and the above loss function.

[0101] 4. Apply the trained visual cue diffusion generation model to generate medical images in the absence of any modality

[0102] After the single-source domain is trained in step three, the learned visual cue diffusion generation model can be directly applied to medical image generation in any modality-missing situation.

[0103] Specifically, the acquired medical images (where one or more modalities are missing) obs And the corresponding modal acquisition code v a The present invention directly converts the collected medical image X obs and the corresponding modal acquisition code v a Send it to the trained visual cue extraction subnetwork F vp In the above example, we extract a unified visual cue F vp (X obs ). For the missing mode m to be generated i The present invention collects the visual prompt F vp (X obs ) and missing modality coding c i Send it to the trained residual enhanced U-Net network f θ As a condition, sample random pure Gaussian noise And starting from it, using the conditional distribution in formula (3) Iterative denoising to obtain a series of noisy images in, That is, the missing mode m generated by the present invention i The above sampling process is performed on all missing modalities one by one to obtain all missing modality medical images. In addition, the sampling process of the present invention can be accelerated by using the Denoising Diffusion Implicit Models (DDIM), which can greatly reduce the time cost of the sampling process.

[0104] In the numerical experiment, the present invention uses the complete multimodal medical MR imaging data of 335 subjects, and randomly divides them into 218 samples for model training, 6 samples for model verification, and 111 samples for calculating the test accuracy. The data of each subject contains four different modal images, namely T1 images, T1 enhanced images, T2 images and FLAIR images, with a total of 14 random missing situations. For each missing situation, the present invention generates all missing modal images based on the collected medical images, and the test accuracy under this missing situation is the average accuracy over all test samples and all generated images.

[0105] As shown in Table 1, the visual cue diffusion generation model of the present invention is compared with a single-to-single medical image generation method (HyperGAN), a multi-to-single medical image generation method (CollaGAN) and three unified medical image generation methods (MM-GAN, ResViT, HyperGAE) in fourteen modality missing situations. The visual cue diffusion generation model designed by the present invention achieves the best average reconstruction accuracy. Figure 5 (a), (b) and (c) are visualization results of the missing modality medical image generation of the present invention. It can be seen that the present invention can accurately generate the missing modality medical image.

[0106] Table 1 Comparison results of different medical image generation methods on the test set under different modality missing conditions

[0107]

[0108]

[0109] Referring to Table 1, it can be seen that the medical image generation accuracy of the present invention is high.

[0110] The present invention provides a unified visual cue diffusion generation model, which can handle the problem of medical image generation in the case of any missing modality, and achieves higher quality of missing modality image generation. Compared with the existing cross-modality medical image generation method, the method of the present invention trains a unified medical image generation model based on the diffusion generation model, which can handle the problem of medical image generation in the case of any missing modality. Compared with the existing unified generation method of missing modality medical images, the method of the present invention is based on a more advanced diffusion generation model, adopts a visual cue extraction subnetwork and a residual enhanced U-Net network, and realizes the unified generation of medical images in the case of any missing modality. It can generate higher quality medical images, and the sampling generation process is more intuitive and visible.

[0111] See also Figure 7 In another embodiment of the present invention, a unified diffusion generation system for medical images with random modality loss is provided, comprising:

[0112] Acquire one or more missing modality medical images and corresponding modality acquisition codes;

[0113] The acquired medical images and modality acquisition codes are fed into a pre-trained visual cue diffusion generation model to obtain the missing modality medical images;

[0114] Among them, the pre-trained visual cue diffusion generation model includes a visual cue extraction subnetwork and a U-Net network. The visual cue extraction subnetwork is used to extract visual cues from medical images according to the modality acquisition coding; the residual enhanced U-Net network is used to iteratively denoise the acquired random Gaussian noise based on visual cues and target modality coding until a medical image with missing modality is generated.

[0115] All relevant contents of each step involved in the embodiment of the aforementioned unified diffusion generation method for medical images with random modality missing can be referred to the functional description of the functional modules corresponding to the unified diffusion generation system for medical images with random modality missing in the embodiment of the present invention, and will not be repeated here.

[0116] In another embodiment of the present invention, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the medical image unified diffusion generation method for random modality loss when executing the computer program. The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc., which are the computing core and control core of the terminal, and are suitable for implementing one or more instructions, and are specifically suitable for loading and executing one or more instructions in a computer storage medium to implement the corresponding method flow or corresponding function; the processor described in the embodiment of the present invention can be used for the operation of the epilepsy signal detection method in the electroencephalogram signal.

[0117] In another embodiment of the present invention, a computer-readable storage medium is provided, specifically a computer-readable storage medium (Memory), which is a memory device in a computer device for storing programs and data. It is understandable that the computer-readable storage medium here can include both built-in storage media in the computer device and, of course, extended storage media supported by the computer device. The computer-readable storage medium provides a storage space, which stores the operating system of the terminal. In addition, one or more instructions suitable for being loaded and executed by the processor are also stored in the storage space, and these instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. The processor can load and execute one or more instructions stored in the computer-readable storage medium to implement the corresponding steps of the unified diffusion generation method for medical images with random modality loss in the above embodiment.

[0118] The above description is only for the best embodiment of the present invention, but it should not be understood as limiting the claims. The present invention is not limited to the above embodiments, and its specific structure is allowed to be changed. However, all changes made within the protection scope of the independent claims of the present invention are within the protection scope of the present invention.

[0119] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more related listed items.

Claims

1. A unified diffusion generation method for medical images with random modality loss, characterized in that: The steps include: Acquire one or more missing modality medical images and corresponding modality acquisition codes; The acquired medical images and modality acquisition codes are fed into a pre-trained visual cue diffusion generation model to obtain the missing modality medical images; Among them, the pre-trained visual cue diffusion generation model includes a visual cue extraction subnetwork and a residual enhanced U-Net network. The visual cue extraction subnetwork is used to extract visual cues from medical images according to the modality acquisition coding; the residual enhanced U-Net network is used to iteratively denoise the acquired random Gaussian noise based on visual cues and target modality coding until a missing modality medical image is generated.

2. The unified diffusion generation method for random modality missing medical images according to claim 1, characterized in that: The visual cue extraction subnetwork includes a main network and a modality acquisition code embedding supernetwork, which takes medical images and modality acquisition codes as inputs respectively, where the modality acquisition code is a 0-1 code to indicate the corresponding modality acquisition or absence; the main network includes three 2D convolutional layers connected in sequence, four first repeating units and a first residual module; The modality acquisition coding embedding hypernetwork consists of five fully connected layers connected sequentially; 2D convolutional layer, used to extract features from medical images; A first repeating unit, for extracting multi-scale visual cues based on features, comprising a second residual module and a downsampling layer; A first residual module for extracting the lowest-scale visual cue from the output of the first repeating unit; The fully connected layer is used to generate parameters according to the modal acquisition encoding and adjust the first residual module and the second residual module in the corresponding main network.

3. The unified diffusion generation method for medical images with random modality loss according to claim 1, characterized in that: The residual-enhanced U-Net network consists of a main network and a target modality encoding embedding super network; The target modality coding embedding hypernetwork includes seventeen fully connected layers connected in sequence, taking the target modality coding as input and generating parameters for adjusting the corresponding residual module in the main network, wherein the target modality coding is a one-hot coding, which is used to indicate the missing modality medical image to be generated.

4. The unified diffusion generation method for random modality missing medical images according to claim 3 is characterized in that: The main network in the residual enhanced U-Net network takes noisy images and visual cues as input to produce denoised images; the main network consists of an encoder, four skip connections and a decoder, with an overall U-shaped symmetrical structure; each skip connection contains two convolutional layers, connecting from the residual module of the encoder to the splicing layer of the corresponding resolution of the decoder.

5. The unified diffusion generation method for medical images with random modality loss according to claim 4, characterized in that: The encoder of the main network is used to extract multi-resolution noisy image features, including a 2D convolutional layer, four second repeating units, a concatenation layer, and a third residual module; 2D convolutional layer, used to extract initial features from noisy images; A second repeating unit, for extracting multi-resolution noise image features according to the initial features, comprises a concatenation layer, a fourth residual module and a downsampling layer; A splicing layer, used for splicing the noise image features output by the second repeating unit and the visual prompts of the corresponding resolution; The third residual module is used to extract the lowest resolution noise image features based on the output of the concatenated layer.

6. The unified diffusion generation method for medical images with random modality loss according to claim 4, characterized in that: The decoder of the main network is used to generate a denoised image according to multi-resolution noise image features, including a fifth residual module, a sixth residual module, four third repeating units, a seventh residual module, an eighth residual module, a 2D convolutional layer, and a tanh activation function; a fifth residual module, for extracting features based on the noise image features of the lowest resolution; a sixth residual module, used for extracting denoised image features of the lowest resolution according to the output of the fifth residual module; A third repeating unit, for extracting higher-resolution denoised image features according to the output of the sixth residual module, comprising a ninth residual module, an upsampling layer, and a splicing layer; The seventh residual module is used to extract the denoised image features with the highest resolution according to the output of the third repeating unit, The eighth residual module is used to extract more complete denoising image features according to the output of the seventh residual module. A 2D convolutional layer for extracting an unnormalized denoised image based on the output of the eighth residual module; The tanh activation function is used to map the output of the 2D convolutional layer to between -1 and 1 and output a denoised image.

7. The unified diffusion generation method for random modality missing medical images according to claim 1, characterized in that: Loss function for a pre-trained generative model of diffusion of visual cues for: Among them, ‖·‖2 represents the 2-norm, N is the number of modes, and m i is the ith mode, For a clean image, is the noise image at time t, F vp (X obs ) is a visual cue, X obs For medical images, c i is the target modality encoding, and t is the time.

8. A unified diffusion generation system for medical images with random modality loss, characterized in that: include: Acquire one or more missing modality medical images and corresponding modality acquisition codes; The acquired medical images and modality acquisition codes are fed into a pre-trained visual cue diffusion generation model to obtain the missing modality medical images; Among them, the pre-trained visual cue diffusion generation model includes a visual cue extraction subnetwork and a residual enhanced U-Net network. The visual cue extraction subnetwork is used to extract visual cues from medical images according to the modality acquisition coding; the residual enhanced U-Net network is used to iteratively denoise the acquired random Gaussian noise based on visual cues and target modality coding until a missing modality medical image is generated.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the unified diffusion generation method for medical images with random modality missing as described in any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the unified diffusion generation method for medical images with random modality missing as described in any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Hyperspectral snapshot compression imaging reconstruction method fusing conditional diffusion

    CN120976433A

  • A hyperspectral snapshot compressive imaging reconstruction method fusing conditional diffusion

    CN120976433B