Artificial intelligence-based multi-modal brain 3d-mri image generation method

By constructing an unconditional diffusion model and a noise prediction network, and combining diffusion noise prediction loss and cross-modal consistency constraint loss, the problem of anatomical inconsistency in multimodal MRI image generation is solved, achieving high-quality multimodal image generation, improving the reliability of clinical applications and the ability to support multi-center studies.

CN122265473BActive Publication Date: 2026-08-25INNOVATION ACAD FOR PRECISION MEASUREMENT SCI & TECH CAS
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610727197.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-25
Publication Date
2026-08-25
Estimated Expiration
2046-05-25

AI Technical Summary

Technical Problem

Existing technologies struggle to generate multimodal MRI images with consistent anatomical structures, leading to issues such as inconsistent anatomical structures, boundary misalignment, and morphological drift in multimodal applications. This affects the reliability of clinical interpretation and subsequent automated analysis tasks.

Method used

An AI-based multimodal brain 3D-MRI image generation method is adopted. By constructing an unconditional diffusion model and a noise prediction network, and combining diffusion noise prediction loss and cross-modal consistency constraint loss, the anatomical structure consistency generation of multimodal images is achieved.

Benefits of technology

The generated quadmodal 3D brain MRI images can improve anatomical consistency and stability in scenarios where multimodal data is scarce or acquisition is limited, expand training samples, support multi-center data research, and enhance the training effect and robustness of downstream segmentation and detection tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122265473B_ABST
    Figure CN122265473B_ABST
Patent Text Reader

Abstract

The application discloses a multi-modal brain 3D-MRI image generation method based on artificial intelligence, and relates to the technical field of medical image processing, in particular to a multi-modal brain 3D-MRI image generation method based on artificial intelligence. The method comprises the following steps: preprocessing multi-modal 3D-MRI original images, stacking the multi-modal 3D-MRI original images into corresponding multi-channel 3D data samples through channel dimensions to construct a training set, and constructing an unconditional diffusion model comprising a denoising diffusion probability model and a noise prediction network; constructing a time step loss function comprising a diffusion noise prediction loss and a cross-modal consistency constraint loss, wherein the cross-modal consistency constraint loss comprises a boundary consistency loss and a structure consistency loss. The training set constructed by the multi-channel 3D data samples is used to train the unconditional diffusion model, so that multi-modal 3D-MRI generated data can be directly generated from random noise, and the dependence on external condition input is avoided; through the cross-modal consistency constraint in the whole time step, the structural misplacement, boundary drift and inconsistent morphology among modes are effectively inhibited, and the anatomical consistency and stability of the generated data are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of nuclear magnetic resonance technology, specifically relating to a method for generating multimodal brain 3D-MRI images based on artificial intelligence. Background Technology

[0002] Magnetic resonance imaging (MRI) offers advantages such as high resolution for soft tissues and the absence of ionizing radiation. Multiple sequences or modalities of MRI are typically used to image the same anatomical region, such as non-contrast T1-weighted (T1n), contrast-enhanced T1-weighted (T1c), and T2-weighted (T2w). While different modalities share the same biological origin and spatially reflect highly consistent anatomical structures, they differ in imaging mechanisms, tissue contrast, and lesion visualization characteristics. Therefore, multimodal MRI can characterize the same anatomical structures and pathological regions from complementary perspectives, providing more comprehensive information for research and improving robustness and accuracy in subsequent tasks such as automated segmentation, registration, detection, and quantitative analysis. However, in practical applications, the acquisition of multimodal images is often affected by factors such as scan time, subject tolerance, contrast agent contraindications, differences in equipment protocols, motion artifacts, or cost, leading to difficulties in obtaining multimodal data simultaneously or inconsistent quality. Therefore, utilizing artificial intelligence models to generate high-quality multimodal images, especially generating anatomically consistent multimodal MRI images, is of great significance for improving the availability of clinical data, supporting multicenter studies, and enhancing downstream algorithm training.

[0003] In existing image generation technologies, many methods focus on modeling and generating images for a single modality, such as generating only a sequence of MRI images or only single-channel medical images. While these single-modal generation methods can learn the overall appearance distribution of the target modality, they have significant shortcomings in multimodal applications. On the one hand, single-modal models cannot explicitly model the correspondence and complementary information between different modalities, making it difficult to simultaneously meet the needs of multimodal joint analysis. On the other hand, when multiple modalities are generated by independent single-modal generation models, the lack of unified structural constraints between different modalities can easily lead to problems such as inconsistencies in anatomical structures, boundary misalignments, morphological drift, or differences in local details, thereby reducing the reliability of multimodal images in clinical interpretation and subsequent automated analysis tasks. Therefore, there is an urgent need for a multimodal image generation method that can ensure the consistency of anatomical structures in the generation results of different modalities through an effective cross-modal consistency mechanism, thereby improving the usability and reliability of multimodal generated images and realizing the joint generation of multimodal medical images. Summary of the Invention

[0004] The purpose of this invention is to address the aforementioned problems in the existing technology by providing a method for generating multimodal brain 3D-MRI images based on artificial intelligence.

[0005] The above-mentioned objectives of the present invention are achieved by the following technical means: A method for generating multimodal brain 3D-MRI images based on artificial intelligence includes the following steps: Step 1: For each of the multiple subjects, acquire the corresponding multimodal 3D-MRI raw images; for each subject: spatially align, unify the size and normalize the intensity of the 3D-MRI raw images corresponding to all modalities, assign one modality to one channel, and stack the 3D-MRI raw images of all modalities according to the channel dimension to form a multi-channel 3D data sample; the training set is composed of multiple multi-channel 3D data samples; Step 2: Construct an unconditional diffusion model, including a denoising diffusion probability model and a noise prediction network. The denoising diffusion probability model includes a forward module and a backward module. Step 3: Construct the time step loss function, which includes the spread noise prediction loss and the cross-modal consistency constraint loss. The cross-modal consistency constraint loss includes the boundary consistency loss and the structural consistency loss. Step 4: Train the noise prediction network in the unconditional diffusion model using the training set and the time step loss function to obtain the trained noise prediction network. Step 5: Input the noisy multi-channel 3D data samples to be processed into the inverse module and the trained noise prediction network, perform inverse denoising iteration, and split the corresponding multimodal 3D-MRI generated data into 3D-MRI generated images of each modality according to the channel.

[0006] The specific structure of the unconditional diffusion model described above is as follows: For each preset time step Multi-channel 3D data samples The input to the forward module is to output the corresponding noisy 3D-MRI data; the noise prediction network performs noise prediction on the noisy 3D-MRI data; the backward module, based on the corresponding predicted noise, infers the predicted 3D-MRI data corresponding to the previous time step of the noisy 3D-MRI data; where... , Maximum time step; When time step The predicted 3D-MRI data from the previous time step is multimodal 3D-MRI generated data; During training, noisy 3D-MRI data were taken at the same time step. Noisy 3D-MRI data; During the application process after training, when the time step The corresponding noisy 3D-MRI data is the noisy multi-channel 3D data sample to be processed, when the time step Time step The corresponding noisy 3D-MRI data is: based on time steps Time step inferred from noisy 3D-MRI data The corresponding predicted 3D-MRI data.

[0007] As mentioned above, the prediction loss for diffused noise is: , in, To predict the loss for diffused noise, For time steps The corresponding external noise is denoted as external noise. The noise added at each time step is random noise sampled from a standard normal distribution, and noise is added independently for each channel; To predict noise; Represents the square of the L2 norm. For noisy 3D-MRI data.

[0008] The calculation steps for the cross-modal consistency constraint loss, as described above, are as follows: Step A: Calculate the back-predicted clean body data based on the following formula: , Predict clean body data The predicted clean volume data for each modality is obtained by splitting the data by channel. , Noise scheduling parameters during diffusion for adding noise to 3D-MRI data Noise scheduling parameters during the diffusion process , To preset noise scheduling parameters, ; Serial Number Time step Modal number , The total number of modes; Step B: Calculate the boundary consistency loss according to the following steps: Predicted clean volume data for each mode Calculate the three-dimensional gradient: , , , in, , , They are exist , , Three-dimensional gradient components in three spatial directions; , , They are , , Gradient operators in three spatial directions; Define intermediate quantities : , 3D boundary structure diagram for: , in, A preset constant introduced to prevent the denominator from being zero; This indicates that the average value of all corresponding pixels is taken. The boundary consistency loss is calculated based on the following formula: , in, Norm, intermediate quantity , Indicates from The number of combinations of choosing 2 modes from 3 modes. For the total number of modes, , Indicates the modal number. ; For three-dimensional boundary structure diagrams correspond of Pick ; Three-dimensional boundary structure diagram correspond of Pick ; The structural consistency loss is calculated based on the following steps: Compute structure embedding vector : , For structural encoders; Calculate structural consistency loss based on the following formula. : , , in, Represents the square of the L2 norm. For average structure embedding vectors; Step C: Calculate the cross-modal consistency constraint loss based on the following formula. : , in, , These are the first weighting coefficient and the second weighting coefficient, respectively.

[0009] As described above, the time-step loss function for: , This is the third weighting coefficient.

[0010] As described above, the structure encoder sequentially includes a first three-dimensional convolutional layer, a first normalization layer, a first activation function layer; a second three-dimensional convolutional layer, a second normalization layer, a second activation function layer; a third three-dimensional convolutional layer, a third normalization layer, a third activation function layer; and a global average pooling layer.

[0011] As described above, the noisy 3D-MRI data is calculated based on the following formula: , In the formula, Noise scheduling parameters during diffusion for adding noise to 3D-MRI data Noise scheduling parameters during the diffusion process , To preset noise scheduling parameters, ; Serial Number ; This is a multi-channel 3D data sample; For time steps The corresponding external noise is denoted as external noise. External noise at each time step For random noise sampled from a standard normal distribution, noise is added independently to each channel; The predicted 3D-MRI data is calculated based on the following formula: , in, To predict 3D-MRI data, the corresponding time step ; For noisy 3D-MRI data, For random noise sampled according to a standard normal distribution, the noise scheduling parameters during the diffusion process. , To preset noise scheduling parameters, ; This refers to the standard deviation of the noise added during the reverse denoising process. Predicting 3D-MRI data Generate data for multimodal 3D-MRI; For the training process, time step Corresponding noisy 3D-MRI data ; For the application process after training, when the time step Noisy 3D-MRI data For the noisy multi-channel 3D data samples to be processed, when the time step Noisy 3D-MRI data .

[0012] As described above, the noise prediction network includes a generator encoder, a generator decoder, and skip connections: The generative encoder includes multiple downsampling layers connected in sequence. Each downsampling layer includes a first normalization layer, a first convolutional layer, a second normalization layer, a second convolutional layer, an activation function layer, and a pooling layer connected in sequence. The generator decoder includes multiple upsampling layers connected in sequence. Each upsampling layer includes a first normalization layer, a first convolutional layer, a second normalization layer, a second convolutional layer, an activation function layer, and an upsampling convolutional module connected in sequence. The jump connection is located between the pooling layer output of each downsampling layer and the input of the first normalized layer of the corresponding upsampling layer.

[0013] As described above, step 4 specifically includes the following steps: Several multi-channel 3D data samples were randomly sampled from the training set. This constitutes a training batch for each multi-channel 3D data sample. , corresponding to the set of time steps Each time step For multi-channel 3D data samples Perform forward noise addition to obtain noisy 3D-MRI data. The noisy 3D-MRI data was then processed. With time step The noise is input into the noise prediction network to obtain the predicted noise. ;Calculate time steps Corresponding time-step loss function ; Loss function for time steps Backpropagation is performed to calculate the parameter gradients of the noise prediction network in the unconditional diffusion model, and the optimizer is used to update the parameters of the noise prediction network in the unconditional diffusion model.

[0014] A computer device includes a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement steps 2 to 5 of any of the artificial intelligence-based multimodal brain 3D-MRI image generation methods.

[0015] Compared with the prior art, the present invention has the following advantages: (1) The present invention stacks the four-modal 3D brain MRI images of T1n, T1c, T2f and T2w into a four-channel whole for diffusion modeling, and uses an unconditional diffusion model to directly generate multimodal 3D-MRI generation data from random noise, avoiding dependence on external condition input (such as text, labels or reference modal), and improving the universality and usability of the generation process.

[0016] (2) The present invention performs reverse denoising at all time steps. A cross-modal consistency constraint module is introduced, which constrains the spatial position of the anatomical boundary of different modalities to be consistent through 3D boundary consistency loss. The structural encoder extracts modality-independent structural embeddings to calculate structural consistency loss, thereby significantly reducing cross-modal structural misalignment, boundary drift and morphological inconsistency, and improving the anatomical consistency and stability of the four-modal generated data.

[0017] (3) The four-modal 3D brain MRI images generated by the present invention can be used to expand training samples, support multi-center data research in scenarios where multimodal data is scarce or acquisition is limited, and improve the training effect and robustness of downstream segmentation, detection and other tasks, and have high application value. Attached Figure Description

[0018] Figure 1 This is a schematic diagram of the method flow of the present invention.

[0019] Figure 2 This is a schematic diagram of the forward module and the reverse module of the present invention.

[0020] Figure 3 This is a schematic diagram of the noise prediction network of the present invention.

[0021] Figure 4 This is a schematic diagram of the downsampling layer of the present invention.

[0022] Figure 5 This is a schematic diagram of the upsampling layer of the present invention.

[0023] Figure 6This section compares the performance of the present invention with other methods. T1n, T1c, T2f, and T2w represent four different MRI modalities; VAE, CycleGAN, DDPM, and DDIM represent different image generation methods; VAE is a variational autoencoder; CycleGAN is a recurrent adversarial generative network; DDPM is a denoising diffusion probabilistic model; DDIM is a denoising diffusion implicit model; and ours represents the method of the present invention. (a0) is the original image of the T1c mode, and (a1) to (a5) are the generation results of each image generation method for the original image of the T1c mode; (b0) is the original image of the T1n mode, and (b1) to (b5) are the generation results of each image generation method for the original image of the T1n mode; (c0) is the original image of the T2f mode, and (c1) to (c5) are the generation results of each image generation method for the original image of the T2f mode; (d0) is the original image of the T2w mode, and (d1) to (d5) are the generation results of each image generation method for the original image of the T2w mode. Detailed Implementation

[0024] To facilitate understanding and implementation of the present invention by those skilled in the art, the following description, in conjunction with embodiments, Figures 1-6 Table 1 provides a further detailed description of the present invention. It should be understood that the embodiments described herein are for illustrative and explanatory purposes only and are not intended to limit the present invention.

[0025] Example 1 Artificial intelligence-based multimodal brain 3D-MRI image generation methods, such as Figure 1 As shown, the specific steps include: Step 1: Obtain the corresponding multimodal 3D-MRI raw images for each of the multiple subjects. For each subject: Spatially align, unify the size, and normalize the intensity of the 3D-MRI raw images corresponding to all modalities. Assign one modality to one channel, and stack all the 3D-MRI raw images of all modalities according to the channel dimension to form a multi-channel 3D data sample. The training set is composed of multiple multi-channel 3D data samples. In this embodiment, the multimodal 3D-MRI raw images are 3D-MRI raw images of the brain in four modalities. All multi-channel 3D data samples are divided into training set and test set, specifically including the following steps: Step 1.1: Acquire multimodal 3D-MRI raw images of multiple subjects. In this embodiment, the total number of modalities is 4, and the modal types include T1n modality, T1c modality, T2f modality, and T2w modality.

[0026] Step 1.2: Preprocess the multimodal 3D-MRI raw images of each subject into multi-channel 3D data samples, specifically including the following steps: Step 1.2.1: Register or spatially align the original 3D-MRI images of all modalities of the same subject to the same spatial coordinate system to obtain spatially aligned 3D-MRI images.

[0027] This embodiment performs registration or spatial alignment processing on the original 3D-MRI images of the same subject in four modalities: T1n, T1c, T2f, and T2w.

[0028] Step 1.2.2: For the same subject, the spatially aligned 3D-MRI images are cropped, resampled, or filled to the same voxel spacing and the same size to obtain 3D-MRI images of consistent size.

[0029] In this embodiment, the spatially aligned 3D-MRI images of the same subject are cropped, resampled, or filled to ensure that the voxel spacing and size of the 3D-MRI images of each modality are consistent. In this embodiment, the spatial size of the 3D-MRI images of each modality is uniformly 128×128×128.

[0030] Step 1.2.3: For the same subject, intensity standardization or normalization is performed on 3D-MRI images of the same size to obtain preprocessed 3D-MRI images of all modalities.

[0031] This embodiment performs intensity normalization or intensity standardization on 3D-MRI images of uniform size across four modalities to reduce intensity distribution differences caused by different scanning protocols or imaging conditions.

[0032] Step 1.2.4: For the same subject, stack the preprocessed 3D-MRI images corresponding to all modalities in steps 1.2.1 to 1.2.3 according to channel dimension to obtain a multi-channel 3D data sample. ;in, For the total number of modes, Indicates the number of slices. Indicates the number of vertical pixels in the image. This indicates the number of horizontal pixels in the image.

[0033] This embodiment obtains multi-channel 3D data samples. The four channels correspond to the T1n mode, T1c mode, T2f mode, and T2w mode, respectively.

[0034] Step 1.3: The training set is composed of multiple multi-channel 3D data samples.

[0035] This embodiment includes multi-channel 3D data samples from all subjects. The samples are randomly divided into training and test sets, and the samples are used as input to the unconditional diffusion model.

[0036] Step 2: Construct an unconditional diffusion model with noise prediction as the objective. The unconditional diffusion model includes a denoising diffusion probability model (DDPM) and a noise prediction network. The denoising diffusion probability model includes a forward module and a backward module. For each preset time step Multi-channel 3D data samples The input to the forward module is the corresponding noisy 3D-MRI data; the noise prediction network predicts noise for the noisy 3D-MRI data; the backward module, based on the predicted noise, infers the predicted 3D-MRI data from the previous time step corresponding to the noisy 3D-MRI data; where... , This represents the maximum time step.

[0037] When time step The predicted 3D-MRI data from the previous time step is multimodal 3D-MRI generated data.

[0038] During training, noisy 3D-MRI data were taken at the same time step. Noisy 3D-MRI data (i.e., multi-channel 3D data samples) (Data with added noise).

[0039] During the application process after training, when the time step The corresponding noisy 3D-MRI data are noisy multi-channel 3D data samples, when the time step and Time step The corresponding noisy 3D-MRI data is: based on time steps Time step inferred from noisy 3D-MRI data The corresponding predicted 3D-MRI data.

[0040] 1. Denoising diffusion probability model: 1.1 Forward module in the denoising diffusion probability model: For each preset time step The forward module processes multi-channel 3D data samples. Time step Add noise and obtain the corresponding noisy 3D-MRI data When time steps Reach the preset maximum time step The corresponding noisy 3D-MRI data For multi-channel pure noise volume data Among them, noisy 3D-MRI data : .

[0041] In the formula, the noise scheduling parameters during the diffusion process are... Noise scheduling parameters during the diffusion process , To preset noise scheduling parameters, ; Serial Number Time step ; This is a multi-channel 3D data sample; For time steps The corresponding external noise is denoted as external noise. External noise at each time step This refers to random noise sampled from a standard normal distribution, with noise added independently to each channel. It is an identity matrix.

[0042] 1.2 The inverse module in the denoising diffusion probability model: For each time step Noisy 3D-MRI data The input includes a reverse module and a noise prediction network, which is based on noisy 3D-MRI data. and time step Predict the current time step Predicted noise The reverse module is based on the corresponding pre-defined noise. For the corresponding noisy 3D-MRI data Noise removal is performed to obtain the time step. Predictive 3D-MRI data : , in, This represents random noise sampled according to a standard normal distribution. and These are noise scheduling parameters during the diffusion process. , To preset noise scheduling parameters, ; This is the noise standard deviation added during the reverse denoising process, used to control the randomness of sampling.

[0043] During the back diffusion process, the time step starts from... Gradually decrease to When time steps At that time, the previous time step predicted Predictive 3D-MRI data Generate data for multimodal 3D-MRI.

[0044] For the training process, time step Corresponding noisy 3D-MRI data ,like Figure 2 As shown; for the application process after training, when time step Noisy 3D-MRI data For the noisy multi-channel 3D data samples to be processed, when the time step Noisy 3D-MRI data .

[0045] 2. Noise prediction network: The noise prediction network adopts a three-dimensional U-Net architecture, such as... Figure 3 As shown, it includes a generator encoder and a generator decoder in sequence, with skip connections set between the generator encoder and the generator decoder; both the generator encoder and the generator decoder are composed of three-dimensional convolutional layers.

[0046] As one possible implementation, the generator encoder includes multiple sequentially connected downsampling layers, such as... Figure 4 As shown, each downsampling layer includes a first normalization layer, a first convolutional layer, a second normalization layer, a second convolutional layer, an activation function layer, and a pooling layer connected in sequence; the generator decoder includes multiple upsampling layers connected in sequence, such as... Figure 5 As shown, each upsampling layer includes a first normalization layer, a first convolutional layer, a second normalization layer, a second convolutional layer, an activation function layer, and an upsampling convolutional module connected in sequence. Skip connections are located between the pooling layer output of each downsampling layer and the input of the first normalization layer of the corresponding upsampling layer, and are used to concatenate or fuse the features output by each downsampling layer with the features of the upsampling layer at the corresponding scale.

[0047] Noisy 3D-MRI data The output of the pooling layer of the first downsampling layer is fed into the first normalization layer of the first upsampling layer; the output of the pooling layer of each downsampling layer (excluding the last downsampling layer) is fed into the first normalization layer of the corresponding upsampling layer; the output features of the upsampling convolutional modules of each upsampling layer (excluding the last upsampling layer) are fed into the first normalization layer of the next upsampling layer; and the output of the upsampling convolutional module of the last upsampling layer is the predicted noise output by the noise prediction network. .

[0048] In this embodiment, the generator decoder includes four downsampling layers connected in sequence, and the generator decoder also includes four upsampling layers connected in sequence. The generator encoder takes samples as input. In the first downsampling layer: the first normalization layer normalizes the input features using a batch normalization function; the first convolutional layer has a kernel size of 3×3, a stride of 1, a padding stride of 1, 4 input channels, 16 output channels, an input feature size of 4×128×128×128, and an output feature size of 16×128×128×128; the second normalization layer normalizes the input features using a batch normalization function; the second convolutional layer has a kernel size of 3×3, a stride of 1, a padding stride of 1, and 1 input channel. 6. The output channel is 32, the input feature size is 16×128×128×128, and the output feature size is 32×128×128×128. The activation function layer uses the ReLU function to perform non-linear activation on the input features. The pooling layer has a convolution kernel size of 2×2, a stride of 2, and no padding. The input channel is 32, the output channel is 32, the input feature size is 32×128×128×128, and the output feature size is 32×64×64×64. The output feature of the pooling layer is used as the output feature of the first downsampling layer.

[0049] The output features of the first downsampling layer of the generator encoder are used as the input features of the normalization layer of the second downsampling layer of the generator encoder. In the second downsampling layer: the first normalization layer uses a batch normalization function to normalize the input features; the first convolutional layer has a kernel size of 3×3, a stride of 1, a padding stride of 1, 32 input channels, and 32 output channels; the input feature size is 32×64×64×64, and the output feature size is 32×64×64×64. The second normalization layer uses a batch normalization function to normalize the input features; the second convolutional layer has a kernel size of 3×3, a stride of 1, a padding stride of 1, and 32 input channels. The input feature size is 32×64×64×64, the output feature size is 64×64×64×64, the activation function layer uses the ReLU function to perform non-linear activation on the input features, the pooling layer has a convolution kernel size of 2×2, a stride of 2, no padding, an input feature size of 64×64×64×64, and an output feature size of 64×32×32×32. The output features of the pooling layer are used as the output features of the first downsampling layer.

[0050] The output feature map of the second downsampling layer of the generator encoder is used as the input feature map of the normalization layer of the third downsampling layer of the generator encoder. In the third downsampling layer: the first normalization layer uses a batch normalization function to normalize the input features; the first convolutional layer has a kernel size of 3×3, a stride of 1, a padding stride of 1, 64 input channels, 64 output channels, and an input feature size of 64×32×32×32 and an output feature size of 64×32×32×32; the second normalization layer uses a batch normalization function to normalize the input features; the second convolutional layer has a kernel size of 3×3, a stride of 1, a padding stride of 1, and 64 input channels. The input feature size is 64×32×32×32, and the output feature size is 128×32×32×32. The activation function layer uses the ReLU function to perform non-linear activation on the input features. The pooling layer has a convolution kernel size of 2×2, a stride of 2, and no padding. The input feature size is 128, and the output feature size is 128×16×16×16. The output features of the pooling layer are used as the output features of the first downsampling layer.

[0051] The output features of the third downsampling layer of the generative encoder are used as the input features of the normalization layer of the fourth downsampling layer of the generative encoder. In the fourth downsampling layer: the first normalization layer uses a batch normalization function to normalize the input features; the first convolutional layer has a kernel size of 3×3, a stride of 1, a padding stride of 1, 128 input channels, and 128 output channels; the input feature size is 128×16×16×16, and the output feature size is 128×16×16×16. The second normalization layer uses a batch normalization function to normalize the input features; the second convolutional layer has a kernel size of 3×3, a stride of 1, a padding stride of 1, and 128 input channels. The input feature size is 128×16×16×16, the output feature size is 256×16×16×16, the activation function layer uses the ReLU function to perform non-linear activation on the input features, the pooling layer has a convolution kernel size of 2×2, a stride of 2, no padding, an input feature size of 256×16×16×16, and an output feature size of 256×8×8×8. The output features of the pooling layer are used as the output features of the first downsampling layer.

[0052] In this embodiment, the generator decoder includes four sequentially connected upsampling layers. Each of the four upsampling layers includes a first normalization layer, a first convolutional layer, a second normalization layer, a second convolutional layer, an activation function layer, and an upsampling convolutional module, all connected in sequence.

[0053] The output of the fourth downsampling layer of the generator encoder is used as the input to the first normalization layer of the first upsampling layer of the generator decoder. In the first upsampling layer: the first normalization layer normalizes the input features using a batch normalization function; the first convolutional layer has a 3×3 kernel size, a stride of 1, a padding stride of 1, 256 input channels, 256 output channels, an input feature size of 256×8×8×8, and an output feature size of 256×8×8×8. The second normalization layer normalizes the input features using a batch normalization function; the second convolutional layer has a 3×3 kernel size, a stride of 1, a padding stride of 1, and 256 input channels. 6. The output channel is 128, the input feature size is 256×8×8×8, and the output feature size is 128×8×8×8. The activation function layer uses the ReLU function to non-linearly activate the input. The upsampling convolution module has a kernel size of 2×2, a stride of 2, and no padding. The input channel is 128, the output channel is 128, the input feature size is 128×8×8×8, and the output feature size is 128×16×16×16. The output features of the upsampling convolution module are used as the output features of the generator decoder.

[0054] The output features of the first upsampling layer of the generator decoder serve as the input features of the first normalization layer of the second upsampling layer of the generator decoder. In the second upsampling layer: the first normalization layer uses a batch normalization function to normalize the input features; the first convolutional layer has a 3×3 kernel size, a stride of 1, a padding stride of 1, 128 input channels, 128 output channels, an input feature size of 128×16×16×16, and an output feature size of 128×16×16×16. The second normalization layer uses a batch normalization function to normalize the input features; the convolutional layer has a 3×3 kernel size, a stride of 1, a padding stride of 1, and 1 input channel. 28, with 64 output channels, an input feature size of 128×16×16×16, and an output feature size of 64×16×16×16. The activation function layer uses the ReLU function for non-linear activation of the input. The upsampling convolutional module has a kernel size of 2×2, a stride of 2, and no padding. It has 64 input channels, 64 output channels, an input feature size of 64×16×16×16, and an output feature size of 64×32×32×32. The output features of the upsampling convolutional module are used as the output features of the generator decoder.

[0055] The output features of the second upsampling layer of the generator decoder serve as the input features of the first normalization layer of the third upsampling layer of the generator decoder. In the third upsampling layer: the first normalization layer uses a batch normalization function to normalize the input features; the first convolutional layer has a 3×3 kernel size, a stride of 1, a padding stride of 1, 64 input channels, 64 output channels, and an input feature size of 64×32×32×32, and an output feature size of 64×32×32×32. The second normalization layer uses a batch normalization function to normalize the input features; the second convolutional layer has a 3×3 kernel size, a stride of 1, a padding stride of 1, and 6 input channels. 4. The output channel is 32, the input feature size is 64×32×32×32, and the output feature size is 32×32×32×32. The activation function layer uses the ReLU function to non-linearly activate the input. The upsampling convolution module has a kernel size of 2×2, a stride of 2, and no padding. It has 32 input channels and 32 output channels. The input feature size is 32×32×32×32, and the output feature size is 32×64×64×64. The output features of the upsampling convolution module are used as the output features of the generator decoder.

[0056] The output features of the third upsampling layer of the generator decoder serve as the input features of the first normalization layer of the fourth upsampling layer of the generator decoder. In the fourth upsampling layer: the first normalization layer uses a batch normalization function to normalize the input features; the first convolutional layer has a 3×3 kernel size, a stride of 1, a padding stride of 1, 32 input channels, and 16 output channels; the input feature size is 32×64×64×64, and the output feature size is 16×64×64×64. The second normalization layer uses a batch normalization function to normalize the input features; the second convolutional layer has a 3×3 kernel size, a stride of 1, a padding stride of 1, and 32 input channels. The input feature size is 16×64×64×64, the output feature size is 8×64×64×64, and the activation function layer uses the ReLU function to non-linearly activate the input. The upsampling convolutional module has a kernel size of 2×2, a stride of 2, no padding, 8 input channels, 4 output channels, an input feature size of 8×64×64×64, and an output feature size of 4×128×128×128. The output features of the upsampling convolutional module are used as the output features of the generator decoder.

[0057] Skip connections concatenate the features output by each downsampling layer to the features output by the corresponding upsampling layer.

[0058] Step 3: Construct the time-step loss function The time-step loss function It includes at least the spread noise prediction loss and the cross-modal consistency constraint loss, which includes the boundary consistency loss and the structural consistency loss.

[0059] 1. Predicted loss of diffused noise: At each time step of the reverse process Calculate the predicted noise External noise added at the corresponding time step during the forward process The differences between them are used to optimize noise prediction capabilities; diffuse noise prediction loss Calculated based on the following formula: .

[0060] 2. Cross-modal consistency constraint loss: To ensure the consistency of the anatomical structure of the four-modal generation results, at each time step of the reverse process... Based on prediction noise The predicted clean volume data of the four channels is reversed, and the boundary consistency loss and structural consistency loss are calculated.

[0061] Step A: Calculate the back-predicted clean body data based on the following formula: .

[0062] Predict clean body data The predicted clean volume data for each modality is obtained by splitting the data by channel. Among them, the modal number , In this embodiment, the total number of modes is [number]. , respectively corresponding to T1n, T1c, T2f, and T2w.

[0063] Step B: Calculate boundary consistency loss and structural consistency loss: 2.1 Boundary Consistency Loss: Predicted clean volume data for each mode Calculate the 3D gradient and obtain the 3D boundary structure map. Then, calculate the boundary consistency loss for the 3D boundary structure diagrams of different modalities: Predicted clean volume data for each mode Calculate the three-dimensional gradient: , , , in, , , They are exist , , The three-dimensional gradient components in three spatial directions are used to characterize the intensity changes of the modal image in the corresponding directions; , , They are , , Gradient operators in three spatial directions. Among them, , , These represent the horizontal direction, vertical direction, and slice thickness direction of the 3D-MRI image, respectively.

[0064] Define intermediate quantities : , Among them, squaring, summing, and square root operations are all pixel-by-pixel operations, resulting in intermediate values. It is three-dimensional data.

[0065] 3D boundary structure diagram for: , in, A preset constant is introduced to prevent the denominator from being zero, and is used to improve the numerical stability of the calculation process. This means taking the average value of all corresponding pixels.

[0066] The boundary consistency loss is calculated based on the following formula: , in, Norm, intermediate quantity , Indicates from The number of combinations of choosing 2 modes from 3 modes. For the total number of modes, , Indicates the modal number. In this embodiment, In the four-mode case For three-dimensional boundary structure diagrams correspond of Pick ; Three-dimensional boundary structure diagram correspond of Pick .

[0067] 2.2 Structural consistency loss: Structural consistency loss through structural encoder Extract modality-independent structure embedding vectors and constrain consistency; the structure encoder predicts clean volume data for each mode. Output structure embedding vector ,satisfy: , The structural encoder is implemented using a 3D convolutional network, which sequentially includes a first 3D convolutional layer, a first normalization layer, and a first activation function layer; a second 3D convolutional layer, a second normalization layer, and a second activation function layer; a third 3D convolutional layer, a third normalization layer, and a third activation function layer; and a global average pooling layer. In this embodiment, the kernel size of the first 3D convolutional layer is 3×3×3, with 32 channels and a stride of 1; the kernel size of the second 3D convolutional layer is 3×3×3, with 64 channels and a stride of 2; and the kernel size of the third 3D convolutional layer is 3×3×3, with 128 channels and a stride of 2. Global average pooling is performed on the 3D spatial dimension to output the structural embedding vector. The normalization layer is either GroupNorm or InstanceNorm.

[0068] Structural consistency loss is calculated based on the following formula: , , Represents the square of the L2 norm. For average structure embedding vectors, Step C, the cross-modal consistency constraint loss is calculated based on the following formula: , Time step loss function Calculated based on the following formula: , In the formula, , , These are the first weight coefficient, the second weight coefficient, and the third weight coefficient, respectively; and during the training phase, for all time steps... Calculate the loss function for the above time steps. .

[0069] Step 4: Utilize the training set and time-step loss function The noise prediction network in the unconditional diffusion model is trained to obtain the trained noise prediction network.

[0070] The unconditional diffusion model constructed in step 2 is trained end-to-end using the training set generated in step 1, at all time steps. Calculate the loss function The model parameters are then updated in reverse; after training, the trained parameters are saved to obtain the trained unconditional diffusion model. The specific training process is as follows: First, several multi-channel 3D data samples were randomly sampled from the training set. This constitutes a training batch, and the batch size can be set from 1 to 16 depending on the available memory. For each multi-channel 3D data sample... , corresponding to the set of time steps Each time step And based on the noise scheduling parameters of the diffusion process, multi-channel 3D data samples... Perform forward noise addition to obtain noisy 3D-MRI data. The noisy 3D-MRI data was then processed. With time step The noise is input into the noise prediction network and the predicted noise is obtained through forward propagation. Based on the predicted noise The difference between the predicted noise and the actual noise is used to calculate the spread noise prediction loss. Simultaneously, based on the predicted clean data, cross-modal consistency constraint losses are calculated, including boundary consistency loss and structural consistency loss. These losses are then weighted and combined to obtain the time step. Corresponding time-step loss function loss function for time steps Backpropagation is performed to calculate the parameter gradients of the noise prediction network in the unconditional diffusion model, and an optimizer is used to update the parameters of the noise prediction network in the unconditional diffusion model. This is done for each multi-channel 3D data sample. Each time a time step is selected For random sampling, the final time step set is traversed. Each time step The optimizer employs stochastic gradient descent (SGD) or an adaptive optimization algorithm, such as the Adam algorithm or the AdamW algorithm, with the learning rate set at... Within a certain range, the learning rate can be adjusted using a decay strategy. During training, the training set is divided into multiple batches, and the forward propagation, loss calculation, and backpropagation processes described above are executed sequentially for each batch to complete one full training epoch. The number of training epochs can be set from 50 to 500. Training stops when the preset number of training epochs is reached or the model loss function converges, and the current model parameters are saved, resulting in a trained unconditional diffusion model for subsequent multimodal 3D MRI image generation.

[0071] Step 5: Input the noisy multi-channel 3D data samples to be processed into the inverse module and the trained noise prediction network, perform inverse denoising iteration, and split the corresponding multimodal 3D-MRI generated data into 3D-MRI generated images of each modality according to the channel.

[0072] This embodiment uses multi-channel pure noise volume data sampled from a standard normal distribution. The noisy multi-channel 3D data samples are input into the inverse module of the trained unconditional diffusion model for inverse denoising iteration to obtain the corresponding multimodal 3D-MRI generated data. The multimodal 3D-MRI generated data is split by channel to obtain four-modal 3D-MRI generated images: T1n, T1c, T2f, and T2w. The samples in the test set are used to perform quantitative or qualitative evaluation of the generated results.

[0073] Compare the segmentation results obtained in step 5 with the data obtained in step 1, such as... Figure 6 As shown, the first row contains images of the T1c modality, the second row contains images of the T1n modality, the third row contains images of the T2f modality, and the fourth row contains images of the T2w modality. The first column contains images from the original dataset, the second column contains images generated using the VAE method, the third column contains images generated using the CycleGAN method, the fourth column contains images generated using the DDPM method, the fifth column contains images generated using the DDIM method, and the sixth column contains images generated using our method. The quantitative images clearly demonstrate that our method generates higher quality images with better consistency across the generated modalities.

[0074] The proposed method was compared with other methods, using the Fraser Initial Distance (FID) as the evaluation metric, as shown in Table 1. Under the four modes T1c, T1n, T2f, and T2w, the scores using the VAE method were 270.18, 310.23, 246.89, and 283.19, respectively; the scores using the CycleGAN method were 210.28, 181.02, 217.28, and 228.29, respectively; the scores using the DDPM method were 173.02, 163.12, 165.29, and 148.89, respectively; the scores using the DDIM method were 172.39, 153.43, 132.68, and 147.28, respectively; and the scores using the method of this invention were 102.98, 110.28, 108.892, and 115.83, respectively. Compared with other methods, the artificial intelligence-based multimodal brain 3D MR image generation method provided by this invention can generate multimodal images better.

[0075] Table 1. Comparison of the Invention with Other Methods

[0076] Example 2 An AI-based multimodal brain 3D-MRI image generation device is used to implement the AI-based multimodal brain 3D-MRI image generation method in Example 1, including: The model building module is used to implement step 2 in Example 1; The total loss function construction module is used to implement step 3 in Example 1; The training module is used to implement step 4 in Example 1; The application module is used to implement step 5 in embodiment 1.

[0077] Example 3 A computer device includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement steps 2 to 5 of Embodiment 1 above.

[0078] Example 4 A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements steps 2 to 5 of Embodiment 1 above.

[0079] Example 5 A computer program product includes a computer program that, when executed by a processor, implements steps 2 to 5 of embodiment 1 described above.

[0080] The specific embodiments described herein are merely illustrative of the spirit of the invention. Those skilled in the art to which this invention pertains may make various modifications or additions to the described embodiments or use similar methods to replace them, without departing from the spirit of the invention or exceeding the scope defined by the appended claims.

Claims

1. A method for generating multimodal brain 3D-MRI images based on artificial intelligence, characterized in that, Includes the following steps: Step 1: For each of the multiple subjects, acquire the corresponding multimodal 3D-MRI raw images; for each subject: spatially align, unify the size and normalize the intensity of the 3D-MRI raw images corresponding to all modalities, assign one modality to one channel, and stack the 3D-MRI raw images of all modalities according to the channel dimension to form a multi-channel 3D data sample; the training set is composed of multiple multi-channel 3D data samples; Step 2: Construct an unconditional diffusion model, including a denoising diffusion probability model and a noise prediction network. The denoising diffusion probability model includes a forward module and a backward module. Step 3: Construct the time step loss function, which includes the spread noise prediction loss and the cross-modal consistency constraint loss. The cross-modal consistency constraint loss includes the boundary consistency loss and the structural consistency loss. Step 4: Train the noise prediction network in the unconditional diffusion model using the training set and the time step loss function to obtain the trained noise prediction network. Step 5: Input the noisy multi-channel 3D data samples to be processed into the inverse module and the trained noise prediction network, perform inverse denoising iteration, and split the corresponding multimodal 3D-MRI generated data into 3D-MRI generated images of each modality according to the channel. The calculation steps for the cross-modal consistency constraint loss are as follows: Step A: Calculate the back-predicted clean body data based on the following formula: , Predict clean body data The predicted clean volume data for each modality is obtained by splitting the data by channel. Noise scheduling parameters during the diffusion process Noise scheduling parameters during the diffusion process , To preset noise scheduling parameters, ; Serial Number Time step Modal number , The total number of modes; To predict noise; For noisy 3D-MRI data; Step B: Calculate the boundary consistency loss according to the following steps: Predicted clean volume data for each mode Calculate the three-dimensional gradient: , , , in, , , They are exist , , Three-dimensional gradient components in three spatial directions; , , They are , , Gradient operators in three spatial directions; Define intermediate quantities : , 3D boundary structure diagram for: , in, A preset constant introduced to prevent the denominator from being zero; This indicates that the average value of all corresponding pixels is taken. The boundary consistency loss is calculated based on the following formula: , in, Norm, intermediate quantity , Indicates from The number of combinations of choosing 2 modes from 3 modes. The total number of modes, , Indicates the modal number. ; For three-dimensional boundary structure diagrams correspond of Pick ; Three-dimensional boundary structure diagram correspond of Pick ; The structural consistency loss is calculated based on the following steps: Compute structure embedding vector : , For structural encoders; Calculate structural consistency loss based on the following formula. : , , in, Represents the square of the L2 norm. For average structure embedding vectors; Step C: Calculate the cross-modal consistency constraint loss based on the following formula. : , in, , These are the first weighting coefficient and the second weighting coefficient, respectively.

2. The method for generating multimodal brain 3D-MRI images based on artificial intelligence according to claim 1, characterized in that, The specific structure of the unconditional diffusion model is as follows: For each preset time step Multi-channel 3D data samples The input to the forward module is to output the corresponding noisy 3D-MRI data; the noise prediction network performs noise prediction on the noisy 3D-MRI data; the backward module, based on the corresponding predicted noise, infers the predicted 3D-MRI data corresponding to the previous time step of the noisy 3D-MRI data; where... , Maximum time step; When time step The predicted 3D-MRI data from the previous time step is generated from multimodal 3D-MRI data; During training, noisy 3D-MRI data were taken at the same time step. Noisy 3D-MRI data; During the application process after training, when the time step The corresponding noisy 3D-MRI data is the noisy multi-channel 3D data sample to be processed, when the time step Time step The corresponding noisy 3D-MRI data is: based on time steps Time step inferred from noisy 3D-MRI data The corresponding predicted 3D-MRI data.

3. The method for generating multimodal brain 3D-MRI images based on artificial intelligence according to claim 1, characterized in that, The predicted loss for the diffused noise is: , in, To predict the loss for diffused noise, For time step The corresponding external noise is denoted as external noise. The noise added at each time step is random noise sampled from a standard normal distribution, and noise is added independently for each channel; To predict noise; Represents the square of the L2 norm. For noisy 3D-MRI data.

4. The method for generating multimodal brain 3D-MRI images based on artificial intelligence according to claim 3, characterized in that, The time step loss function for: , This is the third weighting coefficient.

5. The method for generating multimodal brain 3D-MRI images based on artificial intelligence according to claim 1, characterized in that, The structural encoder sequentially includes a first three-dimensional convolutional layer, a first normalization layer, and a first activation function layer; a second three-dimensional convolutional layer, a second normalization layer, and a second activation function layer; and a third three-dimensional convolutional layer, a third normalization layer, and a third activation function layer. And a global average pooling layer.

6. The method for generating multimodal brain 3D-MRI images based on artificial intelligence according to claim 2, characterized in that, The noisy 3D-MRI data were calculated based on the following formula: , In the formula, Noise scheduling parameters during diffusion for adding noise to 3D-MRI data Noise scheduling parameters during the diffusion process , To preset noise scheduling parameters, ; Serial Number ; This is a multi-channel 3D data sample; For time step The corresponding external noise is denoted as external noise. External noise at each time step For random noise sampled from a standard normal distribution, noise is added independently to each channel; The predicted 3D-MRI data is calculated based on the following formula: , in, To predict 3D-MRI data, the corresponding time step ; For noisy 3D-MRI data, For random noise sampled according to a standard normal distribution, the noise scheduling parameters during the diffusion process. , To preset noise scheduling parameters, ; The standard deviation of the noise added during the reverse denoising process; Predicting 3D-MRI data Generate data for multimodal 3D-MRI; For the training process, time step Corresponding noisy 3D-MRI data ; For the application process after training, when the time step Noisy 3D-MRI data For the noisy multi-channel 3D data samples to be processed, when the time step Noisy 3D-MRI data .

7. The method for generating multimodal brain 3D-MRI images based on artificial intelligence according to claim 1, characterized in that, The noise prediction network includes a generator encoder, a generator decoder, and skip connections: The generative encoder includes multiple downsampling layers connected in sequence. Each downsampling layer includes a first normalization layer, a first convolutional layer, a second normalization layer, a second convolutional layer, an activation function layer, and a pooling layer connected in sequence. The generator decoder includes multiple upsampling layers connected in sequence. Each upsampling layer includes a first normalization layer, a first convolutional layer, a second normalization layer, a second convolutional layer, an activation function layer, and an upsampling convolutional module connected in sequence. The jump connection is located between the pooling layer output of each downsampling layer and the input of the first normalized layer of the corresponding upsampling layer.

8. The method for generating multimodal brain 3D-MRI images based on artificial intelligence according to claim 1, characterized in that, Step 4 specifically includes the following steps: Several multi-channel 3D data samples were randomly sampled from the training set. This constitutes a training batch for each multi-channel 3D data sample. , corresponding to the set of time steps Each time step For multi-channel 3D data samples Perform forward noise addition to obtain noisy 3D-MRI data. ; Then the noisy 3D-MRI data With time step The noise is input into the noise prediction network to obtain the predicted noise. ;Calculate time steps Corresponding time-step loss function ; Loss function for each time step Backpropagation is performed to calculate the parameter gradients of the noise prediction network in the unconditional diffusion model, and the optimizer is used to update the parameters of the noise prediction network in the unconditional diffusion model.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements steps 2 to 5 of the artificial intelligence-based multimodal brain 3D-MRI image generation method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Multi-modal brain MRI image bidirectional conversion method based on multi-generation and multi-confrontation

    CN110444277A