A method for generating an MRI missing modality based on a learnable frequency domain module
By generating MRI missing modalities through a learnable frequency domain module and a conditional diffusion module with modal attention, the problems of model instability and non-convergence in existing technologies are solved, achieving high-precision and robust missing modal generation and improving the accuracy of tumor diagnosis and treatment.
Patent Information
- Application Number
- CN202511131148.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-08-13
AI Technical Summary
Existing methods for generating missing modalities in MRI images suffer from model instability, non-convergence, and high sensitivity to hyperparameters, making them unsuitable for real-world applications with random modal missingness and affecting the accuracy of tumor segmentation and treatment planning.
We employ a learnable frequency domain module and a conditional diffusion module based on modal attention to adaptively separate high-frequency and low-frequency features by extracting frequency domain information. We also utilize forward noise diffusion and reverse noise reduction generation modules to generate images of missing modalities and combine them with a joint loss function to optimize model parameters.
It improves the accuracy and robustness of missing modality generation, can handle different combinations of missing modalities, and enhances the performance of tumor diagnosis tasks and the quality of treatment.
Smart Images

Figure CN120634918B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of digital image processing, specifically to a method, system, device, and storage medium for generating MRI missing modalities based on a learnable frequency domain module. Background Technology
[0002] In medical imaging, multimodal imaging typically provides more comprehensive information about diseased tissue, thereby improving the accuracy of disease diagnosis and the effectiveness of treatment planning. Particularly in the diagnosis of brain tumors, multiple modalities of MRI images (T1, T1ce, T2, FLAIR) are crucial for determining the localization and nature of the tumor.
[0003] However, modality loss is unavoidable across different clinical centers. For example, modalities are frequently lost due to limitations such as scan time, scan corruption, artifacts, motion, and contrast agent intolerance. These factors limit the universality of multimodal imaging, impacting various downstream tasks such as segmentation, detection, and quantification. The problem of missing modalities significantly reduces the diagnostic performance of multimodal imaging, affecting subsequent tumor segmentation and treatment planning. Therefore, synthesizing missing multimodal MRI, as a means to overcome the limitations of modality inadequacy in clinical practice and research, has become a hot research topic.
[0004] Traditional image segmentation and detection methods typically rely on complete multimodal data, and their accuracy drops significantly when faced with missing modalities. With the development of deep learning, many deep learning-based models have been proposed, but most methods can only achieve one-to-one or many-to-one model training, failing to adapt to real-world applications with randomly missing modalities. While generative adversarial networks (GANs) can achieve many-to-many model training, these models suffer from modality collapse, non-convergence, instability, and high sensitivity to hyperparameters.
[0005] Therefore, researching a unified method for generating missing modalities in MRI images is of great significance for improving the performance of downstream tumor diagnostic tasks such as segmentation, detection, and quantification, as well as the quality of subsequent treatment. Summary of the Invention
[0006] This application provides a method, system, device, and storage medium for generating missing MRI modalities based on learnable frequency domain modules, enabling the reconstruction and synthesis of arbitrary missing modalities in multimodal medical images of gliomas.
[0007] To achieve the above objectives, this application adopts the following technical solution:
[0008] In a first aspect, this application provides a method for generating MRI missing modalities based on learnable frequency domain modules, the method comprising:
[0009] Training and test sets of multimodal MRI images of brain tumors were obtained. The training and test sets consisted of multiple samples, and each sample included four original modal images, each of which was either an available modality or a missing modality.
[0010] A missing modality brain tumor generation network model is constructed, comprising a learnable frequency domain module and a conditional diffusion module based on modal attention. The learnable frequency domain module is used to extract the frequency domain information of the available modalities in the i-th sample and adaptively separate high-frequency and low-frequency features to obtain the high-frequency and low-frequency information of the i-th sample. The conditional diffusion module based on modal attention includes a forward noise-adding diffusion module and a backward noise reduction generation module based on modal attention. The forward noise-adding diffusion module is used to add Gaussian noise to the missing modalities in the i-th sample step by step until T noisy images are generated. The backward noise reduction generation module based on modal attention is used to reduce the noise of the T-th noisy image generated by the forward noise-adding diffusion module step by step under the guidance of the high-frequency or low-frequency information of the i-th sample until T noisy images are generated, where T is a positive integer greater than 1 and i is a positive integer greater than or equal to 1.
[0011] The training set is input into the model for training, and a joint loss function is set. The parameters of the model are optimized by using the joint loss function for each sample to obtain the optimized model.
[0012] The multimodal MRI images from the test set are input into the optimized model to obtain the generated images corresponding to the missing modalities.
[0013] Secondly, this application provides an MRI missing modality generation system based on a learnable frequency domain module, the system comprising:
[0014] The data acquisition module is used to acquire training and test sets of multimodal MRI images of brain tumors. The training and test sets include multiple samples, and each sample includes four original modal images, each of which is either an available modality or a missing modality.
[0015] The model building module is used to construct a missing modality brain tumor generation network model. This model includes a learnable frequency domain module and a conditional diffusion module based on modal attention. The learnable frequency domain module extracts the frequency domain information of the available modalities in the i-th sample and adaptively separates high-frequency and low-frequency features to obtain the high-frequency and low-frequency information of the i-th sample. The conditional diffusion module based on modal attention includes a forward noise-adding diffusion module and a backward noise reduction generation module based on modal attention. The forward noise-adding diffusion module adds Gaussian noise to the missing modalities in the i-th sample step by step until T noisy images are generated. The backward noise reduction generation module based on modal attention, guided by the high-frequency or low-frequency information of the i-th sample, progressively reduces the noise of the T-th noisy image generated by the forward noise-adding diffusion module until T denoised images are generated, where T is a positive integer greater than 1 and i is a positive integer greater than or equal to 1.
[0016] The training module is used to input the training set into the model for training, and to set the joint loss function. The parameters of the model are optimized by using the joint loss function for each sample to obtain the optimized model.
[0017] The generation module is used to input multimodal MRI images from the test set into the optimized model to obtain generated images corresponding to the missing modalities.
[0018] Thirdly, an MRI missing modality generation apparatus based on a learnable frequency domain module is provided, the apparatus including a module for performing the method of the first aspect described above.
[0019] In one possible design, the MRI missing modality generation device based on learnable frequency domain modules of the third aspect may further include a transceiver. This transceiver can be a transceiver circuit or an interface circuit. The transceiver can be used for communication between the MRI missing modality generation device based on learnable frequency domain modules of the third aspect and other devices.
[0020] In one possible design, the MRI missing modality generation apparatus based on learnable frequency domain modules of the third aspect may further include a memory. This memory may be integrated with the processor or disposed separately. The memory may be used to store instructions related to the method of the first aspect.
[0021] Fourthly, an MRI missing modality generation apparatus based on a learnable frequency domain module is provided. The MRI missing modality generation apparatus based on a learnable frequency domain module includes: a processor coupled to a memory, the processor executing instructions stored in the memory to cause the MRI missing modality generation apparatus based on the learnable frequency domain module to perform the method of the first aspect.
[0022] In one possible design, the MRI missing modality generation device based on learnable frequency domain modules of the fourth aspect may further include a transceiver. This transceiver can be a transceiver circuit or an interface circuit. The transceiver can be used for communication between the MRI missing modality generation device based on learnable frequency domain modules of the fourth aspect and other devices.
[0023] Fifthly, an MRI missing modality generation apparatus based on a learnable frequency domain module is provided, comprising: a processor and a memory; the memory is used to store instructions, which, when executed by the processor, cause the MRI missing modality generation apparatus based on the learnable frequency domain module to perform the method of the first aspect.
[0024] In one possible design, the MRI missing modality generation device based on learnable frequency domain modules of the fifth aspect may further include a transceiver. This transceiver can be a transceiver circuit or an interface circuit. The transceiver can be used for communication between the MRI missing modality generation device based on learnable frequency domain modules of the fifth aspect and other devices.
[0025] In a sixth aspect, a computer-readable storage medium is provided, the computer-readable storage medium including a computer program or instructions that, when executed, cause the MRI missing modality generation method based on a learnable frequency domain module of the first aspect to be performed.
[0026] Based on the above technical solution, this application can achieve the following technical effects:
[0027] The missing modality brain tumor modality generation network model in this application learns to capture the relevant features between different modalities by using high-frequency and low-frequency information obtained from a learnable frequency domain module, and a modality attention module introduced in the modality attention-based reverse denoising generation module, thereby promoting feature information complementarity. Then, the model is optimized by using the noisy image obtained by the forward noise diffusion module and the denoised image obtained by the modality attention-based reverse denoising generation module using high-frequency and low-frequency information and the relevant features between different modalities. This improves the accuracy and robustness of the missing modality brain tumor modality generation network model for generating missing modalities, and only a single model is needed to handle different combinations of missing modalities.
[0028] Other features and advantages of this application will be described in detail in the following detailed description section. Attached Figure Description
[0029] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0030] Figure 1 A flowchart illustrating the MRI missing mode generation method based on a learnable frequency domain module provided in this application embodiment;
[0031] Figure 2 This is a schematic diagram of the structure of the missing modality brain tumor generation network model provided in the embodiments of this application;
[0032] Figure 3 This is a schematic diagram of the learnable frequency domain module in the missing modality brain tumor generation network model provided in the embodiments of this application.
[0033] Figure 4 A schematic diagram of the MRI missing modality generation device based on a learnable frequency domain module provided in this application embodiment. Figure 1 ;
[0034] Figure 5 A schematic diagram of the MRI missing modality generation device based on a learnable frequency domain module provided in this application embodiment. Figure 2 . Detailed Implementation
[0035] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. At the same time, in the description of the embodiments of this application, the terms "first," "second," etc., are only used to distinguish descriptions and should not be construed as indicating or implying relative importance. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more features. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0036] Figure 1 This is a flowchart illustrating the MRI missing modality generation method based on a learnable frequency domain module provided in an embodiment of this application.
[0037] The workflow of the MRI missing modality generation method based on learnable frequency domain modules is as follows:
[0038] Step S101: Obtain the training set and test set of multimodal MRI images of brain tumors.
[0039] The training set and the test set consist of multiple samples, and each sample consists of four original modality images, each of which is either an available modality or a missing modality.
[0040] Furthermore, the detailed steps for obtaining the training and test sets of multimodal MRI images of brain tumors in step S101 can be understood by referring to steps S201-S204, as follows:
[0041] Step S201: Obtain a multimodal MRI image dataset of brain tumors.
[0042] The multimodal MRI image dataset includes multiple samples, and each sample includes four original modal images, each of which is either an available modality or a missing modality.
[0043] Furthermore, each of the above samples includes four original modal images: Fluid Attenuation Inversion Recovery Imaging (FLAIR), T1-weighted imaging (T1), T1-weighted contrast enhancement imaging (T1ce), and T2-weighted imaging (T2). These four original modal images may be available modalities (also known as non-missing modalities) or missing modalities; there is no specific limitation. This application uses FLAIR and T2 as missing modalities and T1ce and T1 as available modalities as an example, which will not be emphasized further here. However, in actual applications, it depends on the specific circumstances. It should also be noted that in this application, FLAIR and T2, which are missing modalities in the training set, are simulated missing modalities after processing the available modalities.
[0044] Step S202: Preprocess the multimodal MRI image dataset to obtain a two-dimensional multimodal MRI slice image dataset, and divide the two-dimensional multimodal MRI slice image dataset into a first dataset and a second dataset according to a preset ratio.
[0045] The multimodal MRI image dataset in step S201 is a three-dimensional image dataset, and each sample can be represented as... Where C represents the number of modalities of the multimodal MRI image, and in this application C is 4. W, H and D represent the width, height and number of slices of the multimodal MRI image, respectively. The number of slices can be understood as the number of two-dimensional images obtained after slicing a three-dimensional image.
[0046] In this application, the middle two-dimensional slice is selected from multiple two-dimensional slice images, specifically as follows: ,in, For each sample, a two-dimensional slice is selected from the middle position of the four modal 3D images, with its length and width matching those of the 3D images. The reason for selecting the middle 2D slice is that tumors typically exhibit a three-dimensional structure. To obtain complete tumor slice information, it is usually advisable to sample from the middle of the tumor, as this position can more comprehensively reflect the overall morphology and tissue characteristics of the tumor, avoiding insufficient sample representativeness caused by sampling only the surface or edge parts.
[0047] Following the above processing, a two-dimensional multimodal MRI slice image dataset is obtained. Each sample in the two-dimensional multimodal MRI slice image dataset includes two-dimensional slice images of four modalities. Then, the two-dimensional multimodal MRI slice image dataset is divided into a first dataset and a second dataset according to a preset ratio.
[0048] The default ratio is usually 8:2, but in practical applications, 6:4, 7:3, etc. can also be selected. The specific ratio is determined according to the actual situation and no specific restrictions are made here.
[0049] Step S203: Perform data augmentation on the two-dimensional multimodal MRI slice images in the first dataset to obtain the third dataset.
[0050] Data augmentation is a technique that generates new data samples by transforming, amplifying, or perturbing the original data. It can effectively increase the diversity of training data, improve the model's generalization ability, and prevent overfitting. Common image data augmentation methods include random rotation, random flipping, random cropping, and random noise injection. Random rotation involves randomly rotating the image by a specific angle, typically within a certain range (e.g., -30° to +30°); random flipping involves randomly flipping the image horizontally or vertically; random cropping involves randomly cropping a region of a specified size from the image. The size and position of the cropped region can be randomly chosen, and the cropped image retains its original label. Random noise injection involves randomly adding noise (such as Gaussian noise, salt-and-pepper noise, etc.) to the image, making it blurry or introducing random interference.
[0051] Step S204: Generate status codes for the augmented two-dimensional multimodal MRI slice images in the third dataset, obtain the corresponding status codes, and then stitch the status codes with the corresponding original modal images in the multimodal MRI image dataset to obtain the fourth dataset.
[0052] Status codes are generated for the augmented two-dimensional multimodal MRI slice images in the third dataset. Specifically, the augmented two-dimensional multimodal MRI slice images in the third dataset are processed by a missing status code random generator to generate status codes containing only "0" and "1". "0" corresponds to the missing modality of the two-dimensional multimodal MRI slice image, and "1" corresponds to the available modality of the two-dimensional multimodal MRI slice image.
[0053] That is, the missing state code random generator independently samples binary state variables for each modality, satisfying the following condition:
[0054] ,i=1,...,N and ,
[0055] Among them, symbols It means "to obey"; In this application, "Bernoulli distribution" is used to represent this distribution. show There are only two possible values, 0 or 1, and the probability of taking the value 0 or 1 is equal; Indicates the number of modes; A value of 0 indicates a missing mode, and a value of 1 indicates a available mode.
[0056] For example, in a data-augmented two-dimensional multimodal MRI slice image from the third dataset, FLAIR and T2 represent missing modalities, while T1ce and T1 represent available modalities, resulting in a status code of 0110.
[0057] The status codes are then concatenated with the corresponding original modal images from the multimodal MRI image dataset to obtain the fourth dataset, which can be represented as follows: ,
[0058] in, This represents the fourth dataset (also known as the set of images with random missing modalities). This represents a multimodal MRI image dataset; The status code obtained above; {;} indicates a concatenation operation.
[0059] In addition, the fourth dataset obtained above is the training set for multimodal MRI images of brain tumors, and the second dataset obtained above is the test set for multimodal MRI images of brain tumors.
[0060] Step S102: Construct a missing modality brain tumor generation network model, wherein the missing modality brain tumor generation network model includes a learnable frequency domain module and a conditional diffusion module based on modality attention.
[0061] The learnable frequency domain module is used to extract the frequency domain information of the available modes in the i-th sample and perform adaptive separation of high-frequency and low-frequency features to obtain the high-frequency and low-frequency information of the i-th sample.
[0062] The modal attention-based conditional diffusion module includes a forward noise-adding diffusion module and a modal attention-based inverse noise reduction generation module. The forward noise-adding diffusion module adds Gaussian noise to the missing modality in the i-th sample step by step until T noisy images are generated. The modal attention-based inverse noise reduction generation module, guided by the high-frequency or low-frequency information of the i-th sample, reduces the noise of the T-th noisy image generated by the forward noise-adding diffusion module step by step until T denoised images are generated, where T is a positive integer greater than 1 and i is a positive integer greater than or equal to 1.
[0063] It should be noted that the i-th sample mentioned above refers to any single sample, or each sample, which will not be elaborated further hereafter.
[0064] Missing modality brain tumor generation network model, such as Figure 2 As shown, the learnable frequency domain module is as follows: Figure 3 As shown, please refer to the details. Figure 2 , Figure 3 Please understand the following statements.
[0065] In the embodiment, the learnable frequency domain module is used to extract the frequency domain information of the available modes in the i-th sample and perform adaptive separation of high-frequency features and low-frequency features to obtain the high-frequency information and low-frequency information of the i-th sample. For specific implementation, please refer to steps S201-S206.
[0066] Step S201: Perform a Fourier transform on the available modes in the i-th sample to transform the spatial domain information into the frequency domain, thereby obtaining the frequency domain information of the available modes in the i-th sample; wherein, if there is more than one available mode in the i-th sample, the available mode with higher priority is selected according to priority for processing to obtain the frequency domain information of the i-th sample, and the priority is preset.
[0067] In the example above, it is illustrated that FLAIR and T2 are missing modes, and T1ce and T1 are available modes in this application. It can be seen that when there is more than one available mode in the i-th sample, the available mode with the higher priority is selected for processing to obtain the frequency domain information of the i-th sample. The priority is preset.
[0068] In the preset priorities, T1 is the first priority, T2 is the second priority, FLAIR is the third priority, and T1ce is the fourth priority. It can be seen that by performing a Fourier transform on the available modes of T1, the spatial domain information is transformed into the frequency domain to obtain the frequency domain information of the available modes in the i-th sample.
[0069] Specifically, T1 can be represented by a modal image as follows: (Where H represents the height of the image and W represents the width of the image), perform a Fourier transform (DFT) using the following formula:
[0070] In the formula, It is the frequency domain transformation function, which is the frequency domain information mentioned above; u and v are frequency variables, representing the spatial frequency of the image in the horizontal (x-axis) and vertical (y-axis) directions, respectively; For spatial domain images at location Pixel values; These are Fourier basis functions.
[0071] Step S202: Based on the frequency domain information of the i-th sample, obtain the amplitude spectrum and phase spectrum of the i-th sample.
[0072] That is, based on the frequency domain information of T1 in the i-th sample, the Fast Fourier Transform (FFT) is used for decomposition, thereby accelerating the Fourier Transform computation. The FFT output... It is a complex matrix, after The amplitude and phase calculation formulas are used to obtain the amplitude spectrum. and phase spectrum ,as follows:
[0073]
[0074]
[0075] The above, amplitude spectrum As the primary basis for subsequent frequency band allocation, the phase spectrum For inverse transform image reconstruction, please refer to the following description.
[0076] Step S203: Downsample the amplitude spectrum of the i-th sample and normalize it to a preset range to obtain the processed amplitude spectrum of the i-th sample.
[0077] Amplitude spectrum Downsampling (e.g., from H×W to H / 4×W / 4) is performed to reduce computational load, alleviating the problem of large number of diffusion model parameters and long training time caused by high resolution in medical images; and the amplitude spectrum value is normalized to the range of [0,1] or other preset range to obtain the amplitude spectrum of the i-th sample after processing.
[0078] Step S204: Input the amplitude spectrum of the processed i-th sample into the dynamic mask generation module to obtain the frequency band mask of the i-th sample.
[0079] The processed spectral features are fed into a dynamic mask generation module, which consists of convolutional layers, global pooling layers, and fully connected layers, and outputs a frequency band mask. .
[0080] It should also be noted that traditional frequency domain information separation methods (such as Gaussian filters) use fixed thresholds to divide low-frequency and high-frequency components. This method cannot adapt to the frequency domain characteristics of different modes and is difficult to achieve reliable adaptive division of high and low frequency information under different available modes. Therefore, this application designs a module that can dynamically capture the effective frequency domain information of different modes and adaptively adjust the weights of high and low frequency features to improve the rationality and reliability of frequency domain information. The goal of dynamic mask generation is to learn the frequency band division of the input mode through a neural network and generate an adaptive mask. (where regions close to 0 represent low-frequency components (such as smooth regions), and regions close to 1 represent high-frequency components (such as edges and textures)), in order to achieve adaptive extraction of frequency domain information of different modalities and separation of high and low frequency features.
[0081] Step S205: Use the frequency band mask of the i-th sample for weighting to separate the low-frequency amplitude spectrum and high-frequency amplitude spectrum of the i-th sample.
[0082] Using a mask Masking and weighting are performed to separate low-frequency and high-frequency components, as follows:
[0083] Low-frequency amplitude spectrum: High-frequency amplitude spectrum: .
[0084] Step S206: Combine the low-frequency amplitude spectrum and high-frequency amplitude spectrum of the i-th sample with the phase spectrum of the i-th sample, and use inverse fast Fourier transform to obtain the high-frequency information and low-frequency information of the i-th sample.
[0085] The generated mask is obtained by training the network through supervised learning. It can effectively separate low-frequency and high-frequency components. The separated amplitude spectrum is compared with the original phase spectrum. By combining these methods, the spatial domain image is reconstructed using inverse FFT (IFFT) to obtain the high-frequency image (also known as the high-frequency information of the i-th sample) and the low-frequency image (also known as the low-frequency information of the i-th sample) of the i-th sample, as detailed below:
[0086] Low-frequency images: High-frequency images: .
[0087] In this embodiment, the modal attention-based conditional diffusion module includes a forward noise-adding diffusion module and a modal attention-based reverse noise reduction generation module. The forward noise-adding diffusion module adds Gaussian noise to the missing modalities in the i-th sample step by step until T noisy images are generated. The modal attention-based reverse noise reduction generation module, guided by the high-frequency or low-frequency information of the i-th sample, progressively reduces the noise in the T-th noisy image generated by the forward noise-adding diffusion module until T denoised images are generated. Here, T is a positive integer greater than 1, and i is a positive integer greater than or equal to 1.
[0088] For details on the aforementioned forward noise diffusion module, please refer to [link / reference]. Figure 2 Please understand the following content in detail:
[0089] It should be noted that this application uses FLAIR and T2 as missing modes and T1ce and T1 as available modes as examples. Therefore, both the missing modes FLAIR and T2 need to undergo noise addition and noise reduction processing. The following explanation only focuses on the processing of T2; the processing of FLAIR is the same and can be used for reference.
[0090] It should also be noted that the forward diffusion module of the diffusion model can be regarded as a Markov process. In this forward propagation process, noise is added to the given input image (an image simulating the missing modality, such as T2) step by step, so that it is transformed from a clear image into a completely random patch of noisy images.
[0091] Let the number of noise addition steps be t. For each step, t noisy image patches are generated, where The original image without noise. This is the image with completely Gaussian noise obtained after t steps of noise addition.
[0092] Then at time t, through the latent variables After adding Gaussian noise, we get The process can be represented as:
[0093] ,
[0094] in, express exist The distribution under the given conditions represents the distribution from Noise addition The process follows a Gaussian distribution , Indicates a Gaussian distribution; The predefined variance is used to control the noise weights in each diffusion time step.
[0095] Based on the above conditions, the distribution It can be deduced that The expression is:
[0096] ,
[0097] in, This represents the Gaussian noise added at time t; This represents the variance corresponding to Gaussian distributed noise.
[0098] In addition, there is another way to obtain The expression is as follows:
[0099] Given the original image Then at time t, through the initial variable Conditional distribution obtained after adding Gaussian noise for:
[0100] ,
[0101] in, ; (The sum from time t=1 to time t=i).
[0102] Then the initial variables at time 0 The output is obtained after random t-step noise addition. The expression is:
[0103] .
[0104] The 'I' that appears multiple times above refers to the identity matrix, and will be explained here for clarity.
[0105] In this embodiment, the number of noise addition steps is T, and the T noise addition maps of the missing modes FLAIR and T2 are generated by the forward noise addition diffusion module.
[0106] For details on the modal attention-based inverse noise reduction generation module, please refer to [link / reference]. Figure 2 Please understand the following content in detail:
[0107] The modal attention-based inverse noise reduction generation module consists of T sub-generation modules, each of which includes a modal attention module and a U-Net encoding-decoding module;
[0108] The modal attention module is used to, at time t, obtain the t-th embedded feature map of the four modal images in the i-th sample based on the four modal images in the i-th sample, then determine the t-th attention weight between any two modalities in the i-th sample based on the t-th embedded feature map of the four modal images in the i-th sample and the status code of the i-th sample, and then obtain the t-th new feature map of the four modal images in the i-th sample based on the t-th attention weight between any two modalities in the i-th sample.
[0109] The U-Net encoding-decoding module is used to reconstruct and generate available modalities and generate noise reduction of missing modalities at time t, based on the t-th new feature map of the four modal images in the i-th sample and guided by the high-frequency or low-frequency information of the i-th sample, to obtain the reconstructed image of available modalities and the noise reduction image of missing modalities at time t.
[0110] Wherein, when t is T, the four modal images are the original modal images of the available modalities and the Tth noise image of the missing modalities; when t is not T, the four modal images are the reconstructed images of the available modalities and the denoised images of the missing modalities at the previous moment; and t is a positive integer greater than or equal to 1 and less than or equal to T.
[0111] It is understandable that the effective feature interaction of the aforementioned modality attention module is key to modeling complementary features between available modalities. Although the inverse denoising generation module utilizes the correlation between available and missing modalities to model the missing modality, the presence of a large amount of redundant feature information in each modality information is due to the fact that each modality is a different morphological annotation of the same sample. To further extract effective relevant features and suppress redundant noise, the inverse denoising generation module adds a modality attention module before each U-Net encoding-decoding module from time step t to t-1. Leveraging the adaptive characteristics of graph structures in feature learning, this module further extracts relevant information between different modalities to obtain more useful feature maps and suppress redundant noise.
[0112] Specifically, the modal attention module is a graph structure. Here, V represents the graph nodes corresponding to the four modal features, and E represents the adjacency edge matrix representing the relationships between nodes. In graph edge computation, we will start from the node modal features... arrive The message being transmitted is defined as a node characteristic. and edge weights The product of.
[0113] This edge weight Features used to learn complementary nodes are represented as follows:
[0114]
[0115] in, Indicates a mode missing status code Compared with the original feature map The result after splicing, i.e. . The state vectors are the four original modal images, representing the missing modalities and available modalities of modal information. Similarly, we can obtain the following. Indicates a parameter as The nonlinear function G, consisting of two linear mapping layers and a LeakyReLU activation function layer, is used to estimate the attention weights between each pair of modes in G. Based on the edge weights obtained from G, the softmax function is used to merge information from all available modes and update the nodes. Feature Mapping Updated features It can be represented as:
[0116] ,
[0117] Where N=4.
[0118] when At time t, based on the t-th new feature map of the four modal images in the i-th sample, and guided by the low-frequency information of the i-th sample, the available modal reconstruction and the noise reduction of the missing modal are performed to obtain the reconstructed image of the available modal and the noise reduction image of the missing modal at time t.
[0119] when At time t, based on the t-th new feature map of the four modal images in the i-th sample, and guided by the high-frequency information of the i-th sample, the available modalities are reconstructed and the missing modalities are denoised, resulting in the reconstructed image of the available modalities and the denoised image of the missing modalities at time t. Where the true value of T is odd, when calculating... When the value of T is 1, the value of T is incremented by 1.
[0120] Specifically, for each of the T noise maps generated by the aforementioned forward noise diffusion module for the missing modes FLAIR and T2, noise reduction is performed on each of them. Taking the noise reduction of one of them (T2) as an example.
[0121] The core of the specific generation process is through a modal attention generation network. Predicted noise And combined with the diffusion coefficient of the forward process , Calculate the mean of the conditional distribution of the generation process. ,variance .
[0122] Let the condition at time t be... Then, from the image at time t Perform noise reduction to generate the image from the previous time step. The conditional probability distribution can be expressed as:
[0123] ,
[0124] in, The noise mean predicted by the modal attention generation network of the inverse noise reduction module is expressed as:
[0125]
[0126] in, for The variance of the forward diffusion process , and The calculation yielded:
[0127] ,
[0128] Afterwards, according to exist Probability distribution formula under certain conditions The generation process can be represented as:
[0129] ,
[0130] in, The generated noise components are obtained by the modal attention generation module at each time step.
[0131] In addition, such as Figure 1 As shown, in the modal attention-based inverse denoising generation module, the value of t in the subscript ranges from T to 0, while in the forward noise diffusion module, the value of t ranges from 0 to T. In this application, when simultaneously describing time t in both the modal attention-based inverse denoising generation module and the forward noise diffusion module, it is as follows: Figure 1 The corresponding moments above and below.
[0132] Step S103: Input the training set into the model for training, and set the joint loss function. Optimize the model parameters through the joint loss function of each sample to obtain the optimized model.
[0133] The joint loss function consists of the learning loss of the learnable frequency domain module and the inverse noise reduction generation loss based on modal attention.
[0134] The learning loss of the learnable frequency domain module is:
[0135] ,
[0136] in, This represents the learning loss of the learnable frequency domain module for the i-th sample. and Let these represent the low-frequency amplitude spectrum and the high-frequency amplitude spectrum of the i-th sample, respectively. and These are the preset low-frequency amplitude spectrum and the preset high-frequency amplitude spectrum of the i-th sample, which are divided according to preset fixed thresholds.
[0137] The inverse noise reduction generation loss based on modal attention is:
[0138] ,
[0139] Among them, the The inverse noise reduction generation loss based on modal attention for the i-th sample is represented by the following: The forward noise diffusion module at time t obtains the noise image of the missing mode. The image at time t is the denoised image of the missing modality obtained by the modality attention-based inverse denoising generation module. (It should be noted that...) and At time t, such as Figure 1 (corresponding times above and below); the aforementioned The high-frequency or low-frequency information of the i-th sample at time t; Represents joint random variables , The expectation; the Represents Euclidean distance; the stated This represents the predicted noise of the modal attention-based inverse noise reduction generation module. Noise from the forward noise diffusion module The mean square error; m represents the multimodal mode, which has no specific value.
[0140] The joint loss function is:
[0141] ,
[0142] in, Let λ represent the joint loss function for the i-th sample, and let λ represent the preset hyperparameters.
[0143] It should be noted that the joint loss function obtained for each sample above can be used to optimize the model individually, or the joint loss functions of all samples in the training set can be added together to optimize the model. The specific method depends on the actual situation and is not restricted here.
[0144] Step S104: Input the multimodal MRI images from the test set into the optimized model to obtain the generated images corresponding to the missing modalities.
[0145] In summary, the missing modality brain tumor modality generation network model in this application learns to capture relevant features between different modalities by using high-frequency and low-frequency information obtained from the learnable frequency domain module, and a modality attention module introduced in the modality attention-based reverse denoising generation module, thereby promoting feature information complementarity. Then, by using the noisy image obtained from the forward noise diffusion module and the denoised image obtained from the modality attention-based reverse denoising generation module, the model is optimized using a set loss function, thereby improving the accuracy and robustness of the missing modality brain tumor modality generation network model for generating missing modalities, and only a single model is needed to handle different modality combinations.
[0146] The above combination Figures 1-3 This application provides a detailed description of the MRI missing modality generation method based on a learnable frequency domain module, as described in the embodiments of this application. The following details the MRI missing modality generation system based on a learnable frequency domain module provided in the embodiments of this application.
[0147] The system specifically includes: a data acquisition module, a model building module, a training module, and a generation module, as shown below.
[0148] The data acquisition module is used to acquire training and test sets of multimodal MRI images of brain tumors. The training and test sets include multiple samples, and each sample includes four original modal images, each of which is either an available modality or a missing modality.
[0149] The model building module is used to construct a missing modality brain tumor generation network model. This model includes a learnable frequency domain module and a conditional diffusion module based on modal attention. The learnable frequency domain module extracts the frequency domain information of the available modalities in the i-th sample and adaptively separates high-frequency and low-frequency features to obtain the high-frequency and low-frequency information of the i-th sample. The conditional diffusion module based on modal attention includes a forward noise-adding diffusion module and a backward noise reduction generation module based on modal attention. The forward noise-adding diffusion module adds Gaussian noise to the missing modalities in the i-th sample step by step until T noisy images are generated. The backward noise reduction generation module based on modal attention, guided by the high-frequency and low-frequency information of the i-th sample, progressively reduces the noise of the T-th noisy image generated by the forward noise-adding diffusion module until T denoised images are generated, where T is a positive integer greater than 1 and i is a positive integer greater than or equal to 1.
[0150] The training module is used to input the training set into the model for training, and to set the joint loss function. The parameters of the model are optimized by using the joint loss function for each sample to obtain the optimized model.
[0151] The generation module is used to input multimodal MRI images from the test set into the optimized model to obtain generated images corresponding to the missing modalities.
[0152] Furthermore, the specific implementation of the above system is basically similar to the method implementation, so the description is relatively simple. For relevant details, please refer to the description of the method implementation. Moreover, it should be noted that in the various modules of the system of this application, the components are logically divided according to the functions they are to perform. However, this application is not limited to this and can re-divide or combine the components as needed.
[0153] The above describes the method and system for generating MRI missing modalities based on learnable frequency domain modules provided in the embodiments of this application. The following, in conjunction with... Figures 4-5 This document provides a detailed description of the MRI missing modality generation apparatus based on a learnable frequency domain module provided in the embodiments of this application.
[0154] Figure 4 This is a schematic diagram of the MRI missing modality generation device based on a learnable frequency domain module provided in the embodiments of this application. Figure 1 For example, such as Figure 4 As shown, the MRI missing modality generation device 400 based on a learnable frequency domain module includes: a transceiver module 401 and a processing module 402. For ease of explanation, Figure 4 Only the main components of this MRI missing modality generation device based on a learnable frequency domain module are shown.
[0155] The transceiver module 401 is used to perform the transceiver function of the above-mentioned MRI missing modality generation method based on the learnable frequency domain module, and the processing module 402 is used to perform other functions of the above-mentioned MRI missing modality generation method based on the learnable frequency domain module besides the transceiver function.
[0156] Optionally, the transceiver module 401 may include a transmitting module ( Figure 4 (not shown in the image) and receiving module ( Figure 4 (Not shown in the image). The transmitting module is used to implement the transmitting function of the MRI missing modality generation device 400 based on the learnable frequency domain module, and the receiving module is used to implement the receiving function of the MRI missing modality generation device 400 based on the learnable frequency domain module.
[0157] Optionally, the MRI missing modality generation device 400 based on a learnable frequency domain module may further include a storage module. Figure 4(Not shown in the image), the storage module stores programs or instructions. When the processing module 402 executes the program or instructions, the MRI missing modality generation device 400 based on the learnable frequency domain module can execute the MRI missing modality generation method based on the learnable frequency domain module in the embodiments of this application.
[0158] The following is combined Figure 5 Each component of the MRI missing modality generation device 500 based on learnable frequency domain modules is described in detail:
[0159] The processor 501 is the control center of the MRI missing modality generation device 500 based on a learnable frequency domain module. It can be a single processor or a collective term for multiple processing elements. For example, the processor 501 can be one or more central processing units (CPUs), application-specific integrated circuits (ASICs), or one or more integrated circuits configured to implement the embodiments of this application, such as one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs).
[0160] Optionally, the processor 501 can execute various functions of the MRI missing modality generation device 500 based on the learnable frequency domain module by running or executing software programs stored in the memory 502 and calling data stored in the memory 502, such as executing the MRI missing modality generation method based on the learnable frequency domain module in the embodiments of this application.
[0161] In a specific implementation, as one example, the processor 501 may include one or more CPUs, for example... Figure 5 CPU0 and CPU1 are shown in the diagram.
[0162] In a specific implementation, as one example, the MRI missing modality generation device 500 based on a learnable frequency domain module may also include multiple processors, for example... Figure 5The processors 501 and 504 are shown in the diagram. Each of these processors can be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). Here, "processor" can refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions). The memory 502 is used to store the software program executing the scheme of this application, and its execution is controlled by the processor 501. Specific implementation methods can be found in the above method embodiments, and will not be repeated here.
[0163] Optionally, the memory 502 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed discs, laser discs, optical discs, digital universal discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. The memory 502 may be integrated with the processor 501 or exist independently, and may be connected via the interface circuit of the MRI missing modality generation device 500 based on a learnable frequency domain module. Figure 5 (Not shown in the image) is coupled to processor 501, but this embodiment does not specifically limit this.
[0164] Transceiver 503 is used for communication with other communication devices. For example, in the case of an MRI missing modality generation device 500 based on a learnable frequency domain module as the first device, transceiver 503 can be used to communicate with a second device or a third device.
[0165] Alternatively, transceiver 503 may include a receiver and a transmitter. Figure 5 (Not shown separately). The receiver is used to implement the receiving function, and the transmitter is used to implement the transmitting function.
[0166] Optionally, the transceiver 503 can be integrated with the processor 501 or exist independently, and can be connected via the interface circuit of the MRI missing modality generation device 500 based on the learnable frequency domain module. Figure 5(Not shown in the image) is coupled to processor 501, and this embodiment does not specifically limit this.
[0167] Understandable Figure 5 The structure of the MRI missing modality generation device 500 based on a learnable frequency domain module shown in the figure does not constitute a limitation on the MRI missing modality generation device based on a learnable frequency domain module. The actual MRI missing modality generation device based on a learnable frequency domain module may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0168] Furthermore, the technical effects of the MRI missing modality generation device 500 based on the learnable frequency domain module can be referred to the technical effects of the method described in the above method embodiments, and will not be repeated here.
[0169] It should be understood that the processor in the embodiments of this application can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.
[0170] It should also be understood that the memory in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0171] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.
Claims
1. A method for generating MRI missing modalities based on learnable frequency domain modules, characterized in that, The method includes: A training set and a test set for acquiring multimodal MRI images of brain tumors are obtained, wherein the training set and the test set include multiple samples, and each sample includes four original modal images, each original modal image being either an available modality or a missing modality; A missing modality brain tumor generation network model is constructed, comprising a learnable frequency domain module and a conditional diffusion module based on modal attention. The learnable frequency domain module extracts the frequency domain information of available modalities in the i-th sample and adaptively separates high-frequency and low-frequency features to obtain the high-frequency and low-frequency information of the i-th sample. The conditional diffusion module based on modal attention comprises a forward noise-adding diffusion module and a backward noise-reducing generation module based on modal attention. The forward noise-adding diffusion module adds Gaussian noise to the missing modalities in the i-th sample step by step until T noisy images are generated. The backward noise-reducing generation module based on modal attention, guided by the high-frequency or low-frequency information of the i-th sample, progressively reduces the noise of the T-th noisy image generated by the forward noise-adding diffusion module until T denoised images are generated, where T is a positive integer greater than 1 and i is a positive integer greater than or equal to 1. The training set is input into the model for training, and a joint loss function is set. The parameters of the model are optimized by the joint loss function for each sample to obtain the optimized model. The multimodal MRI images in the test set are input into the optimized model to obtain the generated images corresponding to the missing modalities.
2. The method for generating MRI missing modes based on a learnable frequency domain module according to claim 1, characterized in that, The training and test sets for acquiring multimodal MRI images of brain tumors include: A multimodal MRI image dataset of brain tumors is obtained, wherein the multimodal MRI image dataset includes multiple samples, and each sample includes four original modal images, each original modal image being either an available modality or a missing modality; The multimodal MRI image dataset is preprocessed to obtain a two-dimensional multimodal MRI slice image dataset, and the two-dimensional multimodal MRI slice image dataset is divided into a first dataset and a second dataset according to a preset ratio. Data augmentation was performed on the two-dimensional multimodal MRI slice images in the first dataset to obtain the third dataset; Status codes are generated for the augmented two-dimensional multimodal MRI slice images in the third dataset to obtain the corresponding status codes. The status codes are then stitched together with the corresponding original modal images in the multimodal MRI image dataset to obtain the fourth dataset. The fourth dataset is a training set of multimodal MRI images of brain tumors, and the second dataset is a test set of multimodal MRI images of brain tumors.
3. The method for generating MRI missing modes based on learnable frequency domain modules according to claim 2, characterized in that, The data augmentation process includes: random image rotation, random flipping, random cropping, and random addition of noise.
4. The method for generating MRI missing modes based on a learnable frequency domain module according to claim 1, characterized in that, The learnable frequency domain module is used to extract the frequency domain information of the available modes in the i-th sample, and to perform adaptive separation of high-frequency and low-frequency features to obtain the high-frequency and low-frequency information of the i-th sample, including: Perform a Fourier transform on the available modes in the i-th sample to transform the spatial domain information to the frequency domain, thereby obtaining the frequency domain information of the available modes in the i-th sample. If there is more than one available mode in the i-th sample, select the available mode with higher priority according to priority for processing to obtain the frequency domain information of the i-th sample. The priority is preset. Based on the frequency domain information of the i-th sample, the amplitude spectrum and phase spectrum of the i-th sample are obtained; The amplitude spectrum of the i-th sample is downsampled and normalized to a preset range to obtain the processed amplitude spectrum of the i-th sample. The amplitude spectrum of the processed i-th sample is input into the dynamic mask generation module to obtain the frequency band mask of the i-th sample. By using the frequency band mask of the i-th sample for weighting, the low-frequency amplitude spectrum and high-frequency amplitude spectrum of the i-th sample are separated. The low-frequency amplitude spectrum and high-frequency amplitude spectrum of the i-th sample are combined with the phase spectrum of the i-th sample, and the high-frequency information and low-frequency information of the i-th sample are obtained by using inverse fast Fourier transform.
5. The method for generating MRI missing modes based on learnable frequency domain modules according to claim 2, characterized in that, The modal attention-based inverse denoising generation module is used to progressively denoise the T-th noise image generated by the forward denoising diffusion module under the guidance of the high-frequency or low-frequency information of the i-th sample, until T denoised images are generated, including: The modal attention-based inverse noise reduction generation module consists of T sub-generation modules, each of which includes a modal attention module and a U-Net encoding-decoding module; The modal attention module is used to, at time t, obtain the t-th embedded feature map of the four modal images in the i-th sample based on the four modal images in the i-th sample, then determine the t-th attention weight between any two modalities in the i-th sample based on the t-th embedded feature map of the four modal images in the i-th sample and the status code of the i-th sample, and then obtain the t-th new feature map of the four modal images in the i-th sample based on the t-th attention weight between any two modalities in the i-th sample. The U-Net encoding-decoding module is used to reconstruct and generate available modalities and generate noise reduction of missing modalities at time t, based on the t-th new feature map of the four modal images in the i-th sample and guided by the high-frequency or low-frequency information of the i-th sample, to obtain the reconstructed image of available modalities and the noise reduction image of missing modalities at time t. Wherein, when t is T, the four modal images are the original modal images of the available modalities and the Tth noise image of the missing modalities; when t is not T, the four modal images are the reconstructed images of the available modalities and the denoised images of the missing modalities at the previous moment; and t is a positive integer greater than or equal to 1 and less than or equal to T.
6. The method for generating MRI missing modalities based on a learnable frequency domain module according to claim 5, characterized in that, At time t, based on the t-th new feature map of the four modal images in the i-th sample, and guided by the high-frequency or low-frequency information of the i-th sample, the available modalities are reconstructed and the missing modalities are denoised, resulting in the reconstructed image of the available modalities and the denoised image of the missing modalities at time t, including: when At time t, based on the t-th new feature map of the four modal images in the i-th sample, and guided by the low-frequency information of the i-th sample, the available modal reconstruction and the noise reduction of the missing modal are performed to obtain the reconstructed image of the available modal and the noise reduction image of the missing modal at time t. when At time t, based on the t-th new feature map of the four modal images in the i-th sample, and guided by the high-frequency information of the i-th sample, the available modalities are reconstructed and the missing modalities are denoised, resulting in the reconstructed image of the available modalities and the denoised image of the missing modalities at time t. Where the true value of T is odd, when calculating... When the value of T is 1, the value of T is incremented by 1.
7. The method for generating MRI missing modalities based on a learnable frequency domain module according to claim 4, characterized in that, The joint loss function consists of the learning loss of the learnable frequency domain module and the inverse noise reduction generation loss based on modal attention; The learning loss of the learnable frequency domain module is: , Among them, the The learning loss of the learnable frequency domain module for the i-th sample is represented by the following: and stated Let the low-frequency amplitude spectrum and high-frequency amplitude spectrum of the i-th sample be represented respectively. and stated These are the preset low-frequency amplitude spectrum and preset high-frequency amplitude spectrum of the i-th sample, which are divided according to a preset fixed threshold. The inverse noise reduction generation loss based on modal attention is: , Among them, the The inverse noise reduction generation loss based on modal attention for the i-th sample is represented by the following: The forward noise diffusion module at time t obtains the noise image of the missing mode. The denoised image of the missing modality obtained by the modality attention-based inverse denoising generation module at time t; The high-frequency or low-frequency information of the i-th sample at time t; Represents joint random variables , The expectation; the stated Represents Euclidean distance; the stated This represents the predicted noise of the modal attention-based inverse noise reduction generation module. Noise from the forward noise diffusion module The mean square error; m represents the multimodal mode, which has no specific value; The joint loss function is: , Among them, the Let λ represent the joint loss function for the i-th sample, where λ represents a preset hyperparameter.
8. A system for generating missing MRI modalities based on a learnable frequency domain module, characterized in that, The system includes: The data acquisition module is used to acquire training and test sets of multimodal MRI images of brain tumors. The training and test sets include multiple samples, and each sample includes four original modal images, each of which is an available modality or a missing modality. A model building module is used to construct a missing modality brain tumor generation network model. This model includes a learnable frequency domain module and a conditional diffusion module based on modal attention. The learnable frequency domain module extracts the frequency domain information of the available modalities in the i-th sample and adaptively separates high-frequency and low-frequency features to obtain the high-frequency and low-frequency information of the i-th sample. The conditional diffusion module based on modal attention includes a forward noise-adding diffusion module and a backward noise-reducing generation module based on modal attention. The forward noise-adding diffusion module adds Gaussian noise to the missing modalities in the i-th sample step by step until T noisy images are generated. The backward noise-reducing generation module, guided by the high-frequency or low-frequency information of the i-th sample, progressively reduces the noise of the T-th noisy image generated by the forward noise-adding diffusion module until T denoised images are generated, where T is a positive integer greater than 1 and i is a positive integer greater than or equal to 1. The training module is used to input the training set into the model for training, and to set the joint loss function to optimize the parameters of the model through the joint loss function of each sample, so as to obtain the optimized model. The generation module is used to input the multimodal MRI images from the test set into the optimized model to obtain the generated images corresponding to the missing modalities.
9. An MRI missing modality generation device based on a learnable frequency domain module, characterized in that, The apparatus includes a module for performing the method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a computer program or instructions that, when executed, cause the method as described in any one of claims 1-7 to be performed.
Citation Information
Patent Citations
Noisy CS-MRI reconstruction method for pyramid decomposition and dictionary learning
CN103632341A
MRI (Magnetic Resonance Imaging) tumor segmentation method in missing mode based on feature-mode double-level fusion
CN118038054A