Diffusion model training method, functional magnetic resonance imaging data generation method and device

Through the diffusion model training method, multi-view coding and noise splicing technology are used, and the diffusion model is trained in combination with text prompt information, which solves the problem of low accuracy in the generation of fMR data, and achieves high-quality and reliable data generation.

CN120356030APending Publication Date: 2025-07-22ARTIFICIAL INTELLIGENCE RES INST OF HEFEI COMPREHENSIVE NAT SCI CENT (ANHUI ARTIFICIAL INTELLIGENCE LAB)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510391019.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The prior art generates fMR imaging data with low accuracy and high limitations, and fails to effectively utilize the high-dimensional and complex spatial and spatial-temporal patterns of the original fMR imaging data, resulting in inaccurate generation results.

Method used

The diffusion model training method is adopted to process the sample initial fMRI data through multi-view coding, and the sample noise is processed using the noise splicing module and the image segmentation convolution module. The diffusion model is trained in combination with the sample text prompt information to generate high-quality fMRI data.

Benefits of technology

It realizes the generation of fMRI data of any length, high quality and reliable, which can accurately capture the spatial information and timing information of the data, and improves the accuracy and reliability of the generated results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356030A_ABST
    Figure CN120356030A_ABST
Patent Text Reader

Abstract

The invention provides a diffusion model training method and a functional magnetic resonance imaging data generation method and device, and can be applied to the technical field of neural image calculation. The method comprises the following steps: acquiring a training sample, wherein the training sample comprises sample initial functional magnetic resonance imaging data and sample text prompt information of a sample target object; performing multi-view coding processing on the sample initial functional magnetic resonance imaging data to obtain sample time sequence characteristic data; processing the sample noise and the sample time sequence characteristic data by using a noise splicing module to obtain sample noise splicing data; based on sample text prompt information, processing the sample noise spliced data by using an image segmentation convolution module to obtain sample prediction noise, the sample text prompt information being used for guiding a noise prediction process; and training a diffusion model according to the sample noise and the sample prediction noise to obtain a trained diffusion model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of neuroimaging computing technology, and particularly to a diffusion model training method, a functional magnetic resonance imaging data generation method, and an apparatus therefor. Background Art

[0002] The diffusion model is a generative model that generates high-quality samples by gradually adding noise to the data distribution and reverse denoising. In various generation tasks such as images, three-dimensional scenes, and videos in the medical field, the diffusion model has shown a powerful ability to synthesize high-quality samples. However, due to the characteristics of functional magnetic resonance imaging, such as high dimensionality, complex spatial and temporal patterns, and few sample data, in existing research, the method for synthesizing functional magnetic resonance imaging data does not focus on the original functional magnetic resonance imaging data level, resulting in low accuracy and high limitations of the generated or predicted functional magnetic resonance imaging data. Summary of the Invention

[0003] In view of the above technical problems, the present disclosure provides a diffusion model training method, a functional magnetic resonance imaging data generation method, and an apparatus therefor.

[0004] According to a first aspect of the present disclosure, there is provided a diffusion model training method. The diffusion model includes a noise splicing module and an image segmentation convolutional module. The diffusion model training method includes: obtaining training samples, where the training samples include sample initial functional magnetic resonance imaging data of a sample target object and sample text prompt information; performing multi-view encoding processing on the sample initial functional magnetic resonance imaging data to obtain sample temporal feature data; using the noise splicing module to process sample noise and the sample temporal feature data to obtain sample noise splicing data; based on the sample text prompt information, using the image segmentation convolutional module to process the sample noise splicing data to obtain sample predicted noise, where the sample text prompt information is used to guide the predicted noise process; and training the diffusion model according to the sample noise and the sample predicted noise to obtain a trained diffusion model.

[0005] According to an embodiment of the present disclosure, there are F pieces of sample time-series feature data and F pieces of sample noise splicing data; wherein, using a noise splicing module to process the sample noise and the sample time-series feature data, the obtained sample noise splicing data includes: based on the sample noise, respectively performing positive noise addition on the F pieces of sample time-series feature data to obtain sample noise-added data corresponding to each of the F pieces of sample time-series feature data; performing splicing processing on the sample initial embedding data and the first piece of sample noise-added data to obtain the first piece of sample noise splicing data corresponding to the first piece of sample noise-added data; performing splicing processing on the (f - 1)-th piece of sample time-series feature data and the f-th piece of sample noise-added data to obtain the f-th piece of sample noise splicing data corresponding to the f-th piece of sample noise-added data, where the f-th piece of sample noise-added data represents the sample noise-added data corresponding to the f-th piece of sample time-series feature data, and the (f - 1)-th piece of sample time-series feature data represents the sample time-series feature data corresponding to the (f - 1)-th acquisition moment, 1 f F.

[0006] According to an embodiment of the present disclosure, the sample initial functional magnetic resonance imaging data includes sample signal data corresponding to each of the F acquisition moments; wherein, performing multi-view encoding processing on the sample initial functional magnetic resonance imaging data to obtain the sample time-series feature data includes: for each piece of sample signal data, using an encoder to perform feature extraction from the coronal plane view to obtain sample coronal plane feature data; using an encoder to perform feature extraction from the sagittal plane view to obtain sample sagittal plane feature data; using an encoder to perform feature extraction from the transverse plane view to obtain sample transverse plane feature data; merging the sample coronal plane feature data, the sample sagittal plane feature data, and the sample transverse plane feature data to obtain the sample time-series feature data corresponding to each piece of sample signal data.

[0007] According to an embodiment of the present disclosure, the diffusion model further includes a denoising module, and the diffusion model training method further includes: using the denoising module to process the sample noise splicing data and the sample predicted noise to obtain sample denoised data.

[0008] According to an embodiment of the present disclosure, training the diffusion model according to the sample noise and the sample predicted noise to obtain the trained diffusion model includes: using a loss function to calculate the loss value between the sample noise and the sample predicted noise to obtain a target loss value; training the diffusion model according to the target loss value to obtain the trained diffusion model.

[0009] The second aspect of the present disclosure provides a method for generating functional magnetic resonance imaging data, including: obtaining initial embedding data and text prompt information related to an image generation task, where the text prompt information is used to guide a prediction noise process; based on a preset time series length S, using a diffusion model to process the initial embedding data, the text prompt information, and noise, and outputting S denoised data, where the diffusion model is trained according to the above method; using a decoder to process the S denoised data to generate target functional magnetic resonance imaging data.

[0010] According to an embodiment of the present disclosure, based on a preset time series length S, using a diffusion model to process the initial embedding data, the text prompt information, and noise, and outputting S denoised data includes: using a noise splicing module to process the noise and the initial embedding data to obtain a first noise-spliced data; based on the text prompt information, using an image segmentation convolution module to process the first noise-spliced data to obtain a first predicted noise; using a denoising module to process the first noise-spliced data and the first predicted noise to obtain a first denoised data; based on a preset time series length s, using the noise splicing module to process the noise and the (s - 1)-th denoised data to obtain an s-th noise-spliced data, where the (s - 1)-th denoised data is obtained by denoising the (s - 1)-th noise-spliced data, 1 ; based on the text prompt information, using an image segmentation convolution module to process the s-th noise-spliced data to obtain an s-th predicted noise; using a denoising module to process the s-th noise-spliced data and the s-th predicted noise to obtain an s-th denoised data.

[0011] According to an embodiment of the present disclosure, the method for generating functional magnetic resonance imaging data further includes: using a classifier to process the denoised data to obtain a classification result, where the classification result characterizes that the target object is a patient with a mental disorder or a non-mental disorder patient.

[0012] The third aspect of the present disclosure provides a diffusion model training device, including: a first acquisition module, configured to acquire training samples, where the training samples include sample initial functional magnetic resonance imaging data of a sample target object and sample text prompt information; an encoding module, configured to perform multi-view encoding processing on the sample initial functional magnetic resonance imaging data to obtain sample time series feature data; a noise splicing module, configured to use the noise splicing module to process the sample noise and the sample time series feature data to obtain sample noise-spliced data; a first prediction module, configured to, based on the sample text prompt information, use an image segmentation convolution module to process the sample noise-spliced data to obtain sample predicted noise, where the sample text prompt information is used to guide the prediction noise process; a training module, configured to train a diffusion model according to the sample noise and the sample predicted noise to obtain a trained diffusion model.

[0013] The fourth aspect of the present disclosure provides a functional magnetic resonance imaging data generation device, including: a second acquisition module, configured to acquire initial embedding data and text prompt information related to an image generation task, where the text prompt information is used to guide a prediction noise process; a second prediction module, configured to process the initial embedding data, the text prompt information, and noise based on a preset time series length S by using a diffusion model, and output S denoised data, where the diffusion model is obtained by training according to the above-mentioned diffusion model; and a generation module, configured to process the S denoised data by using a decoder to generate target functional magnetic resonance imaging data.

[0014] According to the diffusion model training method, the functional magnetic resonance imaging data generation method and device provided by the present disclosure, by performing multi-perspective encoding processing on sample initial functional magnetic resonance imaging data, sample time series feature data is obtained; a sample noise splicing module is used to process sample noise and sample time series feature data to obtain sample noise splicing data; based on sample text prompt information, an image segmentation convolution module is used to process the sample noise splicing data to obtain sample prediction noise; and according to the sample noise and the sample prediction noise, the diffusion model is trained to obtain a trained diffusion model. Since the sample initial functional magnetic resonance imaging data is mapped to a low-dimensional latent space, the spatial information and time series information of the data are effectively captured from different spatial perspectives; using the sample text prompt information as a pathological condition to control the diffusion model to perform forward noise addition diffusion and reverse denoising on the sample time series feature data, learning the conditional distribution between continuous subsequences in the sample time series feature data, evaluating the loss between the sample prediction noise and the sample noise, and training the diffusion model, the trained diffusion model can predict the spatio-temporal sequence signal of the entire brain, thereby realizing the generation of functional magnetic resonance imaging data with arbitrary length, high quality, and reliability. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, the above-mentioned content and other objects, features, and advantages of the present disclosure will become clearer.

[0016] Figure 1 FIG. shows a flowchart of a diffusion model training method according to an embodiment of the present disclosure.

[0017] Figure 2 FIG. shows a schematic diagram of an encoder and a decoder according to an embodiment of the present disclosure.

[0018] Figure 3 FIG. shows a flowchart of a functional magnetic resonance imaging data generation method according to an embodiment of the present disclosure.

[0019] Figure 4 FIG. shows a schematic diagram of a diffusion model according to an embodiment of the present disclosure.

[0020] Figure 5Shows a schematic diagram for verifying the quality of functional magnetic resonance imaging data according to an embodiment of the present disclosure.

[0021] Figure 6 Shows a structural block diagram of a diffusion model training device according to an embodiment of the present disclosure.

[0022] Figure 7 Shows a structural block diagram of a functional magnetic resonance imaging data generation device according to an embodiment of the present disclosure. Detailed implementation manners

[0023] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, many specific details are set forth in order to provide a thorough understanding of the embodiments of the present disclosure. However, obviously, one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present disclosure.

[0024] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising", etc. used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0025] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0026] In the case of using expressions such as "at least one of A, B, and C", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C).

[0027] In the process of implementing the present disclosure, it is found that in the related art, due to the characteristics of high dimensionality, complex spatio-temporal patterns, and small sample data in functional magnetic resonance imaging, in existing research, the method of synthesizing functional magnetic resonance imaging data does not focus on the level of original functional magnetic resonance imaging data, resulting in low accuracy and high limitations of the generated or predicted functional magnetic resonance imaging data.

[0028] In view of this, embodiments of the present disclosure provide a diffusion model training method, including: obtaining training samples, where the training samples include sample initial functional magnetic resonance imaging data of a sample target object and sample text prompt information; performing multi-view encoding processing on the sample initial functional magnetic resonance imaging data to obtain sample temporal feature data; using a noise splicing module to process sample noise and sample temporal feature data to obtain sample noise splicing data; based on the sample text prompt information, using an image segmentation convolution module to process the sample noise splicing data to obtain sample predicted noise, where the sample text prompt information is used to guide the predicted noise process; and training a diffusion model according to the sample noise and the sample predicted noise to obtain a trained diffusion model.

[0029] In the technical solution of the present disclosure, the involved user information (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties. Moreover, the processing of relevant data, such as collection, storage, use, processing, transmission, provision, disclosure, and application, all complies with relevant laws, regulations, and standards, takes necessary confidentiality measures, does not violate public order and good customs, and provides corresponding operation entrances for users to choose to authorize or refuse.

[0030] Figure 1 The flowchart of the diffusion model training method according to an embodiment of the present disclosure is shown.

[0031] As Figure 1 shown, the diffusion model training method of this embodiment includes operations S110 to S150.

[0032] In operation S110, training samples are obtained.

[0033] In operation S120, multi-view encoding processing is performed on the sample initial functional magnetic resonance imaging data to obtain sample temporal feature data.

[0034] In operation S130, a noise splicing module is used to process sample noise and sample temporal feature data to obtain sample noise splicing data.

[0035] In operation S140, based on the sample text prompt information, an image segmentation convolution module is used to process the sample noise splicing data to obtain sample predicted noise.

[0036] In operation S150, a diffusion model is trained according to the sample noise and the sample predicted noise to obtain a trained diffusion model.

[0037] According to an embodiment of the present disclosure, the training sample includes sample initial functional magnetic resonance imaging data of a sample target object and sample text prompt information. The sample target object is an object to be studied in the field of neurology. For example, the sample target object can be a patient with a mental disorder.

[0038] According to an embodiment of the present disclosure, the sample initial functional magnetic resonance imaging data (Functional Magnetic Resonance Imaging, fMRI) measures the blood oxygen signal data in blood flow by mapping neuron activities related to energy use in the brain. The preprocessed sample initial functional magnetic resonance imaging data is downloaded from a proprietary database platform.

[0039] According to an embodiment of the present disclosure, the sample text prompt information is used to guide the prediction noise process.

[0040] According to an embodiment of the present disclosure, the sample initial functional magnetic resonance imaging data includes blood oxygen signal data of brain slices or voxel points continuously collected within a certain time series. An encoder is used to encode and fuse the blood oxygen signal data at each moment in the sample initial functional magnetic resonance imaging data from different spatial perspectives to obtain sample time series feature data.

[0041] According to an embodiment of the present disclosure, the sample time series feature data is a highly aggregated structural feature information that fuses the blood oxygen signals at each moment.

[0042] According to an embodiment of the present disclosure, the diffusion model includes a noise splicing module and an image segmentation convolution module. The image segmentation convolution module can be constructed based on a convolutional neural network (U-Net). A cross-attention mechanism (Cross Attention) is embedded in the convolutional neural network (U-Net).

[0043] According to an embodiment of the present disclosure, the sample noise is random noise generated in the noise splicing module. The sample time series feature data is forward-denoised based on the sample noise to obtain sample noise splicing data. During the forward denoising process, the sample time series feature data is divided into subsequences and the conditional distribution between the subsequences is learned for continuously generating the sample noise splicing data of the subsequent subsequence.

[0044] According to an embodiment of the present disclosure, the sample noise splicing data is the denoised data corresponding to each of multiple moments with a time series dependence relationship.

[0045] According to an embodiment of the present disclosure, the sample text prompt information is used to guide the denoising process of the image segmentation convolution module. The sample text prompt information can be pathological information, such as the age information, gender information, and neurological diagnosis information of the sample target object.

[0046] In the technical solution of the present disclosure, the user information involved (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties. Moreover, the processing of relevant data, such as collection, storage, use, processing, transmission, provision, disclosure, and application, complies with relevant laws, regulations, and standards, takes necessary confidentiality measures, does not violate public order and good customs, and provides corresponding operation entrances for users to choose to authorize or reject.

[0047] According to an embodiment of the present disclosure, the sample text prompt information and the sample noise splicing data are input into the image segmentation convolution module to obtain sample predicted noise. The sample predicted noise is the noise added to the sample time series feature data during the prediction forward denoising process.

[0048] According to an embodiment of the present disclosure, the loss function is used to calculate the loss value between the sample noise and the sample predicted noise. When the loss value does not meet the preset loss value threshold, the parameters of the diffusion model are adjusted until the preset loss value threshold is met, and then the training stops to obtain the trained diffusion model.

[0049] According to an embodiment of the present disclosure, by mapping the sample initial functional magnetic resonance imaging data to a low-dimensional latent space, the spatial information and time series information of the data can be effectively captured from different spatial perspectives; using the sample text prompt information as the pathological condition to control the diffusion model to perform forward denoising diffusion and reverse denoising on the sample time series feature data, learning the conditional distribution between consecutive subsequences in the sample time series feature data, evaluating the loss between the sample predicted noise and the sample noise, and training the diffusion model. The trained diffusion model can predict the spatio-temporal sequence signals of the entire brain, thereby realizing the generation of functional magnetic resonance imaging data with arbitrary length, high quality, and reliability.

[0050] According to an embodiment of the present disclosure, there are F sample time series feature data and F sample noise splicing data; among them, using the noise splicing module to process the sample noise and the sample time series feature data, the obtained sample noise splicing data includes: based on the sample noise, performing forward denoising on the F sample time series feature data respectively to obtain the sample denoised data corresponding to each of the F sample time series feature data; splicing the sample initial embedding data and the first sample denoised data to obtain the first sample noise splicing data corresponding to the first sample denoised data; splicing the (f - 1)th sample time series feature data and the fth sample denoised data to obtain the fth sample noise splicing data corresponding to the fth sample denoised data, where the fth sample denoised data represents the sample denoised data corresponding to the fth sample time series feature data, and the (f - 1)th sample time series feature data represents the sample time series feature data corresponding to the (f - 1)th acquisition moment. 。

[0051] According to an embodiment of the present disclosure, the sample time-series feature data corresponds one-to-one with the acquisition time, and the F sample time-series feature data corresponding to each of the consecutive F acquisition times are 。 representing the f-th sample time-series feature data.

[0052] According to an embodiment of the present disclosure, the F sample time-series feature data are subjected to sequence partitioning to obtain multiple groups of subsequences. For example, the sample time-series feature data corresponding to each of the consecutive 8 times are taken as a group of subsequences. During the forward noise addition and denoising processes, time dependence modeling in the latent space is performed with each group of subsequences as a unit to learn the conditional distribution between two sample time-series feature data in each group of subsequences to obtain the conditional distribution corresponding to each group of subsequences 。

[0053] According to an embodiment of the present disclosure, the sample time-series feature data are sequential data. Therefore, during the sequence partitioning process, it is necessary to ensure the continuity of the partitioning segments and the constancy of the partitioning length.

[0054] According to an embodiment of the present disclosure, the sample noise can be random noise, and the random noise satisfies a Gaussian distribution. Based on the sample noise, forward noise addition is performed on the F sample time-series feature data respectively to obtain sample noise-added data corresponding to the F sample time-series feature data respectively 。 representing the f-th sample noise-added data.

[0055] According to an embodiment of the present disclosure, the sample initial embedding data can be a matrix with elements of 0.

[0056] According to an embodiment of the present disclosure, the sample initial embedding data and the first sample noise-added data are subjected to splicing processing to obtain the first sample noise splicing data ; the first sample time-series feature data and the second sample noise-added data are subjected to splicing processing to obtain the second sample noise splicing data 。

[0057] According to an embodiment of the present disclosure, the (f - 1)-th sample time-series feature data and the f-th sample noise-added data are subjected to splicing processing to obtain the f-th sample noise splicing data corresponding to the f-th sample noise-added data 。

[0058] According to an embodiment of the present disclosure, the (f - 1)-th sample time-series feature data represents the sample time-series feature data corresponding to the (f - 1)-th acquisition time.

[0059] For example, the first group of subsequence packets includes 8 sample time series feature data , and learn the conditional distribution corresponding to the first group of subsequences . Add noise to and splice the 8 sample time series feature data based on the conditional distribution to obtain sample noise spliced data corresponding to each of the 8 sample time series feature data .

[0060] According to an embodiment of the present disclosure, use an image segmentation convolutional module to process the 8 sample noise spliced data to obtain sample predicted noise, and remove the sample predicted noise from the 8 sample noise spliced data respectively to obtain 8 denoised data.

[0061] According to an embodiment of the present disclosure, add noise to and splice the sample time series feature data, and use an image segmentation convolutional module to denoise each sample noise spliced data, so as to perform latent space time dependence modeling on the temporally continuous sample time series feature data, so that the diffusion model can better understand and generate time series data, and predict and generate the sample time series feature data of the next moment based on the learned conditional distribution and the known sample signal feature data at the current moment, so as to realize the generation of long time series.

[0062] According to an embodiment of the present disclosure, the sample initial functional magnetic resonance imaging data includes sample signal data corresponding to each of the F acquisition times; wherein, performing multi-view encoding processing on the sample initial functional magnetic resonance imaging data to obtain sample time series feature data includes: for each sample signal data, using an encoder to extract features from the coronal plane view to obtain sample coronal plane feature data; using an encoder to extract features from the sagittal plane view to obtain sample sagittal plane feature data; using an encoder to extract features from the transverse plane view to obtain sample transverse plane feature data; merging the sample coronal plane feature data, sample sagittal plane feature data and sample transverse plane feature data to obtain sample time series feature data corresponding to each sample signal data.

[0063] According to an embodiment of the present disclosure, use an encoder to perform multi-view encoding processing on each sample signal data respectively to obtain sample time series feature data corresponding to each sample signal data.

[0064] According to an embodiment of the present disclosure, the encoder is composed of three two-dimensional convolutional network projections with the same structure, and use the three two-dimensional convolutional networks to extract features from the coronal plane view, sagittal plane view, and transverse plane view for each sample signal data respectively: extract features from the coronal plane view to obtain two-dimensional sample coronal plane feature data; extract features from the sagittal plane view to obtain two-dimensional sample sagittal plane feature data; extract features from the transverse plane view to obtain two-dimensional sample transverse plane feature data.

[0065] According to an embodiment of the present disclosure, the sample coronal plane feature data characterizes the spatial information in the coronal direction, the sample sagittal plane feature data characterizes the spatial information in the sagittal direction, and the sample transverse plane feature data characterizes the spatial information in the transverse direction.

[0066] According to an embodiment of the present disclosure, the sample coronal plane feature data, the sample sagittal plane feature data, and the sample transverse plane feature data are spliced to obtain the sample temporal feature data corresponding to this sample signal data.

[0067] Figure 2 A schematic diagram of an encoder and a decoder according to an embodiment of the present disclosure is shown.

[0068] As Figure 2 shown, the encoder includes a downsampling residual block, the decoder includes an upsampling residual block, the structures of the encoder and the decoder are symmetric, and both the downsampling residual block and the upsampling residual block are constructed based on a convolutional network. The encoder and the decoder are trained separately: the number of training rounds is 400, the learning rate is set to 0.00001, each sample signal data in the sample initial functional magnetic resonance imaging data is input into the encoder, and latent space representation learning is performed from the coronal plane perspective, the sagittal plane perspective, and the transverse plane perspective to obtain the sample coronal plane feature data, the sample sagittal plane feature data, and the sample transverse plane feature data; then the sample coronal plane feature data, the sample sagittal plane feature data, and the sample transverse plane feature data are respectively input into the sampling module to obtain the low-dimensional sample temporal feature data, and the sample temporal feature data is input into the decoder for dimension recovery, and the recovered functional magnetic resonance imaging data is output, and then the encoder and the decoder are trained according to the loss value between the recovered functional magnetic resonance imaging data and the sample initial functional magnetic resonance imaging data to obtain the trained encoder and decoder. The trained encoder can fully capture the spatial and temporal characteristics of the data, so as to accurately map each sample signal data in the sample initial functional magnetic resonance imaging data to the sample temporal feature data.

[0069] According to an embodiment of the present disclosure, the diffusion model further includes a denoising module, and the diffusion model training method further includes: using the denoising module to process the sample noise splicing data and the sample predicted noise to obtain the sample denoised data.

[0070] According to an embodiment of the present disclosure, the sample predicted noise is removed from each sample noise splicing data to obtain the sample denoised data. The sample denoised data is the predicted denoised data.

[0071] According to an embodiment of the present disclosure, training the diffusion model according to the sample noise and the sample predicted noise to obtain the trained diffusion model includes: using a loss function to calculate the loss value between the sample noise and the sample predicted noise to obtain the target loss value; training the diffusion model according to the target loss value to obtain the trained diffusion model.

[0072] According to an embodiment of the present disclosure, during the process of training a diffusion model, a plurality of sample time-series feature data is divided into two parts. One part of the sample time-series feature data is used for the diffusion model to learn the conditional distribution, and the other part of the sample time-series feature data is used for the diffusion model to learn the unconditional distribution. Freeze the parameters of the encoder, and train the diffusion model alone. Use Adam with Weight Decay (AdamW) as the optimizer, set the number of training epochs to 400, and set the learning rate to 0.00001.

[0073] In one embodiment, the loss function is as shown in formula (1):

[0074] (1);

[0075] where is a control parameter. When , the unconditional distribution is trained. When , the conditional distribution is trained. represents the sample signal data at the first acquisition moment. represents the sample signal data at the second acquisition moment. represents the sample time-series feature data at the first acquisition moment. represents the sample time-series feature data at the second acquisition moment. represents the sample noise splicing data at the second acquisition moment. , , represents the encoder, t represents the number of times of sample noise addition. represents the actual added sample noise when obtaining the sample noise splicing data at the second acquisition moment based on the sample time-series feature data at the first acquisition moment during the t-th noise addition process. represents the predicted sample noise when using the noise predictor to predict the sample noise splicing data at the second acquisition moment based on the sample time-series feature data at the first acquisition moment during the t-th noise addition process.

[0076] Figure 3 Shows a flowchart of a functional magnetic resonance imaging data generation method according to an embodiment of the present disclosure.

[0077] As Figure 3 shown, the functional magnetic resonance imaging data generation method includes operations S310 to S330.

[0078] In operation S310, obtain initial embedding data and text prompt information related to the image generation task.

[0079] In operation S320, based on the preset time series length S, the diffusion model is used to process the initial embedding data, text prompt information, and noise, and S denoised data are output.

[0080] In operation S330, the decoder is used to process the S denoised data to generate the target functional magnetic resonance imaging data.

[0081] According to an embodiment of the present disclosure, the image generation task is to generate high-quality target functional magnetic resonance imaging data with a preset time series length. The initial embedding data can be a matrix with elements of 0.

[0082] According to an embodiment of the present disclosure, the text prompt information is used to guide the prediction noise process.

[0083] According to an embodiment of the present disclosure, the preset time series length is the length for generating the target functional magnetic resonance imaging data preset. For example, the preset time series length is 10, and the target functional magnetic resonance imaging data includes signal data at 10 consecutive moments respectively.

[0084] According to an embodiment of the present disclosure, the trained diffusion model can better understand and generate time series data, predict and generate the signal data at the next moment based on the learned conditional distribution and the initial embedding data, and realize the generation of a long time series.

[0085] According to an embodiment of the present disclosure, based on the preset time series length S, the diffusion model is used to process the initial embedding data, text prompt information, and noise, and S denoised data are output.

[0086] According to an embodiment of the present disclosure, the decoder is used to perform dimensionality recovery processing on the S low-dimensional denoised data respectively to generate S signal data with time series dependence, which constitute the target functional magnetic resonance imaging data.

[0087] According to an embodiment of the present disclosure, the trained diffusion model can better understand and generate time series data, and based on the learned conditional distribution, realize the generation of target functional magnetic resonance imaging data with any sequence length, high quality, and reliability.

[0088] According to an embodiment of the present disclosure, based on a preset time series length S, a diffusion model is used to process initial embedding data, text prompt information, and noise, and S denoised data are output, including: processing the noise and the initial embedding data using a noise splicing module to obtain a first noise splicing data; based on the text prompt information, processing the first noise splicing data using an image segmentation convolution module to obtain a first predicted noise; processing the first noise splicing data and the first predicted noise using a denoising module to obtain a first denoised data; based on the preset time series length s, processing the noise and the (s - 1)-th denoised data using a noise splicing module to obtain an s-th noise splicing data, where the (s - 1)-th denoised data is obtained by denoising the (s - 1)-th noise splicing data. ; based on the text prompt information, processing the s-th noise splicing data using an image segmentation convolution module to obtain an s-th predicted noise; processing the s-th noise splicing data and the s-th predicted noise using a denoising module to obtain an s-th denoised data.

[0089] According to an embodiment of the present disclosure, a random noise is added to the initial embedding data using a noise splicing module to obtain a first noise splicing data; the text prompt information can be the pathological information of the target object, and based on the text prompt information, an image segmentation convolution module is guided to process the first noise splicing data to obtain a first predicted noise; a denoising module is used to remove the first predicted noise from the first noise splicing data to obtain a first denoised data. The first predicted noise is the noise added when predicting the generation of the first noise splicing data based on the initial embedding data.

[0090] According to an embodiment of the present disclosure, the (s - 1)-th denoised data is obtained by denoising the (s - 1)-th noise splicing data; a noise is added to the (s - 1)-th denoised data using a noise splicing module to obtain an s-th noise splicing data; an image segmentation convolution module is guided to process the s-th noise splicing data based on the text prompt information to obtain an s-th predicted noise; a denoising module is used to remove the s-th predicted noise from the s-th noise splicing data to obtain an s-th denoised data.

[0091] According to an embodiment of the present disclosure, S temporally consecutive denoised data are generated based on the preset time series length S.

[0092] According to an embodiment of the present disclosure, the functional magnetic resonance imaging data generation method further includes: processing the denoised data using a classifier to obtain a classification result, where the classification result characterizes whether the target object is a patient with mental disorder or a non-mental disorder patient.

[0093] According to an embodiment of the present disclosure, the text prompt information can be the pathological information of the target object in neurodiagnosis, and the denoising process is guided by the text prompt information to generate denoised data, and the denoised data is processed using a classifier to obtain the classification result of the target object.

[0094] According to an embodiment of the present disclosure, Autism Spectrum Disorder features are highly heterogeneous, which makes it a major challenge to learn the conditional distribution containing class information in the diffusion model. Therefore, a classifier is additionally added during the diffusion model training phase. The classifier can be constructed based on the Residual Network (ResNet), and then the cross-entropy loss function is used to optimize the classifier to ensure that the diffusion model can fully learn the conditional distribution, thereby obtaining high-quality and accurate denoised data and improving the classification accuracy of the classifier.

[0095] According to an embodiment of the present disclosure, the classification results of the classifier can be used to verify the consistency between the generated target functional magnetic resonance imaging (fMRI) data and the initial fMRI data of autism, and prove the reliability of the present disclosure embodiment in generating target fMRI data of any sequence length.

[0096] Figure 4 FIG. shows a schematic diagram of a diffusion model according to an embodiment of the present disclosure.

[0097] As Figure 4 shown, the sample initial fMRI data is input into the encoder 410 to obtain multiple sample temporal feature data; the multiple sample temporal feature data is input into the diffusion model 420 to obtain the f-th sample noise concatenated data which may include: in the noise concatenation module 421, forward noise addition is performed on the F sample temporal feature data based on the sample noise respectively to obtain F sample noise-added data, and the sample initial embedding data Z0 and the first sample noise-added data are concatenated to obtain the first sample noise concatenated data ; the (f - 1)-th sample temporal feature data and the f-th sample noise-added data are concatenated to obtain the f-th sample noise concatenated data ; cross-attention is embedded in the image segmentation convolutional module 422, and the F sample noise concatenated data and the sample text prompt information are input into the image segmentation convolutional module 422 to obtain the sample predicted noise. The diffusion model is trained according to the loss value between the sample predicted noise and the sample noise ; in the denoising module 423, the sample predicted noise is removed from the F sample noise concatenated data respectively to obtain F sample denoised data. In addition, the F sample denoised data is input into the classifier 430 to obtain the sample classification result, and the classifier is trained according to the loss value between the sample classification result and the sample label ; the sample denoised data is input into the decoder 440 to generate the restored fMRI data.

[0098] Figure 5A schematic diagram showing the verification of the quality of target functional magnetic resonance imaging data according to an embodiment of the present disclosure.

[0099] As Figure 5 shown, on the generated target functional magnetic resonance imaging data, different brain region division templates or atlases (such as Fair34, AAL90, Dosenbach160, Yeo7, RSN28, and Neuromark53) are used to divide and extract the functional connectivity matrix data FC. The FC change of the target functional magnetic resonance imaging data is compared with the sample initial functional magnetic resonance imaging data, and the Pearson correlation coefficient is measured for consistency. ASD is the experimental data based on objects with Autism Spectrum Disorder, and HC is the experimental data based on objects in the Healthy Control group. Since the FC matrix is symmetric, the upper right triangular region in the functional connectivity matrix FC represents the enhanced region, and the lower left triangular region in the functional connectivity matrix FC represents the original unenhanced region. In part A: In the node-based FC analysis, the average functional connectivity of the Fair34, AAL90, and Dosenbach160 atlases was evaluated. The correlation coefficient range of the unenhanced region and the enhanced region is 0.92 - 0.96, indicating that the target functional magnetic resonance imaging data generated by this method has consistent and stable functional connectivity. In part B: In the network-based FC analysis, the average functional connection changes of the Yeo7, RSN28, and Neuromark53 networks were detected, and the correlation coefficient range is 0.96 - 0.99, further confirming the global consistency of the target functional magnetic resonance imaging data generated by this method at the network level.

[0100] Based on the above diffusion model training method, the present disclosure also provides a diffusion model training device. The following will be combined with Figure 6 to describe this device in detail.

[0101] Figure 6 A structural block diagram showing the diffusion model training device according to an embodiment of the present disclosure.

[0102] As Figure 6 shown, the diffusion model training device 600 of this embodiment includes a first acquisition module 610, an encoding module 620, a noise addition module 630, a first prediction module 640, and a training module 650.

[0103] The first acquisition module 610 is used to acquire training samples, and the training samples include the sample initial functional magnetic resonance imaging data of the sample target object and the sample text prompt information. In one embodiment, the first acquisition module 610 can be used to perform the operation S110 described above, which will not be elaborated here.

[0104] The encoding module 620 is configured to perform multi-perspective encoding processing on the sample initial functional magnetic resonance imaging data to obtain sample time-series feature data. In one embodiment, the encoding module 620 may be configured to perform the operation S120 described above, which will not be elaborated here.

[0105] The noise adding module 630 is configured to process the sample noise and the sample time-series feature data by using a noise splicing module to obtain sample noise spliced data. In one embodiment, the noise adding module 630 may be configured to perform the operation S130 described above, which will not be elaborated here.

[0106] The first prediction module 640 is configured to process the sample noise spliced data by using an image segmentation convolutional module based on the sample text prompt information to obtain sample predicted noise, where the sample text prompt information is used to guide the predicted noise process. In one embodiment, the first prediction module 640 may be configured to perform the operation S140 described above, which will not be elaborated here.

[0107] The training module 650 is configured to train a diffusion model according to the sample noise and the sample predicted noise to obtain a trained diffusion model. In one embodiment, the training module 650 may be configured to perform the operation S150 described above, which will not be elaborated here.

[0108] According to an embodiment of the present disclosure, the noise adding module 630 includes a first noise adding sub-module, a second noise adding sub-module, and a third noise adding sub-module.

[0109] The first noise adding sub-module is configured to perform forward noise adding on F sample time-series feature data respectively based on the sample noise to obtain sample noise added data corresponding to the F sample time-series feature data respectively.

[0110] The second noise adding sub-module is configured to perform splicing processing on the sample initial embedding data and the first sample noise added data to obtain the first sample noise spliced data corresponding to the first sample noise added data.

[0111] The third noise adding sub-module is configured to perform splicing processing on the (f - 1)-th sample time-series feature data and the f-th sample noise added data to obtain the f-th sample noise spliced data corresponding to the f-th sample noise added data, where the f-th sample noise added data represents the sample noise added data corresponding to the f-th sample time-series feature data, and the (f - 1)-th sample time-series feature data represents the sample time-series feature data corresponding to the (f - 1)-th acquisition moment. .

[0112] According to an embodiment of the present disclosure, the encoding module 620 includes a first encoding sub-module, a second encoding sub-module, a third encoding sub-module, and a fourth encoding sub-module.

[0113] The first encoding sub-module is used to extract features from the coronal plane perspective for each sample signal data by using an encoder, so as to obtain sample coronal plane feature data.

[0114] The second encoding sub-module is used to extract features from the sagittal plane perspective by using an encoder, so as to obtain sample sagittal plane feature data.

[0115] The third encoding sub-module is used to extract features from the transverse plane perspective by using an encoder, so as to obtain sample transverse plane feature data.

[0116] The fourth encoding sub-module is used to merge the sample coronal plane feature data, the sample sagittal plane feature data and the sample transverse plane feature data to obtain sample temporal sequence feature data corresponding to each sample signal data.

[0117] According to an embodiment of the present disclosure, the diffusion model training device 600 further includes an obtaining module.

[0118] The obtaining module is used to process the sample noise splicing data and the sample predicted noise by using the denoising module to obtain sample denoised data.

[0119] Figure 7 The structural block diagram of a functional magnetic resonance imaging data generation device according to an embodiment of the present disclosure is shown.

[0120] As Figure 7 shown, the functional magnetic resonance imaging data generation device 700 of this embodiment includes a second obtaining module 710, a second prediction module 720 and a generation module 730.

[0121] The second obtaining module 710 is used to obtain initial embedding data and text prompt information related to an image generation task, wherein the text prompt information is used to guide the predicted noise process. In one embodiment, the second obtaining module 710 may be used to perform the operation S310 described above, which will not be elaborated here.

[0122] The second prediction module 720 is used to output S denoised data by using the diffusion model to process the initial embedding data, the text prompt information and the noise based on a preset time sequence length S, wherein the diffusion model is trained according to the above-mentioned diffusion model training device. In one embodiment, the second prediction module 720 may be used to perform the operation S320 described above, which will not be elaborated here.

[0123] The generation module 730 is used to process the S denoised data by using a decoder to generate target functional magnetic resonance imaging data. In one embodiment, the generation module 730 may be used to perform the operation S330 described above, which will not be elaborated here.

[0124] According to an embodiment of the present disclosure, the second prediction module 720 includes a first prediction sub-module, a second prediction sub-module, a third prediction sub-module, a fourth prediction sub-module, a fifth prediction sub-module, and a sixth prediction sub-module.

[0125] The first prediction sub-module is configured to process noise and initial embedding data by using a noise splicing module to obtain the first noise splicing data.

[0126] The second prediction sub-module is configured to process the first noise splicing data by using an image segmentation convolution module based on text prompt information to obtain the first predicted noise.

[0127] The third prediction sub-module is configured to process the first noise splicing data and the first predicted noise by using a denoising module to obtain the first denoised data.

[0128] The fourth prediction sub-module is configured to process noise and the (s - 1)-th denoised data by using a noise splicing module based on a preset time series length s to obtain the s-th noise splicing data, where the (s - 1)-th denoised data is obtained by denoising the (s - 1)-th noise splicing data, 1 .

[0129] The fifth prediction sub-module is configured to process the s-th noise splicing data by using an image segmentation convolution module based on text prompt information to obtain the s-th predicted noise.

[0130] The sixth prediction sub-module is configured to process the s-th noise splicing data and the s-th predicted noise by using a denoising module to obtain the s-th denoised data.

[0131] According to an embodiment of the present disclosure, the functional magnetic resonance imaging data generation device 700 further includes a classification module.

[0132] The classification module is configured to process the denoised data by using a classifier to obtain a classification result, where the classification result represents that the target object is a patient with mental disorder or a non-patient with mental disorder.

[0133] According to embodiments of the present disclosure, any plurality of modules among modules, sub-modules, units, and sub-units may be combined and implemented in one module, or any one of them may be split into multiple modules. Or, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to embodiments of the present disclosure, at least one of the modules, sub-modules, units, and sub-units may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or may be implemented by any other reasonable means such as hardware or firmware for integrating or packaging circuits, or may be implemented in any one of the three implementation manners of software, hardware, and firmware, or in any suitable combination of several of them. Or, at least one of the modules, sub-modules, units, and sub-units may be at least partially implemented as a computer program module, and when the computer program module is run, corresponding functions may be executed.

[0134] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in an order different from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0135] Those skilled in the art can understand that the features described in the various embodiments of the present disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features described in the various embodiments of the present disclosure can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present disclosure.

[0136] The embodiments of the present disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications shall fall within the scope of the present disclosure.

Claims

1. A diffusion model training method, characterized in that, The diffusion model includes a noise splicing module and an image segmentation convolution module, and the method includes: Obtain training samples, where the training samples include sample initial functional magnetic resonance imaging data of the sample target object and sample text prompt information; Perform multi-view encoding processing on the sample initial functional magnetic resonance imaging data to obtain sample temporal feature data; Use the noise splicing module to process the sample noise and the sample temporal feature data to obtain sample noise splicing data; Based on the sample text prompt information, use the image segmentation convolution module to process the sample noise splicing data to obtain sample predicted noise, where the sample text prompt information is used to guide the predicted noise process; Train the diffusion model according to the sample noise and the sample predicted noise to obtain a trained diffusion model.

2. The method according to claim 1, wherein There are F pieces of the sample temporal feature data, and there are F pieces of the sample noise splicing data; Among them, the using the noise splicing module to process the sample noise and the sample temporal feature data to obtain sample noise splicing data includes: Based on the sample noise, perform forward noise addition on the F pieces of the sample temporal feature data respectively to obtain sample noise-added data corresponding to each of the F pieces of the sample temporal feature data; Perform splicing processing on the sample initial embedding data and the first sample noise-added data to obtain the first sample noise splicing data corresponding to the first sample noise-added data; Perform splicing processing on the (f - 1)-th sample time-series feature data and the f-th sample noise-added data to obtain the f-th sample noise-spliced data corresponding to the f-th sample noise-added data, where the f-th sample noise-added data represents the sample noise-added data corresponding to the f-th sample time-series feature data, the (f - 1)-th sample time-series feature data represents the sample time-series feature data corresponding to the (f - 1)-th acquisition moment, 1 f F。 3. The method according to claim 1, characterized in that, The sample initial functional magnetic resonance imaging data includes sample signal data at each of the F acquisition times; Among them, the performing multi-view encoding processing on the sample initial functional magnetic resonance imaging data to obtain sample temporal feature data includes: For each sample signal data, use an encoder to extract features from the coronal plane view to obtain sample coronal plane feature data; Use an encoder to extract features from the sagittal plane view to obtain sample sagittal plane feature data; Use an encoder to extract features from the transverse plane view to obtain sample transverse plane feature data; Merge the sample coronal plane feature data, the sample sagittal plane feature data, and the sample transverse plane feature data to obtain sample temporal feature data corresponding to each sample signal data.

4. The method according to claim 3, characterized in that, The diffusion model further includes a denoising module, and the method further includes: Use the denoising module to process the sample noise splicing data and the sample predicted noise to obtain sample denoised data.

5. The method according to claim 1, characterized in that, The training the diffusion model according to the sample noise and the sample predicted noise to obtain a trained diffusion model includes: Use a loss function to calculate the loss value between the sample noise and the sample predicted noise to obtain a target loss value; Train the diffusion model according to the target loss value to obtain a trained diffusion model.

6. A method for generating functional magnetic resonance imaging data, characterized in that, The method includes: Obtain initial embedding data and text prompt information related to an image generation task, where the text prompt information is used to guide the predicted noise process; Based on a preset time sequence length S, use the diffusion model to process the initial embedding data, the text prompt information, and noise, and output S pieces of denoised data, where the diffusion model is trained according to the method described in claims 1 to 5. Process the S denoised data using a decoder to generate target functional magnetic resonance imaging data.

7. The method according to claim 6, wherein Based on the preset time series length S, processing the initial embedding data, the text prompt information, and noise using a diffusion model, and outputting S denoised data includes: Process the noise and the initial embedding data using a noise splicing module to obtain the first noise-spliced data; Based on the text prompt information, process the first noise-spliced data using an image segmentation convolutional module to obtain the first predicted noise; Process the first noise-spliced data and the first predicted noise using a denoising module to obtain the first denoised data; Based on a preset timing length s, the noise splicing module processes the noise and the (s - 1)-th denoised data to obtain the s-th noise splicing data, where the (s - 1)-th denoised data is obtained by denoising the (s - 1)-th noise splicing data, 1 ; Based on the text prompt information, process the s-th noise-spliced data using an image segmentation convolutional module to obtain the s-th predicted noise; Process the s-th noise-spliced data and the s-th predicted noise using the denoising module to obtain the s-th denoised data.

8. The method according to claim 6, characterized in that The method further includes: Process the denoised data using a classifier to obtain a classification result, where the classification result indicates whether the target object is a patient with a mental disorder or a non-mental disorder patient.

9. A diffusion model training device, characterized in that, The device includes: A first acquisition module, configured to acquire training samples, where the training samples include sample initial functional magnetic resonance imaging data of a sample target object and sample text prompt information; An encoding module, configured to perform multi-view encoding processing on the sample initial functional magnetic resonance imaging data to obtain sample time series feature data; A noise addition module, configured to process sample noise and the sample time series feature data using a noise splicing module to obtain sample noise-spliced data; A first prediction module, configured to process the sample noise-spliced data using an image segmentation convolutional module based on the sample text prompt information to obtain sample predicted noise, where the sample text prompt information is used to guide the predicted noise process; A training module, configured to train the diffusion model according to the sample noise and the sample predicted noise to obtain a trained diffusion model.

10. A functional magnetic resonance imaging data generation device, characterized in that, The device includes: A second acquisition module, configured to acquire initial embedding data and text prompt information related to an image generation task, where the text prompt information is used to guide the predicted noise process; A second prediction module, configured to output S denoised data by processing the initial embedding data, the text prompt information, and noise using a diffusion model based on a preset time series length S, where the diffusion model is trained according to the device described in claim 9; A generation module, configured to process the S denoised data using a decoder to generate target functional magnetic resonance imaging data.