A multimodal medical image reconstruction method based on adversarial diffusion model
By combining the adversarial diffusion model and multimodal conditional encoder, the problem of low-quality reconstruction caused by ignoring multimodal information in the prior art is solved, and high-quality SPET image reconstruction and diagnostic utility are improved.
Patent Information
- Application Number
- CN202410853360.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-28
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-06-28
AI Technical Summary
The prior art ignores important information in the multimodal input when reconstructing standard dose PET (SPET) images from low-dose PET (LPET) images, resulting in low reconstruction quality and limited diagnostic utility.
Adversarial diffusion model is adopted, combining a multimodal conditional encoder and a multimodal mask text reconstruction module to extract information from LPET images and clinical table data, and high-quality SPET images are reconstructed through the conditional diffusion process.
By combining multimodal input, the quality of the reconstruction image is significantly improved, the diagnostic utility is enhanced, and the semantic distortion is reduced, ensuring semantic consistency during the reconstruction process.
Smart Images

Figure CN118628602B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and in particular relates to a multi-modal medical image reconstruction method based on a counter-diffusion model. Background Art
[0002] Positron emission tomography (PET) is an advanced nuclear imaging technique that can visualize and quantify metabolic activities in the human body. In clinical practice, standard-dose PET (SPET) images are of high quality and rich diagnostic information, which is what doctors need. However, the high radiation exposure associated with SPET imaging has raised health concerns. The radiation hazards associated with standard-dose positron emission tomography (SPET) images remain a concern, while the quality of low-dose PET (LPET) images does not meet clinical requirements. To address this issue, the injected tracer dose can be reduced, but this may induce unexpected noise and artifacts, resulting in reduced image quality and limited diagnostic value.
[0003] To address this challenge, people have focused on reconstructing SPET images from LPET images. However, previous studies have focused only on image data, ignoring important complementary information from other modalities, such as the patient's clinical form, resulting in impaired reconstruction and limited diagnostic utility. In addition, they often ignore the semantic consistency between the real SPET and the reconstructed images, resulting in distorted semantic context. Summary of the invention
[0004] In order to solve the above problems, the present invention proposes a multimodal medical image reconstruction method based on a diffusion model to reconstruct SPET images from multimodal inputs to solve the above problems.
[0005] To achieve the above object, the technical solution adopted by the present invention is: a multimodal medical image reconstruction method based on a diffusion-resistant model, comprising the steps of:
[0006] S10, obtaining LPET images and clinical form sample data;
[0007] S20, establishing a medical image reconstruction model based on conditional diffusion, including an adversarial diffusion network, a multimodal conditional encoder and a multimodal masked text reconstruction module; during the training process, the adversarial diffusion network combines the diffusion generator and the diffusion discriminator to diffuse on the real SPET image, and the multimodal conditional encoder extracts information from the LPET image and clinical form data input and injects it into the diffusion generator in a hierarchical manner; at the same time, the clinical form is also input into the multimodal masked text reconstruction module for reconstruction;
[0008] S30, given the LPET image and clinical table detection data, input them into the medical image reconstruction model based on conditional diffusion, and use the adversarial diffusion network and multimodal conditional encoder to iteratively convert the noise into an estimated PET image.
[0009] Furthermore, the adversarial diffusion network adopts a conditional diffusion framework, including a forward diffusion process and a reverse diffusion process; a 4-level diffusion generator and a 5-level diffusion discriminator are used to accelerate the reverse sampling process in an adversarial manner.
[0010] Furthermore, the forward diffusion process: in t time steps, a fixed Gaussian noise is gradually added to the SPET image y0 through a Markov chain to generate a series of noise images y1 to y t .
[0011] Furthermore, the reverse diffusion process:
[0012] By gradually deriving the noise from y at t time steps t Restore to y0; at each time step t, provided with the current noise data y t and conditional information c, the neural network parameterized by θ predicts the denoised data y t-1 The conditional probability distribution of ; the expected prediction distribution p θ (y t-1 ly t ,c) will be close to its actual corresponding distribution q(y t-1 ly t ,c);
[0013] Using GAN network, GAN network uses conditional diffusion generator G θ To predict p θ (y t-1 ly t ,c), and use the diffusion discriminator To minimize the predicted p θ (y t-1 ly t ,c) and the actual distribution q(y t-1 ly t ,c).
[0014] Furthermore, in the GAN network, the conditional diffusion generator G θ Receive data pair (y t ,c) as input to predict the uncorrupted image y'0; then, the noise is reintroduced into y'0 using the posterior distribution, from p θ (y t-1 ly t ,c) sampling y′ t-1 .
[0015] For adversarial training, ({y′ t-1 or t-1}, yt, t) aims to distinguish the estimated posterior distribution p θ (y t-1 ly t ,c) The sample extracted is used as the pseudo sample and the real corresponding sample q(y t-1 ly t ,c) as true samples.
[0016] Furthermore, the multimodal conditional encoder is used to extract multimodal features, which are then introduced into a diffusion generator for guidance; in order to balance the heterogeneous modes of images and tables, an optimal multimodal transfer co-attention module is used to integrate table features and image features while inferring their correlation.
[0017] Furthermore, in the table branch of the multimodal conditional encoder, the clinical table detection data x tb is fed into the CXR-BERT model to generate the tabular features p tb ;
[0018] In the image branch of the multimodal conditional encoder, corresponding to the encoding part, the image branch involves 4 levels. Given an LPET image x l , extract multi-level features x cnd .
[0019] Further, in the optimal multimodal transfer co-attention module, after each level of the image branch, an optimal multimodal transfer co-attention module is integrated to achieve active interaction between image and table features;
[0020] The optimal multimodal transfer co-attention module exploits the matching ability of optimal transfer to reconcile image-table differences, while leveraging the selective focusing ability of the attention mechanism to emphasize image regions related to more relative attributes in the table.
[0021] Furthermore, the multimodal mask text reconstruction module includes the following steps in its processing:
[0022] Given clinical table test data x tb , encode the sentences based on the CXR-BERT module and construct a real text prompt P T ;
[0023] Extract clinical table attributes in sentences, mask the attributes, and encode the masked sentences into masked text prompts P based on the CXR-BERT module m middle;
[0024] Apply global average pooling to the denoised image y'0 to obtain the semantic embedding z'0; m Connected with z'0, the masked text reconstructor R m Perform joint decoding to reconstruct the mask attributes and obtain the reconstructed text hint;
[0025] Minimize the difference between the reconstructed text prompt and the true text prompt during training.
[0026] Furthermore, the objective function of the medical image reconstruction model based on conditional diffusion is:
[0027] L totat =L adv +λ1L image +λ 2text ;
[0028] Among them, λ1, λ2 and λ3 are weighting coefficients, L adv is the countervailing diffusion loss, L image is the image estimation loss, L text is the mask reconstruction loss.
[0029] The beneficial effects of adopting this technical solution are:
[0030] The method constructed by the present invention reconstructs high-quality SPET images from LPET images and clinical table data. Unlike previous methods that rely only on image data, the present invention combines multimodal conditions, using LPET images to fundamentally guide reconstruction and non-image clinical tables, providing a complementary perspective for further improving image quality.
[0031] The present invention integrates a multimodal conditional encoder based on a multimodal conditional adversarial diffusion model to extract multimodal features, then mixes noise with the multimodal features through a conditional diffusion process, and gradually maps the mixed features to a target SPET image.
[0032] In order to fully utilize the information in the multimodal input and balance the multimodal input, an optimal multimodal transfer co-attention is embedded in the multimodal conditional encoder to narrow the heterogeneity gap between images and tables while capturing their interactions to provide sufficient guidance for reconstruction.
[0033] In addition, in order to alleviate semantic distortion and ensure the consistency of high-level semantics that are easily distorted during the reconstruction process, the present invention establishes a multimodal masked text reconstruction, which utilizes the semantic knowledge extracted from the denoised PET images to restore the masked clinical forms. Since the precise semantics in the denoised images can effectively promote the accurate recovery of the masked clinical attributes, the multimodal masked text reconstruction in turn strengthens the retention of semantics in the reconstruction, thereby forcing the maintenance of accurate semantics during the reconstruction process.
[0034] To speed up the diffusion process, an adversarial diffusion network with reduced diffusion steps is further introduced to reduce the number of diffusion steps to speed up the diffusion process.
[0035] The experiments in the embodiments of the present invention show that the proposed method achieves the most advanced performance in terms of both quality and quantity. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 A schematic diagram of a process flow of a multimodal medical image reconstruction method based on a diffusion-resistant model according to the present invention;
[0037] Figure 2 This is a comparison chart of visual results of reconstructing images using different methods in a comparative embodiment of the present invention.
[0038] Figure 3 4 is a comparison chart of the results of ablation experiments in an embodiment of the present invention. DETAILED DESCRIPTION
[0039] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is further described below with reference to the accompanying drawings.
[0040] In this embodiment, see Figure 1 As shown, the present invention proposes a multimodal medical image reconstruction method based on a counter-diffusion model, comprising the steps of:
[0041] S10, obtaining LPET images and clinical form sample data;
[0042] S20, establish a medical image reconstruction model based on conditional diffusion, including an adversarial diffusion network, a multimodal conditional encoder, and a multimodal masked text reconstruction module; during the training process, the adversarial diffusion network combines the diffusion generator and the diffusion discriminator to diffuse on real SPET images, aiming to accelerate the diffusion process. In order to achieve a more controllable diffusion process and enhance the reconstruction quality, the multimodal conditional encoder extracts information from the LPET image and clinical table data input and injects it into the diffusion generator in a hierarchical manner; at the same time, the clinical table is also input into the multimodal masked text reconstruction module for reconstruction;
[0043] S30, given the LPET image and clinical table detection data, input them into the medical image reconstruction model based on conditional diffusion, and use the adversarial diffusion network and multimodal conditional encoder to iteratively convert the noise into an estimated PET image.
[0044] For the optimization scheme of the above embodiment, the adversarial diffusion network adopts a conditional diffusion framework, including a forward diffusion process and a reverse diffusion process; a 4-level diffusion generator and a 5-level diffusion discriminator are used to accelerate the reverse sampling process in an adversarial manner.
[0045] The forward diffusion process is as follows: at t time steps, a fixed Gaussian noise is gradually added to the SPET image y0 through a Markov chain to generate a series of noise images y1 to y t .
[0046] The reverse diffusion process is as follows: by gradually deriving the noise from y at t time steps, t Restore to y0; at each time step t, provided with the current noise data y t and conditional information c, the neural network parameterized by θ predicts the denoised data y t-1 The conditional probability distribution of ; the expected prediction distribution p θ (y t-1 ly t ,c) will be close to its actual corresponding distribution q(y t- 1ly t ,c);
[0047] In the traditional diffusion model, p θ (y t-1 ly t ,c) follows a Gaussian distribution, because the step size from t to t-1 is small enough to produce a Gaussian distribution of q(y t-1 ly t ,c); however, when the step size becomes larger, q(y t-1 ly t ,c); becomes non-Gaussian. Therefore, for fast reverse sampling (i.e., larger step size), the present invention breaks the traditional Gaussian assumption,
[0048] Using GAN network, GAN network uses conditional diffusion generator G θ To predict p θ (y t-1 ly t ,c), and use the diffusion discriminator To minimize the predicted p θ (y t-1 ly t ,c) and the actual distribution q(y t-1 ly t ,c).
[0049] Preferably, in the GAN network, a conditional diffusion generator G is used θ Receives a data pair (yt,c) as input to predict the uncorrupted image y'0; then, the noise is reintroduced into y'0 using the posterior distribution, from p θ (y t-1 ly t ,c) sampling y′ t-1 .
[0050] The specific formula is:
[0051] y′ t-1 ~p θ (y t-1 |y t ):=q(y t-1 |y t , y′0=G θ (y t ,c,t));
[0052] For adversarial training, ({y′ t-1 or t-1}, yt, t) aims to distinguish the estimated posterior distribution p θ (y t- 1ly t ,c) The sample extracted is used as the pseudo sample and the real corresponding sample q(y t-1 ly t ,c) as true samples.
[0053] The above process achieves mutual benefits in reducing the diffusion time step T and guiding the network to capture the complex distribution in real SPET images.
[0054] As an optimization scheme for the above embodiment, the multimodal conditional encoder is used to extract multimodal features, which are then introduced into the diffusion generator for guidance; in order to balance the heterogeneous modes of images and tables, the optimal multimodal transfer co-attention (OMTA) module is used to integrate table features and image features, and their correlation is inferred at the same time.
[0055] Considering the complex metabolic distribution in PET images, it is challenging to accurately restore SPET images only based on LPET images. Therefore, the present invention introduces a multi-modal conditional input x m ={x l ,x tb}, containing the LPET image x l , also contains clinical table test data x tb , to guide more precise diffusion processes. LPET images provide basic guidance for guiding reconstruction, while clinical tables provide complementary metabolism-related perspectives for further image quality refinement.
[0056] Among them, in the table branch of the multimodal conditional encoder, the clinical table detection data x tb is fed into the CXR-BERT model to generate the tabular features p tb ;
[0057] Pre-trained on biomedical texts, the CXR-BERT model can impart biomedical context to PTB, providing it with clinical cues for high-quality reconstruction.
[0058] In the image branch of the multimodal conditional encoder, corresponding to the encoding part, the image branch involves 4 levels. Given an LPET image x l , extract multi-level features x cnd .
[0059] Wherein, in the optimal multimodal transfer co-attention module (OMTA), an optimal multimodal transfer co-attention module is integrated after each level of the image branch to achieve active interaction between image and table features;
[0060] The optimal multimodal transfer co-attention module exploits the matching ability of optimal transfer to reconcile image-table differences, while leveraging the selective focusing ability of the attention mechanism to emphasize image regions related to more relative attributes in the table.
[0061] Specifically, after the i-th level in the image branch, the image feature vi is flattened to v i * ; At the same time, the table feature p tb is projected to match the dimensions of the flattened image, yielding p tb * The optimal transport between them is then defined via the discrete Kantorovich formula, which searches for the overall optimal matching flow between v and pth.
[0062] As an optimization solution of the above embodiment, the multimodal masked text reconstruction module retains basic semantic information during the reconstruction process, and the processing process includes the following steps:
[0063] Given clinical table test data x tb , encode the sentences based on the CXR-BERT module and construct a real text prompt P T ;
[0064] Extract clinical table attributes in sentences, mask the attributes, and encode the masked sentences into masked text prompts P based on the CXR-BERT module m middle;
[0065] Apply global average pooling to the denoised image y'0 to obtain the semantic embedding z'0; m Connected with z'0, the masked text reconstructor R m Perform joint decoding to reconstruct the mask attributes and obtain the reconstructed text hint;
[0066] Minimize the difference between the reconstructed text prompt and the true text prompt during training.
[0067] The accurate semantics captured in y'0 can effectively guide the retrieval of masked attributes through their corresponding global semantic embedding z'0. Therefore, by minimizing the difference between the reconstructed textual cues and the true textual cues during training, we can in turn promote semantic consistency in the reconstructed PET images, thereby reducing distortion. The masked reconstruction loss is then introduced to supervise the above process.
[0068] As an optimization solution of the above embodiment, the objective function of the medical image reconstruction model based on conditional diffusion is:
[0069] L totat =L adv +λ1L image +λ2L text ;
[0070] Among them, λ1, λ2 and λ3 are weighting coefficients, L adv is the countervailing diffusion loss, L image is the image estimation loss, L text is the mask reconstruction loss.
[0071] The image estimation loss implements an L1 loss to minimize the gap between the true SPET image and the EPET image while encouraging less blur.
[0072] The masked reconstruction loss is calculated as the ground-truth text hint P T and its reconstructed counterpart Rm(P m ,z’0).
[0073] The advantages of the present invention are explained below by experimental analysis:
[0074] Dataset: The proposed method is trained and evaluated on a publicly available UDPET dataset. This dataset selects 160 18F-FDGPET imaging subjects scanned by Siemens Biograph Vision Quadra, and uses 130, 10, and 20 for training, validation, and testing, respectively. LPET images are generated by subsampling SPET scans with a dose reduction factor of 100 to simulate acquisition at 1 / 100 of the standard dose. The 3D scan size of each brain is 128X128X128 and is cut into 2D slices of size 128X128. Three standard metrics including peak signal-to-noise ratio (PSNR), similarity index (SSIM), and normalized mean square error (NMSE) are used for performance evaluation. The 2D slices are restacked into a full 3D PET scan for evaluation.
[0075] Comparison with state-of-the-art methods:
[0076] The proposed method is compared with six major reconstruction methods, including regression-based methods: (1) Auto-Context; (2) TriDo-Former; GAN-based methods: (3) Ea-GAN, (4) AR-GAN, (5) CVT-GAN; and diffusion-based methods: (6) CDM, on the UDPET dataset.
[0077] The comparison results are shown in Table 1. It can be seen that the present invention is superior to the compared methods in the three evaluation criteria with moderate parameters. In particular, compared with the current leading CDM method, the method proposed by the present invention still improves PSNR and SSIM by 0.687dB and 0.019db respectively. In addition, the method of the present invention only contains 30M parameters, while CDM has 34M parameters, which further verifies the speed and feasibility of the method of the present invention.
[0078] Table 1
[0079]
[0080] The visual comparison results of the method of the present invention and the comparative method are as follows: Figure 2 The areas where significant improvements are shown are highlighted with red arrows and circles in the magnified area. It can be observed that the present invention produces the best visual results with the finest details and the smallest reconstruction errors. Overall, these results show that the present invention outperforms the state-of-the-art methods.
[0081] To verify the effectiveness of the key components of the present invention, ablation studies were performed on the following variants:
[0082] (1) The inverse diffusion network conditioned only on the LPET image (i.e., baseline).
[0083] (2) Connect clinical tables to LPET images (i.e., baseline + text).
[0084] (3) Integrate LPET images and clinical tables (i.e., baseline+text+CA) using conventional attention (CA).
[0085] (4) Replace CA with the OMTA of the present invention (ie, baseline+text+OMTA).
[0086] (5) Inject M3TRec (i.e., baseline + text + OMTA + M3TRec, recommended).
[0087] The qualitative results are shown in Figure 3Compared with other ablation variants, it is obvious that the images reconstructed by this method are closest to the real SPET images and retain richer details, especially in the areas highlighted by the red boxes.
[0088] The quantitative results are shown in Table 2. It can be observed that the performance of the proposed model gradually improves with the introduction of each component. In particular, the introduction of the clinical table significantly improves the PSNR by 0.605dB, which shows its potential in providing important complementary information to improve the reconstruction quality. In addition, compared with the traditional CA, the proposed OMTA obtains better performance, demonstrating its ability to align and model the interactions between heterogeneous images and tabular data. In addition, the addition of the proposed model greatly improves the performance of the model, confirming its effectiveness in preserving semantic information.
[0089] Table 2
[0090]
[0091] The above shows and describes the basic principles and main features of the present invention and the advantages of the present invention. It should be understood by those skilled in the art that the present invention is not limited to the above embodiments. The above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention to be protected. The scope of protection of the present invention is defined by the attached claims and their equivalents.
Claims
1. A multimodal medical image reconstruction method based on a counter-diffusion model, characterized in that: Includes steps: S10, obtaining LPET images and clinical form sample data; S20, establish a medical image reconstruction model based on conditional diffusion, including an adversarial diffusion network, a multimodal conditional encoder and a multimodal masked text reconstruction module; During the training process, the adversarial diffusion network adopts the conditional diffusion framework, combining the conditional diffusion generator and the diffusion discriminator to diffuse on the real SPET images. The multimodal conditional encoder extracts information from the LPET image and clinical form sample data input and injects it into the conditional diffusion generator in a hierarchical manner; at the same time, the clinical form samples are also input into the multimodal masked text reconstruction module for reconstruction; S30, given the LPET image and clinical table detection data, input them into a medical image reconstruction model based on conditional diffusion, and use an adversarial diffusion network and a multimodal conditional encoder to iteratively convert the noise into an estimated PET image; The multimodal mask text reconstruction module, the processing process includes the steps of: Given clinical table test data x tb , encode the sentences based on the CXR-BERT module and construct a real text prompt P T ; Extract clinical table attributes in sentences, mask the attributes, and encode the masked sentences into masked text prompts P based on the CXR-BERT module m middle; Apply global average pooling to the denoised image y'0 to obtain the semantic embedding z'0; m Connected with z'0, the masked text reconstructor R m Perform joint decoding to reconstruct the mask attributes and obtain the reconstructed text hint; Minimize the difference between the reconstructed text prompt and the true text prompt during training.
2. The method for multimodal medical image reconstruction based on a diffusion-resistant model according to claim 1, characterized in that: The adversarial diffusion network includes a forward diffusion process and a reverse diffusion process; a 4-level diffusion generator and a 5-level diffusion discriminator are used to accelerate the reverse sampling process in an adversarial manner.
3. The method for multimodal medical image reconstruction based on the anti-diffusion model according to claim 2, characterized in that: The forward diffusion process: In t time steps, fixed Gaussian noise is gradually added to the SPET image y0 through the Markov chain to generate a series of noise images y1 to y t .
4. The method for multimodal medical image reconstruction based on a diffusion-resistant model according to claim 2, characterized in that: The reverse diffusion process: By gradually deriving the noise from y at t time steps t Restore to y0; at each time step t, provided with the current noise data y t and conditional information c, the neural network parameterized by θ predicts the denoised data y t-1 The conditional probability distribution of ; the expected prediction distribution p θ (y t-1 ly t ,c) will be close to its actual corresponding distribution q(y t-1 ly t ,c); Using GAN network, GAN network uses conditional diffusion generator G θ To predict p θ (y t-1 ly t ,c), and use the diffusion discriminator To minimize the predicted p θ (y t-1 ly t ,c) and the actual distribution q(y t-1 ly t ,c).
5. The method for multimodal medical image reconstruction based on the anti-diffusion model according to claim 4, characterized in that: In the GAN network, the conditional diffusion generator G is used θ Receive data pair (y t ,c) as input to predict the uncorrupted image y'0; then, the noise is reintroduced into y'0 using the posterior distribution, from p θ (y t-1 ly t ,c) Sample y' t-1 ; For adversarial training, ({y' t-1 or t-1 }, yt, t) aims to distinguish the estimated posterior distribution p θ (y t-1 ly t ,c) The sample extracted is used as the pseudo sample and the real corresponding sample q(y t-1 ly t ,c) as true samples.
6. The method for multimodal medical image reconstruction based on a diffusion-resistant model according to claim 1, characterized in that: The multimodal conditional encoder is used to extract multimodal features, which are then introduced into the conditional diffusion generator for guidance; in order to balance the heterogeneous modes of images and tables, the optimal multimodal transfer co-attention module is used to integrate table features and image features, and their correlation is inferred at the same time; The optimal multimodal transfer co-attention module exploits the matching ability of the optimal transfer to reconcile image-table differences, while leveraging the selective focusing ability of the attention mechanism to emphasize image regions in the table that are associated with more relative attributes; the optimal transfer is defined by the discrete Kantorovich formula.
7. The method for multimodal medical image reconstruction based on the anti-diffusion model according to claim 6, characterized in that: In the table branch of the multimodal conditional encoder, the clinical table detection data x tb is fed into the CXR-BERT model to generate the tabular features p tb ; In the image branch of the multimodal conditional encoder, corresponding to the encoding part, the image branch involves 4 levels. Given an LPET image x l , extract multi-level features x cnd .
8. The method for multimodal medical image reconstruction based on the anti-diffusion model according to claim 7, characterized in that: After each level of the image branch, an optimal multimodal transfer co-attention module is integrated to enable active interaction between image and table features.
9. The method for multimodal medical image reconstruction based on a diffusion-resistant model according to claim 1, characterized in that: The objective function of the medical image reconstruction model based on conditional diffusion is: L total =L adv +λ1L image +λ2L text ; Where λ1 and λ2 are weighting coefficients, L adv is the countervailing diffusion loss, L image is the image estimation loss, L text is the mask reconstruction loss.
Citation Information
Patent Citations
Multi-modal image super-resolution reconstruction method based on structured knowledge distillation
CN117911246A
Method for improving CBCT image quality based on generative adversarial diffusion model
CN117911355A