Diffusion prior synthesis and optimization method for cross-modal medical image synthesis

Through the diffusion prior synthesis and optimization method, using the diffusion model and conditional diffusion model trained by the target modality image, the problem of dependence on the source modality in cross-modal medical image synthesis is solved, and high-fidelity and stable image generation is achieved, which is suitable for the clinical application of multimodal medical images.

CN120634951APending Publication Date: 2025-09-12WUHAN TEXTILE UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510540930.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing cross-modal medical image synthesis methods rely on source modality images, which limits their application in actual clinical scenarios with limited resources. They also have problems such as uncontrollable synthesis paths, large differences between samples, and insufficient detail fidelity.

Method used

A method based on diffusion prior synthesis and optimization is adopted. A deterministic generation path is constructed through probability flow ordinary differential equations. The diffusion model is trained using the target modal image. Combined with the conditional diffusion model and singular value decomposition, high-fidelity image synthesis and optimization without the need for source modal data is achieved.

Benefits of technology

It achieves stable, controllable, high-fidelity cross-modal image synthesis in the case of passive modal data, improves the stability of image generation and detail retention capabilities, and is suitable for clinical applications of multimodal medical images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120634951A_ABST
    Figure CN120634951A_ABST
Patent Text Reader

Abstract

The invention relates to a diffusion prior synthesis and optimization method for cross-modal medical image synthesis, which is used for solving the problems of strong dependence on source modal data, high acquisition cost and the like in the existing method. According to the method, source modal data is not needed, training is carried out only based on single target modal data, and a diffusion prior synthesis (DPS) module and a diffusion prior optimization (DPO) module are included. The DPS encodes the image to a potential space guided by general diffusion prior through a probability flow ordinary differential equation, and decodes the image into an initial image through a target diffusion model; the DPO optimizes an initial image through a linear inverse problem, recovers high-frequency details and corrects discrete errors, thereby improving image fidelity and stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing, and in particular to a diffusion prior synthesis and optimization method for cross-modality medical image synthesis. The method is suitable for generating high-fidelity target modality images in scenarios where source modality images are lacking, and has broad clinical application prospects. Background Art

[0002] Medical imaging plays a vital role in modern clinical diagnosis and treatment. Multimodal medical imaging, in particular, (such as CT, CBCT, PET, T1 / T2 / PD-MRI, etc.), significantly improves the diagnostic accuracy and scientificity of treatment planning by integrating structural or functional features from different modalities. For example, CT images offer the advantage of clear structures, while MRI performs better in soft tissue contrast. The complementary nature of these two is particularly important in preoperative tumor assessment.

[0003] However, despite the clinical value of multimodal image fusion, its large-scale application is still limited by many practical issues, including individual patient differences (such as incompatibility between pacemakers and MRI equipment), high image acquisition costs, and insufficient coverage of advanced imaging equipment in primary healthcare institutions. This leads to challenges such as delayed diagnosis and incomplete treatment plans when clinicians are faced with missing multimodal data, thus promoting the development of cross-modal image synthesis technology.

[0004] Existing cross-modal image synthesis methods mainly fall into two categories: supervised learning methods based on paired data, such as GANs, Transformers, or diffusion models. These rely on strictly aligned data pairs between the source and target modalities. While they can achieve high-quality synthesis results, they place extremely high demands on data acquisition conditions. Unsupervised methods based on unpaired data relax data alignment constraints by introducing cycle consistency or latent space mapping mechanisms, resulting in greater versatility. Representative methods include CycleGAN and its variants. However, both of these methods still rely on source modality images for training, limiting their application in real-world clinical scenarios where source data is missing, privacy is limited, and data acquisition is difficult.

[0005] In addition, diffusion models have performed well in image synthesis tasks in recent years, but their standard forms are mostly based on stochastic differential equations (SDE) to implement sampling, which brings problems such as large differences between samples and uncontrollable synthesis paths; at the same time, the numerical discrete errors in the sampling process will also affect the ability to restore key details in medical images.

[0006] Therefore, there is an urgent need for a method that can achieve cross-modal conversion without relying on source modality images. This method should have the characteristics of high stability, strong detail fidelity, and high computational efficiency to promote its application in real medical environments with limited resources. Summary of the Invention

[0007] The technical problem to be solved by the present invention is to provide a cross-modal medical image generation method based on diffusion prior synthesis and optimization (DPSO) in response to the above problems and requirements. This method does not require the use of source modality images in both the training and inference stages, but only relies on the diffusion prior model constructed by the target modality image to complete the synthesis and detail optimization of the target image.

[0008] In order to solve the above technical problems, the present invention adopts the following technical solutions:

[0009] A diffusion prior synthesis and optimization method for cross-modal medical image synthesis includes the following steps:

[0010] Step 1: Diffusion prior synthesis step: Input the source modal image into a universal domain diffusion model and perform forward solution on the probability flow ordinary differential equation to obtain the potential representation of the modal image. The universal diffusion model is a denoised diffusion probability model pre-trained on a natural image dataset.

[0011] Step 2: Image decoding step: Input the latent representation of the modal image obtained in the previous step into the target modal diffusion model, and generate the initial target modal image by solving the inverse PF ODE. The target modal diffusion model is a model obtained by pre-training the target modal image dataset.

[0012] Step 3, diffusion prior optimization step: taking the initial image as a degraded observation, constructing a linear inverse problem, using a conditional diffusion model to refine the image, restore high-frequency details and improve image fidelity, and outputting a synthesized image. The parameters of the conditional diffusion model adopt the parameters of the target modal diffusion model.

[0013] Furthermore, the universal domain diffusion model is a denoising diffusion probability model pre-trained on the ImageNet image dataset.

[0014] Furthermore, the probability flow ordinary differential equation is used to transform the diffusion process from stochastic modeling to deterministic modeling, thereby improving the stability and controllability of the synthesized image.

[0015] Furthermore, a variational inference method is used in the diffusion prior optimization step to optimize the consistency between the synthesized image and the target image by maximizing the lower bound of evidence.

[0016] Furthermore, the linear inverse problem is analyzed in the frequency domain using singular value decomposition to improve the ability to reconstruct image details.

[0017] Furthermore, the step 2 specifically includes the following steps:

[0018] Step 2.1, from the image prior probability p θ(V) and the observed conditional probability p(Y|V) together form the posterior distribution p θ (V|Y);

[0019] Step 2.2, conditional diffusion model for posterior distribution p θ (V|Y) is modeled and approximated, and the conditional diffusion model takes the observed image Y as a condition and starts from Gaussian noise in the latent space. First, a Markov sampling path for stepwise denoising is constructed to eventually restore the target image V0;

[0020] Furthermore, in the reverse sampling process of the conditional diffusion model, the system uses a neural network to perform conditional denoising at each time step t to gradually restore the real image. Specifically, the model starts from the current perturbation state V t+1 Starting from the observation image Y and the time code t, the intermediate state V after denoising is predicted t .

[0021] After adopting the above technical solution, the present invention has the following advantages compared with the prior art:

[0022] (1) Achieve completely passive modality data training and improve the deployability of the method in actual clinical settings;

[0023] (2) Using PF ODE to model the deterministic diffusion process and improve the stability and controllability of the synthetic image;

[0024] (3) Combining the conditional diffusion model with the frequency domain reconstruction method, the detail retention and structural consistency of the synthesized image can be effectively improved.

[0025] The present invention is described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 This is an overall process framework diagram of the method of the present invention, including a diffusion priori synthesis module and a diffusion priori optimization module. DETAILED DESCRIPTION

[0027] The principles and features of the present invention are described below with reference to the accompanying drawings. The examples given are only used to explain the present invention and are not used to limit the scope of the present invention.

[0028] like Figure 1 As shown in the figure, this paper proposes a cross-modal medical image generation method based on Diffusion Prior Synthesis and Optimization (DPSO). This method does not require the use of source modality images for model training. Instead, it uses only the diffusion model trained on the target modality images to achieve high-fidelity cross-modal synthesis from the source modality images to the target modality images.

[0029] Specifically, the present invention includes two key modules that cooperate with each other: a Diffusion Prior Synthesis module (DPS) and a Diffusion Prior Optimization module (DPO), which respectively complete preliminary image synthesis and subsequent image refinement optimization.

[0030] (1) Diffusion Prior Synthesis (DPS) stage

[0031] The diffusion prior synthesis stage aims to use the general diffusion model and the target-specific diffusion model to complete the preliminary cross-modal synthesis process from the source modality image to the target modality image.

[0032] The general diffusion model used in this invention / project is the 256×256 diffusion model (non class-conditional), and its pre-trained weight file is 256×256_diffusion_uncond.pt, which was released by OpenAI in its Guided Diffusion project and is denoted as Corresponding to Figure 1The diffusion prior synthesizes the "Diffusion Model Trained on the ImageNet Dataset" in purple. This model is an unconditional diffusion model for image generation, designed for 256×256 resolution images. It is often used to evaluate the basic generative capabilities of diffusion models or as a general backbone for various conditional generative systems. The model is built on the Denoising Diffusion Probabilistic Models (DDPM) framework. Its core mechanism is a reversible "noising-denoising" generative process: in the forward pass, Gaussian noise is gradually added to the real image, eventually degenerating into a nearly isotropic pure noise image. In the backward pass, a neural network is trained to gradually denoise and reconstruct the image, achieving the goal of generating natural-looking images from random noise. The model backbone uses a symmetrical U-Net structure, consisting of four deep encoder and decoder layers (with the number of channels increasing from 128 to 256 to 512 to 1024). A multi-head self-attention mechanism is introduced at multiple resolution scales to enhance long-range modeling capabilities. Group Normalization, Swish activation function, positional encoding, and up / downsampling residual blocks are used to achieve efficient multi-scale feature learning. During training, a 1000-step linear β schedule (increasing linearly from 1e-4 to 0.02) was used as the noise schedule for the diffusion process. The goal was to predict the additive noise ε at each step, and the mean squared error (MSE) was used as the loss function. The optimizer was Adam (β1 = 0.9, β2 = 0.999), the initial learning rate was set to 2e-4, the batch size was 128, and the total number of training steps was approximately 1 million. Training was performed on ImageNet or a high-quality subset of OpenAI's internal images, with images uniformly scaled to 256×256 resolution.

[0033] In this paper, a target-specific diffusion model is further constructed, which keeps the network structure and training process of the general model unchanged and only performs retraining (fine-tuning) on ​​the new target modality dataset, denoted as Corresponding to Figure 1The yellow area above the diffusion prior synthesis module in the figure shows the "Diffusion model trained on the target dataset." The model still uses a U-Net-based denoising architecture, integrating multi-resolution residual blocks, a multi-head self-attention module, Group Normalization, and the Swish activation function to ensure the synergistic expressiveness of structure preservation and semantic modeling. The training configuration remains the same as the original design: 1000 diffusion steps (--diffusion_steps 1000), linear noise scheduling (--noise_schedule linear), MSE prediction with an ε target, and the Adam optimizer. Specific training parameters are: batch size = 4, learning rate = 1e-4, no learning rate decay (--lr_anneal_steps 0), EMA decay rate set to 0.9999, mixed precision training enabled (--use_fp16 = True, --fp16_scale_growth = 1e-3), and a micro-batch size of 1. The model uses an input image size of 256 and 256 channels. An attention mechanism is introduced at the 32-, 16-, and 8-resolution layers (--attention_resolutions 32,16,8). Each layer contains two residual blocks (--num_res_blocks 2) and uses 64 attention channel heads (--num_head_channels64). The `--resblock_updown` and `--use_scale_shift_norm` flags are enabled to enhance training stability and generation quality. Furthermore, `--learn_sigma=True` is set to simultaneously predict the image mean and variance, further improving the fidelity and diversity of image detail.

[0034] Step 1: Source modality image encoding

[0035] 1. Technical core and theoretical basis of source modality image compression conversion

[0036] The technical core of this step is to map the input source modality image into the latent state space by constructing a deterministic path that evolves along time (corresponding to the probability flow ordinary differential equation trajectory in the diffusion model). This mapping not only realizes the structured compression representation of complex image information in the high-dimensional latent space, but also constructs a continuous generation trajectory with path stability and reversibility, thereby providing solid structural prior support for the subsequent reconstruction of the target modality image. Compared with traditional image generation methods (such as variational autoencoders (VAE), autoencoders (AE), or generative adversarial networks (GAN), these methods usually compress the image into a static low-dimensional latent vector and have three core technical bottlenecks: first, the latent variables lack clear structural semantics, making it difficult to effectively express the geometric contours or anatomical structures of the image; second, the lack of temporal continuity and controllability of the generation path in the latent space leads to an unstable generation process; third, it is difficult to establish a unified and consistent expression mapping between different modalities, which limits the effect of cross-modal image synthesis. To overcome the above problems, the present invention proposes a latent state modeling method based on a diffusion model. By constructing a latent state trajectory that continuously evolves over time, it achieves high-fidelity retention in structure and has determinism and reversibility in the path, thereby constructing a unified intermediate representation space between different modalities, significantly enhancing the stability, structural consistency and expressiveness of cross-modal image synthesis.

[0037] 2. Definition of potential states in diffusion trajectories

[0038] In the present invention, the “latent state” refers to: the input image X (s) On the continuous time path defined by the diffusion model, the final state X is reached by smoothly evolving from the starting time point (t=0) to the end point (t=1). (l) . Its pixel distribution has been significantly disturbed by Gaussian noise, but it still fully retains the core structural features of the original image, especially in key visual elements such as edge contours, shape structures and anatomical relationships, and still has good recognizability. This latent state plays a key bridging role in the entire diffusion prior synthesis path. Its technical value is mainly reflected in the following aspects: First, it can effectively retain key visual clues such as the edges, contours and anatomical structures of the image, providing a solid semantic foundation for subsequent image reconstruction; second, as the end point of the forward diffusion path, this state has good reversibility and path continuity, and can serve as the natural starting point of the target modality decoding process, effectively reducing the reconstruction ambiguity and uncertainty caused by information loss. By introducing this latent representation, the system not only realizes the transfer of structural consistency between modalities, but also significantly improves the generation controllability and image fidelity in the cross-modal image synthesis process.

[0039] 3. Modeling of Continuously Generated Paths: Probabilistic Flow Ordinary Differential Equations (PF ODEs)

[0040] In order to realize the construction of the aforementioned latent state trajectory, the present invention introduces the Probability Flow Ordinary Differential Equation (PF ODE) as the core theoretical basis for image generation path modeling. This equation is derived from the mathematical principles of the diffusion model and provides a deterministic modeling method for characterizing image perturbations and evolution processes in the time continuous domain. Through this method, a smooth, stable and controllable latent trajectory can be constructed in the high-dimensional image space to ensure that the generation process has path uniqueness and dynamic interpretability, thereby providing solid modeling support for subsequent reversible decoding and cross-modal mapping tasks.

[0041] Compared with the traditional stochastic diffusion equation (SDE), PF ODE completely eliminates the random noise term in the system, and therefore has a series of outstanding technical advantages: first, its generation path is unique under the same input and is not affected by changes in the random seed, which greatly improves the controllability and determinism of the image generation process; second, the path is more stable during the numerical solution process and can effectively avoid structural damage caused by noise disturbances, which is particularly suitable for medical image processing tasks with high requirements for image fidelity; finally, while retaining the powerful distribution modeling capabilities of the diffusion model, PF ODE improves the interpretability and traceability of the path, laying a solid foundation for multimodal structural consistency.

[0042]

[0043] Where S(t) represents the state of the image at time t, f(S,t) is the basic dynamic term of the system, which is usually set to zero or a constant to simplify path modeling; g(t) is the noise amplitude function that varies with time and is used to control the disturbance intensity. Represents the score function, that is, the gradient direction of the logarithmic probability density under the current image state, specifically the neural network s θ (X,t) is estimated.

[0044] The differential equation is numerically integrated using DDIM's ordinary differential equation solver (ODE Solve), and its expression is as follows:

[0045]

[0046] Among them, X (s) Represents the input source modality image, X (l) is the potential representation obtained after encoding by the probability flow ordinary differential equation (PF ODE), are the diffusion model parameters pre-trained in the general domain. The variable t0 = 0 represents the initial time step of the image state, corresponding to the original clear image; while t1 = 1 represents the final time step of the state evolution, corresponding to the fully perturbed latent state. ODESolve represents a numerical solver for ordinary differential equations, used to continuously compute the image state over this time interval, generating a deterministic state evolution trajectory.

[0047] This encoding path realizes a continuous transition process from the clear image state (t=0) to the perturbation state (t=1), thereby constructing a smooth, structurally consistent deterministic trajectory in the latent space, providing a stable theoretical basis and structural guarantee for subsequent image reversible decoding and cross-modal mapping tasks.

[0048] 4. Latent State Image Features and Technical Value

[0049] The generated potential state X (l) Still in the image space, it is not an abstract low-dimensional vector, but a structure-preserving and perturbation-enhanced image representation with several significant features. First, in terms of distribution standardization, the pixel values ​​in this state tend to be standard Gaussian, which helps to achieve unified mapping between modalities in subsequent processing; second, in terms of structural integrity, despite the introduction of perturbation components, its key edge structure and tissue texture information are still preserved, ensuring the identifiability of basic semantics; third, in terms of reversible decoding capability, this state can serve as an effective starting point for diffusion decoding of the target modality, supporting high-quality reverse reconstruction of the image; finally, in terms of modality unification intermediary function, X (l) It can serve as a bridge state between different modalities to achieve consistency control and path controllability in the process of style transfer and semantic alignment.

[0050] In summary, this step successfully achieves the structural transformation from static images to continuous latent trajectories by introducing the latent trajectory modeling mechanism established by PF ODE, providing a solid theoretical support and implementation foundation for cross-modal image synthesis and structural consistency restoration.

[0051] Step 2: Latent Space Decoding

[0052] After completing the forward encoding of the source modality image through the universal diffusion model, the system obtains the potential state X in the middle perturbation stage of the diffusion trajectory (l)Next, in order to achieve the reconstruction of the target modality image, the system needs to map the potential state to the target modality image space and complete the decoding conversion from the potential trajectory to the target distribution. This process is essentially a reverse diffusion process. Its modeling method maintains structural symmetry with the compression stage, but performs exclusive modeling for the characteristic distribution of the target modality. In this step, the present invention introduces the probability flow ordinary differential equation (PF ODE) constructed based on the target modality diffusion model, and transforms the potential state X (l) As the initial condition, reverse integration is performed along the time dimension from the end point t = 1 to the starting point t = 0 to generate a high-fidelity image X that conforms to the target modal distribution. (t) .

[0053] To achieve the above reverse evolution path, the present invention first constructs the diffusion vector field of the target mode The vector field is defined by the logarithmic gradient (i.e., score function) of the probability density function of the target mode at each time point. Its mathematical form is as follows:

[0054]

[0055] Among them, f(X,t) represents the basic dynamic term of the system, which is usually set to zero to simplify the modeling; g(t) is the noise intensity function that changes with time and controls the disturbance amplitude at each moment; and is the probability distribution gradient of the target modality image at time t, which is used to guide the direction of image evolution. This score function is not given explicitly, but is a neural network trained specifically for the target modality. The network training phase is based on the score matching principle. By perturbing a large number of target modal samples and learning their diffusion reverse trajectory, the optimal probability gradient direction at each time step is approximated.

[0056] In the actual execution process, the system calls DDIM's ordinary differential equation solver (ODE Solve) to perform numerical integration on the differential equation. Its mathematical expression is as follows:

[0057]

[0058] Among them, X (t) represents the final generated target modality image, X (l) is the potential representation of the input, The variable t0 = 1 represents the start time of the solution process (corresponding to the latent space state), and t1 = 0 represents the end time of the solution process (corresponding to the clear image output state).

[0059] The deterministic PF ODE decoding mechanism adopted in the present invention effectively avoids the uncertainty problem in the traditional random diffusion process and significantly improves the stability and structural fidelity of image generation.

[0060] In the specific implementation, the solver starts at time t=1 with the potential representation X (l) As the initial input, it gradually advances to t = 0 under the control of time step dt, and the final state is X (t) The target modality image is closely distributed with the statistical characteristics of the target modality training set, and at the structural level, it restores the edge contours, anatomical structure, and grayscale texture information of the real image as much as possible.

[0061] The technical advantage of the decoding process lies not only in its path continuity and controllability, but also in its symmetric mapping mechanism with the forward compression path. By embedding structurally consistent intermediate latent states within the encoding-decoding chain, this method ensures structural consistency and semantic alignment between modalities. This approach is particularly well-suited for common modality conversion tasks in medical images (e.g., T1→PD, T1→T2, etc.), significantly reducing distortion issues caused by inconsistent modal features.

[0062] In summary, the target modality PF ODE decoding module achieves a high-fidelity mapping from the latent space to the target modality image by accurately modeling the inverse diffusion trajectory of the target modality, combining high-precision numerical integration methods with a neural network-driven scoring function fitting mechanism. This method not only significantly improves the overall performance of the cross-modality image synthesis system in terms of image reconstruction quality, trajectory stability, and clinical adaptability, but also enables high-quality image reconstruction based solely on target modality data, providing solid technical support for downstream tasks such as medical image generation, synthesis, and enhancement.

[0063] (2) Diffusion Prior Optimization (DPO) stage

[0064] The image optimization reconstruction phase (DPO) proposed in this paper aims to optimize the target modality image X generated by the diffusion prior synthesis module (DPS). (t) Perform high-precision reconstruction and detail compensation. Although the system has used the general diffusion model in the DPS stage Target Modal Diffusion Model Implemented from source image X (s) To potential state X (l) Then to image state X (t) However, due to the integral calculation of PF ODE in discrete time steps, the approximation error of the score function, etc., the generated target image X (t) Frequently, problems such as blurred edges, missing textures, and insufficient high-frequency content occur. Therefore, in order to further improve the DPS synthesized image X(t) The present invention models the process as a linear inverse problem of image restoration and uses the conditional diffusion model (see Figure 1 The green part in the diffusion prior optimization is used for joint optimization.

[0065] 1) Linear inverse problem modeling

[0066] In this stage, the image X synthesized by DPS is first (t) Considered as a potential degradation process from the ideal image The observation image generated in The process can be described by the following formula:

[0067] Y=HV+Z (5)

[0068] in, The degradation matrix of the image usually contains information such as blur kernel, downsampling operator, occlusion mask or compression mapping matrix; is additive Gaussian noise, which obeys the distribution The above expression is a classic linear degradation model, which is widely used in image restoration, super-resolution, image inpainting and other tasks.

[0069] In practical applications, Y corresponds to the initial image result generated by DPS, and V is the high-quality target modality image we hope to restore (the final output of the diffusion prior optimization). Since H is usually irreversible and has rank deficiency (such as occlusion, downsampling, etc.), the inverse problem is ill-posed and must introduce a strong prior to be effectively solved.

[0070] 2) Bayesian modeling and conditional diffusion reconstruction

[0071] To solve the above inverse problem, the present invention models image restoration as a posterior probability estimation problem based on the Bayesian inference framework, aiming to maximize the posterior distribution p θ (V|Y), infer the original image V that is most likely to correspond to the observed image Y. The posterior is determined by the image prior probability p θ (V) and the observation conditional probability p(Y|V) together form:

[0072] p θ (V∣Y)∝p θ (V)·p(Y∣V) (6)

[0073] Among them, p θ(V) is the prior distribution of the image, which is used to characterize the structure and distribution law of the image in the absence of observation information, and is usually modeled by a diffusion model; and p(Y|V) is the observation likelihood term, which is used to characterize the conditional probability distribution of the observed image Y generated after the degradation matrix H and noise perturbation given the image V, corresponding to the statistical modeling form of the degradation model Y=HV+Z.

[0074] Since the degradation matrix H is usually irreversible, and degradation operations such as downsampling, occlusion, and blurring will cause partial loss of original image information, the posterior distribution p θ (V|Y) is difficult to express explicitly and solve directly in high-dimensional space. Therefore, it is necessary to introduce a generative model with strong modeling capabilities to indirectly approximate the posterior distribution p from the perspective of sampling. θ (V|Y), thereby achieving high-quality image reconstruction and optimized restoration.

[0075] To this end, the present invention designs a conditional diffusion model to be used for the posterior distribution p θ (V|Y) is modeled and approximated. The model takes the observed image Y as a condition and generates Gaussian noise in the latent space. First, we construct a Markov sampling path for step-by-step denoising and finally restore the target image V0:

[0076]

[0077] In the reverse sampling process of the conditional diffusion model, the system uses a neural network to perform conditional denoising at each time step t to gradually restore the real image. Specifically, the model starts from the current perturbation state V t+1 Starting from the observation image Y and the time code t, the intermediate state V after denoising is predicted t , which is updated as follows:

[0078]

[0079] Among them, μ θ Represents the mean term of the neural network output, integrating the current noise state V t+1 , the conditional information of the observed image Y and time step t; σ t is the sampling standard deviation corresponding to the current time step, which is used to adjust the amplitude of random disturbance.

[0080] To train the model and ensure it learns the correct conditional denoising strategy at each time step, we introduce a variational inference mechanism and construct a forward denoising process in the opposite direction of the sampling path to generate training data. This "dual path" starts from the real image V0 and gradually adds Gaussian noise to generate intermediate states at each time step:

[0081]

[0082] This variational distribution is used to simulate the gradual degradation of an image from a clear state to Gaussian noise, providing a supervisory signal for each time step during training. By modeling the sampling process as a posterior distribution and the forward noise as an inference approximation distribution, the system further constructs an evidence lower bound (ELBO) as an optimization objective to approximate the true posterior of the conditional diffusion model:

[0083]

[0084] The above loss function consists of two parts: the first is the KL divergence loss at each time step, which encourages the sampling distribution of the model output to be close to the true posterior transfer distribution; the second is the final reconstruction error, which ensures that the image finally restored from the noise is consistent with the original image in structure and semantics.

[0085] This variational framework provides a theoretically rigorous training path for the conditional diffusion model, enabling the model to converge stably under high-dimensional conditions and achieve consistent alignment of image quality with conditional constraints at multiple time scales.

[0086] 3) Singular Value Decomposition Enhancement Module (SVD Regularization)

[0087] In order to improve the structural stability and high-frequency detail preservation ability in the image reconstruction process, the present invention introduces a regularization enhancement mechanism based on singular value decomposition (SVD) to achieve effective separation of image signals and degradation noise in the frequency domain.

[0088] The design motivation of this mechanism is that during the image degradation process, due to the existence of blur, occlusion, downsampling and other operations, the degradation matrix It usually exhibits a rank-deficient structure or low-pass characteristics, which causes the high-frequency details in the image (such as edges and textures) to be compressed into a set of directions that are easily dominated by noise during the degenerate transformation, making them difficult to recover and becoming a key factor limiting the reconstruction quality.

[0089] To solve this problem, the system first performs singular value decomposition on the degenerate matrix H, which is mathematically expressed as follows:

[0090] H=U∑P T (11)

[0091] Here, U and P are left and right orthogonal matrices, representing the orthogonal transformation basis of the degradation matrix in the signal space and observation space, respectively. ∑ is a diagonal singular value matrix, where each singular value represents the information retention capacity of the corresponding direction in the degradation mapping. In practice, this singular value spectrum often exhibits rapid decay: larger singular values ​​correspond to the principal component structure of the image, while smaller singular values ​​often correspond to high-frequency texture, detailed edges, or even directions dominated by noise.

[0092] Based on the above observations, the present invention introduces a spectrum enhancement strategy in the singular value domain during the reconstruction phase, including:

[0093] Preserve the principal component channels: For channels corresponding to large singular values, their reconstruction contributions are fully retained to ensure the restoration effect of key structural areas; Suppress weak channel noise: For channels with small singular values, introduce spectral domain filtering, soft thresholding or reweighting mechanisms to weaken the misleading recovery caused by degradation instability; Impose subspace constraints: The system can dynamically set an upper limit on the dimension of the reconstruction subspace, and only allow structural information to be reconstructed within this dimension, thereby improving the overall robustness and reconstruction consistency.

[0094] This enhancement mechanism can be used independently for SVD-guided recovery and can also be integrated into the diffusion model's stepwise sampling process as an external regularization term. During each denoising step, the system performs singular value domain transformation and filtering on the current image state, forming a synergistic "diffusion modeling + frequency domain manipulation" mechanism. In complex scenes such as medical images, this module excels in improving edge sharpness, detail consistency, and anatomical fidelity.

[0095] Compared with traditional Tikhonov regularization terms or frequency domain filters, this method has the following technical advantages: strong structural interpretability: the intervention is based on the singular spectrum expansion of the degradation operator and has a clear geometric / frequency physical meaning; high compatibility: there is no need to modify the diffusion model structure, it can be plugged into the sampling process or added to the training path as a loss term; good adaptability: in different task scenarios, the singular value reweighting, dimensionality restriction, and gating adjustment mechanism can all be adaptively configured, which is convenient for deployment and generalization.

[0096] In summary, the SVD regularization enhancement module provides stability compensation and structural alignment capabilities in the frequency domain for the image optimization reconstruction path proposed in this invention, further expanding the adaptability and practical value of the diffusion model in dealing with non-uniform degradation and multimodal image restoration.

[0097] 4) Pre-training model reuse and parameter sharing mechanism

[0098] In order to improve the practicality and deployment efficiency of the image optimization reconstruction stage (DPO), the present invention proposes a pre-trained model reuse mechanism based on simplified optimization objectives. Under the premise of meeting specific mathematical conditions, the image optimization objectives of the DPO stage can be equivalently converted into the image sampling process in the standard diffusion model (such as DDPM or DDIM), which makes it unnecessary to retrain the structure or parameters of the diffusion model in any form. By reusing the weights of the "diffusion model trained on the target modality dataset", this mechanism can efficiently achieve the optimized generation of the target image Y, greatly reducing the deployment cost and training complexity of the model, while improving the adaptability and engineering usability of the system.

[0099] A key advantage of this process is that it requires no modifications to the existing diffusion model structure, nor does it rely on task-specific fine-tuning or retraining. This significantly reduces development and deployment cycles and reduces computing resource consumption. Compared to traditional conditional diffusion methods, the reuse mechanism proposed in this paper retains the powerful generative capabilities of the original diffusion model while enabling rapid adaptation to a variety of tasks, such as image restoration, structure completion, and deblurring. Through this mechanism, the system achieves a high degree of versatility, computational efficiency, and engineering scalability, enabling it to flexibly adapt to image optimization needs in a variety of real-world application scenarios.

[0100] Specifically, this mechanism supports rapid switching between different tasks and modalities without having to train each specialized model from scratch. By leveraging pre-trained model parameters, the system can directly optimize images without having to relearn the underlying generation model, greatly improving the efficiency of task execution. This "plug-and-play" model is ideal for multi-modal, multi-task image generation and optimization applications, particularly in areas such as medical image processing and cross-modal image synthesis, and can cope with the wide differences and complex requirements between image types and modalities.

[0101] Furthermore, the pre-trained model reuse mechanism significantly enhances the system's flexibility and scalability, enabling rapid deployment and adaptation to diverse application scenarios while reducing the difficulty of migration between different devices and platforms. This mechanism ensures efficient and stable operation, whether on low-resource edge devices or in high-performance computing environments, and has broad engineering application value.

[0102] 5) Technical value and clinical application significance

[0103] The image optimization reconstruction stage (Diffusion Prior Optimization, DPO) proposed in this invention is a key component of the DPSO framework, which aims to perform high-precision reconstruction and detail compensation on the initial target modality image generated by the diffusion prior synthesis stage. In order to solve the problems such as edge blurring and high-frequency texture loss that may be caused by the discrete integral error of PF ODE and the approximation error of the scoring function, DPO models the optimization process as a typical linear inverse problem, introduces the degradation matrix and additive Gaussian noise to construct the observation model, and uses the conditional diffusion model to sample and approximate the posterior distribution based on the Bayesian inference framework. The system uses U-Net as the core backbone network, introduces conditional normalization, attention mechanism and multi-scale fusion in the diffusion modeling process to achieve high-fidelity reconstruction under the guidance of conditional images. At the same time, DPO integrates the singular value decomposition enhancement module to perform spectral separation and subspace filtering on the degradation matrix, effectively strengthening the structural edges and improving the texture retention ability. At the engineering deployment level, this paper further proposes an optimization target simplification mechanism, enabling the DPO stage to directly reuse the unconditional diffusion model weights pre-trained for the target modality, provided that specific mathematical conditions are met, without requiring structural modifications or fine-tuning, significantly reducing computational resource consumption and model deployment costs. Overall, the DPO stage possesses a solid theoretical foundation, a flexible network structure, and excellent practical adaptability. It can be widely applied to medical image optimization needs in a variety of real-world scenarios, including image restoration, completion, and deblurring, promoting the development and application of portable, scalable, and plug-and-play generative medical imaging systems.

[0104] The present invention proposes a cross-modal image generation and optimization reconstruction framework based on a diffusion model, named DPSO (Diffusion Prior Synthesis and Optimization), which consists of two parts: the "Diffusion Prior Synthesis (DPS) stage" and the "Diffusion Prior Optimization (DPO) stage". It systematically solves key problems such as structural distortion, loss of details and high optimization cost in the cross-modal image synthesis process. In the DPS stage, the present invention uses a general diffusion model and a target modality-specific diffusion model to construct a structurally continuous, reversible and controllable latent trajectory through the probability flow ordinary differential equation (PF ODE), compresses and maps the source modality image to the perturbed latent state, and performs high-fidelity decoding based on the diffusion vector field of the target modality to complete the preliminary image synthesis from the source modality to the target modality. This stage emphasizes the structural fidelity and path symmetry of the latent state, laying a structural foundation for subsequent optimization. In the DPO stage, the present invention further models the image optimization process as a linear inverse problem, combines Bayesian posterior modeling with conditional diffusion model construction, and accurately repairs problems such as edge blur and high-frequency texture loss in the preliminary synthetic image; at the same time, the singular value decomposition enhancement module is introduced to achieve spectral selective retention, effectively improving the image reconstruction quality and structural clarity. In order to improve deployment efficiency, the DPO stage also supports the reuse of pre-trained diffusion model weights under specific conditions, without the need for structural modification and task fine-tuning, and has high engineering usability and system scalability. In summary, the DPSO framework realizes a closed-loop structure from potential trajectory generation to inverse problem optimization in theoretical modeling, and completes fine control from cross-modal structure alignment to high-fidelity image reconstruction in the technical path. It can be widely used in multiple practical application scenarios such as medical image generation, modality conversion, image enhancement and diagnostic assistance, and has significant academic innovation value and clinical promotion prospects.

[0105] (3) Experimental verification and performance evaluation

[0106] Dataset selection and composition:

[0107] Evaluation is performed on two publicly available multimodal medical image datasets, the IXI dataset and the SynthRAD2023 dataset, to comprehensively verify the versatility and effectiveness of the proposed method under different imaging modalities and anatomical structures.

[0108] 1. IXI dataset:

[0109] This dataset is provided by three hospitals in London (Guy's Hospital, Hammersmith Hospital and Institute of Psychiatry). It contains MRI scans of about 600 healthy subjects, covering three common sequence types: T1-weighted (T1w), T2-weighted (T2w) and proton density-weighted (PD). To maintain consistency between images, we preprocess all images to a uniform resolution of 256×256 pixels. This dataset is suitable for cross-sequence image synthesis tasks (such as T1→

[0110] PD, T2→T1) provides a standard baseline with good modality alignment and structural consistency, which is suitable for evaluating the performance of the model in brain structure detail recovery and modality conversion capabilities.

[0111] 2. SynthRAD2023 dataset:

[0112] This dataset was released by the AAPM RT-MAC conference and is designed for synthetic image research. It contains highly structured synthetic MR, CT, and CBCT (cone beam CT) image pairs to simulate cross-modal imaging scenarios in real clinical settings. The dataset contains 540 brain image samples (including MR-CT and MR-CBCT pairings) and 540 pelvic image samples (including MR-CT pairings). Each sample contains complete modality pairing, unified registration and alignment, and image mask information, and has rich anatomical differences and intensity distribution, which greatly enhances the test depth of the model's generalization ability. To ensure data consistency, the present invention normalizes all images and resamples them to 256×256 resolution.

[0113] Experimental setup

[0114] To comprehensively evaluate the performance of the proposed DPSO framework, we compared it with four state-of-the-art image synthesis methods: ResViT, SynDiff, I2I-Mamba, and SelfRDB. These baseline methods represent the current state of the art in medical image synthesis, and detailed implementation details are provided in their original papers.

[0115] To ensure a fair comparison, we adhered to the configurations described in each method's original paper, using the same hyperparameter settings and training strategies as the baseline methods. Specifically, all methods were trained on the same computing environment, with the data split using a 70% training set, 15% validation set, and 15% test set ratio. This configuration ensures that all methods are tested on the same data distribution, eliminating potential evaluation bias due to differences in data partitioning. Furthermore, all experiments were conducted on the same hardware environment to ensure a fair performance comparison.

[0116] Evaluation Metrics

[0117] To comprehensively measure the model's performance in medical image synthesis tasks, we employed six mainstream evaluation metrics to quantitatively analyze the quality of generated images across multiple dimensions, including pixel-level accuracy, structural fidelity, perceptual similarity, and statistical distribution consistency. Specifically, the Peak Signal-to-Noise Ratio (PSNR) assesses pixel differences in image reconstruction, with higher values ​​indicating less distortion. The Structural Similarity Index (SSIM) measures the similarity of images in brightness, contrast, and structure, with values ​​closer to 1 indicating greater structural fidelity. The Root Mean Square Error (RMSE) reflects the overall deviation in pixel intensity, with lower values ​​indicating better results. The Perceptual Image Difference (LPIPS) assesses perceptual similarity based on deep features, more closely resembling human visual judgment, with lower values ​​indicating less discrepancy. The Fréchet Inception Distance (FID), which measures the statistical difference between generated and real images by measuring the distance in feature distribution, is a key metric for measuring image authenticity and diversity, with lower values ​​indicating better results. Finally, the Natural Image Quality Evaluation (NIQE) is a no-reference evaluation method that relies solely on the statistical characteristics of the image itself to determine naturalness, with lower values ​​indicating a more realistic and natural image. In subsequent experiments, we uniformly report the mean and standard deviation of the above six indicators to comprehensively reflect the performance of each method in image reconstruction accuracy, structure preservation ability and perceptual quality.

[0118] Synthesis performance analysis

[0119] To comprehensively evaluate the performance of the proposed DPSO framework in multimodal medical image synthesis tasks, we designed and executed eight cross-modal image synthesis tasks on two representative public datasets, IXI and SynthRAD2023. These tasks are:

[0120] · Six brain modality conversion tasks based on the IXI dataset: T1→PD, T2→PD, T2

[0121] →T1, PD→T1, T1→T2, and PD→T2

[0122] · Two cross-modal organ-level tasks based on the SynthRAD2023 dataset: MR→CT

[0123] (brain) and CBCT→CT (pelvis)

[0124] The above tasks cover a variety of scenarios from sequence-level conversion to modality-level conversion, verifying the model's adaptability and generalization performance in structure preservation, detail recovery, and modality transfer.

[0125] In each task, we systematically compare DPSO with four state-of-the-art representative image synthesis methods: ResViT, SynDiff, I2I-Mamba, and SelfRDB. All methods operate under consistent data partitioning and training settings, and the synthesized images are evaluated using six mainstream quantitative metrics (PSNR, SSIM, RMSE, LPIPS, FID, NIQE). These metrics comprehensively measure the generation performance of the model from multiple perspectives, including pixel-level accuracy, structural consistency, perceptual quality, and statistical distribution. The quantitative results of the experiments are summarized in Table 1, showing the average performance and standard deviation of each method across all tasks.

[0126] Table 1

[0127]

[0128] Progress in pixel-level reconstruction tasks:

[0129] Experimental results on various tasks (Table 1) show that DPSO consistently maintains competitive pixel-level accuracy (e.g., PSNR of 30.35±0.49 for T1→PD). Although slightly behind top models such as ResViT and SynDiff in terms of average PSNR and SSIM metrics, DPSO demonstrates significantly higher stability and significantly lower standard deviation across all tasks (e.g., in the T1→PD task, DPSO's PSNR standard deviation is 0.49, while ResViT's is 0.86). This stability is crucial in medical imaging scenarios, where consistency and reliability are particularly important across diverse patient scans.

[0130] Structural information and perceptual quality:

[0131] DPSO effectively maintains structural integrity and perceptual quality in simpler translation tasks such as T2→PD, achieving an SSIM of 0.87±0.09 and a LPIPS of 0.07±0.05. However, in more complex tasks such as CBCT→CT, performance degrades significantly, with SSIM dropping significantly to 0.49±0.06 and LPIPS increasing substantially to 0.35±0.07. This suggests that DPSO struggles to capture detailed structural relationships when there are significant differences between the source and target modalities. In contrast, competing models such as ResViT and SynDiff demonstrate higher adaptability, suggesting that DPSO's structural learning mechanism needs to be further enhanced in the future.

[0132] Image quality and distribution similarity:

[0133] DPSO's performance showed mixed results in terms of FID and NIQE. For example, in the T1→PD task, DPSO achieved an FID and NIQE score of 18.56 and 5.45, respectively, lagging behind the significantly superior performance of SynDiff (FID: 5.79; NIQE: 4.36). Interestingly, DPSO performed well in the CBCT→CT task, achieving a better FID score (77.60), surpassing ResViT (89.89) and SynDiff (83.82). However, its NIQE remained high (9.00), highlighting the gap between capturing the overall image distribution and the perceived quality of individual images, which is an issue worthy of further research in the future.

[0134] Ablation studies

[0135] To quantify the contribution of the proposed optimization module (DPO) in DPSO, we performed ablation comparisons between DPS (without the optimization module) and DPSO. The test results show that the overall performance of SSIM is improved by 3.2% and the FID is reduced by 10.8% (see Table 2). However, in the more structurally demanding cross-modal tasks (MR→CT, CBCT→CT), the module shows some limitations, with SSIM and LPIPS scores deteriorating. This shows that although the optimization module improves perceptual quality, its structure preservation ability in complex cross-modal translation still needs further improvement.

[0136] Table 2

[0137]

[0138] Summarize

[0139] The method of the present invention was experimentally validated on multiple public medical image datasets, covering multiple common modality conversion tasks from MRI to CT and between different MRI sequences. Experimental results show that the DPSO method of the present invention excels in pixel-level reconstruction performance and stability, especially in standard inter-sequence tasks. Its source-independent training method offers significant advantages, enabling stable and efficient generation of high-quality cross-modal medical images. It is particularly suitable for clinical applications in data-scarce and privacy-critical scenarios.

[0140] Those skilled in the art may make appropriate adjustments or modifications to the method of the present invention based on the above embodiments in combination with actual needs and specific application scenarios, and such adjustments or modifications should be deemed to be within the scope of protection of the present invention.

Claims

1. A diffusion prior synthesis and optimization method for cross-modal medical image synthesis, characterized in that: The following steps are involved: Step 1: Diffusion prior synthesis step: Input the source modal image into a universal domain diffusion model and solve the probability flow ordinary differential equation forward to obtain the potential representation of the modal image. The universal diffusion model is a denoised diffusion probability model pre-trained on a natural image dataset. Step 2: Image decoding step: Input the latent representation of the modal image obtained in the previous step into the target modal diffusion model, and generate the initial target modal image by solving the inverse PF ODE. The target modal diffusion model is a model obtained by pre-training the target modal image dataset. Step 3, diffusion prior optimization step: taking the initial image as a degraded observation, constructing a linear inverse problem, using a conditional diffusion model to refine the image, restore high-frequency details and improve image fidelity, and outputting a synthesized image. The parameters of the conditional diffusion model adopt the parameters of the target modal diffusion model.

2. The diffusion prior synthesis and optimization method for cross-modal medical image synthesis according to claim 1, characterized in that: The general domain diffusion model is a denoising diffusion probability model pre-trained on the ImageNet image dataset.

3. The diffusion prior synthesis and optimization method for cross-modal medical image synthesis according to claim 1, characterized in that: The probability flow ordinary differential equation is used to transform the diffusion process from stochastic modeling to deterministic modeling, thereby improving the stability and controllability of the synthesized image.

4. The diffusion prior synthesis and optimization method for cross-modal medical image synthesis according to claim 1, characterized in that: The variational inference method is used in the diffusion prior optimization step to optimize the consistency between the synthesized image and the target image by maximizing the lower bound of evidence.

5. The diffusion prior synthesis and optimization method for cross-modal medical image synthesis according to claim 1, characterized in that: The linear inverse problem is analyzed in the frequency domain using singular value decomposition to improve the detail reconstruction capability of the image.

6. The diffusion prior synthesis and optimization method for cross-modal medical image synthesis according to claim 1, characterized in that: The step 2 specifically includes the following steps: Step 2.1, from the image prior probability p θ (V) and the observed conditional probability p(Y|V) together form the posterior distribution p θ (V|Y); Step 2.2, conditional diffusion model for posterior distribution p θ (V|Y) is modeled and approximated, and the conditional diffusion model takes the observed image Y as a condition and starts from Gaussian noise in the latent space. Start by building a Markov sampling path for step-by-step denoising and finally restore the target image V 0。 7. The diffusion prior synthesis and optimization method for cross-modal medical image synthesis according to claim 6, characterized in that: In the reverse sampling process of the conditional diffusion model, the system uses a neural network to perform conditional denoising at each time step t to gradually restore the real image. Specifically, the model starts from the current perturbation state V t+1 Starting from the observation image Y and the time code t, the intermediate state V after denoising is predicted t .

Citation Information

Cited By

  • Image event multi-mode semantic segmentation method, device and equipment

    CN121010757A

  • Medical image cross-modal synthesis method, terminal equipment and storage medium

    CN121148619A