Image synthesis method and system based on contour guide diffusion model, and storage medium

By using a contour-guided accelerated diffusion model in medical image synthesis, combining non-uniform sampling and multi-frequency enhancement attention modules, the problems of anatomical structure consistency and noise distinction are solved, and an efficient image synthesis process is achieved.

CN120147152APending Publication Date: 2025-06-13CHINA WEST NORMAL UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510303200.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

During the post-contrast medical image synthesis process, it is difficult to maintain consistency of anatomical structures, distinguish between random noise and high-frequency characteristics, and there is a long training process.

Method used

A contour-guided accelerated diffusion model is adopted, and the non-uniform sampling strategy and multi-frequency enhancement attention module is combined with contour information as conditional constraints to guide the generation process, maintain anatomical consistency, and learn global key features through frequency channels.

Benefits of technology

Effectively solve the problems of anatomical deformation and noise interference, reducing training time and calculation costs, while retaining high-quality image details.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147152A_ABST
    Figure CN120147152A_ABST
Patent Text Reader

Abstract

The invention discloses an image synthesis method and system based on a contour-guided diffusion model, and a storage medium. The method comprises the following steps: acquiring a medical image before comparison; constructing and training a contour-guided accelerated diffusion model; the trained contour-guided accelerated diffusion model is utilized to generate a compared medical image, the contour-guided accelerated diffusion model comprises two processes, namely a forward diffusion process and a reverse diffusion process, the whole time step length of the contour-guided accelerated diffusion model is divided into a change area and a stable area, and the change area is divided into a forward diffusion process and a reverse diffusion process; the time sampling interval of the change area is smaller than the time sampling interval of the stationary area, in the denoising process of any time step in the reverse diffusion process, a guide condition is determined based on the contour of the medical image before comparison, denoising is carried out based on the guide condition, and therefore the medical image after comparison is synthesized. According to the method, the quality of the synthesized image can be improved, and the number of iterations and the calculation cost are effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical image processing, and particularly relates to an image synthesis method, system and storage medium based on a contour-guided diffusion model. Background Art

[0002] With the continuous progress of medical imaging, modern medical imaging technologies have been able to provide various modalities of medical images. Angiography has high clinical value. The post-contrast images collected after injecting a contrast agent can clearly show blood vessels and lesion areas, which are crucial for the imaging diagnosis of diseases, tumor localization and efficacy evaluation. For example, gadolinium-based contrast agents are widely used in magnetic resonance imaging (MRI) as a common contrast agent to enhance tissue contrast and better identify active lesions in multiple sclerosis and tumors. Iodine-based contrast agents utilize their X-ray impermeability to enhance the contrast between lesions and surrounding tissues in computed tomography (CT), which helps to detect lesions that may be overlooked by ordinary CT and provides the scope and nature of the lesions. However, injecting contrast agents poses potential risks to health and safety. The potential toxicity of the metal compounds contained in gadolinium-based contrast agents will remain in the human body and may cause diseases such as nephrogenic fibrosis that cannot be effectively treated. Iodine-based contrast agents may cause allergic reactions and nephrotoxicity in clinical applications. In addition, factors such as time-consuming procedures, high costs, and cumbersome contrast medium injections together bring unnecessary burdens to patients. Therefore, it is very necessary and meaningful to synthesize post-contrast images under the condition of not injecting contrast agents for clinical diagnosis.

[0003] With the progress of deep learning, deep generative models can be trained to synthesize post-contrast images without actually injecting contrast agents. Although existing methods based on Generative Adversarial Networks (GANs) have made some progress, the complex adversarial architecture design, loss function design for different tasks, and high sensitivity to hyperparameters may lead to mode collapse in GANs, affecting the performance and stability of the model. Compared with GANs, diffusion models can generate highly realistic images from noise, but there are still some limitations when applied to contrast-enhanced medical image synthesis. Specifically, due to the randomness of the sampling and inversion processes, the predicted noise may deviate from the standard Gaussian distribution, which increases the accumulation of latent codes and causes anatomical structure deformation. Then, high-frequency features, such as edges and textures, are particularly vulnerable to noise interference because when noise is introduced into the image, they usually appear as random high-frequency information, making it impossible to distinguish between noise and frequency features, thus resulting in the damage of edge and texture details. Moreover, to obtain high-quality generation results, a large number of iterations are required, which leads to a large amount of computation during the training and sampling processes. Summary of the Invention

[0004] The present invention provides a method for synthesizing pre- and post-contrast medical images based on contour-guided accelerated diffusion models, and the technical problem to be solved is: to achieve the consistency of medical image anatomical structures and effectively distinguish random noise and high-frequency features during the synthesis of post-contrast medical images, while solving the long training process.

[0005] One aspect of the present invention provides a method for synthesizing pre- and post-contrast medical images based on contour-guided accelerated diffusion models, including:

[0006] Obtain pre-contrast medical images

[0007] Construct and train a contour-guided accelerated diffusion model;

[0008] Use the trained contour-guided accelerated diffusion model to generate post-contrast medical images corresponding to the pre-contrast medical images wherein, the contour-guided accelerated diffusion model includes two processes, a forward diffusion process and a reverse diffusion process. The entire time step of the contour-guided accelerated diffusion model is divided into a changing region and a stable region, and the time sampling interval of the changing region is smaller than that of the stable region. During the denoising process at any time step in the reverse diffusion process, determine the guiding condition c based on the contour of the pre-contrast medical image, and thus perform denoising based on the guiding condition c to synthesize the post-contrast medical image.

[0009] ​​

[0010] Furthermore, the division of the changing region and the stable region is carried out according to a set time threshold, and the setting method of the time threshold is as follows:

[0011] In the forward diffusion process, determine the image x after adding noise at each time step t t ;

[0012] Calculate the process increment δ t = x t+1 - x t variance;

[0013] If the variance of the process increment obtained at a time step t is greater than the preset variance value, then take this time step t as the time threshold.

[0014] Furthermore, based on the contour of the pre - comparison medical image to determine the guiding condition c includes:

[0015] Use an edge detector to determine the contour representation of the pre - comparison medical image

[0016] Concatenate the pre - comparison medical image and the contour representation in the channel dimension to generate the guiding condition c.

[0017] Furthermore, the denoising process at any time step in the reverse diffusion process includes estimating the noise using a noise prediction model, where the noise estimation model includes an encoder, a multi - frequency attention module, and a decoder,

[0018] The encoder receives the noise image at the current time step t ∈ {0, 1,..., T} and the condition c as inputs, and after being encoded by the encoder, obtains the tensor X, where T represents the entire time step length of the diffusion process;

[0019] The multi - frequency attention module processes the tensor X in the spatial branch and the channel branch and then performs frequency - domain processing to generate the frequency - domain processing result A(X);

[0020] The decoder decodes A(X) to generate the predicted noise corresponding to the time step t;

[0021] By predicting the noise for all time steps and removing the noise, the post - comparison image is obtained.

[0022] Furthermore, the multi - frequency attention module processes the tensor X in the channel branch including:

[0023] Perform the first convolution and the second convolution on the tensor X; ​

[0024] Process the result of the second convolution through the Softmax function to generate a first processing result;

[0025] Perform matrix multiplication on the first processing result and the first convolution result, and sequentially process the multiplication result through a third convolution, layer normalization, and the Softmax function to generate a channel branch result.

[0026] Furthermore, the multi-frequency attention module processes the tensor X in the spatial branch, including:

[0027] Perform a fourth convolution and a fifth convolution on the tensor X;

[0028] Sequentially process the result of the fourth convolution through a global pooling operation and the Softmax function to generate a second processing result;

[0029] Perform matrix multiplication on the second processing result and the fifth convolution result, and process the multiplication result through the Softmax function again to generate a spatial branch result.

[0030] Furthermore, the method further includes:

[0031] After performing spatial branch and channel branch processing, perform dot product operations on the spatial branch result and the channel branch result with the tensor X respectively, and add the dot product results to obtain an enhanced feature map

[0032] The enhanced feature map Perform frequency domain processing, specifically including:

[0033] The enhanced feature map Divide it into two or more features along the channel dimension;

[0034] For each feature, perform a two-dimensional discrete cosine transform;

[0035] Perform a fully connected operation on the multi-frequency vectors composed of the corresponding two-dimensional discrete cosine transform results of each part and process them through the Softmax function to generate a third processing result;

[0036] Perform matrix multiplication on the third processing result and the enhanced feature map The multiplication result is denoted as the frequency domain processing result A(X).

[0037] Furthermore, when training the model, use the mean square error between the noise added at each time step and the predicted noise as the loss function to optimize the model.

[0038] The present invention also provides a before-and-after comparison medical image synthesis system based on a contour-guided accelerated diffusion model, including:

[0039] A memory configured to store a computer program;

[0040] A processor configured to execute the computer program to implement the method for synthesizing pre- and post-contrast medical images based on a contour-guided accelerated diffusion model as described above.

[0041] The present invention also provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the method for synthesizing pre- and post-contrast medical images based on a contour-guided accelerated diffusion model as described above is implemented.

[0042] The beneficial effects of the present invention are as follows:

[0043] The present invention proposes a new method for synthesizing post-contrast medical images, which consists of a non-uniform sampling strategy, contour guidance, and a multi-frequency enhancement module. Among them, the non-uniform sampling strategy uses a jumping time step according to the changes in the mean and variance to reduce the number of iterations. The contour guidance adds generation constraints at each time step to maintain structural consistency. Finally, the multi-frequency enhancement attention module enhances the feature representation in parallel branches in space and channels, and aggregates to learn global key features through frequency channels. This method can solve the problems of anatomical structure deformation and the difficulty in distinguishing random noise and high-frequency features during the generation of post-contrast images, and maintain a short training time during this process. This method innovates on the problems existing in the current mainstream methods for post-contrast medical images, uses contour information as a conditional constraint, reduces non-Gaussian prediction noise, and guides the generation process to maintain anatomical consistency. The multi-frequency enhancement attention module is used to capture key high-frequency features, distinguish noise, and further retain fine edge details. The non-uniform sampling method is used to improve the training sequence. The non-uniform time step sequence reduces the number of iterations and computational costs.

[0044] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent description, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the following description. Description of the Drawings

[0045] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings, where:

[0046] Figure 1 is the architecture diagram of the diffusion model provided by the embodiment of the present invention;

[0047] Figures 2(A) and 2(B) are the post-contrast image results generated from brain magnetic resonance imaging (MRI) provided by the embodiment of the present invention;

[0048] Figures 3(A) and 3(B) are the post-contrast image results generated from the computed tomography (CT) images of the nasopharynx, larynx, and pharynx provided by the embodiments of the present invention. Detailed implementation manners

[0049] The preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the preferred embodiments are only for illustrating the present invention and not for limiting the protection scope of the present invention.

[0050] The present invention provides a method for synthesizing pre- and post-contrast medical images based on a contour-guided accelerated diffusion model, including:

[0051] Obtaining pre-contrast medical images

[0052] Constructing and training a contour-guided accelerated diffusion model;

[0053] Using the trained contour-guided accelerated diffusion model to generate post-contrast medical images corresponding to the pre-contrast medical images The entire time step of the contour-guided accelerated diffusion model is divided into a changing region and a stable region, and the time sampling interval in the changing region is smaller than that in the stable region.

[0054] During the denoising process at any time step in the reverse diffusion process, based on the contour of the pre-contrast medical image, a guiding condition c is determined, and denoising is performed based on this guiding condition c to synthesize the post-contrast medical image. That is Figure 1 in

[0055] Referring to Figure 1 , the diffusion model proposed by the present invention, namely the contour-guided accelerated diffusion model, includes two processes, a forward diffusion process and a reverse diffusion process. The entire time step of this diffusion model is divided into a changing region and a stable region, and the division is based on a preset threshold τ. When the time step t < τ (i.e., the changing region A), it means that the diffusion process is in the changing stage; when τ < t ≤ T, it means that the diffusion process is in the stable stage. Among them, the time sampling interval in the changing region A is S1, and the time sampling interval in the stable region B is S2, and S1 is smaller than S2.

[0056] The changing region corresponds to the changing stage in the diffusion process, and the stable region corresponds to the stable stage in the diffusion process. The changing stage refers to the stage where the mean and variance of the process increment change significantly, and the stable stage refers to the stage where the mean and variance of the process increment tend to be stable and approach the noise distribution.

[0057] In the diffusion process, x at any time step tt can be regarded as the original data x 0 and the linear combination of the random noise ε, which can be expressed as Due to the process increment δ t = x t+1 - x t , so the mean is and the variance is where I represents the identity matrix, and α t = 1 - β t ; β t represents the noise level at time step t; As the time step increases, the mean ω t of the process increment δ t decreases and approaches 0, and the variance t of the process increment δ shows a trend of increasing first and then decreasing and finally stabilizing, and finally stabilizes at a level close to pure noise (that is, finally approaches 2I). Therefore, according to the change trend of the mean and variance of the process increment δ t , the entire diffusion process can be divided into a changing stage and a stable stage. The range of the stable stage is represented by the scale r. When , the time step t that meets the condition is in the stationary stage, where r is a given value and greater than 1,

[0058] Therefore, through the above description, the setting of the time threshold τ can be expressed as:

[0059] In the forward diffusion process, determine the image x t after adding noise at each time step t;

[0060] Calculate the variance of the process increment δ t = x t+1 - x t ;

[0061] If the variance of the process increment obtained at a time step t is greater than the preset variance value, then take this time step t as the time threshold.

[0062] In some embodiments of the present invention, determining the guiding condition c based on the contour of the pre - comparison medical image may include the following steps:

[0063] Use an edge detector to determine the contour representation of the pre - comparison medical image The edge detector can be a Canny edge detector;

[0064] Then, the pre - comparison medical image and the contour representation Concatenate on the channel dimension to generate the guiding condition c, that is

[0065] Take the guiding condition c as the input of the denoising network (for example, it can be a U-Net denoising network). Then, in each synthesis stage (i.e., the reverse diffusion process), the model adjusts the synthesized content according to this contour information and keeps the data manifold of the intermediate features within appropriate boundaries, avoiding the deviation of the anatomical structure of the synthesized image from reality and the deformation of the anatomical structure.

[0066] Furthermore, the denoising process at any time step in the reverse diffusion process includes using a noise prediction model to estimate the noise. Among them, the noise estimation model includes an encoder, a multi-frequency attention module, and a decoder. The multi-frequency attention module enhances the feature representation in the channel and spatial dimensions through channel and spatial selective filtering and dynamic range enhancement, and then selectively emphasizes specific frequency components in the frequency domain in the frequency channel to strengthen the attention to the channel and comprehensively capture key features.

[0067] The encoder receives the noise image at the current time step t∈{0,1,...,T} and the condition c as inputs, and after being encoded by the encoder, a tensor X is obtained. That is En represents the encoder, and the noise image represents the image obtained by removing the predicted noise at time step t + 1 from the noise image at time step t + 1.

[0068] The multi-frequency attention module processes the tensor X in the spatial branch and the channel branch and then performs frequency domain processing to generate the frequency domain processing result A(X);

[0069] The decoder decodes A(X) to generate the predicted noise corresponding to time step t;

[0070] By predicting the noise for all time steps and removing the noise, the comparison image is obtained.

[0071] Among them, the processing of the tensor X by the multi-frequency attention module in the channel branch includes:

[0072] Perform the first convolution and the second convolution on the tensor X;

[0073] Perform the Softmax function processing on the result of the second convolution to generate the first processing result;

[0074] Perform matrix multiplication on the first processing result and the result of the first convolution, and sequentially pass the multiplication result through the third convolution, layer normalization processing, and Softmax function processing to generate the channel branch result.

[0075] The multi-frequency attention module processes the tensor X in the spatial branch, including:

[0076] Performing the fourth convolution and the fifth convolution on the tensor X;

[0077] Sequentially passing the result of the fourth convolution through a global pooling operation and a Softmax function to generate a second processing result;

[0078] Performing matrix multiplication on the second processing result and the result of the fifth convolution, and then performing the Softmax function on the multiplication result again to generate the spatial branch result.

[0079] The processing of the tensor X by the multi-frequency attention module in the channel branch can be expressed by the formula:

[0080]

[0081] The processing of the tensor X by the multi-frequency attention module in the spatial branch can be expressed by the formula:

[0082]

[0083] Among them, represents matrix multiplication, W z , W v and W q respectively represent 1x1 convolutional layers, λ 1 , λ 2 and λ 3 respectively represent tensor reshaping operators, f sg (·) represents the Sigmoid function, f sm (·) represents the Softmax function, f gp (·) represents the global pooling operation, M ch (X) and M sp (X) respectively represent the channel branch result and the spatial branch result.

[0084] It should be noted that the above formula representations for the processing of the tensor X by the multi-frequency attention module in the channel branch and the spatial branch are only illustrative, and do not mean that only these processes shown in the formula are performed during the processing. The formula only shows the key steps in the processing.

[0085] The channel branch only operates on the channel dimension to capture the correlation between different channels, and the spatial branch only operates on the spatial dimension to measure the correlation between different spatial positions.

[0086] Referring to Figure 1 , after performing the spatial branch and channel branch processing, the spatial branch result and the channel branch result are respectively dot-multiplied with the tensor X, and the dot-multiplication result Z chand Z sp Add them together and obtain an enhanced feature map through fusion That is where ⊙ represents dot product

[0087] After completing the above steps, process the aggregated features through a two-dimensional discrete cosine transform (DCT) layer in the frequency channel Inside it, select specific frequency components to enhance channel attention. Specifically, for the enhanced feature map The frequency domain processing includes

[0088] Divide the enhanced feature map along the channel dimension into two or more parts of features, which can be expressed as where and C′ = C / n, n represents the number of divided parts, and its value can be dynamically determined by separately evaluating the performance of each frequency component according to experiments and then selecting the n components with the highest performance; C represents the number of channels of the original input image, H represents the height of the original input image, and W represents the width of the original input image

[0089] Next, for each part of the features, perform a two-dimensional discrete cosine transform, that is 2DDCT is the two-dimensional discrete cosine transform, [u i , v i is the two-dimensional index corresponding to the frequency component

[0090] Conduct a fully connected operation on the multi-frequency vector Y composed of the two-dimensional discrete cosine transform results of each part and perform Softmax function processing to generate the third processing result

[0091] Multiply the third processing result with the enhanced feature map The multiplication result is denoted as the frequency domain processing result A(X)

[0092] The above process is expressed by the formula as

[0093]

[0094] where cat represents concatenation along the channel dimension, f fc (·) represents the fully connected operation, f sg (·) represents the Sigmoid function

[0095] The high resolution calculated by A(X) obtained through multi-frequency enhanced attention helps to capture the details of the image, while retaining the overall structure of the image, also retaining the spatial features and complex texture details of the image

[0096] Then, the decoder is used to predict the noise.

[0097] For any time step t, the predicted noise image is ε θ = Dn(A(X)), where Dn is the decoder.

[0098] The noise prediction loss between the predicted noise and the noise added at each step of forward diffusion can be expressed by the mean squared error through the following formula:

[0099]

[0100] where ε θ (x t , t|c) is the predicted noise generated by the decoder at time step t, and θ is the denoising network parameter. Subsequently, this predicted noise image is compared with the actual real noise image to guide the optimization of the model.

[0101] The present invention also provides a before-and-after contrast medical image synthesis system based on a contour-guided accelerated diffusion model, including:

[0102] A memory configured to store a computer program;

[0103] A processor configured to execute the computer program to implement the before-and-after contrast medical image synthesis method based on the contour-guided accelerated diffusion model as described above.

[0104] The present invention also provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, it implements the before-and-after contrast medical image synthesis method based on the contour-guided accelerated diffusion model as described above.

[0105] Figures 2(A) and 2(B) show the generated results after comparison of the present invention in brain magnetic resonance imaging, and Figures 2(A) and 2(B) represent different instances. Figures 3(A) and 3(B) show the generated results after comparison of the present invention in nasopharyngeal computed tomography imaging, and Figures 3(A) and 3(B) represent different instances. The content corresponding to the second row of each instance is the locally enlarged area. In addition, the difference between the generated image and the real comparison image is highlighted with a MAE (i.e., Masked Autoencoder Error) heat map.

[0106] UNIT, StarGANv2, DDPM, Reg-GAN, DC-CycleGAN, and DCE-MRI in FIGS. 2(A) and 2(B) and FIGS. 3(A) and 3(B) are comparative methods. Among them, the relevant content of UNIT can be referred to "Unsupervised image-to-image translation networks" published by Liu et al. in NeurIPs 2017; the relevant content of StarGANv2 can be referred to "Starganv2: Diverse image synthesis for multiple domains" published by Choi et al. in CVPR 2020; the relevant content of DDPM can be referred to "Denoising diffusion probabilistic models" published by Ho et al. in NeurIPS 2020; the relevant content of Reg-GAN can be referred to "Breaking the dilemma of medical image-to-image translation" published by Ho et al. in NeurIPS 2020; the relevant content of DC-CycleGAN can be referred to "DC-cycleGAN: bidirectional CT-to-MR synthesis from unpaired data" published by Wang et al. in CMIG 2023; the relevant content of DCE-MRI can be referred to "Pre-to post-contrast breast MRI synthesis for enhanced tumour segmentation" published by Osuala et al. in MIP 2024.

[0107] As can be seen from FIGS. 2(A) and 2(B) and FIGS. 3(A) and 3(B), the results of the method proposed by the present invention are closer to the real results and more pathological information is diagnosed.

[0108] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the present technical solution, and they should all be covered by the scope of the claims of the present invention.

Claims

1. A method for synthesizing before-after contrast medical images based on a contour-guided accelerated diffusion model, characterized in that: include: Obtaining pre-contrast medical images Build and train a contour-guided accelerated diffusion model; Generate and compare pre-clinical medical images using the trained contour-guided accelerated diffusion model Corresponding contrasted medical images The contour-guided accelerated diffusion model includes two processes, the forward diffusion process and the reverse diffusion process. The entire time step of the contour-guided accelerated diffusion model is divided into a changing region and a stable region. The time sampling interval of the changing region is smaller than the time sampling interval of the stable region. In the denoising process at any time step in the reverse diffusion process, based on the pre-contrast medical image The contour of the image is used to determine the guiding condition c, and then denoising is performed based on the guiding condition c to synthesize the contrasted medical image.

2. The method for synthesizing before-after contrast medical images based on contour-guided accelerated diffusion model according to claim 1, characterized in that: The division of the changing area and the stable area is based on the set time threshold. The time threshold is set as follows: In the forward diffusion process, determine the image x after adding noise at each time step t t ; Calculate the process increment δ t =x t+1 -x t The variance of If the variance of the process increment obtained at a time step t is greater than the preset variance value, the time step t is used as the time threshold.

3. The method for synthesizing before-after contrast medical images based on contour-guided accelerated diffusion model according to claim 1, characterized in that: Based on pre-contrast medical images The contour determination guide conditions c include: Using edge detectors, determine the pre-contrast medical image The contour representation Medical images before comparison and contour representation Concatenate in the channel dimension to generate the guided condition c.

4. The method for synthesizing before-after contrast medical images based on contour-guided accelerated diffusion model according to claim 1, characterized in that: The denoising process at any time step in the reverse diffusion process involves estimating the noise using a noise prediction model. The noise estimation model includes an encoder, a multi-frequency attention module and a decoder. The encoder receives the noisy image at the current time step t∈{0,1,...,T} and condition c as input, and after being encoded by the encoder, the tensor X is obtained, where T represents the entire time step of the diffusion process; The multi-frequency attention module processes the tensor X in the spatial branch and the channel branch and then processes it in the frequency domain to generate the frequency domain processing result A(X); The decoder decodes A(X) to generate prediction noise corresponding to time step t; The contrasted image is obtained by predicting the noise for all time steps and removing the noise.

5. The method for synthesizing before-after contrast medical images based on contour-guided accelerated diffusion model according to claim 4, characterized in that: The multi-frequency attention module processes the tensor X in the channel branch including: Perform the first and second convolutions on the tensor X; The result of the second convolution is processed by the Softmax function to generate a first processing result; The first processing result is matrix multiplied with the first convolution result, and the multiplication result is sequentially processed by the third convolution, layer normalization and Softmax function to generate a channel branch result.

6. The method for synthesizing before-after contrast medical images based on contour-guided accelerated diffusion model according to claim 4, characterized in that: The multi-frequency attention module processes the tensor X in the spatial branch including: Perform the fourth and fifth convolutions on the tensor X; The result of the fourth convolution is processed by the global pooling operation and the Softmax function in sequence to generate a second processing result; The second processing result and the fifth convolution result are matrix multiplied, and the multiplication result is processed by the Softmax function again to generate a spatial branch result.

7. The method for synthesizing before-after contrast medical images based on contour-guided accelerated diffusion model according to claim 4, characterized in that: The method further comprises: After processing the spatial branch and channel branch, the spatial branch result and the channel branch result are respectively multiplied with the tensor X, and the dot product results are added to obtain the enhanced feature map The enhanced feature map Perform frequency domain processing, including: The enhanced feature map Features that are divided into two or more parts along the channel dimension; For each feature, perform a two-dimensional discrete cosine transform; Performing a full connection operation and a Softmax function processing on the multi-frequency vector composed of each corresponding two-dimensional discrete cosine transform result to generate a third processing result; The third processing result is combined with the enhanced feature map Perform matrix multiplication, and the multiplication result is recorded as the frequency domain processing result A(X).

8. The method for synthesizing before-after contrast medical images based on contour-guided accelerated diffusion model according to claim 4, characterized in that: When training the model, the mean square error between the noise added at each time step and the predicted noise is used as the loss function to optimize the model.

9. A before-after contrast medical image synthesis system based on contour-guided accelerated diffusion model, characterized in that: include: a memory configured to store a computer program; A processor is configured to execute the computer program to implement the before-after contrast medical image synthesis method based on the contour-guided accelerated diffusion model as described in any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the before-after contrast medical image synthesis method based on the contour-guided accelerated diffusion model according to any one of claims 1 to 8 is implemented.