Wavelet diffusion model CTA image generation method based on adaptive coordinate attention mechanism

By employing a wavelet diffusion model with an adaptive coordinate attention mechanism, the risks of nephrotoxicity and radiation associated with CTA examinations are mitigated, resulting in high-quality CTA images that meet the needs for early diagnosis of aortic diseases.

CN121564149APending Publication Date: 2026-02-24ZHEJIANG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511581627.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-31
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Current CTA examinations pose risks of nephrotoxicity and high radiation doses. Plain CT scans lack sufficient sensitivity and specificity in the screening and diagnosis of aortic diseases. Furthermore, CTA images generated by existing deep learning models have blurred vascular features and significant noise, affecting diagnostic accuracy.

Method used

A wavelet diffusion model based on an adaptive coordinate attention mechanism is adopted to generate high-quality CTA images through three-dimensional wavelet transform and adaptive coordinate attention mechanism, thereby reducing the memory overhead during the training phase and improving the accuracy of image detail and lesion identification.

Benefits of technology

It generates high-quality CTA images in a short time, avoids the problem of excessive video memory usage, improves the accuracy of image detail information and lesion generation, and meets the needs of clinical diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121564149A_ABST
    Figure CN121564149A_ABST
Patent Text Reader

Abstract

A wavelet diffusion model CTA image generation method based on an adaptive coordinate attention mechanism comprises the following steps: firstly, performing registration and data preprocessing on CT and CTA images, and decomposing the images into a low-frequency sub-band and a plurality of high-frequency sub-bands through three-dimensional discrete wavelet transform so as to obtain multi-scale feature representation; then modeling and iterative generation are carried out on wavelet sub-bands in a conditional diffusion probability model, meanwhile, an adaptive coordinate attention mechanism is introduced, a context dependency relationship is modeled in the depth direction, the height direction and the width direction respectively, a direction-sensitive attention weight map is generated, wavelet coefficients are modulated element by element, and the wavelet coefficients are modulated element by element; therefore, the blood vessel area is enhanced and background noise is suppressed. And finally, reconstructing the generated wavelet coefficient through three-dimensional inverse wavelet transform to obtain a high-quality CTA image. According to the method, the focus form can be more accurately generated, the quality of the generated image is improved, meanwhile, the video memory overhead and the calculation complexity of the diffusion model in the training and reasoning stages are remarkably reduced, and the generation speed is increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of medical imaging and cardiovascular surgery assistance under computer graphics, and in particular to the early diagnosis of cardiovascular diseases such as aortic dissection, aortic aneurysm, aortoiliac artery stenosis or occlusion, and aortic atherosclerosis based on plain CT images. Background Technology

[0002] The aorta is the main conduit for blood delivery to all parts of the body and is the most important arterial vessel in the human body. Therefore, aortic diseases have a rapid onset, progress quickly, and have a high mortality rate; early and accurate diagnosis is crucial for reducing mortality and improving prognosis. In current clinical practice, computed tomography angiography (CTA) has become the preferred imaging method for diagnosing aortic diseases due to its high temporal resolution, spatial resolution, and three-dimensional reconstruction advantages. It is widely used in the diagnosis and postoperative evaluation of diseases such as aortic dissection, aortic aneurysm, aortic stenosis and occlusion, and aortic atherosclerosis.

[0003] However, the clinical application of CTA in my country still faces many challenges: ① CTA relies on intravenous injection of iodine contrast agents, posing potential risks of nephrotoxicity and adverse contrast agent reactions, limiting its promotion in high-risk populations and primary hospitals. ② The radiation dose from CTA imaging is relatively high, making it unsuitable for frequent follow-ups or large-scale screening. In contrast, computed tomography (CT) has significant advantages such as ease of operation and no need for contrast agent injection; however, its sensitivity and specificity in aortic disease screening and diagnosis remain significantly limited.

[0004] To address the aforementioned issues, current technologies typically utilize deep learning models to convert CT images into corresponding CTA images. These deep learning models often employ adversarial network frameworks or diffusion models. However, adversarial generative networks suffer from issues such as pattern collapse during training, while diffusion models have drawbacks including limited controllability of image details and lesion morphology, high storage and GPU overhead, and slow inference speed. Arterial disease diagnosis relies on accurate identification of vessel morphology (such as dissection rupture and aneurysm size). If the generated CTA images contain blurred vessel features and significant noise, it directly impacts diagnostic accuracy. Furthermore, existing models lack the ability to adaptively distinguish between "vessels and background," potentially leading to blurred vessel edges and unintended noise in the generated images, increasing the risk of missed or misdiagnosed diagnoses. Summary of the Invention

[0005] To overcome the problems of high memory overhead, slow inference time, blurred details in generated images, weak feature focusing ability, and inability to meet clinical requirements for vascular enhancement and noise suppression in diffusion models during the training and inference phases, this invention provides a CTA image generation method based on an adaptive coordinate attention mechanism wavelet diffusion model. This method, by applying a diffusion model based on wavelet transform, can significantly reduce the memory overhead during the training phase. Simultaneously, we propose an adaptive coordinate attention mechanism based on wavelet coefficients. This mechanism, through three-dimensional context modeling and adaptive weight allocation, allows the neural network to autonomously focus on the detailed information of the region of interest, thereby generating lesion morphology more accurately. Therefore, it becomes a necessary technical means to improve the quality of generated images and meet clinical diagnostic needs.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0007] A wavelet diffusion model CTA image generation method based on adaptive coordinate attention mechanism includes the following steps:

[0008] Step 1, Dataset Preparation: Model training and validation are performed using large-scale data. The dataset is obtained by organizing existing CT and CTA images in the hospital database.

[0009] Step 2, Discrete Wavelet Transform (DWT): Perform three-dimensional wavelet decomposition on both CT and CTA images, decomposing them into low-frequency subbands (LF) and multiple high-frequency subbands (HF);

[0010] Step 3, Design and training of diffusion probability model network: Construct a conditional diffusion model (based on 3D U-Net), with the input being the multi-scale sub-band after wavelet decomposition. The model generates predicted wavelet coefficients through an iterative diffusion process, and the conditional information ensures the consistency of vascular structure.

[0011] Step 4, Inverse Discrete Wavelet Transform (IDWT): The generated wavelet coefficients are reconstructed into a complete CTA image through a three-dimensional inverse wavelet transform;

[0012] Step 5, Prediction Generation: Use the model trained in Step 3 to generate and predict CTA images from clinical CT data.

[0013] Furthermore, the process of step 1 is as follows:

[0014] 1.1) Data Sources and Parameters:

[0015] The data comes from the First Affiliated Hospital of Wenzhou Medical University and includes CT and CTA modalities. The data parameters are as follows: CT uses 5mm thick mediastinal window images, and CTA uses 2.5mm thick mediastinal window images.

[0016] 1.2) Image Registration

[0017] To ensure that the images have the same slice thickness and the corresponding organs are aligned, CT and CTA images from the same patient are registered. The Ants tool is used to attach the CT image to the CTA image to align the two images in spatial position.

[0018] Furthermore, the process of step 2 is as follows:

[0019] For a one-dimensional signal, such as audio, where the value is a single independent variable, s = {s j} j∈Z DWT decomposes it into a low-frequency signal s1 = {s 1k} k∈Z and high-frequency signal d1={d 1k} k∈Z ,in

[0020] s 1k =∑ j l j-2k s j , (1);

[0021] d 1k =∑ j h j-2k s j , (2);

[0022] Where L={l k} k∈Z , h={h k} k∈Z These are low-pass and high-pass filters, respectively. In this invention, a three-dimensional signal, i.e., volume data from a CT scan, is used. Similarly, the three-dimensional signal DWT is decomposed into low-frequency and high-frequency signals. The discrete wavelet transform is composed of a low-pass filter and a high-pass filter with a step size of 2, expressed as... and This is applied in three spatial dimensions, which decompose the registered 3D image from step 1 into 8 wavelet coefficients (x... lll x llh x lhl x lhh x hll x hlh x hhl x hhh ).

[0023] Furthermore, the process of step 3 is as follows:

[0024] 3.1) Adaptive Coordinate Attention (ACA) Mechanism for Wavelet Coefficients: To better utilize the sub-band features obtained from 3D wavelet decomposition, an adaptive coordinate attention mechanism is proposed. Its core idea is to model contextual dependencies in the depth (D), height (H), and width (W) directions respectively, and then fuse them to generate a spatially correlated attention weight map, thereby enhancing the vascular structure and suppressing background noise. The first step is directional feature extraction, where the feature map F∈R of the 3D sub-band is processed. 8×D×H×W Global average pooling is performed in the depth, width, and height directions respectively. Then, feature compression and fusion are performed, concatenating the context features from the three directions to form a joint representation. A feature compression module then compresses the dimensionality of the concatenated features to generate a global low-dimensional representation. Finally, this global representation is mapped to the depth, height, and width directions respectively to generate corresponding weight maps, A. D A H A w These weight maps, after being normalized using the Sigmoid function, can reflect the importance of different positions; finally, attention fusion is performed, expanding the weight maps in the three directions to the same spatial dimension as the original feature F, and then multiplying them element-wise: ☉ indicates element-wise multiplication. The fusion operation, which represents the orientation weights, uses an adaptive coordinate attention mechanism to obtain the corresponding feature map from the CT image.

[0025] 3.2) Diffusion Model: The diffusion model consists of two processes: a forward process and a reverse process. The forward process is also called the diffusion process or the noise-adding process, and the reverse process is also called the noise-removing process. Given a sample x0 in the real data distribution, the noise-adding process gradually adds Gaussian noise to the sample according to a series of normal distributions within a specified time step T.

[0026]

[0027] t∈{1,…,T},β 1:T It is a defined variance table. Reverse process modeling is called a Markov chain:

[0028] Each step of the process follows a Gaussian distribution, with its mean being... It is determined by the parameters ε of the neural network. θ The decision can predict noise added to an image; where α t :=1-β t , From a noisy image x with a time step of t t Perform sampling:

[0029]

[0030] By training ε θ To predict the denoised image x0 = ε θ (x t During training, the wavelet coefficients of the conditional images, namely CT images and CTA images, are concatenated according to channels, and the training model obtains CTA images from CT images. During the prediction stage, the input image consists only of CT images, and CTA images are generated through the U-net network.

[0031] Furthermore, the process of step 4 is as follows:

[0032] During the IDWT process, the generated wavelet coefficients are reconstructed into a complete CTA image using a three-dimensional inverse wavelet transform. The image is then reconstructed using the data from s.

[0033] s j =∑ k (l j-2ks1k +h j-2k d 1k (6);

[0034] In the overall network training process, the forward diffusion process is trained, allowing the network to obtain the corresponding CTA image from the input CT image. Mean Squared Error (MSE) loss is used to determine the difference between the generated result and the real result.

[0035] The CTA image is generated by the neural network, and x0 is the real CTA image; then the generated CTA image is predicted through a backsampling process.

[0036] In step 5, the clinical data used for testing is input into the network trained in step 3. By inputting CT images, CTA images are accurately generated in the wavelet diffusion model network.

[0037] The beneficial effects of this invention are: 1. It can generate corresponding CTA images from CT images in a very short time; 2. It avoids the problems of network training difficulties and large memory usage during model training; 3. The adaptive coordinate attention mechanism provides rich detailed information for the generated images, ensuring the accuracy of lesion generation and making the generated images more accurate and realistic. Attached Figure Description

[0038] Figure 1This is a flowchart of the overall network training process. Detailed Implementation

[0039] The present invention will be further described below.

[0040] Reference Figure 1 A wavelet diffusion model CTA image generation method based on adaptive coordinate attention mechanism includes the following steps:

[0041] Step 1: Dataset preparation, the process is as follows:

[0042] 1.1) Data Sources and Parameters

[0043] The data comes from the First Affiliated Hospital of Wenzhou Medical University and includes CT and CTA modalities. The data parameters are as follows: CT uses 5mm thick mediastinal window images, and CTA uses 2.5mm thick mediastinal window images.

[0044] 1.2) Image Registration

[0045] To ensure that the images have the same slice thickness and the corresponding organs are aligned, CT and CTA images from the same patient are registered. The Ants tool is used to attach the CT image to the CTA image to align the two images in spatial position.

[0046] Step 2, Discrete Wavelet Transform: For a one-dimensional signal s = {s j} j∈Z DWT can decompose it into low-frequency signals s1 = {s 1k} k∈Z and high-frequency signal d1={d 1k} k∈Z ,in,

[0047] s 1k =∑ j l j-2k s j , (1);

[0048] d 1k =∑ j h j-2 ks j , (2);

[0049] Where L={l k} k∈Z , h={h k} k∈Z These are a low-pass filter and a high-pass filter, respectively. The same principle applies to three-dimensional signals. The discrete wavelet transform is composed of a low-pass filter and a high-pass filter with a step size of 2, expressed as: and This is applied to three spatial dimensions. They decompose the registered 3D image from step one into eight wavelet coefficients (x... lll x llh x lhl x lhl x hll x hlh x hhl x hhh ).

[0050] Step 3: Design and training of the diffusion probability model network, the process is as follows:

[0051] 3.1) Diffusion Model: The diffusion model consists of two processes: a forward process and a reverse process. The forward process is also called the diffusion process, and the reverse process is also called the denoising process. Given a sample x0 in the real data distribution, this denoising process adds Gaussian noise to the sample step by step within a specified time step T according to a series of normal distributions.

[0052]

[0053] t∈{1,…,T},β 1:T It is a defined variance table. Reverse process modeling is called a Markov chain:

[0054] Each step of the process follows a Gaussian distribution, with its mean being... It is determined by the parameters ε of the neural network. θ This decision allows it to predict noise added to an image. Where α... t :=1-β t , From a noisy image x with a time step of t t Perform sampling:

[0055]

[0056] By training ε θ To predict the denoised image x0 = ε θ (x t ,t).

[0057] 3.2) Adaptive Coordinate Attention (ACA) Mechanism for Wavelet Coefficients: To better utilize the sub-band features obtained from 3D wavelet decomposition, an adaptive coordinate attention mechanism is proposed. Its core idea is to model context dependencies in the depth (D), height (H), and width (W) directions respectively, and then fuse them to generate a spatially correlated attention weight map, thereby enhancing the vascular structure and suppressing background noise. The first step is directional feature extraction, where the feature map F∈R of the 3D sub-band is processed. 8×D×H×W Global average pooling is performed in the depth, width, and height directions respectively. Then, feature compression and fusion are performed, concatenating the context features from the three directions to form a joint representation. A feature compression module then compresses the dimensionality of the concatenated features to generate a global low-dimensional representation. Finally, this global representation is mapped to the depth, height, and width directions respectively to generate corresponding weight maps, A. D A H A w These weight maps, after being normalized using the Sigmoid function, can reflect the importance of different positions; finally, attention fusion is performed, expanding the weight maps in the three directions to the same spatial dimension as the original feature F, and then multiplying them element-wise: ☉ indicates element-wise multiplication. This represents the fusion operation of directional weights.

[0058] Step 4, Inverse Wavelet Transform: During the IDWT process, the generated wavelet coefficients are reconstructed into a complete CTA image using a three-dimensional inverse wavelet transform. The image is then reconstructed using the data from s.

[0059] s j =∑ k (l j-2k S 1k +h j-2k d 1k (6);

[0060] In the overall network training process, the forward diffusion process is trained, allowing the network to obtain the corresponding CTA image from the input CT image. Mean Squared Error (MSE) loss is used to determine the difference between the generated result and the real result.

[0061] The CTA image is generated by the neural network, where x0 is the real CTA image; then, the generated CTA image is predicted through a backsampling process.

[0062] Step 5, Prediction Generation: Input the clinical data used for testing into the network trained in Step 3, input the CT images into the wavelet diffusion model network, and generate CTA images.

[0063] The solution in this embodiment can generate corresponding CTA images from CT images in a very short time. At the same time, it avoids the problems of difficult network training and large memory consumption during model training. The adaptive coordinate attention mechanism in this embodiment provides rich detail information for the generated images, ensuring the accuracy of lesion generation and making the generated images more accurate and realistic.

[0064] The embodiments described in this specification are merely examples of implementations of the inventive concept and are for illustrative purposes only. The scope of protection of this invention should not be considered limited to the specific forms described in these embodiments; rather, it extends to equivalent technical means conceived by those skilled in the art based on the inventive concept.

Claims

1. A wavelet diffusion model CTA image generation method based on adaptive coordinate attention mechanism, characterized in that, The method includes the following steps: Step 1, Dataset Preparation: Model training and validation are performed using large-scale data. The dataset is obtained by organizing existing CT and CTA images in the hospital database. Step 2, Discrete Wavelet Transform (DWT): Perform three-dimensional wavelet decomposition on both CT and CTA images, decomposing them into a low-frequency subband (LF) and multiple high-frequency subbands (HF). Step 3, Design and training of diffusion probability model network: Construct a conditional diffusion model based on 3D U-Net. The input is the multi-scale sub-band after wavelet decomposition. The model generates predicted wavelet coefficients through an iterative diffusion process. The conditional information ensures the consistency of vascular structure. Step 4, Inverse Wavelet Transform IDWT: The generated wavelet coefficients are reconstructed into a complete CTA image through three-dimensional inverse wavelet transform; Step 5, Prediction Generation: Use the model trained in Step 3 to generate and predict CTA images from clinical CT data.

2. The CTA image generation method based on an adaptive coordinate attention mechanism wavelet diffusion model as described in claim 1, characterized in that, The process of step 1 is as follows: 1.1) Data Sources and Parameters: The data consisted of CT and CTA modalities; 1.2) Image Registration To ensure that the images have the same slice thickness and the corresponding organs are aligned, CT and CTA images from the same patient are registered. The Ants tool is used to attach the CT image to the CTA image to align the two images in spatial position.

3. The wavelet diffusion model CTA image generation method based on adaptive coordinate attention mechanism as described in claim 1 or 2, characterized in that, The process of step 2 is as follows: For a one-dimensional signal s = {s j } j∈Z DWT decomposes it into a low-frequency signal s1 = {s 1k } k∈z and high-frequency signal d1={d 1k } k∈Z ,in s 1k =∑ j l j - 2k s j , (1); d 1k =∑ j h j-2k s j , (2); Where L={l k } k∈z , h={h k } k∈Z These are a low-pass filter and a high-pass filter, respectively. The same principle applies to three-dimensional signals. The discrete wavelet transform is composed of a low-pass filter and a high-pass filter with a step size of 2, expressed as: and This is applied in three spatial dimensions, which decompose the registered 3D image from step 1 into 8 wavelet coefficients (x... lll x llh x lhl x lhh x hll x hlh x hhl x hhh ).

4. The wavelet diffusion model CTA image generation method based on adaptive coordinate attention mechanism as described in claim 1 or 2, characterized in that: The process of step 3 is as follows: 3.1) Diffusion Model: The diffusion model consists of two processes: a forward process and a reverse process. The forward process is also called the diffusion process or the noise-adding process, and the reverse process is also called the noise-removing process. Given a sample x0 in the real data distribution, the noise-adding process is to gradually add Gaussian noise to the sample according to a series of normal distributions within a specified time step T. t∈{1,…,T},β 1:T It is a defined variance table; the reverse process modeling is called a Markov chain: Each step of the process follows a Gaussian distribution, with its mean being... It is determined by the parameters ε of the neural network. θ The decision can predict noise added to an image, where α t :=1-β t , From a noisy image x with a time step of t t Perform sampling: By training ε θ To predict the denoised image x0 = ε θ (x t ,t).

5. The wavelet diffusion model CTA image generation method based on adaptive coordinate attention mechanism as described in claim 4, characterized in that: Step 3 also Includes the following processes: 3.2) Adaptive Coordinate Attention Mechanism (ACA) for Wavelet Coefficients: First, directional feature extraction is performed on the feature map F∈R of the three-dimensional sub-band. 8×D×H×W Global average pooling is performed in the depth, width, and height directions respectively. Then, feature compression and fusion are performed, concatenating the context features from the three directions to form a joint representation. A feature compression module then compresses the dimensionality of the concatenated features to generate a global low-dimensional representation. Finally, this global representation is mapped to the depth, height, and width directions respectively to generate corresponding weight maps, A. D A H A w These weight maps, after being normalized using the Sigmoid function, can reflect the importance of different positions; finally, attention fusion is performed, expanding the weight maps in the three directions to the same spatial dimension as the original feature F, and then multiplying them element-wise: ☉ indicates element-wise multiplication. This represents the fusion operation of directional weights.

6. The CTA image generation method based on an adaptive coordinate attention mechanism wavelet diffusion model as described in claim 1 or 2, characterized in that, The process of step 4 is as follows: During the IDWT process, the generated wavelet coefficients are reconstructed into a complete CTA image using a three-dimensional inverse wavelet transform. The image is then reconstructed using the data from s. s j =∑ k (l j-2k s 1k +h j-2k d 1k ) (6); In the overall network training process, the forward diffusion process is trained, allowing the network to obtain the corresponding CTA image from the input CT image. Mean squared error (MSE) loss is used to determine the difference between the generated result and the true result. The CTA image is generated by the neural network, and x0 is the real CTA image; then the generated CTA image is predicted through a backsampling process.

7. The wavelet diffusion model CTA image generation method based on adaptive coordinate attention mechanism as described in claim 1 or 2, characterized in that, In step 5, the clinical data used for testing is input into the network trained in step 3. By inputting CT images, CTA images are accurately generated in the wavelet diffusion model network.