Low-dose CT denoising method based on tight frame wavelet residual diffusion model
The tight-frame wavelet residual diffusion model of low-dose CT images is used to denoise the problem that image denoising is not fine enough in the prior art, and a more efficient and accurate noise removal effect is achieved.
Patent Information
- Application Number
- CN202510123213.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-24
AI Technical Summary
The existing low-dose CT image denoising methods are difficult to effectively remove noise while retaining image quality, and the fixed basis function of traditional wavelet transform is difficult to adapt to image characteristics, resulting in insufficient denoising results.
The denoising method based on the tight frame wavelet residual diffusion model is adopted, and the image is decomposed into low-frequency and high-frequency components through GTF transformation, the low-frequency components are denoised by using the tight frame diffusion model, and the high-frequency components are recovered through the group high-frequency enhancement network, and finally the denoising image is generated through the inverse GTF transformation.
It improves the denoising accuracy and certainty of low-dose CT images, reduces the possibility of additional interference information after denoising, and significantly improves the noise removal effect of CT images.
Smart Images

Figure CN119991491A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image denoising, and in particular to a low-dose CT image denoising method. Background Art
[0002] In the field of medical imaging, computed tomography (CT) plays a very important role in modern clinical medicine. However, X-rays may cause irreversible damage to the human body, including increasing the risk of cancer. When performing X-ray scans, medical staff will formulate strategies to reduce the X-ray dose based on actual conditions. For patients, lower X-ray intensity means lower cancer risk. But for medical staff, the decline in CT images caused by reduced X-ray intensity means the need for more accurate diagnosis. Therefore, it is crucial to improve LDCT images as much as possible while reducing X-ray doses.
[0003] Existing LDCT image reconstruction methods are mainly divided into three categories: denoising based on sinusoidal filtering, iterative reconstruction denoising, and image post-processing denoising. The first two methods require access to the difficult-to-obtain CT raw sinusoidal data, which limits their application. Denoising methods based on image post-processing are mainly divided into two categories: traditional machine learning and deep learning (DL). In recent years, many advances have been made in the field of image processing through the use of deep learning (DL), and the results achieved are far superior to traditional machine learning algorithms in many aspects.
[0004] In recent years, many studies on low-dose CT reconstruction have shown that the diffusion model using a denoising codec network has attracted much attention by mapping high-quality images from randomly sampled Gaussian noise through an iterative diffusion process. However, the reverse process starts from full-noise Gaussian noise, which makes the image restoration results full of diversity, while image denoising requires certainty. Secondly, the traditional wavelet transform shows certain shortcomings when dealing with non-stationary signals or strong noise interference. Its fixed wavelet basis function is difficult to adaptively adjust according to the specific characteristics of the image, resulting in insufficient denoising results. With the development of deep learning technology, wavelet-related methods based on deep learning have gradually attracted attention. By combining the multi-scale characteristics of wavelet transform and the powerful feature learning ability of deep neural networks, more flexible and accurate denoising effects can be achieved in complex image features and noise environments.
[0005] In summary, in the existing technology, traditional wavelet functions are difficult to meet the growing demand for high CT image quality, and need to be combined with deep learning networks; at the same time, the diversity of the inverse process of the diffusion model causes the restored image to have a lot of interfering information out of thin air. Summary of the invention
[0006] In order to solve the above technical problems, the present invention provides a low-dose CT denoising method based on a tight-frame wavelet residual diffusion model, which can provide users with a CT image denoising service that is more consistent with the observation effect of the human eye.
[0007] Specifically, the method comprises the following steps:
[0008] S1: Collect low-dose CT images and conventional-dose CT images to construct CT data sets and convert them into Numpy matrix data;
[0009] S2: Perform GTF (Geometric Tight Framelet) transformation on the image data to obtain 1 low-frequency component and 8 high-frequency components. The high-frequency components are grouped, and the number of high-frequency components in each group is 2, 3, and 3;
[0010] S3: Create a tight frame diffusion model to denoise the low-frequency components;
[0011] S4: Create a group high-frequency enhancement network to restore high-frequency components;
[0012] S5: Finally, the enhanced and denoised components are subjected to inverse GTF transformation to obtain a denoised image.
[0013] Preferably, S1 comprises the following steps:
[0014] S1.1: Acquire CT raw data in dicom format containing normal dose and low dose;
[0015] S1.2: Label the normal-dose CT in the CT raw data as “target” and the low-dose CT as “input”;
[0016] S1.3: Convert the labeled normal-dose CT and low-dose CT in dicom format into NPY matrix format data.
[0017] Preferably, S2 comprises the following steps:
[0018] S2.1: Construct GTF transform and its inverse transform filter;
[0019] S2.2: Perform GTF transformation on the Numpy matrix to obtain 1 low-frequency component and 8 high-frequency components;
[0020] S2.3: Group the transformed components, and divide the high-frequency components into three groups, with the numbers in each group being 2, 3, and 3 respectively.
[0021] Preferably, S2.1 comprises the following steps:
[0022] Construct the basic matrices W0, W1 and W2, which will serve as the key starting modules for generating a series of complex matrix structures; the three basic matrices are in the following form:
[0023]
[0024] The size of the matrix matches the slice size of the image.
[0025] Now defined Where i, j = 0, 1, 2, Represents the Kronecker product, which is combined to generate the GTF transform filter matrix group:
[0026]
[0027] Where D0 is a low-pass filter, and the other components are high-pass filters obtained at different angles, both of which satisfy:
[0028] D L T D H =E
[0029] D L Denotes a low-pass filter, D H represents a high-pass filter and E represents the identity matrix.
[0030] Preferably, S2.2 comprises the following steps:
[0031] Given a single-channel preprocessed Numpy image data matrix X, after GTF transformation, 1 low-frequency component and 8 high-frequency components are obtained. The description of GTF transformation is as follows:
[0032] [X L ,X H ]=Conv([D L ,D H ],X)
[0033] X L is the low frequency component, X H is the high frequency component. Now define the inverse transform of GTF transform as IGTF, which is described as follows:
[0034] X'=Conv(Transposed[D L ,D H ],[X L ,X H ])
[0035] X' is the Numpy matrix of the image output after IGTF, Conv is the convolution operation, and Transposed is the transposition operation.
[0036] Preferably, S2.3 comprises the following steps:
[0037] According to the connection between high-frequency components, the transformation components with the same first digit of the basic matrix subscript number obtained by Kronecker product are grouped into one group, and the grouping is as follows:
[0038] D1 and D2 are grouped together; D4, D5 and D6 are grouped together; D7, D8 and D9 are grouped together.
[0039] Preferably, S3 comprises the following steps:
[0040] S3.1: Introduce a tight frame diffusion model to perform noise diffusion and residual diffusion on low-frequency components;
[0041] S3.2: Construct residual prediction network and noise prediction network;
[0042] S3.3: reverse sampling the diffused degraded image to gradually remove the residual and noise in the components;
[0043] Preferably, S3.1 comprises the following steps:
[0044] Order I t-1 is the low-frequency component of the previous state, I t is the low-frequency component in the current state, T is the total time step of diffusion, and the diffusion process is as follows:
[0045]
[0046] In the above formula, I res is the low-frequency component of the low-dose image after GTF transformation I in The difference between the low-frequency component I0 and the normal dose image after GTF transformation (i.e., I res =I in -I0). α t and β t Control residual diffusion and noise diffusion respectively, t∈{1,…,T}, When t = T,
[0047] Preferably, S3.2 comprises the following steps:
[0048] Use U-Net network to build residual prediction network and the noise prediction network ∈ θ (I t ,t,I in ), use the L1 loss function to guide the learning of the network, the loss functions of the two networks are as follows:
[0049]
[0050] L ∈ (θ): = E[λ ∈ |∈-∈ θ (I t ,t,I in )| 2 ]
[0051] The residual loss is L res (θ), the noise loss is L ∈ (θ). λ res ,λ ε ∈{0,1}, when λ res =1 and λ ε When =1, both residual and noise are predicted.
[0052] Preferably, S3.3 comprises the following steps:
[0053] I at each time step t It has been derived from the diffusion process. Subtracting them from the trained prediction network can get the reverse sampling process as follows:
[0054]
[0055] Finally, the restored low-frequency component and the low-frequency component of the normal dose after GTF transformation are subjected to MSE loss:
[0056]
[0057] is the low-frequency component obtained through the model, low is the low-frequency component obtained through GTF transformation of normal dose, L low is the MSE loss of the two, and N is the number of images.
[0058] Preferably, S4 comprises the following steps:
[0059] S4.1: Construct a shallow feature extraction network for high-frequency components;
[0060] S4.2: Construct a spatial frequency enhancement module for high-frequency components;
[0061] S4.3: construct a group high frequency enhancement network to enhance each high frequency group component;
[0062] Preferably, S4.1 comprises the following steps:
[0063] Design a feature extraction module High-frequency Feature Extract Module (HFEM). This network module has four branches. The first three branches are 1×1, 3×3 and 5×5 multi-scale convolutions used to extract the features of high-frequency components. The last branch is a learnable Sobel operator Learnable Sobel Operator (LSO) used to extract the edge information of high-frequency components. Then the features extracted by the shallow network are flattened and normalized. It is described as follows:
[0064]
[0065] f e =Relu(Conv 3×3 (LSO(x)))
[0066]
[0067] f o =Flatten(LayerNorm(f m ))
[0068] Relu(·) represents the Relu activation function, Conv n×n (·) represents a convolution kernel of size n, represents the concatenation operation, LSO(·) represents the edge detection algorithm, Flatten(·) represents the linear representation of flattening the two-dimensional space into one dimension, LayerNorm(·) represents layer normalization, and f g represents the local features extracted from multiple scales, f e represents the edge features extracted by the learnable Sobel operator, f m It is a combination of local features and edge features. Finally, f m Flatten and normalize to get the output f of the shallow network o .
[0069] Preferably, S4.2 comprises the following steps:
[0070] Construct a spatial-frequency enhancement module Spatial-Frequency Enhance Block (SFEB), use the state space model State Spatial Model (SSM) and convolution operation to extract the characteristics of the component space domain, which is described as follows:
[0071] f ssm1 =SiLU(x)
[0072] f ssm2=DWConv(SiLU(SSM(LayerNorm(x))))
[0073]
[0074] SiLU(·) represents SiLU activation function, LayerNorm(·) represents layer normalization, SSM(·) represents spatial state model, DWConv(·) represents 1×1 separable depthwise convolution, It is a matrix point-by-point multiplication operation.
[0075] The phase spectrum and amplitude obtained by two-dimensional Fourier transform of the high-frequency component are convolved to obtain the characteristics of the frequency domain, which is described as follows:
[0076] f input =Conv 1×1 (x)
[0077] Amp,Phase=FFT(f input )
[0078] f amp =Conv 3×3 (Amp)
[0079] f phase =Conv 3×3 (Phase)
[0080] f fft =IFFT(f amp ,f phase )
[0081] Conv n×n It represents an n-dimensional convolution operation, FFT represents a two-dimensional Fourier transform, Amp and Phase represent amplitude and phase spectra, and IFFT represents the inverse two-dimensional Fourier transform.
[0082] Preferably, S4.3 comprises the following steps:
[0083] A Grouped High-Frequency Enhancement Network (GHFEN) is designed to extract the features of each high-frequency component, fuse the features with other high-frequency components, and enhance the components. Its structure is as follows:
[0084] When the number of high-frequency components in a group is 2, HFEM is used to extract the features of each component, and then the extracted features are spliced and fused with the other component corresponding to each component, and then enhanced by a U-shaped network structure of SFFB to obtain the enhanced high-frequency components.Figure 3 .
[0085] When the number of high-frequency components in a group is 3, HFEM is used to extract the features of each component, and then the extracted features are spliced and fused with the other two components corresponding to each component, and then enhanced by a U-shaped network structure of SFFB to obtain the enhanced high-frequency components. Figure 4 .
[0086] Finally, MSE loss is used to guide GHFEN to learn, and the expression is as follows:
[0087]
[0088] N represents the number of images, M represents the number of high-frequency components in each group, and hight i,j Represents the high-frequency component extracted from a normal dose, Represents the learning results of GHFEN.
[0089] Preferably, S5 comprises the following steps:
[0090] S5.1: merging the enhanced high-frequency component with the enhanced low-frequency component;
[0091] S5.2: performing an inverse GTF transformation on the combined components so that the components are transformed back into the image domain;
[0092] S5.3: Evaluate the denoising effect of the model based on the calculated quantitative indicators PSNR, SSIM and qualitative human visual effect.
[0093] The second technical solution adopted by the present invention is: a low-dose CT denoising device based on a tight frame wavelet residual diffusion model, comprising:
[0094] Preprocessing module: used to collect data, label the data and convert it into Numpy matrix;
[0095] GTF transformation module: used to perform wavelet transformation on the image. The low-frequency component represents the image contour, and the high-frequency component represents the edge and details of the image.
[0096] Tight frame diffusion model module: used to enhance low-frequency components;
[0097] GHFEN: used to perform fusion enhancement on high-frequency components;
[0098] HFEM: used for feature extraction of high-frequency components;
[0099] SFEB: used to enhance the features of high-frequency components;
[0100] Image reconstruction module: used to perform inverse GTF transformation on the enhanced low-frequency components and high-frequency components to generate a reconstructed CT image.
[0101] The beneficial effect of the method of the present invention is as follows: the present invention first obtains 1 low-frequency component and 8 high-frequency components by performing GTF transformation on the Numpy matrix, restores and enhances the low-frequency components using a tight frame diffusion model, groups the high-frequency components and fuses and enhances them respectively, and finally merges the components and performs inverse GTF transformation to obtain the denoised result. At the same time, the present invention can effectively cope with the challenges commonly encountered by existing neural network models in CT denoising. On the one hand, the tight frame wavelet transform is combined with the deep learning network, which not only retains the advantages of the wavelet transform in multi-scale feature extraction and frequency analysis, but also utilizes the powerful nonlinear modeling ability of deep learning to process complex image features and noise patterns; on the other hand, the residual is introduced into the diffusion model to guide the denoising of the diffused image, thereby improving the certainty of denoising and reducing the possibility of extra interference information after denoising, which is beneficial to the noise removal of CT images. BRIEF DESCRIPTION OF THE DRAWINGS
[0102] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the prior art and the drawings required for use in the embodiments. The following drawings are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0103] Figure 1 It is a flow chart of a low-dose CT denoising method based on a tight frame wavelet residual diffusion model of the present invention;
[0104] Figure 2 It is a schematic diagram of a specific embodiment of a low-dose CT denoising method based on a tight frame wavelet residual diffusion model of the present invention;
[0105] Figure 3 It is a schematic diagram of the structure of a three-branch GHFEN module of a low-dose CT denoising method based on a tight frame wavelet residual diffusion model of the present invention;
[0106] Figure 4 It is a schematic diagram of the structure of a double-branch GHFEN module of a low-dose CT denoising method based on a tight frame wavelet residual diffusion model of the present invention;
[0107] Figure 5 It is a schematic diagram of the HFEM module structure of a low-dose CT denoising method based on a tight frame wavelet residual diffusion model of the present invention;
[0108] Figure 6It is a schematic diagram of the structure of the SFEB module of a low-dose CT denoising method based on a tight frame wavelet residual diffusion model of the present invention; Specific implementation plan
[0109] In order to make the purpose, features and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. It should be noted that the following detailed descriptions are illustrative and are intended to provide further explanation of the present application. Unless otherwise specified, all other embodiments obtained by ordinary technicians in this field without creative work based on the embodiments of the present invention are within the scope of protection of the present invention.
[0110] The embodiment of the present invention provides a method for LDCT denoising based on a tight frame wavelet residual diffusion model, which is used to reduce the noise level caused by photon starvation in CT at the algorithm level, thereby improving the visualization effect of low-dose CT.
[0111] A typical embodiment of the present invention, taking an animal data set as an example, refers to Figure 1 , Figure 2 , the method comprises the following steps:
[0112] S1: Collect low-dose CT images and conventional-dose CT images to construct CT data sets and convert them into Numpy matrix data;
[0113] S2: Perform GTF (Geometric Tight Framelet) transformation on the image data to obtain 1 low-frequency component and 8 high-frequency components. The high-frequency components are grouped, and the number of high-frequency components in each group is 2, 3, and 3;
[0114] S3: Create a tight frame diffusion model to denoise the low-frequency components;
[0115] S4: Create a group high-frequency enhancement network to restore high-frequency components;
[0116] S5: Finally, the enhanced and denoised components are subjected to inverse GTF transformation to obtain a denoised image.
[0117] Further, as a preferred embodiment of the method, S1 comprises the following steps:
[0118] S1.1: Acquire CT raw data in dicom format containing normal dose and low dose;
[0119] Further, the step of acquiring the DICOM format CT raw data including normal dose and low dose specifically includes:
[0120] Animal data sets were acquired at a source voltage of 100 kVp and a slice thickness of 0.625 mm, with tube current adjusted from 300 mAs down to 15 mAs.
[0121] S1.2. Label the normal-dose CT in the CT raw data with the label “target” and the low-dose CT with the label “input”;
[0122] Furthermore, the step of marking the normal-dose CT in the CT raw data with a "target" label and the low-dose CT with an "input" label specifically includes:
[0123] The CT images acquired under the imaging conditions of tube current 300mAs and source voltage 100kVp were labeled with the label “target”;
[0124] The CT images acquired under the imaging conditions of tube current 15mAs and source voltage 100kVp are labeled with the label “input”;
[0125] Convert normal-dose CT and low-dose CT in dicom format with labels into NPY matrix format data.
[0126] S1.3: Convert the labeled normal-dose CT and low-dose CT in dicom format into Numpy matrix format data.
[0127] Further, refer to Figure 2 , S2 includes the following steps:
[0128] S2.1: Construct GTF transform and its inverse transform filter;
[0129] S2.2: Perform GTF transformation on the Numpy matrix to obtain 1 low-frequency component and 8 high-frequency components;
[0130] S2.3: Group the transformed components, and divide the high-frequency components into three groups, with the numbers in each group being 2, 3, and 3 respectively.
[0131] Further, S2.1 includes the following steps:
[0132] Construct the basic matrices W0, W1 and W2, which will serve as the key starting modules for generating a series of complex matrix structures; the three basic matrices are in the following form:
[0133]
[0134] The size of the matrix matches the slice size of the image.
[0135] Now defined Where i, j = 0, 1, 2, Represents the Kronecker product, which is combined to generate the GTF transform filter matrix group:
[0136]
[0137]
[0138] Where D0 is a low-pass filter, and the other components are high-pass filters obtained at different angles, both of which satisfy:
[0139] D L T D H =E
[0140] D L Denotes a low-pass filter, D H represents a high-pass filter and E represents the identity matrix.
[0141] Further, S2.2 includes the following steps:
[0142] Given a single-channel preprocessed Numpy image data matrix X, after GTF transformation, 1 low-frequency component and 8 high-frequency components are obtained. The description of GTF transformation is as follows:
[0143] [X L ,X H ]=Conv([D L ,D H ],X)
[0144] X L is the low frequency component, X H is the high frequency component. Now define the inverse transform of GTF transform as IGTF, which is described as follows:
[0145] X'=Conv(Transposed[D L ,D H ],[X L ,X H ])
[0146] X' is the Numpy matrix of the image output after IGTF, Conv is the convolution operation, and Transposed is the transposition operation.
[0147] Further, S2.3 includes the following steps:
[0148] According to the connection between high-frequency components, the transformation components with the same first digit of the basic matrix subscript number obtained by Kronecker product are grouped into one group, and the grouping is as follows:
[0149] D1 and D2 are grouped together; D4, D5 and D6 are grouped together; D7, D8 and D9 are grouped together.
[0150] Further, refer to Figure 2 , S3 includes the following steps:
[0151] S3.1: Introduce a tight frame diffusion model to perform noise diffusion and residual diffusion on low-frequency components;
[0152] S3.2: Construct residual prediction network and noise prediction network;
[0153] S3.3: Perform reverse sampling on the diffused degraded image to gradually remove the residuals and noise in the components.
[0154] Further, S3.1 includes the following steps:
[0155] Order I t-1 is the low-frequency component of the previous state, I t is the low-frequency component in the current state, T is the total time step of diffusion, and the diffusion process is as follows:
[0156] I t =I t-1 +α t I res +β t ∈ t-1
[0157]
[0158] In the above formula, I res is the low-frequency component of the low-dose image after GTF transformation I in The difference between the low-frequency component I0 and the normal dose image after GTF transformation (i.e., I res =I in -I0). α t and β t Control residual diffusion and noise diffusion, t∈{1,…,T}, ∈ t-1 ,…,∈~N(0,1). When t=T,
[0159] Further, S3.2 includes the following steps:
[0160] Use U-Net network to build residual prediction network and the noise prediction network ∈ θ (I t ,t,I in ), use the L1 loss function to guide the learning of the network, the loss functions of the two networks are as follows:
[0161]
[0162] L ∈ (θ): = E[λ ∈ |∈-∈ θ (I t ,t,I in )| 2 ]
[0163] The residual loss is L res (θ), the noise loss is L ∈ (θ). λ res ,λ ε ∈{0,1}, when λ res =1 and λ ε When =1, both residual and noise are predicted.
[0164] Further, S3.3 includes the following steps:
[0165] I at each time step t It has been derived from the diffusion process. Subtracting them from the trained prediction network can get the reverse sampling process as follows:
[0166]
[0167] Finally, the restored low-frequency component and the low-frequency component of the normal dose after GTF transformation are subjected to MSE loss:
[0168]
[0169] is the low-frequency component obtained through the model, low is the low-frequency component obtained through GTF transformation of normal dose, L low is the MSE loss of the two, and N is the number of images.
[0170] Further, refer to Figure 3 , Figure 4 and Figure 5 , S4 comprises the following steps:
[0171] S4.1: Construct feature extraction module for high frequency components;
[0172] S4.2: Construct a spatial frequency enhancement module for high-frequency components;
[0173] S4.3: Construct a high-frequency feature enhancement fusion module to enhance each high-frequency group component;
[0174] Further, S4.1 includes the following steps:
[0175] Design a feature extraction module High-frequency Feature Extract Module (HFEM). This network module has four branches. The first three branches are 1×1, 3×3 and 5×5 multi-scale convolutions used to extract the features of high-frequency components. The last branch is a learnable Sobel operator Learnable Sobel Operator (LSO) used to extract the edge information of high-frequency components. Then the features extracted by the shallow network are flattened and normalized. It is described as follows:
[0176]
[0177] f e =Relu(Conv 3×3 (LSO(x)))
[0178]
[0179] f o =Flatten(LayerNorm(f m ))
[0180] Relu(·) represents the Relu activation function, Conv n×n (·) represents a convolution kernel of size n, represents the concatenation operation, LSO(·) represents the edge detection algorithm, Flatten(·) represents the linear representation of flattening the two-dimensional space into one dimension, LayerNorm(·) represents layer normalization, and f g represents the local features extracted from multiple scales, f e represents the edge features extracted by the learnable Sobel operator, f m It is a combination of local features and edge features. Finally, f m Flatten and normalize to get the output f of the shallow network o .
[0181] Further, S4.2 includes the following steps:
[0182] Construct a spatial-frequency enhancement module Spatial-Frequency Enhance Block (SFEB), use the state space model State Spatial Model (SSM) and convolution operation to extract the characteristics of the component space domain, which is described as follows:
[0183] f ssm1 =SiLU(x)
[0184] f ssm2=DWConv(SiLU(SSM(LayerNorm(x))))
[0185]
[0186] SiLU(·) represents SiLU activation function, LayerNorm(·) represents layer normalization, SSM(·) represents spatial state model, DWConv(·) represents 1×1 separable depthwise convolution, It is a matrix point-by-point multiplication operation.
[0187] The phase spectrum and amplitude obtained by two-dimensional Fourier transform of the high-frequency component are convolved to obtain the characteristics of the frequency domain, which is described as follows:
[0188] f input =Conv 1×1 (x)
[0189] Amp,Phase=FFT(f input )
[0190] f amp =Conv 3×3 (Amp)
[0191] f phase =Conv 3×3 (Phase)
[0192] f fft =IEFFT(f amp ,f phase )
[0193] Conv n×n It represents an n-dimensional convolution operation, FFT represents a two-dimensional Fourier transform, Amp and Phase represent amplitude and phase spectra, and IFFT represents the inverse two-dimensional Fourier transform.
[0194] Further, S4.3 includes the following steps:
[0195] A Grouped High-Frequency Enhancement Network (GHFEN) is designed to extract the features of each high-frequency component, fuse the features with other high-frequency components, and enhance the components. Its structure is as follows:
[0196] When the number of high-frequency components in a group is 2, HFEM is used to extract the features of each component, and then the extracted features are spliced and fused with the other component corresponding to each component, and then enhanced by a U-shaped network structure of SFFB to obtain the enhanced high-frequency components. Figure 3 .
[0197] When the number of high-frequency components in a group is 3, HFEM is used to extract the features of each component, and then the extracted features are spliced and fused with the other two components corresponding to each component, and then enhanced by a U-shaped network structure of SFFB to obtain the enhanced high-frequency components. Figure 4 .
[0198] Finally, MSE loss is used to guide GHFEN to learn, and the expression is as follows:
[0199]
[0200] N represents the number of images, M represents the number of high-frequency components in each group, and high i,j Represents the high-frequency component extracted from a normal dose, Represents the learning results of GHFEN.
[0201] Further, S5 comprises the following steps:
[0202] S5.1: merging the enhanced high-frequency component with the enhanced low-frequency component;
[0203] S5.2: performing an inverse GTF transformation on the combined components so that the components are transformed back into the image domain;
[0204] S5.3: Evaluate the denoising effect of the model based on the calculated quantitative indicators PSNR, SSIM and qualitative human visual effect.
[0205] The above is a specific description of the preferred implementation of the present invention, but the invention is not limited to the embodiments. Those skilled in the art may make other equivalent modifications or substitutions without violating the spirit of the invention, and these equivalent modifications or substitutions are included in the scope defined by the application claims.
Claims
1. A low-dose CT denoising method based on a tight frame wavelet residual diffusion model, characterized in that: The method comprises the following steps: S1: Collect low-dose CT images and conventional-dose CT images to construct CT data sets and convert them into Numpy matrix data; S2: Perform GTF (Geometric Tight Framelet) transformation on the image data to obtain 1 low-frequency component and 8 high-frequency components. The high-frequency components are grouped, and the number of high-frequency components in each group is 2, 3, and 3; S3: Create a tight frame diffusion model to denoise the low-frequency components; S4: Create a group high-frequency enhancement network to restore high-frequency components; S5: Finally, the enhanced and denoised components are subjected to inverse GTF transformation to obtain a denoised image.
2. The low-dose CT denoising method based on a tight frame wavelet residual diffusion model according to claim 1, characterized in that: S1 includes the following steps: S1.1: Acquire CT raw data in dicom format containing normal dose and low dose; S1.2: Label the normal-dose CT in the CT raw data as "target" and the low-dose CT as "input"; S1.3: Convert the labeled normal-dose CT and low-dose CT in dicom format into NPY matrix format data.
3. The low-dose CT denoising method based on a tight frame wavelet residual diffusion model according to claim 1, characterized in that: S2 includes the following steps: S2.1: Construct GTF transform and its inverse transform filter; S2.2: Perform GTF transformation on the Numpy matrix to obtain 1 low-frequency component and 8 high-frequency components; S2.3: Group the transformed components, and divide the high-frequency components into three groups, with the numbers of components in each group being 2, 3, and 3 respectively; Among them, S2.1 includes the following steps: Construct the basic matrices W0, W1 and W2, which will serve as the key starting modules for generating a series of complex matrix structures; the three basic matrices are in the following form: The size of the matrix is consistent with the slice size of the image; Current definition where i, j = 0, 1, 2, Represents the Kronecker product, which is combined to generate the GTF transform filter matrix group: D0=W 0,0 , D5=W 1,2 D7=W 2,1 , Where D0 is a low-pass filter, and the other components are high-pass filters obtained at different angles, both of which satisfy: D L T D H =E D L Denotes a low-pass filter, D H represents a high-pass filter, and E represents the identity matrix; Among them, S2.2 includes the following steps: Given a single-channel preprocessed Numpy image data matrix X, after GTF transformation, 1 low-frequency component and 8 high-frequency components are obtained. The description of GTF transformation is as follows: [X L ,X H ]=Conv([D L ,D H ],X) X L is the low frequency component, X H It is the high frequency component. Now we define the inverse transform of GTF transform as IGTF, which is described as follows: X′=Conv(Transposed[D L ,D H ],[X L ,X H ]) X′ is the Numpy matrix of the image output after IGTF, Conv is the convolution operation, and Transposed is the transposition operation; Among them, S2.3 includes the following steps: According to the connection between high-frequency components, the transformation components with the same first digit of the basic matrix subscript number obtained by Kronecker product are grouped into one group, and the grouping is as follows: D1 and D2 are grouped together; D4, D5 and D6 are grouped together; D7, D8 and D9 are grouped together.
4. The low-dose CT denoising method based on a tight frame wavelet residual diffusion model according to claim 1, characterized in that: S3 includes the following steps: S3.1: Introduce a tight frame diffusion model to perform noise diffusion and residual diffusion on low-frequency components; S3.2: Construct residual prediction network and noise prediction network; S3.3: reverse sampling the diffused degraded image to gradually remove the residual and noise in the components; Among them, S3.1 includes the following steps: Order I t-1 is the low-frequency component of the previous state, I t is the low-frequency component in the current state, T is the total time step of diffusion, and the diffusion process is as follows: In the above formula, I res is the low-frequency component of the low-dose image after GTF transformation I in The difference between the low-frequency component I0 and the normal dose image after GTF transformation (i.e., I res =I in -I0). α t and β t Control residual diffusion and noise diffusion, t∈{1,...,T}, ∈ t-1 , ..., ∈~N(0,1), when t=T, Among them, S3.2 includes the following steps: Use U-Net network to build residual prediction network and the noise prediction network ∈ θ (I t ,t,I in ), use the L1 loss function to guide the learning of the network, and the loss functions of the two networks are as follows: L ∈ (θ):=E[λ ∈ |∈-∈ θ (I t ,t,I in )| 2 ] The residual loss is L res (θ), the noise loss is L ∈ (θ), λ res ,λ ε ∈{0,1}, when λ res =1 and λ ε =1, both residual and noise are predicted; Among them, S3.3 includes the following steps: I at each time step t It has been derived from the diffusion process. Subtracting them from the trained prediction network can get the reverse sampling process as follows: Finally, the restored low-frequency component and the low-frequency component of the normal dose after GTF transformation are subjected to MSE loss: is the low-frequency component obtained through the model, low is the low-frequency component obtained through GTF transformation of normal dose, L low is the MSE loss of the two, and N is the number of images.
5. The low-dose CT denoising method based on a tight frame wavelet residual diffusion model according to claim 1, characterized in that: S4 includes the following steps: S4.1: Construct feature extraction module for high frequency components; S4.2: Construct a spatial frequency enhancement module for high-frequency components; S4.3: construct a group high frequency enhancement network to enhance each high frequency group component; Among them, S4.1 includes the following steps: A feature extraction module High-frequency Feature Extract Module (HFEM) is designed. This network module has four branches. The first three branches are 1×1, 3×3 and 5×5 multi-scale convolutions used to extract the features of high-frequency components. The last branch is a learnable Sobel operator Learnable Sobel Operator (LSO) used to extract the edge information of high-frequency components. The features extracted by the shallow network are then flattened and normalized. Its description is as follows: f e =Relu(Conv 3×3 (LSO(x))) f o =Flatten(LayrerNorm(f m )) Relu(·) represents the Relu activation function, Conv n×n (·) represents a convolution kernel of size n, represents the concatenation operation, LSO(·) represents the edge detection algorithm, Flatten(·) represents the linear representation of flattening the two-dimensional space into one dimension, LayerNorm(·) represents layer normalization, and f g represents the local features extracted from multiple scales, f e represents the edge features extracted by the learnable Sobel operator, f m It is a combination of local features and edge features. Finally, f m Flatten and normalize to get the output f of the shallow network o ; Among them, S4.2 includes the following steps: Construct a spatial-frequency enhancement module Spatial-Frequency Enhance Block (SFEB), use the state space model State SpatialModel (SSM) and convolution operation to extract the characteristics of the component space domain, which is described as follows: f ssm1 =SiLU(x) f ssm2 =DWConv(SiLU(SSM(LayerNorm(x)))) SiLU(·) represents SiLU activation function, LayerNorm(·) represents layer normalization, SSM(.) represents spatial state model, DWConv(·) represents 1×1 separable depthwise convolution, It is a matrix point-by-point multiplication operation; The phase spectrum and amplitude obtained by two-dimensional Fourier transform of the high-frequency component are convolved to obtain the characteristics of the frequency domain, which can be described as follows: f input =Conv 1×1 (x) Amp,Phase=FFT(f input ) f amp =Conv 3×3 (Amp) f phase =Conv 3×3 (Phase) f fft =IFFT(f amp , f phase ) Conv n×n Represents an n-dimensional convolution operation, FFT represents a two-dimensional Fourier transform, Amp and Phase represent amplitude and phase spectra, and IFFT represents an inverse two-dimensional Fourier transform; Among them, S4.3 includes the following steps: A Grouped High-Frequency Enhancement Network (GHFEN) is designed to extract the features of each high-frequency component, fuse the features with other high-frequency components, and enhance the components. Its structure is as follows: When the number of high-frequency components in a group is 2, HFEM is used to extract the features of each component, and then the extracted features are spliced and fused with the other component corresponding to each component, and then enhanced by a U-shaped network structure of SFFB to obtain the enhanced high-frequency component. The detailed structure is shown in Figure 3; When the number of high-frequency components in a group is 3, HFEM is used to extract the features of each component, and then the extracted features are spliced and fused with the other two components corresponding to each component, and then enhanced by a U-shaped network structure of SFFB to obtain the enhanced high-frequency components. The detailed structure is shown in Figure 4; Finally, MSE loss is used to guide GHFEN to learn, and the expression is as follows: N represents the number of images, M represents the number of high-frequency components in each group, and high i,j Represents the high-frequency component extracted from a normal dose, Represents the learning results of GHFEN.
6. The low-dose CT denoising method based on a tight frame wavelet residual diffusion model according to claim 1, characterized in that: S5 includes the following steps: S5.1: merging the enhanced high-frequency component with the enhanced low-frequency component; S5.2: performing an inverse GTF transformation on the combined components so that the components are transformed back into the image domain; S5.3: Evaluate the denoising effect of the model based on the calculated quantitative indicators PSNR, SSIM and qualitative human visual effect.
7. The low-dose CT denoising device based on the tight frame wavelet residual diffusion model according to claim 1, characterized in that: include: Preprocessing module: used to collect data, label the data and convert it into Numpy matrix; GTF transformation module: used to perform wavelet transformation on the image. The low-frequency component represents the image contour, and the high-frequency component represents the edge and details of the image. Tight frame diffusion model module: used to enhance low-frequency components; HFEM: used to extract high-frequency component features; SFFB: used to enhance the spatial and frequency domains of high-frequency components; GHFEN: used to perform fusion enhancement between high-frequency components; Image reconstruction module: used to perform inverse GTF transformation on the enhanced low-frequency components and high-frequency components to generate a reconstructed CT image.
Citation Information
Patent Citations
Medical CT image super-resolution algorithm based on dual-frequency-domain denoising enhancement
CN116777749A
Unsupervised low-dose CT (Computed Tomography) denoising model training method, denoising method and device
CN117094902A
Image shadow removal method based on wavelet non-uniform diffusion model
CN118134800A
Two-stage low-dose CT denoising method based on residual diffusion model and medium
CN118396881A
Method and apparatus for low-dose x-ray computed tomography image processing based on efficient unsupervised learning using invertible neural network
US20220414954A1
Cited By
FFA image generation system and method based on double-domain constraint Mama diffusion model, medium and device
CN121095098A