CT image denoising method and device based on prior knowledge guidance
Through the CT image denoising method guided by prior knowledge, the problem of difficulty in retention of inter-tissue hierarchical relationships and boundary details in low-dose CT image denoising is solved, and more efficient denoising performance and image quality are achieved, assisting more accurate diagnosis and treatment.
Patent Information
- Application Number
- CN202510117079.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
AI Technical Summary
Existing low-dose CT image denoising methods are difficult to effectively retain the hierarchical relationships and boundary details between tissues, resulting in the image being too smooth and lacking details, which affects diagnostic and treatment decisions.
The CT image denoising method based on prior knowledge is adopted, and the low-dose CT image data is acquired and preprocessed, and converted into a prior mask image is converted using a discrete encoding and decoding network, and the denoising network is introduced through the knowledge fusion module to construct a negative sample set and joint loss function to improve the denoising performance.
It significantly improves the denoising performance of low-dose CT images, retains clearer tissue boundaries, and assists doctors in more accurate diagnosis and treatment.
Smart Images

Figure CN120047346A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a CT image denoising method and device guided by prior knowledge, belonging to the technical field of computer vision processing. Background Art
[0002] Computed tomography (CT) is a technology widely used in disease screening and diagnosis, with significant advantages such as non-invasiveness and high-resolution imaging. However, the ionizing radiation generated during CT scanning may pose a certain risk to the health of patients. To reduce the radiation risk, low-dose CT technology has gradually attracted attention. However, this method inevitably introduces more noise and artifacts when acquiring image data, significantly hindering the accurate quantitative assessment of body tissues (such as cerebral hemorrhage, abdominal stones, etc.), which may thus impose important limitations on diagnosis, pathological monitoring, and treatment decisions.
[0003] In this context, reconstructing low-dose CT images into normal-dose CT images has become an urgent problem to be solved. Early image denoising methods mainly relied on manually designed regularization terms. With the rapid development of deep learning, researchers began to adopt an encoder-decoder structure with residual connections and optimize it through mean squared error loss. Although such structures have significantly improved the denoising performance in low-dose CT image denoising, they often result in overly smooth images lacking details. To solve this problem, researchers have tried to use a learnable Sobel operator to extract tissue edge details to guide the model to pay more attention to the boundary information of organs, thereby alleviating the over-smoothing phenomenon. In recent years, methods based on the Transformer architecture have gradually been applied to the field of low-dose CT denoising due to their excellent long-distance dependence capture ability. In addition, as a novel likelihood-based generative model, the diffusion probability model has begun to attract attention in low-dose CT denoising due to its powerful pattern coverage ability.
[0004] Although the above algorithms have made significant progress in the field of low-dose CT denoising, they still face some challenges. First, medical images usually contain complex structures, such as fine blood vessels, tissues, and boundaries. Although existing methods have introduced a trainable Sobel operator to improve the boundary retention ability or used a diffusion probability model to model the real distribution, these methods have failed to effectively guide the generation of hierarchical relationships between tissues. Second, existing models usually optimize the objective through mean squared error loss or combined loss to reduce the difference between the reconstructed image and the real image. However, their ability to approximate normal-dose CT images is still limited, and this limitation often leads to overly smooth images and the loss of detailed textures. Summary of the Invention
[0005] To solve the above problems, the present invention proposes a CT image denoising method and device guided by prior knowledge, which can improve the denoising performance of low-dose CT images and retain clearer tissue boundaries, thereby assisting doctors in more accurate diagnosis and treatment.
[0006] The technical solution adopted by the present invention to solve its technical problems is as follows:
[0007] In a first aspect, a CT image denoising method guided by prior knowledge provided by an embodiment of the present invention includes the following steps:
[0008] Step S1, obtain low-dose CT image data and perform preprocessing;
[0009] Step S2, convert the preprocessed low-dose CT image into a prior mask image;
[0010] Step S3, process the prior mask image through a discrete encoding-decoding network to obtain an enhanced prior mask image;
[0011] Step S4, introduce the enhanced prior mask into the denoising network through a knowledge fusion module to obtain a denoising model;
[0012] Step S5, construct a negative sample set and propose a joint loss function;
[0013] Step S6, obtain a conventional-dose CT image paired with the low-dose CT image, add Gaussian noise and splice it with the low-dose CT image on the channel to obtain a prior mask image of the test image;
[0014] Step S7, input the prior mask image of the test image into the denoising model to output the final denoised CT image.
[0015] As a possible implementation manner of this embodiment, the step S1 includes the following steps:
[0016] Step S11, collect low-dose CT image data by using a CT scanning method, and define the low-dose CT image as I LD ∈R C ×H×W , and the conventional-dose CT image I ND ∈R C×H×W , where C represents the image dimension, H represents the height of the image, and W represents the width of the image;
[0017] Step S12, perform preprocessing operations on the collected low-dose CT image:
[0018] S LD (c,h,w)=clip(I LD ,-1024,3072)
[0019]
[0020] S ND (c, h, w) = clip(I ND , -1024, 3072)
[0021]
[0022] where clip(·) represents truncating the HU values of CT images, and S LD (c, h, w) and I LD (c, h, w) are the pixel values of the low-dose CT image (c, h, w) respectively, representing the maximum and minimum values of all pixel specific values in the low-dose CT image respectively, and S ND (c, h, w) and I ND (c, h, w) are the pixels of the conventional-dose CT image (c, h, w) respectively, representing the maximum and minimum values of all pixel specific values in the conventional-dose CT image respectively;
[0023] Step S13, the HU value range of the preprocessed low-dose CT image is intercepted as [-1024, 3072] and normalized to between [0, 1].
[0024] As a possible implementation manner of this embodiment, the said step S2 includes the following steps:
[0025] Step S21, input the preprocessed low-dose CT image into the pre-trained segmentation large model based on the VIT-H architecture to generate multiple mask images of the low-dose CT image I LD ;
[0026] Step S22, screen the generated mask images and filter out the masks with a total area less than 50 pixels:
[0027] M′ = {m i ∈M|A(m i ) ≥ 50}
[0028] where M′ is the set of effective mask images after screening, m i is the i-th mask image in the set, and A(m i ) represents the area of the mask image m i ∈R 1×H×W ;
[0029] Step S23, add all the screened effective mask images at the channel level and normalize them to the range of [0, 1] to obtain the final mask image m ∈ R 1×H×W :
[0030] m' = Add(m 1 , m 2 , …, m n )
[0031]
[0032] where Add(·) represents adding the pixel values of the effective mask image in the channel dimension, m' min , m' max represent the minimum and maximum values of all pixel values in the concatenated mask image respectively, and n represents the number of mask images generated by SAM (Split-Attention Mechanism).
[0033] As a possible implementation of this embodiment, step S3 includes the following steps:
[0034] Step S31: Input the obtained prior mask image into the discrete encoding network to obtain the latent space feature representation m z :
[0035] m z = VQEncoder(m)
[0036] where VQEncoder(·) represents the discrete encoding network respectively;
[0037] Step S32: Use the nearest neighbor algorithm to map the latent space feature representation m z of the prior mask image to the codebook vector E = [e 1 , e 2 , …, e K to obtain the discretized feature representation m q :
[0038] m q = e k , where k = argmin j ||m z - e j || 2 ;
[0039] Step S33: Obtain the enhanced prior mask image through the discrete decoding network
[0040]
[0041] where VQDecoder(·) represents the discrete decoding network respectively;
[0042] Step S34, construct the loss function of the discrete encoding and decoding network:
[0043]
[0044] where sg[·] represents the gradient clipping operation, and m z represents the feature vector obtained by encoding m through VQDecoder(·). Optimize the model parameters through the gradient descent algorithm until the model converges.
[0045] As a possible implementation of this embodiment, the step S4 includes the following steps:
[0046] Step S41, perform a linear mapping on the obtained discretized prior mask and the deep feature z in the denoising network latent :
[0047] Q = z latent W q , where W q , W k , W v are linear transformation matrices respectively;
[0048] Step S42, calculate the cross-attention activation map between the query Q and the key K, and weight it into the value V encoding vector:
[0049]
[0050] where CA(·) represents the cross-attention mechanism, and softmax(·) represents an activation function that normalizes a numerical vector into a probability distribution, represents the dimension of the encoding vector;
[0051] Step S43, perform layer normalization LN on the feature representation weighted by the cross-attention mechanism, add it to the time encoding t, and then input the result into the feed-forward neural network FFN and the self-attention module MHSA to obtain the feature representation o after fusing prior knowledge:
[0052] o = MHSA(FFN(LN(CA) + t)).
[0053] As a possible implementation of this embodiment, the specific process of constructing the negative sample set is as follows:
[0054] Denote the conventional-dose CT image as x 0 , and construct a negative sample set by using the method of randomly adding Gaussian noise and mean filtering:
[0055] n(x 0 ) = x 0 + ∈
[0056]
[0057] Among them, n(·) uniformly represents the method of constructing the negative sample set, ∈ is Gaussian noise, which follows the standard Gaussian distribution, k represents the size of the filter, and x 0 (r, c) represents the pixel value of the image at the position (r, c), and (i, j) is the position of the filter.
[0058] As a possible implementation manner of this embodiment, the joint loss function is:
[0059]
[0060] L total = L diff + λL r-cl
[0061] Among them, L total is the total loss value, L r-cl is the denoising network loss, L diff is the discrete encoding and decoding network loss, λ is a coefficient, y is the preprocessed low-dose CT test image I LD , the time encoding t represents an integer within the range of [0, 1000], and m q is the discrete prior mask.
[0062] As a possible implementation manner of this embodiment, the step S6 includes the following steps:
[0063] Step S61, obtain the low-dose CT test image and perform preprocessing;
[0064] Step S62, randomly sample the Gaussian noise ∈, and gradually increase the Gaussian noise for the conventional-dose CT image I ND to generate a new data sample x t :
[0065]
[0066] Among them, ∈ represents Gaussian noise, and α t represents a gradually decreasing parameter used to control the intensity of the noise, and the time encoding t represents an integer within the range of [0, 1000];
[0067] Step S63, use the preprocessed low-dose CT test image I LD as the condition y, and perform dimensional concatenation on the channels with the new data sample x t ;
[0068] Step S64, perform conversion and enhancement processing on the concatenated test image to obtain the discrete prior mask representation mq .
[0069] As a possible implementation of this embodiment, step S7 includes the following steps:
[0070] Step S71, obtaining the discrete prior mask representation m q and introducing it into the denoising network;
[0071] Step S72, obtaining the denoised image by iterating the denoising network multiple times:
[0072]
[0073] where ∈ θ (·) is the output of the denoising network, is the discretized prior mask, x 0 is the conventional-dose CT image, x t is the generated data sample, y is the preprocessed low-dose CT test image I LD , the time encoding t starts from 1000 and decreases sequentially in steps of 100 until t = 1.
[0074] In a second aspect, a CT image denoising device based on prior knowledge guidance provided by an embodiment of the present invention includes:
[0075] An image acquisition module, configured to acquire low-dose CT image data and perform preprocessing;
[0076] An image conversion module, configured to convert the preprocessed low-dose CT image into a prior mask image;
[0077] An image enhancement module, configured to process the prior mask image through a discrete encoding-decoding network to obtain an enhanced prior mask image;
[0078] A model acquisition module, configured to introduce the enhanced prior mask into the denoising network through a knowledge fusion module to obtain a denoising model;
[0079] A negative sample construction module, configured to construct a negative sample set and propose a joint loss function;
[0080] An image splicing module, configured to acquire a conventional-dose CT image paired with the low-dose CT image, add Gaussian noise and splice it with the low-dose CT image on the channel to obtain the prior mask image of the test image;
[0081] An image denoising module, configured to input the prior mask image of the test image into the denoising model and output the final denoised CT image.
[0082] The beneficial effects of the technical solution of the embodiment of the present invention are as follows:
[0083] A CT image denoising method guided by prior knowledge according to the technical solution of the embodiment of the present invention includes the following steps: Step S1, obtaining low-dose CT image data and performing preprocessing; Step S2, converting the preprocessed low-dose CT image into a prior mask image; Step S3, processing the prior mask image through a discrete encoding-decoding network to obtain an enhanced prior mask image; Step S4, introducing the enhanced prior mask into a denoising network through a knowledge fusion module to obtain a denoising model; Step S5, constructing a negative sample set and proposing a joint loss function; Step S6, obtaining a conventional-dose CT image paired with the low-dose CT image, adding Gaussian noise and splicing it with the low-dose CT image on the channel to obtain a prior mask image of the test image; Step S7, inputting the prior mask image of the test image into the denoising model to output the final denoised CT image. The present invention significantly improves the denoising performance of low-dose CT images, retains clearer tissue boundaries, thereby assisting doctors in more accurate diagnosis and treatment. The present invention models the hierarchical relationship between tissues through the mask generated by SAM, introduces this hierarchical prior information into the diffusion model, guides the model to generate clearer tissue boundaries, and also proposes a simple and efficient joint optimization loss to narrow the gap between the reconstructed image and the real image. The present invention demonstrates excellent performance and theoretical advantages in CT image denoising, significantly improving the quality of the reconstructed image.
[0084] The present invention proposes a controllable diffusion model based on SAM prior knowledge, which is specifically used for low-dose CT denoising. Thanks to the prior knowledge provided by SAM, the present invention significantly improves the effect of low-dose CT denoising in clinical application scenarios, generates images with clearer boundaries, and the noise is more easily suppressed.
[0085] A CT image denoising device guided by prior knowledge according to the technical solution of the embodiment of the present invention has the same beneficial effects as a CT image denoising method guided by prior knowledge according to the technical solution of the embodiment of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0086] Figure 1 is a flowchart of a CT image denoising method guided by prior knowledge shown according to an exemplary embodiment;
[0087] Figure 2 is a schematic structural diagram of a CT image denoising device guided by prior knowledge shown according to an exemplary embodiment;
[0088] Figure 3 is a specific implementation flowchart of the present invention for low-dose CT image denoising. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0089] To more clearly illustrate the technical features of the solution of the present invention, the present invention will be described in detail below through specific embodiments and in conjunction with its accompanying drawings.
[0090] As Figure 1 shown, a CT image denoising method based on prior knowledge guidance provided by an embodiment of the present invention includes the following steps:
[0091] Step S1, obtaining low-dose CT image data and performing preprocessing;
[0092] Step S2, converting the preprocessed low-dose CT image into a prior mask image;
[0093] Step S3, processing the prior mask image through a discrete encoding-decoding network to obtain an enhanced prior mask image;
[0094] Step S4, introducing the enhanced prior mask into the denoising network through a knowledge fusion module to obtain a denoising model;
[0095] Step S5, constructing a negative sample set and proposing a joint loss function;
[0096] Step S6, obtaining a conventional-dose CT image paired with the low-dose CT image, adding Gaussian noise and splicing it with the low-dose CT image on the channel to obtain a prior mask image of the test image;
[0097] Step S7, inputting the prior mask image of the test image into the denoising model to output the final denoised CT image.
[0098] As a possible implementation manner of this embodiment, the step S1 includes the following steps:
[0099] Step S11, collecting low-dose CT image data by means of CT scanning and defining the low-dose CT image as I LD ∈R C ×H×W , and the conventional-dose CT image I ND ∈R C×H×W , where C represents the image dimension, H represents the height of the image, and W represents the width of the image;
[0100] Step S12, performing preprocessing operations on the collected low-dose CT image:
[0101] S LD (c, h, w) = clip(I LD , -1024, 3072)
[0102]
[0103] S ND(c, h, w) = clip(I ND , -1024, 3072)
[0104]
[0105] where clip(·) represents truncating the HU value of the CT image, S LD (c, h, w) and I LD (c, h, w) are the pixel values of the low-dose CT image (c, h, w) respectively, representing the maximum and minimum values of all pixel specific values in the low-dose CT image respectively, S ND (c, h, w) and I ND (c, h, w) are the pixels of the conventional-dose CT image (c, h, w) respectively, representing the maximum and minimum values of all pixel specific values in the conventional-dose CT image respectively;
[0106] Step S13, the HU value range of the preprocessed low-dose CT image is intercepted as [-1024, 3072] and normalized to between [0, 1].
[0107] As a possible implementation of this embodiment, step S2 includes the following steps:
[0108] Step S21, input the preprocessed low-dose CT image into the pre-trained SAM model based on the VIT-H architecture to generate multiple mask images of the low-dose CT image I LD ;
[0109] Step S22, screen the generated mask images and filter out the masks with a total area less than 50 pixels:
[0110] M′ = {m i ∈M | A(m i ) ≥ 50}
[0111] where M′ is the set of valid mask images after screening, m i is the i-th mask image in the set, and A(m i ) represents the area of the mask image m i ∈R 1×H×W ;
[0112] Step S23, add all the screened valid mask images at the channel level and normalize them to the range of [0, 1] to obtain the final mask image m ∈ R 1×H×W :
[0113] m′ = Add(m 1 , m 2 , …, mn )
[0114]
[0115] Among them, Add(·) represents adding the pixel values of the effective mask image in the channel dimension, and m′ min , m′ max respectively represent the minimum and maximum values of all pixel values in the spliced mask image, and n represents the number of mask images generated by SAM.
[0116] As a possible implementation of this embodiment, step S3 includes the following steps:
[0117] Step S31: Input the obtained prior mask image into the discrete coding network to obtain the latent space feature representation m of the prior mask image z :
[0118] m z = VQEncoder(m)
[0119] Among them, VQEncoder(·) respectively represents the discretization coding network;
[0120] Step S32: Use the nearest neighbor algorithm to map the latent space feature representation m of the prior mask image z to the codebook vector E = [e 1 , e 2 , …, e K to obtain the discretized feature representation m q :
[0121] m q = e k , where k = argmin j ||m z - e j || 2 ;
[0122] Step S33: Obtain the enhanced prior mask image through the discrete decoding network
[0123]
[0124] Among them, VQDecoder(·) respectively represents the discrete decoding network;
[0125] Step S34: Construct the discrete coding and decoding network loss function:
[0126]
[0127] Among them, sg[·] represents the gradient truncation operation, and mz Denote the feature vector of \(m\) after being encoded by VQEncoder(·), and optimize the model parameters through the gradient descent algorithm until the model converges.
[0128] As a possible implementation manner of this embodiment, the step S4 includes the following steps:
[0129] Step S41, perform a linear mapping on the obtained discretized prior mask and the deep feature \(z\) in the denoising network: latent Denote as:
[0130] \(Q = zW^1\), latent \(W^1\) q ,
[0131] where \(W^1\), q , \(W^2\), k , \(W^3\) v are linear transformation matrices respectively;
[0132] Step S42, calculate the cross-attention activation map between the query \(Q\) and the key \(K\), and weight it into the value \(V\) encoding vector:
[0133]
[0134] where \(CA(·)\) represents the cross-attention mechanism, and \(softmax(·)\) represents an activation function that normalizes a numerical vector into a probability distribution, denotes the dimension of the encoding vector;
[0135] Step S43, perform layer normalization LN on the feature representation weighted by the cross-attention mechanism, add it to the time encoding \(t\), and then input the result into the feed-forward neural network FFN and the self-attention module MHSA to obtain the feature representation \(o\) after fusing the prior knowledge:
[0136] \(o = MHSA(FFN(LN(CA)+t))\).
[0137] As a possible implementation manner of this embodiment, the specific process of constructing the negative sample set is as follows:
[0138] Denote the conventional-dose CT image as \(x\) 0 , and construct a negative sample set by using the method of randomly adding Gaussian noise and mean filtering:
[0139] \(n(x)\) 0 \(= x + \epsilon\) 0
[0140]
[0141] Among them, n(·) uniformly represents the method of constructing the negative sample set, ∈ is Gaussian noise, which follows the standard Gaussian distribution, k represents the size of the filter, and x 0 (r, c) represents the pixel value of the image at the position (r, c), (i, j) is the position of the filter, and N is the number of negative samples.
[0142] As a possible implementation manner of this embodiment, the joint loss function is:
[0143]
[0144] L total = L diff + λL r-cl
[0145] Among them, L total is the total loss value, L r-cl is the denoising network loss, L diff is the discrete encoding and decoding network loss, λ is a coefficient, y is the preprocessed low-dose CT test image I LD , the time encoding t represents an integer within the range of [0, 1000], and m q is the discrete prior mask.
[0146] As a possible implementation manner of this embodiment, the step S6 includes the following steps:
[0147] Step S61, obtain the low-dose CT test image and perform preprocessing;
[0148] Step S62, randomly sample the Gaussian noise ∈, and gradually increase the Gaussian noise for the conventional-dose CT image I ND to generate new data samples x t :
[0149]
[0150] Among them, ∈ represents Gaussian noise, α t represents a gradually decreasing parameter used to control the intensity of the noise, and the time encoding t represents an integer within the range of [0, 1000];
[0151] Step S63, use the preprocessed low-dose CT test image I LD as the condition y, and perform dimensional concatenation on the channels with the new data samples x t ;
[0152] Step S64, perform conversion and enhancement processing on the concatenated test image to obtain the discrete prior mask representation m q .
[0153] As a possible implementation of this embodiment, step S7 includes the following steps:
[0154] Step S71: Obtain the discrete prior mask representation m q Introduce it into the denoising network;
[0155] Step S72: Obtain the denoised image by iterating the denoising network multiple times:
[0156]
[0157] where ∈ θ (·) is the output of the denoising network, is the discretized prior mask, x 0 is the conventional-dose CT image, x t is the generated data sample, y is the preprocessed low-dose CT test image I LD , the time encoding t starts from 1000 and decreases sequentially in steps of 100 until t = 1.
[0158] As Figure 2 shown, a CT image denoising device based on prior knowledge guidance provided by an embodiment of the present invention includes:
[0159] An image acquisition module, configured to acquire low-dose CT image data and perform preprocessing;
[0160] An image conversion module, configured to convert the preprocessed low-dose CT image into a prior mask image;
[0161] An image enhancement module, configured to process the prior mask image through a discrete encoding-decoding network to obtain an enhanced prior mask image;
[0162] A model acquisition module, configured to introduce the enhanced prior mask into the denoising network through a knowledge fusion module to obtain a denoising model;
[0163] A negative sample construction module, configured to construct a negative sample set and propose a joint loss function;
[0164] An image splicing module, configured to obtain a conventional-dose CT image paired with the low-dose CT image, add Gaussian noise and splice it with the low-dose CT image on the channel to obtain a prior mask image of the test image;
[0165] An image denoising module, configured to input the prior mask image of the test image into the denoising model and output the final denoised CT image.
[0166] As Figure 3 shown, the specific process of denoising low-dose CT images by the present invention is as follows.
[0167] Step 1: Obtain and preprocess the low-dose CT image data.
[0168] Step 11: Use a CT scanning device to collect image data, and define the low-dose CT image as I LD ∈R C×H×W , and the conventional-dose CT image I ND ∈R C×H×W , where C represents the image dimension, H represents the height of the image, and W represents the width of the image;
[0169] Step 12: Perform preprocessing operations on the image, intercept the HU value range of the CT image as [-1024, 3072], and normalize it to between [0, 1]; specifically, it can be expressed as:
[0170] S LD (c, h, w) = clip(I LD , -1024, 3072)
[0171]
[0172] S ND (c, h, w) = clip(I ND , -1024, 3072)
[0173]
[0174] where clip(·) represents truncating the HU value of the CT image, and S LD (c, h, w) and I LD (c, h, w) are the pixel values of the low-dose CT image (c, h, w), represent the maximum and minimum values of all pixel specific values in the low-dose CT image respectively, and S ND (c, h, w) and I ND (c, h, w) are the pixels of the conventional-dose CT image (c, h, w), represent the maximum and minimum values of all pixel specific values in the conventional-dose CT image respectively.
[0175] Step 2: Input the preprocessed image into the hierarchical prior knowledge extraction module to generate a prior mask image.
[0176] Step 21: Input the preprocessed image into a pre-trained segmentation large model based on the VIT-H architecture, and generate multiple mask images of the low-dose CT image I LD through the SamAutomaticMaskGenerator interface;
[0177] Step 22, screen the generated mask images, and filter out the masks with a total area less than 50 pixels to ensure the validity of the masks:
[0178] M′ = {m i ∈M | A(m i ) ≥ 50}
[0179] where M′ is the set of valid mask images after screening, m i is the i-th mask image in the set, and A(m i ) represents the area of the mask image m i ∈R 1×G×W ;
[0180] Step 23, add all the valid mask images at the channel level and normalize them to the range of [0, 1] to obtain the final mask image m ∈ R 1×H×W :
[0181] m′ = Add(m 1 , m 2 , …, m n )
[0182]
[0183] where Add(·) represents adding the pixel values of the valid mask images in the channel dimension, m′ min , m′ max represent the minimum and maximum values of all pixel values in the concatenated mask image respectively, and m represents the number of mask images generated by SAM.
[0184] Step 3, process the prior mask image through a discrete encoding and decoding network to obtain an enhanced prior mask image.
[0185] Step 31, input the obtained prior mask image into the discrete encoding network to obtain the latent space feature representation m z :
[0186] m z = VQEncoder(m)
[0187] where VQEncoder(·) represents the discrete encoding network respectively;
[0188] Step 32, use the nearest neighbor algorithm to map m z to the codebook vector E = [e 1 , e 2 , …, e K to obtain the discretized feature representation m q :
[0189] mq = e k , where k = argmin j ||m z - e j || 2 ;
[0190] Step 33, obtain the enhanced prior mask image through the discrete decoding network
[0191]
[0192] where VQDecoder(·) represents the discrete decoding network respectively;
[0193] Step 34, construct the loss function of the discrete encoding-decoding network:
[0194]
[0195] where sg[·] represents the gradient clipping operation, and m z represents the feature vector obtained by encoding m through VQEncoder(·). Optimize the model parameters through the gradient descent algorithm until the model converges.
[0196] Step 4, obtain the conventional-dose CT image paired with the low-dose CT image and add Gaussian noise with a certain intensity, and splice it with the low-dose CT image on the channel.
[0197] Step 41, gradually increase the Gaussian noise for the conventional-dose CT image I ND to generate new data samples x t :
[0198]
[0199] where ∈ represents the Gaussian noise, and α t represents the gradually decreasing parameter used to control the intensity of the noise, and t represents an integer in the range of [0, 1000];
[0200] Step 42, use the preprocessed low-dose CT image I LD as the condition y, and after splicing it with the data sample x obtained in Step 41 t on the channel dimension, use it as the input of the denoising network;
[0201] Step 5, introduce the enhanced prior mask into the denoising network through the knowledge fusion module.
[0202] Step 51, perform a linear mapping on the obtained discrete prior mask and the deep feature z in the denoising network latent denoted as:
[0203] Q = z latent W q ,
[0204] where W q , W k , W v are linear transformation matrices respectively;
[0205] Step 52, calculate the cross-attention activation map between the query Q and the key K, and weight it into the value V encoding vector:
[0206]
[0207] where CA(·) represents the cross-attention mechanism, and softmax(·) represents an activation function that normalizes a numerical vector into a probability distribution, represents the dimension of the encoding vector;
[0208] Step 53, perform layer normalization LN on the feature representation weighted by the cross-attention mechanism, and add it to the temporal encoding t. Subsequently, input the result into the feed-forward neural network FFN and the self-attention module MHSA to obtain the feature representation o after fusing prior knowledge:
[0209] o = MHSA(FFN(LN(CA) + t)).
[0210] Step 6, construct a negative sample set and propose a joint loss function to enhance the ability of the denoising model while retaining boundary information.
[0211] Step 61, input the conventional-dose CT image, denoted as x 0 , and randomly add Gaussian noise, apply mean filtering, or perform both operations simultaneously to construct a negative sample set:
[0212] n(x 0 ) = x 0 + ∈
[0213]
[0214]
[0215] where n(·) uniformly represents the method of constructing the negative sample set, the Gaussian noise ∈ follows the standard Gaussian distribution, k represents the size of the filter, and x 0 (r, c) represents the pixel value of the image at the position (r, c), and (i, j) is the position of the filter;
[0216] Step 62, construct the loss function:
[0217]
[0218] L total = L diff + λL r-cl
[0219] Optimize the model parameters through the gradient descent algorithm until the model converges.
[0220] Step 7: Concatenate the randomly generated Gaussian noise with the preprocessed test data on the channel and obtain the prior mask image.
[0221] Step 71: Preprocess the low-dose CT test image through Step 1 and denote it as y;
[0222] Step 72: Randomly sample Gaussian noise ∈ and denote it as x t , where t is set to 1000, and concatenate it with the preprocessed test image on the channel through Step 4;
[0223] Step 73: Obtain the discrete prior mask representation m by passing the preprocessed test image through Steps 2 and 3 q .
[0224] Step 8: Introduce the prior mask image of the test image into the denoising network and output the final denoised image after multiple iterations.
[0225] Step 81: Obtain the discrete prior mask representation m q and introduce it into the denoising network through Step 5;
[0226] Step 82: Obtain the denoised image finally through multiple iterations of the denoising network:
[0227]
[0228]
[0229] where ∈ θ (·) is the output of the denoising network; is the discretized prior mask, x 0 is the conventional-dose CT image, x t is the generated data sample, y is the preprocessed low-dose CT test image I LD , the parameter t starts from 1000 and decreases sequentially in steps of 100 until t = 1.
[0230] Taking the medical low-dose CT image as the input, the low-dose CT denoising method of the present invention is used for image enhancement. During training, the registered low-dose CT images and conventional-dose CT images are used. First, the low-dose CT image data is acquired and preprocessed; then, the preprocessed image is input into the prior knowledge extraction module to generate the corresponding prior mask image. Next, the discrete coding network is used to discretize the prior mask, and a negative sample set is constructed based on this, and at the same time, a joint loss function is designed to enhance the ability of the denoising network to retain key boundary information. Finally, the prior mask of the test image is introduced into the denoising network through the knowledge fusion module, and after multiple iterative processes, the denoised image is finally output.
[0231] The present invention proposes a controllable diffusion model based on SAM prior knowledge, which is specifically used for low-dose CT denoising; its basic process includes: first, acquiring and preprocessing the low-dose CT image data, and then inputting it into the prior knowledge extraction module to generate the corresponding prior mask image; then, discretizing the prior mask through the discrete coding network, and constructing a negative sample set and designing a joint loss function based on this to enhance the ability of the denoising network to retain key boundary information; finally, introducing the prior mask of the test image into the denoising network through the knowledge fusion module, and through multiple iterations, the denoised image is finally output. Thanks to the prior knowledge provided by SAM, the present invention significantly improves the effect of low-dose CT denoising in clinical application scenarios, the generated image boundaries are clearer, and the noise is more easily suppressed. By implementing the technical solution of the present invention, the denoising performance of low-dose CT images can be significantly improved, and clearer tissue boundaries can be retained, so as to assist doctors in making more accurate diagnoses and treatments.
[0232] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that the specific implementation manners of the present invention can still be modified or equivalently replaced. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.
Claims
1. A CT image denoising method based on prior knowledge, characterized in that: The steps include: Step S1, acquiring low-dose CT image data and performing preprocessing; Step S2, converting the preprocessed low-dose CT image into a priori mask image; Step S3, processing the priori mask image through a discrete coding and decoding network to obtain an enhanced priori mask image; Step S4, introducing the enhanced priori mask into the denoising network through the knowledge fusion module to obtain a denoising model; Step S5, construct a negative sample set and propose a joint loss function; Step S6, obtaining a conventional-dose CT image paired with the low-dose CT image, adding Gaussian noise to the conventional-dose CT image, and splicing the conventional-dose CT image with the low-dose CT image on the channel to obtain a priori mask image of the test image; Step S7, inputting the priori mask image of the test image into the denoising model, and outputting the final denoised CT image.
2. The CT image denoising method based on prior knowledge guidance according to claim 1, characterized in that: The step S1 comprises the following steps: Step S11, using CT scanning to collect low-dose CT image data, and defining the low-dose CT image as I LD ∈R C×H×W , conventional dose CT image I ND ∈R C×H×W , where C represents the image dimension, H represents the height of the image, and W represents the width of the image; Step S12, preprocessing the acquired low-dose CT image: S LD (c,h,w)=clip(I LD ,-1024,3072) S ND (c,h,w)=clip(I ND ,-1024,3072) Where, clip(·) represents the HU value of the truncated CT image, S LD (c,h,w) and I LD (c,h,w) are the pixel values of the low-dose CT image (c,h,w), They represent the maximum and minimum values of all pixels in the low-dose CT image, respectively. ND (c,h,w) and I ND (c,h,w) are the pixels of the conventional dose CT image (c,h,w), They represent the maximum and minimum values of all pixel specific values in conventional dose CT images respectively; Step S13, the HU value range of the preprocessed low-dose CT image is intercepted to [-1024, 3072], and normalized to [0, 1].
3. The CT image denoising method based on prior knowledge guidance according to claim 1, characterized in that: The step S2 comprises the following steps: Step S21, input the pre-processed low-dose CT image into the pre-trained segmentation model based on the VIT-H architecture to generate multiple low-dose CT images I LD The mask image of Step S22, filter the generated mask image and filter out the masks with a total area less than 50 pixels: M′={m i ∈M|A(m i )≥50} Among them, M′ is the set of effective mask images after screening, m i is the i-th mask image in the set, and A(m i ) represents the mask image m i ∈R 1×H×W area; Step S23: add all the filtered valid mask images at the channel level and normalize them to the range of [0,1] to obtain the final mask image m∈R 1×H×W : m′=Add(m1,m2,…,m n ) Where Add(·) means adding the effective mask image pixel values in the channel dimension, m′ min , m′ max They represent the minimum and maximum values of all pixel values in the spliced mask image respectively, and n represents the number of mask images generated by SAM.
4. The CT image denoising method based on prior knowledge guidance according to claim 1, characterized in that: The step S3 comprises the following steps: Step S31, input the acquired priori mask image into the discrete coding network to obtain the latent space feature representation m of the priori mask image z : m z =VQEncoder(m) Among them, VQEncoder(·) represents the discretized encoding network; Step S32, using the nearest neighbor algorithm to represent the latent space feature m of the prior mask image z Mapped to the codebook vector E = [e1, e2, ..., e K ], the discretized feature representation m is obtained q : m q =e k ,where k=argmin j ||m z -e j ||2; Step S33, obtaining the enhanced priori mask image through the discrete decoding network Among them, VQDecoder(·) represents the discretized decoding network; Step S34, constructing a discrete encoding and decoding network loss function: Among them, sg[·] represents the gradient truncation operation, m z It represents the feature vector of m after being encoded by VQEncoder(·). The model parameters are optimized by gradient descent algorithm until the model converges.
5. The CT image denoising method based on prior knowledge guidance according to claim 4, characterized in that: The step S4 comprises the following steps: Step S41, obtaining the discretized priori mask and the deep features z in the denoising network latent Represents a linear mapping: Among them, W q , W k , W v are linear transformation matrices respectively; Step S42, calculate the cross-attention activation map between the query Q and the key K, and weight it into the value V encoding vector: Among them, CA(·) represents the cross attention mechanism, softmax(·) represents an activation function that normalizes the numerical vector into a probability distribution, represents the dimension of the encoding vector; Step S43, the feature representation after the cross attention mechanism weighting is layer normalized LN, and added to the time code t, and then the result is input into the feedforward neural network FFN and the self-attention module MHSA to obtain the feature representation o after integrating the prior knowledge: o = MHSA (FFN (LN (CA) + t)).
6. The CT image denoising method based on prior knowledge guidance according to claim 1, characterized in that: The specific process of constructing the negative sample set is: The conventional dose CT image is recorded as x0, and the negative sample set is constructed by randomly adding Gaussian noise and mean filtering method: n(x0)=x0+∈ Among them, n(·) uniformly represents the method of constructing a set of negative samples, ∈ is Gaussian noise, which obeys the standard Gaussian distribution, k represents the size of the filter, x0(r,c) represents the pixel value of the image at position (r,c), and (i,j) is the position of the filter.
7. The CT image denoising method based on prior knowledge guidance according to claim 6, characterized in that: The joint loss function is: THE total =L diff +λL r-cl Among them, L total is the total loss value, L r-cl is the denoising network loss, L diff is the discrete coding and decoding network loss, λ is the coefficient, and y is the preprocessed low-dose CT test image I LD , the time code t represents an integer in the range [0,1000], m q is the discrete prior mask.
8. The CT image denoising method based on prior knowledge guidance according to any one of claims 1 to 7, characterized in that: The step S6 comprises the following steps: Step S61, acquiring a low-dose CT test image and performing preprocessing; Step S62, randomly sampling Gaussian noise ∈, for the conventional dose CT image I ND Gradually add Gaussian noise to generate new data samples x t : Among them, ∈ represents Gaussian noise, α t represents a gradually decreasing parameter, and the time code t represents an integer in the range of [0,1000]; Step S63: The pre-processed low-dose CT test image I LD As condition y, with the new data sample x t Dimensional splicing on channels; Step S64, converting and enhancing the spliced test image to obtain a discrete priori mask representation m q .
9. The CT image denoising method based on prior knowledge guidance according to claim 8, characterized in that: The step S7 comprises the following steps: Step S71, obtain the discrete priori mask representation m q Introduced into the denoising network; Step S72, by iterating the denoising network multiple times, a denoised image is obtained: Among them, ∈ θ (·) is the output of the denoising network, is the discretized priori mask, x0 is the conventional dose CT image, x t is the generated data sample, y is the preprocessed low-dose CT test image I LD , the time code t starts from 1000 and decreases in steps of 100 until t=1.
10. A CT image denoising device based on prior knowledge guidance, characterized in that: include: An image acquisition module, used for acquiring low-dose CT image data and performing preprocessing; An image conversion module, used for converting the preprocessed low-dose CT image into a priori mask image; An image enhancement module is used to process the prior mask image through a discrete encoding and decoding network to obtain an enhanced prior mask image; A model acquisition module is used to introduce the enhanced prior mask into the denoising network through the knowledge fusion module to obtain a denoising model; Negative sample construction module, used to construct negative sample sets and propose joint loss functions; An image stitching module is used to obtain a conventional dose CT image paired with a low dose CT image, add Gaussian noise and stitch it with the low dose CT image on the channel to obtain a priori mask image of the test image; The image denoising module is used to input the prior mask image of the test image into the denoising model and output the final denoised CT image.
Citation Information
Cited By
Contrast-agent-free angiography generation system and method based on mask guidance
CN120316282A