Low-dose CT denoising method and device for enhancing diffusion posterior sampling
By combining UNet and the latent diffusion model, the low-dose CT image denoising scheme is optimized, which solves the problems of poor denoising effect and insufficient generalization ability in the existing technology, and realizes low-radiation and high-quality imaging examination.
Patent Information
- Application Number
- CN202510560251.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-09-26
AI Technical Summary
Existing low-dose CT image denoising schemes are unable to cope with complex practical situations, especially when image misalignment is caused by physiological movement. It is difficult to achieve high-quality denoising effects, and supervised learning has insufficient generalization ability in the distribution difference between training data and real low-dose CT images.
A hybrid low-dose CT denoising mechanism with UNet-enhanced diffuse posterior sampling is designed. The UNet model is used for preliminary denoising, and the latent diffusion model is used to perform feature diffusion and reconstruction in the latent feature space. The variational quantization autoencoder is combined for image reconstruction to optimize the generalization performance of the denoising model.
It achieves high-quality denoising of low-dose CT images in complex practical situations, improves the generalization ability of the model, and is able to generate low-radiation and high-quality imaging examination results.
Smart Images

Figure CN120707414A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing, and in particular to a low-dose CT denoising method and device with enhanced diffusion posterior sampling. Background Art
[0002] As an important imaging method for tumor diagnosis and treatment, positron emission tomography combined with computed tomography (PET-CT) plays an irreplaceable role in the diagnosis, staging, and efficacy evaluation of malignant tumors. However, its radiation safety issues are receiving increasing clinical attention. Based on the principle of radiation protection optimization, how to minimize the radiation dose of PET / CT examinations while ensuring diagnostic effectiveness has become a major issue in the imaging field.
[0003] In recent years, the clinical application of whole-body PET / CT systems has brought new opportunities for reducing radiation dose. Thanks to their excellent spatiotemporal resolution and detection sensitivity, the dose of radioactive tracers can be reduced to 50% of the conventional dose. However, it is worth noting that CT scans are still the main source of radiation in PET / CT examinations. CT images have a dual role in PET / CT examinations. On the one hand, they provide attenuation correction data for PET images, and on the other hand, they can independently provide high-resolution anatomical images. Clinical practice has shown that diagnostic CT images require a higher X-ray dose, while CT images used only for attenuation correction can adopt an ultra-low-dose scanning scheme.
[0004] Therefore, researching and developing new low-radiation-dose whole-body PET / CT imaging methods to improve the quality of low-dose CT (LDCT) images to that of conventional-dose CT (NDCT) images can not only ensure imaging quality but also reduce harmful radiation doses. This has important scientific significance and application prospects for the clinical diagnosis of whole-body PET / CT of the subjects.
[0005] Low-dose CT imaging denoising methods based on deep learning are gaining increasing attention. Initially, convolutional neural networks (CNNs), such as the RED-CNN model, were the primary approach for low-dose CT image denoising. To further improve visual quality, generative adversarial networks (GANs) have been applied to low-dose CT image denoising. With the successful application of visual transformers in various tasks, several Transformer-based methods have also been applied to low-dose CT image denoising.
[0006] However, most current deep learning-based methods for low-dose CT imaging rely on image pairs of structurally aligned low-dose CT images and conventional-dose CT images. This requirement is often difficult to achieve in actual clinical practice because even in the case of two consecutive low-dose scans, physiological movements such as breathing, heartbeat, and gastrointestinal motility can cause inherent misalignment between images.
[0007] Currently, a commonly used solution is to use a pseudo-paired dataset. However, the inventors of this application found that although certain results have been achieved, its effectiveness in actual clinical applications is still significantly limited due to the insufficient generalization ability of supervised learning when dealing with the distribution differences between training data and real low-dose CT images.
[0008] In addition, some existing supervised training models often exhibit problems of blurred edges and over-smoothing of images.
[0009] From this, we can see that the existing low-dose CT image denoising solutions are difficult to cope with complex actual situations and are difficult to meet the high-quality denoising requirements for low-dose CT images. Summary of the Invention
[0010] The present application provides a low-dose CT denoising method and device with enhanced diffuse posterior sampling. By designing a novel hybrid low-dose CT denoising mechanism with UNet enhanced diffuse posterior sampling and designing a corresponding model training scheme for this purpose, the CT image denoising model obtained by such training can achieve better denoising effect and stronger generalization, can well cope with complex actual situations, and well meet the high-quality denoising requirements for low-dose CT images, thereby helping to provide low-radiation and high-quality imaging examination services for subjects in practical applications.
[0011] In a first aspect, the present application provides a low-dose CT denoising method with enhanced diffusion posterior sampling, the method comprising:
[0012] Acquiring a sample LDCT image, wherein the sample LDCT image is specifically a CT image corresponding to a low-level radiation dose condition;
[0013] Configuring a corresponding sample NDCT image for the sample LDCT image as a label of the sample LDCT image, wherein the sample NDCT image is specifically a CT image corresponding to a conventional level radiation dose condition;
[0014] The CT image denoising model is trained using the labeled sample LDCT images, wherein the CT image denoising model is used to process the CT images input by the overall model corresponding to low-level radiation dose conditions into images corresponding to conventional radiation dose conditions. The CT image denoising model specifically includes a UNet model and a latent diffusion model. The UNet model is used to perform preliminary denoising on the CT images input by the overall model to predict the denoising prior. The latent diffusion model is used to encode the image features output by the UNet model and the image features of the LDCT images input by the overall model into the latent feature space for feature diffusion using a variational quantization autoencoder, and then reconstruct the latent feature representation back to the pixel space using a variational quantization autoencoder to generate the corresponding conventional radiation dose image as the output of the overall model.
[0015] In a second aspect, the present application provides a low-dose CT denoising device with enhanced diffuse posterior sampling, the device comprising:
[0016] an acquisition unit, configured to acquire a sample LDCT image, wherein the sample LDCT image is specifically a CT image corresponding to a low-level radiation dose condition;
[0017] a labeling unit, configured to configure a corresponding sample NDCT image for the sample LDCT image as a label for the sample LDCT image, wherein the sample CDCT image is specifically a CT image corresponding to a conventional level radiation dose condition;
[0018] A training unit is used to train a CT image denoising model with the labeled sample LDCT images, wherein the CT image denoising model is used to process the CT image input by the overall model corresponding to the low-level radiation dose condition into an image corresponding to the conventional level radiation dose condition. The CT image denoising model specifically includes a UNet model and a latent diffusion model. The UNet model is used to perform preliminary denoising on the CT image input by the overall model to predict the denoising prior. The latent diffusion model is used to encode the image features output by the UNet model and the image features of the LDCT image input by the overall model into the latent feature space for feature diffusion using a variational quantization autoencoder, and then reconstruct the latent feature representation back to the pixel space using the variational quantization autoencoder to generate the corresponding conventional level radiation dose image as the output of the overall model.
[0019] In a third aspect, the present application provides a processing device comprising a processor and a memory, wherein a computer program is stored in the memory, and when the processor calls the computer program in the memory, the method provided in the first aspect of the present application or any possible implementation of the first aspect of the present application is executed.
[0020] In a fourth aspect, the present application provides a computer-readable storage medium, which stores multiple instructions, and the instructions are suitable for a processor to load to execute the method provided in the first aspect of the present application or any possible implementation of the first aspect of the present application.
[0021] From the above content, it can be concluded that this application has the following beneficial effects:
[0022] Aiming at the goal of low-dose CT denoising, this application designs a novel hybrid low-dose CT denoising mechanism of UNet enhanced diffusion posterior sampling, and designs a corresponding model training scheme for this purpose. The CT image denoising model obtained by such training can achieve better denoising effect and stronger generalization, can cope well with complex actual situations, and meet the high-quality denoising requirements for low-dose CT images, thereby helping to provide low-radiation and high-quality imaging examination services to the examinees in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0024] Figure 1 A flowchart of a low-dose CT denoising method based on enhanced diffuse posterior sampling in this application;
[0025] Figure 2 A schematic diagram of a scenario for the overall framework of this application plan;
[0026] Figure 3 A schematic diagram of a low-dose CT denoising device for enhanced diffuse posterior sampling in this application;
[0027] Figure 4 This is a structural diagram of the processing equipment for this application. DETAILED DESCRIPTION
[0028] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.
[0029] The terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or modules is not necessarily limited to those steps or modules clearly listed, but may include other steps or modules that are not clearly listed or that are inherent to these processes, methods, products or devices. The naming or numbering of steps in this application does not mean that the steps in the method flow must be executed in the time / logical sequence indicated by the naming or numbering. The process steps that have been named or numbered can be changed in the execution order according to the technical purpose to be achieved, as long as the same or similar technical effects can be achieved.
[0030] The division of modules in this application is a logical division. In actual application, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection between modules can be electrical or other similar forms, which are not limited in this application. Moreover, the modules or submodules described as separate components may or may not be physically separated, may or may not be physical modules, or may be distributed into multiple circuit modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this application.
[0031] Before introducing the low-dose CT denoising method with enhanced diffusion posterior sampling provided by this application, the background content involved in this application is first introduced.
[0032] The low-dose CT denoising method, device and computer-readable storage medium with enhanced diffuse posterior sampling provided in this application can be applied to processing equipment. By designing a novel hybrid low-dose CT denoising mechanism with UNet enhanced diffuse posterior sampling and designing a corresponding model training scheme for this purpose, the CT image denoising model obtained by such training can achieve better denoising effect and stronger generalization, can well cope with complex actual situations, and well meet the high-quality denoising requirements for low-dose CT images, thereby helping to provide low-radiation and high-quality imaging examination services for subjects in practical applications.
[0033] The low-dose CT denoising method for enhanced diffuse posterior sampling mentioned in this application can be implemented by a low-dose CT denoising device for enhanced diffuse posterior sampling, or by various types of processing devices such as a server, physical host, or user equipment (UE) that integrates the low-dose CT denoising device for enhanced diffuse posterior sampling. The low-dose CT denoising device for enhanced diffuse posterior sampling can be implemented in hardware or software, and the UE can specifically be a terminal device such as a smartphone, tablet computer, laptop computer, desktop computer, or personal digital assistant (PDA). The processing device can be arranged in a device cluster.
[0034] It can be understood that, in actual situations, the data processing corresponding to the present application solution is usually performed on the basis of existing scanned images. Therefore, the processing equipment that executes the low-dose CT denoising method of the present application with enhanced diffusion posterior sampling or is equipped with the corresponding application service of the low-dose CT denoising method of the present application with enhanced diffusion posterior sampling usually only needs to have the required data processing capabilities, and the specific equipment type and form of the equipment are very flexible.
[0035] If direct acquisition of CT images is not involved, the processing device may involve specific methods such as retrieving from the system, retrieving locally, forwarding from other devices, manual entry, or triggering an external scanning device / system for real-time scanning.
[0036] If the direct acquisition capability of CT images is also involved, the processing device needs to make further adaptive configuration work in terms of software and hardware. For example, the processing device can incorporate existing scanning equipment / systems into its own equipment components in the form of equipment clusters, or the existing scanning equipment / systems can be further modified to integrate them into the processing device.
[0037] In addition, data applications corresponding to CT images, such as issuing reports, conducting further data analysis or presenting results, also require further adaptation and configuration of the processing equipment.
[0038] Taking result display as an example, the processing device can be equipped with a display screen (including a touch screen) to complete the result display work, or it can achieve the purpose of result display through an external display screen device or other device with a display screen.
[0039] Next, the low-dose CT denoising method with enhanced diffusion posterior sampling provided by this application is introduced.
[0040] First, see Figure 1 , Figure 1 A flow chart of the low-dose CT denoising method for enhanced diffuse posterior sampling of the present application is shown. The low-dose CT denoising method for enhanced diffuse posterior sampling provided by the present application may specifically include the following steps S101 to S103:
[0041] Step S101, obtaining a sample LDCT image, wherein the sample LDCT image is specifically a CT image corresponding to a low-level radiation dose condition;
[0042] It can be understood that the acquisition process of the sample LDCT image here is usually the acquisition of the on-site image. Of course, in some cases, it can also be the real-time scanning process of the image.
[0043] Among them, for low-level radiation dose conditions or low-dose configurations, there is usually a specific range in clinical work, which can be considered to be within the scope of existing technology. Usually, the CT dose is less than 1mSV (a unit of radiation dose). Of course, the specific radiation dose can be adaptively adjusted based on actual conditions and is usually much less than the radiation dose involved in routine / normal imaging examinations.
[0044] Step S102, configuring a corresponding sample NDCT image for the sample LDCT image as a label of the sample LDCT image, wherein the sample NDCT image is specifically a CT image corresponding to a conventional level radiation dose condition;
[0045] It can be understood that, corresponding to the model training requirements, on the basis of configuring the sample LDCT image, it is also necessary to configure the corresponding annotation, which can also be understood as a model prediction true value, thereby guiding the model training during the model training process.
[0046] In the present application, corresponding to the processing logic of the CT image denoising model, the annotation of the sample LDCT image can specifically be a sample NDCT image corresponding to the conventional level radiation dose condition. It can also be seen from here that for the sample LDCT image, after denoising is completed, it can be converted into a CT image corresponding to the conventional level radiation dose condition. Therefore, the denoising processing of the LDCT image can also be understood as the image reconstruction processing corresponding to the conventional level radiation dose condition.
[0047] As an example, the CT images that can be used in this application can specifically come from diagnostic conventional-dose CT images and low-dose ACCT images for attenuation correction acquired in the uEXPLORER children's whole-body PET / CT scanning protocol.
[0048] For low-dose ACCT images, the scanning tube voltage is 100kVp, the tube current is 20mA, and the scanning range is the whole body; for diagnostic conventional-dose CT images, the scanning tube voltage is 100kVp, a dynamically modulated tube current mode is used, the tube current range is 80-200mA, and the scanning range is from the skull base to the mid-femur.
[0049] The thickness of low-dose ACCT images and diagnostic conventional-dose CT images was set to 3 mm, and the reconstructed image size was 512 × 512. Both were obtained through two scans.
[0050] It should be noted that, in specific applications, the sample NDCT image required to be configured as the annotation of the sample LDCT image can be configured after the sample LDCT image is acquired, or it can be configured together with the sample LDCT image, or even the sample LDCT image can be processed after being configured to the sample NDCT image.
[0051] For the latter, since the sample LDCT image is directly processed based on the sample NDCT image, it is obviously easier to be easy to operate and highly adaptable, which helps to configure higher quality training samples and promote better model training effects.
[0052] In this regard, as an exemplary embodiment, the above step of acquiring the sample LDCT image may specifically include:
[0053] Get sample NDCT image;
[0054] Based on the sample NDCT image, a sample LDCT image is generated.
[0055] It can be understood that the main difference between the two sample CT images is the radiation dose level involved. Therefore, a corresponding algorithm can be conveniently used to downsample the high-quality sample BDCT image into a lower-quality sample LDCT image.
[0056] Furthermore, with respect to the downsampling algorithm configuration work involved, as an exemplary embodiment, generating a sample LDCT image based on the sample NDCT image may specifically include:
[0057] 1) Obtaining projection sinogram data of the sample NDCT image scan using a fan-beam scanning geometry;
[0058] It can be understood that the fan beam scanning geometry itself is a conventional concept in the fan beam imaging process, and the same is true for the projected sinusoidal data.
[0059] 2) Based on the projection sinusoidal data, a mixed noise model of Poisson noise and Gaussian noise is introduced to generate noisy projection data as a sample LDCT image. The quantization formula involved is specifically expressed as follows:
[0060] b=Possion(k*e -p +n)+Guassian(∈)
[0061]
[0062] Where b is the intermediate variable, Poission is the Poisson noise generating function, k represents the X-ray intensity, p is the projection sinusoidal data, n is the readout noise, Gaussian is the Gaussian noise generating function, ∈ is the background noise variance, is the noisy projection data.
[0063] As an example, the k value can be set to 2×10^5, 4×10^5 and 6×10^5, and the corresponding low-dose radiation levels are recorded as level 1 to level 3 respectively.
[0064] It can be seen that in the embodiment here, starting from a specific quantization formula, a specific implementation scheme is given for how this application generates a sample LDCT image based on a sample NDCT image, which has better practical significance.
[0065] In the case where the sample LDCT image is obtained by processing the sample NDCT image, it can be understood that the annotation processing performed in step S102 here is the configuration / pairing processing of the sample LDCT image and the sample NDCT image obtained in step S101 in terms of annotation, which corresponds to the configuration work requirements of the training sample.
[0066] Step S103 is to train a CT image denoising model on the labeled sample LDCT image, wherein the CT image denoising model is used to process the CT image input by the overall model corresponding to the low-level radiation dose condition into an image corresponding to the conventional level radiation dose condition. The CT image denoising model specifically includes a UNet model, a VQ-VAE model and a latent diffusion model. The UNet model is used to perform preliminary denoising on the CT image input by the overall model to predict the denoising prior. The latent diffusion model is used to encode the image features output by the UNet model and the image features of the LDCT image input by the overall model into the latent feature space for feature diffusion using the encoder of the VQ-VAE model, and then use the decoder of the VQ-VAE model to reconstruct the latent feature representation back to the pixel space to generate the corresponding conventional level radiation dose image as the output of the overall model.
[0067] Among them, for the three models specifically involved in the CT image denoising model, this application also provides specific model structure configuration content.
[0068] Correspondingly, as an exemplary embodiment, the UNet model (or U-Net model) of the present application may specifically include the following configuration contents:
[0069] The UNet model includes a first downsampling module and a first upsampling module. A skip connection is introduced between the first downsampling module and the first upsampling module. The first downsampling module and the first upsampling module each contain four convolution modules. Each convolution module is followed by a first downsampling layer or a first upsampling layer. Each convolution module contains two two-dimensional convolution layers. A batch normalization layer and an activation function layer are added after each two-dimensional convolution layer.
[0070] The Vector Quantized-Variational Auto Encoder (VQ-VAE) model of this application may specifically include the following configuration contents:
[0071] The encoder and decoder of the VQ-VAE model adopt a symmetrical structural design and are both composed of an input layer, a two-layer cascaded second downsampling layer, a second upsampling layer, a second intermediate module, and a second output layer. The second downsampling layer of each level contains two second residual blocks, and the first second residual block contained in the second downsampling layer of each level is followed by a corresponding downsampling operation. The second upsampling layer of each level contains two second residual blocks, and the first second residual block contained in the second upsampling layer of each level is followed by a corresponding upsampling operation.
[0072] The latent diffusion model of this application may specifically include the following configuration contents:
[0073] The model input first passes through a 3×3 convolutional layer to expand the feature dimension to 64, and then enters 5 cascaded third downsampling modules. Each level of the third downsampling module contains 3 third residual blocks and a third downsampling layer. Each third residual block consists of two 3×3 convolutional layers. The two 3×3 convolutional layers contain group normalization and SiLU activation functions. During the downsampling process, the number of channels is gradually expanded according to the ratio of [1, 2, 4, 4, 8]. The third intermediate module consists of two third residual blocks and an attention module. The third upsampling module is symmetrical with the third downsampling module and gradually restores to the input resolution. Finally, it passes through a third output layer to output features with the same dimension as the input. The third output layer consists of group normalization, SiLU activation layer and 3×3 convolutional layer.
[0074] It can be understood that in the embodiment here, for the specific model structure of the three models specifically involved in this application, a set of practical and specific implementation supporting solutions is provided from the practical operation level, which has better practical significance.
[0075] It can be understood that the latent diffusion model has higher computational efficiency and scalability advantages compared to the direct diffusion method. Based on this, the present application introduces the latent diffusion model to generate conventional dose CT images in a low-dimensional latent space.
[0076] Among them, the model processing task of the potential diffusion model can also be referred to as the posterior diffusion sampling processing.
[0077] Specifically, the encoder of the VQ-VAE model can be used to first encode the image features into the latent space. After the latent feature diffusion is completed, the decoder of the VQ-VAE model can be used to reconstruct the latent feature representation back to the pixel space. This method can not only significantly reduce the computational complexity, but also maintain the quality of image reconstruction.
[0078] Based on the previous specific model structure content, we will now start from the formula aspect to provide a more specific explanation of the corresponding model processing logic.
[0079] As an exemplary embodiment, for the latent diffusion model, assuming that the sample NDCT image x (i.e., the sample NDCT image is denoted as x), the feature after the encoder ε of the VQ-VAE model is z0, and according to the forward diffusion process, z0 is gradually denoised, which can be expressed as:
[0080]
[0081] Among them, z t is the feature of adding noise to t steps, t∈(0,T),∈ t is Gaussian noise sampled from Gaussian distribution N(0,1), β1,…,βt is a predefined parameter sequence that satisfies β t ∈(0,1), and β t <β t+1 , through the transformation of the Markov chain, the noise result of any number of steps can be directly obtained from z0, which is expressed as:
[0082] Among them, α t is the intermediate quantity, α t =1-β t , is α t The cumulative amount of
[0083] When performing back diffusion, a z is randomly sampled from the Gaussian distribution N(0,1) T ~N(0,1), gradually denoising to get z0, given the sample z at the current time t t , z t-1 The conditional distribution p (i.e., the probability distribution of the reverse diffusion model) θ (z t-1 |z t ) can be expressed as:
[0084] p θ (z t-1 |z t ):=N(z t-1 ;μ θ (z t , t), σ 2 I),
[0085] Among them, σ 2 I is the variance in the reverse diffusion process, σ 2 is the noise intensity of the diffusion process, for the unknown noise distribution μ θ (z t ,t), denoted by f θ (.) is predicted by the potential diffusion model, which is expressed as:
[0086]
[0087] Next, regarding the post-diffusion sampling method designed in this application, it can be understood that when facing the need to denoise low-dose CT images, the inventors of this application considered that the general noise inverse problem can be expressed as the following formula:
[0088] y=H(x0)+ε,
[0089] Where x0 represents the clean image, H(.) represents the forward degradation operator, and ε is Gaussian noise.
[0090] However, since the forward degradation operator H(.) is unknown in the low-dose CT image denoising task, this application introduces the UNet model. First, the pre-trained UNet model performs preliminary denoising on the low-dose CT image and predicts the denoising prior without estimating H(.). The preliminary denoising result is then added as a posteriori information to the diffusion inverse process to constrain sampling.
[0091] Thus, as an exemplary embodiment, there may be:
[0092] According to the reverse process of the classical diffusion model, f θ (.) Input the current time step t and feature z t , calculate μ θ (z t ,t), we can roughly estimate the characteristics at t=0 It is expressed as follows:
[0093]
[0094] Then we can combine the noise of any number of steps in the previous step t Quantify the formula and get the result of step t-1, which is expressed as follows:
[0095]
[0096] Different from this, in this application, the above results are the intermediate results of step t-1. On this basis, this application further introduces an update strategy to update the output results of the UNet model. Adjustments are made to constrain the structure of the generated image and adapt it to the denoising task, as shown below:
[0097]
[0098] Among them, z t-1 is the final sampling result of step t-1, z y is the feature obtained by encoding the sample LDCT image through the encoder of the VQ-VAE model, z x U is the feature obtained by encoding the sample LDCT image after preliminary denoising by the UNet model and then by the encoder of the VQ-VAE model. λ is the update step size, μ is an adjustable parameter that controls the ratio of the two, and ▽ represents the gradient operator. is the symbol of the two-norm, and then z t-1 Input to f θ (.) Perform the next step of sampling, iterate repeatedly, and finally predict z 0 .
[0099] At the same time, this application also designs an intermediate step initialization strategy to effectively reduce the number of sampling steps required for the posterior diffusion sampling processing (model processing of the potential diffusion model) and speed up the sampling process, instead of starting from pure Gaussian noise, thereby reducing the required number of sampling steps and stabilizing the sampling process.
[0100] In this regard, as an exemplary embodiment, unlike the traditional diffusion model that starts with pure noise and diffuses backward, for the potential diffusion model, the present application can specifically add noise to the image features output by the UNet model for a specified number of steps under the design of the intermediate step initialization strategy, and use the noise addition result as the starting point for sampling.
[0101] The specified number of steps can be recorded as N, which is much smaller than the sampling steps of the conventional diffusion model (DDPM). The specific value can be set according to the task effect.
[0102] For the overall model working logic, there can be:
[0103]
[0104] In this way, the model fusion architecture of the UNet model, VQ-VAE model and latent diffusion model involved in the above-mentioned application, by introducing a latent diffusion model that performs self-supervised learning only on conventional-dose CT images, combines the improved posterior sampling algorithm to optimize the preliminary denoising results of the supervised trained UNet model, so that it has the powerful denoising ability of UNet while having stronger generalization performance, and can well achieve the denoising goal of low-dose CT images.
[0105] In the specific model training process, it can be understood that it involves the training of both the UNet model and the potential diffusion model.
[0106] In the specific configuration of training samples, as an example, there are:
[0107] For the UNet model, given the training data set D1, D1={(y1,x1),(y2,x2),…,(y n ,x n )}, where n is the total number of training samples, x is a sample in the conventional dose CT image dataset, and x={x1,x2,…,x n}, y is a sample in the low-dose CT image dataset generated based on conventional-dose CT images, y={y1,y2,…,y n}.
[0108] For the latent diffusion model, self-supervised training is performed given a training dataset D2, D2 = {x1, x2, ..., x n}.
[0109] For the loss function involved in the training process, this application also provides further optimization configuration work.
[0110] Correspondingly, as an exemplary embodiment, the loss function adopted by the UNet model during the training process may specifically be the L1 loss function, and the loss function adopted by the latent diffusion model during the training process may specifically be the L2 loss function.
[0111] Among them, the L1 loss function and the L2 loss function are existing loss functions themselves. This application specifically selects these two loss functions from a large number of existing loss functions that can be adopted to achieve the purpose of being highly adapted to the application scenarios of this application, which can well assist the training of the UNet model and the potential diffusion model.
[0112] Specifically, the L1 loss function can be expressed as:
[0113] L1=‖U(y)-x‖,
[0114] On the other hand, the L2 loss function can be specifically expressed as:
[0115]
[0116] As an example, for the latent diffusion model f θ The training of (.) can be implemented as follows:
[0117] At the beginning of each training iteration, a training sample z0 is drawn and a time step t is randomly selected from the time step sequence. Add noise to sample z0, and then add noise result z t The input is processed into the potential diffusion model to realize forward propagation. Then, based on the output of the model and the added noise ∈, the loss function is calculated. The model parameters are optimized according to the loss function calculation results to realize back propagation. In this way, a round of model training is formed in which the supervised model learns to predict the noise distribution of the corresponding time step. In this way, when the preset model training requirements such as training time, number of training times, and prediction accuracy are met, the training of the potential diffusion model can be completed (the training model can predict its noise distribution based on the sample of the current time step), and it can be put into practical use later.
[0118] In contrast, the training of the UNet model is relatively simple. The core lies in calculating the loss function based on the image features of the LDCT image denoising results and the NDCT image paired with the LDCT image before denoising to promote model parameter optimization.
[0119] Furthermore, for the above-mentioned solutions, Figure 2 A scenario diagram of the overall framework of the present application solution is shown for a more vivid understanding.
[0120] After completing the preliminary training / preliminary configuration work for the CT image denoising model, it is easy to understand and can be put into subsequent practical applications. In this regard, this application can also involve subsequent practical application links.
[0121] Correspondingly, as an exemplary embodiment, the method of the present application may further include:
[0122] Acquiring an initial CT image corresponding to a low-level radiation dose condition;
[0123] Input the initial CT image into the CT image denoising model;
[0124] A target CT image corresponding to a conventional level radiation dose condition is obtained as output by the CT image denoising model.
[0125] It is understandable that in the model application phase, the initial CT image is similar to the sample LDCT image, and the acquisition path or image source can be configured very flexibly. It can be obtained by scanning with a local device, or it can be retrieved from the system, retrieved locally, forwarded from other devices, manually entered, or triggered by an external scanning device / system for real-time scanning.
[0126] In this way, after obtaining the initial CT image that needs to be denoised in the current application scenario, it can be input and imported into the previously trained CT image denoising model to carry out the corresponding denoising processing. After the CT image denoising model completes the denoising processing of the initial CT image, it can output the target CT image corresponding to the conventional level radiation dose condition. In this case, the target CT image can be extracted.
[0127] At the same time, it is understandable that for the target CT image, further data application links may be involved, such as local storage, remote storage, result display, result forwarding, output of completion processing prompts, or further data analysis (typically such as image analysis, disease diagnosis), etc.
[0128] As for the result display setting, as described above, it can be processed through the display screen of the local device itself, or through an external display screen device or other device with a display screen.
[0129] It can be understood that specific data application processing can be adaptively configured according to pre-configured and real-time configured data application strategies, and is highly flexible.
[0130] Finally, regarding the above solution content, in general, for the low-dose CT denoising goal, this application designs a novel hybrid low-dose CT denoising mechanism of UNet enhanced diffusion posterior sampling, and designs a corresponding model training scheme for this purpose. The CT image denoising model trained in this way can achieve better denoising effect and stronger generalization, can cope with complex actual situations well, and meet the high-quality denoising requirements for low-dose CT images, and thus help to provide low-radiation and high-quality imaging examination services to the examinees in practical applications.
[0131] The above is an introduction to the low-dose CT denoising method with enhanced diffuse posterior sampling provided in this application. In order to facilitate better implementation of the low-dose CT denoising method with enhanced diffuse posterior sampling provided in this application, this application also provides a low-dose CT denoising device with enhanced diffuse posterior sampling from the perspective of functional modules.
[0132] See Figure 3 , Figure 3 This is a schematic diagram of the structure of a low-dose CT denoising device with enhanced diffuse posterior sampling in this application. In this application, the low-dose CT denoising device with enhanced diffuse posterior sampling 300 may specifically include the following structure:
[0133] An acquisition unit 301 is configured to acquire a sample LDCT image, wherein the sample LDCT image is specifically a CT image corresponding to a low-level radiation dose condition;
[0134] The labeling unit 302 is configured to configure a corresponding sample NDCT image for the sample LDCT image as a label of the sample LDCT image, wherein the sample NDCT image is specifically a CT image corresponding to a conventional level radiation dose condition;
[0135] The training unit 303 is used to train a CT image denoising model using the labeled sample LDCT images, wherein the CT image denoising model is used to process the CT image input by the overall model corresponding to the low-level radiation dose condition into an image corresponding to the conventional level radiation dose condition. The CT image denoising model specifically includes a UNet model, a VQ-VAE model and a latent diffusion model. The UNet model is used to perform preliminary denoising on the CT image input by the overall model to predict the denoising prior. The latent diffusion model is used to encode the image features output by the UNet model and the image features of the LDCT image input by the overall model into the latent feature space for feature diffusion using the encoder of the VQ-VAE model, and then reconstruct the latent feature representation back to the pixel space using the decoder of the VQ-VAE model to generate the corresponding conventional level radiation dose image as the output of the overall model.
[0136] In an exemplary embodiment, the acquiring unit 301 is specifically configured to:
[0137] Get sample NDCT image;
[0138] Based on the sample NDCT image, a sample LDCT image is generated.
[0139] In another exemplary embodiment, the acquiring unit 301 is specifically configured to:
[0140] Using fan-beam scanning geometry, projected sinogram data of the sample NDCT image scan is acquired;
[0141] Based on the projection sinusoidal data, a mixed noise model of Poisson noise and Gaussian noise is introduced to generate noisy projection data as a sample LDCT image. The quantization formula involved is specifically expressed as follows:
[0142] b=Possion(k*e -p +n)+Guassian(∈)
[0143]
[0144] Where b is the intermediate variable, Poission is the Poisson noise generating function, k represents the X-ray intensity, p is the projection sinusoidal data, n is the readout noise, Gaussian is the Gaussian noise generating function, ∈ is the background noise variance, is the noisy projection data.
[0145] In another exemplary embodiment, the UNet model specifically includes the following configuration content:
[0146] The UNet model includes a first downsampling module and a first upsampling module. A skip connection is introduced between the first downsampling module and the first upsampling module. The first downsampling module and the first upsampling module each contain four convolution modules. Each convolution module is followed by a first downsampling layer or a first upsampling layer. Each convolution module contains two two-dimensional convolution layers. A batch normalization layer and an activation function layer are added after each two-dimensional convolution layer.
[0147] The VQ-VAE model specifically includes the following configurations:
[0148] The encoder and decoder of the VQ-VAE model adopt a symmetrical structural design and are both composed of an input layer, a two-layer cascaded second downsampling layer, a second upsampling layer, a second intermediate module, and a second output layer. The second downsampling layer of each level contains two second residual blocks, and the first second residual block contained in the second downsampling layer of each level is followed by a corresponding downsampling operation. The second upsampling layer of each level contains two second residual blocks, and the first second residual block contained in the second upsampling layer of each level is followed by a corresponding upsampling operation.
[0149] The potential diffusion model specifically includes the following configuration contents:
[0150] The model input first passes through a 3×3 convolutional layer to expand the feature dimension to 64, and then enters 5 cascaded third downsampling modules. Each level of the third downsampling module contains 3 third residual blocks and a third downsampling layer. Each third residual block consists of two 3×3 convolutional layers. The two 3×3 convolutional layers contain group normalization and SiLU activation functions. During the downsampling process, the number of channels is gradually expanded according to the ratio of [1, 2, 4, 4, 8]. The third intermediate module consists of two third residual blocks and an attention module. The third upsampling module is symmetrical with the third downsampling module and gradually restores to the input resolution. Finally, it passes through a third output layer to output features with the same dimension as the input. The third output layer consists of group normalization, SiLU activation layer and 3×3 convolutional layer.
[0151] In another exemplary embodiment, for the latent diffusion model, assuming that the sample NDCT image x has a feature z0 after passing through the encoder ε of the VQ-VAE model, z0 is gradually denoised according to the forward diffusion process, which can be expressed as:
[0152]
[0153] Among them, t∈(0,T), z t is the feature of adding noise to t steps, ∈ t is Gaussian noise sampled from Gaussian distribution N(0,1), β1,…,β t is a predefined parameter sequence that satisfies β t ∈(0,1), and β t <β t+1 , through the transformation of the Markov chain, the noise result of any number of steps can be directly obtained from z0, which is expressed as:
[0154]
[0155] Among them, α t is the intermediate quantity, α t =1-β t , is α t The cumulative amount of
[0156] When performing back diffusion, a z is randomly sampled from the Gaussian distribution N(0,1) T ~N(0,1), gradually denoising to get z0, given the sample z at the current time t t , z t-1 The conditional distribution p θ (z t-1|z t ) is expressed as:
[0157] p θ (z t-1 |z t ):=N(z t-1 ;μ θ (z t , t), σ 2 I),
[0158] Among them, σ 2 I is the variance in the reverse diffusion process, σ 2 is the noise intensity of the diffusion process, for the unknown noise distribution μ θ (z t ,t), denoted by f θ (.) is predicted by the potential diffusion model, which is expressed as:
[0159]
[0160] In yet another exemplary embodiment, θ (.) Input the current time step t and feature z t , calculate μ θ (z t ,t), roughly estimate the characteristics at t=0 It is expressed as follows:
[0161]
[0162] Then we get the result of step t-1, which is expressed as follows:
[0163]
[0164] The output of the UNet model is Adjustments are made to constrain the structure of the generated image and adapt it to the denoising task, as shown below:
[0165]
[0166] Among them, z t-1 is the final sampling result of step t-1, z y is the feature obtained by encoding the sample LDCT image through the encoder of the VQ-VAE model, z x U is the feature obtained by encoding the sample LDCT image after preliminary denoising by the UNet model and then by the encoder of the VQ-VAE model, λ is the update step size, μ is an adjustable parameter, ▽ represents the gradient operator, is the symbol of the two norm, and then z t-1 Input to f θ(.) Perform the next step of sampling, iterate repeatedly, and finally predict z 0 .
[0167] In another exemplary embodiment, for the latent diffusion model, under the design of the intermediate step initialization strategy, the image features output by the UNet model are denoised for a specified number of steps, and the denoised results are used as the starting point for sampling.
[0168] In another exemplary embodiment, the apparatus further includes an application unit 304 configured to:
[0169] Acquiring an initial CT image corresponding to a low-level radiation dose condition;
[0170] Input the initial CT image into the CT image denoising model;
[0171] A target CT image corresponding to a conventional level radiation dose condition is obtained as output by the CT image denoising model.
[0172] This application also provides a processing device from the perspective of hardware structure, which can be either a single device or a device cluster. For the convenience of explanation, the devices that may be involved in different situations are described as a single processing device. Figure 4 , Figure 4 The schematic diagram of the structure of the processing device of the present application is shown. Specifically, the processing device of the present application may include a processor 401, a memory 402 and an input / output device 403. The processor 401 is used to execute the computer program stored in the memory 402 to implement the following Figure 1 The steps of the low-dose CT denoising method for enhanced diffusion posterior sampling in the corresponding embodiment; or, when the processor 401 is used to execute the computer program stored in the memory 402, the following is implemented: Figure 3 The memory 402 is used to store the functions of each unit in the embodiment, and the processor 401 executes the above Figure 1 The computer program required for the low-dose CT denoising method with enhanced diffusion posterior sampling in the corresponding embodiment.
[0173] For example, the computer program may be divided into one or more modules / units, one or more of which are stored in memory 402 and executed by processor 401 to implement the present application. One or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in a computer device.
[0174] The processing device may include, but is not limited to, a processor 401, a memory 402, and an input / output device 403. Those skilled in the art will appreciate that the illustrations are merely examples of processing devices and do not limit the processing device. The processing device may include more or fewer components than shown, or a combination of certain components, or different components. For example, the processing device may also include a network access device, a bus, etc., and the processor 401, the memory 402, the input / output device 403, etc. are connected via the bus.
[0175] The processor 401 may be a central processing unit (CPU), or other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor. The processor is the control center of the processing device and connects various parts of the entire device using various interfaces and lines.
[0176] Memory 402 can be used to store computer programs and / or modules. Processor 401 implements various functions of the computer device by running or executing computer programs and / or modules stored in memory 402 and accessing data stored in memory 402. Memory 402 may primarily include a program storage area and a data storage area. The program storage area may store an operating system, at least one application required for a function, and the like; the data storage area may store data created based on the use of the processing device. Furthermore, memory may include high-speed random access memory and non-volatile memory, such as a hard disk, internal memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state storage device.
[0177] When the processor 401 is used to execute the computer program stored in the memory 402, it can specifically implement the following functions:
[0178] Acquiring a sample LDCT image, wherein the sample LDCT image is specifically a CT image corresponding to a low-level radiation dose condition;
[0179] Configuring a corresponding sample NDCT image for the sample LDCT image as a label of the sample LDCT image, wherein the sample NDCT image is specifically a CT image corresponding to a conventional level radiation dose condition;
[0180] The CT image denoising model is trained with the labeled sample LDCT images, wherein the CT image denoising model is used to process the CT images input by the overall model corresponding to low-level radiation dose conditions into images corresponding to conventional radiation dose conditions. The CT image denoising model specifically includes a UNet model, a VQ-VAE model and a latent diffusion model. The UNet model is used to perform preliminary denoising on the CT images input by the overall model to predict the denoising prior. The latent diffusion model is used to encode the image features output by the UNet model and the image features of the LDCT images input by the overall model into the latent feature space for feature diffusion using the encoder of the VQ-VAE model, and then use the decoder of the VQ-VAE model to reconstruct the latent feature representation back to the pixel space to generate the corresponding conventional radiation dose image as the output of the overall model.
[0181] Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working process of the low-dose CT denoising device, processing equipment and corresponding units of the enhanced diffusion posterior sampling described above can be referred to as follows: Figure 1 The description of the low-dose CT denoising method with enhanced diffuse posterior sampling in the corresponding embodiment will not be repeated here in detail.
[0182] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be accomplished by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.
[0183] To this end, the present application provides a computer-readable storage medium, which stores a plurality of instructions, which can be loaded by a processor to execute the present application as follows: Figure 1 The steps of the low-dose CT denoising method for enhanced diffusion posterior sampling in the corresponding embodiment, the specific operations can be referred to as follows Figure 1 The description of the low-dose CT denoising method with enhanced diffuse posterior sampling in the corresponding embodiment will not be repeated here.
[0184] The computer-readable storage medium may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0185] Due to the instructions stored in the computer readable storage medium, the present application can be executed as follows: Figure 1The steps of the low-dose CT denoising method for enhancing diffusion posterior sampling in the corresponding embodiment, therefore, the present application can be realized as follows Figure 1 The beneficial effects that can be achieved by the low-dose CT denoising method with enhanced diffuse posterior sampling in the corresponding embodiment are detailed in the previous description and will not be repeated here.
[0186] The above is a detailed introduction to the low-dose CT denoising method, device, processing equipment and computer-readable storage medium for enhanced diffuse posterior sampling provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the core idea of the present application. At the same time, for those skilled in the art, based on the ideas of the present application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present application.
Claims
1. A low-dose CT denoising method with enhanced diffuse posterior sampling, characterized in that: The method comprises: Acquiring a sample LDCT image, wherein the sample LDCT image is specifically a CT image corresponding to a low-level radiation dose condition; Configuring a corresponding sample NDCT image for the sample LDCT image as a label for the sample LDCT image, wherein the sample NDCT image is specifically a CT image corresponding to a conventional level radiation dose condition; A CT image denoising model is trained using the labeled sample LDCT image, wherein the CT image denoising model is used to process the LDCT image input by the overall model corresponding to the low-level radiation dose condition into an image corresponding to the conventional level radiation dose condition. The CT image denoising model specifically includes a UNet model, a VQ-VAE model and a latent diffusion model. The UNet model is used to perform preliminary denoising on the LDCT image input by the overall model to predict the denoising prior. The latent diffusion model is used to encode the image features output by the UNet model and the image features of the LDCT image input by the overall model into the latent feature space for feature diffusion using the encoder of the VQ-VAE model, and then reconstruct the latent feature representation back to the pixel space using the decoder of the VQ-VAE model to generate the corresponding conventional level radiation dose image as the output of the overall model.
2. The method according to claim 1, characterized in that The acquiring of the sample LDCT image comprises: Acquiring the sample NDCT image; The sample LDCT image is generated based on the sample NDCT image.
3. The method according to claim 2, characterized in that Generating the sample LDCT image based on the sample NDCT image includes: Acquiring projection sinogram data of the sample NDCT image scan using a fan beam scanning geometry; On the basis of the projection sinusoidal data, a mixed noise model of Poisson noise and Gaussian noise is introduced to generate noisy projection data as the sample LDCT image, wherein the quantization formula involved is specifically expressed as follows: b=Possion(k*e -p +n)+Guassian(∈) Wherein, b is an intermediate variable, Poission is a Poisson noise generating function, k represents the X-ray intensity, p is the projection sinusoidal data, n is the readout noise, Gaussian is a Gaussian noise generating function, ∈ is the background noise variance, is the noisy projection data.
4. The method according to claim 1, wherein The UNet model specifically includes the following configuration contents: The UNet model includes a first downsampling module and a first upsampling module, wherein a skip connection is introduced between the first downsampling module and the first upsampling module, wherein the first downsampling module and the first upsampling module each include four convolution modules, each of which is followed by a first downsampling layer or a first upsampling layer, and each of which includes two two-dimensional convolution layers, and each of which is followed by a batch normalization layer and an activation function layer; The VQ-VAE model specifically includes the following configuration contents: The encoder of the VQ-VAE model and the decoder of the VQ-VAE model adopt a symmetrical structural design and are both composed of an input layer, a two-layer cascaded second downsampling layer, a second upsampling layer, a second intermediate module and a second output layer, wherein the second downsampling layer of each level includes two second residual blocks, and the first second residual block included in the second downsampling layer of each level is followed by a corresponding downsampling operation; the second upsampling layer of each level includes two second residual blocks, and the first second residual block included in the second upsampling layer of each level is followed by a corresponding upsampling operation; The potential diffusion model specifically includes the following configuration contents: The model input first passes through a 3×3 convolution layer to expand the feature dimension to 64, and then enters 5 cascaded third downsampling modules. The third downsampling module at each level contains 3 third residual blocks and a third downsampling layer. Each of the third residual blocks consists of two 3×3 convolution layers. The two 3×3 convolution layers contain group normalization and SiLU activation functions. During the downsampling process, the number of channels is gradually expanded according to the ratio of [1, 2, 4, 4, 8]. The third intermediate module consists of two third residual blocks and an attention module. The third upsampling module is symmetrical with the third downsampling module and gradually restores to the input resolution. Finally, it passes through a third output layer to output features with the same dimension as the input. The third output layer consists of group normalization, SiLU activation layer and 3×3 convolution layer.
5. The method according to claim 1, wherein For the latent diffusion model, assuming that the sample NDCT image x has a feature z0 after passing through the encoder ε of the VQ-VAE model, z0 is gradually denoised according to the forward diffusion process, which is expressed as: Among them, t∈(0,T), z t is the feature of adding noise to t steps, ∈ t is Gaussian noise sampled from Gaussian distribution N(0,1), β1,…,β t is a predefined parameter sequence that satisfies β t ∈(0,1), and β t <β t+1 , through the transformation of the Markov chain, the noise result of any number of steps can be directly obtained from z0, which is expressed as: Among them, α t is the intermediate quantity, α t =1-β t , is α t The cumulative amount of When performing back diffusion, a z is randomly sampled from the Gaussian distribution N(0,1) T ~N(0,1), gradually denoising to get z0, given the sample z at the current time t t , z t-1 The conditional distribution p θ (z t-1 |z t ) is expressed as: p θ (With t-1 |from t ):=N(z t-1 ;μ θ (With t ,t),σ 2 AND), Among them, σ 2 I is the variance in the reverse diffusion process, σ 2 is the noise intensity of the diffusion process, for the unknown noise distribution μ θ (z t ,t), denoted by f θ (.) is predicted by the potential diffusion model, which is expressed as:
6. The method according to claim 5, characterized in that Towards f θ (.) Input the current time step t and feature z t , calculate μ θ (z t ,t), roughly estimate the characteristics at t=0 It is expressed as follows: Then we get the result of step t-1, which is expressed as follows: The output results of the UNet model are Adjustments are made to constrain the structure of the generated image and adapt it to the denoising task, as shown below: Among them, z t-1 is the final sampling result of step t-1, z y is the feature obtained by encoding the sample LDCT image through the encoder of the VQ-VAE model, z x U is the feature obtained by encoding the sample LDCT image after preliminary denoising by the UNet model and then by the encoder of the VQ-VAE model, λ is the update step size, μ is an adjustable parameter, represents the gradient operator, is the symbol of the two-norm, and then z t-1 Input to f θ (.) Perform the next step of sampling, iterate repeatedly, and finally predict z 0 .
7. The method according to claim 1, characterized in that For the latent diffusion model, under the intermediate step initialization strategy design, the image features output by the UNet model are denoised for a specified number of steps, and the denoised results are used as the starting point for sampling.
8. A low-dose CT denoising device with enhanced diffuse posterior sampling, characterized in that: The device comprises: an acquisition unit, configured to acquire a sample LDCT image, wherein the sample LDCT image is specifically a CT image corresponding to a low-level radiation dose condition; a labeling unit, configured to configure a corresponding sample NDCT image for the sample LDCT image as a label for the sample LDCT image, wherein the sample NDCT image is specifically a CT image corresponding to a conventional level radiation dose condition; A training unit is used to train a CT image denoising model using the labeled sample LDCT image, wherein the CT image denoising model is used to process the CT image input by the overall model corresponding to the low-level radiation dose condition into an image corresponding to the conventional level radiation dose condition. The CT image denoising model specifically includes a UNet model and a latent diffusion model. The UNet model is used to perform preliminary denoising on the CT image input by the overall model to predict a denoising prior. The latent diffusion model is used to encode the image features output by the UNet model and the image features of the LDCT image input by the overall model into a latent feature space for feature diffusion using a variational quantization autoencoder, and then reconstruct the latent feature representation back to the pixel space using a variational quantization autoencoder to generate a corresponding conventional level radiation dose image as the output of the overall model.
9. A processing device, characterized in that The method comprises a processor and a memory, wherein a computer program is stored in the memory, and when the processor calls the computer program in the memory, the method according to any one of claims 1 to 7 is executed.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the method according to any one of claims 1 to 7.