Low-dose CT image denoising method and system based on U-Net GAN and deformation field registration
Through the U-Net GAN and deformation field registration method, the problem that the unaligned area in low-dose CT image denoising is misjudged as noise is solved, and rapid denoising and tissue structure are achieved, which significantly improves training stability.
Patent Information
- Application Number
- CN202510326959.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-13
AI Technical Summary
The existing low-dose CT image denoising method in actual clinical data is due to the problem of not being strictly aligned, resulting in the misaligned area being miscalculated as noise and being removed, seriously destroying the original structure of the CT image, and the training process is prone to instability problems.
The low-dose CT image denoising method based on U-Net GAN and deformation field registration is adopted. The generator is trained simultaneously with the deformation field network through the generator, and image alignment is used using STN affine transformation to introduce registration loss and smoothing loss to ensure that the original tissue structure of the unaligned area is maintained during the denoising process.
Fast noise denoising end-to-end is achieved, maintaining the organizational structure of the unaligned area, avoiding new artifacts, and significantly improving training stability.
Smart Images

Figure CN120147177A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of deep learning and medical low-dose CT applications, and particularly relates to a low-dose CT image denoising method and system based on U-NetGAN and deformation field registration. Background Art
[0002] Computed Tomography (CT) has become an important issue in the global public health field due to the ionizing radiation of X-rays it uses. The research and development of Low-Dose Computed Tomography (LDCT) has received extensive attention, aiming to reduce the radiation dose while minimizing the impact on image quality. However, reducing the radiation dose leads to a decrease in the signal-to-noise ratio, thereby introducing noise in CT images, manifested as streaky artifacts not seen in Normal-Dose Computed Tomography (NDCT) images and blurring of edge and corner features. These image quality problems may lead to misdiagnosis or missed diagnosis by doctors during diagnosis.
[0003] In recent years, scholars have been dedicated to the research of LDCT denoising algorithms, which can be basically classified into projection domain chordogram data correction algorithms, statistical iterative reconstruction algorithms, and image domain post-processing denoising algorithms. Among them, the image domain post-processing denoising algorithm has received extensive attention because it does not depend on the projection domain, can achieve end-to-end training, and has a relatively fast denoising speed. In this field, traditional denoising algorithms such as BM3D mostly use smoothing filtering techniques to reduce noise, but often result in significant loss of image details and even change the original CT values. With the continuous development of deep learning technology, deep networks mainly based on supervised learning, such as REDCNN and DnCNN, have gradually been applied in LDCT denoising research, and these networks have shown better denoising performance than traditional methods on multiple simulated datasets.
[0004] However, in actual clinical data, due to the influence of factors such as motion and breathing, the alignment between low-dose CT images and normal-dose CT images is often not strict enough. When using supervised learning methods, unaligned regions may be misjudged as noise and removed, thus seriously damaging the original structure of CT images. In response to this phenomenon, some scholars have proposed unsupervised learning methods, such as WGAN and CycleGAN, to learn the mapping relationship between noisy images and clear images through generative adversarial networks. However, in practical applications, these methods can still only alleviate the denoising error in unaligned regions to a certain extent, and in most scenarios, unaligned regions may still be wrongly denoised, and instability problems are also likely to occur during the training process.
[0005] The Chinese patent with the prior art publication number CN115760651A provides an improved RegGAN low-dose CT image denoising method. By dynamically measuring the similarity between the denoised image and the noise-free image during the training process in the GAN framework, the low-dose CT image can fully participate in the deep network training without the guidance of high-quality pixel-by-pixel aligned images. The Sobel convolution operator is added to enhance the edge information of the CT image in the model, and the channel attention mechanism is used to assign different weights to the channels of the feature map to model the correlation between features. A deeper convolutional network is used to replace the original model discriminator, and the self-attention mechanism is added to make the discriminator pay more attention to the important information in the image. However, this method has limitations in the design of the generator network structure and is difficult to handle the detailed information of complex medical images from the perspective of multi-scale processing ability. Although the self-attention mechanism is added to its discriminator, the architecture design is difficult to simultaneously meet the discriminant requirements of global consistency and local details, especially when dealing with non-aligned regions, the effect is limited. Moreover, the use of a single-scale consistency training strategy limits the generalization ability and training stability of the model. In addition, in the design of the loss function, there is a lack of identity loss as a constraint, which cannot effectively guarantee the retention of the original anatomical structure during the denoising process, and the use of L1 loss will lead to overemphasis on the differences between pixels, resulting in the unaligned regions being regarded as noise during the denoising process, affecting the final image quality and clinical application value. Summary of the Invention
[0006] The present invention aims to propose a low-dose CT image denoising method based on U-Net GAN and deformation field registration. Aiming at the problem of non-rigid alignment existing in actual clinical data, it can achieve end-to-end fast denoising, while maintaining the original tissue structure of the non-aligned regions, avoiding the generation of new artifacts, and significantly improving the training stability.
[0007] In the first aspect, the present invention provides a low-dose CT image denoising method based on U-Net GAN and deformation field registration, including: inputting the original low-dose CT image into a trained low-dose CT image denoising model During the downsampling stage, feature extraction stage and upsampling stage, the downsampling stage extracts features through edge padding and multi-scale convolution; the feature extraction stage performs deep processing using a residual module with batch normalization and ReLU activation; the upsampling stage gradually restores the image resolution using deconvolution combined with normalization operations, and finally outputs the denoised low-dose CT image.
[0008] Furthermore, the training of the generator includes the following steps: Collect chest low-dose CT images and their paired but non-rigidly aligned chest normal-dose CT images in clinical practice as training samples; Randomly divide the paired CT image training samples into a training set and a test set, and construct a dynamic training set TrainLoader with different sizes and a static test set TestLoader with the original size; Construct a two-stage training strategy. In the first several iteration rounds, use the dynamic training set TrainLoader to crop images with a size of for the first-stage training. In the subsequent several iteration rounds, use full-size images for the second-stage training; The generator and the deformation field network are trained and updated simultaneously, and are trained non-consistently and alternately with the discriminator; Through the two-stage training strategy and non-consistent alternating training, denoising is performed while the structure of the non-aligned region is not affected, and a low-dose CT image denoising model is obtained; Evaluate the denoising performance of the denoiser on the test set. If the evaluation result meets the requirements, the final low-dose CT image denoising model is obtained; Use the trained low-dose CT image denoising model to perform denoising processing and image restoration on the low-dose CT image.
[0009] Furthermore, the discriminator is based on the U-Net architecture. In the downsampling stage, the denoised image and the normal-dose CT image first pass through multiple downsampling blocks. Each downsampling block performs two convolutional and LeakyReLU activation operations, and a residual connection is applied after each convolution to add the input and output. At the same time, the last convolutional layer performs downsampling of the image features; Then, through a global pooling layer and a fully connected layer to generate the discriminant output; In the upsampling stage, the size of the feature map is restored through the upsampling block deconvolution, convolution, and LeakyReLU activation operations. The input is concatenated with the output of the corresponding downsampling layer through a skip connection, and finally the discriminant result is obtained after passing the feature through the convolutional layer.
[0010] Furthermore, the deformation field network is used to geometrically align low-dose CT images with normal-dose CT images. By learning the geometric differences between the images, a deformation field is generated to adjust the pixel positions of the low-dose images to make their structures consistent with those of the normal-dose images. Based on the ResUnet architecture, it is divided into a downsampling stage, a residual module stage, and an upsampling stage. In the downsampling stage, the denoised image first undergoes multiple convolutional operations, followed by the LeakyReLU activation function after each convolution, along with residual connections. In the residual module stage, feature information is further extracted and processed through multiple residual blocks, and the residual blocks are transformed through convolutional and activation function layers. In the upsampling stage, the spatial dimension of the feature map is restored through transposed convolution operations. During each transposed convolution, it is concatenated with the corresponding feature map in the downsampling stage through skip connections, and a two-dimensional deformation field is output through a convolutional layer, which is used for spatial transformation and alignment of the input image through STN affine transformation to obtain the registered denoised image.
[0011] Furthermore, the paired CT image training samples are randomly divided into a training set and a test set, and different-sized dynamic training sets TrainLoader and the original-sized static test set TestLoader are constructed, including: The paired CT image training samples are read and converted into pydicom.dataset.FileDataset objects suitable for the Python language, and the information is anonymized. The image data pixel_array in the pydicom.dataset.FileDataset objects is pairwise converted into the Tensor data type suitable for the Pytorch framework, and the data is divided into a training set and a test set. The training set data is converted into a torch.utils.data.Dataset object, and torch.utils.data.DataLoader is used to construct different-sized dynamic training sets TrainLoader and the original-sized test set TestLoader. Construction of the TrainLoader dynamic training set. In each iteration round, image pairs are randomly sampled from the training set, and then the randomly sampled images are dynamically cropped to obtain 4 cropped images. Furthermore, the input for each training set is The image; Construction of the TestLoader static test set. The test set does not require cropping operations and directly uses the original-sized Images for verification.
[0012] Furthermore, the paired CT images are obtained by pairing the collected low-dose chest CT images with normal-dose CT images one by one, ensuring that their quantities and sizes are the same, and the time interval between the front and back acquisitions is as short as possible. The final images are saved in a unified format.
[0013] Furthermore, the generator and the deformation field network are trained simultaneously to update the parameters, and the training with the discriminator is performed alternately in a non-uniform manner, including the following steps: Set the training parameters. In each round of training, the generator and the deformation field network are iterated 5 times, and the discriminator is iterated 1 time. Discriminator training: In the U-Net-based discriminator, during the downsampling stage, the input image is globally compressed to capture the global deep information, and the output is ; Through upsampling, the local discriminant information at the pixel level of the original image size is obtained. Similar to PatchGAN, the output is ; Use Hinge Loss as the loss function. Randomly crop a matrix from a part of the real image, and replace the corresponding area of the matrix with the corresponding part of the generated image; at the same time, the labels of the corresponding area of the matrix are proportionally mixed according to the size of the occluded area, and the network parameters of the discriminator are updated through backpropagation gradients. Generator and deformation field network training: The input low-dose CT image passes through the generator to obtain the denoised image , which is input into the discriminator together with the normal-dose CT image to obtain the adversarial loss of Hinge Loss . Then, and are input into the deformation field network to obtain the two-dimensional deformation field . Combine the deformation field Through the affine transformation of STN, the denoised image after deformation field registration is obtained . Calculate the loss function after registration , the smoothing loss , the edge loss and the identity loss; Construct the generator loss as:
[0014] Wherein, is the loss function after registration, is the adversarial loss of Hinge Loss is the smoothing loss, is the edge loss, is the identity loss They are the weights of various losses. Different hyperparameters are set, and through backpropagation gradient, the network parameters of the generator and the deformation field network are updated iteratively.
[0015] In a second aspect, the present invention provides a low-dose CT image denoising system based on U-Net GAN and deformation field registration, including a data acquisition module and a denoising module; The data acquisition module is used to acquire the original low-dose CT image; Based on the low-dose CT image generator, the denoising module, through the downsampling stage, the feature extraction stage, and the upsampling stage. In the downsampling stage, the input image first undergoes edge padding and convolution operations, and then enters the multi-scale convolution module to extract multi-scale features. The multi-scale features are processed by normalization and ReLU activation, and an image is output through a convolutional layer; in the feature extraction stage, based on the residual network, two edge padding and convolution operations are performed in each residual module, along with normalization and activation function layers, and then residual connection is performed to add the input and output; the upsampling stage restores the image size through two deconvolutions, and normalization and ReLU activation operations are added after each deconvolution, and a denoised low-dose CT image is output through a convolutional layer; the low-dose CT image denoising model During training, the low-dose CT image is input into the generator for generative adversarial training. Based on the deformation field network, STN affine transformation is used to introduce registration loss and smoothing loss to align the low-dose CT image and the normal-dose CT image in geometric structure.
[0016] In a third aspect, the present invention can also provide a computer device, including a processor and a memory. Among them, the memory is used to store computer-executable programs, and the processor reads and executes the computer-executable programs from the memory. When the processor executes the computer-executable programs, the low-dose CT image denoising method based on U-Net GAN and deformation field registration of the present invention can be realized.
[0017] Meanwhile, the present invention provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the low-dose CT image denoising method based on U-Net GAN and deformation field registration of the present invention can be realized.
[0018] Compared with the prior art, the present invention has at least the following beneficial effects: In the present invention, multi-scale feature extraction and a deep residual network are used as the generator, which has the advantages of fewer parameters, diverse levels of extracted features, end-to-end inference, etc. Compared with other methods, it can better handle the noise and structure understanding of low-dose CT images with different resolutions; In the present invention, a deformation field network is used to construct a two-dimensional deformation field, and STN affine transformation is used for registration. By introducing registration loss and smoothing loss, it can minimize the mis-denoising caused by non-rigid alignment, maintain the tissue structure in the unaligned area, and perform denoising simultaneously.
[0019] Furthermore, in the present invention, a discriminator based on the Encoder-Decoder architecture of U-Net is used, which can focus on information at different levels globally (distribution) and locally (texture), so as to alleviate the misalignment problem to a certain extent.
[0020] Furthermore, in the present invention, dynamic random cropping in the dynamic training set TrainLoader is constructed, and it comes from multiple occurrences of the same pair of images, which can enhance the randomness of the data and contribute to the generalization ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 It is a training architecture diagram of a low-dose CT image denoising method based on U-Net GAN and deformation field registration according to the present invention; Figure 2 It is a normal-dose CT image (a) and a low-dose CT image (b) of an embodiment of the present invention Figure 3 It is a normal-dose CT image (a) and a low-dose CT image (b) of an embodiment of the present invention with unaligned areas marked; Figure 4 It is a schematic diagram of the overall structure of the generator for multi-scale feature extraction of an embodiment of the present invention; Figure 5 is a schematic diagram of the overall structure of the discriminator based on U-Net (a) and the overall structure of the deformation field network based on ResUnet (b) of an embodiment of the present invention; Figure 6 It is a low-dose CT image, a normal-dose CT image, and a CT image after denoising of an embodiment of the present invention, with a total of six groups.
[0022] Figure 7 It is a low-dose CT image, a normal-dose CT image, and a CT image after denoising by removing the deformation field network and restoring the traditional discriminator of an embodiment of the present invention, with a total of six groups. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0024] This application provides a low-dose CT image denoising network based on U-Net GAN and deformation field registration. The concept is to input the low-dose CT image into the generator for generative adversarial training. The "generated" image is the denoised image, and the generator is used as a denoiser to achieve an end-to-end denoising effect. The design idea and specific architecture are as follows: In order to better capture the features of low-dose CT images at different levels (under different resolutions) and generate CT images with rich details, high quality, and consistency, the generator combines multi-scale feature extraction and deep residual networks, while the discriminator is based on the Encoder-Decoder architecture of U-Net, replacing the original Encoder-Only architecture, to increase the "deception difficulty" of the generator from both the global (distribution information) and local (texture information) aspects and obtain higher-quality denoised images.
[0025] 1) Generator (denoiser) : The network structure of the generator is divided into a downsampling stage, a feature extraction stage, and an upsampling stage. In the downsampling stage, the input image first undergoes edge padding and convolution operations, and then enters the multi-scale convolution module, where multi-scale features are extracted in parallel through and convolution kernels. Subsequently, the features are processed through normalization and ReLU activation, and then downsampled through a convolutional layer. In the feature extraction stage, the network consists of nine residual modules. Each residual module includes two edge padding and convolution operations, accompanied by normalization and activation function layers, and the residual connection adds the input and output. In the upsampling stage, the image size is gradually restored through two transposed convolutions. After each transposed convolution, normalization and ReLU activation operations are added, and finally, the final image is output through a convolutional layer, using the Tanh activation function to ensure the output range.
[0026] 2) Discriminator : The Encoder-Only architecture of the original GAN is replaced with the Encoder-Decoder architecture based on U-Net. Its specific network architecture is based on the U-Net architecture and is divided into a downsampling stage and an upsampling stage. In the downsampling stage, the input image first passes through multiple downsampling blocks (DownBlock). Each block includes two convolutional and LeakyReLU activation operations, and a residual connection is applied after each convolution to add the input and output. At the same time, the last convolutional layer is used to further downsample the image features. After the features are extracted to the last layer, the network generates a discriminant output through a global pooling layer and a fully connected layer. In the upsampling stage, the network gradually restores the size of the feature map through upsampling blocks (UpBlock). Each upsampling block includes a transposed convolution, a convolution, and LeakyReLU activation operations. The input is concatenated with the output of the corresponding downsampling layer through a skip connection, and finally, the features pass through a convolutional layer to obtain the output result.
[0027] 3) Deformation Field Network : The deformation field network is used to geometrically align low-dose CT images with normal-dose CT images. It learns the geometric differences between images and generates a deformation field to adjust the pixel positions of low-dose images to be consistent with the structure of normal-dose images. Its specific network architecture is based on the ResUnet architecture, which is divided into a downsampling stage, a residual module stage, and an upsampling stage. In the downsampling stage, the input image first undergoes multiple convolutional operations, followed by a LeakyReLU activation function after each convolution. At the same time, residual connections are accompanied to retain the input information, and the spatial dimension of the feature map is gradually reduced. In the residual module stage, the network further extracts and processes feature information through multiple residual blocks. The residual blocks are transformed through convolution and activation function layers to maintain the compactness and expressiveness of the features. The upsampling stage gradually restores the spatial dimension of the feature map through transposed convolution operations. Each time a transposed convolution is performed, it is concatenated with the corresponding feature map in the downsampling stage through skip connections to ensure the retention of multi-scale information. Finally, the network outputs a two-dimensional deformation field through a convolutional layer, which is used for spatial transformation and alignment of the input image through STN affine transformation.
[0028] Example 1, the training of the low-dose CT image denoising model of the present invention includes the following steps: Step 1, collect low-dose chest CT images and their paired but not strictly aligned normal-dose chest CT images in clinical practice as training samples, where sample examples are as shown in Figure 2 and Figure 3 ; The collected low-dose chest CT images and normal-dose CT images need to be paired one by one, and the corresponding quantities and sizes are the same. It is required that the time interval between acquisitions is as short as possible to ensure that there are no significant differences in the physical state. Finally, the final images are saved in the.dcm format.
[0029] Step 2, randomly divide the paired CT image training samples into a training set and a test set, and construct a dynamic training set TrainLoader with different sizes and a static test set TestLoader with the original size; the specific process is as follows: Step 2.1, read and convert the paired CT image training samples into pydicom.dataset.FileDataset objects suitable for the Python language, and anonymize the information of the pydicom.dataset.FileDataset objects; Step 2.2, convert the pixel_array of the image data in pairs into the Tensor data type suitable for the Pytorch framework, and divide the data into a training set and a test set.
[0030] Step 2.3: Convert the training set data into a torch.utils.data.Dataset object, and use torch.utils.data.DataLoader to construct a training set TrainLoader with dynamically different sizes and a test set TestLoader with the original size. Step 2.3.1: Construction of the dynamic training set TrainLoader. In each Epoch, randomly extract image pairs from the training set, and then dynamically crop these images to obtain 4 cropped images. As an example, the cropping size can be selected or , and thus the input of the training set each time is pairs of images; Step 2.3.2: Construction of the static test set TestLoader. The test set does not require cropping operations and directly uses the original-sized images for verification.
[0031] Step 3: Construct the generator and discriminator of the U-Net GAN. The generator combines multi-scale feature extraction and a deep residual network. Its network structure is divided into a downsampling stage, a feature extraction stage, and an upsampling stage. In the downsampling stage, the input image first undergoes padding and convolution operations, and then enters the multi-scale convolution module, where multi-scale features are extracted in parallel using 3×3, 5×5, and 7×7 convolutional kernels. Subsequently, the features are processed through normalization and ReLU activation, and then downsampling is achieved through a convolutional layer. In the feature extraction stage, the network consists of nine residual modules, each of which includes two padding and convolution operations, accompanied by normalization and activation function layers, and the residual connection adds the input and output. The upsampling stage gradually restores the image size through two transposed convolutions. After each transposed convolution, additional normalization and ReLU activation operations are performed, and finally, the final image is output through a convolutional layer. The Tanh activation function is used to ensure the output range, as shown in Figure 4 ; the discriminator is based on the Encoder-Decoder architecture of U-Net and is divided into global and local discrimination, as shown in Figure 5aAs shown in the figure, the Encoder-Only architecture of the original GAN is replaced with an Encoder-Decoder architecture based on U-Net. The specific architecture of the network is based on the U-Net architecture and is divided into a downsampling stage and an upsampling stage. In the downsampling stage, the input image first passes through multiple downsampling blocks (DownBlock). Each block includes two convolutional operations and LeakyReLU activation operations, and a residual connection is applied after each convolution to add the input and the output. At the same time, the last convolutional layer is used to further downsample the image features. After the feature extraction reaches the last layer, the network generates a discriminant output through a global pooling layer and a fully connected layer. In the upsampling stage, the network gradually restores the size of the feature map through upsampling blocks (UpBlock). Each upsampling block includes a transposed convolution, a convolution, and LeakyReLU activation operations. The input is concatenated with the output of the corresponding downsampling layer through a skip connection, and finally the features pass through a convolutional layer to obtain the output result. Construct a deformation field network , the deformation field network is constructed based on ResUnet and can generate a two-dimensional deformation field for STN affine transformation to obtain a registered image. The structure is as Figure 5b shown. Its specific network architecture is based on the ResUnet architecture and is divided into a downsampling stage, a residual module stage, and an upsampling stage. In the downsampling stage, the input image first passes through multiple convolutional operations. After each convolution, a LeakyReLU activation function is connected, and at the same time, a residual connection is accompanied to retain the input information, and the spatial dimension of the feature map is gradually reduced. In the residual module stage, the network further extracts and processes feature information through multiple residual blocks. The residual blocks are transformed through 1×1 convolutions and activation function layers to maintain the compactness and expressive power of the features. The upsampling stage gradually restores the spatial dimension of the feature map through transposed convolution operations. Each time a transposed convolution is performed, it is concatenated with the corresponding feature map in the downsampling stage through a skip connection to ensure the retention of multi-scale information. Finally, the network outputs a two-dimensional deformation field through a convolutional layer, which is used for spatial transformation and alignment of the input image through STN affine transformation. The normal-dose CT image is set as , and the low-dose CT image is set as , and then the trained network is used for denoising to obtain a denoised image . The specific process is as follows: Step 3.1: Construct a two-stage training strategy. In the first 50,000 epochs, the dynamic training set TrainLoader is used to crop the size to for the first-stage training; in the subsequent 50,000 epochs, the full size is used for the second-stage training; Step 3.2. The generator (denoiser) and the deformation field network are trained to update the parameters simultaneously, and the training of the two and the discriminator is carried out alternately in a non-consistent manner; the discriminator is fixed, and then the generator and the deformation field network are trained to obtain the adversarial loss, the registration loss, etc., and then the generator is updated by backpropagation. Then the generator is fixed, and the discriminator is trained to obtain the discriminant loss for updating the discriminator. Set Batch Size = 8, d_iter = 5, g_iter = 1, that is, the generator and the deformation field network are iterated 5 times per round, and the discriminator is iterated 1 time. The number of input image pairs is . The generator and the deformation field network are more complex and require more iterative updates to converge as soon as possible. Therefore, as much training as possible is needed to avoid the situation where the generator and the deformation field network have not converged yet while the discriminator has learned some other methods to cause the generator mode to collapse. The specific process is as follows: Step 3.2.1. Discriminator training. In the U-Net-based discriminator, the downsampling stage is similar to the traditional discriminator. The input image is globally compressed to capture the global deep information, and the output is ; while the upsampling obtains the local discriminant information at the pixel level of the original image size, which is similar to PatchGAN, and the output is ; at the same time, the Hinge Loss is used as the loss function, and its loss function form can be expressed as:
[0032] where 1 is the boundary value of the Hinge Loss, is the normal-dose CT image, is the low-dose CT image, is the generator (denoiser), and E represents the expectation. To enhance the robustness of the discriminator, the CutMix enhancement method is adopted. The CutMix enhancement method randomly crops a matrix from a part of the real image and replaces this matrix area with the corresponding part of the generated image; at the same time, its corresponding label is proportionally mixed according to the size of the occluded area. The mixed image form is as follows:
[0033] The network parameters of the discriminator are updated iteratively by backpropagation of the gradient.
[0034] Step 3.2.2. Generator and deformation field network training. The input low-dose CT image , passes through the generator to obtain the denoised image , and then is input into the discriminator together with the normal-dose CT image to obtain the adversarial loss of the Hinge Loss:
[0035] In order to better handle the case of non-rigid alignment, and the input deformation field network , a two-dimensional deformation field network can be obtained . Combined with the deformation field network Through the affine transformation of STN, the denoised image after registration of the deformation field network can be obtained , thereby calculating the loss function after registration :
[0036] At the same time, considering the smoothness of the deformation field network, the preservation of edge information of the denoised image, and the loss of its own information of the denoised image, an additional smoothness loss , edge loss and identity loss are respectively set
[0037] Among them, Sobel refers to calculating the edge information of the image through the Sobel operator. The specific calculation process of the Sobel operator is as follows: 1) The Sobel operator convolution kernels, and are the convolution kernels in the horizontal and vertical directions respectively:
[0038] 2) Gradient edge calculation, is the input image, represents the convolution operation:
[0039] Thus, the final generator loss is:
[0040] Among them are the weights of each loss respectively. By setting different hyperparameters, different emphases are placed on the information. Through backpropagation gradient, the network parameters of the generator and the deformation field network are updated iteratively.
[0041] Step 3.3: Through a two-stage training strategy and non-uniform alternating training, the generator can better adapt to multi-scale feature changes and can better handle image structures and denoising tasks with different resolutions. Using the registration loss of the deformation field network, denoising can be performed while ensuring that the structure of the non-aligned region is not affected, thereby obtaining the final low-dose CT image denoising model 。
[0042] Step 4: Finally, evaluate the denoising performance of the denoiser on the test set, and then use the trained denoiser model to perform high-quality denoising processing and image restoration on the low-dose CT images.
[0043] Based on the method of the present invention, the actually collected medical images are processed. 18 sets of paired chest normal-dose CT images and low-dose CT images of clinical patients (with privacy protection) are collected. The low dose is approximately 1 / 5 of the normal dose as a whole. There are 12,873 pairs of images in total, among which 11,997 pairs of images are selected as the training set and 876 pairs of images are selected as the test set. When training, a generator structure based on multi-scale feature fusion and residual network is selected as Figure 4 shown. The low-dose CT images are shown in the first column of Figure 6 , the normal-dose CT images are shown in the second column of Figure 6 , and the denoised CT images are shown in the third column of Figure 6 . To illustrate the effect of the present invention, a control group is trained. The network in the control group does not use deformation field network registration and uses a traditional discriminator for training, and the obtained results are as shown in Figure 7 .
[0044] Comparing Figure 6 and Figure 7 for the network inputs, outputs and the corresponding normal-dose CT images, it can be seen that in the aligned regions, both the control group and the denoising model of the present invention have relatively excellent denoising effects; but for the unaligned regions, obvious large artifacts appear in the control group, and the unaligned tissues are denoised as noise, as shown in the last group of images of Figure 7 . Comparing with the denoising model proposed by the present invention, the original tissue structure is still maintained in the unaligned regions, and the denoising performance is the same as that in the aligned regions. Therefore, the denoising network can achieve good results for denoising low-dose CT images that are not strictly aligned. Through the above method, the present invention discloses a low-dose CT image denoising method based on U-Net GAN and deformation field registration, which is suitable for clinical scenarios with non-strict alignment, can meet the actual use requirements, and has good end-to-end denoising ability.
[0045] Embodiment 2: Based on the technical concept of the method of the present invention, the present invention also provides a low-dose CT image denoising system based on U-Net GAN and deformation field registration, including a data acquisition module and a denoising module; The data acquisition module is used to acquire the original low-dose CT images; The denoising module is based on a low-dose CT image generator. Through the downsampling stage, feature extraction stage, and upsampling stage, in the downsampling stage, the input image first undergoes edge padding and convolution operations, and then enters the multi-scale convolution module to extract multi-scale features. The multi-scale features are processed through normalization and ReLU activation, and an image is output through a convolutional layer; in the feature extraction stage, based on the residual network, two edge padding and convolution operations are performed in each residual module, along with normalization and activation function layers, and then residual connection is performed to add the input and output; the upsampling stage restores the image size through two transposed convolutions. After each transposed convolution, additional normalization and ReLU activation operations are attached, and a denoised low-dose CT image is output through a convolutional layer; low-dose CT image denoising model During training, the low-dose CT image is input into the generator for generative adversarial training. Based on the deformation field network, STN affine transformation is used to introduce registration loss and smoothing loss to align the low-dose CT image and the normal-dose CT image in terms of geometric structure.
[0046] The present invention also provides a computer device, which includes a processor and a memory. The memory is used to store computer-executable programs, and the processor reads and executes the computer-executable programs from the memory. When the processor executes the computer-executable programs, it can implement the low-dose CT image denoising method based on U-Net GAN and deformation field registration of the present invention.
[0047] On the other hand, the present invention provides a computer-readable storage medium in which a computer program is stored. When the computer program is executed by a processor, it can implement the low-dose CT image denoising method based on U-Net GAN and deformation field registration of the present invention.
[0048] The computer device can be a laptop computer, a desktop computer, or a workstation.
[0049] The processor of the present invention can be a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), or a field-programmable gate array (FPGA).
[0050] For the memory of the present invention, it can be an internal storage unit of a laptop computer, a desktop computer, or a workstation, such as a memory or a hard disk; it can also use an external storage unit, such as a mobile hard disk or a flash card.
[0051] A computer-readable storage medium may include a computer storage medium and a communication medium. The computer storage medium includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. The computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), solid state drives (SSD, Solid State Drives), or optical discs, etc. Among them, the random access memory may include resistive random access memory (ReRAM, Resistance Random Access Memory) and dynamic random access memory (DRAM, Dynamic Random Access Memory).
[0052] The present invention provides a low-dose CT image denoising method based on U-Net GAN and deformation field registration and its related implementation schemes. It should be noted that those skilled in the art can modify and equivalently replace each feature of the present invention according to actual needs and technical backgrounds without departing from the technical idea of the present invention. All these modifications and alternative schemes, if not deviating from the substantial content of the present invention, shall be regarded as the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the claims, and all contents falling within the claims are the protected objects of the present invention.
Claims
1. A low-dose CT image denoising method based on U-Net GAN and deformation field registration, characterized in that: include: The original low-dose CT image is input into the trained low-dose CT image denoising model In the process, after the downsampling stage, feature extraction stage and upsampling stage, the downsampling stage extracts features through edge filling and multi-scale convolution; feature In the extraction stage, batch normalization and ReLU activated residual modules are applied for deep processing; In the upsampling stage, deconvolution and normalization operations are used to gradually restore the image resolution, and finally a denoised low-dose CT image is generated as output; Low-dose CT image denoising model During training, the low-dose CT images are input into the generator for generative adversarial training. Based on the deformation field network, the STN affine transformation is used, and the registration loss and smoothing loss are introduced to align the low-dose CT images with the normal-dose CT images in terms of geometric structure.
2. The low-dose CT image denoising method based on U-Net GAN and deformation field registration according to claim 1, characterized in that: Low-dose CT image denoising model The training includes the following steps: Collect clinical low-dose chest CT images and their paired but not strictly aligned normal-dose chest CT images as training samples; The paired CT image training samples are randomly divided into training sets and test sets, and dynamic training sets TrainLoader of different sizes and static test sets TestLoader of original sizes are constructed; Construct a two-stage training strategy. In the first few iterations, the dynamic training set TrainLoader is cut to size The first stage of training is carried out on images of The second stage of training is performed on the images of the generator. The generator and the deformation field network are trained and updated simultaneously, and the training is performed alternately with the discriminator inconsistently. Through a two-stage training strategy and non-uniform alternating training, the structure of the non-aligned area is not affected while denoising, and a low-dose CT image denoising model is obtained. ; The denoising performance of the denoiser is evaluated on the test set. If the evaluation result meets the requirements, the final low-dose CT image denoising model is obtained. ; Using the trained low-dose CT image denoising model De-noising and image restoration are performed on low-dose CT images.
3. The low-dose CT image denoising method based on U-Net GAN and deformation field registration according to claim 2, characterized in that: The discriminator is based on the U-Net architecture. In the downsampling stage, the denoised image and the normal dose CT image are first passed through multiple downsampling blocks. Each downsampling block performs two convolutions and LeakyReLU activation operations, and applies a residual connection after each convolution to add the input and output. At the same time, the last convolution layer downsamples the image features. Then, a global pooling layer and a fully connected layer are used to generate the discriminant output. In the upsampling stage, the size of the feature map is restored through deconvolution, convolution and LeakyReLU activation operations of the upsampling block. The input is concatenated with the corresponding downsampling layer output through a jump connection, and finally the feature is passed through the convolution layer to obtain the discriminant result.
4. The low-dose CT image denoising method based on U-Net GAN and deformation field registration according to claim 2, characterized in that: The deformation field network is based on the ResUnet architecture and is divided into a downsampling stage, a residual module stage, and an upsampling stage. In the downsampling stage, the denoised image is first subjected to multi-layer convolution operations, each layer of convolution is followed by a LeakyReLU activation function, and is accompanied by a residual connection. In the residual module stage, feature information is further extracted and processed through multiple residual blocks, which are transformed through convolution and activation function layers. In the upsampling stage, the spatial dimension of the feature map is restored through a deconvolution operation. Each deconvolution is spliced with the corresponding feature map of the downsampling stage through a jump connection, and a two-dimensional deformation field is output through a convolution layer. The STN affine transformation is used to spatially transform and align the input image to obtain the registered denoised image.
5. According to claim 2, a low-dose CT image denoising method based on U-Net GAN and deformation field registration is characterized in that: The paired CT image training samples are randomly divided into training sets and test sets, and dynamic training sets TrainLoader of different sizes and static test sets TestLoader of original sizes are constructed, including: Read the paired CT image training samples and convert them into pydicom.dataset.FileDataset objects suitable for Python language, and anonymize the information; Convert the image data pixel_array in the pydicom.dataset.FileDataset object into pairs of Tensor data types suitable for the Pytorch framework, and divide the data into training and test sets; Convert the training set data into a torch.utils.data.Dataset object, and use torch.utils.data.DataLoader to construct a training set TrainLoader of dynamic different sizes and a test set TestLoader of original size; TrainLoader dynamically constructs a training set. In each iteration, it randomly extracts image pairs from the training set, and then dynamically crops the randomly extracted images to obtain 4 cropped images. The input of each training set is for images; TestLoader static test set construction, the test set does not need to be cut, and the original size is used directly in sequence Image verification.
6. The low-dose CT image denoising method based on U-Net GAN and deformation field registration according to claim 5, characterized in that: Paired CT images are: paired one by one with the collected low-dose CT images of the chest and normal-dose CT images, with the same number and size, and the interval between the previous and subsequent acquisitions is as short as possible, and the final images are saved in a unified format.
7. The low-dose CT image denoising method based on U-Net GAN and deformation field registration according to claim 2, characterized in that: The generator and the deformation field network are trained and updated simultaneously, and the training is performed inconsistently and alternately with the discriminator. The following steps are included: Set the training parameters, iterate the generator and deformation field network 5 times in each round of training, and iterate the discriminator once; Discriminator training,In the U-Net based discriminator, the downsampling stage globally compresses the input image,captures the global deep information, and outputs ; Upsampling obtains local discriminant information at the pixel level of the original image size, similar to PatchGAN, and the output is ; Using Hinge Loss as the loss function, randomly crop a matrix from part of the real image, and replace the area corresponding to the matrix with the corresponding part in the generated image; at the same time, the labels of the corresponding areas of the matrix are mixed proportionally according to the size of the occluded area, and the network parameters of the iterative discriminator are updated by back-propagating the gradient; Generator and deformation field network training, input low-dose CT image , the denoised image is obtained through the generator , compared with normal dose CT images Input the discriminator and get the adversarial loss of Hinge Loss ,Will and Input deformation field network , and obtain the two-dimensional deformation field , Combined deformation field Through the affine transformation of STN, the denoised image after deformation field registration is obtained. , calculate the loss function after registration , smoothing loss , edge loss and loss of identity; Constructing the generator loss for: in, is the loss function after registration, Hinge Loss is the adversarial loss To smooth the loss, is the edge loss, Loss of identity Different hyperparameters are set for the weights of each loss respectively, and the network parameters of the iterative generator and the deformation field network are updated by back-propagating the gradient.
8. A low-dose CT image denoising system based on U-Net GAN and deformation field registration, characterized in that: It includes a data acquisition module and a denoising module; The data acquisition module is used to acquire original low-dose CT images; The denoising module is based on the low-dose CT image generator. It goes through the downsampling stage, feature extraction stage and upsampling stage. In the downsampling stage, the input image first undergoes edge padding and convolution operations, and then enters the multi-scale convolution module to extract multi-scale features. The multi-scale features are processed by normalization and ReLU activation, and then the image is output through a convolution layer. In the feature extraction stage, based on the residual network, two edge padding and convolution operations are performed in each residual module, normalization and activation function layers are performed at the same time, and then residual connections are performed to add the input and output; In the upsampling stage, the image size is restored through two deconvolutions. Normalization and ReLU activation operations are added after each deconvolution, and denoised low-dose CT images are generated through the output of the convolution layer. Low-dose CT image denoising model During training, the low-dose CT images are input into the generator for generative adversarial training. Based on the deformation field network, the STN affine transformation is used, and the registration loss and smoothing loss are introduced to align the low-dose CT images with the normal-dose CT images in terms of geometric structure.
9. A computer device, characterized in that: The invention comprises a processor and a memory, wherein the memory is used to store a computer executable program, the processor reads the computer executable program from the memory and executes it, and when the processor executes the computer executable program, the low-dose CT image denoising method based on U-Net GAN and deformation field registration as described in any one of claims 1 to 7 can be implemented.
10. A computer-readable storage medium, characterized in that: A computer program is stored in a computer-readable storage medium. When the computer program is executed by a processor, the low-dose CT image denoising method based on U-NetGAN and deformation field registration as described in any one of claims 1 to 7 can be implemented.
Citation Information
Patent Citations
Improved RegGAN low-dose CT image denoising method and related device
CN115760651A