Denoising Method and System for Low-Dose CT Images, and Electronic Device
Through the dual-channel neural network model, the sparse Transformer block and residual attention module are used to denoise the low-dose CT images, which solves the problems of high noise and blurred details in the low-dose CT images, and achieves the effect of efficient denoising and detail retention.
Patent Information
- Application Number
- CN202510595399.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-05-09
AI Technical Summary
Due to the reduction of radiation dose, low-dose CT images lead to high image noise and blurred image details, which affect the accurate identification of lesions.
The two-channel neural network model is used to denoise low-dose CT images. The first channel is based on the U-Net architecture composed of sparse Transformer blocks, and the second channel is based on the U-Net architecture composed of residual attention modules, and the output results are fused through convolution processing.
It realizes efficient denoising of low-dose CT images, improves the retention ability and visual quality of image details, and provides doctors with a more accurate diagnostic basis.
Smart Images

Figure CN120107110B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to a denoising method and system for low-dose CT images, and an electronic device. Background Art
[0002] Computed Tomography (CT) technology plays a crucial role in modern medical diagnosis. CT technology scans through X-rays and reconstructs tomographic images of objects, and has become an important tool for non-invasive diagnosis.
[0003] Since CT technology uses X-rays, it may have a certain impact on the human body. For infants or pregnant women, low-dose CT images have become the choice for this group. Low-dose CT reduces radiation by reducing X-rays, but it also reduces the number of photons received by the detector, thereby increasing the noise of the image. The noise sources of low-dose CT images include the reduction of radiation dose, the sensitivity of the detector, the working state of electronic components, and external mechanical vibrations. These factors may all cause the image details to be blurred, thereby affecting the accurate identification of lesions. Summary of the Invention
[0004] The present invention provides a denoising method and system for low-dose CT images, and an electronic device, to solve the defects of large noise and blurred image details in low-dose CT images. The solution of the present application can automatically denoise low-dose CT images through a neural network, and the denoising effect is better.
[0005] The present invention provides a denoising method for low-dose CT images, including:
[0006] Obtain a target image, where the target image is a low-dose CT image to be processed;
[0007] Input the target image into a pre-trained dual-channel neural network model for denoising processing to obtain an output result;
[0008] Wherein, the dual-channel neural network model includes a first channel and a second channel. The first channel is a U-Net architecture based on a plurality of sparse Transformer blocks, and the second channel is a U-Net architecture based on a plurality of residual attention modules; the output result is obtained by convolution processing after fusing the first output of the first channel and the second output of the second channel.
[0009] According to the denoising method of low-dose CT images provided by the present invention, the first channel includes a first sparse Transformer block, a second sparse Transformer block, a third sparse Transformer block, a fourth sparse Transformer block, and a fifth sparse Transformer block. Among them, the first sparse Transformer block, the second sparse Transformer block, and the third sparse Transformer block constitute an encoder, and the fourth sparse Transformer block and the fifth sparse Transformer block constitute a decoder;
[0010] The second channel includes a first residual attention module, a second residual attention module, a third residual attention module, a fourth residual attention module, and a fifth residual attention module. Among them, the first residual attention module, the second residual attention module, and the third residual attention module constitute an encoder, and the fourth residual attention module and the fifth residual attention module constitute a decoder.
[0011] According to the denoising method of low-dose CT images provided by the present invention, the sparse Transformer block includes a mixed-scale feed-forward network layer and a Top-K sparse attention layer;
[0012] The mixed-scale feed-forward network layer is used to capture and integrate multi-scale features and identify noise;
[0013] The Top-K sparse attention layer is used to retain the feature information of the target image and reduce the interference of irrelevant information on the denoising process.
[0014] According to the denoising method of low-dose CT images provided by the present invention, the residual attention module includes a residual block and a spatial attention module;
[0015] The residual block includes several standard convolutional layers and activation function layers, and the residual block is used to extract local features;
[0016] The spatial attention module is used to extract the statistical information of the target image through global maximum pooling and global average pooling to supplement the global feature information.
[0017] According to the denoising method of low-dose CT images provided by the present invention, the training process of the dual-channel neural network model includes:
[0018] Inputting the training set data into a pre-constructed initial model to obtain a first output image. The training set data includes low-dose CT images and standard-dose CT images, and the first output image is the image after denoising the low-dose CT image;
[0019] Calculate the similarity between the first output image and the standard-dose CT image through a loss function;
[0020] Update the initial model based on the similarity.
[0021] According to the denoising method of low-dose CT images provided by the present invention, the loss function includes a reconstruction loss, a perceptual loss, and a structural similarity loss;
[0022] The reconstruction loss is used to calculate the pixel value similarity between the first output image and the standard-dose CT image;
[0023] The perceptual loss is used to calculate the feature similarity between the first output image and the standard-dose CT image;
[0024] The structural similarity loss is used to calculate the structural similarity between the first output image and the standard-dose CT image.
[0025] According to the denoising method of low-dose CT images provided by the present invention, the loss function is a weighted sum of the reconstruction loss, the perceptual loss, and the structural similarity loss.
[0026] According to the denoising method of low-dose CT images provided by the present invention, after inputting the target image into a pre-trained dual-channel neural network model for denoising processing to obtain an output result, it further includes:
[0027] Evaluate the denoised image using a preset evaluation metric. If the evaluation is unqualified, re-perform the denoising process.
[0028] The present invention also provides a denoising system for low-dose CT images, including:
[0029] An image loading module for loading a target image, where the target image is a low-dose CT image to be processed;
[0030] A denoising module for inputting the target image into a pre-trained dual-channel neural network model for denoising processing to obtain an output result;
[0031] Among them, the dual-channel neural network model includes a first channel and a second channel. The first channel is a U-Net architecture based on a number of sparse Transformer blocks, and the second channel is a U-Net architecture based on a number of residual attention modules; the output result is obtained by fusing the first output of the first channel and the second output of the second channel and then performing convolution processing.
[0032] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the denoising method of any one of the above-mentioned low-dose CT images is implemented.
[0033] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the denoising method of any one of the above-mentioned low-dose CT images is implemented.
[0034] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the denoising method of any one of the above-mentioned low-dose CT images is implemented.
[0035] In the denoising method of the low-dose CT images provided by the present invention, the low-dose CT images can be automatically denoised based on a pre-trained neural network. At the same time, the applied neural network is a dual-channel neural network model, which integrates the advantages of high denoising efficiency of the sparse Transformer mechanism and high denoising accuracy of the residual attention mechanism. Further, both the first channel and the second channel are of the U-Net architecture, which can also achieve excellent image data analysis effects when the data volume is small. Especially when applied to the field of CT image denoising in this application, accurate image denoising processing can be realized with very little data volume, providing accurate diagnostic basis for doctors. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0037] Figure 1 is a schematic flowchart of the denoising method of the low-dose CT images provided by the embodiments of the present invention;
[0038] Figure 2 is a schematic structural diagram of the dual-channel neural network model provided by the embodiments of the present invention;
[0039] Figure 3 is a schematic structural diagram of the sparse Transformer block provided by the embodiments of the present invention;
[0040] Figure 4 is a schematic structural diagram of the residual attention module provided by the embodiments of the present invention;
[0041] Figure 5 is a schematic structural diagram of the denoising system of the low-dose CT images provided by the embodiments of the present invention;
[0042] Figure 6 It is a schematic diagram of the physical structure of the electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0043] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.
[0044] Figure 1 It is a schematic flowchart of a method for denoising low-dose CT images provided by an embodiment of the present invention.
[0045] As Figure 1 shown, this embodiment provides a method for denoising low-dose CT images, including:
[0046] Step 101, obtain a target image, where the target image is a low-dose CT image to be processed;
[0047] Step 102, input the target image into a pre-trained dual-channel neural network model for denoising processing to obtain an output result;
[0048] Among them, the dual-channel neural network model includes a first channel and a second channel. The first channel is a U-Net architecture based on a number of sparse Transformer blocks, and the second channel is a U-Net architecture based on a number of residual attention modules; the output result is obtained by convolution processing after fusing the first output of the first channel and the second output of the second channel.
[0049] As Figure 2 shown, the target image input into the dual-channel neural network model can first be processed by a Sobel operator to obtain an intermediate image. After extracting edge features through Sobel, the intermediate image can be concatenated with the original image information. Then, the concatenated information is projected by an initial feature through a convolutional layer and then input into the first channel and the second channel. This convolutional layer can be a 3×3 convolutional kernel, and the process of feature projection can be from 2 channels to 96 channels, that is, the 2 channels of the original image information are projected and converted into 96 channels.
[0050] In practical applications, as Figure 2As shown, the first channel and the second channel can both be U-Net architectures composed of 5-level modules. In the U-Net architecture, the encoder performs 2 downsampling operations through 3×3 convolutional downsampling with a stride of 2, and the decoder restores the resolution through two 3×3 transposed convolutions with a stride of 2 and performs skip connections with the corresponding hierarchical features.
[0051] Furthermore, as Figure 2 shown, after being processed by five sparse Transformer blocks in the first channel, it can then pass through a convolutional layer for projection from 96 channels to 1 channel. Similarly, after being processed by five residual attention modules in the second channel, it can then pass through a convolutional layer for projection from 96 channels to 1 channel. The output after projection by the convolutional layer in the first channel and the output after projection by the convolutional layer in the second channel can be concatenated and fused, and the fused features are mapped to a single-channel output layer to effectively retain the low-frequency information of the original image, thereby ensuring that the denoised image not only has higher visual quality but also can accurately reflect the pathological state.
[0052] From the perspective of cognitive resonance, the residual attention module (Residual Attention Block, RAB) and the sparse transformer model (Sparse Transformer, STB) maximize the extraction of meaningful features through collaborative interaction, and at the same time balance detail retention and image integrity in low-dose CT image processing, achieving significant improvement in the effect and detail retention ability of medical image denoising while improving computational efficiency.
[0053] Specifically, the first channel includes a first sparse Transformer block, a second sparse Transformer block, a third sparse Transformer block, a fourth sparse Transformer block, and a fifth sparse Transformer block. Among them, the first sparse Transformer block, the second sparse Transformer block, and the third sparse Transformer block constitute the encoder, and the fourth sparse Transformer block and the fifth sparse Transformer block constitute the decoder;
[0054] The second channel includes a first residual attention module, a second residual attention module, a third residual attention module, a fourth residual attention module, and a fifth residual attention module. Among them, the first residual attention module, the second residual attention module, and the third residual attention module constitute the encoder, and the fourth residual attention module and the fifth residual attention module constitute the decoder.
[0055] Figure 3 is a schematic structural diagram of the sparse Transformer block provided by an embodiment of the present invention.
[0056] As Figure 3 shown, the sparse Transformer block includes a mixed-scale feed-forward network layer and a Top-K sparse attention layer;
[0057] The mixed-scale feed-forward network layer is used to capture and integrate multi-scale features and identify noise;
[0058] The Top-K sparse attention layer is used to retain the feature information of the target image and reduce the interference of irrelevant information on the noise reduction process.
[0059] Specifically, the sparse Transformer block is fully called the Sparse Transformer module (STB). STB takes Top-k Sparse Attention as the core. Through the sparse attention mechanism of the Top-K sparse attention layer, it focuses on the important local features in the image, avoiding the interference of irrelevant information that may exist in traditional denoising methods. The sparse attention mechanism can dynamically select the features most significant for diagnosis and aggregate them, thus effectively retaining key medical details during the denoising process. To further improve the denoising effect, a mixed-scale feed-forward network layer (Mixed-Scale Feed-Forward Network, MSFN) is introduced. Through the collaborative action of the multi-scale convolutional kernels of MSFN, it captures feature information at different scales, helps accurately identify and remove high-frequency noise, and at the same time maintains the structural integrity of low-frequency artifacts. Finally, through the design of the double residual connection of Top-k Sparse Attention and MSFN, STB can achieve multi-scale feature fusion of local perception and global sparse attention.
[0060] Figure 4 is a schematic structural diagram of the residual attention module provided by an embodiment of the present invention.
[0061] As Figure 4 shown, the residual attention module includes a residual block and a spatial attention module;
[0062] The residual block includes a number of standard convolutional layers and activation function layers, and the residual block is used to extract local features;
[0063] The spatial attention module is used to extract the statistical information of the target image through global max pooling and global average pooling to supplement global feature information.
[0064] The Residual Attention Block (RAB) combines residual learning and the attention mechanism. Among them, the spatial attention module can weight the feature map through global maximum pooling (GMP) and global average pooling (GAP) to extract the global context information of the image. Further, the residual attention module strengthens the learning and enhancement of multi-level features through cross-level skip connections. These skip connections not only help capture deeper features but also retain shallow-level information, enabling the network to more accurately adjust the feature weights at different scales, thereby achieving spatially adaptive feature fusion and noise suppression. Specifically, the statistical information extracted by the GMP and GAP modules helps the module focus on information-dense regions in the image. By means of residual learning, the problem of information loss in traditional convolutional networks is avoided, and fine enhancement and adaptive optimization of features are achieved.
[0065] In an exemplary embodiment, the training process of the dual-channel neural network model includes:
[0066] Inputting the training set data into a pre-constructed initial model to obtain a first output image. The training set data includes low-dose CT images and standard-dose CT images, and the first output image is the image after denoising the low-dose CT image;
[0067] Calculating the similarity between the first output image and the standard-dose CT image through a loss function;
[0068] Updating the initial model based on the similarity.
[0069] In practical applications, the data in the training set can be the dataset publicly released by the 2016 NIH-AAPM-Mayo Clinic Low-Dose CT Grand Challenge, and the 3 mm slice data in this dataset is used for training and testing. Specifically, the 3 mm slice data includes 2378 3 mm low-dose (one-quarter dose) and standard-dose CT image slices from 10 anonymous patients. Slices are selected from the first 9 patients, with a total of 2167 slices for training, and 211 slices are selected from the 10th patient for testing. And data augmentation and reduction of the computational burden are performed by cutting into small pieces (patches).
[0070] In an exemplary embodiment, the loss function includes a reconstruction loss, a perceptual loss, and a structural similarity loss;
[0071] The reconstruction loss is used to calculate the pixel value similarity between the first output image and the standard-dose CT image;
[0072] The perceptual loss is used to calculate the feature similarity between the first output image and the standard-dose CT image;
[0073] The structural similarity loss is used to calculate the structural similarity between the first output image and the standard-dose CT image.
[0074] In an exemplary embodiment, the loss function is a weighted sum of the reconstruction loss, the perceptual loss, and the structural similarity loss.
[0075] In practical applications, the reconstruction loss is the main loss function of the model, which is used to ensure the pixel-level similarity between the denoising result and the reference full-dose CT image. The mean squared error (MSE) is used to measure the difference between the predicted image and the real image, which conforms to the following formula (1):
[0076] (1)
[0077] Where, N is the total number of pixels in the image, and are the pixel values of the normal-dose CT image and the output image, respectively.
[0078] The perceptual loss is used to improve the visual quality of the output image. The perceptual loss captures more distinguishable features by calculating the difference between the output image and the normal-dose CT image in the high-level feature space, which conforms to the following formula (2):
[0079] (2)
[0080] Where, represents the feature extraction function of the pre-trained VGG network at the l layer, is the size of the feature map of this layer, l is taken as the 16th layer.
[0081] SSIM is an evaluation index for evaluating the similarity of images in terms of brightness, contrast, and structure, with a range of [-1, 1], and 1 indicates exactly the same. The SSIM loss conforms to the following formula (3):
[0082] (3)
[0083] Where, and are the average values of the images and , and are the variances, is the covariance. and are small constants for stability.
[0084] After separately calculating the reconstruction loss, perceptual loss, and structural similarity loss, the total loss function can be calculated through the following formula (4):
[0085] (4)
[0086] where 、 and are the weight coefficients for adjusting the influence of each loss term. Exemplarily, , , 。
[0087] In an exemplary embodiment, the target image is input into a pre-trained two-channel neural network model for denoising to obtain an output result. After that, it further includes:
[0088] Evaluating the denoised image using a preset evaluation metric. If the evaluation is unqualified, re-perform the denoising process until a qualified denoised image is obtained.
[0089] In practice, the evaluation metrics can include Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index (SSIM), and Root Mean Square Error (RMSE).
[0090] In practical applications, the denoising method provided by the solution of the present application can be compared with five existing methods, namely RED-CNN, WGAN-VGG, CTFormer, EDCNN, and HFormer. At the same time, it can also be compared with the initial low-dose CT image and the standard-dose CT image.
[0091] Table 1 below shows the comparison results of the average PSNR, SSIM, and RMSE values of the present application and different models. As shown in Table 1, compared with the existing denoising methods, the method of the present application achieves the highest denoising effect, and at the same time, a very high degree of retention of the image detail structure is achieved. The method of the present invention provides higher reliability for clinical diagnosis.
[0092] Table 1 Comparison Table of Denoising Effects
[0093]
[0094] Table 2 below shows the loss ablation comparison of the solution of this application. It can be seen from Table 2 that, compared with using only the reconstruction loss, the addition of the SSIM loss and the perceptual loss can better guide the learning of the network, achieving a better denoising effect and retaining more detailed structures.
[0095] Table 2 Loss Ablation Comparison Table
[0096]
[0097] Table 3 below shows the structure ablation comparison of the solution of this application. It can be seen from Table 3 that the effective fusion of the sparse attention block and the residual attention module further enhances the denoising ability of the network.
[0098] Table 3 Structure Ablation Comparison Table
[0099]
[0100] The denoising system for low-dose CT images provided by the present invention will be described below. The denoising system for low-dose CT images described below can be mutually corresponding and referred to the denoising method for low-dose CT images described above.
[0101] Figure 5 It is a schematic structural diagram of the denoising system for low-dose CT images provided by an embodiment of the present invention.
[0102] As Figure 5 shown, the denoising system for low-dose CT images provided in this embodiment includes:
[0103] An image loading module 501, configured to load a target image, where the target image is a low-dose CT image to be processed;
[0104] A denoising module 502, configured to input the target image into a pre-trained dual-channel neural network model for denoising processing to obtain an output result;
[0105] Wherein, the dual-channel neural network model includes a first channel and a second channel. The first channel is a U-Net architecture based on a plurality of sparse Transformer blocks, and the second channel is a U-Net architecture based on a plurality of residual attention modules; the output result is obtained by convolution processing after the first output of the first channel and the second output of the second channel are fused.
[0106] In an exemplary embodiment, the first channel includes a first sparse Transformer block, a second sparse Transformer block, a third sparse Transformer block, a fourth sparse Transformer block, and a fifth sparse Transformer block. Among them, the first sparse Transformer block, the second sparse Transformer block, and the third sparse Transformer block constitute an encoder, and the fourth sparse Transformer block and the fifth sparse Transformer block constitute a decoder;
[0107] The second channel includes a first residual attention module, a second residual attention module, a third residual attention module, a fourth residual attention module, and a fifth residual attention module. Among them, the first residual attention module, the second residual attention module, and the third residual attention module constitute an encoder, and the fourth residual attention module and the fifth residual attention module constitute a decoder.
[0108] In an exemplary embodiment, the sparse Transformer block includes a mixed-scale feed-forward network layer and a Top-K sparse attention layer;
[0109] The mixed-scale feed-forward network layer is used to capture and integrate multi-scale features and identify noise;
[0110] The Top-K sparse attention layer is used to retain the feature information of the target image and reduce the interference of irrelevant information on the noise reduction process.
[0111] In an exemplary embodiment, the residual attention module includes a residual block and a spatial attention module;
[0112] The residual block includes a plurality of standard convolutional layers and activation function layers, and the residual block is used to extract local features;
[0113] The spatial attention module is used to extract the statistical information of the target image through global max pooling and global average pooling to supplement the global feature information.
[0114] In an exemplary embodiment, the training process of the dual-channel neural network model includes:
[0115] Inputting the training set data into a pre-constructed initial model to obtain a first output image. The training set data includes low-dose CT images and standard-dose CT images, and the first output image is the image after noise reduction of the low-dose CT image;
[0116] Calculating the similarity between the first output image and the standard-dose CT image through a loss function;
[0117] Update the initial model based on the similarity.
[0118] In an exemplary embodiment, the loss function includes a reconstruction loss, a perceptual loss, and a structural similarity loss;
[0119] The reconstruction loss is used to calculate the pixel value similarity between the first output image and the standard-dose CT image;
[0120] The perceptual loss is used to calculate the feature similarity between the first output image and the standard-dose CT image;
[0121] The structural similarity loss is used to calculate the structural similarity between the first output image and the standard-dose CT image.
[0122] In an exemplary embodiment, the loss function is a weighted sum of the reconstruction loss, the perceptual loss, and the structural similarity loss.
[0123] In an exemplary embodiment, an evaluation module is further included. The evaluation module is specifically configured to: evaluate the denoised image using a preset evaluation metric, and if the evaluation is unqualified, perform denoising processing again.
[0124] Figure 6 An exemplary physical structure diagram of an electronic device is shown as Figure 6 shown. The electronic device may include: a processor 610, a communication interface 620, a memory 630, and a communication bus 640. Among them, the processor 610, the communication interface 620, and the memory 630 communicate with each other through the communication bus 640. The processor 610 can call the logical instructions in the memory 630 to execute the denoising method for low-dose CT images, and the method includes:
[0125] Obtain a target image, where the target image is a low-dose CT image to be processed;
[0126] Apply a pre-trained dual-channel neural network model to perform denoising processing on the target image;
[0127] Input the target image into a pre-trained dual-channel neural network model for denoising processing to obtain an output result;
[0128] Among them, the dual-channel neural network model includes a first channel and a second channel. The first channel is a U-Net architecture based on a number of sparse Transformer blocks, and the second channel is a U-Net architecture based on a number of residual attention modules. The output result is obtained by performing convolution processing after fusing the first output of the first channel and the second output of the second channel.
[0129] In addition, when the logical instructions in the above-mentioned memory 630 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs, etc., which can store program codes.
[0130] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the denoising method for low-dose CT images provided by the above-mentioned various methods. The method includes:
[0131] Obtain a target image, where the target image is a low-dose CT image to be processed;
[0132] Input the target image into a pre-trained dual-channel neural network model for denoising processing to obtain an output result;
[0133] Among them, the dual-channel neural network model includes a first channel and a second channel. The first channel is a U-Net architecture based on a number of sparse Transformer blocks, and the second channel is a U-Net architecture based on a number of residual attention modules. The output result is obtained by performing convolution processing after fusing the first output of the first channel and the second output of the second channel.
[0134] On the other hand, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the denoising method for low-dose CT images provided by the above-mentioned various methods. The method includes:
[0135] Obtain a target image, where the target image is a low-dose CT image to be processed;
[0136] Input the target image into a pre-trained dual-channel neural network model for denoising processing to obtain an output result;
[0137] Wherein, the dual-channel neural network model includes a first channel and a second channel. The first channel is a U-Net architecture based on a number of sparse Transformer blocks, and the second channel is a U-Net architecture based on a number of residual attention modules; the output result is obtained by convolution processing after fusing the first output of the first channel and the second output of the second channel.
[0138] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.
[0139] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods of each embodiment or some parts of the embodiments.
[0140] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.
Claims
1. A method for denoising low-dose CT images, characterized in that: include: Acquiring a target image, wherein the target image is a low-dose CT image to be processed; Inputting the target image into a pre-trained dual-channel neural network model for denoising to obtain an output result; The dual-channel neural network model includes a first channel and a second channel, the first channel is a U-Net architecture based on a plurality of sparse Transformer blocks, and the second channel is a U-Net architecture based on a plurality of residual attention modules; the output result is obtained by fusion of the first output of the first channel and the second output of the second channel through convolution processing; The first channel includes a first sparse Transformer block, a second sparse Transformer block, a third sparse Transformer block, a fourth sparse Transformer block, and a fifth sparse Transformer block, wherein the first sparse Transformer block, the second sparse Transformer block, and the third sparse Transformer block constitute an encoder, and the fourth sparse Transformer block and the fifth sparse Transformer block constitute a decoder; The second channel includes a first residual attention module, a second residual attention module, a third residual attention module, a fourth residual attention module and a fifth residual attention module, wherein the first residual attention module, the second residual attention module and the third residual attention module constitute an encoder, and the fourth residual attention module and the fifth residual attention module constitute a decoder; The sparse Transformer block includes a mixed-scale feed-forward network layer and a Top-K sparse attention layer; The mixed-scale feedforward network layer is used to capture and integrate multi-scale features and identify noise; The Top-K sparse attention layer is used to retain the characteristic information of the target image and reduce the interference of irrelevant information on the noise reduction process; The residual attention module includes a residual block and a spatial attention module; The residual block includes several standard convolutional layers and activation function layers, and the residual block is used to extract local features; The spatial attention module is used to extract statistical information of the target image through global maximum pooling and global average pooling to supplement the global feature information.
2. The method for denoising low-dose CT images according to claim 1, characterized in that: The training process of the dual-channel neural network model includes: Inputting the training set data into a pre-built initial model to obtain a first output image, wherein the training set data includes a low-dose CT image and a standard-dose CT image, and the first output image is a denoised image of the low-dose CT image; Calculating the similarity between the first output image and the standard dose CT image by using a loss function; The initial model is updated based on the similarity.
3. The method for denoising low-dose CT images according to claim 2, characterized in that: The loss function includes reconstruction loss, perceptual loss and structural similarity loss; The reconstruction loss is used to calculate the pixel value similarity between the first output image and the standard dose CT image; The perceptual loss is used to calculate the feature similarity between the first output image and the standard dose CT image; The structural similarity loss is used to calculate the structural similarity between the first output image and the standard-dose CT image.
4. The method for denoising low-dose CT images according to claim 3, characterized in that: The loss function is a weighted sum of reconstruction loss, perceptual loss and structural similarity loss.
5. The method for denoising low-dose CT images according to claim 1, characterized in that: The target image is input into a pre-trained dual-channel neural network model for denoising to obtain an output result, and then the method further includes: The denoised image is evaluated using the preset evaluation index. If the evaluation is unqualified, the denoising process is repeated until a qualified denoised image is obtained.
6. A low-dose CT image denoising system, used to execute the low-dose CT image denoising method according to any one of claims 1 to 5, characterized in that: include: An image loading module, used for loading a target image, wherein the target image is a low-dose CT image to be processed; A denoising module, used for inputting the target image into a pre-trained dual-channel neural network model for denoising to obtain an output result; Among them, the dual-channel neural network model includes a first channel and a second channel, the first channel is a U-Net architecture based on several sparse Transformer blocks, and the second channel is a U-Net architecture based on several residual attention modules; the output result is the first output of the first channel and the second output of the second channel fused together and obtained through convolution processing.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that: When the processor executes the program, the low-dose CT image denoising method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Motion recognition system and method based on time sequence aggregation and gating Transform
CN116824694A
Lung CT image segmentation method based on Transform and U-Net
CN119722707A