Edge enhancement-based CNN and Transform mixed low-dose CT noise reduction method
By constructing the E-CTformer model and combining it with the U-Net structure and edge enhancement technology, the problems of blurred image details and loss of edge information in low-dose CT image denoising are solved, achieving efficient image denoising and diagnostic support.
Patent Information
- Application Number
- CN202510955448.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-11
- Publication Date
- 2025-09-12
AI Technical Summary
Existing low-dose CT image denoising technologies have problems such as blurred image details, loss of edge information, and unsatisfactory denoising effects, as well as insufficient computational efficiency and generalization performance.
A low-dose CT denoising method based on edge enhancement and hybrid CNN and Transformer is adopted. By constructing the E-CTformer model, combining the U-Net structure, edge enhancement stage, encoding stage and decoding stage, and using convolutional layers, residual blocks, dense blocks, window attention and channel attention techniques, the image edge features are extracted and enhanced, and the model training process is optimized.
It significantly enhances image edge details, improves image quality and diagnostic accuracy, and improves computational efficiency and model generalization performance.
Smart Images

Figure CN120634899A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and more particularly, to a low-dose CT denoising method based on edge enhancement using a hybrid CNN and Transformer. Background Art
[0002] Low-dose CT (Computed Tomography) technology occupies a core position in the field of medical imaging, playing an irreplaceable role in early screening and accurate diagnosis. However, a reduction in radiation dose is often accompanied by an increase in image noise, which poses a severe challenge to image quality and diagnostic accuracy. Traditional image filtering algorithms, such as mean filtering, median filtering, and Gaussian filtering, can reduce noise to a certain extent, but inevitably lose image details and may even lead to the loss of critical medical information. In addition, these algorithms may introduce artifacts during the processing process, further distorting the true appearance of the image.
[0003] With the continuous advancement of technology, denoising methods based on physical and statistical principles have gradually emerged. These methods achieve optimization by establishing complex image generation and noise statistical models. However, these methods are not only computationally complex and slow, but also often lack satisfactory denoising results when processing extremely low-dose CT images. In recent years, the rise of deep learning technology has opened up new avenues for low-dose CT denoising. Models such as convolutional neural networks (CNNs), generative adversarial networks (GANs), and transformers have been widely used in low-dose CT image denoising. For example, CNN-based models such as RIDNet, EDCNN, RED-CNN, and ADNet; GAN-based models such as WGAN, CNCL, and Q-AE; and transformer-based models such as CTformer and TransCT, with their powerful learning and generalization capabilities, strike a balance between denoising and detail preservation. Trained on large amounts of data, these models can accurately identify and effectively remove noise while maximally preserving key image details.
[0004] However, although deep learning technology has made remarkable achievements in low-dose CT noise reduction, it still faces two core challenges: one is how to better retain the edge information of the image while further reducing noise; the other is how to improve the computational efficiency and generalization performance of the model to cope with more diverse low-dose CT image challenges.
[0005] Aiming at the common problems of blurred image details, loss of edge information and unsatisfactory noise reduction effect in existing low-dose CT image noise reduction technologies, a method is proposed that can effectively remove image noise while significantly enhancing image edge details, thereby improving the overall image quality and diagnostic accuracy. Summary of the Invention
[0006] In view of the shortcomings of the existing technology, the purpose of the present invention is to provide a low-dose CT denoising method based on edge enhancement, which is a hybrid CNN and Transformer and can effectively remove image noise while significantly enhancing image edge details, thereby improving the overall image quality and diagnostic accuracy.
[0007] To solve the above technical problems, the present invention is achieved through the following technical solutions:
[0008] The present invention provides a low-dose CT denoising method based on edge enhancement using a hybrid CNN and Transformer, comprising the following steps:
[0009] S110, data acquisition: collect low-dose CT (LDCT) image data of patients and form training and test sets with normal-dose CT image (NDCT) data. The data format is DICOM.
[0010] S120, data preprocessing: preprocessing the LDCT image, setting pixel values outside the effective area of the CT scan (i.e., outside the boundary) to zero, and performing normalization within the range of [-1024, 3072], and then saving the image data in npy format;
[0011] S130, Model Construction: Construct E-CTformer (Edgeenhanced-hybrid CNN-Transformer) based on edge enhancement for denoising low-dose CT images;
[0012] S140, Model training: Train the E-CTformer model on the training set to learn the features and mapping relationships required for low-dose CT image denoising;
[0013] S150, Model Testing: Evaluate the performance of the trained E-CTformer model on the test set to obtain denoising results for low-dose CT images.
[0014] As a preferred technical solution of the present invention, the specific steps of data preprocessing in step S120 are as follows:
[0015] S210, removing invalid areas in the LDCT image (boundary processing): identifying valid areas (i.e., scanning areas) in the image, and setting pixel values in the image that exceed the valid scanning area to zero;
[0016] S220, Data Normalization: Grayscale values in CT images usually have a certain range, but different equipment or scanning conditions may result in different grayscale ranges. Therefore, all grayscale values are mapped to a unified range of [-1024, 3072];
[0017] S230, saving data: saving in npy format: after completing the above preprocessing, the processed image data is saved in NumPy format (.npy file).
[0018] As a preferred technical solution of the present invention, the normalization formula in step S220 is as follows:
[0019]
[0020] Among them, X norm is the normalized image data, X is the original image data, X max The maximum value is 3072, X min The minimum value is -1024.
[0021] As a preferred technical solution of the present invention, in step S130, a low-dose CT denoising method based on edge enhancement hybrid CNN-Transformer (E-CTformer) is constructed. The network structure uses U-Net as the backbone network, including an input end, an edge enhancement stage, an encoding stage (Encoder), a decoding stage (Decoder), and an output end, as follows:
[0022] S310: At the input end, the image after data preprocessing is used as the input of the network;
[0023] S320, in the edge enhancement stage, edge features are extracted from the EdgeNet network, which consists of two convolutional layers, two residual blocks (RB) and one dense block (DB);
[0024] S330, encoding stage (Encoder), including three basic modules (Basic Block), three CNN-Transformer hybrid modules (CNN-Transformer hybrid Block, CTHB) and two Concat operations;
[0025] 340. Decoder stage, including two basic blocks, three CNN-Transformer hybrid blocks (CTHB), and a subtraction operation;
[0026] S350: At the output end, the output result of the decoding stage in S340 is used as the final noise reduction result.
[0027] As a preferred technical solution of the present invention, the convolution layer in step S320 is used to change the spatial dimension of the image, and RB and DB are used to extract image features; each RB is composed of two 3×3 convolutional layers and one ReLu, and each DB is composed of N groups of 3×3 convolutional layers and ReLu and one 1×1 convolutional layer; the specific calculation process is: the input image passes through the first convolution (Conv) block, the second residual block (RB), the third dense block (DB), the fourth residual block (RB), and the fifth convolution (Conv) block in sequence, and then subtracts it from the input image to finally obtain the edge feature map extracted by EdgeNet.
[0028] As a preferred technical solution of the present invention, each Basic Block in step S330 is composed of two 3×3 convolutional layers and two ReLu; each CTHB is composed of a window attention mechanism (Window Attention), two convolutional layers, a CBAM, a multi-layer perceptron (MLP), and two layer normalization (Layer Norm);
[0029] Among them, CBAM consists of channel attention and spatial attention; channel attention consists of maximum pooling, average pooling and multi-layer perceptron; spatial attention consists of maximum pooling, average pooling and convolutional layer.
[0030] As a preferred technical solution of the present invention, the specific calculation process of step S330 is: concat the input image and the edge image, sequentially passing through the first Basic Block, the second CTHB, the third CTHB, the fourth downsampling, the fifth edge image downsampling, the sixth concat the fourth output and the fifth output, the seventh Basic Block, the eighth CTHB, and the ninth BasicBlock to obtain the encoder feature map.
[0031] As a preferred technical solution of the present invention, the specific calculation process of the decoding stage of step 340 is as follows: the feature map output in the decoding stage passes through the first CTHB in sequence, the second is the first output and the seventh output of the encoder pass through the Basic Block in the decoder together, the third is upsampling, the fourth CTHB, the fifth is the first output and the second output of the encoder pass through the CTHB in the decoder together, the sixth Basic Block, and the seventh is subtracted from the input image to obtain the output result image.
[0032] As a preferred technical solution of the present invention, in step S140, the model is trained using the low-dose CT image dataset and the normal-dose CT image dataset prepared in step S120; the specific execution process is: inputting the training set into the network, and obtaining the prediction result of the current iteration through forward propagation calculation; comparing this prediction result with the true label (normal-dose CT image), applying the loss function to calculate the loss value, and evaluating the gap between the prediction result and the true value; using the stochastic gradient descent optimization algorithm to calculate the gradient of the loss value with respect to the network parameters, and updating the network weights through back propagation; this process is repeated until the preset error standard is met or other stopping conditions are reached, thereby obtaining a trained model; the low-dose CT test dataset and the normal-dose CT test dataset are used to verify and evaluate the performance of the model.
[0033] As a preferred technical solution of the present invention, the loss function in step S140 is the MSE loss function L mse , MSE is used to calculate the mean square error of each pixel between the generated image and the clean image, and achieves accurate measurement through comparison and matching; its expression is:
[0034]
[0035] Among them, X i represents the input image, f(X i ) represents the predicted image, Y i is the target image, and N represents the number of images.
[0036] The advantages of the present invention are:
[0037] This invention overcomes the defects of image detail loss caused by traditional filtering algorithms and the high computational complexity and slow processing speed of denoising methods based on physical and statistical models. By combining the hybrid CNN-Transformer structure in deep learning with edge enhancement technology, it achieves high-quality denoising processing of low-dose CT images, providing clearer and more accurate image support for medical imaging diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1This is a comparison diagram of the different algorithms of the present invention and E-CTformer.
[0039] Figure 2 It is the overall flow chart of the present invention.
[0040] Figure 3 This is a flow chart of the data preprocessing method of the present invention.
[0041] Figure 4 This is the network structure diagram of the E-CTformer of the present invention.
[0042] Figure 5 This is the RB structure diagram of the present invention.
[0043] Figure 6 This is the DB structure diagram of the present invention.
[0044] Figure 7 This is the Basic Block structure diagram of the present invention.
[0045] Figure 8 This is a structural diagram of the CTHB module of the present invention.
[0046] Figure 9 This is the structural diagram of the Window Attention module of the present invention.
[0047] Figure 10 This is a structural diagram of the CBAM module of the present invention.
[0048] Figure 11 This is the structural diagram of the channel attention module of the present invention.
[0049] Figure 12 This is the structural diagram of the spatial attention module of the present invention.
[0050] Figure 13 The following is a specific flow chart for calculating an embodiment of the present invention. DETAILED DESCRIPTION
[0051] The principles and features of the present invention are described below with reference to the accompanying drawings. The examples provided are intended only to illustrate the present invention and are not intended to limit the scope of the present invention. The following paragraphs describe the present invention in more detail by way of example with reference to the accompanying drawings. It should be noted that the drawings are all in a very simplified form and are not to exact scale, and are only used for the purpose of conveniently and clearly illustrating the embodiments of the present invention.
[0052] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this invention pertains. The terms used in this specification of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0053] The present invention provides a low-dose CT denoising method based on edge enhancement hybrid CNN-Transformer. The overall process is as follows: Figure 2 As shown, steps S110-S150 are included, which are specifically as follows:
[0054] S110, data acquisition: collect low-dose CT (LDCT) image data of patients and form training and test sets with normal-dose CT image (NDCT) data. The data format is DICOM.
[0055] S120, data preprocessing: preprocessing the LDCT image, setting pixel values outside the effective area of the CT scan (i.e., outside the boundary) to zero, and performing normalization within the range of [-1024, 3072], and then saving the image data in npy format;
[0056] S130, Model Construction: Construct E-CTformer (Edgeenhanced-hybrid CNN-Transformer) based on edge enhancement for denoising low-dose CT images;
[0057] S140, Model training: Train the E-CTformer model on the training set to learn the features and mapping relationships required for low-dose CT image denoising;
[0058] S150, Model Testing: Evaluate the performance of the trained E-CTformer model on the test set to obtain denoising results for low-dose CT images.
[0059] The data preprocessing method in step S120 is as follows: Figure 3 As shown, it includes steps S210-S240. Specifically, S210 removes invalid areas in the LDCT image; S220 normalizes the data in the range of [-1024, 3072]; S230 stores the normalized data in npy format. The following is a description of each module:
[0060] S210: Remove invalid areas (boundary processing): Identify the valid area (i.e., scanning area) in the image and set the pixel values in the image that exceed the valid scanning area to zero;
[0061] S220: Data normalization: Grayscale values in CT images usually have a certain range, but different equipment or scanning conditions may result in different grayscale ranges. Therefore, all grayscale values are mapped to a unified range of [-1024, 3072];
[0062] The normalization formula is as follows:
[0063]
[0064] Among them, X norm is the normalized image data, X is the original image data, X max The maximum value is 3072, X min The minimum value is -1024;
[0065] S230: Save data: Save in NumPy format: After completing the above preprocessing, save the processed image data in NumPy format (.npy file). By saving in NumPy format, batch data loading can be conveniently performed, avoiding the need to reprocess the original image data each time training.
[0066] In step S130, a low-dose CT denoising method based on edge enhancement hybrid CNN-Transformer (E-CTformer) is constructed, and the network structure is as follows: Figure 4 shown.
[0067] U-Net is used as the backbone network, including the input end, edge enhancement stage, encoding stage (Encoder), decoding stage (Decoder) and output end, as follows:
[0068] S310: At the input end, the image after data preprocessing is used as the input of the network.
[0069] S320, in the edge enhancement stage, edge features are extracted from the EdgeNet network, which consists of two convolutional layers, two residual blocks (RB) and a dense block (DB). Among them, the convolutional layer is used to change the spatial dimension of the image, and RB and DB are used to extract image features. Figure 5 As shown, each RB consists of two 3×3 convolutional layers and a ReLu, and DB is as follows Figure 6 As shown in the figure, each DB consists of N groups of 3×3 convolutional layers and ReLu and one 1×1 convolutional layer. The specific calculation process is as follows: the input image passes through the first convolutional (Conv) block, the second residual block (RB), the third dense block (DB), the fourth residual block (RB), and the fifth convolutional (Conv) block in sequence, and then subtracts the input image to obtain the edge feature map extracted by EdgeNet.
[0070] S330, the encoding stage (Encoder), includes three basic modules (Basic Block), three CNN-Transformer hybrid modules (CNN-Transformer hybrid Block, CTHB) and two Concat operations. Basic Block is as follows Figure 7 As shown, each Basic Block consists of two 3×3 convolutional layers and two ReLu layers. Figure 8 As shown, each CTHB consists of a window attention mechanism (Window Attention) as shown in Figure 9 As shown in the figure, it consists of two convolutional layers, a CBAM, a multi-layer perceptron (MLP), and two layer normalization layers. Among them, CBAM consists of channel attention and spatial attention, as shown in the figure. Figure 10 ; Channel attention consists of maximum pooling, average pooling and multi-layer perceptron, such as Figure 11 As shown in ; spatial attention consists of maximum pooling, average pooling and convolutional layers, as shown in Figure 12 shown.
[0071] Specific calculation process: Concat the input image with the edge image, pass through the first Basic Block, the second CTHB, the third CTHB, the fourth downsampling, the fifth edge image downsampling, the sixth concat the fourth output with the fifth output, the seventh Basic Block, the eighth CTHB, and the ninth Basic Block to obtain the encoder feature map.
[0072] 340. The decoding stage (Decoder) includes two basic modules (Basic Block), three CNN-Transformer hybrid modules (CNN-Transformer hybrid Block, CTHB) and a subtraction operation. The above modules are the same as those in the encoder structure. Basic Block is as follows Figure 7 As shown; CTHB as Figure 8 As shown; Window Attention Figure 9 CBAM Figure 10 ; Channel attention such as Figure 11 As shown; spatial attention as Figure 12 shown.
[0073] Specific calculation process: The feature map output in the decoding stage passes through the first CTHB in sequence, the second is the first output and the seventh output of the encoder pass through the Basic Block in the decoder, the third is upsampling, the fourth CTHB, the fifth is the first output and the second output of the encoder pass through the CTHB in the decoder, the sixth Basic Block, and the seventh is subtracted from the input image to obtain the output result image.
[0074] S350: At the output end, the output result of the decoding stage in S340 is used as the final noise reduction result.
[0075] In step S140, the model is trained using the low-dose CT image dataset and the normal-dose CT image dataset prepared in step S120. The specific execution process is: first, the training set is input into the network, and the prediction result of the current iteration is obtained by forward propagation calculation. Then, this prediction result is compared with the true label (normal-dose CT image), and the loss value is calculated by applying the loss function to evaluate the gap between the prediction result and the true value. Next, the gradient of the loss value to the network parameters is calculated using the stochastic gradient descent optimization algorithm, and the weights of the network are updated through back propagation. This process is repeated until the preset error standard is met or other stopping conditions are reached, thereby obtaining a trained model. Finally, the test set is used to verify and evaluate the performance of the model. The loss function in the above steps is the MSE loss function L mse ,MSE is used to calculate the mean square error of each pixel between the generated image and the clean image, and achieve accurate measurement through comparison and matching.
[0076] Its expression is:
[0077]
[0078] Among them, X i represents the input image, f(X i ) represents the predicted image, Y i is the target image, and N represents the number of images.
[0079] In conjunction with the accompanying drawings and implementation cases, the present invention will describe in detail the implementation process of the AAPM-Mayo dataset used in the "NIH-AAPM-Mayo Clinic Low-dose CT Challenge" held by the Mayo Clinic in 2016.
[0080] Example:
[0081] The AAPM-Mayo dataset is used for low-dose CT image denoising. The specific process in this embodiment is as follows: Figure 13 The specific steps are as follows:
[0082] The AAPM-Mayo dataset covers NDCT (normal-dose CT) and low-dose CT image data of 10 anonymous patients. The NDCT images were acquired by scanning at a tube voltage of 120 kV and a tube current of 200 mAs, and the slice thickness of all images was uniformly 3 mm. The images were stored in DICOM format with a resolution of 512 × 512 pixels, and the CT value range was set between -140 HU and 260 HU. In this embodiment, 1816 pairs of images were selected for the experiment, of which 1612 pairs of images were used as training sets, and the remaining 204 pairs of images constituted the test set. In addition, the data was preprocessed to ensure its suitability for subsequent analysis and processing.
[0083] Taking an abdominal low-dose CT image as an input, the E-CTformer model is constructed and trained. The specific steps are as follows:
[0084] First, an edge enhancement network is constructed. Edge features are extracted in the EdgeNet network. The network consists of two convolutional layers, two residual blocks (RB) and a dense block (DB). Figure 5 As shown, each RB consists of two 3×3 convolutional layers and a ReLu, and DB is as follows Figure 6 As shown in the figure, each DB consists of N groups of 3×3 convolutional layers and ReLu and one 1×1 convolutional layer. The specific calculation process is as follows: the input image passes through the first convolutional (Conv) block, the second residual block (RB), the third dense block (DB), the fourth residual block (RB), and the fifth convolutional (Conv) block in sequence, and then subtracts the input image to obtain the edge feature map extracted by EdgeNet.
[0085] Secondly, the encoding phase is constructed. It includes three basic blocks (Basic Block), three CNN-Transformer hybrid blocks (CNN-Transformer hybrid Block, CTHB) and two Concat operations. Basic Block is as follows Figure 7 As shown, each Basic Block consists of two 3×3 convolutional layers and two ReLu layers. Figure 8 As shown, each CTHB consists of a window attention mechanism (Window Attention) as shown in Figure 9 As shown in the figure, it consists of two convolutional layers, a CBAM, a multi-layer perceptron (MLP), and two layer normalization layers. Among them, CBAM consists of channel attention and spatial attention, as shown in the figure. Figure 10 ; Channel attention consists of maximum pooling, average pooling and multi-layer perceptron, such as Figure 11As shown in ; spatial attention consists of maximum pooling, average pooling and convolutional layers, as shown in Figure 12 The specific calculation process is as follows: the input image is concat with the edge image, and then passes through the first Basic Block, the second CTHB, the third CTHB, the fourth downsampling, the fifth edge image downsampling, the sixth concat of the fourth output and the fifth output, the seventh Basic Block, the eighth CTHB, and the ninth BasicBlock to obtain the encoder feature map.
[0086] Next, the decoding stage (Decoder) is constructed, which includes two basic modules (Basic Block), three CNN-Transformer hybrid modules (CNN-Transformer hybrid Block, CTHB) and a subtraction operation. The above modules are the same as those in the encoder structure. Basic Block is as follows Figure 7 As shown; CTHB as Figure 8 As shown; Window Attention Figure 9 CBAM Figure 10 ; Channel attention such as Figure 11 As shown; spatial attention as Figure 12 The specific calculation process is as follows: the feature map output in the decoding stage passes through the first CTHB, the second is the first output and the seventh output of the encoder pass through the Basic Block in the decoder, the third is upsampling, the fourth CTHB, the fifth is the first output and the second output of the encoder pass through the CTHB in the decoder, the sixth Basic Block, and the seventh is subtracted from the input image to obtain the output result image.
[0087] Finally, the Mayo dataset was used to train the model. The specific steps are as follows: the LDCT image is input into the network to calculate the result of this round of iterative calculation, the result is compared with the corresponding NDCT image, and the loss value is calculated using the loss function. The loss function uses MSE Loss. At the same time, the network model parameters are optimized, and then the gradient is calculated using the stochastic gradient descent optimizer and the weights in the network are updated through back propagation. The above process is iterated until the error requirements are met to obtain the network training model. Finally, an abdominal low-dose CT image is input into the model to obtain the final denoising result. The experimental results of the E-CTformer algorithm in this embodiment are shown in Table 1.
[0088] Table 1 Experimental results of E-CTformer on the Mayo dataset
[0089]
[0090] Based on this structure, the PSNR improvement of E-CTformer demonstrates that it more effectively suppresses noise during denoising while avoiding detail loss caused by oversmoothing. Compared to using only CNN or pure Transformer models, the hybrid structure combined with edge enhancement strategies significantly improves pixel-level reconstruction accuracy.
[0091] E-CTformer better preserves the edge and texture details of the image during the denoising process through the edge enhancement module (EdgeNet) and the hybrid attention mechanism (CBAM in CTHB), thereby improving the structural similarity.
[0092] The improvement in VIF indicates that E-CTformer better preserves visual information (such as edges and textures) that the human eye is sensitive to while reducing noise. This is due to the edge enhancement module's protection of high-frequency details and the hybrid structure's balance of global and local features.
[0093] In summary, edge features are extracted through convolutional layers, residual blocks (RBs), and dense blocks (DBs), and then concatenated with the input image, forcing the network to focus on edge information at an early stage. Compared to methods that do not explicitly design edge enhancement modules (such as CTformer), E-CTformer improves SSIM and VIF by 0.0051 and 0.0126, respectively, directly demonstrating its improved ability to preserve edge detail.
[0094] 3×3 convolution extracts local features, preventing the Transformer from neglecting local details. Window attention captures global context, compensating for the limited field of view of CNNs. Compared to pure CNNs (EDCNN) and pure Transformers (CTformer), E-CTformer achieves PSNR improvements of 0.66dB and 0.48dB, respectively, validating the advantages of the hybrid architecture.
[0095] Global information is aggregated through maximum pooling and average pooling, and channel weights are adaptively adjusted. Spatial position perception is enhanced through convolutional layers, focusing on key areas. The introduction of CBAM further improves the model in SSIM and VIF, indicating that its refined adjustment of feature maps effectively improves structural consistency and visual quality.
[0096] The above are only specific embodiments of the present invention, but the technical features of the present invention are not limited thereto. Any simple changes, equivalent substitutions, or modifications based on the present invention to solve substantially the same technical problems and achieve substantially the same technical effects are all included in the scope of protection of the present invention.
Claims
1. A low-dose CT denoising method based on edge enhancement using a hybrid CNN and Transformer, characterized in that: The steps include: S110, data acquisition: collect low-dose CT (LDCT) image data of patients and form training and test sets with normal-dose CT image (NDCT) data. The data format is DICOM. S120, data preprocessing: preprocessing the LDCT image, setting pixel values outside the effective area of the CT scan (i.e., outside the boundary) to zero, and performing normalization within the range of [-1024, 3072], and then saving the image data in npy format; S130, Model Construction: Construct E-CTformer (Edgeenhanced-hybrid CNN-Transformer) based on edge enhancement for denoising low-dose CT images; S140, Model training: Train the E-CTformer model on the training set to learn the features and mapping relationships required for low-dose CT image denoising; S150, Model Testing: Evaluate the performance of the trained E-CTformer model on the test set to obtain denoising results for low-dose CT images.
2. The low-dose CT denoising method based on edge enhancement using a hybrid CNN and Transformer according to claim 1, characterized in that: The specific steps of data preprocessing in step S120 are as follows: S210, removing invalid areas in the LDCT image (boundary processing): identifying valid areas (i.e., scanning areas) in the image, and setting pixel values in the image that exceed the valid scanning area to zero; S220, Data Normalization: Grayscale values in CT images usually have a certain range, but different equipment or scanning conditions may result in different grayscale ranges. Therefore, all grayscale values are mapped to a unified range of [-1024, 3072]; S230, saving data: saving in npy format: after completing the above preprocessing, the processed image data is saved in NumPy format (.npy file).
3. The low-dose CT denoising method based on edge enhancement using a hybrid CNN and Transformer according to claim 2, characterized in that: The normalization formula in step S220 is as follows: Among them, X norm is the normalized image data, X is the original image data, X max The maximum value is 3072, X min The minimum value is -1024.
4. The low-dose CT denoising method based on edge enhancement using a hybrid CNN and Transformer according to claim 1, characterized in that: In step S130, a low-dose CT denoising method based on edge enhancement hybrid CNN-Transformer (E-CTformer) is constructed. The network structure uses U-Net as the backbone network, including an input end, an edge enhancement stage, an encoding stage (Encoder), a decoding stage (Decoder), and an output end, as follows: S310: At the input end, the image after data preprocessing is used as the input of the network; S320, in the edge enhancement stage, edge features are extracted from the EdgeNet network, which consists of two convolutional layers, two residual blocks (RB) and one dense block (DB); S330, encoding stage (Encoder), including three basic modules (Basic Block), three CNN-Transformer hybrid modules (CNN-Transformer hybrid Block, CTHB) and two Concat operations; 340. Decoder stage, including two basic blocks, three CNN-Transformer hybrid blocks (CTHB), and a subtraction operation; S350: At the output end, the output result of the decoding stage in S340 is used as the final noise reduction result.
5. The low-dose CT denoising method based on edge enhancement using a hybrid CNN and Transformer according to claim 4, characterized in that: In step S320, the convolution layer is used to change the spatial dimension of the image, and the RB and DB are used to extract image features; each RB is composed of two 3×3 convolution layers and one ReLu, and each DB is composed of N groups of 3×3 convolution layers and ReLu and one 1×1 convolution layer; Specific calculation process: The input image passes through the first convolution (Conv) block, the second residual block (RB), the third dense block (DB), the fourth residual block (RB), and the fifth convolution (Conv) block in sequence, and then subtracts it from the input image to finally obtain the edge feature map extracted by EdgeNet.
6. The low-dose CT denoising method based on edge enhancement hybrid CNN and Transformer according to claim 4, characterized in that: In step S330, each Basic Block consists of two 3×3 convolutional layers and two ReLu layers; each CTHB consists of a window attention mechanism (Window Attention), two convolutional layers, a CBAM, a multi-layer perceptron (MLP), and two layer normalization (Layer Norm); Among them, CBAM consists of channel attention and spatial attention; channel attention consists of maximum pooling, average pooling and multi-layer perceptron; spatial attention consists of maximum pooling, average pooling and convolutional layer.
7. The low-dose CT denoising method based on edge enhancement using a hybrid CNN and Transformer according to claim 6, characterized in that: The specific calculation process of step S330 is as follows: The input image is concat with the edge image, and passes through the first Basic Block, the second CTHB, the third CTHB, the fourth downsampling, the fifth edge image downsampling, the sixth concat of the fourth output and the fifth output, the seventh Basic Block, the eighth CTHB, and the ninth Basic Block to obtain the encoder feature map.
8. The low-dose CT denoising method based on edge enhancement using a hybrid CNN and Transformer according to claim 4, characterized in that: The specific calculation process of the decoding stage in step 340 is as follows: The feature map output in the decoding stage passes through the first CTHB in sequence, the second is the first output and the seventh output of the encoder passing through the Basic Block in the decoder, the third is upsampling, the fourth CTHB, the fifth is the first output and the second output of the encoder passing through the CTHB in the decoder, the sixth Basic Block, and the seventh is subtracted from the input image to obtain the output result image.
9. The low-dose CT denoising method based on edge enhancement hybrid CNN and Transformer according to claim 1, characterized in that: In step S140, the model is trained using the low-dose CT image dataset and the normal-dose CT image dataset prepared in step S120; The specific execution process is as follows: input the training set into the network, and obtain the prediction result of the current iteration through forward propagation calculation; Compare this prediction result with the true label (normal dose CT image), apply the loss function to calculate the loss value, and evaluate the gap between the prediction result and the true value; Use the stochastic gradient descent optimization algorithm to calculate the gradient of the loss value with respect to the network parameters, and update the network weights through backpropagation; This process is repeated until the preset error standard is met or other stopping conditions are reached, thus obtaining a trained model. A low-dose CT test dataset was used to validate and evaluate the model's performance.
10. The low-dose CT denoising method based on edge enhancement hybrid CNN and Transformer according to claim 9, characterized in that: The loss function in step S140 is the MSE loss function L mse ,MSE is used to calculate the mean square error of each pixel between the generated image and the clean image, and achieve accurate measurement through comparison and matching; Its expression is: Among them, X i represents the input image, f(X i ) represents the predicted image, Y i is the target image, and N represents the number of images.