Low-dose CT image denoising method based on window mixed attention

Through the low-dose CT image denoising method based on window mixed attention, the U-shaped architectural model of window attention module and channel self-attention mechanism is used to solve the problems of low-dose CT image noise and artifacts, achieving high-quality image denoising processing, and improving diagnostic accuracy.

CN119991492APending Publication Date: 2025-05-13HANGZHOU NORMAL UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510133484.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Due to the reduction of radiation dose, image quality decays, and more noise and artifacts appear, affecting the accuracy of diagnosis, especially the diagnosis of early lesions with small areas and subtle shapes.

Method used

Using a low-dose CT image denoising method based on window mixed attention, a U-shaped architecture image denoising model is constructed by constructing a model sample set containing low-dose CT images, including the input projection layer, the encoder, the bottleneck layer, the decoder and the output layer, and the multi-scale feature information is captured using the window attention module and the channel self-attention mechanism.

Benefits of technology

Effectively denoising low-dose CT images, improve image quality, preserve image structure and texture details, obtain images with similar quality to conventional dose CT images, and improve the diagnostic accuracy of small-area lesions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991492A_ABST
    Figure CN119991492A_ABST
Patent Text Reader

Abstract

The invention discloses a low-dose CT (Computed Tomography) image denoising method based on window mixed attention. According to the image denoising method, noise of an image is removed by constructing an image denoising model; the image denoising model adopts a U-shaped framework and comprises an input projection layer, an encoder, a bottleneck layer, a decoder and an output layer, in the encoder part, the provided module is used in combination with a down-sampling module to complete extraction of image information, and in the decoder part, image features output by all layers are efficiently fused, and in combination with jump connection, full restoration of the overall structure and texture details of an image is achieved. Windowing operation is carried out through the window attention module forming the encoder and the decoder, the capturing range of features is reasonably planned, context information can be efficiently captured, global features and local features are captured through two paths in the window attention module respectively, and efficient processing of image information of different scales is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of image denoising, and in particular relates to a low-dose CT image denoising method based on window mixed attention. Background Art

[0002] Computed tomography (CT) is a reliable and non-invasive medical imaging method that helps to detect pathological abnormalities in the human body, such as tumors, cardiovascular diseases, lung nodules, internal injuries and fractures. In addition to diagnosis, CT is also very useful in guiding various clinical treatments, such as radiotherapy and surgery. However, during repeated CT scans, X-ray radiation may cause point mutations, chromosome translocations and gene fusions, increasing the risk of cancer. Therefore, low-dose CT scanning schemes are now widely used. The purpose is to reduce the X-ray radiation dose to a smaller range than conventional dose CT scans by optimizing scanning parameters while ensuring the quality of CT imaging, thereby maintaining high-quality CT imaging quality levels while protecting the patient's physiological health.

[0003] The reduction of X-ray radiation dose can be achieved by setting the scanning parameters of the CT scanner, which can effectively reduce the X-ray radiation dose penetrating the human body. The radiation dose of CT images generated in this way is lower than that of conventional dose images, and such CT images are called low-dose computed tomography (LDCT) images.

[0004] However, reducing the radiation dose of X-rays may lead to image quality degradation, which is mainly reflected in the appearance of more noise and artifacts in the image. The quality degradation of CT images will seriously affect the accuracy of diagnosis, especially the diagnosis of early lesions with small areas and subtle shapes. Therefore, it is very necessary to study and analyze medical image denoising technology, so that the noise area and subtle structure texture can be accurately analyzed at a lower radiation dose, and the image structure and texture details can be retained as much as possible while the noise area is efficiently denoised, so as to obtain CT images with similar quality to conventional dose CT images.

[0005] How to obtain clear LDCT images using less X-ray doses is of great significance for medical diagnosis. However, the multiple noise points in LDCT images are unavoidable at the current hardware level. Therefore, the research on denoising tasks of LDCT images is of great significance, but it also faces a series of challenges.

[0006] With the rapid development of deep learning technology, image processing methods based on deep learning have achieved excellent results due to their data-driven and high-performance characteristics, showing a broad space for development. The development of CNN and the rise of Transformer have had a profound impact on the field of image denoising. The application of deep learning technology to LDCT image denoising tasks faces several major challenges, such as blurred texture details caused by over-smoothing, scarcity of real data sets, and low differentiation between local texture and noise artifacts of organ tissues. Based on the above content, the current LDCT image denoising research still has the following major challenges:

[0007] First, CT images have the characteristics of high similarity and high precision, which requires the denoising algorithm to be able to distinguish between the original content of the image and noise artifacts and perform efficient denoising operations.

[0008] Second, influenced by the CT imaging principle, various types of noise and artifacts will inevitably be generated during the imaging process. That is, the factors that cause the deterioration of CT image quality are complex, which poses a challenge to the denoising task of LDCT images.

[0009] Third, there is a lack of sufficient real data sets. Due to the reasons stated previously, it is usually difficult to obtain LDCT images corresponding to normal-dose computed tomography (NDCT) images, and artificially adding noise can only simulate LDCT images that are close to those generated by low-dose CT scanning schemes in real situations. Therefore, the denoising performance of deep learning denoising methods based on simulated data sets in the clinical application stage remains to be tested. Summary of the invention

[0010] The object of the present invention is to provide a low-dose CT image denoising method based on window mixed attention, which is used to effectively denoise LDCT images.

[0011] The present invention provides a low-dose CT image denoising method based on window mixed attention, which comprises the following steps: Step 1: construct a model sample set containing low-dose computed tomography images, and label the images in the model sample set; Step 2: construct an image denoising model; the image denoising model adopts a U-shaped architecture, which includes an input projection layer, an encoder, a bottleneck layer, a decoder and an output layer; the input projection layer processes the image input to the image denoising model and then inputs it into the encoder; The encoder includes multiple layers of encoding units, and the decoder includes multiple layers of decoding units; the encoding units and decoding units in the same layer are connected by skipping; the bottleneck layer is connected in series between the last layer of encoding units and the first layer of decoding units; the output layer is used to process the feature map output by the decoder to obtain the output image of the image denoising model; the encoding unit, the decoding unit and the bottleneck layer each include two serially connected window attention modules; Step 3: Use the model sample set to train the image denoising model, and use the trained image denoising model to perform image denoising on the low-dose computed tomography images.

[0012] Preferably, the window attention module includes a parallel window attention path, a channel attention path and a fusion module; the window attention path splits the feature map of the input window attention module into multiple windows, and extracts the features of each window to obtain local features; the channel attention path extracts the global features of the feature map of the input window attention module, and splices them with the local features and then inputs them into the fusion module.

[0013] Preferably, the fusion module specifically performs ordinary convolution and deep convolution processing on the input features in sequence, and then adds them to the feature map input by the window attention module to obtain intermediate features; the intermediate features are processed by layer normalization and multi-scale fusion feedforward network in sequence, and then added to the intermediate features to obtain the output features of the fusion module.

[0014] Preferably, the window attention path includes a windowing module, a channel self-attention module, a multi-head self-attention module and a de-windowing module connected in sequence.

[0015] Preferably, the channel self-attention module includes three branches with the same structure; the image input to the channel self-attention module is subjected to layer normalization processing and then input into the three branches respectively, and ordinary convolution and deep convolution processing are performed in sequence to obtain a query vector, a key vector and a value vector; the query vector and the key vector are multiplied and the result after softmax activation function processing is multiplied with the value vector to obtain a preliminary fusion result; the preliminary fusion result is convolved and added to the input image of the channel self-attention module to obtain the output features of the channel self-attention module.

[0016] Preferably, the channel attention path comprises a downsampling layer, a first convolutional attention module, a second convolutional attention module and an upsampling layer connected in sequence; the output features of the channel self-attention module are sequentially de-windowed and down-sampled, and the addition result of the features and the output features of the first convolutional attention module is used as the input of the second convolutional attention module.

[0017] Preferably, the first convolutional attention module and the second convolutional attention module have the same structure, both including layer normalization, a first convolutional layer, a depth convolutional layer, an activation function, a channel attention module and a second convolutional layer connected in sequence; the output of the convolutional attention module is the result of adding the output features of the second convolutional layer to the input features of the convolutional attention module.

[0018] Preferably, the multi-scale fusion feedforward network comprises three branches connected in parallel; each branch comprises a convolution layer, a first depth convolution layer, a first GELU activation function, a second depth convolution layer and a second GELU activation function connected in sequence; Each branch uses a convolution layer, a first deep convolution layer and a GELU activation function in turn to process features and fuses the processed features with the features processed by the other two branches to obtain intermediate features of each branch; each branch continues to process the intermediate features using a second deep convolution layer and a second GELU activation function to obtain output features; the output features of the three branches are fused and input into the convolution layer to obtain the output features of the multi-scale fusion feedforward network.

[0019] Preferably, the encoding units in the encoder except the last layer of encoding units all input their output feature maps to the next layer of encoding units after downsampling; the last layer of encoding units inputs their output feature maps to the bottleneck layer after downsampling; The input of the decoding units in the decoder except the first-layer decoding units is the concatenation of the output feature map of the previous layer decoding unit after upsampling and the processing result of the corresponding encoding unit; the input of the first-layer decoding unit is the concatenation of the output feature map of the bottleneck layer after upsampling and the processing result of the corresponding encoding unit.

[0020] Preferably, the output layer comprises a serially connected window attention module and an output projection layer; the output of the output layer is connected to the input of the image denoising model via a residual connection.

[0021] The present invention has the following beneficial effects: 1. The present invention learns the local features in each non-overlapping window through the window attention module, and uses the channel self-attention mechanism and the spatial self-attention mechanism to more efficiently capture multi-scale feature information, thereby solving the defect that the original VisionTransformer cannot capture local features well; at the same time, each time the feature map passes through a downsampling layer, its number of channels will be expanded to twice the original, and the length and width will be reduced to half of the original, thereby solving the problem that the Transformer module is difficult to capture multi-scale features due to the same size of the input feature map and the output feature map.

[0022] 2. The present invention uses the window attention module as a component of the encoder and decoder, and uses the windowing module to window the features of the input window attention module through the window attention path in the window attention module, and further divides the features into multiple non-overlapping windows, so that the channel self-attention module can implement the self-attention mechanism in each window, thereby enhancing the model's ability to capture local features; at the same time, the global interaction of the entire feature map range is captured through the channel attention path, thereby improving the model's ability to process the global context, which helps to maintain the stability of the overall image structure. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 It is a structural diagram of the image denoising model in the present invention.

[0024] Figure 2 This is a structural diagram of the window attention module in the present invention.

[0025] Figure 3 Figure 2 is the structural diagram of the channel self-attention module in the window attention path.

[0026] Figure 4 Figure 2 is the structural diagram of the convolutional attention module in the channel attention path.

[0027] Figure 5 This is the structural diagram of the multi-scale fusion feed-forward network in the window attention module; Figure 6 Normal dose computed tomography scan for true label.

[0028] Figure 7 This is the denoising result image output by the image denoising model. DETAILED DESCRIPTION

[0029] The present invention will be further described below in conjunction with the accompanying drawings.

[0030] A low-dose CT image denoising method based on window mixed attention includes the following steps: Step 1: Data preprocessing The AAPM dataset and the QIN-LUNG-CT dataset were used to construct the model sample set. The AAPM dataset was used as the real dataset, which included 2378 abdominal CT scan images with a size of 512x512. The QIN-LUNG-CT dataset was used as the simulation dataset, which included 3841 conventional dose CT images with a size of 512×512 provided by 46 patients. The images in the model sample set were flipped, scaled, and cropped, and the image size was unified. The processed model sample set was divided into a training set and a test set at a ratio of 9:1.

[0031] Step 2: Build the image denoising model WHformer like Figure 1 As shown in the figure, the image denoising model adopts a U-shaped architecture, which includes an input projection layer (Input ProjectionLayer), an encoder, a bottleneck layer, a decoder and an output layer; the input projection layer is used to convert the LDCT image of the input image denoising model into an unfolded two-dimensional patch sequence, and input the output result into the encoder to achieve the dimensional mapping of the input image to meet the input standard of the encoder. The encoder includes three layers of encoding units connected in sequence, which can reduce the width of the feature map input to the encoder by eight times; except for the last layer of encoding units, the other encoding units input their respective output feature maps to the next layer of encoding units after downsampling; the last layer of encoding units inputs the output feature map into the bottleneck layer after downsampling, so as to achieve full fusion of abstract features.

[0032] The decoder includes three layers of decoding units corresponding to the encoding units one by one; the input of the decoding units other than the first layer of decoding units is the concatenation of the output feature map of the previous layer of decoding units after upsampling and the processing result of the corresponding encoding unit; the input of the first layer of decoding units is the concatenation of the output feature map of the bottleneck layer after upsampling and the processing result of the corresponding encoding unit; the encoding unit and the decoding unit each include two series-connected window attention modules (WHTB). The output layer is used to process the feature image output by the decoder, complete the final feature restoration process, and obtain the denoised image; the output layer includes a series-connected window attention module and an output projection layer; the output of the output layer is connected to the input of the image denoising model through a residual connection, so that the network can learn the residual image between the LDCT image and the NDCT image during the training process, thereby improving the training efficiency of the network. Finally, by subtracting the LDCT image from the residual image learned by the network, a denoised image highly similar to the NDCT image can be generated.

[0033] like Figure 2 As shown in Figure 1, the window attention module includes the window attention path (WHAP), the channel attention path (GCAP) and the fusion module. The window attention path includes the window partition module (Window Partition), the CMSA module (Channel Self-Attention Module), the multi-head self-attention module (MSA) and the anti-window module (Window Reversion) connected in sequence. The window module processes the input feature map as follows: If the size of the input feature map is , set the window size to , then the number of acquisitions after windowing is , size is Features .

[0034] like Figure 3 As shown, the CMSA module includes three branches with the same structure; the three branches are characterized by Features after layer normalization As input, and the features Carry out 1 in sequence 1 convolution and 3 3 Deep convolution processing to obtain (query vector), (key vector), (value vector); 1 1 Convolution is used to aggregate and fuse pixel-level feature information across channels, 3 3D convolution is used to encode spatial feature information channel by channel. The relevant formula is as follows: in, They are the query vector, key vector and value vector output by the CMSA module respectively; Indicates 1 on different branches 1. Convolution operation; Indicates 3 on different branches 3 depth-wise convolution operations.

[0035] The query vector and key vector The result after multiplication and softmax activation function processing and value vector Multiply to get the initial fusion result ; 1 1 Fusion results after convolution processing With features Add and obtain the local features of the input image as the output features of the channel self-attention module , and its related formula is as follows: in, represents a learnable scaling parameter used to control size.

[0036] Features output by the channel self-attention module The size is , its computational complexity will not increase quadratically with the increase of the spatial resolution of the feature map, thus improving the performance of the model to a certain extent.

[0037] The multi-head self-attention module is used to adjust the features output by the channel self-attention module. Processing, get the size Features , the relevant formula definition is as follows: in, They are the query vector, key vector, and value vector output by the multi-head self-attention module respectively; The projection matrices representing the query vector, key vector, and value vector respectively; Represents the key vector Length; Indicates relative position deviation.

[0038] like Figure 4 As shown in Figure 1, the channel attention path includes the downsampling layer, the first convolutional attention module (CAB), the second convolutional attention module, and the upsampling layer. The downsampling layer is connected by 4 layers with a stride of 2 and a padding of 1. 4 convolutions input the features of the channel attention path (Size is ) is converted into features (Size is ); The input of the second convolutional attention module is the feature The result of the de-windowing (Window Reversion) and down-sampling processing is added to the output result of the first convolutional attention module; the upsampling layer is used to obtain the same features as the input Global features of the same size, i.e., output features of the channel attention path .

[0039] The two convolutional attention modules have the same structure, which consists of layer normalization, the first convolutional layer (convolution kernel size is 1×1), 3×3 depthwise convolution, GELU activation function, CA (channel attention module), and the second convolutional layer (convolution kernel size is 1×1). First, a 1×1 convolution is performed to aggregate and fuse the pixel-level feature information across channels. Then, a 3×3 deep convolution is performed to encode the spatial feature information channel by channel. The feature information is obtained through a GELU activation function. , and then sent to the channel attention module to capture global information, and finally use 1×1 convolution to aggregate channel information and add it to the features of the input convolution attention module to obtain the output features ; In the channel attention module, first the feature The multi-layer perceptron (MLP) in the channel attention module is then used to calculate the cross-channel attention, and then the obtained attention is combined with the feature Weighted multiplication is performed on the channel; the convolutional attention module can capture larger-scale image features while reducing the amount of calculation. The flow process of features in the first convolutional attention module can be defined as:

[0040] in, represents the GELU activation function; Representation 3 Depth convolution of 3; and Represents the first convolutional layer and the second convolutional layer; Representation layer normalization; is the Sigmoid nonlinear activation function; and represents two fully convolutional layers; represents the global average pooling operation; Represents the channel attention module.

[0041] The flow process of the input features of the second convolutional attention module can be defined as: in, Features The result after de-windowing and downsampling; It means that by using 4 with a stride of 2 and a fill of 1 4 downsampling layers implemented by convolution; Indicates the de-windowing operation; The features output by the second convolutional attention module; This means that by using 4 4 Upsampling layer implemented by transposed convolution; Represents the channel attention module.

[0042] The fusion module includes 1 connected in sequence 1 convolutional layer, 3 3 deep convolutional layers, layer normalization, and multi-scale fusion feedforward network (MFN); the output of the window attention path and the channel attention path is in the channel dimension ( ) is used as the input of the fusion module; after 1 1 convolution and 3 3. Deep convolution fuses the feature information of the channel and performs dimensionality reduction processing to restore the dimension of the feature to After that, it is added to the feature map input by the window attention module to obtain the intermediate feature; the intermediate feature is processed by layer normalization and multi-scale fusion feedforward network in turn and then added to the intermediate feature to obtain the output feature of the fusion module. In the process, the flow process of features can be defined as:

[0043] Among them, LN represents layer normalization; Represents a splicing operation; Representation 1 1 convolution; Indicates size 3 Depth-wise convolution of 3.

[0044] like Figure 5 As shown in the figure, the multi-scale fusion feedforward network includes three parallel branches; each branch includes a convolutional layer, a first deep convolutional layer, a first GELU activation function, a second deep convolutional layer, and a second GELU activation function connected in sequence; in each branch, the convolutional layer, the first deep convolutional layer, and the GELU activation function are used in sequence to process the features to obtain the intermediate features of each branch, and the intermediate features of the other two branches are extracted and spliced ​​in the channel dimension to achieve preliminary fusion of features. At this time, the output features of the three branches all contain local features of three scales. Each branch continues to process the intermediate features using the second deep convolutional layer and the second GELU activation function. Finally, a 1 1 The convolution layer finally fuses the output features of the three branches and restores the channel dimension to the same dimension as the input feature. Among the three parallel branches, the convolution kernel sizes of the deep convolution layer are 3 3, 5 5 and 7 7. The feature fusion process of the multi-scale fusion feedforward network can be defined as:

[0045] in, express convolution; , and Respectively indicate the size , and Depth convolution; represents the GELU activation function; Represents the concatenation operation on the channel dimension; , and It represents the intermediate features after the i-th deep convolution processing in the three branches; Represents the features of the multi-scale fusion feed-forward network output.

[0046] The features output by the multi-scale fusion feedforward network are used as the features output by the window attention module.

[0047] Step 3: Model loss function selection and model training The image denoising model is trained using the training set. Each CT image used for training is randomly cropped into four 64×64 patches before being input into the image denoising model. The Charbonnier loss, also known as the Smooth L1 Loss, is used as the loss function of the image denoising model. The formula definition of Charbonnier is as follows: in, Indicates the number of pixels; Indicates the first pixel values, and denote the NDCT image and the generated image, respectively. pixel value; is a constant, which is set to 1e in this example. -3 .

[0048] The advantage of Charbonnier loss over standard L1 loss is that its curve is smoother and can converge faster. The advantage over MSE loss is that it is less sensitive to outliers and can provide a smoother optimization surface. Secondly, L1 loss is not differentiable at the lowest point, so there is a problem of difficulty in converging to the optimal solution, while Charbonnier loss is differentiable, which is important for model training using optimization algorithms such as gradient descent. These characteristics make it more robust in image denoising tasks.

[0049] The Adam optimizer is used to optimize all parameters of the model; the weight parameters are updated through continuous iterations of back-propagation and gradient descent algorithms.

[0050] Step 4: Use the LDCT images in the test set for testing; input any LDCT image into the verified model and output the denoised LDCT image: The label results of the NDCT image are as follows: Figure 6 As shown in the figure, the model denoises the LDCT image. Figure 7 As shown, it can be clearly seen that the output result of the present invention can well perform image denoising and retain a large amount of detail information, and the difference in details between the output result of the present invention and the label result of NDCT is very small, so the present invention is of great significance to image denoising.

[0051] Step 5: Model Evaluation The method of the present invention and the existing denoising methods are compared through three indicators: peak signal-to-noise ratio (PSNR), root mean square error (RMSE), and structural similarity (SSIM) to evaluate the denoising effects of different denoising methods, and the CT images under different noise levels are compared to verify the denoising level of the method of the present invention.

[0052] The peak signal-to-noise ratio (PSNR) definition formula is as follows: Where MSE is the mean square error; H and W are the height and width of the image respectively; and They are the denoised image and the original image with normal dose respectively; n is the number of bits.

[0053] The root mean square error (RMSE) definition formula is as follows: Where X represents the generated restored image; Y represents the corresponding NDCT image.

[0054] The structural similarity (SSIM) definition formula is as follows: in, and are the means of image X and image Y respectively; and is a constant; and are the standard deviations, respectively; is the covariance.

[0055] The quantitative comparison results of the denoising performance of different denoising methods on the AAPM dataset are shown in Table 1.

[0056] Table 1 Quantitative comparison of denoising performance of different denoising methods The larger the PSNR and SSIM values ​​are, the better, and the smaller the RMSE value is, the better. It can be seen from Table 1 that the comprehensive performance of the three evaluation indicators of the present invention is better than that of other denoising methods. A large number of comparative experiments with classic denoising methods and the latest methods have proved the effectiveness of the denoising method proposed in the present invention.

[0057] The drawings of the embodiments disclosed in the present invention only involve structures related to the embodiments disclosed in the present invention, but the above is only a preferred embodiment of the present invention, which can be fully applied to various fields suitable for the present invention. Therefore, without departing from the general concept defined by the claims and the equivalent scope, the present invention is not limited to the specific details and the illustrations shown and described here.

Claims

1. A low-dose CT image denoising method based on window mixed attention, characterized in that: The following steps are involved: Step 1: construct a model sample set containing low-dose computed tomography images, and label the images in the model sample set; Step 2: construct an image denoising model; the image denoising model adopts a U-shaped architecture, which includes an input projection layer, an encoder, a bottleneck layer, a decoder and an output layer; the input projection layer processes the image input to the image denoising model and then inputs it into the encoder; The encoder includes multiple layers of encoding units, and the decoder includes multiple layers of decoding units; the encoding units and decoding units in the same layer are connected by skipping; the bottleneck layer is connected in series between the last layer of encoding units and the first layer of decoding units; the output layer is used to process the feature map output by the decoder to obtain the output image of the image denoising model; the encoding unit, the decoding unit and the bottleneck layer each include two serially connected window attention modules; Step 3: Use the model sample set to train the image denoising model; Step 4: Use the trained image denoising model to denoise the low-dose computed tomography images.

2. The low-dose CT image denoising method based on window mixed attention according to claim 1, characterized in that: The window attention module includes a parallel window attention path, a channel attention path and a fusion module; the window attention path splits the feature map of the input window attention module into multiple windows, extracts the features of each window, and obtains local features; the channel attention path extracts the global features of the feature map of the input window attention module, and splices them with the local features and inputs them into the fusion module.

3. The low-dose CT image denoising method based on window mixed attention according to claim 2, characterized in that: The fusion module specifically performs ordinary convolution and deep convolution on the input features in sequence, and then adds them to the feature map input by the window attention module to obtain intermediate features; the intermediate features are processed by layer normalization and multi-scale fusion feedforward network in sequence, and the processed features are added to the intermediate features before processing to obtain the output features of the fusion module.

4. The low-dose CT image denoising method based on window mixed attention according to claim 3 is characterized in that: The window attention path includes a windowing module, a channel self-attention module, a multi-head self-attention module and a de-windowing module connected in sequence.

5. The low-dose CT image denoising method based on window mixed attention according to claim 4, characterized in that: The channel self-attention module includes three branches with the same structure; the image input to the channel self-attention module is subjected to layer normalization processing and then input into the three branches respectively, and ordinary convolution and deep convolution processing are performed in sequence to obtain a query vector, a key vector and a value vector; the query vector and the key vector are multiplied and the result after being processed by a softmax activation function is multiplied with the value vector to obtain a preliminary fusion result; the preliminary fusion result is convolved and added to the input image of the channel self-attention module to obtain the output features of the channel self-attention module.

6. The low-dose CT image denoising method based on window mixed attention according to claim 4, characterized in that: The channel attention path includes a downsampling layer, a first convolutional attention module, a second convolutional attention module and an upsampling layer connected in sequence; the output features of the channel self-attention module are sequentially de-windowed and down-sampled, and the addition result of the output features of the first convolutional attention module is used as the input of the second convolutional attention module.

7. The low-dose CT image denoising method based on window mixed attention according to claim 6, characterized in that: The first convolutional attention module and the second convolutional attention module have the same structure, both of which include layer normalization, a first convolutional layer, a depth convolutional layer, an activation function, a channel attention module and a second convolutional layer connected in sequence; the output of the convolutional attention module is the result of adding the output features of the second convolutional layer to the input features of the convolutional attention module.

8. The low-dose CT image denoising method based on window mixed attention according to claim 6, characterized in that: The multi-scale fusion feedforward network includes three branches connected in parallel; the three branches have the same structure; Each branch uses the convolution layer, the first deep convolution layer and the GELU activation function to process the features in turn and fuses the processed features with the features processed by the other two branches to obtain the intermediate features of each branch; each branch uses the second deep convolution layer and the second GELU activation function to continue processing the intermediate features to obtain output features; The output features of the three branches are fused and input into the convolutional layer to obtain the output features of the multi-scale fusion feedforward network.

9. The low-dose CT image denoising method based on window mixed attention according to claim 1, characterized in that: All the encoding units except the last layer in the encoder input their output feature maps to the next layer of encoding units after downsampling; the last layer of encoding units input their output feature maps to the bottleneck layer after downsampling; The input of the decoding units in the decoder except the first-layer decoding units is the concatenation of the output feature map of the previous layer decoding unit after upsampling and the processing result of the corresponding encoding unit; the input of the first-layer decoding unit is the concatenation of the output feature map of the bottleneck layer after upsampling and the processing result of the corresponding encoding unit.

10. The low-dose CT image denoising method based on window mixed attention according to claim 1, characterized in that: The output layer includes a serially connected window attention module and an output projection layer; the output of the output layer is connected to the input of the image denoising model through a residual connection.

Citation Information

Cited By

  • Low-dose CT image reconstruction method and system, electronic equipment and storage medium

    CN121708157A