CT image processing method based on attention-enhanced residual UNet network

By adopting an attention-enhanced residual UNet network method in CT image processing, the dynamic residual block and channel attention mechanism module are used to solve the problem of noise and artifact interference in CT image processing, and the image segmentation accuracy and scanning quality are improved.

CN120107290AInactive Publication Date: 2025-06-06ANHUI UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510269208.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, CT image processing is disturbed by noise and artifacts, resulting in the problem of low image segmentation accuracy.

Method used

Using the CT image processing method based on attention-enhanced residual UNet network, the dynamic residual block and the introduction of channel attention mechanism module are used to adapt to feature maps of different intensities and forms to reduce noise and artifact interference.

Benefits of technology

It effectively improves the segmentation accuracy and scanning quality of CT images, reduces the influence of noise and artifacts, and makes image segmentation more accurate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107290A_ABST
    Figure CN120107290A_ABST
Patent Text Reader

Abstract

The invention discloses a CT image processing method based on an attention-enhanced residual UNet network, and belongs to the technical field of image processing. The method comprises the following steps: acquiring a CT image data set, and dividing the data set into a training set, a verification set and a test set; constructing a CT image processing model based on the improved UNet network, namely replacing convolutional layers in an encoder and a decoder of the original UNet network with dynamic residual blocks, and fusing an up-sampling feature map of the decoder with a feature map of the encoder in the original UNet network; training the CT image processing model based on the training set, and evaluating the CT image processing model in the training process based on the verification set; and inputting the test set into the trained CT image processing model for CT image segmentation to obtain a segmentation result. Through the dynamic residual module, the UNet network can adapt to different features, important features are better focused, and the image segmentation precision is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and more specifically, to a CT image processing method based on an attention enhanced residual UNet network. Background Art

[0002] Computed tomography (CT) uses highly penetrating X-rays to perform cross-sectional scans of different human body tissues, and uses detectors to receive the X-ray electrical signals that pass through the layers, projecting the electrical signals into a cross-sectional image of the human body. Since the invention of the first tomography machine in 1972, doctors can use the CT scans to assist in clinical diagnosis, examination, radiotherapy, and other tasks. CT is used in a variety of scenarios, such as tumor diagnosis and nodule detection, and has become an indispensable part of medical diagnosis and treatment. However, over time, CT's use of X-rays to perform cross-sectional scans of the human body may increase the patient's potential health risks related to the probability of cancer, which has attracted widespread attention.

[0003] Considering the risks of radiation exposure in conventional CT images, reducing radiation dose as much as possible has become a trend in CT-related research in the past few decades. There are methods to reduce radiation dose in specific CT studies. Generally speaking, lower radiation dose can be obtained by reducing the tube current or shortening the exposure time of the X-ray tube.

[0004] In the existing technology, the advantages of using LDCT scanning (Low Dose CT, LDCT) are: 1) It can effectively reduce various potential risks of X-rays to the human body, especially it can well protect infants, pregnant women, cancer patients and other people who are not suitable for conventional dose CT scanning and people who need multiple CT scans in a short period of time; 2) Compared with conventional dose CT scanning, reduced radiation dose scanning can alleviate equipment aging, reduce equipment loss, extend the service life of CT scanners, and reduce the hospital's cost of updating CT equipment.

[0005] Although lowering the radiation dose can reduce potential risks, low-dose CT scans will add a lot of noise and artifacts to the reconstructed images and reduce the signal-to-noise ratio, which will seriously affect the diagnostic accuracy. Summary of the invention

[0006] 1. Technical issues to be solved

[0007] In view of the problem in the prior art that the image segmentation accuracy is low due to the interference of noise and artifacts in the CT image processing process, the present invention provides a CT image processing method based on the attention enhanced residual UNet network, which can effectively reduce the interference of noise and artifacts and improve the segmentation accuracy and scanning quality of CT images.

[0008] 2. Technical solution

[0009] The purpose of the present invention is achieved through the following technical solutions.

[0010] A CT image processing method based on an attention-enhanced residual UNet network comprises the following steps:

[0011] Collecting a CT image data set, and dividing the CT image data set into a training set, a validation set, and a test set;

[0012] Construct a CT image processing model based on the improved UNet network, including:

[0013] 1) Replace the convolutional layers in the original UNet network encoder and decoder with dynamic residual blocks;

[0014] 2) In the original UNet network, the upsampled feature map of the decoder is fused with the feature map of the encoder;

[0015] The CT image processing model is trained based on the training set, and the CT image processing model in the training process is evaluated based on the validation set to obtain a trained CT image processing model;

[0016] The test set is input into the trained CT image processing model for CT image segmentation to obtain the segmentation result.

[0017] As a further improvement of the present invention, the dynamic residual block includes two 3×3 convolutional layers and a channel attention mechanism module.

[0018] As a further improvement of the present invention, the collected CT image x is subjected to feature extraction through a 3×3 convolutional layer to obtain a feature map x 1 , introduce the ReLU activation function to the feature map x 1 Perform nonlinear processing to obtain the feature map F(x 1 );

[0019] The feature map F(x 1 ) After performing a 3×3 convolutional layer for feature extraction, the feature map x is obtained. 2 , and then use the ReLU activation function to activate the feature map x 2 Perform nonlinear processing to obtain the feature map F(x 2 ).

[0020] As a further improvement of the present invention, the weight of each channel is calculated by the channel attention mechanism module, and the weight is applied to the feature map F(x 2 ), we get the weighted feature map F(x 3 ).

[0021] As a further improvement of the present invention, the channel attention mechanism module calculates the weight of each channel, including a squeezing operation and an excitation operation.

[0022] As a further improvement of the present invention, the squeezing operation includes: 2 ) performs a global average pooling operation and compresses it into a global descriptor. The output of each channel is the average value of all spatial position features of the channel. The calculation formula is:

[0023]

[0024] Among them, z c represents the global descriptor of the c-th channel, H and W represent the feature map F(x 2 ), c represents the number of channels, i represents the row index, j represents the column index, and x represents the spatial dimension of ijc Represents the feature map F(x 2 ) is the value of the cth channel at position (i,j).

[0025] As a further improvement of the present invention, in the excitation operation, the weight of each channel is learned by the fully connected layer, including: converting the global descriptor z c Input into the first fully connected layer and compress its dimension. The calculation formula is:

[0026]

[0027] in, represents the compressed output, ReLU represents the ReLU activation function, W 1 represents the weight of the first fully connected layer, b 1 Represents the bias term of the first fully connected layer.

[0028] As a further improvement of the present invention, the compressed output is mapped back to the original number of channels through the second fully connected layer, and the weight of each channel is obtained through the Sigmoid activation function; the weight of each channel is expressed as:

[0029]

[0030] Among them, Sc represents the weight of each channel, Sigmoid represents the Sigmoid activation function, and W 2 represents the weight of the second fully connected layer, b 2 Represents the bias term of the second fully connected layer.

[0031] As a further improvement of the present invention, the CT image x and the weighted feature map F(x 3 ) performs residual connection to obtain the feature map x 4, through the maximum pooling operation, the feature map x 4 The spatial resolution is downsampled to obtain the downsampled feature map x pooled , in each layer of the encoder, the dynamic residual block and the maximum pooling operation are repeatedly used to output the encoder feature map x encoded .

[0032] As a further improvement of the present invention, a deconvolution operation is used at each layer of the decoder to convert the encoder feature map x encoded The spatial resolution of the upsampled feature map x is increased. upsampied , through concatenation, the upsampled feature map x upsampied With the encoder feature map x encoded Fusion is performed to obtain the fusion feature map x fused , and then fuse the feature map x fused Convolution activation via dynamic residual blocks;

[0033] After each layer of the decoder repeats the operation, the final feature map x is output final .

[0034] 3. Beneficial effects

[0035] Compared with the prior art, the advantages of the present invention are:

[0036] A CT image processing method based on an attention-enhanced residual UNet network of the present invention can better focus on important features and improve image segmentation accuracy by constructing a dynamic residual block to adapt to feature maps of different intensities and forms, especially in areas with complex details such as CT image boundaries. In addition, a channel attention mechanism module is introduced to adjust the control weights. The present invention can effectively combine the advantages of the channel attention mechanism, dynamic residual blocks and UNet networks, is suitable for image segmentation tasks in CT image processing, effectively reduces noise and artifact interference, and improves the segmentation accuracy and scanning quality of CT images. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 This is a flow chart of a CT image processing method according to an embodiment of the present invention;

[0038] Figure 2 Schematic diagram of the dynamic residual block structure according to an embodiment of the present invention;

[0039] Figure 3 This is a schematic diagram of the AE-ResUNet network structure according to an embodiment of the present invention;

[0040] Figure 4 Graphs showing processing results of different CT image processing methods according to embodiments of the present invention; Figure 4 (a) is the reference CT image Figure 4(b) is the result image after UNet network processing. Figure 4 (c) is the result image after processing by the AE-ResUNet network. DETAILED DESCRIPTION

[0041] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0042] Example

[0043] like Figure 1 As shown, a CT image processing method based on an attention-enhanced residual UNet network provided in this embodiment includes the following steps: acquiring a CT image data set, dividing the CT image data set into a training set, a validation set and a test set; constructing a CT image processing model based on an improved UNet network, including: 1) replacing the convolutional layers in the original UNet network encoder and decoder with dynamic residual blocks; 2) in the original UNet network, fusing the upsampled feature map of the decoder with the feature map of the encoder; training the CT image processing model based on the training set, and evaluating the CT image processing model during the training process based on the validation set to obtain a trained CT image processing model; inputting the test set into the trained CT image processing model to perform CT image segmentation to obtain a segmentation result.

[0044] Specifically in this embodiment, first, a CT image dataset is collected. In this embodiment, the collected CT image dataset comes from the dataset of the "2016 NIH-AAPM-Mayo LDCT Image Processing Competition" authorized by Mayo Clinic (hereinafter referred to as Mayo data). The collected CT image dataset is divided into a training set, a validation set, and a test set.

[0045] Furthermore, a CT image processing model based on the improved UNet network is constructed, including replacing the convolutional layers in the original UNet network encoder and decoder with dynamic residual blocks, and fusing the upsampled feature map of the decoder with the encoder feature map in the original UNet network.

[0046] Specifically, a dynamic residual block is constructed. On the basis of retaining the traditional convolution operation of the residual block, a ReLU activation function is added to help introduce nonlinearity. At the same time, a channel attention mechanism is introduced so that the residual block can adjust the weight of the feature according to the importance of each channel. The weight is applied to the feature map of the convolution output to achieve weighted adjustment, so that the features of important channels are enhanced.

[0047] In this embodiment, the dynamic residual block includes two 3×3 convolutional layers and a channel attention mechanism module.

[0048] Each 3×3 convolutional layer is used to extract high-level information of the input features. A ReLU activation function is added after each convolution operation to help introduce nonlinearity and improve the expressiveness of the network.

[0049] The channel attention mechanism module is used to adjust the weight of each channel. In this embodiment, the channel attention mechanism module first uses global average pooling to summarize the global information of each channel, and then calculates the weight of each channel through two fully connected layers (ReLU activation function and Sigmoid activation function). Finally, these weights are applied to the feature map of the convolution output to achieve weighted adjustment, so that the features of important channels are enhanced.

[0050] Input CT image x, CT image x is extracted through a 3×3 convolution layer to obtain feature map x 1 , introduce the ReLU activation function to the feature map x 1 Perform nonlinear processing to obtain the feature map F(x 1 ). In this embodiment, the feature map F(x 1 ) is expressed as:

[0051] F 1 (x) = ReLU(x 1 )

[0052] Among them, ReLU represents the ReLU activation function.

[0053] Feature map F(x 1 ) After performing a 3×3 convolutional layer for feature extraction, the feature map x is obtained. 2 , and then use the ReLU activation function to activate the feature map x 2 Perform nonlinear processing to obtain the feature map F(x 2 ). In this embodiment, the feature map F(x 2 ) is expressed as:

[0054] F 2 (x) = ReLU(x 2 )

[0055] Then, the weight of each channel is calculated through the channel attention mechanism module (SE module), and the weight is applied to the feature map F(x 2 ), we get the weighted feature map F(x 3 ). In this embodiment, the weighted feature map F(x 3 ) is expressed as:

[0056] F 3 (x) = Sc·F 2 (x)

[0057] Among them, Sc represents the weight of each channel.

[0058] It should be noted that, in this embodiment, Figure 2 As shown in Figure 2, the calculation process of the channel attention mechanism module includes a squeeze operation and an excitation operation.

[0059] The squeezing operation involves applying the input feature map F(x 2 ) performs a global average pooling operation. The pooling operation averages the spatial dimensions of each channel and compresses them into a global descriptor. The output of each channel is the average value of all spatial position features of the channel. In this embodiment, the calculation formula is:

[0060]

[0061] Among them, z c represents the global descriptor of the c-th channel, H and W represent the feature map F(x 2 ), c represents the number of channels, i represents the row index, j represents the column index, and x represents the spatial dimension of ijc Represents the feature map F(x 2 ) is the value of the cth channel at position (i,j).

[0062] After the squeezing operation, the global descriptor z of each channel is obtained c In the excitation operation, the weight of each channel is learned through the fully connected layer. This process is to learn a low-dimensional representation (through compression) and then increase the dimension to generate the weight of each channel. Specifically, the global descriptor z c Input into the first fully connected layer and compressed to a smaller dimension, the calculation formula is:

[0063]

[0064] in, represents the compressed output, ReLU represents the ReLU activation function, W 1 represents the weight of the first fully connected layer, b 1 Represents the bias term of the first fully connected layer.

[0065] The compressed output is mapped back to the original number of channels through the second fully connected layer, and the weight of each channel is obtained through the Sigmoid activation function, which controls the importance of each channel in subsequent calculations. In this embodiment, the weight of each channel is expressed as:

[0066]

[0067] Among them, Sc represents the weight of each channel, the range is [0,1], Sigmoid represents the Sigmoid activation function, W 2represents the weight of the second fully connected layer, b 2 Represents the bias term of the second fully connected layer.

[0068] like Figure 3 As shown in the figure, the constructed dynamic residual block is integrated into the encoder and decoder of the UNet network to obtain the AE-ResUNet network structure, thereby improving the UNet network's ability to learn and segment CT image details.

[0069] Specifically, the CT image x is input into the encoder of the AE-ResUNet network after normalization. The dynamic residual block in the encoder performs a convolution operation on the CT image x, and then passes through the channel attention mechanism module to weight the features of each channel through global average pooling to obtain the weighted feature map F(x 3 ), the CT image x and the weighted feature map F(x 3 ) performs residual connection to obtain the feature map x 4 In this embodiment, the feature map x 4 It is expressed as:

[0070] x 4 =F 3 (x)+x

[0071] Then, the feature map x is transformed into 4 The spatial resolution is downsampled to obtain the downsampled feature map x pooled In this embodiment, the downsampled feature map x pooled It is expressed as:

[0072] x pooled =MaxPool(x 4 )

[0073] Among them, MaxPool represents the maximum pooling operation.

[0074] At each layer of the encoder, the dynamic residual block and the maximum pooling operation are repeatedly used to output the encoder feature map x encoded In this embodiment, the encoder feature map x encoded It is expressed as:

[0075] x encoded =Encoded(x pooled )

[0076] Among them, Encoded represents the encoder output operation.

[0077] In each layer of the decoder, deconvolution operation is used to transform the encoder feature map x encoded The spatial resolution of the upsampled feature map x is increased. upsampied In this embodiment, the upsampled feature map xupsampied It is expressed as:

[0078] x upsampled =Upsample(x encoded )

[0079] Among them, Upsample represents an upsampling operation.

[0080] By concatenating the upsampled feature map x obtained by skip connection upsampied With the encoder feature map x encoded Fusion is performed to obtain the fusion feature map x fused In this embodiment, the fusion feature map x fused It is expressed as:

[0081] x fused =Concat(x upsampled , x encoded )

[0082] Among them, Concat represents a fusion operation.

[0083] The fused feature map x fused Through dynamic residual block convolution activation, the decoder repeats the operation at each layer and outputs the final feature map x final In this embodiment, changing the network's mapping relationship prediction residual can allow the UNet network to ignore the same subject, thereby focusing on the change in output, so that even slight differences can have a greater impact on the weights, so that the degradation problem of the UNet network can be solved.

[0084] Therefore, in the encoder and decoder of the CT image processing model based on the improved UNet network, dynamic residual blocks are used to replace ordinary convolutional layers, so that the UNet network can more accurately capture important image details in the CT image segmentation task. In the decoder, the upsampled feature map x upsampied With the encoder feature map x encoded Fusion is performed to help restore the spatial resolution of CT images.

[0085] In this embodiment, the output layer consists of a 3×3 convolutional layer, which converts the final feature map x output by the decoder into final Map to the desired number of categories.

[0086] In this embodiment, the improved UNet network is used to obtain a more comprehensive and discriminative feature representation. On the basis of retaining the traditional residual connection, a channel attention mechanism module is introduced for each dynamic residual block, so that the residual network can dynamically adjust the weight of the feature.

[0087] like Figure 4 Shown are the results of Mayo lung CT image processing. Figure 4 (a) is the reference CT image. By comparing and analyzing the lung CT image, the display window is [-150HU, 250HU]. The comparison shows that Figure 4 (b) is the result image after UNet network processing, which has obvious improvement in reducing noise and artifacts in CT images, but there are still some residual noise and artifacts. Figure 4 (c) is the result image after processing by the AE-ResUNet network. The AE-ResUNet network has a good suppression effect on strip artifacts at various locations, which can not only maintain the details of the image but also reduce noise, so that the uniformity of the CT image is better preserved.

[0088] In this embodiment, the processing results of different CT image processing methods are evaluated and analyzed from two aspects: subjective visual effect and objective quantitative index. The objective quantitative index uses the existing PSNR value (Peak Signal-to-Noise Ratio) and SSIM value (Structural Similarity Index) to evaluate the quality of the processed CT image. The larger the PSNR value and SSIM value, the higher the image quantization quality. The calculation formula of PSNR value and SSIM value is:

[0089]

[0090] Where MSE stands for mean square error, which is used to calculate the average square difference between the pixel values ​​of the original CT image and the processed CT image, i stands for row index, j stands for column index, I(i,j) stands for the pixel value of the original CT image at coordinate (i,j), K(i,j) stands for the pixel value of the processed CT image at coordinate (i,j), M stands for the number of image rows, N stands for the number of image columns, MAX stands for the maximum possible value of the pixel value, x stands for a local image block in the original CT image, y stands for the local image block corresponding to the position x in the processed CT image, and μ x and μ y Both represent the mean value of the image block (brightness comparison), σ x and σ y represents the standard deviation (contrast comparison), σ xy represents covariance (structural similarity comparison), C 1 and C 2 Represents a constant that prevents the denominator from being zero.

[0091] As shown in Table 1, the quantitative results of different CT image processing methods after using Mayo test data for processing. In order to quantify and intuitively compare the Mayo data results, this embodiment uses PSNR values ​​and SSIM values ​​to evaluate the processing results.

[0092] Table 1

[0093] Metrics CT UNET AE-ResUnet PSNR 33.753 37.843 38.412 SSIM 0.7617 0.878 0.8927

[0094] It can be seen from Table 1 that compared with the existing CT image processing method, the CT image processing method based on the attention enhanced residual UNet network provided in this embodiment has larger values ​​of the objective quantitative indicators PSNR value and SSIM value, and higher image quantization quality.

[0095] As shown in Table 2, it is a subjective quality evaluation scoring table. The project research needs to be closely integrated with clinical practice, and different CT image processing results are evaluated manually. The blind evaluation method is mainly used (erasing system information, scanning information and reconstruction information, etc.), and the scoring standard adopts a five-point system: 1 to 5 points, representing unacceptable, below average, acceptable, above average and very good results. Subjective scoring is performed according to the processing results of each type of denoising algorithm to obtain the "average + standard deviation" score. The specific scoring results are shown in Table 2. Compared with the original UNET network, the AE-ResUNet network used in this embodiment has a higher score in the noise artifact suppression and contrast preservation evaluation of the CT image processed by the AE-ResUNet network, indicating that this embodiment can suppress the noise in the CT image to a lower level, and thus can obtain image quality close to that of conventional dose CT images.

[0096] Table 2

[0097] Refer to the figure CT UNET AE-ResUnet Noise Suppression 3.23+0.360 4.521+0.245 4.531+0.249 Contrast distinction 3.13+0.637 4.458+0.256 4.462+0.246 Organizational differentiation 3.15+0.501 4.493+0.222 4.499+0.221 Overall image quality 3.17+0.623 4.481+0.241 4.487+0.241

[0098] Therefore, the present embodiment provides a CT image processing method based on the attention-enhanced residual UNet network, which adjusts the control weights by introducing a channel attention mechanism module to construct a dynamic residual block to adapt to feature maps of different intensities and forms. Based on the existing UNet network model, a customized AE-ResUNet network structure is provided. Compared with the prior art, the present embodiment provides a CT image processing method based on the attention-enhanced residual UNet network, which can suppress the noise in the CT image to a lower level and obtain image quality close to that of a conventional dose image.

[0099] The above schematically describes the invention and its implementation methods, which is not restrictive. Without departing from the spirit or basic features of the invention, the invention can be implemented in other specific forms. What is shown in the accompanying drawings is only one of the implementation methods of the invention. The actual structure is not limited thereto, and any figure mark in the claims should not limit the claims involved. Therefore, if a person of ordinary skill in the art is inspired by it, without departing from the purpose of the invention, a structural method and an embodiment similar to the technical solution are designed without creativity, which should all belong to the protection scope of the present invention. In addition, the word "including" does not exclude other elements or steps, and the word "one" before the element does not exclude the inclusion of "multiple" elements. The multiple elements stated in the product claim can also be implemented by one element through software or hardware. The words first, second, etc. are used to indicate the name, and do not indicate any specific order.

Claims

1. A CT image processing method based on an attention-enhanced residual UNet network, comprising the following steps: Collecting a CT image data set, and dividing the CT image data set into a training set, a validation set, and a test set; Construct a CT image processing model based on the improved UNet network, including: 1) Replace the convolutional layers in the original UNet network encoder and decoder with dynamic residual blocks; 2) In the original UNet network, the upsampled feature map of the decoder is fused with the feature map of the encoder; The CT image processing model is trained based on the training set, and the CT image processing model in the training process is evaluated based on the validation set to obtain a trained CT image processing model; The test set is input into the trained CT image processing model for CT image segmentation to obtain the segmentation result.

2. According to claim 1, a CT image processing method based on attention enhanced residual UNet network is characterized in that: The dynamic residual block consists of two 3×3 convolutional layers and a channel attention mechanism module.

3. According to claim 2, a CT image processing method based on attention enhanced residual UNet network is characterized in that: After the collected CT image x is passed through a 3×3 convolutional layer for feature extraction, a feature map x1 is obtained. The ReLU activation function is introduced to perform nonlinear processing on the feature map x1 to obtain the feature map F(x1); The feature map F(x1) is subjected to a 3×3 convolutional layer for feature extraction to obtain a feature map x2, and then the feature map x2 is subjected to nonlinear processing through a ReLU activation function to obtain a feature map F(x2).

4. A CT image processing method based on attention enhanced residual UNet network according to claim 3, characterized in that: The weight of each channel is calculated through the channel attention mechanism module, and the weight is applied to the feature map F(x2) to obtain the weighted feature map F(x3).

5. A CT image processing method based on attention enhanced residual UNet network according to claim 4, characterized in that: The channel attention mechanism module calculates the weight of each channel, including a squeeze operation and an excitation operation.

6. A CT image processing method based on attention enhanced residual UNet network according to claim 5, characterized in that: The squeezing operation includes: performing a global average pooling operation on the input feature map F(x2) to compress it into a global descriptor, and the output of each channel is the average value of all spatial position features of the channel, and the calculation formula is: Among them, z c represents the global descriptor of the c-th channel, H and W represent the spatial dimensions of the feature map F(x2), c represents the number of channels, i represents the row index, j represents the column index, and x ijc Represents the value of the cth channel at position (i, j) in the feature map F(x2).

7. A CT image processing method based on attention enhanced residual UNet network according to claim 6, characterized in that: In the excitation operation, the weight of each channel is learned through the fully connected layer, including: converting the global descriptor z c Input into the first fully connected layer and compress its dimension. The calculation formula is: in, ReLU represents the compressed output, W1 represents the weight of the first fully connected layer, and b1 represents the bias term of the first fully connected layer.

8. The CT image processing method based on the attention-enhanced residual UNet network according to claim 7 is characterized in that: The compressed output is mapped back to the original number of channels through the second fully connected layer, and the weight of each channel is obtained through the Sigmoid activation function; the weight of each channel is expressed as: Among them, Sc represents the weight of each channel, Sigmoid represents the Sigmoid activation function, W2 represents the weight of the second fully connected layer, and b2 represents the bias term of the second fully connected layer.

9. A CT image processing method based on attention-enhanced residual UNet network according to claims 4-8, characterized in that: Perform residual connection on the CT image x and the weighted feature map F(x3) to obtain the feature map x4. Downsample the spatial resolution of the feature map x4 through the maximum pooling operation to obtain the downsampled feature map x pooled , At each layer of the encoder, the dynamic residual block and the maximum pooling operation are repeatedly used to output the encoder feature map x encoded .

10. A CT image processing method based on attention enhanced residual UNet network according to claim 9, characterized in that: In each layer of the decoder, deconvolution operation is used to transform the encoder feature map x encoded The spatial resolution of the upsampled feature map x is increased. upsampied , through concatenation, the upsampled feature map x upsampied With the encoder feature map x encoded Fusion is performed to obtain the fusion feature map x fused , and then fuse the feature map x fused Convolution activation via dynamic residual blocks; After each layer of the decoder repeats the operation, the final feature map x is output final .