LDCT image super-resolution enhancement method and device based on residual convolutional neural network
Through the LDCT image super-resolution enhancement method based on residual convolutional neural network, the problem of low LDCT image resolution is solved using the improved U-Net network, and image detail enhancement and efficient detection of multi-disease screening is achieved.
Patent Information
- Application Number
- CN202111510075.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-10
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-12-10
AI Technical Summary
The prior art has problems in low-dose CT (LDCT) image processing with low imaging resolution, high cost, high radiation, and the inability of existing artificial intelligence methods to adapt to complex situations.
Using the LDCT image super-resolution enhancement method based on residual convolutional neural network, a new method of super-resolution low-resolution LDCT image is constructed by improving the hybrid cascade task U-Net deep neural network to realize multi-disease screening, detection and analysis of chest LDCT scans.
It improves the resolution of LDCT images, increases image details, and makes the display effect clearer, improves the fineness of the image, and realizes high-resolution detection without disturbance and lossless image.
Smart Images

Figure CN114255168B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical LDCT image processing, and in particular to a method and device for super-resolution enhancement of LDCT images based on a residual convolutional neural network. Background Art
[0002] In 1970-1980, CRX (chest X-ray) and sputum cytology were used as tools for lung cancer screening. Four randomized controlled trials failed to reduce lung cancer mortality. In the 1990s, CT was introduced, and LDCT (low-dose CT) was able to detect early lung cancer, but the mortality rate did not decrease. In the 2000s, low-dose CT screening could reduce the lung cancer mortality rate of the subjects by 20% compared with chest X-ray. After 2011, LDCT became popular worldwide. Artificial intelligence can be applied to all stages of accurate diagnosis and treatment of lung nodules, including early screening of the population, intelligent early diagnosis, accurate early treatment, full follow-up management, and translational research.
[0003] The main method currently used in the medical community is to use high-dose CT to screen, detect and analyze diseases such as lung nodules, chronic obstructive pulmonary disease, and coronary heart disease. The main disadvantages of these methods are slow imaging speed, high cost, and high radiation. LDCT uses low-dose CT and has low imaging resolution. In particular, existing artificial intelligence methods only perform analysis on low-resolution LDCT and are no longer able to meet the technical requirements of complex situations. Summary of the invention
[0004] In order to overcome the shortcomings of the prior art, the present invention proposes a method and device for super-resolution enhancement of LDCT images based on residual convolutional neural networks. Aiming at the low-resolution characteristics of LDCT images, the present invention uses an improved hybrid cascade task U-Net deep neural network to construct a new method for super-resolution of low-resolution LDCT images. The present invention can be used to screen, detect and analyze the three major diseases (pulmonary nodules, chronic obstructive pulmonary disease, coronary heart disease) and the spine in a one-time chest LDCT scan, without the need for additional high-precision CT scans of a certain part of the chest with higher spatial resolution requirements.
[0005] A method for super-resolution enhancement of LDCT images based on residual convolutional neural network comprises the following steps:
[0006] Step 1) Create training sets and test sets;
[0007] Step 2) LDCT initial image preprocessing;
[0008] Step 3) Determine whether training is done, if yes, go to step 4), if not, go to step 8);
[0009] Step 4) Improve the hybrid cascade task U-Net for feature extraction;
[0010] Step 5) Error calculation;
[0011] Step 6) Error back propagation;
[0012] Step 7) Determine whether the error meets the requirements, if yes, go to step 8), if not, return to step 4);
[0013] Step 8) Output the image super-resolution model;
[0014] Step 9) generating a super-resolution CT image;
[0015] Step 10) End.
[0016] The steps of step 1) to prepare the training set and the test set are as follows:
[0017] Step 1-1) Find a large number of low-resolution LDCT images and their corresponding true high-resolution CT images, convert the CT images in DICOM format into grayscale images in PNG format, randomly intercept 128×128 images of the low-resolution LDCT images and 256×256 images of the corresponding positions of the true high-resolution CT images, rotate them 90°, 180°, and 270° respectively, and flip them accordingly to obtain variants of each image, a total of 8 images; randomly intercept 10 different areas of each LDCT image, collect 50,000 different cropped low-resolution LDCT images and true high-resolution CT images as training sets, and 5,000 different cropped low-resolution LDCT images and true high-resolution CT images as test sets.
[0018] The steps of step 2) LDCT initial image preprocessing are as follows:
[0019] Step 2-1) Before image super-resolution enhancement, the low-resolution LDCT image is subjected to standardization preprocessing to obtain a standardized image;
[0020] Step 2-2) interpolating the standardized image using a bicubic upsampling method to obtain a high-resolution image so that it has the same resolution size as the true high-resolution CT image;
[0021] Step 2-3) Input the high-resolution image into a convolutional layer with a convolution kernel size of 3×3, an input channel number of 1, and an output channel number of 64 to transform the image into an original feature map with 64 channels.
[0022] The steps of improving the hybrid cascade task U-Net for feature extraction and constructing the spatial context branch network are as follows:
[0023] Step 4-1-1) Input the original feature map into a convolutional layer with a kernel size of 3×3, 64 input and output channels, a stride of 2, and a padding of 1. After three such convolutional layers, each layer outputs a feature map with a resolution reduced by half relative to the input feature map. After the four feature maps pass through a convolutional layer with a kernel size of 1×1 and 64 input and output channels, the low-resolution feature maps are upsampled by deconvolutional layers with kernel sizes of 2×2, 4×4, and 8×8, 64 input and output channels, and strides of 2, 4, and 8, respectively, to output a feature map with the same resolution as the original feature map.
[0024] Step 4-1-2) Add the four feature maps of the same resolution to output a feature map, and then extract the features through four depthwise separable convolutions to obtain a feature map; each depthwise separable convolution consists of a convolution layer with a size of 3×3, 64 input and output channels, 64 groups, and a padding of 1, and a point convolution layer with a size of 1×1 and 64 input and output channels;
[0025] Step 4-1-3) After the obtained feature map is input into the convolution layer with a convolution kernel size of 1×1 and 64 input and output channels, the spatial context branch feature map is obtained.
[0026] The steps of improving the hybrid cascade task U-Net for feature extraction and constructing the Fourier channel attention residual module are as follows:
[0027] Step 4-2-1) The original feature map is passed through two depth-wise separable convolutional layers and a Swish activation layer to obtain a new feature map; the depth-wise separable convolutional layer consists of a convolutional layer with a size of 3×3, 64 input and output channels, 64 groups, and a padding of 1, and a point convolutional layer with a size of 1×1 and 64 input and output channels;
[0028] Step 4-2-2) The new feature map is input into the Fourier attention module to obtain the attention feature. The Fourier attention module is composed of fast Fourier transform, downsampling layer, global average pooling layer, upsampling layer, depth-separable convolution layer, and Sigmoid layer. The downsampling layer includes a convolution layer with a convolution kernel size of 1×1, 64 input channels, and 16 output channels, and a Swish activation function layer. The upsampling layer includes a convolution layer with a convolution kernel size of 1×1, 16 input channels, and 64 output channels, and a Swish activation function layer. The 64 feature maps obtained are passed through the Sigmoid layer to calculate the attention value of each channel, and the attention value is multiplied with each channel of the input feature map to obtain the attention feature map.
[0029] Step 4-2-3) Take the obtained attention feature map as the residual and add it to the original feature map at the element level to obtain the attention residual feature map.
[0030] The steps of improving the hybrid cascade task U-Net for feature extraction and constructing the U-Net network structure based on the Fourier channel attention residual module are as follows:
[0031] Step 4-3-1) The U-Net network structure based on the Fourier channel attention residual module includes 2 downsampling layers and 2 upsampling layers; the original feature map is extracted by the Fourier channel attention residual module in the downsampling layer, and then passes through a convolution layer with a convolution kernel size of 3×3, 64 input channels, 128 output channels, a step size of 2, and a padding of 1, and then passes through a Swish activation function layer and a depthwise separable convolution layer to obtain a downsampled feature map, where the depthwise separable convolution layer is composed of a convolution layer with a convolution layer size of 3×3, 128 input and output channels, 128 groups, and a padding of 1, and a point convolution layer with a convolution layer size of 1×1 and 128 input and output channels;
[0032] Step 4-3-2) The downsampled feature map is downsampled again, and the features are extracted by the Fourier channel attention residual module. Then, after passing through a convolution layer with a convolution kernel size of 3×3, an input channel number of 128, an output channel number of 256, a step size of 2, and a padding of 1, it passes through a Swish activation function layer and a depthwise separable convolution layer to obtain a new downsampled feature map; the depthwise separable convolution layer is composed of a convolution layer with a convolution layer size of 3×3, an input and output channel number of 256, a group number of 256, a padding of 1, and a point convolution layer with a convolution layer size of 1×1 and an input and output channel number of 256;
[0033] Step 4-3-3) The new downsampled feature map passes through two upsampling layers in sequence to improve the resolution of the feature map; after upsampling, the Fourier channel attention residual module extracts features, passes through a convolution layer with a convolution kernel size of 3×3, an input channel number of 256, and an output channel number of 1024, and then passes through a PixelShuffle layer to convert the number of channels to 256, and at the same time, the length and width of the feature map are doubled; then it passes through a Swish activation function layer, and then passes through a convolution layer with a convolution kernel size of 3×3, an input channel number of 256, and an output channel number of 128, and then it is spliced with the 128 feature maps obtained after the first downsampling layer of the original feature map to obtain a feature map with 256 channels;
[0034] Step 4-3-4) After the feature map of 128 channels is extracted by the Fourier channel attention residual module, it passes through a convolution layer with a convolution kernel size of 3×3, an input channel number of 256, and an output channel number of 1024, and then passes through a PixelShuffle layer to convert the number of channels to 256, and at the same time, the length and width of the feature map are doubled; then it passes through a Swish activation function layer, and then a convolution layer with a convolution kernel size of 3×3, an input channel number of 256, and an output channel number of 64, and then it is spliced with the original feature map of 64 channels to obtain a feature map of 128 channels; then it passes through a convolution layer with a convolution kernel size of 3×3, an input channel number of 128, and an output channel number of 64, and outputs a U-Net feature map of 64 channels.
[0035] The steps of improving the hybrid cascade task U-Net for feature extraction and constructing a feature refinement module based on U-Net are as follows:
[0036] Step 4-4) The U-Net-based feature refinement module consists of channel-level splicing, convolutional layer and U-Net in sequence; the channel-level splicing splices the U-Net feature map of 64 channels with the original feature map of 64 channels to obtain a feature map of 128 channels; then it passes through a convolutional layer with a convolution kernel size of 3×3, 128 input channels and 64 output channels, and then passes through a U-Net network based on the Fourier channel attention residual module to obtain a refined feature map.
[0037] The above step 4) improves the hybrid cascade task U-Net for feature extraction, and the steps for constructing a network structure based on the improved hybrid cascade task are as follows:
[0038] Step 4-5) The U-Net feature map is sequentially input into three feature refinement modules to obtain three refined feature maps of different levels, from coarse to fine; at the same time, the original feature map is input into three feature refinement modules for channel-level splicing; at the same time, the spatial context feature map and the three refined feature maps of different levels are input into the pixel branch network to generate a super-resolution CT image, where the pixel branch network contains element-level addition, 4 consecutive depth-separable convolutional layers and reconstruction layers; the pixel branch network is a residual network, and the input is the spatial context feature map, the refined feature map and the residual of the previous pixel branch network Output special diagnosis map; the input is added element-wise to obtain 64 feature maps, and then passes through 4 convolutional layers with a size of 3×3, 64 input and output channels, 64 groups, and a padding of 1, and a point convolutional layer with a convolutional layer size of 1×1 and 64 input and output channels to obtain a reconstructed feature map; the reconstructed feature map is input as the residual to the next pixel branch network, and after passing through a convolutional layer with a size of 3×3, 64 input channels, and 1 output channel, a super-resolution CT image is obtained after denormalization; 3 pixel branches output 3 super-resolution CT images with different levels of fineness.
[0039] The steps of error calculation in step 5) are as follows:
[0040] Step 5) The mean square error and structural similarity error of the true high-resolution image after bicubic downsampling 4 times and bicubic upsampling 4 times are calculated with the first super-resolution CT image as the first stage error, where the structural similarity error is multiplied by a coefficient of 0.1; the mean square error and structural similarity error of the true high-resolution image after bicubic downsampling 2 times and bicubic upsampling 2 times are calculated with the second super-resolution CT image as the second stage error, where the structural similarity error is multiplied by a coefficient of 0.1; the mean square error and structural similarity error of the true high-resolution image and the third super-resolution CT image are calculated as the third stage error, where the structural similarity error is multiplied by a coefficient of 0.1; the errors of the first two stages are multiplied by 0.3 and added to the error of the last stage to obtain the final error value;
[0041] The steps of step 6) error back propagation are as follows:
[0042] Step 6) The error value is back-propagated, the learning rate is specified to be 0.0001, the optimizer is ADAM, and the learning rate adopts a staged decrease strategy to continuously reduce the loss between the super-resolution CT image and the true high-resolution image.
[0043] The device using the method is a CT that generates super-resolution CT images based on the image super-resolution enhancement method of the residual convolutional neural network; or a microscope that generates super-resolution microscopic images based on the image super-resolution enhancement method of the residual convolutional neural network.
[0044] Beneficial effects of the present invention:
[0045] (1) The LDCT super-resolution image effect of the present invention is better, the image details are increased, the display effect is clearer, and the image fineness is improved.
[0046] (2) The LDCT image super-resolution method of the present invention can not only enlarge the image size, but also has the function of image restoration in a certain sense, and can improve the image signal-to-noise ratio and image structural similarity to a certain extent.
[0047] (3) The LDCT image super-resolution method of the present invention can achieve high-resolution detection without disturbing or damaging the image.
[0048] (4) The method of the present invention can be applied to CT devices and microscopes (with different image sources). BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 It is a flow chart of the present invention.
[0050] Figure 2 It is the structural diagram of the improved mixed cascade task U-Net.
[0051] Figure 3 It is the structural diagram of the spatial context branch network.
[0052] Figure 4 It is the structural diagram of the Fourier channel attention residual module.
[0053] Figure 5 This is the structural diagram of the U-Net network based on the Fourier channel attention residual module. DETAILED DESCRIPTION
[0054] The technical solution of the present invention will be further described below in conjunction with the accompanying drawings and implementation examples.
[0055] As attached Figure 1 As shown, it is a schematic diagram of a process of the present invention, comprising the following steps:
[0056] Step 1) Create training sets and test sets;
[0057] Step 2) LDCT initial image preprocessing;
[0058] Step 3) Determine whether training is completed, if yes, proceed to step 4); if no, proceed to step 8);
[0059] Step 4) Improve the hybrid cascade task U-Net for feature extraction;
[0060] Step 5) Error calculation;
[0061] Step 6) Error back propagation;
[0062] Step 7) Determine whether the error meets the requirements. If yes, proceed to step 8). If not, return to step 4);
[0063] Step 8) Output the image super-resolution model;
[0064] Step 9) generating a super-resolution CT image;
[0065] Step 10) End.
[0066] The steps for making training sets and test sets are as follows:
[0067] Step 1-1: Find a large number of low-resolution LDCT images and their corresponding true high-resolution CT images, convert the CT images in DICOM format into grayscale images in PNG format, randomly intercept 128×128 images of the low-resolution LDCT images and 256×256 images of the corresponding positions of the true high-resolution CT images, rotate them 90°, 180°, and 270° respectively, and flip them accordingly to obtain variants of each image, a total of 8 images; randomly intercept 10 different areas of each LDCT image, collect 50,000 different cropped low-resolution LDCT images and true high-resolution CT images as training sets, and 5,000 different cropped low-resolution LDCT images and true high-resolution CT images as test sets.
[0068] The steps of LDCT initial image preprocessing are as follows:
[0069] Step 2-1) Before image super-resolution enhancement, the low-resolution LDCT image is subjected to standardization preprocessing to obtain a standardized image;
[0070] Step 2-2) interpolating the standardized image using a bicubic upsampling method to obtain a high-resolution image so that it has the same resolution size as the true high-resolution CT image;
[0071] Step 2-3) Input the high-resolution image into a convolutional layer with a convolution kernel size of 3×3, an input channel number of 1, and an output channel number of 64 to transform the image into an original feature map with 64 channels.
[0072] As attached Figure 3 As shown in Figure 2, the steps to construct a spatial context branch network are as follows:
[0073] Step 4-1-1) Input the original feature map into a convolutional layer with a kernel size of 3×3, 64 input and output channels, a stride of 2, and a padding of 1. After three such convolutional layers, each layer outputs a feature map with a resolution reduced by half relative to the input feature map. After the four feature maps pass through a convolutional layer with a kernel size of 1×1 and 64 input and output channels, the low-resolution feature maps are upsampled by deconvolutional layers with kernel sizes of 2×2, 4×4, and 8×8, 64 input and output channels, and strides of 2, 4, and 8, respectively, to output a feature map with the same resolution as the original feature map.
[0074] Step 4-1-2) Add the four feature maps of the same resolution to output a feature map, and then extract the features through four depthwise separable convolutions to obtain a feature map; each depthwise separable convolution consists of a convolution layer with a size of 3×3, 64 input and output channels, 64 groups, and a padding of 1, and a point convolution layer with a size of 1×1 and 64 input and output channels;
[0075] Step 4-1-3) After the obtained feature map is input into the convolution layer with a convolution kernel size of 1×1 and 64 input and output channels, the spatial context branch feature map is obtained.
[0076] As attached Figure 4 As shown, the steps to construct the Fourier channel attention residual module are as follows:
[0077] Step 4-2-1) The original feature map is passed through two depth-wise separable convolutional layers and a Swish activation layer to obtain a new feature map; the depth-wise separable convolutional layer consists of a convolutional layer with a size of 3×3, 64 input and output channels, 64 groups, and a padding of 1, and a point convolutional layer with a size of 1×1 and 64 input and output channels;
[0078] Step 4-2-2) The new feature map is input into the Fourier attention module to obtain the attention feature. The Fourier attention module is composed of fast Fourier transform, downsampling layer, global average pooling layer, upsampling layer, depth-separable convolution layer, and Sigmoid layer. The downsampling layer includes a convolution layer with a convolution kernel size of 1×1, 64 input channels, and 16 output channels, and a Swish activation function layer. The upsampling layer includes a convolution layer with a convolution kernel size of 1×1, 16 input channels, and 64 output channels, and a Swish activation function layer. The 64 feature maps obtained are passed through the Sigmoid layer to calculate the attention value of each channel, and the attention value is multiplied with each channel of the input feature map to obtain the attention feature map;
[0079] Step 4-2-3) Take the obtained attention feature map as the residual and add it to the original feature map at the element level to obtain the attention residual feature map.
[0080] As attached Figure 5 As shown in the figure, the steps to construct the U-Net network structure based on the Fourier channel attention residual module are as follows:
[0081] Step 4-3-1) The U-Net network structure based on the Fourier channel attention residual module includes 2 downsampling layers and 2 upsampling layers; the original feature map is extracted by the Fourier channel attention residual module in the downsampling layer, and then passes through a convolution layer with a convolution kernel size of 3×3, 64 input channels, 128 output channels, a step size of 2, and a padding of 1, and then passes through a Swish activation function layer and a depthwise separable convolution layer to obtain a downsampled feature map, where the depthwise separable convolution layer is composed of a convolution layer with a convolution layer size of 3×3, 128 input and output channels, 128 groups, and a padding of 1, and a point convolution layer with a convolution layer size of 1×1 and 128 input and output channels;
[0082] Step 4-3-2) The downsampled feature map is downsampled again, and the features are extracted by the Fourier channel attention residual module. Then, after passing through a convolution layer with a convolution kernel size of 3×3, an input channel number of 128, an output channel number of 256, a step size of 2, and a padding of 1, it passes through a Swish activation function layer and a depthwise separable convolution layer to obtain a new downsampled feature map; the depthwise separable convolution layer is composed of a convolution layer with a convolution layer size of 3×3, an input and output channel number of 256, a group number of 256, a padding of 1, and a point convolution layer with a convolution layer size of 1×1 and an input and output channel number of 256;
[0083] Step 4-3-3) The new downsampled feature map passes through two upsampling layers in sequence to improve the resolution of the feature map; after upsampling, the Fourier channel attention residual module extracts features, passes through a convolution layer with a convolution kernel size of 3×3, an input channel number of 256, and an output channel number of 1024, and then passes through a PixelShuffle layer to convert the number of channels to 256, and at the same time, the length and width of the feature map are doubled; then it passes through a Swish activation function layer, and then passes through a convolution layer with a convolution kernel size of 3×3, an input channel number of 256, and an output channel number of 128, and then it is spliced with the 128 feature maps obtained after the first downsampling layer of the original feature map to obtain a feature map with 256 channels;
[0084] Step 4-3-4) After the feature map of 128 channels is extracted by the Fourier channel attention residual module, it passes through a convolution layer with a convolution kernel size of 3×3, an input channel number of 256, and an output channel number of 1024, and then passes through a PixelShuffle layer to convert the number of channels to 256, and at the same time, the length and width of the feature map are doubled; then it passes through a Swish activation function layer, and then a convolution layer with a convolution kernel size of 3×3, an input channel number of 256, and an output channel number of 64, and then it is spliced with the original feature map of 64 channels to obtain a feature map of 128 channels; then it passes through a convolution layer with a convolution kernel size of 3×3, an input channel number of 128, and an output channel number of 64, and outputs a U-Net feature map of 64 channels.
[0085] As attached Figure 2 As shown in the figure, the steps to build a feature refinement module based on U-Net are as follows:
[0086] Step 4-4) The U-Net-based feature refinement module consists of channel-level splicing, convolutional layer, and U-Net. The channel-level splicing concatenates the 64-channel U-Net feature map with the 64-channel original feature map to obtain a 128-channel feature map; then it passes through a convolutional layer with a convolution kernel size of 3×3, 128 input channels, and 64 output channels, and then passes through a U-Net network based on the Fourier channel attention residual module to obtain a refined feature map.
[0087] As attached Figure 2 As shown, the steps to build a network structure based on the improved hybrid cascade task are as follows:
[0088] Step 4-5) The U-Net feature map is sequentially input into three feature refinement modules to obtain three refined feature maps of different levels, from coarse to fine. At the same time, the original feature map is input into three feature refinement modules for channel-level splicing; at the same time, the spatial context feature map and the three refined feature maps of different levels are input into the pixel branch network to generate a super-resolution CT image, where the pixel branch network contains element-level addition, 4 consecutive depth-separable convolutional layers and reconstruction layers; the pixel branch network is a residual network, and the input is the spatial context feature map, the refined feature map and the residual output feature map of the previous pixel branch network. The input is added element-wise to obtain 64 feature maps, which are then passed through 4 convolutional layers with a size of 3×3, 64 input and output channels, 64 groups, and a padding of 1, and a point convolutional layer with a size of 1×1 and 64 input and output channels to obtain a reconstructed feature map; the reconstructed feature map is input as the residual to the next pixel branch network, and after passing through a convolutional layer with a size of 3×3, 64 input channels, and 1 output channel, a super-resolution CT image is obtained after denormalization; the three pixel branches output three super-resolution CT images with different levels of fineness.
[0089] The steps for error calculation are as follows:
[0090] Step 5: The mean square error and structural similarity error of the true high-resolution image after bicubic downsampling 4 times and bicubic upsampling 4 times are calculated with the first super-resolution CT image as the first stage error, where the structural similarity error is multiplied by a coefficient of 0.1; the mean square error and structural similarity error of the true high-resolution image after bicubic downsampling 2 times and bicubic upsampling 2 times are calculated with the second super-resolution CT image as the second stage error, where the structural similarity error is multiplied by a coefficient of 0.1. The mean square error and structural similarity error of the true high-resolution image and the third super-resolution CT image are calculated as the third stage error, where the structural similarity error is multiplied by a coefficient of 0.1. The errors of the first two stages are multiplied by 0.3 and added to the error of the last stage to obtain the final error value.
[0091] The steps of error back propagation are as follows:
[0092] Step 6) Back propagation of the error value. The learning rate is set to 0.0001, the optimizer is ADAM, and the learning rate adopts a staged decrease strategy to continuously reduce the loss between the super-resolution CT image and the true high-resolution image.
[0093] The trained model is used to enhance the new LDCT image. The steps to obtain the super-resolution CT image are as follows:
[0094] Step 9) The low-resolution LDCT image is input into the trained model according to training step 4 to obtain a super-resolution CT image.
[0095] The embodiments described above can be further combined or replaced, and the embodiments are only descriptions of preferred embodiments of the present invention, and do not limit the concept and scope of the present invention. Without departing from the design concept of the present invention, various changes and improvements made by ordinary technicians in this field to the technical solution of the present invention belong to the protection scope of the present invention. The protection scope of the present invention is given by the attached claims and any equivalents thereof.
Claims
1. A LDCT image super-resolution enhancement method based on residual convolutional neural network, characterized in that: The following steps are involved: Step 1) Create training sets and test sets; Step 2) LDCT initial image preprocessing; Step 3) Determine whether training is done, if yes, go to step 4), if not, go to step 8); Step 4) Improve the hybrid cascade task U-Net for feature extraction; Step 5) Error calculation; Step 6) Error back propagation; Step 7) Determine whether the error meets the requirements, if yes, go to step 8), if not, return to step 4); Step 8) Output the image super-resolution model; Step 9) generating a super-resolution CT image; Step 10) End; The steps of improving the hybrid cascade task U-Net for feature extraction and constructing the spatial context branch network are as follows: Step 4-1-1) Input the original feature map to a convolution layer with a kernel size of 3×3, 64 input and output channels, a stride of 2, and a padding of 1. After three such convolution layers, each layer outputs a feature map with a resolution reduced by half relative to the input feature map. After the four feature maps pass through a convolution layer with a kernel size of 1×1 and 64 input and output channels, the low-resolution feature maps are upsampled by deconvolution layers with kernel sizes of 2×2, 4×4, and 8×8, 64 input and output channels, and strides of 2, 4, and 8, respectively, to output a feature map with the same resolution as the original feature map. Step 4-1-2) Add the four feature maps of the same resolution to output a feature map, and then extract the features through four depthwise separable convolutions to obtain a feature map; each depthwise separable convolution consists of a convolution layer with a size of 3×3, 64 input and output channels, 64 groups, and a padding of 1, and a point convolution layer with a size of 1×1 and 64 input and output channels; Step 4-1-3) After inputting the obtained feature map into the convolution layer with a convolution kernel size of 1×1 and 64 input and output channels, the spatial context branch feature map is obtained; The steps of improving the hybrid cascade task U-Net for feature extraction and constructing the Fourier channel attention residual module are as follows: Step 4-2-1) The original feature map is passed through two depth-wise separable convolutional layers and a Swish activation layer to obtain a new feature map; the depth-wise separable convolutional layer consists of a convolutional layer with a size of 3×3, 64 input and output channels, 64 groups, and a padding of 1, and a point convolutional layer with a size of 1×1 and 64 input and output channels; Step 4-2-2) The new feature map is input into the Fourier attention module to obtain the attention feature. The Fourier attention module is composed of fast Fourier transform, downsampling layer, global average pooling layer, upsampling layer, depth-separable convolution layer, and Sigmoid layer. The downsampling layer includes a convolution layer with a convolution kernel size of 1×1, 64 input channels, and 16 output channels, and a Swish activation function layer. The upsampling layer includes a convolution layer with a convolution kernel size of 1×1, 16 input channels, and 64 output channels, and a Swish activation function layer. The 64 feature maps obtained are passed through the Sigmoid layer to calculate the attention value of each channel, and the attention value is multiplied with each channel of the input feature map to obtain the attention feature map. Step 4-2-3) Take the obtained attention feature map as the residual and add it to the original feature map element-wise to obtain the attention residual feature map; The steps of improving the hybrid cascade task U-Net for feature extraction and constructing the U-Net network structure based on the Fourier channel attention residual module are as follows: Step 4-3-1) The U-Net network structure based on the Fourier channel attention residual module includes 2 downsampling layers and 2 upsampling layers; the original feature map is extracted by the Fourier channel attention residual module in the downsampling layer, and then passes through a convolution layer with a convolution kernel size of 3×3, 64 input channels, 128 output channels, a step size of 2, and a padding of 1, and then passes through a Swish activation function layer and a depthwise separable convolution layer to obtain a downsampled feature map, where the depthwise separable convolution layer is composed of a convolution layer with a convolution layer size of 3×3, 128 input and output channels, 128 groups, and a padding of 1, and a point convolution layer with a convolution layer size of 1×1 and 128 input and output channels; Step 4-3-2) The downsampled feature map is downsampled again, and the features are extracted by the Fourier channel attention residual module. Then, after passing through a convolution layer with a convolution kernel size of 3×3, an input channel number of 128, an output channel number of 256, a step size of 2, and a padding of 1, it passes through a Swish activation function layer and a depthwise separable convolution layer to obtain a new downsampled feature map; the depthwise separable convolution layer is composed of a convolution layer with a convolution layer size of 3×3, an input and output channel number of 256, a group number of 256, a padding of 1, and a point convolution layer with a convolution layer size of 1×1 and an input and output channel number of 256; Step 4-3-3) The new downsampled feature map passes through two upsampling layers in sequence to improve the resolution of the feature map; after upsampling, the Fourier channel attention residual module extracts features, passes through a convolution layer with a convolution kernel size of 3×3, an input channel number of 256, and an output channel number of 1024, and then passes through a PixelShuffle layer to convert the number of channels to 256, and at the same time, the length and width of the feature map are doubled; then it passes through a Swish activation function layer, and then passes through a convolution layer with a convolution kernel size of 3×3, an input channel number of 256, and an output channel number of 128, and then it is spliced with the 128 feature maps obtained after the first downsampling layer of the original feature map to obtain a feature map with 256 channels; Step 4-3-4) After the 128-channel feature map is extracted by the Fourier channel attention residual module, it passes through a convolution layer with a convolution kernel size of 3×3, an input channel number of 256, and an output channel number of 1024, and then passes through a PixelShuffle layer to convert the number of channels to 256, and at the same time, the length and width of the feature map are doubled; then it passes through a Swish activation function layer, and then through a convolution layer with a convolution kernel size of 3×3, an input channel number of 256, and an output channel number of 64, and then it is spliced with the original feature map of 64 channels to obtain a feature map of 128 channels; then it passes through a convolution layer with a convolution kernel size of 3×3, an input channel number of 128, and an output channel number of 64, and outputs a U-Net feature map of 64 channels; The steps of improving the hybrid cascade task U-Net for feature extraction and constructing a feature refinement module based on U-Net are as follows: Step 4-4) The U-Net-based feature refinement module consists of channel-level splicing, convolutional layer, and U-Net in sequence; the channel-level splicing splices the 64-channel U-Net feature map with the 64-channel original feature map to obtain a 128-channel feature map; then it passes through a convolutional layer with a convolution kernel size of 3×3, 128 input channels, and 64 output channels, and then passes through a U-Net network based on a Fourier channel attention residual module to obtain a refined feature map; The above step 4) improves the hybrid cascade task U-Net for feature extraction, and the steps for constructing a network structure based on the improved hybrid cascade task are as follows: Step 4-5) The U-Net feature map is sequentially input into three feature refinement modules to obtain three refined feature maps of different levels, from coarse to fine; at the same time, the original feature map is input into three feature refinement modules for channel-level splicing; at the same time, the spatial context feature map and the three refined feature maps of different levels are input into the pixel branch network to generate a super-resolution CT image, where the pixel branch network contains element-level addition, 4 consecutive depth-separable convolutional layers and reconstruction layers; the pixel branch network is a residual network, and the input is the spatial context feature map, the refined feature map and the residual of the previous pixel branch network Output special diagnosis map; input is added element-wise to obtain 64 feature maps, and then passed through 4 convolutional layers with a size of 3×3, 64 input and output channels, 64 groups, and a padding of 1, and a point convolutional layer with a size of 1×1 and 64 input and output channels to obtain a reconstructed feature map; the reconstructed feature map is input as a residual to the next pixel branch network, and after passing through a convolutional layer with a size of 3×3, 64 input channels, and 1 output channel, a super-resolution CT image is obtained after denormalization; 3 pixel branches output 3 super-resolution CT images with different degrees of refinement; The steps of error calculation in step 5) are as follows: Step 5) The true high-resolution image is downsampled 4 times by bicubic and upsampled 4 times by bicubic, and the mean square error and structural similarity error are calculated with the first super-resolution CT image as the first stage error, where the structural similarity error is multiplied by a coefficient of 0.1; The true high-resolution image is downsampled by 2 times and upsampled by 2 times by bicubic, and the mean square error and structural similarity error are calculated with the second super-resolution CT image as the second stage error, where the structural similarity error is multiplied by a coefficient of 0.1; The mean square error and structural similarity error of the true high-resolution image and the third super-resolution CT image are calculated as the third stage error, where the structural similarity error is multiplied by a coefficient of 0.1; the errors of the first two stages are multiplied by 0.3 and added to the error of the last stage to obtain the final error value; The steps of step 6) error back propagation are as follows: Step 6) The error value is back-propagated, the learning rate is specified to be 0.0001, the optimizer is ADAM, and the learning rate adopts a staged decrease strategy to continuously reduce the loss between the super-resolution CT image and the true high-resolution image.
2. The method according to claim 1, characterized in that The steps of step 1) to prepare the training set and the test set are as follows: Step 1-1) Find a large number of low-resolution LDCT images and their corresponding true high-resolution CT images, convert the CT images in DICOM format into grayscale images in PNG format, randomly intercept 128×128 images of the low-resolution LDCT images and 256×256 images of the corresponding positions of the true high-resolution CT images, rotate them 90°, 180°, and 270° respectively, and flip them accordingly to obtain variants of each image, a total of 8 images; randomly intercept 10 different areas of each LDCT image, collect 50,000 different cropped low-resolution LDCT images and true high-resolution CT images as training sets, and 5,000 different cropped low-resolution LDCT images and true high-resolution CT images as test sets.
3. The method according to claim 1, characterized in that The steps of step 2) LDCT initial image preprocessing are as follows: Step 2-1) Before image super-resolution enhancement, the low-resolution LDCT image is subjected to standardization preprocessing to obtain a standardized image; Step 2-2) interpolating the standardized image using a bicubic upsampling method to obtain a high-resolution image so that it has the same resolution size as the true high-resolution CT image; Step 2-3) Input the high-resolution image into a convolutional layer with a convolution kernel size of 3×3, an input channel number of 1, and an output channel number of 64 to transform the image into an original feature map with 64 channels.
4. A device for applying the method according to claim 1, characterized in that: A CT that generates a super-resolution CT image using an image super-resolution enhancement method based on a residual convolutional neural network; or a microscope that generates a super-resolution microscopic image using an image super-resolution enhancement method based on a residual convolutional neural network.
Citation Information
Patent Citations
Full-network low-dose CT imaging method and apparatus based on convolution residual network
CN109102550A
Low-dose CT reconstruction method based on interpolation convolutional neural network
CN112489156A