A medical image super-resolution reconstruction method based on detail enhancement
By introducing the Gabor detail feature extraction module into the generative adversarial network, the problems of low resolution and blurred details in the traditional Chinese medicine in the prior art are solved, and efficient detail enhancement and resolution improvement of lung CT images are achieved.
Patent Information
- Application Number
- CN202510199201.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-02-24
AI Technical Summary
Existing super-resolution algorithms are difficult to effectively improve the resolution of medical images, especially in lung CT images, which leads to low image resolution, blurred details, and difficult to identify micro lesions.
Using a generative adversarial network based on detail enhancement, a network structure including a generator and a discriminator is built, combined with the Gabor detail feature extraction module, is trained to improve image resolution and enhance texture details.
The detailed characteristics enhancement of low-resolution lung CT images is achieved, and the generated high-resolution image has clear texture details, which can more clearly display micro lesions, and improve the imaging quality of the image.
Smart Images

Figure CN119671856B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image super-resolution reconstruction, and in particular to a medical image super-resolution reconstruction method based on detail enhancement. Background Art
[0002] Lung CT is a commonly used lung examination method. During the acquisition process, X-rays are used to penetrate the human body. Considering factors such as cost control and reducing side effects on the human body, the X-ray dose needs to be strictly controlled during the acquisition process. The obtained medical images have problems such as low resolution, poor imaging quality, and blurred details. It is difficult to effectively identify tiny lesions such as nodules on CT images. Therefore, it is of great significance to obtain high-resolution lung CT images using a detail-enhanced super-resolution reconstruction algorithm.
[0003] Most of the current mainstream super-resolution algorithms are aimed at color natural images, and few are aimed at medical images. In addition, the current mainstream super-resolution algorithms generally do not consider texture detail information and context information, and the generated images have problems such as overall blur, loss of texture details, and image distortion. Summary of the invention
[0004] In order to overcome the shortcomings of the prior art, the present invention provides a medical image super-resolution reconstruction method based on detail enhancement, which can enhance texture details while improving the resolution of lung CT images, so that tiny lesions can be displayed more clearly.
[0005] In order to realize the above technology, the present invention provides a medical image super-resolution reconstruction method based on detail enhancement, comprising the following steps:
[0006] S1. Collect high-resolution lung CT images, perform normalization and other preprocessing on the images, and then perform downsampling operations to obtain low-resolution images. The high-resolution images before and after downsampling are matched one-to-one with the low-resolution images to form a training set.
[0007] S2. Construct a generative adversarial network structure including a generator and a discriminator, and train it based on the image data before and after downsampling; the training process is as follows:
[0008] The discriminator parameters are fixed, and the downsampled low-resolution CT images are passed into the generator. The discriminator is used to determine whether the generated high-resolution images are real and accurate to train the generator.
[0009] The generator parameters are fixed, and the discriminator is trained by inputting the high-resolution CT images before downsampling and the generator-generated images into the discriminator respectively.
[0010] Repeat the above operations to train the generator and discriminator, and update the parameters of the generator and discriminator through the loss function until the discriminator determines that the generated image is real and accurate, and the model training is terminated.
[0011] S3. The lung CT image obtained during image acquisition is passed into the trained generator for testing to obtain a lung CT reconstructed image with enhanced detail features.
[0012] Preferably, the generator described in step S2 adopts a residual neural network, including five convolutions, a Gabor detail feature extraction module, a nearest neighbor upsampling layer, an average pooling layer and a residual connection. The specific operation includes the following steps:
[0013] In step S21, the first three convolutions are connected to preliminarily extract image information, followed by a Gabor detail feature extraction module to extract texture detail information.
[0014] Step S22, clipping the output of the first convolution to the output of the Gabor detail feature extraction module through a residual connection to obtain a feature map.
[0015] In step S23, the feature map resolution is improved through a nearest neighbor upsampling layer, followed by two convolutions.
[0016] In step S24, the feature map is finally reduced in dimension through the average pooling layer to ensure that the input and output dimensions are consistent.
[0017] Preferably, the Gabor detail feature extraction module is composed of two convolutions, a Gabor convolution, a primitive extraction layer and an average pooling layer, and specifically includes the following steps:
[0018] The input feature map is recorded as , after a convolution and activation function activation, the image information is initially extracted, and then a convolution is performed to extract the overall image features and a Gabor convolution is performed to extract the image texture details features after activation function activation, and the multi-scale feature map is obtained by splicing. .
[0019] The feature map Perform primitive extraction to further enhance the extracted texture details. The central area of size n×n is denoted by ,Will and According to the synthesis formula, the primitive information is synthesized, that is, the position weight is recorded as , and then and Multiplication further enhances the texture detail features and obtains the output feature map The synthesis formula is recorded as:
[0020]
[0021] The sigmoid function plays a normalization role. The function represents and The measure of composability between them is denoted as:
[0022]
[0023] The output feature map After an average pooling, the dimension is reduced and the number of parameters is reduced.
[0024] Preferably, the discriminator in step S2 adopts a dual-path U-Net network, including a main path and a detail extraction path, and both paths are divided into an encoder and a decoder. The specific operation of the main path includes the following steps:
[0025] In step S25, the encoder part extracts image features by a convolution, which is activated by an activation function, followed by a maximum pooling layer for downsampling.
[0026] Step S26, repeat the operation in step S25 m times to reduce the image size and obtain a shallow feature map ps of the image.
[0027] In step S27, the decoder part upsamples the feature map ps in step S26 by a deconvolution and a convolution.
[0028] Step S28, repeat the operation described in step S27 m times to restore the image size and obtain the deep feature map pd of the image.
[0029] The structure of the detail extraction pathway is consistent with that of the main pathway, and the specific operation includes the following steps:
[0030] Step S29, the Gabor detail feature extraction module described in the encoder part extracts image detail features, which are activated by an activation function, followed by a maximum pooling layer for downsampling.
[0031] Step S210, repeating the operation described in step S29 m times to reduce the image size and obtain a shallow detail feature map qs of the image.
[0032] Step S211, same as the main channel decoder, repeats the operation described in step S27 m times to restore the image size and obtain the deep detail feature map qd of the image.
[0033] Step S212, in order to prevent the loss of original feature information, the shallow feature map ps described in step S26 and the deep feature map pd described in step S28 are spliced with corresponding sizes through jump connections to obtain the overall feature map p, and the shallow feature map qs described in step S210 and the deep feature map qd described in step S211 are spliced with corresponding sizes through jump connections to obtain the detail feature map q, and then the detail feature map q is spliced to the overall feature map p through a jump connection and merged and reduced in dimension through a 1×1 convolution. The obtained discriminator output is the authenticity value corresponding to each pixel value and whether the input image is true and accurate.
[0034] Preferably, the loss function of the discriminator is Including encoder perceptual loss Loss Function and Decoder Perceptual Loss The loss function is recorded as:
[0035]
[0036] represents the input of the discriminator, represents the random noise input to the generator, represents the expectation on all samples, represents the encoder of the discriminator, represents the decoder of the discriminator, refers to the discriminator's judgment result at this pixel, Represents a generator.
[0037] Preferably, the loss function of the generator in step S6 is recorded as:
[0038]
[0039] in is the context loss function, is the weight, is the random noise input to the generator. The calculation formula is as follows:
[0040]
[0041] Beneficial effects of the invention: In view of the low resolution of lung CT images and the loss of texture detail information during the reconstruction process of the common image super-resolution reconstruction methods on the market, the invention designs a Gabor detail feature extraction module and uses the module to construct a detail enhancement generative adversarial network. Low-resolution lung CT images with unclear details can be converted into high-resolution lung CT images with clear texture details through the network, which is conducive to presenting the lesions more clearly. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 The overall flow chart of the lung CT image super-resolution reconstruction method based on detail enhancement generative adversarial network designed for the present invention;
[0043] Figure 2 The Gabor detail feature extraction module structure diagram designed for the present invention;
[0044] Figure 3 A structural diagram of the detail enhancement generative adversarial network generator designed for the present invention;
[0045] Figure 4 A diagram of the structure of the detail-enhanced generative adversarial network discriminator designed for the present invention;
[0046] Figure 5 It is a comparison diagram of the lung CT image before and after super-resolution reconstruction in a specific example of the present invention. DETAILED DESCRIPTION
[0047] The lung CT image super-resolution reconstruction method based on detail enhancement generative adversarial network of the present invention is described in detail below with reference to the accompanying drawings. The embodiments are only for explanation of the present invention and are not intended to limit it.
[0048] Combination Figures 1 to 3 This embodiment provides a medical image super-resolution reconstruction method based on detail enhancement, which includes the following steps:
[0049] S1. Collect high-resolution lung CT images, perform normalization and other preprocessing on the images, and then perform downsampling operations to obtain low-resolution images. The high-resolution images are matched one-to-one with the low-resolution images to form a training set.
[0050] S2. Construct a Gabor detail feature extraction module.
[0051] S3, based on the Gabor detail feature extraction module constructed in step S2, construct a generative adversarial network structure, the components of which include a generator and a discriminator embedded with the Gabor detail feature extraction module, and perform training. The training process is as follows:
[0052] The discriminator parameters are fixed, and the low-resolution images are passed into the generator. The discriminator is used to determine whether the generated high-resolution images are real and accurate to train the generator.
[0053] The generator parameters are fixed, and the discriminator is trained by feeding the high-resolution image and the generator-generated image into the discriminator respectively.
[0054] Repeat the above steps and update the parameters of the generator and discriminator through the loss function until the discriminator determines that the generated image is real and accurate, and the model training is terminated.
[0055] S4. The low-resolution lung CT image obtained during image acquisition is passed into the trained generator to obtain a high-resolution lung CT image with enhanced detail features.
[0056] Through the method in this embodiment, low-resolution lung CT images can be preprocessed, and then the preprocessed images can be input into the detail enhancement generative adversarial network for super-resolution reconstruction to obtain detail-enhanced high-resolution lung CT images, so that tiny lesions in the lung CT images can be displayed more clearly, helping doctors diagnose or carry out subsequent treatment.
[0057] The detail-enhanced super-resolution network includes a generator and a discriminator that use a Gabor detail feature extraction module. The generator is responsible for generating detail-enhanced high-resolution images from noisy images and low-resolution images, and the discriminator is responsible for distinguishing between real images and generated images. By establishing a low-resolution-high-resolution image pair dataset to train a generative adversarial network, the optimal network parameters can be better obtained, thereby better performing detail-enhanced super-resolution reconstruction on the subsequently acquired low-resolution lung CT images.
[0058] The lung CT images described in step S1 of this embodiment are from the Large COVID-19 CT scanslice dataset of the Kaggle project, and a total of 10,000 images are selected. The preprocessing specifically includes the following steps:
[0059] In step S11, the collected high-resolution lung CT images are uniformly converted into grayscale images and cropped to a size of 512×512 as a high-resolution image training set.
[0060] Step S12, using bilinear interpolation to perform a 4-fold downsampling operation on the high-resolution image, and adding Gaussian noise to serve as a low-resolution image training set.
[0061] Step S13: The high-resolution images and the low-resolution images are matched one-to-one to form a complete image pair training set.
[0062] In this embodiment, through step S11, the original images can be standardized to facilitate subsequent processing. Through steps S12 and S13, low-resolution images can be generated from high-resolution images and form data pairs to form a training data set.
[0063] In this embodiment, all images in the training set in step S13 are grayscale images with a size of 512×512, and there are 10,000 data pairs in total.
[0064] Furthermore, combined with Figure 2As shown, in step S2 of this embodiment, a Gabor detail feature extraction module is constructed, and the direction parameter U of the Gabor convolution kernel is selected as 4 (U=4), which specifically includes the following steps:
[0065] Step S21, record the input feature map as , a 3×3 convolution and a Leaky ReLU activation function are used to initially extract image information, and then a 3×3 convolution is used to extract the overall image features and a 3×3 Gabor convolution is used to extract the image texture details. The multi-scale feature map is obtained by splicing. .
[0066] Step S22, the feature map of step S21 Perform primitive extraction to further enhance the extracted texture details. The central area of size 4×4 is denoted as ,Will and According to the synthesis formula, the primitive information is synthesized, that is, the position weight is recorded as , and then and Multiplication further enhances the texture detail features and obtains the output feature map The synthesis formula is recorded as:
[0067]
[0068] The sigmoid function plays a normalization role. The function represents and The measure of composability between them is denoted as:
[0069]
[0070] Combination Figure 3 As shown, the generator described in step S3 of this embodiment adopts a residual neural network, including five convolutions, a Gabor detail feature extraction module constructed in step S2, a 4-fold nearest neighbor upsampling layer, an average pooling layer and a residual connection. The specific operation includes the following steps:
[0071] Step S31, the first three convolutions are connected and the size is 5×5 to preliminarily extract image information, followed by a Gabor detail feature extraction module to extract texture detail information.
[0072] Step S32, clipping the output of the first convolution to the output of the Gabor detail feature extraction module through a residual connection to obtain a feature map.
[0073] In step S33, the feature map resolution is improved through a 4-fold nearest neighbor upsampling layer, followed by two 3×3 convolutions.
[0074] In step S34, the feature map is finally reduced in dimension through the average pooling layer to ensure that the input and output dimensions are consistent.
[0075] Combination Figure 4 As shown, the discriminator described in step S3 of this embodiment adopts a dual-path U-Net network, including a main path and a detail extraction path, and both paths are divided into an encoder and a decoder. The specific operation of the main path includes the following steps:
[0076] In step S35, the encoder part extracts image features by a 3×3 convolution, which is activated by a Leaky ReLU activation function, followed by a 2×2 maximum pooling layer for downsampling.
[0077] Step S36, repeat the operation described in step S35 three times to reduce the image size and obtain the shallow feature map ps of the image.
[0078] In step S37, the decoder part upsamples the feature map ps in step S36 by a deconvolution of size 2×2 and a convolution of size 3×3.
[0079] Step S38, repeat the operation described in step S37 three times to restore the image size and obtain the deep feature map pd of the image.
[0080] The structure of the detail extraction pathway is consistent with that of the main pathway, and the specific operation includes the following steps:
[0081] Step S39, the encoder part extracts image detail features by the Gabor detail feature extraction module, activates by the Leaky ReLU activation function, and then performs downsampling by a maximum pooling layer of size 2×2.
[0082] Step S310, repeating the operation described in step S39 three times to reduce the image size and obtain a shallow detail feature map qs of the image.
[0083] Step S311, same as the main channel decoder, repeats the operation described in step S37 three times to restore the image size and obtain the deep detail feature map qd of the image.
[0084] Step S312, in order to prevent the loss of original feature information, the shallow feature map ps described in step S36 and the deep feature map pd described in step S38 are spliced with corresponding sizes through jump connections to obtain the overall feature map p, and the shallow feature map qs described in step S310 and the deep feature map qd described in step S311 are spliced with corresponding sizes through jump connections to obtain the detail feature map q, and then the detail feature map q is spliced to the overall feature map p through a jump connection and merged and reduced in dimension through a 1×1 convolution. The obtained discriminator output is the authenticity value corresponding to each pixel value and whether the input image is true and accurate.
[0085] In this embodiment, the discriminator parameters are fixed, and the loss function of the generator is minimized through the back-propagation algorithm, thereby better realizing the training of the generator.
[0086] In this embodiment, the generator parameters are fixed, and the loss function of the discriminator is minimized through the back propagation algorithm, thereby better realizing the training of the discriminator.
[0087] In this embodiment, the loss function of the discriminator is Including encoder Loss Function and Decoder The loss function is recorded as:
[0088]
[0089]
[0090]
[0091] in represents the input of the discriminator, represents the random noise input to the generator, represents the expectation on all samples, represents the encoder of the discriminator, represents the decoder of the discriminator, refers to the discriminator's judgment result at this pixel, Represents a generator.
[0092] In this embodiment, the loss function of the generator is recorded as:
[0093]
[0094] in is the context loss function, is the weight, is the random noise input to the generator. The calculation formula is as follows:
[0095]
[0096] In this embodiment, the loss function of the generator introduces context loss, so that the generated image retains as much context information as possible, reduces the distortion of the generated image, and improves the authenticity and accuracy of the generated image.
[0097] More specifically, the parameters for generative adversarial network training in this embodiment are set as follows: batch size is 8, the learning rate of the generator is 0.0002, the learning rate of the discriminator is 0.0005, the total number of iterations is 25000, and the optimizer adopts Adam optimizer.
[0098] In this embodiment, the acquired low-resolution lung CT image is passed into the trained generator to obtain a high-resolution lung CT image with enhanced details.
[0099] Combination Figure 5 As shown in the figure, it is a comparison diagram of the effect of the lung CT image super-resolution reconstruction method based on detail enhancement generative adversarial network described in this embodiment. More specifically, the left figure is a low-resolution lung CT image obtained during a medical examination, and the right figure is a high-resolution lung CT image generated by super-resolution reconstruction of the method described in the present invention. As shown in the effect comparison diagram, the present invention has a good effect in terms of image clarity, image authenticity, image texture detail enhancement, and image context information retention.
[0100] The embodiments of the present invention are described in detail above with reference to the accompanying drawings, but the present invention should not be limited to the embodiments described. For those skilled in the art, any modification, substitution, variation and improvement made without departing from the principle and spirit of the present invention should be included in the protection scope of the present invention.
Claims
1. A medical image super-resolution reconstruction method based on detail enhancement, characterized in that: The steps include: S1, collect lung CT images, perform downsampling operation on the images after preprocessing, and form a training set by corresponding the images before and after downsampling one by one; S2. Construct a generative adversarial network structure including a generator and a discriminator, and perform training based on image data before and after downsampling; the generator adopts a residual neural network, including five convolutions, a Gabor detail feature extraction module, a nearest neighbor upsampling layer, an average pooling layer, and a residual connection; the specific implementation includes the following steps: Step S21, the first three convolutions are connected to initially extract image information, followed by a Gabor detail feature extraction module to extract texture detail information; the Gabor detail feature extraction module consists of two convolutions, a Gabor convolution, a primitive extraction layer and an average pooling layer, and is specifically implemented as follows: The input feature map is recorded as I i , a convolution and activation function are used to initially extract image information, and then a convolution is used to extract the overall image features and a Gabor convolution is used and an activation function is used to extract the image texture detail features, and the multi-scale feature map I is obtained by splicing. t ; For feature map I t Perform primitive extraction and select I t The central area of size n×n is denoted as X tc , will I t With X tc According to the synthesis formula, the primitive information is synthesized, and the position weight is recorded as P w , and then P w with I t Multiply and enhance the texture detail features to obtain the output feature map I o ; The synthesis formula is recorded as: P w =sigmoid(Φ(I t ,X tc )); The Φ function represents the t With X tc The measure of composability between them is denoted as: Φ(I t ,X tc )=-(I t -X tc ) 2 ; The feature map I o After an average pooling, the dimension is reduced and the number of parameters is reduced; Step S22, clipping the output of the first convolution to the output of the Gabor detail feature extraction module through a residual connection to obtain a feature map H; Step S23, the resolution of the feature map H is improved through a nearest neighbor upsampling layer, followed by two convolutions; Step S24, finally, the feature map is reduced in dimension through the average pooling layer to ensure that the input and output dimensions are consistent; S3. The lung CT image obtained during image acquisition is passed into the trained generator for testing to obtain a lung CT reconstructed image with enhanced detail features.
2. The method for medical image super-resolution reconstruction based on detail enhancement according to claim 1, characterized in that: The preprocessing in step S1 is specifically implemented as follows: Step S11, converting the collected lung CT images into grayscale images and cropping them to a consistent size as an image training set before downsampling; Step S12, downsampling the high-resolution image using bilinear interpolation, adding Gaussian noise, and using the downsampled image as a training set; Step S13, CT images before and after downsampling are matched one by one to form a complete image pair training set.
3. The method for medical image super-resolution reconstruction based on detail enhancement according to claim 2, characterized in that: The discriminator described in step S2 adopts a dual-path U-Net network, including a main path and a detail extraction path, and both paths are divided into an encoder and a decoder part; The main path is specifically implemented as follows: Step S25, the encoder extracts image features by a convolution, passes an activation function, and then performs downsampling by a maximum pooling layer; Step S26, repeating the operation in step S25 m times to obtain a shallow feature map ps of the image; Step S27, the decoder upsamples the feature map ps by a deconvolution and a convolution; Step S28, repeating the operation in step S27 m times to obtain a deep feature map pd of the image; The detail extraction path is specifically implemented as follows: Step S29, the encoder part extracts image detail features by a Gabor detail feature extraction module, activates it through an activation function, and then performs downsampling through a maximum pooling layer; Step S210, repeating the operation in step S29 m times to obtain a shallow detail feature map qs of the image; Step S211, same as the main channel decoder, repeat the operation in step S27 m times to obtain a deep detail feature map qd of the image; Step S212, the feature map ps and the feature map pd are spliced with corresponding sizes through jump connections to obtain the overall feature map p; the feature map qs and the feature map qd are spliced with corresponding sizes through jump connections to obtain the detail feature map q; the detail feature map q is then spliced to the overall feature map p through a jump connection and merged and reduced in dimension through a 1×1 convolution. The discriminator output obtained is the authenticity value corresponding to each pixel value and whether the input image is true and accurate.
4. The method for medical image super-resolution reconstruction based on detail enhancement according to claim 3, characterized in that: The training process of the generative adversarial network structure is as follows: The discriminator parameters are fixed, and the downsampled CT images are passed into the generator. The discriminator is used to determine whether the generated high-resolution images are real and accurate to train the generator. The generator parameters are fixed, and the discriminator is trained by inputting the CT image before downsampling and the image generated by the generator into the discriminator respectively; Repeat the above operations to train the generator and discriminator, and update the parameters of the generator and discriminator through the loss function until the discriminator determines that the generated image is real and accurate, and the training is terminated.
5. The method for medical image super-resolution reconstruction based on detail enhancement according to claim 4, characterized in that: During training, the loss function of the discriminator is Including encoder perceptual loss Loss Function and Decoder Perceptual Loss The loss function is recorded as: The loss function of the generator is: in is the context loss function, represents the encoder of the discriminator, represents the decoder of the discriminator, Refers to the discriminator's judgment result at this pixel, x represents the input of the discriminator, E represents the expectation on all samples, G represents the generator, μ is the weight, and z is the random noise input to the generator. The calculation formula is as follows:
Citation Information
Patent Citations
Progressive generative adversarial network for low-dose CT image noise reduction and artifact removal
CN112837244A
Multi-scale connection generative adversarial network medical image super-resolution reconstruction method
CN116612009A