An ultra-low-dose CT image reconstruction method based on image domain

Through the ultra-low dose CT image reconstruction method based on the image domain end, the SRResNet model and wavelet transformation are used to optimize image quality, and the problem of CT image reconstruction relying on high-end devices and compatibility is solved, high-quality ULDCT image reconstruction is achieved, and the effect of lung cancer screening is improved.

CN120047563BActive Publication Date: 2025-08-12PEKING UNIVERSITY THIRD HOSPITAL (THE THIRD CLINICAL MEDICAL SCHOOL OF PEKING UNIVERSITY) +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510123065.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-08-12
Estimated Expiration
2045-01-26

AI Technical Summary

Technical Problem

The existing CT image reconstruction methods rely on high-end CT scanning equipment, with poor image quality and CT scanners from different manufacturers are incompatible, which limits the widespread application of ultra-low dose CT lung cancer screening.

Method used

Based on the ultra-low dose CT image reconstruction method at the image domain end, the training sample set and model training unit is constructed, and the feature extraction and reconstruction is performed using the SRResNet model, residual density block module and convolution module, and iterative training is performed by combining the generator composite loss and discriminator loss, and the image quality is optimized by wavelet transform.

Benefits of technology

It significantly improves the clarity and quality of ultra-low dose CT images, reduces artifacts, and improves the detection rate of lung nodules, especially for non-solid and small-sized nodules, supporting the large-scale application of ULDCT lung cancer screening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047563B_ABST
    Figure CN120047563B_ABST
Patent Text Reader

Abstract

The present invention relates to an ultra-low-dose CT image reconstruction method based on an image domain. The method comprises: obtaining high- and low-dose paired CT images of multiple patients, constructing a training sample set using the ultra-low-dose CT images in the paired CT images as input samples and the low-dose CT images as labels; preliminarily constructing a CT image reconstruction unit and a model training unit, and training the preliminarily constructed CT image reconstruction unit based on the training sample set and the model training unit; the model training unit is configured to receive the reconstructed image and the corresponding low-dose CT image, and iteratively train the CT image reconstruction unit based on a generator composite loss and a discriminator loss fused from multiple loss functions to obtain a converged CT image reconstruction unit; and inputting the ultra-low-dose CT image to be processed into the trained CT image reconstruction unit to obtain a reconstructed CT image. The present invention solves the problem that CT image reconstruction methods in the prior art rely on high-end CT scanning equipment and have poor image reconstruction quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of tomography, and in particular relates to an ultra-low-dose CT image reconstruction method based on an image domain end. Background Art

[0002] Lung cancer is currently the leading cause of death from malignant tumors, and delayed diagnosis is the primary factor impacting the reduction in lung cancer mortality. Currently, low-dose computed tomography (LDCT) is the primary recommended method for lung cancer screening. Although guidelines recommend a 1.5 mSv LDCT dose for lung cancer screening, this is still significantly higher than the 0.03-0.1 mSv radiation dose of chest X-rays, and the cumulative radiation risk from repeated examinations over multiple years cannot be ignored. Therefore, an increasing number of researchers are attempting to use ultra-low-dose CT (ULDCT) for lung cancer screening. ULDCT is a sub-mSv chest CT scanning method that emerged with the development of iterative reconstruction (IR) algorithms. Its radiation dose is comparable to that of chest X-rays and is generally considered to be below 0.2 mSv.

[0003] However, due to the extremely low radiation dose of the scan, noise in ULDCT images can significantly increase, reducing image quality. Therefore, ULDCT examinations must rely on advanced image reconstruction algorithms to reduce image noise and meet minimum clinical diagnostic requirements. The IR algorithm, by introducing a mathematical model to selectively identify and remove image noise, has achieved some improvement in image quality. However, excessive noise reduction can introduce new wax artifacts that interfere with diagnosis. Therefore, further improvement in the image quality of ULDCT based on the IR algorithm is necessary to reduce the rate of missed nodules.

[0004] With the advancement of computer technology and the rapid development of artificial intelligence, deep learning (DL) algorithms, which are more suitable for ULDCT image reconstruction, have emerged. Compared with IR algorithms, DL algorithms can further reduce image noise and improve image quality. However, DL post-processing algorithms are currently only available on high-end CT scanners from major manufacturers. These CT scanners are typically used for more complex examinations and are rarely used for physical examinations. At the same time, a large number of grassroots hospitals are rarely equipped with these high-end CT scanners, and the vast majority of CT scanners they currently have cannot be upgraded to use DL post-processing algorithms, further limiting the coverage of ULDCT scans. In addition, the DL post-processing algorithms provided by major manufacturers all perform image post-processing and reconstruction on the raw data obtained from CT scans from the projection domain. Raw data from different manufacturers is incompatible with the DL algorithms of other manufacturers.

[0005] In summary, these factors significantly limit the widespread application of ULDCT lung cancer screening. In current medical imaging diagnostics, the universal image data format standard is the Digital Imaging and Communications in Medicine (DICOM). Therefore, the development of a broadly applicable DL image reconstruction technique for post-processing all DICOM images, independent of specific CT scanner types, is a prerequisite for the large-scale clinical application of ULDCT lung cancer screening. With the advancement of artificial intelligence research, it has become possible to enhance the quality of CT images in the image domain based on DICOM images. Summary of the Invention

[0006] In view of the above analysis, the present invention aims to disclose an ultra-low-dose CT image reconstruction method based on the image domain end, which solves the problem that the CT image reconstruction method in the prior art depends on high-end CT scanning equipment or has poor image reconstruction quality.

[0007] The purpose of the present invention is mainly achieved through the following technical solutions:

[0008] The present invention discloses an ultra-low-dose CT image reconstruction method based on an image domain end, comprising:

[0009] Acquire high- and low-dose paired CT images of multiple patients at the image domain end, wherein the high- and low-dose paired CT images include ultra-low-dose CT images and low-dose CT images of the same part of the corresponding patients, and construct a training sample set using the ultra-low-dose CT images as input samples and the low-dose CT images as labels;

[0010] A CT image reconstruction unit and a model training unit are preliminarily constructed, and the preliminarily constructed CT image reconstruction unit is trained based on the training sample set and the model training unit; the model training unit is used to receive the reconstructed image output by the CT image reconstruction unit and the corresponding low-dose CT image in the training sample set, and iteratively train the CT image reconstruction unit based on the generator composite loss and the discriminator loss fused by multiple loss functions to obtain a converged CT image reconstruction unit;

[0011] The ultra-low-dose CT image to be processed is obtained and input into the trained CT image reconstruction unit to obtain a reconstructed CT image.

[0012] Furthermore, the CT image reconstruction unit is constructed based on the SRResNet model, including a feature extraction layer, a residual density block module and a convolution module;

[0013] The feature extraction layer extracts features from the input ultra-low-dose CT image through a convolution operation to obtain a primary feature map;

[0014] The residual density block module extracts a high-level feature map based on the primary feature map by fusing a multi-level residual network with dense connections;

[0015] The convolution module is used to fuse the high-level feature maps and adjust the dimensions to obtain a reconstructed image.

[0016] Furthermore, the residual density block module includes multiple RRDB modules arranged in sequence, each RRDB module is composed of five residual blocks arranged in sequence, and each residual block also contains a sub-residual block; each sub-residual block is composed of multiple densely connected 3×3 convolutional layers, the convolution step size is 1, and LeakyReLU is used as the activation function.

[0017] Furthermore, the discriminator is a PatchGAN architecture; comprising n convolutional layers arranged in sequence;

[0018] Among them, the 2nd to nth convolutional layers also include batch normalization layers; the 1st to n-1th convolutional layers use the LeakyReLU activation function, and the nth convolutional layer uses the Sigmoid activation function to output the discrimination result of each patch.

[0019] Furthermore, the generator composite loss includes the first adversarial loss, content loss, SSIM loss and first perceptual loss; the generator composite loss is expressed as:

[0020] L G =αL Gadv +β·L content +γ·L SSIM +δL perceptual ;

[0021] Among them, L Gadv is the first adversarial loss, L content is the content loss, L SSIM is the SSIM loss; L perceptual is the first perceptual loss, and α, β, γ, and δ are weight coefficients.

[0022] Furthermore, the first adversarial loss is expressed as:

[0023] L Gadv = -E[log(D(G(x)))];

[0024] Where E represents averaging, G(x) is the reconstructed image, and D(G(x)) is the discrimination result of the reconstructed image;

[0025] The content loss is expressed as:

[0026]

[0027] Among them, G(x) i is the i-th pixel value of the reconstructed image, y i is the ith pixel value of the real low-dose CT image, and N is the total number of image pixels.

[0028] The SSIM loss is expressed as:

[0029] L SSIM =1-SSIM(G(x),y);

[0030] Where G(x) is the reconstructed image and y is the real low-dose CT image;

[0031] The first perceptual loss is expressed as:

[0032] L perceptual =E[DISTS(G(x),y)];

[0033] Among them, DISTS() is the image structure and texture similarity evaluation model;

[0034] The discriminator is iteratively optimized using the following adversarial loss:

[0035] L D =-E[logD(y)]-E[log(1-D(G(x)))];

[0036] Where E[] represents averaging, D represents the discriminator model, G represents the generator model, y represents the real LDCT image, and G(*) represents the output of the generator.

[0037] Furthermore, after receiving the reconstructed image output by the CT image reconstruction unit and the corresponding low-dose CT image in the training sample set, the model training unit further performs wavelet transform on the reconstructed image and the corresponding low-dose CT image, and iteratively trains the CT image reconstruction unit based on the high-frequency subband and low-frequency subband obtained by the wavelet transform by the following method;

[0038] Inputting the ultra-low-dose CT images in the training sample set into the CT image reconstruction unit to obtain a reconstructed image;

[0039] Based on the multiple high-frequency sub-bands corresponding to the reconstructed image and the low-dose CT image, using a discriminator to distinguish true from false on the reconstructed image, and generating a discrimination result;

[0040] The CT image reconstruction unit is iteratively updated based on the low-frequency sub-band and high-frequency sub-band corresponding to the reconstructed image and the low-dose CT image and the discrimination result of the discriminator to obtain a converged CT image reconstruction unit.

[0041] Furthermore, the discriminator is iteratively optimized using the discriminator loss constructed by the high-frequency sub-band, and the discriminator and the CT image reconstruction unit generate an adversarial architecture for training the CT image reconstruction unit;

[0042] The discriminator loss is expressed as:

[0043] L D = -E[logD(SWT(y)) * ]-E[log(1-D(SWT(G(x)) * ))];

[0044] Where E[] represents averaging, x is an ultra-low-dose CT image, G(x) is a reconstructed image generated by a CT image reconstruction unit, y is a low-dose CT image, and SWT(y) is * represents the detail subband connection after wavelet transform of low-dose CT image, D represents the discriminator model, SWT(G(x)) * It represents the connection of detail subbands after wavelet transformation of the reconstructed image, that is, connecting each high-frequency subband in the image channel dimension, and G represents the generator model.

[0045] Furthermore, based on the low-frequency subband and high-frequency subband corresponding to the reconstructed image and the low-dose CT image and the discrimination result of the discriminator, the CT image reconstruction unit is iteratively updated by a generator composite loss that fuses the wavelet domain loss, the second adversarial loss, and the second perceptual loss; its loss function is expressed as:

[0046] L=L SWT +λ adv ·L adv,Gg +λ perc ·L perc ;

[0047] Among them, L SWT is the wavelet domain loss, L adv,G is the second adversarial loss, L perc is the second perceptual loss; perc and λ adv is the weight coefficient.

[0048] Furthermore, the wavelet domain loss is expressed as:

[0049] L SWT =E[∑ j λ j (α1‖SWY(G(x)) j -SWT(y) j ‖1+(1-α1)(1-SSIM(x(G(x)) j ,SWt(y) j)))];

[0050] Where E[] represents averaging, x is the ultra-low-dose CT image, G(x) is the reconstructed image generated by the CT image reconstruction unit, y is the low-dose image, α1 is the weight, and SWT(G(x)) j represents the jth high-frequency subband and low-frequency subband obtained by performing stationary wavelet transform on the reconstructed image, SWT(y) j represents the jth high-frequency subband or low-frequency subband obtained after the stationary wavelet transform of the low-dose CT image; λ is the scaling factor, and SSIM() is the structural similarity index between the reconstructed image and the low-dose CT image after wavelet transform;

[0051] The second adversarial loss is expressed as:

[0052] L adv,G = -E[log(D(SWT(G(x))*))];

[0053] Where G(x) is the reconstructed image, D(SWT(G(x))*)) is the discrimination result of the detail subband of the reconstructed image, SWT represents wavelet transform, and * is the connection of detail subbands, that is, connecting each high-frequency subband in the image channel dimension;

[0054] The second perceptual loss is expressed as:

[0055] L perc =E[DISTS(G(x),y)];

[0056] Among them, DISTS() is an image structure and texture similarity evaluation model.

[0057] The present invention can achieve at least the following beneficial effects:

[0058] The present invention proposes an ultra-low-dose CT image reconstruction method based on the image domain. It uses a large number of paired ultra-low-dose and low-dose CT images for model training, integrates the RRDB module in the generator to implement the application of global residual connection, and uses the generator's adversarial loss, content loss, SSIM loss, and perceptual loss to better preserve the edge and detail information of the image and capture the overall structural characteristics and overall visual perception characteristics of the image; or through wavelet domain loss optimization, the reconstructed image has fewer artifacts and higher clarity than the original ultra-low-dose CT image; and the CT image reconstruction unit of the present invention removes the upsampling layer in the traditional model, avoids artifacts and retains the original information, speeds up image processing, and provides strong technical support for the large-scale application of ULDCT lung cancer screening; it has been verified in specific cases and helps doctors observe the lung tissue structure and pathological characteristics more accurately. It significantly improves the detection rate of lung nodules, especially for non-solid and small-sized nodules, and reduces missed diagnoses. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] The accompanying drawings are only used for the purpose of illustrating specific embodiments and are not to be considered as limiting the present invention. Throughout the drawings, the same reference symbols denote the same components.

[0060] Figure 1 This is a flow chart of an ultra-low-dose CT image reconstruction method based on an image domain end in an embodiment of the present invention;

[0061] Figure 2 Schematic diagram of the structure of a CT image reconstruction unit in an embodiment of the present invention;

[0062] Figure 3 Schematic diagram of the structure of the residual dense block (RRDB) in an embodiment of the present invention;

[0063] Figure 4 Schematic diagram of the structure of the residual module in the residual dense module (RRDB) in an embodiment of the present invention;

[0064] Figure 5 is a chest CT image of a patient in an embodiment of the present invention, wherein: Figure 5 (a) is a low-dose CT image; Figure 5 (b) is the ultra-low-dose CT image (ULDCT) processed based on the iterative reconstruction (IR) algorithm; Figure 5 (c) is a schematic diagram of the ULDCT image (ULDCT-DL) processed by the CT image reconstruction algorithm of this embodiment. DETAILED DESCRIPTION

[0065] The preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, which constitute a part of this application and are used to illustrate the principles of the present invention together with the embodiments of the present invention.

[0066] The embodiment of the present invention discloses an ultra-low-dose CT image reconstruction method based on the image domain end, such as Figure 1 As shown, the following steps are included:

[0067] Step S1: Acquire high- and low-dose paired CT images of multiple patients at the image domain end, wherein the high- and low-dose paired CT images include ultra-low-dose CT images and low-dose CT images of the same part of the corresponding patient, and construct a training sample set using the ultra-low-dose CT images as input samples and the low-dose CT images as labels;

[0068] Specifically, in this embodiment, ultra-low-dose CT images and low-dose CT images of the same part of multiple patients are obtained respectively to form high- and low-dose paired CT images, which are converted into image domain end images in DICOM format. After preprocessing operations such as cropping and normalization, a training sample set is constructed; the dose of ultra-low-dose CT images (ULDCT) is generally 0.17mSv, and the dose of low-dose CT images (LDCT) is generally 0.81mSv; for patients undergoing lung nodule screening, low-dose screening is usually recommended. With the advancement of equipment, the quality of low-dose CT images has been greatly improved and can fully meet diagnostic requirements. Low-dose CT images are clearer and more detailed than ultra-low-dose CT images, and their spatial resolution and accuracy are higher than those of ultra-low-dose CT images, and tiny details and lesions in the lungs can be observed; therefore, in this embodiment, low-dose CT images are used as labels and the corresponding ultra-low-dose CT images are used as input samples to learn the structural details of low-dose CT images and reconstruct ultra-low-dose CT images into high-quality CT images.

[0069] Step S2: Preliminarily constructing a CT image reconstruction unit and a model training unit, and training the preliminarily constructed CT image reconstruction unit based on the training sample set and the model training unit; the model training unit is used to receive the reconstructed image output by the CT image reconstruction unit and the corresponding low-dose CT image in the training sample set, and iteratively train the CT image reconstruction unit based on the generator composite loss and the discriminator loss fused by multiple loss functions to obtain a converged CT image reconstruction unit;

[0070] Specifically, the CT image reconstruction unit in this embodiment is a generator, which forms a generative adversarial architecture with the discriminator of the model training unit. The (generator) is used to reconstruct ultra-low-dose CT images to generate clearer, high-quality CT images. The CT image reconstruction unit is constructed based on the SRResNet model and includes a feature extraction layer, a residual density block module, and a convolution module.

[0071] The feature extraction layer extracts features from the input ultra-low-dose CT image through a convolution operation to obtain a primary feature map;

[0072] The residual density block module extracts a high-level feature map based on the primary feature map by fusing a multi-level residual network with dense connections;

[0073] The convolution module is used to fuse the high-level feature maps and adjust the dimensions to obtain a reconstructed image.

[0074] The residual density block module includes multiple RRDB modules arranged in sequence, each RRDB module is composed of five residual blocks arranged in sequence, and each residual block also contains a sub-residual block; each sub-residual block is composed of multiple densely connected 3×3 convolutional layers, the convolution step size is 1, and LeakyReLU is used as the activation function.

[0075] That is, Figure 2 As shown, the CT image reconstruction unit (generator) of this embodiment is constructed based on the SRResNet framework; and the residual density block module (RRDB) is introduced to perform global feature fusion on the image. First, the convolutional layer is used to extract features of the input ultra-low dose CT (ULDCT) image, and then the extracted features are input into the RRDB module. The RRDB module combines the advantages of multi-level residual networks and dense connections, and can efficiently mine high-level features from images. That is, in this embodiment, the overall generator network is constructed by superimposing multiple RRDB modules, and the original residual block is replaced by the RRDB module, such as Figure 3 As shown in , each RRDB module consists of five closely connected residual blocks, and these residual blocks also contain sub-residual blocks. Figure 4 As shown in Figure 1, each sub-residual block consists of multiple 3×3 convolutional layers, the convolution stride is set to 1, and LeakyReLU is used as the activation function.

[0076] Moreover, the CT image reconstruction unit of this embodiment abandons the upsampling layer in the traditional SRResNet architecture, so that the output reconstructed image size is consistent with the input ULDCT image, avoiding artifacts and blurring caused by additional interpolation or estimation processes, better preserving the original information of the input image, and improving network computing efficiency.

[0077] Furthermore, the discriminator includes n convolutional layers arranged in sequence; wherein, the 2nd to nth convolutional layers also include a batch normalization layer; the 1st to n-1th convolutional layers use a LeakyReLU activation function, and the nth convolutional layer uses a Sigmoid activation function to output the discrimination result of each patch.

[0078] The use of the discriminator to perform true or false discrimination on the reconstructed image includes: using the discriminator to divide the received reconstructed image and the low-dose CT image into multiple patches, and based on each patch of the low-dose CT image, performing true or false discrimination on the patch corresponding to the reconstructed image, obtaining the true or false probability of each patch, taking the average of the true or false probabilities of all patches, and obtaining the discrimination result of the reconstructed image. That is, the discriminator in this embodiment adopts the PatchGAN architecture, the core idea of which is to divide the input image into multiple small blocks (patches), and then perform true or false discrimination on each small block, rather than performing true or false discrimination on the entire image as a whole. It allows the discriminator to pay more attention to the local structure of the image, and is more effective in identifying detailed information in medical images (such as lung nodules, etc.). The discriminator's input layer receives the reconstructed image (ULDCT-DL image) or low-dose CT image output by the CT image reconstruction unit. It then passes through four to five convolutional layers for feature extraction. Each convolutional layer uses a kernel size of 4×4, a stride of 2, and padding of 1 to gradually reduce the spatial resolution of the image and increase the number of feature channels. The negative slope of the LeakyReLU is set to 0.2, preserving negative feature information while addressing the issue of zero gradient on the negative semi-axis of the ReLU function, making network training more stable. Furthermore, except for the first convolutional layer, batch normalization layers are added after each convolutional layer. Batch normalization helps accelerate network convergence and prevent vanishing or exploding gradients. The output of the final convolutional layer passes through a sigmoid activation function, generating a two-dimensional array. Each element in the array corresponds to the discrimination result for a small image patch (0 for false, 1 for true). This allows the discriminator to assess the authenticity of the image from the perspective of local details, thereby guiding the generator to produce images that are more realistic.

[0079] Furthermore, for the training of the CT image reconstruction unit, a specific embodiment of the present invention performs iterative training on the CT image reconstruction unit based on a generator composite loss and a discriminator loss fused with multiple loss functions; the generator composite loss specifically includes a first adversarial loss, a content loss, an SSIM loss, and a first perceptual loss; the generator composite loss is expressed as:

[0080] L G =αL Gadv +β·L content +γ·L SSIM +δL perceptual ;

[0081] Among them, L Gadv is the first adversarial loss, L content is the content loss, L SSIM is the SSIM loss; Lperceptual is the first perceptual loss, and α, β, γ, and δ are weight coefficients.

[0082] The first adversarial loss is expressed as:

[0083] L Gadv = -E[log(D(G(x)))];

[0084] Where E represents averaging, G(x) is the reconstructed image, and D(G(x)) is the discrimination result of the reconstructed image;

[0085] The content loss is expressed as:

[0086]

[0087] Among them, G(x) i is the i-th pixel value of the reconstructed image, y i is the ith pixel value of the real low-dose CT image, and N is the total number of image pixels.

[0088] The SSIM loss is expressed as:

[0089] L SSIM =1-SSIM(G(x),y);

[0090] Where G(x) is the reconstructed image and y is the real low-dose CT image;

[0091] The first perceptual loss is expressed as:

[0092] L perceptual =E[DISTS(G(x),y)];

[0093] Among them, DISTS() is an image structure and texture similarity evaluation model.

[0094] That is, in the generative adversarial architecture, the first adversarial loss aims to enable the generated image to deceive the discriminator and be as close to the real low-dose CT (LDCT) image as possible. By minimizing the generator adversarial loss, the generator is prompted to continuously optimize the generated image so that it is closer to the real image under the judgment of the discriminator. The purpose of this form of adversarial loss function is to encourage the image generated by the generator to be judged as a real image in the discriminator, because D(G(*)) represents the discriminant's judgment result on the generator's output image, that is, the judgment probability. When taking its logarithm and taking the negative expectation, the goal of the generator is to maximize this value, that is, to make the discriminator judge that the generated image is real, thereby promoting the generator to continuously improve the quality of the generated image and make it closer to the real LDCT image distribution.

[0095] For content loss, we use the L1 loss to accurately measure the pixel-level difference between the reconstructed ULDCT image generated by the generator and the true LDCT image. This loss function is not overly sensitive to outliers in the image and, compared to loss functions such as the mean square error (MSE), can better preserve image edges and details, thereby improving the quality of the generated image to a certain extent.

[0096] SSIM loss is used to evaluate the structural similarity between the generated image and the real image. The SSIM indicator comprehensively considers the brightness, contrast and structural information of the image, so that the image generated by the generator is structurally closer to the real LDCT image. Compared with the loss function based only on pixel differences, SSIM loss can better capture the overall structural characteristics of the image, which helps to generate images with more natural visual effects and more similar structures to real images, thereby improving image quality.

[0097] The first perceptual loss uses DISTS to evaluate the generated image G() (the reconstructed ULDCT image generated by the generator) and the ground-truth LDCT image. DISTS is a metric specifically designed to measure the similarity of image structure and texture. Based on deep feature extraction, it comprehensively analyzes the structural and texture information of an image at different scales and semantic levels. Using a specific mathematical model, it calculates the difference between the generated image and the ground-truth image to produce the evaluation result. Its design focuses on capturing the overall visual perceptual characteristics of an image and is sensitive to the combined similarity of local details and global structure. In the medical imaging field, such as lung CT image processing, where precise attention must be paid to the subtle structure, texture variations, and overall morphological similarity of lung tissue, DISTS is better suited to this requirement. It can more accurately assess the similarity between the generated ULDCT image and the ground-truth LDCT image in clinically relevant structural and texture details, helping to improve the accuracy of the model's performance evaluation in medical image reconstruction tasks and ensure that the reconstructed images meet the stringent image quality requirements of clinical diagnosis. By optimizing the perceptual loss, the generator learns more realistic image details, reduces artifacts and spurious structures, and produces high-quality images.

[0098] Furthermore, the task of the discriminator is to distinguish the reconstructed image output by the generator from the real LDCT image. In this embodiment, the discriminator loss is expressed as:

[0099] L D =-E[logD(y)]-E[log(1-D(G(x)))];

[0100] Where E[] represents averaging, D represents the discriminator model, G represents the generator model, y represents the real LDCT image, and G(*) represents the output of the generator.

[0101] By minimizing the discriminator's loss, the discriminator can better learn to distinguish real and fake images, thereby guiding the generator to produce images that are more realistic. During the training process, the generator and discriminator compete with each other and continuously optimize, gradually improving the quality of the generated images.

[0102] In particular, with respect to the training of the CT image reconstruction unit, in another specific embodiment of the present invention, after receiving the reconstructed image output by the CT image reconstruction unit and the corresponding low-dose CT image in the training sample set, the model training unit further performs wavelet transform on the reconstructed image and the corresponding low-dose CT image, and iteratively trains the CT image reconstruction unit based on the high-frequency subband and low-frequency subband obtained by the wavelet transform using the following method;

[0103] Inputting the ultra-low-dose CT images in the training sample set into the CT image reconstruction unit to obtain a reconstructed image;

[0104] Based on the multiple high-frequency sub-bands corresponding to the reconstructed image and the low-dose CT image, using a discriminator to distinguish true from false on the reconstructed image, and generating a discrimination result;

[0105] The CT image reconstruction unit is iteratively updated based on the low-frequency sub-band and high-frequency sub-band corresponding to the reconstructed image and the low-dose CT image and the discrimination result of the discriminator to obtain a converged CT image reconstruction unit.

[0106] Preferably, this embodiment iteratively updates the CT image reconstruction unit by fusing the wavelet domain loss, the second adversarial loss, and the second perceptual loss based on the low-frequency subbands and high-frequency subbands corresponding to the reconstructed image and the low-dose CT image and the discrimination result of the discriminator; the loss function is expressed as:

[0107] L G =L SWT +λ adv ·L adv,G +λ perc ·L perc ;

[0108] Among them, L SWT is the wavelet domain loss, L adv,G is the second adversarial loss, L perc is the second perceptual loss; perc and λ adv is the weight coefficient obtained through training.

[0109] Wherein, the wavelet domain loss is expressed as:

[0110] L SWT =E[∑ j λ j(α1‖SWT(G(x)) j -SWT(y) j ‖1+(1-α1)(1-SSIM(SWT(G(x)) j ,SWT(y) j )))];

[0111] Where E[] represents averaging, x is the ultra-low-dose CT image, G(x) is the reconstructed image generated by the CT image reconstruction unit, y is the low-dose image, α1 is the weight, and SWT(G(x)) j represents the jth high-frequency subband and low-frequency subband obtained by performing stationary wavelet transform on the reconstructed image, SWT(y) j represents the jth high-frequency subband or low-frequency subband obtained after stationary wavelet transform of the low-dose CT image; λ is the scaling factor, and SSIM() is the structural similarity index between the reconstructed image and the low-dose CT image after wavelet transform.

[0112] Specifically, the scaling factor λ is used to effectively control the generated high-frequency details according to the characteristics and importance of different sub-bands, such as avoiding visual artifacts around fine structures such as small blood vessels and nodule edges in the lungs; α is used to adjust L1 (i.e., ‖SWT(G(x)) j -SWT(y) j ‖) is a key parameter of the loss and SSIM loss weight, ranging from 0 to 1. Its value is obtained through training to achieve the best loss function balance effect.

[0113] In this combined loss function, the L1 loss is mainly used to accurately measure the differences in images at the pixel level, strengthen the learning of high-frequency details (such as subtle density changes in the lung parenchyma, nodule edges, etc.), and ensure the basic accuracy and detail integrity of the image; while the SSIM loss comprehensively evaluates image similarity from three aspects: brightness, contrast, and structure, ensuring the integrity and similarity of complex lung structures (such as lung tissue morphology, bronchial branching structure, etc.), and is more in line with the human visual system's perception of image quality, thereby improving the clinical usability of reconstructed images.

[0114] Furthermore, the second adversarial loss is expressed as:

[0115] L adv,G = -E[log(D(SWT(G(x)) * ))];

[0116] Where G(x) is the reconstructed image, D(SWT(G(x))*)) is the discrimination result of the detail subband of the reconstructed image, SWT represents wavelet transform, and * is the connection of detail subbands, that is, connecting each high-frequency subband in the image channel dimension;

[0117] In the second adversarial loss, SWT(G(x)) * represents the concatenation of the detail subbands (LH, HL, and HH) of the reconstructed image G(x). The goal of the generator is to deceive the discriminator into thinking that the generated image is real. In adversarial network training, the output of the discriminator D is the probability that the input image is real, which usually ranges from 0 to 1.

[0118] When the generator is well behaved, it expects D(SWT(G(x)) * )) is close to 1, because when the input of the log function is close to 1, the value of log(1) is 0; so when D(SWT(G(x)) * )) approaches 1, -E[log(D(SWT(G(x)) * ))] approaches 0, indicating that the generator loss is minimal.

[0119] For example, suppose D(SWT(G(x)) * ))=0.9, then log(D(SWT(G(x)) * ))=log(0.9)≈-0.105,-log(D(SWT(G(x)) * ))≈0.105; when D(SWT(G(x)) * )) approaches 1, log(D(SWT(G(x)) * )) approaches 0, -log(D(SWT(G(x)) * )) also approaches 0, which means that the generator successfully deceives the discriminator and its second adversarial loss approaches 0.

[0120] Furthermore, the second perceptual loss is expressed as:

[0121] L perc =E[DISTS(G(x),y)];

[0122] Among them, DISTS() is an image structure and texture similarity evaluation model.

[0123] Preferably, in this embodiment, the discriminator uses the discriminator loss constructed by the high-frequency sub-band to perform iterative optimization, and the discriminator and the CT image reconstruction unit generate an adversarial architecture for training the CT image reconstruction unit;

[0124] The discriminant loss function is expressed as:

[0125] L adv,D = -E[logD(SWT(y) * )]-E[log(1-D(SWT(G(x))* ))];

[0126] Among them, E[] represents averaging, D represents the discriminator model, G represents the generator model, SWT(y) * Represents the detail subband connection after wavelet transform of low-dose CT image, SWT(G(x)) * It represents the connection of detail subbands after wavelet transform of the reconstructed image. The detail subband connection is to connect each high-frequency subband in the image channel dimension. For example, the size of each subband is (N, C, H, W), where N is the number of images in each batch during training, C is the number of image channels, and H and W are the length and width of the image. After connecting the three detail subbands, the size becomes (N, 3C, H, W), which serves as the input of the discriminator.

[0127] In this embodiment, the discriminator is trained only on the high-frequency subbands (LH, HL, and HH) to distinguish high-frequency details in the reconstructed image from those in the corresponding low-dose CT image. The discriminator's task is to learn the "realism" of the generated high-frequency details compared to those in the low-dose CT image. During training, the generator and discriminator compete with each other, continuously optimizing to make the generator's images closer to the low-dose CT image.

[0128] For the discriminator adversarial loss part 1 - E[logD(SWT(y) * )], for the high frequency subband SWT(y) of the real image y * , the goal of the discriminator is to correctly identify it as real, that is, D(SWT(y) * ) should be close to 1, then -E[logD(SWT(y) * )] approaches 0. For example, if D(SWT(y) * )=0.99,logD(SWT(y) * )=log(0.99)≈-0.01,-logD(SWT(y) * )≈0.01; when D(SWT(y) * ) approaches 1, -logD(SWT(y) * ) approaches 0;

[0129] For the second part - E[log(1-D(SWT(g(X)) * ))], for the detail subband of the reconstructed image G(x), the discriminator hopes to judge it as false, so D(SWT(G(x)) * )) should be close to 0, then 1-D(SWT(G(x)) * )) is close to 1, then -E[log(1-D(SWT(G(x)) *))] approaches 0. For example, when D(SWT(G(x)) * ))=0.1, 1-D(SWT(G(x)) * ))=0.9, then log(1-D(SWT(G(x)) * ))=log(0.9)≈-0.105,-log(1-D(SWT(G(x)) * ))=log(0.9)≈0.105;

[0130] When D(SWT(G(x)) * )) approaches 0, 1-D(SWT(G(x)) * )) approaches 1, then -log(1-D(SWT(G(x)) * )) approaches 0. That is, the goal of the discriminator is to make L adv,D Minimize the overall value, that is, try to judge the real image as real and the reconstructed image as false, so that both parts are close to 0.

[0131] Furthermore, the present invention uses the Adam optimization algorithm to train the generative adversarial architecture (i.e., the generator (CT image reconstruction unit) and the discriminator). During the training process, appropriate hyperparameters such as learning rate and momentum are set according to experience, and these parameters are adjusted according to certain strategies (such as learning rate decay) during the training process. During the training process, the performance of the model (generator) is also regularly evaluated on the validation set, using a variety of evaluation indicators such as peak signal-to-noise ratio (PSNR), structural similarity (SSIM), no-reference evaluation indicators (NIQE, NRQM, etc.) and lung nodule detection accuracy. Based on the evaluation results, the generator model with the best performance is selected for storage and subsequent application.

[0132] Step S3: obtaining an ultra-low-dose CT image to be processed and inputting it into a trained CT image reconstruction unit to obtain a reconstructed CT image;

[0133] Specifically, after processing the ultra-low-dose CT images using the trained CT image reconstruction unit, the resulting image quality is significantly better than the original ultra-low-dose CT images, and the image artifacts are fewer.

[0134] Figure 5 This is a chest CT image of a patient, where: Figure 5 (a) is a low-dose CT image; Figure 5 (b) is the ultra-low-dose CT image (ULDCT) processed based on the iterative reconstruction (IR) algorithm; Figure 5(c) shows the ULDCT image (ULDCT-DL) processed using the CT image reconstruction algorithm of this embodiment. Comparison shows that the quality of the ULDCT-DL image processed using the image reconstruction algorithm of this embodiment is significantly superior to that of the unprocessed ultra-low-dose ULDCT image, with fewer image artifacts, validating the effectiveness of the ultra-low-dose CT image reconstruction algorithm of this embodiment.

[0135] In summary, the image-domain-based ultra-low-dose CT image reconstruction method of the present invention utilizes a large number of paired ultra-low-dose and low-dose CT images for model training, integrates the RRDB module in the generator to implement the application of global residual connections, and uses the generator's adversarial loss, content loss, SSIM loss, and perceptual loss to better preserve the image's edge and detail information, and capture the image's overall structural features and overall visual perception features; or through wavelet domain loss optimization, the reconstructed image has fewer artifacts and higher clarity than the original ultra-low-dose CT image; and the CT image reconstruction unit of the present invention removes the upsampling layer in the traditional model, avoids artifacts and retains the original information, speeds up image processing, and provides strong technical support for the large-scale application of ULDCT lung cancer screening; it has been verified in specific cases and helps doctors observe lung tissue structure and pathological characteristics more accurately. It significantly improves the detection rate of lung nodules, especially for non-solid and small-sized nodules, reducing missed diagnoses.

[0136] Those skilled in the art will appreciate that all or part of the process steps of the methods in the above embodiments can be implemented by instructing relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium, such as a magnetic disk, an optical disk, a read-only memory, or a random access memory.

[0137] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by any technician familiar with this technical field within the technical scope disclosed by the present invention should be covered by the scope of protection of the present invention.

Claims

1. An ultra-low-dose CT image reconstruction method based on image domain end, characterized in that: include: Acquire high- and low-dose paired CT images of multiple patients at the image domain end, wherein the high- and low-dose paired CT images include ultra-low-dose CT images and low-dose CT images of the same part of the corresponding patients, and construct a training sample set using the ultra-low-dose CT images as input samples and the low-dose CT images as labels; A CT image reconstruction unit and a model training unit are preliminarily constructed, and the preliminarily constructed CT image reconstruction unit is trained based on the training sample set and the model training unit; the model training unit is used to receive a reconstructed image output by the CT image reconstruction unit and a corresponding low-dose CT image in the training sample set, perform a wavelet transform on the reconstructed image and the corresponding low-dose CT image, respectively, and iteratively train the CT image reconstruction unit based on the high-frequency subband and low-frequency subband obtained by the wavelet transform by the following method: input the ultra-low-dose CT image in the training sample set into the CT image reconstruction unit to obtain a reconstructed image; Based on the multiple high-frequency sub-bands corresponding to the reconstructed image and the low-dose CT image, using a discriminator to distinguish true from false on the reconstructed image, and generating a discrimination result; Based on the low-frequency subband and high-frequency subband corresponding to the reconstructed image and the low-dose CT image and the discrimination result of the discriminator, the CT image reconstruction unit is iteratively updated using the generator composite loss and the discriminator loss fused by multiple loss functions to obtain a converged CT image reconstruction unit; The ultra-low-dose CT image to be processed is obtained and input into the trained CT image reconstruction unit to obtain a reconstructed CT image.

2. The image domain-based ultra-low-dose CT image reconstruction method according to claim 1, characterized in that: The CT image reconstruction unit is constructed based on the SRResNet model, and includes a feature extraction layer, a residual density block module, and a convolution module; The feature extraction layer extracts features from the input ultra-low-dose CT image through a convolution operation to obtain a primary feature map; The residual density block module extracts a high-level feature map based on the primary feature map by fusing a multi-level residual network with dense connections; The convolution module is used to fuse the high-level feature maps and adjust the dimensions to obtain a reconstructed image.

3. The image domain-based ultra-low-dose CT image reconstruction method according to claim 2, characterized in that: The residual density block module includes multiple RRDB modules arranged in sequence, each RRDB module is composed of five residual blocks arranged in sequence, and each residual block also contains a sub-residual block; each sub-residual block is composed of multiple densely connected 3×3 convolutional layers, the convolution step size is 1, and LeakyReLU is used as the activation function.

4. The image domain-based ultra-low-dose CT image reconstruction method according to claim 3, characterized in that: The discriminator is a PatchGAN architecture; it includes n convolutional layers arranged in sequence; Among them, the 2nd to nth convolutional layers also include batch normalization layers; the 1st to n-1th convolutional layers use the LeakyReLU activation function, and the nth convolutional layer uses the Sigmoid activation function to output the discrimination result of each patch.

5. The image domain-based ultra-low-dose CT image reconstruction method according to claim 4, characterized in that: The generator composite loss includes the first adversarial loss, content loss, SSIM loss and first perceptual loss; the generator composite loss is expressed as: L G =αL Gadv +β·L content +γ·L SSIM +δL perceptual ; Among them, L Gadv is the first adversarial loss, L content is the content loss, L SSIM is the SSIM loss; L perceptual is the first perceptual loss, and α, β, γ, and δ are weight coefficients.

6. The image domain-based ultra-low-dose CT image reconstruction method according to claim 5, characterized in that: The first adversarial loss is expressed as: L Gadv =-E[log(D(G(x)))]; Where E represents averaging, G(x) is the reconstructed image, and D(G(x)) is the discrimination result of the reconstructed image; The content loss is expressed as: Among them, G(x) i is the i-th pixel value of the reconstructed image, y i is the ith pixel value of the real low-dose CT image, and N is the total number of image pixels; The SSIM loss is expressed as: L SSIM =1-SSIM(G(x),y); Where G(x) is the reconstructed image and y is the real low-dose CT image; The first perceptual loss is expressed as: L perceptual =E[DISTS(G(x),y)]; Among them, DISTS() is the image structure and texture similarity evaluation model; The discriminator is iteratively optimized using the following adversarial loss: L D =-E[logD(y)]-E[log(1-D(G(x)))]; Where E[] represents averaging, D represents the discriminator model, G represents the generator model, y represents the real LDCT image, and G(*) represents the output of the generator.

7. The image domain-based ultra-low-dose CT image reconstruction method according to claim 1, characterized in that: The discriminator uses the discriminator loss constructed by the high-frequency sub-band to perform iterative optimization, and the discriminator and the CT image reconstruction unit generate an adversarial architecture for training the CT image reconstruction unit; The discriminator loss is expressed as: L D =-E[logD(SWT(y)) * ]-E[log(1-D(SWT(G(x)) * ))]; Where E[] represents averaging, x is an ultra-low-dose CT image, G(x) is a reconstructed image generated by a CT image reconstruction unit, y is a low-dose CT image, and SWT(y) is * represents the detail subband connection after wavelet transform of low-dose CT image, D represents the discriminator model, SWT(G(x)) * It represents the connection of detail subbands after wavelet transformation of the reconstructed image, that is, connecting each high-frequency subband in the image channel dimension, and G represents the generator model.

8. The image domain-based ultra-low-dose CT image reconstruction method according to claim 1, characterized in that: Based on the low-frequency subband and high-frequency subband corresponding to the reconstructed image and the low-dose CT image and the discrimination result of the discriminator, the CT image reconstruction unit is iteratively updated by a generator composite loss that fuses the wavelet domain loss, the second adversarial loss, and the second perceptual loss; its loss function is expressed as: L G =L SWT +λ adv ·L adv,G +λ perc ·L perc ; Among them, L SWT is the wavelet domain loss, L adv,G is the second adversarial loss, L perc is the second perceptual loss; perc and λ adv is the weight coefficient.

9. The image domain-based ultra-low-dose CT image reconstruction method according to claim 8, characterized in that: The wavelet domain loss is expressed as: L SWT =E[∑ j l j (α1‖SWT(G(x)) j -SWT(y) j ‖1+(1-α1)(1-SSIM(SWT(G(x)) j ,SWT(y) j )))]; Where E[] represents averaging, x is the ultra-low-dose CT image, G(x) is the reconstructed image generated by the CT image reconstruction unit, y is the low-dose image, α1 is the weight, and SWT(G(x)) j represents the jth high-frequency subband and low-frequency subband obtained by performing stationary wavelet transform on the reconstructed image, SWT(y) j represents the jth high-frequency subband or low-frequency subband obtained after the stationary wavelet transform of the low-dose CT image; λ is the scaling factor, and SSIM() is the structural similarity index between the reconstructed image and the low-dose CT image after wavelet transform; The second adversarial loss is expressed as: L adv,G =-E[log(D(SWT(G(x))*))]; Where G(x) is the reconstructed image, D(SWT(G(x))*)) is the discrimination result of the detail subband of the reconstructed image, SWT represents wavelet transform, and * is the connection of detail subbands, that is, connecting each high-frequency subband in the image channel dimension; The second perceptual loss is expressed as: L perc =E[DISTS(G(x),y)]; Among them, DISTS() is an image structure and texture similarity evaluation model.

Citation Information

Patent Citations

  • Multi-band low-resolution image synchronous fusion method based on dense network and local brightness traversal operator

    CN113112441A