A Deep Learning-Based Super-Resolution Method for CT Images
By constructing a multi-domain residual self-calibration super-resolution network with projection domain physical consistency loss and multi-scale feature residual calibration loss, the problem of visual clarity but physical inconsistency in existing CT image super-resolution methods is solved, achieving the unification of physical consistency and visual clarity of high-resolution CT images, which is suitable for clinical diagnosis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHAZHOU PROFESSIONAL INST OF TECH
- Filing Date
- 2026-02-05
- Publication Date
- 2026-06-02
AI Technical Summary
Existing deep learning-based CT image super-resolution methods have failed to effectively integrate the physical laws of CT imaging, resulting in visually clear reconstructed images that are inconsistent with the actual projection measurement data, and even artifact problems.
We construct a projection domain physical consistency loss and a multi-scale feature residual calibration loss. Through a multi-domain residual self-calibration super-resolution network, combined with differentiable ray-driven forward projection operation and a multi-stage progressive training algorithm, we ensure that the reconstruction result conforms to the physical constraints of the original projection data.
It achieves a balance between visual clarity and physical realism, with reconstruction results consistent with real scan data, reducing data costs, improving image detail richness and structural accuracy, adapting to different scanning protocols and scenarios, and meeting clinical diagnostic requirements.
Smart Images

Figure CN122134556A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of super-resolution industrial CT image generation technology, and in particular to a deep learning-based super-resolution CT image method. Background Technology
[0002] Existing CT image super-resolution technology is mainly based on deep learning-based CT super-resolution methods. With its powerful feature learning and mapping capabilities, it surpasses traditional methods in image detail restoration and has become the mainstream research direction of current CT super-resolution technology. However, in actual clinical and industrial applications, existing deep learning CT super-resolution methods only rely on image domain data for modeling, completely ignoring the physical origin of the CT imaging system—projection data, and lacking the integration of the physical laws of CT imaging.
[0003] Unlike ordinary natural images, CT images are not directly acquired. Instead, they are reconstructed from projection data (sine waves) collected by detectors through physical imaging models such as Radon transform and filtered back projection. Projection data is the physical measurement source of CT images. However, existing methods treat CT images as ordinary two-dimensional images, directly transferring the network architecture and training strategies of super-resolution natural images without incorporating prior physical knowledge of CT imaging into the deep learning framework. This may result in the network failing to learn the physical laws of CT imaging. The reconstructed high-resolution images often only guarantee visual clarity but deviate from the actually acquired projection measurement data, resulting in visual clarity but physical inconsistencies. It may even lead to artifacts that violate the physical laws of CT imaging.
[0004] In summary, the technical problems existing in the relevant technologies need to be improved. Summary of the Invention
[0005] The purpose of this invention is to provide a deep learning-based super-resolution method for CT images. This invention constructs a projection domain physical consistency loss and a multi-scale feature residual calibration loss to force the network reconstruction results to conform to the physical constraints of the original projection data. This ensures visual clarity while strictly guaranteeing the physical authenticity and consistency of the results, effectively solving the problems of deviation between reconstructed images and physical measurement data, and the generation of artifacts that violate imaging principles in existing technologies.
[0006] To achieve the above objectives, this application proposes a deep learning-based super-resolution method for CT images, which specifically includes: Acquire the original CT image of the same scanned object and the original projection data corresponding to the original CT image; The original CT images are subjected to degradation simulation and data augmentation processing based on a degradation process simulation model to obtain an enhanced training sample set of original CT images; Based on the super-resolution requirements of the original CT images, a multi-domain residual self-calibration super-resolution network is constructed. The multi-domain residual self-calibration super-resolution network is trained using a multi-stage progressive training algorithm based on the enhanced original CT image training sample set and the original projection data corresponding to the original CT image, to obtain the trained super-resolution network. Based on the trained super-resolution network, high-resolution CT images are output.
[0007] In some embodiments, the degradation process simulation model is a neural network model; The neural network model simulates the degradation effect of the target scanning protocol parameters on the image based on the input target scanning protocol parameters, thereby processing the input image into a degraded image consistent with the target scanning protocol parameters; The target scanning protocol parameters include tube current, tube voltage, scanning trajectory, and detector spacing parameters.
[0008] In some embodiments, the original CT image is a set of clear CT images; The original projection data corresponding to the original CT image is a set of projection data acquired based on a low-dose scanning protocol or a sparse view scanning protocol.
[0009] In some embodiments, the spatial resolution of the set of clear CT images is greater than the resolution of the images directly reconstructed by the low-dose scanning protocol or sparse view scanning protocol.
[0010] In some embodiments, the training process of the multi-domain residual self-calibration super-resolution network includes an interactive image domain reconstruction process and a projection domain constraint process. The image domain reconstruction process includes feature extraction and reconstruction of the degraded image generated by the degradation process simulation model to generate a high-resolution CT image. The projection domain constraint process includes mapping the high-resolution CT image to the projection domain using a differentiable forward projection operation to obtain simulated projection data; The error between the simulated projection data and the original projection data corresponding to the original CT image is calculated to obtain the physical consistency loss.
[0011] In some embodiments, the differentiable forward projection operation is implemented based on a ray-driven linear integral model to support backpropagation of gradients during neural network model training.
[0012] In some embodiments, the image domain reconstruction process specifically includes: Multi-scale feature extraction is performed on the input image to obtain multi-level feature maps; Residual self-calibration is performed on the multi-level feature maps to enhance effective detail features; The effective detail features after residual self-calibration are fused and sampled to reconstruct a high-resolution image.
[0013] In some embodiments, the multi-stage progressive training algorithm includes: The first stage involves optimizing the image domain reconstruction process based on an image domain loss function. In the second stage, training is performed in conjunction with the projection domain constraint process, based on the relatively stable parameters of the image domain reconstruction process. In the third stage, the overall network is optimized end-to-end based on the weighted sum of the image domain reconstruction loss and the projection domain physical consistency loss as the total loss function.
[0014] In some embodiments, the third stage includes a residual calibration step based on the scanning protocol, specifically including: The simulated projection data is subjected to learnable protocol degradation processing to obtain degraded simulated data; The degraded simulated data and the original projected data are compared in a multi-scale feature space using residuals. The residual signal generated by the comparison is used as an additional projection domain residual calibration loss, which is fed back and used to optimize the image domain reconstruction process.
[0015] In some embodiments, the weighting coefficients of the image domain reconstruction loss and the projection domain physical consistency loss in the total loss function are dynamically adjusted according to the number of training iterations or network performance metrics.
[0016] Compared with the prior art, this application includes at least the following beneficial effects: 1. This invention embeds differentiable ray-driven forward projection as a network layer and constructs projection domain physical consistency loss and multi-scale feature residual calibration loss. It forces the reconstruction results of the multi-domain residual self-calibration super-resolution network to conform to the physical constraints of the original projection data, thus achieving an inherent unity of visual clarity and physical reality. The reconstruction results have solid physical measurement support, thereby ensuring visual clarity while strictly guaranteeing the physical authenticity and consistency of the results. This effectively solves the problems of deviation between reconstructed images and physical measurement data and the generation of artifacts that violate imaging rules in the prior art.
[0017] 2. This invention constructs a physically parameterized degradation process simulation model. This degradation process simulation model generates simulated low-resolution images with highly consistent physical properties through learnable mapping, which are then expanded into a training set. This solves the problem of the scarcity of training data for CT super-resolution and reduces data costs.
[0018] 3. This invention adopts an improved U-Net architecture with an integrated residual self-calibration module in the image domain. Through adaptive comparison and calibration of grouped features, it can effectively enhance key diagnostic features such as blood vessels and small lesions, while suppressing noise and artifacts, improving the detail richness and structural accuracy of the reconstructed image, which is more conducive to clinical diagnosis.
[0019] 4. This invention designs a three-stage progressive training algorithm and a dynamic weighted loss adjustment mechanism. The adjustment mechanism has an optimization logic from simple to complex and intelligently balances different training objectives, ensuring stable convergence of complex model training and enhancing the model's generalization adaptability to different scanning protocols and scenarios.
[0020] 5. This invention designs a standardized CT value processing flow, including precise linear normalization and reversible inverse normalization. This flow ensures the consistency of network input and output data, enabling the final reconstructed image to have a clear physical meaning of tissue density, which can be directly connected to the clinical diagnostic system, meeting the standardization and practicality requirements of clinical applications. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments or examples of the present invention, the drawings used in the embodiments or examples will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained according to these drawings without creative effort.
[0022] Figure 1 This is a flowchart illustrating a deep learning-based CT image super-resolution method according to the present invention. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.
[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0025] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0026] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0027] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0028] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0029] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0030] Existing CT image super-resolution technologies are mainly divided into two categories: traditional methods and deep learning-based methods. Traditional methods are mostly based on interpolation algorithms (such as bicubic interpolation and bilinear interpolation) or traditional iterative reconstruction algorithms. These methods rely solely on simple pixel value fitting or basic iterative optimization to restore image details. When dealing with complex anatomical structures and delicate industrial components, they suffer from insufficient detail restoration capabilities, image blurring, and artifacts, making it difficult to meet the high-precision requirements of clinical diagnosis and industrial inspection. With the development of deep learning technology, deep learning-based CT super-resolution methods, with their powerful feature learning and mapping capabilities, significantly surpass traditional methods in image detail restoration, becoming the mainstream research direction in current CT super-resolution technology. However, in actual clinical and industrial applications, they still have discrepancies with the actually acquired projection measurement data, resulting in visually clear but physically inconsistent images, and even the technical problem of reconstructing artifacts that violate the physical laws of CT imaging.
[0031] For example, in clinical chest CT super-resolution diagnostic scenarios, for low-resolution CT images used for screening small pulmonary nodules, existing deep learning-based super-resolution methods can reconstruct visually clear and well-defined lung structures, even revealing visible "microvascular branches" and "nodule edge details." From a visual perspective, the reconstruction effect is superior to traditional interpolation methods, seemingly meeting the needs of early screening for small nodules. However, when verified against the physical origin of CT imaging—raw projection data (PRAW)—it is found that the "microvascular branches" presented in the reconstructed image do not have corresponding projection signals in the actual acquired projection data. The core physical logic of CT imaging is that projection data is a record of the intensity attenuation after X-rays penetrate human tissue. Any anatomical structure in the reconstructed image must have a corresponding attenuation signal in the projection data. However, existing methods, because they do not incorporate the physical constraints of the projection domain, merely "fictionalize" visually reasonable but physically non-existent vascular structures through feature fitting in the image domain. This leads to a serious deviation between the reconstructed image and the actual projection measurement data, i.e., "visually clear but physically inconsistent."
[0032] To address the issues in existing CT image processing, such as discrepancies between the image and the actual projected measurement data, resulting in visually clear images that are physically inconsistent, and even artifacts that violate the physical laws of CT imaging, this application provides a deep learning-based CT image super-resolution method. The core technical features include the following methods (see appendix). Figure 1 ): Acquire the original CT image of the same scanned object and the original projection data corresponding to the original CT image; The original CT images are subjected to degradation simulation and data augmentation processing based on a degradation process simulation model to obtain an enhanced training sample set of original CT images; Based on the super-resolution requirements of the original CT images, a multi-domain residual self-calibration super-resolution network is constructed. The multi-domain residual self-calibration super-resolution network is trained using a multi-stage progressive training algorithm based on the enhanced original CT image training sample set and the original projection data corresponding to the original CT images, resulting in a well-trained super-resolution network. Based on the trained super-resolution network, high-resolution CT images are output.
[0033] In this context, the same scan subject refers to the same patient (clinical scenario), and for the same patient undergoing a chest CT scan, the scan interval should not exceed 10 minutes (to avoid data mismatch caused by organ movement); the original CT image (I RAW This refers to clear CT images acquired and reconstructed based on standard scanning protocols; the following three sets of paired data are obtained from the same chest CT scan patient: the original CT image (I RAW ): As a high-resolution reference image, it was acquired and reconstructed using a standard scanning protocol. Original projection data (P... RAW As a source of physical constraints, low-dose or sparse view protocols are used for acquisition; this embodiment uses a sparse view protocol. This also includes the low-resolution CT image to be processed (I...). LR ): As input to network inference, it consists of the aforementioned original projection data (P) RAW The image was reconstructed using the Filtered Back Projection (FBP) algorithm.
[0034] Furthermore, the above data is standardized to ensure consistency; Among them, I RAW I LR The HU value is mapped to the [0,1] interval, providing standardized input normalized pixel values for network training. The formula for the normalized pixel value is: ; In the formula, I norm (x,y) represents the normalized pixel values; HU min HU is the preset normalization lower limit (-1000). max The preset normalization upper limit is (400); (x,y) are the pixel space coordinates of the CT image; I HU (x,y) represents the original HU value of the pixel; This method compresses the HU values of the original CT images, which range from thousands (e.g., -1000 to 400), to the interval of 0, 10, and 1. This improves the numerical stability of subsequent training and accelerates the convergence of the degradation process simulation model. It provides the network with a scale-uniform input, eliminating the influence of differences in absolute HU values between different scanning protocols and devices on the degradation process simulation model, allowing the network to focus on learning general image structure and texture features. Based on the set HU... min-1000, the formula implicitly normalizes the air background to a value close to 0, which helps the model distinguish diagnostically significant tissue regions.
[0035] Furthermore, the normalized high-resolution image output by the network is restored to clinically usable inversely normalized HU value pixels, where the formula for the inversely normalized HU value pixels is: I SR-HU (x,y)=I SR-norm (x,y)·(HU max- HU min )+HU min ; In the formula, I SR-norm (x,y) represents the normalized high-resolution image pixel values output by the network; I SR-HU (x,y) represents the pixels with inversely normalized HU values, directly used for clinical diagnosis; HU max The preset normalization upper limit is (400); HU min The preset normalization lower limit is (-1000); HU max、 HU min The value must be strictly consistent with the normalization parameter in the preprocessing stage; The physical meaning of the inverse normalized HU value pixel formula is that, from I... RAW with I LR The HU value is mapped to the [0,1] interval and converted back to the physical representation (HU value). After this step, the image output by the network has a clear meaning of tissue density, and doctors can make diagnoses based on the standard HU value window. The denormalized HU value pixel formula and the normalized pixel value formula constitute a precise inverse transformation pair. It ensures the data consistency of the entire processing flow (original HU → normalization → network processing → denormalization → HU restoration), so that the model enhancement results can be accurately interpreted back to the original physical scale. Through the formula calculation, I SR -HU ISR The theoretical range of -HU is mapped back to the range of [-1000, 400]HU, which is consistent with the display and diagnostic range of conventional CT images, avoiding unreasonable or undisplayable pixel values.
[0036] Among them, I RAW with I LR Rigid registration is performed using a mutual information algorithm to ensure spatial alignment. Then, the registered I... RAW Downsampled to 256×256 pixels to match I LR Alignment in size is used for subsequent degradation process simulation models. For P RAW Dark current, scattering, and detector offset corrections are performed.
[0037] Furthermore, the low-resolution CT images (I) actually obtained in clinical practice will be processed... LR The data is input into the trained "multi-domain residual self-calibration super-resolution network". The network performs forward propagation, generating a high-resolution image through the image domain reconstruction branch. Because the network is subject to strict physical constraints of the projection domain during training, its output high-resolution CT images are not only visually clearer, but also physically consistent with the patient's actual scan data, thus having greater clinical diagnostic value.
[0038] The degradation process simulation model is a neural network model; The neural network model simulates the degradation effect of the target scanning protocol parameters on the image based on the input target scanning protocol parameters, thereby processing the input image into a degraded image consistent with the target scanning protocol parameters; The target scanning protocol parameters include tube current, tube voltage, scanning trajectory, and detector spacing parameters.
[0039] Specifically, the degradation process simulation model is a 3-layer convolutional neural network (CNN). Its inputs are the downsampled original CT image and the target scanning protocol parameter vector (including tube current, tube voltage, detector spacing, and slice thickness), and its output is the simulated degradation image. The preprocessed original CT image (I... RAW ) and target protocol parameters (i.e., the acquired P) RAW The low-dose / sparse view protocol parameters used are input into a pre-trained degradation process simulation model. The degradation process simulation model, through the learned mapping relationship, "degrades" the high-resolution image into a format similar to the real low-dose / sparse view scan image (I0). LR Images that exhibit consistent characteristics in terms of noise, blur, and artifacts.
[0040] Furthermore, the simulated degraded images are subjected to random rotation (±15°), translation (±10 pixels), scaling (0.9-1.1 times), and the addition of Gaussian noise (standard deviation 0.01-0.03) to generate an "enhanced original CT image training sample set." The images in this sample set are visually similar to the real CT images. LR Similar, but more numerous and more varied.
[0041] Among them, the degradation process simulation model learns and simulates the degradation effect of the target scanning protocol on the image, thereby improving the high-resolution reference image I. RAW Processed into degraded images consistent with low-quality protocols The mathematical relationship is expressed as follows: ; A model for simulating the degradation process (neural network); For the learnable parameters of the model; IRAW The input is a high-resolution reference image; p is the target scanning protocol parameter vector; This is a simulated low-resolution image for output. Furthermore, the degradation process simulation model is the core of this embodiment for solving the training data problem. The degradation process simulation model is based on a learnable mapping. This transforms the generation process of elusive "real low-resolution images" into a controllable image transformation problem based on physical protocol parameters p. This allows for the generation of a large number of well-matched "low-high" training pairs requiring only a high-quality image library and protocol parameters. Furthermore, the degradation process simulation model is protocol-aware. Its input includes not only image content I... RAW This also includes explicit physical parameters, such as p. This allows the degradation process simulation model to learn specific degradation patterns coupled with protocol parameters, such as "decreased tube current primarily introduces noise" and "increased detector spacing leads to decreased spatial resolution," thereby generating more physically realistic and diverse degradation samples. Furthermore, the degradation process simulation model parameters... It is based on pre-trained learning (using a portion of real I) RAW and I LR (Paired data). This enables it to capture more complex and realistic degradation mixtures (such as combinations of noise, blur, and artifacts) than traditional fixed filter kernels (such as Gaussian blur), and it has a certain interpolation and extrapolation capability for unseen protocol parameters, enhancing the generalization of the entire super-resolution system.
[0042] Formulas for the degradation process simulation model This paper summarizes a data-driven, physically parameterized image degradation generator. It combines prior physical knowledge (protocol parameters) with data-driven learning, solving a key data challenge faced by deep learning in CT super-resolution, and laying a solid foundation for building a powerful and general-purpose super-resolution system. Its output I... LRsim It is a key bridge connecting high-quality data with the low-quality physical world.
[0043] The original CT images are a set of clear CT images; In this implementation, the original projection data corresponding to the original CT image is a set of projection data acquired based on a low-dose scanning protocol or a sparse view scanning protocol.
[0044] In this implementation, the spatial resolution of a set of clear CT images is greater than the resolution of images directly reconstructed by low-dose scanning protocols or sparse view scanning protocols.
[0045] In this implementation, the training process of the multi-domain residual self-calibration super-resolution network includes an interactive image domain reconstruction process and a projection domain constraint process. The image domain reconstruction process includes feature extraction and reconstruction of the degraded image generated by the degradation process simulation model to generate a high-resolution CT image. The projection domain constraint process involves mapping a high-resolution CT image to the projection domain using a differentiable forward projection operation to obtain simulated projection data. The error between the simulated projection data and the original projection data corresponding to the original CT image is calculated to obtain the physical consistency loss.
[0046] In this implementation, the differentiable forward projection operation is implemented based on a ray-driven linear integral model to support the backpropagation of gradients during neural network model training.
[0047] In this implementation, the image domain reconstruction process specifically includes: Multi-scale feature extraction is performed on the input image to obtain multi-level feature maps; Residual self-calibration is performed on multi-level feature maps to enhance effective detail features; The effective detail features after residual self-calibration are fused and sampled to reconstruct a high-resolution image.
[0048] Furthermore, the image domain reconstruction process specifically includes: employing an improved U-Net architecture as the image domain reconstruction branch. This includes an encoder (for multi-scale feature extraction), a decoder (for feature fusion and upsampling), and a multi-domain residual self-calibrating super-resolution network connected to the deepest layer of the encoder. The multi-domain residual self-calibrating super-resolution network utilizes the core feature maps extracted by the encoder, splitting them into multiple groups. Feature comparison and adaptive calibration are performed within and between each group to enhance effective detail features such as blood vessels and minute lesions, while suppressing invalid noise. The calibrated features are then fed into the decoder for upsampling, ultimately reconstructing a high-resolution image.
[0049] Specifically, the differentiable forward projection operation is the key physical model embedding layer. The high-resolution image output from the image domain reconstruction branch is used to simulate the physical process of a CT scan based on a ray-driven linear integral model, calculating the resulting projection data, i.e., "simulated projection data (P)". sim The ray-driven linear integral physics model embedding layer is implemented using numerical differentiation to ensure that the gradient can propagate backward. The physics model embedding layer simulates the forward projection physics process of a CT scan, integrating the image domain I... SR Mapping to the projection domain to generate P sim It supports gradient backpropagation, enabling end-to-end network training under physical consistency constraints. The formula for the ray-driven linear integral model's dispersive function is as follows: P sim (θ,t)=∑ s∈S(θ,t) I SR(x(θ,t,s),y(θ,t,s))·Δs; In the formula, P sim (θ,t) represents the simulated projection data; θ is the projection viewpoint; t is the detector channel index; S(θ,t) is the set of ray sampling points; s is the path sampling point index; x(θ,t,s) and y(θ,t,s) are the image coordinates of the sampling points; I SR (x,y) represents the pixel values of the high-resolution image; Δs is the ray sampling step size; ∑ is the discrete summation operation; This ray-driven linear integral model discretization formula is a concrete implementation of the physical principle of forward projection in CT imaging (Radon transform) within a neural network. It tightly integrates deep learning with domain physics knowledge. Based on the continuous integral discretization into a weighted sum of a finite number of sampling points, and the image pixel values I... SR The acquisition of (x,y) is designed using differentiable bilinear interpolation (taking values from the image grid based on the coordinates (x,y)), ensuring that the entire computational graph is smooth and differentiable. This allows the gradient of the physical consistency loss to propagate smoothly back to optimize the generation of I. SR The image domain reconstruction network is constructed by building the projection domain physical consistency loss L. proj The prerequisite step; P calculated using this formula sim Will be compared with the actual collected P RAW The differences between these measurements directly measure whether the network reconstruction results conform to real physical measurements, thus constituting a strong physical constraint on network training. This ray-driven linear integral model's discretization formula employs a discretized engineering implementation, specifying the method for converting from continuous physics to computable code (i.e., "ray-driven" and "sample-summation"), enabling those skilled in the art to directly program and implement it based on given geometric parameters (ray source-detector distance, etc.) and step size Δs.
[0050] Furthermore, calculate P sim Compared with the actual original projection data (P) RAW The difference between the images (such as L2 loss) is used as the "physical consistency loss". This loss directly measures whether the image reconstructed by the network matches the actual physical measurement data.
[0051] In this implementation, the multi-stage progressive training algorithm includes: The first stage involves optimizing the image domain reconstruction process based on the image domain loss function. In the second stage, training is carried out in conjunction with the projection domain constraint process, based on the relatively stable parameters of the image domain reconstruction process. In the third stage, the overall network is optimized end-to-end based on the weighted sum of the image domain reconstruction loss and the projection domain physical consistency loss as the total loss function.
[0052] In this implementation, the third stage includes a residual calibration step based on the scanning protocol, specifically including: Learnable protocol degradation is applied to the simulated projection data to obtain degraded simulated data; The degraded simulated data and the original projected data are compared in a multi-scale feature space using residuals. The residual signal generated by the comparison is used as an additional projection domain residual calibration loss, which is fed back and used to optimize the image domain reconstruction process.
[0053] The specific implementation steps of the multi-stage progressive training algorithm include the following: The first stage (image domain pre-training): The network's image domain reconstruction process (i.e., the U-Net part) is trained using only image domain losses (such as a weighted sum of L1 loss and perceptual loss), allowing the network to initially learn to recover structures from low-quality images. At this stage, the projection domain-related parts are frozen. The second stage (introducing projection domain constraints): The projection domain constraint process is unfrozen. During forward propagation, the output of the image domain branch is passed through a differentiable forward projection layer to generate P. sim And calculate with P RAW The physical consistency loss is then addressed. This stage jointly optimizes the network by combining image domain loss and physical consistency loss, enabling the reconstruction results to balance visual quality and physical realism. The third stage (end-to-end joint optimization and residual calibration) introduces a residual calibration step based on the scanning protocol. First, the P generated in the second stage is... sim The degraded simulation data (P) is obtained by processing the data through a learnable protocol degradation process simulation model. simdeg This degradation process simulation model is used to further simulate the degradation effects introduced in the projection domain by specific scanning protocols (such as low dose). Then, in the multi-scale feature space, P... simdeg With the real P RAW Residual comparisons are performed to generate a more refined "physical consistency error signal." Finally, this error signal is used as an additional projection domain residual calibration loss, fed back and used to optimize the image domain reconstruction process, particularly adjusting the weights of its residual self-calibration, thereby achieving fine-tuning of image detail reconstruction.
[0054] Among them, the image domain reconstruction loss L img , is a measure of the high-resolution image output by the network. SR Compared with reference high-resolution image I RAW The composite loss function is the difference between the two. It consists of two parts: norm loss L1 and perceptual loss L... percepThe loss function is balanced using a weighting coefficient α. This loss function is the sole optimization objective in the first stage of network training (image domain pre-training), aiming to enable the network to initially learn to recover structurally accurate and detail-rich high-quality images from low-quality images. The image domain reconstruction loss L... img The formula is: L img =α·L1+(1-α)·L percep ; In the formula, L img α is the total loss for image domain reconstruction; α is the weighting coefficient of the L1 loss; L1 is the mean absolute error; L percep To perceive loss; Image domain reconstruction loss L img The formula combines low-level pixel precision loss and high-level perceptual similarity loss to guide the network to generate high-resolution images that are both structurally accurate and visually natural during the pre-training stage, providing a high-quality initial solution for subsequent integration of physical constraints and the final high-fidelity reconstruction.
[0055] Among them, the physical consistency loss of the projection domain L proj The formula for P is based on the L2 norm loss, which measures the P-value. sim With P RAW To mitigate the differences in projection domains and ensure that the reconstructed image conforms to the physical imaging principles of CT, the physical consistency loss of the projection domain is L. proj The formula is: ; In the formula, L proj This represents the loss of physical consistency in the projection domain. θ is the set of projection viewpoints; θ is a single projection viewpoint; T is the set of detector channels; t is the index of a single detector channel; P sim (θ,t) represents the simulated projection data values; P RAW (θ,t) represents the original projected data values; ∑ θ∈Θ ∑ t∈T This is a double summation operation; Among them, the image domain reconstruction loss L img The formula is a mathematical expression of the data consistency condition in CT imaging. Loss L proj The value of I, obtained through a differentiable forward projection layer, depends entirely on the output I of the image domain network. SR Therefore, by minimizing L proj The gradient propagates back and guides the image domain reconstruction network to adjust its parameters. This loss is activated and used in the second and third phases of training. In the second phase, it is used in conjunction with the image domain loss L. img The system then begins to incorporate physical constraints into networks that already possess basic reconstruction capabilities. In the third phase, the total loss L... total Part of, with Limg and residual calibration loss L calib We jointly perform end-to-end optimization to ensure the network finds the optimal balance under complex multi-objective conditions. Image domain reconstruction loss L img The formula seamlessly integrates prior knowledge of CT physics into the deep learning framework. Image domain reconstruction loss L... img The formula is not an independent metric, but rather the engine that drives the entire network to optimize towards physical reality. By forcing the network to reconstruct images that meet the consistency of the projection data, it ensures the generation of clinical CT images that are both visually clear and physically reliable; this projection domain physical consistency loss L proj The formula quantifies the overall consistency between the simulated projection and the measured projection of the reconstructed image by averaging the angle and time points, and is used to optimize image reconstruction.
[0056] Among them, based on the additional projection domain residual calibration loss introduced in the third stage, based on P sim-deg With P RAW Multi-scale feature space residual calculation, where the projection domain residual calibration loss L calib The formula is: ; In the formula, L calib The residual calibration loss is for the projection domain; K is the total number of multi-scale values (e.g., K=2). For scale index (e.g.) =1 represents the first scale); represents the weighting coefficient for the k-th scale; (·) represents the feature extraction mapping at the k-th scale; P sim-deg For degraded analog projection data; P RAW For the original projection data; C k W k H k Let be the size parameter of the feature map; Normalization factor; For norm operations, the sum of the absolute values of the differences between corresponding elements in the feature map is calculated, where... and Multiplication means that the feature differences at the k-th scale are weighted according to the strategic importance set by humans, and the resulting final supervision signal is contributed. Multi-scale summation; Among them, C k W k H k The specific meanings are as follows: C kThe physical meaning of c is the number of different "feature filters" or "feature detectors" learned by that layer. For example, in the lower layers (where k is small), channels may correspond to simple edges or texture directions; in the higher layers (where k is large), channels may correspond to more complex patterns or structures. Typically, as the network deepens (k increases), c... k This will increase because the network needs to combine more low-level features to represent complex high-level features.
[0057] W k The physical meaning is the size of the feature map in the horizontal direction (usually corresponding to the "detector unit" dimension of the projected data).
[0058] Same as height, W k It usually decreases as the network deepens.
[0059] H k The physical meaning of H is the size of the feature map in the vertical direction (usually corresponding to the "angle" dimension of the projected data). Due to the presence of pooling layers or strided convolutions in CNNs, H... k It typically decreases as the network deepens (downsampling). When k=1, it may be close to the height of the original projected data, but when k=K, it is much smaller.
[0060] The above formula for projection domain residual calibration loss defines the projection domain residual calibration loss L. calib , is a high-order physical consistency constraint introduced in the third stage of network training for fine-tuning. calib The formula does not directly compare the original projected data values, but instead calculates the protocol-degraded simulated projected data P in a multi-scale feature space. sim-deg Compared with the true original projection data P RAW The difference between them. This loss generates a finer error signal by capturing feature residuals at different scales, which is used to feed back and optimize the image domain reconstruction process, especially residual self-calibration.
[0061] In this embodiment, the total loss function for the entire third stage is a weighted sum of the image domain reconstruction loss, the projection domain physical consistency loss, and the aforementioned projection domain residual calibration loss. The weighting coefficients for these losses are not fixed but are determined based on the number of training iterations or the network's performance metrics on the validation set (such as P). SNR S SIM Dynamic adjustments are made to balance the optimization intensity of each objective at different training stages.
[0062] Wherein, the total loss function L total The algorithm is defined in three stages, which strictly correspond to the multi-stage progressive training algorithm, with the third stage being the core: Phase 1 (Image Domain Pre-training): Optimize only the image domain reconstruction loss. Ltotal-1 =L img ; In the formula, L total-1 This represents the total loss for the first phase; L img Image domain reconstruction loss; L total-1 =L img The total loss function L for the first stage of network training (image domain pre-training) is defined. total-1 At this stage, the network focuses solely on optimizing image quality; therefore, its total loss is equal to and only equal to the image domain reconstruction loss L. img All relevant modules in the projection domain (such as the forward projection layer and the protocol degradation model) are frozen and do not participate in training; Among them, L total-1 =L img Although the formula is simple in form, it strategically isolates the training objective, which is a key design feature of this embodiment for achieving "progressive" training. It ensures that the network has strong basic image reconstruction capabilities before encountering complex physical consistency constraints, thus making the entire multi-stage training process more robust and efficient.
[0063] Phase 2 (Introducing projection domain constraints), combining image domain and physical consistency loss: L total-2 =λ1·L img +λ2·L proj ; In the formula, L total-2 λ1 represents the total loss during the second stage of training; λ1 is the weighting coefficient for the image domain loss; L img λ is the image domain reconstruction loss; λ² is the weighting coefficient of the projection domain loss; L proj This represents the loss of physical consistency in the projection domain. L total-2 =λ1·L img +λ2·L proj The total loss function L for the second stage of network training (introducing projection domain constraints) is defined. total-2 At this stage, the network's optimization objective shifts from solely focusing on image quality to a dual objective of jointly optimizing both image quality and physical realism. Therefore, its total loss is the image domain reconstruction loss L. img Physical consistency loss L with the projection domain proj The weighted sum; in the second phase, those frozen in the first phase (such as the forward projection layer, protocol degradation model) are responsible for calculating P. sim and L proj The parameters of all network layers (i.e., differentiable forward projection layers, etc.) are unlocked, and they begin to participate in training and receive gradients for updates; Among them, L total-2 =λ1·L img+λ2·L proj The formula is a crucial link in the progressive training of this embodiment, bridging the previous and subsequent steps. By unfreezing the physical model and introducing a weighted joint loss, the network is initiated to learn the physical laws, guiding the "visually correct" images obtained in the first stage towards images that are "more credible both visually and physically," ultimately achieving high-quality, high-fidelity CT super-resolution images.
[0064] Phase 3 (End-to-end joint optimization + residual calibration), weighted sum of three losses: L total-3 =λ1·L img +λ2·L proj +λ3·L calib ; In the formula, L total-3 λ1 represents the total loss during the third stage of training; λ1 is the weighting coefficient for the image domain loss; L img λ is the image domain reconstruction loss; λ² is the weighting coefficient for the projection domain physical consistency loss; L proj λ3 represents the physical consistency loss in the projection domain; λ3 is the weighting coefficient for the residual calibration loss in the projection domain; L calib This is the calibration loss for the projection domain residuals; L total-3 The total loss function for the third stage of network training (end-to-end joint optimization and residual calibration) is defined. In this final stage, the network undergoes end-to-end optimization of all parameters, and its objective function integrates constraints from three dimensions: image quality, underlying physical consistency, and fine feature calibration. Therefore, the total loss is the image domain reconstruction loss L. img Physical consistency loss of the projection domain L proj And projection domain residual calibration loss L calib The weighted sum; L total-3 The formula represents the final form of training in this embodiment. It also considers pixel-level accuracy (L). img Physical model consistency (L) proj ) and feature-level protocol matching (L calib The three levels of supervision signals represent the most comprehensive definition of high-quality CT images. total-3 The weighting coefficients λ1, λ2, and λ3 in the formula are not fixed but dynamically adjusted based on the number of training iterations or validation set performance metrics (such as PSNR and SSIM). This allows the network to intelligently allocate attention at different stages of training; for example, focusing more on building the basic structure in the early stages and more on fine-tuning physical details in the later stages. The final loss function is based on a weighted sum of image domain reconstruction loss and projection domain physical consistency loss. This document details the mathematical integration and implementation of end-to-end optimization of the overall network, residual calibration steps, and the dynamic weighting mechanism.
[0065] Furthermore, assuming the initial weighting coefficients are λ1=0.5, λ2=0.3, and λ3=0.2, the design considerations are as follows: 1. Maintain the fundamental position of image quality (λ1=0.5): Image domain reconstruction loss L img To ensure the visual usability of the output results, even while pursuing physical realism, visual clarity should not be severely sacrificed. Assigning it the highest initial weight (0.5) ensures that the network's existing good reconstruction capabilities (from the first and second stages) at the start of joint optimization will not degrade due to the sudden introduction of excessively strong additional constraints, thus maintaining the stability of the optimization.
[0066] 2. Equilibrium fundamental physical constraints (λ²=0.3): Projection domain physical consistency loss L proj This approach is crucial for achieving "physical interpretability," but its optimization objective (projection data matching) and image visual quality objective are inherently at odds. Setting its initial weight to 0.3, lower than λ1, means that the physical consistency constraint is introduced with a "mild" strength at the beginning of the third stage. This avoids drastic and unstable adjustments to image domain features due to excessively strong physical constraints, which helps the network smoothly find a balance point among multiple objectives.
[0067] 3. Gradually introduce higher-order calibration (λ3=0.2): Projection domain residual calibration loss L calib This is a more complex and refined supervisory signal applied to the feature space. In the early stages of training, when the network's fit to fundamental physical laws is not yet perfect, assigning it too high a weight too early may lead to local optima or difficulty in convergence. Setting its initial weight to the lowest of the three (0.2) reflects the gradual approach of "from coarse to fine." It allows the network to first focus on reconciling the two main contradictions of image quality and consistency with fundamental physics, and then, after stabilizing, gradually enhance the effect of fine calibration through subsequent dynamic adjustment mechanisms.
[0068] Normalization and interpretability include setting λ1+λ2+λ3=1, which allows the total loss to be intuitively understood as a weighted average of the three sub-losses. This simplifies the understanding of the loss scale and facilitates the design of dynamic adjustment strategies (e.g., maintaining the sum at 1 during adjustment). In summary, the initial coefficients λ1=0.5, λ2=0.3, λ3=0.2 are a natural extension of the multi-stage progressive training logic in this embodiment. It ensures that the network, in the final joint optimization stage, can start from a robust and balanced starting point and, guided by the dynamic adjustment mechanism, ultimately converge to a solution that achieves optimal visual quality, fundamental physical fidelity, and fine-grained feature matching.
[0069] In this embodiment, during the residual calibration step, P needs to be calculated in the multi-scale feature space. sim-deg With P RAW The residuals, the formula for single-scale feature residuals is: R k =F k (P sim-deg )-F k (P RAW ); In the formula, R k This is the feature residual map at scale k; k is the scale index; F k (·) is the feature extraction function at the k-th scale; P sim-deg For degraded analog projection data; P RAW This is the original projection data; The single-scale characteristic residual formula defines the characteristic residual R calculated at a single scale (the k-th scale) during the projection domain residual calibration step. k It is used to construct the higher-order calibration loss L. calib The basic unit is obtained by comparing the protocol-degraded analog projection data P. sim-deg Compared with the true original projection data P RAW Differences in a specific feature space are used to capture subtle physical inconsistencies.
[0070] The single-scale characteristic residual formula means that the error calculated directly on the projected data values (such as L) is different from the error calculated on the projected data values. proj Unlike other formulas, this one measures the difference in the transformed feature space. Feature extraction network F k It can learn representations that are more meaningful for subsequent tasks, such as distinguishing different protocols and identifying artifacts. Therefore, R... k It captures differences at the semantic or pattern level, rather than simple numerical biases. One of the inputs to the single-scale feature residual formula is P. sim-deg This refers to the simulated data that has undergone the protocol degradation process simulation model. This means that R... k In fact, the comparison—"How do the feature representations of the projected data of a network-reconstructed image differ from those of a real scan after simulating the exact same hardware and protocol limitations"—makes the residual signal highly specific to a particular scanning protocol. Among these, the feature residual R at a single scale... k It is to calculate the residual calibration loss L in the projection domain. calib Intermediate results. L calib By applying R to all scales k k We obtain R by taking the L1 norm and weighted summation; therefore, R k It is the direct source for generating the calibration gradient that is ultimately used for backpropagation.
[0071] In this embodiment, R kAs a feature map itself, its spatial distribution and channel activation patterns can provide interpretable feedback. For example, regions with large residuals may indicate significant discrepancies between the details reconstructed by the network in that region and the physical measurements. This can be used for visualization analysis or to guide targeted adjustments to specific network modules (such as residual self-calibration in the image domain).
[0072] R k =F k (P sim-deg )-F k (P RAW This is the key operation for achieving fine physical calibration in this embodiment. It elevates the physical consistency check from "numerical matching" to "feature matching". By comparing the data simulated by protocol degradation with real data at different semantic scales, it generates information-rich residual signals, thereby driving the network to generate more realistic high-resolution CT images at multiple visual and physical levels.
[0073] In this implementation, the weighting coefficients of the image domain reconstruction loss and the projection domain physical consistency loss in the total loss function are dynamically adjusted according to the number of training iterations or network performance metrics.
[0074] λ1, λ2, and λ3 are dynamically adjusted based on the number of training iterations and network performance metrics (PSNR / SSIM), using a step-by-step adjustment rule. The core rule is as follows: Let the total number of training iterations be E, the current iteration number be e, and the adjustment period be E0 = 50: When e∈[0,E0]: initial coefficients λ1=0.5, λ2=0.3, λ3=0.2; When e∈(E0,2E0]: If the validation set PSNR≥34dB and SSIM≥0.90, then λ1-0.1, λ2+0.05, λ3+0.05; if the standard is not met, then the coefficients of the previous stage are maintained. When e∈(2E0,E]: If the validation set PSNR≥35dB and SSIM≥0.92, then λ1-0.1, λ2+0.05, λ3+0.05; if the standard is not met, then the coefficients of the previous stage are maintained. All stages satisfy λ1+λ2+λ3=1, ensuring that the loss weights are normalized.
[0075] If the root mean square error (RMSE) of the projection domain is within 10 consecutive iterations, then... sim ,P RAW If λ2 > 0.03 (the threshold for physical consistency in the disclosure), then λ2 will be forcibly increased by 0.1 and λ1 will be decreased by 0.1 to prioritize physical consistency.
[0076] In this embodiment, training and inference verification were performed using data from 100 patients undergoing chest CT scans: Image domain performance: The average PSNR on the test set was 35.8dB and the average SSIM was 0.93, which is a significant improvement over the input image (PSNR=28.2dB, SSIM=0.76).
[0077] Projection domain performance: The average RMSE of the simulated projection and the original projection is 0.028, which is lower than the threshold of 0.03, and the physical consistency meets the standard.
[0078] Detailed reconstruction: It can clearly reconstruct blood vessels with a diameter ≥1mm and tiny ground-glass nodules ≥3mm.
[0079] Artifact reduction: The number of artifacts is reduced by 60% compared to the input image.
[0080] This embodiment provides a method for clinical practice to stably and reliably reconstruct high-quality diagnostic images that are both visually clear and physically realistic from low-dose, low-resolution CT scans.
[0081] This embodiment does not rely on specific CT equipment or scanning sites, and can be flexibly adapted to various low-quality scanning protocols such as low dose, sparse view, and thick slice, as well as different clinical CT scanning scenarios such as chest, head, and abdomen. It can also be extended to the field of industrial CT non-destructive testing.
[0082] Another embodiment of the present invention provides a computer device, including: a memory and a processor, and a computer program stored in the memory, wherein when the computer program is executed on the processor, it implements a deep learning-based CT image super-resolution method as described in any of the above methods.
[0083] It is understandable that, such as Figure 1 The content of the embodiment of the deep learning-based CT image super-resolution method shown is applicable to this embodiment, and the specific functions implemented in this embodiment are the same as those shown below. Figure 1 The embodiment of the deep learning-based CT image super-resolution method shown is the same, and the beneficial effects achieved are the same as those described above. Figure 1 The beneficial effects achieved by the embodiment of the deep learning-based CT image super-resolution method shown are also the same.
[0084] Another embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements a deep learning-based CT image super-resolution method as described in any of the above methods.
[0085] It is understandable that, such as Figure 1 The content of the embodiment of the deep learning-based CT image super-resolution method shown is applicable to this embodiment, and the specific functions implemented in this embodiment are the same as those shown below. Figure 1The embodiment of the deep learning-based CT image super-resolution method shown is the same, and the beneficial effects achieved are the same as those described above. Figure 1 The beneficial effects achieved by the embodiment of the deep learning-based CT image super-resolution method shown are also the same.
[0086] It should be noted that the information interaction and execution process between the above systems are based on the same concept as the method embodiments of the present invention. For details on their specific functions and technical effects, please refer to the method embodiments section, which will not be repeated here.
[0087] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0088] This invention also provides a computer device, including: a memory and a processor, and a computer program stored in the memory, wherein when the computer program is executed on the processor, it implements a deep learning-based CT image super-resolution method as described in any of the above methods.
[0089] The computer device may be a desktop computer, laptop, handheld computer, or cloud server, etc. This computer device may include, but is not limited to, a processor and memory.
[0090] The processor referred to can be a central processing unit (CPU), or it can be other general-purpose processors or digital signal processors (DSPs). Application-Specific Integrated Circuit (ASIC), Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.
[0091] In some embodiments, the memory may be an internal storage unit of the computer device, such as a hard drive or RAM. In other embodiments, the memory may be an external storage device of the computer device, such as a plug-in hard drive, SmartMediaCard (SMC), SecureDigital (SD) card, or FlashCard. Furthermore, the memory may include both internal and external storage units of the computer device. The memory is used to store the operating system, applications, bootloader, data, and other programs, such as the program code of the computer program. The memory can also be used to temporarily store data that has been output or will be output.
[0092] This invention also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements a deep learning-based CT image super-resolution method as described in any of the above methods.
[0093] In this embodiment, if the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a photographing device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.
[0094] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0095] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0096] In the embodiments disclosed in this application, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0097] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
Claims
1. A deep learning-based super-resolution method for CT images, characterized in that, The method includes: Acquire the original CT image of the same scanned object and the original projection data corresponding to the original CT image; The original CT images are subjected to degradation simulation and data augmentation processing based on a degradation process simulation model to obtain an enhanced training sample set of original CT images; Based on the super-resolution requirements of the original CT images, a multi-domain residual self-calibration super-resolution network is constructed. The multi-domain residual self-calibration super-resolution network is trained using a multi-stage progressive training algorithm based on the enhanced original CT image training sample set and the original projection data corresponding to the original CT image, to obtain the trained super-resolution network. Based on the trained super-resolution network, high-resolution CT images are output.
2. The deep learning-based CT image super-resolution method according to claim 1, characterized in that, The degradation process simulation model is a neural network model; The neural network model simulates the degradation effect of the target scanning protocol parameters on the image based on the input target scanning protocol parameters, thereby processing the input image into a degraded image consistent with the target scanning protocol parameters; The target scanning protocol parameters include tube current, tube voltage, scanning trajectory, and detector spacing parameters.
3. The deep learning-based CT image super-resolution method according to claim 2, characterized in that, The original CT images are a set of clear CT images; The original projection data corresponding to the original CT image is a set of projection data acquired based on a low-dose scanning protocol or a sparse view scanning protocol.
4. The deep learning-based CT image super-resolution method according to claim 3, characterized in that, The spatial resolution of the set of clear CT images is greater than the resolution of the images directly reconstructed by the low-dose scanning protocol or sparse view scanning protocol.
5. The deep learning-based CT image super-resolution method according to claim 4, characterized in that, The training process of the multi-domain residual self-calibration super-resolution network includes an interactive image domain reconstruction process and a projection domain constraint process. The image domain reconstruction process includes feature extraction and reconstruction of the degraded image generated by the degradation process simulation model to generate a high-resolution CT image. The projection domain constraint process includes mapping the high-resolution CT image to the projection domain using a differentiable forward projection operation to obtain simulated projection data; The error between the simulated projection data and the original projection data corresponding to the original CT image is calculated to obtain the physical consistency loss.
6. The deep learning-based super-resolution method for CT images according to claim 5, characterized in that, The differentiable forward projection operation is implemented based on a ray-driven linear integral model to support the backpropagation of gradients during neural network model training.
7. The deep learning-based CT image super-resolution method according to claim 6, characterized in that, The image domain reconstruction process specifically includes: Multi-scale feature extraction is performed on the input image to obtain multi-level feature maps; Residual self-calibration is performed on the multi-level feature maps to enhance effective detail features; The effective detail features after residual self-calibration are fused and sampled to reconstruct a high-resolution image.
8. The deep learning-based super-resolution method for CT images according to claim 7, characterized in that, The multi-stage progressive training algorithm includes: The first stage involves optimizing the image domain reconstruction process based on an image domain loss function. In the second stage, training is performed in conjunction with the projection domain constraint process, based on the relatively stable parameters of the image domain reconstruction process. In the third stage, the overall network is optimized end-to-end based on the weighted sum of the image domain reconstruction loss and the projection domain physical consistency loss as the total loss function.
9. The deep learning-based CT image super-resolution method according to claim 8, characterized in that, The third stage includes a residual calibration step based on the scanning protocol, specifically including: The simulated projection data is subjected to learnable protocol degradation processing to obtain degraded simulated data; The degraded simulated data and the original projected data are compared in a multi-scale feature space using residuals. The residual signal generated by the comparison is used as an additional projection domain residual calibration loss, which is fed back and used to optimize the image domain reconstruction process.
10. A deep learning-based super-resolution method for CT images according to claim 9, characterized in that, The weighting coefficients of the image domain reconstruction loss and the projection domain physical consistency loss in the total loss function are dynamically adjusted according to the number of training iterations or network performance indicators.