Low-dose CT reconstruction method and device
Through multi-stage processing of deep learning networks in the projection domain, chord diagram domain and image domain, combined with the multi-scale Transformer module and convolutional block attention module, the noise and artifact problems in low-dose CT image reconstruction are solved, and high-quality and efficient image reconstruction is achieved.
Patent Information
- Application Number
- CN202510888206.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-06-30
AI Technical Summary
The existing low-dose CT reconstruction method has degraded image quality after reducing the radiation dose, and has problems such as increased noise, artifacts and blurring. In addition, the existing hybrid method has low reconstruction efficiency and low data quality.
The trained deep learning network is used to process in the projection domain, chord diagram domain and image domain. The projection domain denoising model, chord diagram domain denoising model and image domain denoising model are combined with the multi-scale Transformer module and the convolutional block attention module to perform intermediate supervision and fine-tuning to reconstruct low-dose CT images.
It effectively reduces noise, maintains tissue structure, improves image quality, achieves good reconstruction of low-dose CT images, and enhances reconstruction efficiency and image detail processing capabilities.
Smart Images

Figure CN120807682A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of image processing, and particularly relates to a low-dose CT reconstruction method and device. BACKGROUND
[0002] Computed Tomography (CT) is an important tool for medical diagnosis, but the high radiation dose of traditional CT can cause health risks. Low-dose CT (LDCT) reduces radiation hazards by reducing the X-ray dose, but the dose reduction leads to increased projection data noise, missing information, and degraded image quality (such as artifacts and blurring) in the reconstructed image. The core challenge of low-dose CT reconstruction is how to reconstruct a high-quality image that meets the needs of clinical diagnosis under low-dose conditions.
[0003] In related technologies, low-dose CT reconstruction mainly includes analytical methods, iterative reconstruction, deep learning methods, and hybrid methods. The analytical method includes filtered back projection (FBP), and the core of FBP is Radon inverse transform, which assumes that the projection data is noise-free and complete. In low-dose CT, photon statistical noise is significantly enhanced, and FBP direct reconstruction will amplify noise (especially high-frequency noise), resulting in grainy artifacts in the image and a significant decrease in signal-to-noise ratio (SNR). Iterative reconstruction includes statistical iterative reconstruction (SIR) and model-based iterative reconstruction (MBIR), and iterative reconstruction requires repeated execution of forward projection and back projection operations, with a calculation amount of about several times that of FBP for each iteration, resulting in high computational complexity and poor real-time performance. Deep learning methods include projection domain preprocessing, image preprocessing, end-to-end reconstruction, and hybrid models. In the study of low-dose CT image reconstruction, deep learning methods have made significant discoveries, but data-driven deep learning has certain limitations, such as data dependency and limited generalization, and lack of physical consistency. On this basis, deep learning is combined with iterative reconstruction to form a hybrid method. Traditional CT reconstruction methods can comprehensively describe the physical laws of the CT reconstruction process, and combining deep learning methods with the ability to learn complex statistical priors from a large amount of data makes the hybrid method a trend in recent years. Although the hybrid method has achieved remarkable results, existing reconstruction methods are limited by data acquisition in clinical practice, resulting in low data quality and low reconstruction efficiency.
[0004] Therefore, there is an urgent need to provide a low-dose CT reconstruction method to improve the above-mentioned deficiencies in the prior art. SUMMARY
[0005] To solve the above-mentioned problems in the prior art, the application provides a low-dose CT reconstruction method and device. The technical problem to be solved by the application is solved by the following technical scheme:
[0006] In a first aspect, the present application provides a low-dose CT reconstruction method, comprising:
[0007] obtaining a low-dose CT image;
[0008] processing the low-dose CT image in the projection domain, the chord domain and the image domain respectively by using the trained deep learning network; in the projection domain, processing the low-dose CT image to obtain updated projection data, in the chord domain, processing the updated projection data to obtain updated chord data, and in the image domain, processing the updated chord data to obtain a reconstructed low-dose CT image;
[0009] wherein the trained deep learning network is trained by using data of a preset category as a training data set, and the initial deep learning network is trained, intermediate supervision is performed in the training process, and the trained deep learning network is obtained by fine-tuning the deep learning network in the training process.
[0010] In a second aspect, the present application further provides a low-dose CT reconstruction device, comprising:
[0011] an image acquisition module configured to obtain a low-dose CT image;
[0012] an image processing module configured to process the low-dose CT image in the projection domain, the chord domain and the image domain respectively by using the trained deep learning network; in the projection domain, processing the low-dose CT image to obtain updated projection data, in the chord domain, processing the updated projection data to obtain updated chord data, and in the image domain, processing the updated chord data to obtain a reconstructed low-dose CT image;
[0013] wherein the trained deep learning network is trained by using data of a preset category as a training data set, and the initial deep learning network is trained, intermediate supervision is performed in the training process, and the trained deep learning network is obtained by fine-tuning the deep learning network in the training process.
[0014] The present application has the following beneficial effects:
[0015] The low-dose CT reconstruction method and device provided by the present application can consider the noise in the projection domain, the chord domain and the image domain, can maximize the preservation of the tissue structure in the non-artifact region, and can achieve good reconstruction of the low-dose CT image.
[0016] The present application will be further described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF DRAWINGS
[0017] Figure 1 is a flow chart of a low-dose CT reconstruction method provided by an embodiment of the present application;
[0018] Figure 2 is a schematic diagram of a sinogram domain denoising model provided by an embodiment of the present application;
[0019] Figure 3 is a schematic diagram of a first multi-scale down-sampling fusion module provided by an embodiment of the present application;
[0020] Figure 4 is a schematic diagram of a first up-sampling feature fusion module provided by an embodiment of the present application;
[0021] Figure 5 is a schematic diagram of an image domain denoising model provided by an embodiment of the present application;
[0022] Figure 6 is a flow chart of generating a low-dose CT image corresponding to a regular-dose CT image based on a CT scanning mechanism provided by an embodiment of the present application. DETAILED DESCRIPTION
[0023] The present application will be further described in detail below with specific embodiments, but the embodiments of the present application are not limited thereto.
[0024] Please refer to Figure 1 , Figure 1 is a flow chart of a low-dose CT reconstruction method provided by an embodiment of the present application, and the low-dose CT reconstruction method provided by the present application comprises:
[0025] S101, acquiring a low-dose CT image.
[0026] Specifically, in the present embodiment, the low-dose CT (Low-Dose CT, LDCT) image is a CT image obtained by reducing the X-ray tube current, i.e., reducing the emission amount of photons. The low-dose CT image increases the image noise while reducing the radiation dose, and therefore, it is necessary to reconstruct the low-dose CT image to improve the noise influence.
[0027] It should be noted that the low-dose CT image involves three key data domains, i.e., a projection domain, a sinogram domain and an image domain, in the CT imaging process. These domains correspond to different stages of CT imaging, and the data characteristics and challenges of each domain are different under low-dose conditions.
[0028] The projection domain data is the original signal directly received by the detector during the CT scanning process, that is, the attenuation value of the X-ray after passing through the object (original projection data). These data are one-dimensional or two-dimensional measurements without reconstruction. The projection domain data has the following characteristics: significant noise increase, that is, reducing the X-ray dose will lead to insufficient photon counting, and the quantum noise (Poisson noise) and electronic background noise (Gaussian noise) in the projection data will be enhanced; data sparsity, that is, some low-dose protocols (such as sparse angle scanning) will lead to incomplete projection data, and artifacts will be introduced during reconstruction.
[0029] The sinogram domain data is a two-dimensional matrix arranged according to the scanning angle and detector position (similar to a sine curve, hence the name "sine graph"). The projection values of each detector unit at different angles are organized into a row, reflecting the projection of the object at a certain angle. The sinogram domain data has the following characteristics: streak noise and artifacts, that is, under low-dose conditions, the noise in the sinogram presents a random streak pattern, which may lead to the appearance of streak artifacts in the reconstructed image domain; structure blurring, that is, noise will mask the detailed features (such as the projection of a small lesion) in the sinogram.
[0030] The image domain data is converted from the projection data or the sinogram data into a cross-sectional image (two-dimensional / three-dimensional volume data) by a reconstruction algorithm (such as filtered back projection (FBP), iterative reconstruction algorithm or deep learning reconstruction). The image domain data has the following characteristics: noise and artifacts, that is, low dose leads to mottle noise, streak artifacts and low-contrast structure blurring in the image domain; resolution degradation, that is, noise suppression may sacrifice image details (such as small calcification points or lung nodules).
[0031] S102, using the trained deep learning network to process the low-dose CT image in the projection domain, the sinogram domain and the image domain respectively; in the projection domain, the low-dose CT image is processed to obtain updated projection data, in the sinogram domain, the updated projection data is processed to obtain updated sinogram data, and in the image domain, the updated sinogram data is processed to obtain a reconstructed low-dose CT image;
[0032] The trained deep learning network uses a preset category of data as a training data set to train an initial deep learning network, and the trained deep learning network is obtained by fine-tuning the deep learning network during the training process.
[0033] Specifically, in this embodiment, the trained deep learning network includes a projection domain denoising model; the low-dose CT image is processed in the projection domain to obtain updated projection data, which includes:
[0034] The projection data generation model is constructed and represented as:
[0035] P = W + δ;
[0036] where W denotes the number of quanta received by the detector, δ denotes the electronic background noise,
[0037] The electronic background noise follows a non-stationary Gaussian distribution with zero mean, denoted as:
[0038]
[0039] where σ 2 denotes the variance of the noise;
[0040] Since the electronic background noise follows a Gaussian distribution, the first conditional distribution can be obtained, denoted as:
[0041]
[0042] The quantum noise W i follows a Poisson distribution, denoted as:
[0043]
[0044] where G i denotes the intensity of the rays in an ideal environment, according to the Beer-Lambert law, G i satisfies:
[0045]
[0046] where Y i denotes the ideal string diagram data, G 0i denotes the actual intensity of the rays;
[0047] According to the Poisson distribution followed by the quantum noise W i , combined with the Beer-Lambert law, the second conditional distribution can be obtained, denoted as:
[0048]
[0049] where i denotes the index of the quanta;
[0050] According to the first conditional distribution and the second conditional distribution, the projection domain denoising model can be obtained, denoted as:
[0051]
[0052] According to the maximum a posteriori estimation theory, the posterior distribution of the projection domain denoising model can be obtained, denoted as:
[0053]
[0054] The posterior distribution of the projection domain denoising model is simplified to obtain an updated projection domain denoising model, which is expressed as:
[0055]
[0056] The updated projection domain denoising model is optimized using the block coordinate descent method to obtain the quantum noise W i The objective function is expressed as:
[0057]
[0058] In the first stage, the memory-limited BFGS algorithm is used to calculate the quantum noise W i The objective function is solved. In the second stage, the damped Newton method is used to solve the quantum noise W i Solve the objective function to obtain the optimal quantum noise, that is, the optimal number of quantum W received by the detector;
[0059] According to the quantum number W received by the optimal detector, updated projection data is obtained.
[0060] It should be noted that, in this embodiment, in the first stage, the memory-limited BFGS algorithm is used to calculate the quantum noise W i The objective function is solved to quickly escape the flat region of the objective function and achieve efficient initial convergence using first-order and approximate second-order information. In the second stage, the damped Newton method is used to solve the quantum noise W i The objective function is solved and when approaching the optimal solution, the precise Hessian matrix is used to achieve superlinear convergence and avoid falling into local oscillation.
[0061] This embodiment also includes: a Convolutional Block Attention Module (CBAM), which is used to enhance the attention mechanism of deep learning network feature expression. By dynamically adjusting the channel and spatial weights of the feature map, the deep learning network pays more attention to the feature area, thereby improving network performance.
[0062] The convolutional block attention module is used to process the updated projection data to enhance the capture of key features.
[0063] CBAM consists of channel attention and spatial attention. Channel attention performs global average pooling and global maximum pooling on the input feature map F to obtain a one-dimensional feature vector respectively. The two feature vectors are passed through a weight-sharing multi-layer perceptron (MLP), and then the weights are added. Finally, the channel attention vector M is obtained through the sigmoid activation function. c(F), is represented as:
[0064]
[0065] wherein, sigma represents a sigmoid function, AvgPool(·) represents an average pooling operation, and MaxPool(·) represents a maximum pooling operation;
[0066] The spatial attention part processes the input feature map F' in the channel dimension through global average pooling and global maximum pooling, respectively, to obtain two feature maps with the same size and F', and the channel number is 1; then the two feature maps are merged, and a 7*7 convolution kernel is used for convolution operation, and finally a sigmoid activation function is used to obtain a spatial attention vector M s (F), is represented as:
[0067]
[0068] wherein, f 7×7 represents a 7*7 convolution kernel.
[0069] In the embodiment, please refer to Figure 2 , Figure 2 is a schematic diagram of the chord domain denoising model provided by the embodiment of the present application, the trained deep learning network further includes a chord domain denoising model, the chord domain denoising model includes a noise estimation network and a first multi-scale Transformer network; the updated projection data is processed in the chord domain to obtain updated chord diagram data, including:
[0070] The noise estimation network is used to process the updated projection data to obtain a noise estimation feature map, and the first multi-scale Transformer network is used to process the noise estimation feature map and the updated projection data to obtain the updated chord diagram data.
[0071] In the embodiment, the noise estimation network includes a plurality of convolution blocks, and each convolution block includes a first convolution layer, a first normalization layer, a first activation function, a second convolution layer, a second normalization layer and a second activation function; optionally, four convolution blocks are set in the embodiment, which are a first convolution block, a second convolution block, a third convolution block and a fourth convolution block; the noise estimation network is used to process the updated projection data to obtain a noise estimation feature map, including:
[0072] The updated projection data is processed by the convolution block, and the processing result is added to the updated projection data to obtain the output result of the convolution block.
[0073] In the embodiment, the first multi-scale Transformer network comprises a third convolutional layer, a first Transformer module, a first multi-scale down-sampling fusion module, a second Transformer module, a second multi-scale down-sampling fusion module, a third Transformer module, a first up-sampling feature fusion module, a fourth Transformer module, a second up-sampling feature fusion module, a fifth Transformer module and a fourth convolutional layer, and optionally, the third convolutional layer and the fourth convolutional layer are both 3x3 convolutional kernels; the first multi-scale Transformer network is used to process the noise estimation feature map and the updated projection data to obtain updated chord diagram data, comprising:
[0074] The third convolutional layer is used to process the noise estimation feature map and the updated projection data to obtain high-dimensional features;
[0075] The first Transformer module is used to process the high-dimensional features to obtain first global features;
[0076] The first multi-scale down-sampling fusion module is used to process the first global features to obtain first detail features;
[0077] The second Transformer module is used to process the first detail features to obtain second global features;
[0078] The second multi-scale down-sampling fusion module is used to process the second global features to obtain second detail features;
[0079] The third Transformer module is used to process the second detail features to obtain third global features;
[0080] The first up-sampling feature fusion module is used to process the third global features and the second global features to obtain first spliced features;
[0081] The fourth Transformer module is used to process the first spliced features to obtain fourth global features;
[0082] The second up-sampling feature fusion module is used to process the fourth global features and the first global features to obtain second spliced features;
[0083] The fifth Transformer module is used to process the second spliced features to obtain fifth global features;
[0084] The fourth convolutional layer is used to process the fifth global features to obtain global features;
[0085] The global features, the noise estimation feature map and the updated projection data are added to obtain the updated chord diagram data.
[0086] It can be understood that the first, second and third Transformer modules are for the encoding stage, and the fourth and fifth Transformer modules belong to the decoding stage, in the encoding stage, the first and second multi-scale down-sampling fusion modules can enhance the global feature correlation and weak component features, in the decoding stage, the first and second up-sampling feature fusion modules can gradually restore the resolution, and the channel attention gate mechanism is used to dynamically weight the features in the encoding process and the features in the decoding process to cross-scale level fusion, which is beneficial to noise elimination and construction of global information.
[0087] It should be noted that the first, second, third, fourth and fifth Transformer modules are all composed of two normalization layers, an attention module and a feedforward network.
[0088] In the embodiment, please refer to Figure 3 and Figure 4 , Figure 3 is a schematic diagram of the first multi-scale down-sampling fusion module provided by the embodiment of the application, Figure 4 is a schematic diagram of the first up-sampling feature fusion module provided by the embodiment of the application, the first and second multi-scale down-sampling fusion modules have the same structure and both include a max-pooling layer, a fifth convolutional layer, a first splicing layer and a sixth convolutional layer, optionally, the fifth convolutional layer is a 4x4 convolutional kernel, the sixth convolutional layer is a 1x1 convolutional kernel, and the max-pooling layer is a 2x2; the first multi-scale down-sampling fusion module is used to process the first global feature to obtain the first detail feature, including:
[0089] the max-pooling layer is used to process the first global feature to obtain an edge feature;
[0090] the fifth convolutional layer is used to process the first global feature to obtain a first feature;
[0091] the first splicing layer is used to splice the edge feature and the first feature to obtain a third splicing feature, so as to expand the receptive field;
[0092] the sixth convolutional layer is used to process the third splicing feature to obtain the first detail feature;
[0093] The first up-sampling feature fusion module and the second up-sampling feature fusion module are the same in structure, and each includes a multi-stage transposed convolution layer, a second splicing layer, a global average pooling layer, a sigmoid function and a seventh convolution layer; the first up-sampling feature fusion module is used to process the third global feature and the second global feature to obtain first splicing features, including:
[0094] The third global feature is processed by the multi-stage transposed convolution layer to obtain a second feature through multi-stage up-sampling;
[0095] The second feature and the second global feature are spliced by the second splicing layer to obtain fourth splicing features;
[0096] The global average pooling layer and the sigmoid function are used to generate spatial attention weights;
[0097] The spatial attention weights are multiplied with the fourth splicing features to obtain a third feature;
[0098] The third feature is processed by the seventh convolution layer to obtain the first splicing features.
[0099] It should be noted that the first multi-scale down-sampling fusion module and the second multi-scale down-sampling fusion module can more finely extract weak component region features; the first up-sampling feature fusion module and the second up-sampling feature fusion module use progressive up-sampling 1 fusion to realize cross-level feature fine calibration through a variable up-sampling chain and a channel attention gate.
[0100] In the embodiment, please refer to Figure 5 , Figure 5 is a schematic diagram of an image domain denoising model provided by the embodiment of the application. The trained deep learning network further includes an image domain denoising model, and the image domain denoising model includes a second multi-scale Transformer network and an image refinement module; in the image domain, the updated chord diagram data is processed to obtain a reconstructed low-dose CT image, including:
[0101] The second multi-scale Transformer network is used to process the updated chord diagram data to obtain a fourth feature; the second multi-scale Transformer network is the same in structure as the first multi-scale Transformer network;
[0102] The image refinement module is used to process the fourth feature to obtain the reconstructed low-dose CT image.
[0103] Optionally, the image refinement module can be a U-net network with CNN as the backbone, and the encoding and decoding operations are realized through a Conv Block.
[0104] In the embodiment, the training process of the trained multi-domain deep learning network includes:
[0105] S1, acquire data of a preset category as samples in a training data set, and acquire real labels of the samples in the training data set; wherein the data of the preset category are low-dose CT images, and the real labels are conventional-dose CT images.
[0106] Specifically, the conventional-dose CT images in the public data of the clinical scanning patients in the "2016 NIH-AAPM-Mayo Clinic Low Dose CT Challenge" are used, and the low-dose CT images corresponding to the conventional-dose CT images are generated based on the CT scanning mechanism by using a simulation method, including low-dose projection data, low-dose chord diagram data and low-dose image data, as shown in FIG. 1. Figure 6 Figure 6 is a flowchart for generating low-dose CT images corresponding to conventional-dose CT images based on the CT scanning mechanism provided by the embodiment of the present application, integer Hounsfield unit values are acquired from the conventional-dose CT images, and then dequantization is performed, the attenuation distribution function corresponding to the human slice tissue is calculated according to the dequantized integer Hounsfield unit values, the attenuation distribution function is converted into chord diagram data corresponding to the conventional-dose CT images, and optionally, the Radon transform is used to calculate the chord diagram data corresponding to the conventional-dose CT images; next, the chord diagram data corresponding to the conventional-dose CT images is processed by using a Poisson-Gaussian noise model, the composite Gaussian and Poisson noise is injected into the undamaged projection, the projection data corresponding to the low-dose CT images is obtained, the projection data corresponding to the low-dose CT images is converted into low-dose CT chord diagram data by using a logarithmic transformation, and then the filtered back projection algorithm is used to obtain the low-dose CT image data, that is, to reconstruct the low-dose CT images.
[0107] S2, input part of the samples in the training data set to the jth deep learning network to be trained to perform training, and obtain a prediction result output in the jth training process.
[0108] S3, calculate a loss according to the prediction result output in the jth training process and the real label of the sample of the jth deep learning network to be trained, and use the loss as the loss of the jth training process.
[0109] S4, perform back propagation according to the loss of the jth training process to update the network parameters of the jth deep learning network to be trained, and obtain a (j+1)th deep learning network to be trained; meanwhile, in the training process, a fine-tuning loss is calculated according to an intermediate iteration result, and the deep learning network in the training process is fine-tuned; the iteration is performed until the training times or the convergence degree meet a preset condition, and a trained deep learning network is obtained.
[0110] In the embodiment, the expression of the loss Loss is as follows:
[0111]
[0112] wherein X n represents the updated image data in the last iteration, Y n represents the updated sinogram image in the last iteration, X * represents the image data corresponding to the low-dose CT image, Y * represents the sinogram image corresponding to the low-dose CT image, a, b and g represent the hyperparameters for controlling the weights of different loss terms, L MAE (·) represents the log average absolute error loss function, L MSE (·) represents the log mean square error loss function, L SSIM (·) represents the structural similarity loss function.
[0113] It should be noted that the embodiment assigns the log average absolute error loss to the sinogram data to prevent the processed sinogram data from being excessively smoothed; the log mean square error loss and the log average absolute error loss are used to solve the distance between the reconstructed image and the reference image; finally, the structural similarity loss is used to improve the structural fidelity.
[0114] In the embodiment, the expression of the fine-tuning loss loss avg is as follows:
[0115]
[0116] wherein n represents the number of iteration layers, w i represents the loss weight of the i-th layer in training, output i represents the reconstructed low-dose CT image output by the i-th layer, and target represents the normal-dose CT image.
[0117] It should be noted that the overall network architecture is obtained by cascading the denoising networks; wherein the output result of the previous network is input into the next network every iteration, which provides a mechanism for repeated top-down and bottom-up reasoning for the network, increases the overall structural complexity and training difficulty of the network, and the reconstruction image of the intermediate layer in the network is included in the loss calculation through intermediate supervision to slow down the error of the output layer. The gradient vanishing phenomenon after multiple back propagation ensures the normal update of the bottom layer parameters.
[0118] It should be noted that the deep learning network of the embodiment is implemented based on the PyTorch framework, uses the Adam optimizer, the batch size is set to 5, and the learning rate is gradually decayed from 2x10 -4 to 1x10 -6The total iteration period of network model training is set to 100 epochs, and finally the optimal model is selected according to the best PSNR value of the verification set, and the model is used for visual enhancement of data containing weak components, and the experiment is completed based on an NVIDIA GeForce RTX 4090 GPU cluster.
[0119] In summary, the low-dose CT reconstruction method provided by the application has the following beneficial effects:
[0120] Firstly, in the process of reconstructing the low-dose CT image, the application considers the noise in the projection domain, the chord diagram domain and the image domain, can maximize the preservation of the tissue structure of the non-artifact area, and realizes the good reconstruction of the low-dose CT image.
[0121] Secondly, the application combines the Unet architecture with the Transformer module, processes the image globally and locally, and realizes better detail processing capability.
[0122] Thirdly, the application proposes a structure of a multi-scale down-sampling fusion module combined with a Transformer module, which can effectively focus on weak component areas and noise areas.
[0123] Fourthly, the application proposes a structure of a Transformer module combined with an up-sampling feature fusion module, which can effectively enhance the weak component while weakening the noise.
[0124] Based on the same inventive concept, the application also provides a low-dose CT reconstruction device for realizing the low-dose CT reconstruction method provided by the above-mentioned embodiments of the application.
[0125] The image acquisition module is used to acquire the low-dose CT image.
[0126] The image processing module is used for image processing, and the trained deep learning network is used to process the low-dose CT image in the projection domain, the chord diagram domain and the image domain respectively.
[0127] The trained deep learning network uses data of a preset category as a training data set to train the initial deep learning network, and the deep learning network in the training process is fine-tuned to obtain the trained deep learning network.
[0128] It should be noted that, in this document, relational terms such as first and second are used solely to distinguish one entity or operation from another, and do not necessarily require or imply any actual relationship or order between these entities or operations. Furthermore, the terms "comprise," "include," or any other variations thereof are intended to encompass non-exclusive inclusion, such that an article or device comprising a list of elements includes not only those elements but also other elements not explicitly listed. Without further limitation, an element defined by the phrase "comprising a..." does not preclude the presence of additional identical elements in the article or device comprising the element. Terms such as "connected" or "connected" are not limited to physical or mechanical connections but may include electrical connections, whether direct or indirect. References to orientations or positional relationships, such as "upper," "lower," "left," and "right," are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate description and simplify the description of the present invention. They do not indicate or imply that the device or element referred to must have, be constructed, or operate in a specific orientation, and are therefore not to be construed as limiting the present invention.
[0129] In the description of this specification, the reference terms "one embodiment," "some embodiments," "example," "specific example," or "some examples" mean that the specific features or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described can be combined in any suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification.
[0130] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. For those skilled in the art of the present invention, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should be considered to fall within the scope of protection of the present invention.
Claims
1. A low-dose CT reconstruction method, characterized in that: include: Acquire low-dose CT images; Using a trained deep learning network to process the low-dose CT image in a projection domain, a chord diagram domain, and an image domain, respectively; processing the low-dose CT image in the projection domain to obtain updated projection data, processing the updated projection data in the chord diagram domain to obtain updated chord diagram data, and processing the updated chord diagram data in the image domain to obtain a reconstructed low-dose CT image; The trained deep learning network uses data of preset categories as a training data set to train the initial deep learning network, performs intermediate supervision during the training process, and fine-tunes the deep learning network in training.
2. The low-dose CT reconstruction method according to claim 1, characterized in that: The trained deep learning network includes a projection domain denoising model; processing the low-dose CT image in the projection domain to obtain updated projection data includes: Construct a projection data generation model, expressed as: P = W + δ; Where W represents the number of quanta received by the detector, δ represents the electronic background noise, The electronic background noise follows a non-stationary Gaussian distribution with zero mean, which can be expressed as: Among them, σ 2 represents the variance of the noise; Since the electronic background noise obeys Gaussian distribution, the first conditional distribution can be obtained, which is expressed as: Quantum noise W i It follows the Poisson distribution, which is expressed as: Among them, G i Indicates the ray intensity under ideal conditions. According to the Beer-Lambert law, G i satisfy: Among them, Y i represents the ideal chord diagram data, G 0i Indicates the actual ray intensity; According to the quantum noise W i The Poisson distribution followed, combined with the Beer-Lambert law, gives the second conditional distribution, which is expressed as: Where i represents the index of quantum; According to the first conditional distribution and the second conditional distribution, the projection domain denoising model is obtained, which is expressed as: According to the maximum a posteriori estimation theory, the posterior distribution of the projection domain denoising model is obtained, which is expressed as: The posterior distribution of the projection domain denoising model is simplified to obtain an updated projection domain denoising model, which is expressed as: The updated projection domain denoising model is optimized using the block coordinate descent method to obtain the quantum noise W i The objective function is expressed as: In the first stage, the memory-limited BFGS algorithm is used to calculate the quantum noise W i The objective function is solved. In the second stage, the damped Newton method is used to solve the quantum noise W i Solve the objective function to obtain the optimal quantum noise, that is, the optimal number of quantum W received by the detector; According to the quantum number W received by the optimal detector, updated projection data is obtained.
3. The low-dose CT reconstruction method according to claim 1, characterized in that: The trained deep learning network also includes a chord diagram domain denoising model, and the chord diagram domain denoising model includes a noise estimation network and a first multi-scale Transformer network; The step of processing the updated projection data in the chord diagram domain to obtain updated chord diagram data includes: The updated projection data is processed using the noise estimation network to obtain a noise estimation feature map, and the noise estimation feature map and the updated projection data are processed using the first multi-scale Transformer network to obtain updated chord diagram data.
4. The low-dose CT reconstruction method according to claim 3, characterized in that: The noise estimation network includes a plurality of convolution blocks, each of which includes a first convolution layer, a first normalization layer, a first activation function, a second convolution layer, a second normalization layer, and a second activation function; The adopting the noise estimation network to process the updated projection data to obtain a noise estimation feature map includes: The updated projection data is processed using the convolution block, and a processing result is added to the updated projection data to obtain an output result of the convolution block.
5. The low-dose CT reconstruction method according to claim 3, characterized in that: The first multi-scale Transformer network includes a third convolutional layer, a first Transformer module, a first multi-scale downsampling fusion module, a second Transformer module, a second multi-scale downsampling fusion module, a third Transformer module, a first upsampling feature fusion module, a fourth Transformer module, a second upsampling feature fusion module, a fifth Transformer module, and a fourth convolutional layer; the first multi-scale Transformer network is used to process the noise estimation feature map and the updated projection data to obtain updated chord diagram data, including: Processing the noise estimation feature map and the updated projection data using the third convolutional layer to obtain high-dimensional features; Processing the high-dimensional features using the first Transformer module to obtain a first global feature; Processing the first global feature using the first multi-scale downsampling fusion module to obtain a first detail feature; Using the second Transformer module to process the first detail feature to obtain a second global feature; Processing the second global feature using the second multi-scale downsampling fusion module to obtain a second detail feature; Processing the second detail feature using the third Transformer module to obtain a third global feature; Using the first upsampling feature fusion module to process the third global feature and the second global feature to obtain a first splicing feature; Using the fourth Transformer module to process the first splicing feature to obtain a fourth global feature; Using the second upsampling feature fusion module to process the fourth global feature and the first global feature to obtain a second splicing feature; Using the fifth Transformer module to process the second splicing feature to obtain a fifth global feature; Processing the fifth global feature using the fourth convolutional layer to obtain a global feature; The global feature, the noise estimation feature map and the updated projection data are added to obtain updated chord diagram data.
6. The low-dose CT reconstruction method according to claim 5, characterized in that: The first multi-scale downsampling fusion module and the second multi-scale downsampling fusion module have the same structure, both including a maximum pooling layer, a fifth convolutional layer, a first splicing layer, and a sixth convolutional layer; the first multi-scale downsampling fusion module is used to process the first global feature to obtain a first detail feature, including: Processing the first global feature using the maximum pooling layer to obtain an edge feature; Processing the first global feature using the fifth convolutional layer to obtain a first feature; Using the first splicing layer to splice the edge feature and the first feature to obtain a third splicing feature; Using the sixth convolutional layer to process the third splicing feature to obtain a first detail feature; The first upsampling feature fusion module and the second upsampling feature fusion module have the same structure, both including a multi-stage transposed convolution layer, a second splicing layer, a global average pooling layer, a sigmoid function, and a seventh convolution layer; the first upsampling feature fusion module is used to process the third global feature and the second global feature to obtain a first splicing feature, including: Performing multi-level upsampling processing on the third global feature using the multi-level transposed convolution layer to obtain a second feature; Using the second splicing layer to splice the second feature and the second global feature to obtain a fourth splicing feature; Generate spatial attention weights using the global average pooling layer and the sigmoid function; Multiplying the spatial attention weight by the fourth concatenated feature to obtain a third feature; The seventh convolutional layer is used to process the third feature to obtain a first splicing feature.
7. The low-dose CT reconstruction method according to claim 1, characterized in that: The trained deep learning network also includes an image domain denoising model, which includes a second multi-scale Transformer network and an image refinement module; The step of processing the updated chord diagram data in the image domain to obtain a reconstructed low-dose CT image includes: Processing the updated chord graph data using the second multi-scale Transformer network to obtain a fourth feature; wherein the second multi-scale Transformer network has the same structure as the first multi-scale Transformer network; The image refinement module is used to process the fourth feature to obtain a reconstructed low-dose CT image.
8. The low-dose CT reconstruction method according to claim 1, characterized in that: The training process of the trained deep learning network includes: Obtaining data of a preset category as samples in a training data set, and obtaining true labels of the samples in the training data set; wherein the data of the preset category are low-dose CT images, and the true labels are conventional-dose CT images; Inputting some samples in the training data set into the deep learning network to be trained for the jth time for training, and obtaining the prediction results output during the jth training process; Calculate the loss based on the prediction results output during the j-th training process and the true labels of the samples of the deep learning network to be trained for the j-th training process, and use it as the loss of the j-th training process; Backpropagation is performed based on the loss of the j-th training process to update the network parameters of the deep learning network to be trained for the j-th time, and the deep learning network to be trained for the j+1-th time is obtained; at the same time, during the training process, the fine-tuning loss is calculated based on the intermediate iteration results, and the deep learning network in training is fine-tuned; this iteration is repeated until the number of training times or the degree of convergence meets the preset conditions, and the trained deep learning network is obtained.
9. The low-dose CT reconstruction method according to claim 8, characterized in that: The expression for calculating the loss Loss is: Among them, X n Indicates the updated image data in the last iteration, Y n represents the updated chord diagram image in the last iteration, X * Represents the image data corresponding to the low-dose CT image, Y * represents the chord image corresponding to the low-dose CT image, α, β, and γ represent the hyperparameters that control the weights of different loss terms, L MAE (·) represents the logarithmic mean absolute error loss function, L MSE (·) represents the log mean square error loss function, L SSIM (·) represents the structural similarity loss function.
10. A low-dose CT reconstruction device, characterized in that: include: An image acquisition module, used for acquiring low-dose CT images; an image processing module, configured to process the low-dose CT image in a projection domain, a chord diagram domain, and an image domain using a trained deep learning network; processing the low-dose CT image in the projection domain to obtain updated projection data, processing the updated projection data in the chord diagram domain to obtain updated chord diagram data, and processing the updated chord diagram data in the image domain to obtain a reconstructed low-dose CT image; The trained deep learning network uses data of preset categories as a training data set to train the initial deep learning network, performs intermediate supervision during the training process, and fine-tunes the deep learning network in training.
Citation Information
Patent Citations
Medical low-dose CT image restoration method based on frequency domain Transform
CN117541500A
Low-dose CT image super-resolution method and system based on multi-scale wavelet transform
CN118674623A
Systems and methods for multimodal fusion of missing and unpaired image and tabular data
KR1020240141126A
Cited By
Pre-logarithm field Voronoi decomposition assisted low-dose CT reconstruction method
CN122176118A
A pre-logarithmic domain voronoi decomposition assisted low-dose ct reconstruction method
CN122176118B