Self-supervised CT reconstruction method based on adaptive domain transformation
Through adaptive domain transformation and self-supervised learning, the problems of poor image quality and excessive radiation caused by undersampling are solved, and high-quality and low-radiation CT image reconstruction is achieved, which is suitable for medical diagnosis.
Patent Information
- Application Number
- CN202510411567.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-11
AI Technical Summary
The existing CT reconstruction technology has poor quality in undersampling, lacks interpretability, and has high radiation dose, making it difficult to meet the high-precision needs of medical diagnosis.
The self-supervised CT reconstruction method of adaptive domain transformation is adopted, and the tomographic information matrix is introduced through adaptive domain transformation, combined with the tomographic attention reconstruction module and image optimization network, and the reconstruction results are optimized by self-supervised learning to reduce radiation dose and improve image quality.
Realize high-quality CT image reconstruction under undersampling conditions, reduce radiation dose, improve reconstruction speed, avoid artifacts, and meet medical diagnosis needs.
Smart Images

Figure CN120298523A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of CT reconstruction, and specifically provides a self-supervised CT reconstruction method based on adaptive domain transformation. Background Art
[0002] Computed tomography (CT), as a key technology in the field of medical imaging, can generate high-quality cross-sectional images of the body by the coordinated operation of X-ray equipment and computer processing technology. These images provide extremely detailed and clear information for medical diagnosis and treatment, playing an indispensable and important role in the medical field.
[0003] In the actual process of CT imaging, X-rays penetrate the human body from multiple rotation angles, and the detector is responsible for receiving the X-ray attenuation information at different angles, thereby forming a sinogram that records the changes in ray intensity and angle. Subsequently, CT reconstruction technology is used to convert the sinogram into a cross-sectional tomographic image that describes the internal structure of the body. However, there are many influencing factors in the imaging process. Among them, the undersampling defects caused by insufficient X-ray dose or lack of projection angles will cause artifacts in the reconstructed image, and the imaging quality will also be greatly reduced.
[0004] With the continuous progress of medical technology, the combination of CT reconstruction technology and deep learning has given rise to deep CT reconstruction technology, aiming to solve the problems caused by undersampling. However, there are still many defects in the actual application of this technology. On the one hand, when using an end-to-end network for reconstruction, the sinogram signal is usually directly converted into a CT image. This method makes the reconstruction process lack interpretability, and it is difficult for doctors to deeply understand the principle and basis of image formation. On the other hand, its reconstruction accuracy is also unsatisfactory and cannot meet the strict requirements of clinical diagnosis for high-precision images. In addition, the difficulty of obtaining data in the field of medical images also limits the further development and optimization of deep CT reconstruction technology.
[0005] In view of this, a self-supervised CT reconstruction method based on adaptive domain transformation is provided to overcome the above problems. Summary of the Invention
[0006] The purpose of the present invention is to provide a self-supervised CT reconstruction method based on adaptive domain transformation to solve the problems mentioned in the above background art.
[0007] To solve the above technical problems, a self-supervised CT reconstruction method based on adaptive domain transformation provided by the present invention includes the following steps:
[0008] S1. Adaptive domain transformation:
[0009] S1.1. Sinogram-CT image cross-domain transformation: For any point m of body tissue ( x , y)
[0010] CT projection, based on the scanner rotation projection angle θ and the reconstructed target position coordinates (x, y), through the formula:
[0011]
[0012] Calculate the position t of the target on the detector, where w is the width of the reconstructed target image, H is the height, D is the detector width, and combine the Radon transform theory to determine the signal intensity p of any point (t, θ) in the sinogram; for N projection angles, calculate the reconstructed target point m (x,y) The coordinates (t j , θ j ) and its corresponding signal intensity p(t j , θ j ), introduce the tomographic information matrix M(x, y, θ) ∈ R 3 To represent and calculate the intensity of the sine curve of each reconstructed target pixel;
[0013] S1.2, Tomographic attention reconstruction: Define the tomographic attention matrix G(x, y, θ) ∈ R H×W×N , which is composed of the spatial attention α(x, y) ∈ R H×W and the perspective attention ω(θ) ∈ R N respectively; through:
[0014]
[0015] Calculate the relevant weights; use the attention mechanism to learn and obtain the attention matrix G(x, y, θ), and use the formula:
[0016] m (x,y) = σ(OutConv(G(x, y, θ) × M(x, y, θ)))
[0017] Calculate the intensity of the reconstructed target pixel to obtain the initial reconstruction result m;
[0018] S2. Image optimization network: Use the U-Net encoder-decoder network to optimize the reconstruction result R' ∈ R H×W obtained by the adaptive domain transformation; the encoder consists of 4 groups of convolutional layers, and each group of convolutional layers uses a 3×3 convolutional kernel to extract image features, uses the activation function ReLU and combines the max-pooling operation to obtain a deep feature map with high-level semantic expression ability:
[0019] f ∈ R H’×W’×1024 ;
[0020] The decoder adopts 4 sets of convolutional layers, and maps the high-level semantic features f to the CT image space at high resolution through layer-by-layer upsampling and convolutional operations; at the output end of the decoder, an output layer and a Sigmoid activation function are used to normalize the result between 0 and 1 to obtain the optimized reconstructed image:
[0021] R ∈ R H×W ,
[0022] and a skip connection layer is introduced between the encoder and the decoder;
[0023] S3. Self-supervised learning: Using the Radon transform, according to the formula:
[0024] S'(t,θ) = ∫ L f(x,y)dl, L: t = xcosθ + ysinθ
[0025] project the reconstructed result R output by the network into the sinogram space; in the sinogram space, use the formula:
[0026]
[0027] calculate the loss function, where ‖·‖ 2 represents the Euclidean distance between the predicted sinogram and the observed sinogram, and ε is a decimal number approaching 0. The CT reconstruction model is optimized by minimizing the loss function.
[0028] Furthermore, in the sinogram-CT image cross-domain transformation module, W is the width of the reconstructed target image, H is the height of the reconstructed target image, and D is the width of the detector in the sinogram.
[0029] Furthermore, in the tomographic attention reconstruction module, the calculation of the spatial attention α(x,y) involves the set P of reconstructed target pixels, which is used to reflect the scattering influence weight coefficients between different pixel points in space.
[0030] Furthermore, in the tomographic attention reconstruction module, the view angle attention ω(θ) is for N projection angles, which is used to reflect the contribution weights of the original signals at different angles to the theoretical reconstruction result.
[0031] Furthermore, in the encoder of the image optimization network, the maximum pooling operation is combined to reduce the spatial size of the feature map and increase the number of channels of the feature map at the same time, which is used to reduce the input data volume.
[0032] Furthermore, in the decoder of the image optimization network, through layer-by-layer upsampling and convolutional operations, the details and size of the image are gradually restored.
[0033] Furthermore, in self-supervised learning, the basis for the Radon transform calculation:
[0034] L: t = xcosθ + ysinθ
[0035] For determining the projection relationship.
[0036] Furthermore, in the adaptive domain transformation, when the tomographic attention reconstruction module obtains the attention matrix G(x, y, θ), a deep learning-based attention mechanism is adopted.
[0037] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0038] 1. Improve the quality of the reconstructed image: The existing methods are prone to problems such as artifacts in the reconstructed image and poor imaging quality. However, in this method, the tomographic information matrix is introduced through the adaptive domain transformation, providing rich and interpretable information for cross-domain reconstruction; at the same time, the tomographic attention reconstruction module is used to simulate the scattering effect of X-ray imaging, and the reconstruction result is optimized in combination with the image optimization network; in the under-sampled reconstruction scenarios with sparse and limited angles and other low-radiation conditions, this method can still reconstruct the CT target image well, avoiding the artifact phenomenon in the traditional methods, so as to reduce the radiation of CT diagnosis while ensuring the high-quality reconstruction of the image.
[0039] 2. Reduce the radiation dose: On the premise of ensuring the image quality, this method can reduce the signal acquisition density of CT, thereby reducing the overall radiation dose and exposure time of X-rays to the human body. This benefits from its characteristic of still being able to achieve high-quality reconstruction under under-sampling conditions. When obtaining images with the same diagnostic value, it does not need to rely on high-dose X-rays and a large number of projection angles like the traditional methods, thus reducing the radiation hazard received by patients in medical diagnosis.
[0040] 3. Thanks to the adaptive domain transformation function and lightweight network configuration, this method can quickly reconstruct CT images, and has certain competitiveness in terms of time compared with the existing methods on the premise of ensuring the reconstruction quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 It is the schematic diagram of a self-supervised CT reconstruction method based on adaptive domain transformation of the present invention;
[0042] Figure 2 It is the implementation detail diagram of the sinogram-CT image cross-domain transformation module in a self-supervised CT reconstruction method based on adaptive domain transformation of the present invention;
[0043] Figure 3 It is the implementation detail diagram of the tomographic attention reconstruction module in a self-supervised CT reconstruction method based on adaptive domain transformation of the present invention;
[0044] Figure 4 It is the implementation detail diagram of the CT image optimization network in a self-supervised CT reconstruction method based on adaptive domain transformation of the present invention;
[0045] Figure 5 This is a comparison chart of the reconstruction time, PSNR, and SSIM of each algorithm at different projection angles in a self-supervised CT reconstruction method based on adaptive domain transformation according to the present invention;
[0046] Figure 6 This is another comparison chart of the reconstruction time, PSNR, and SSIM of each algorithm at different projection angles in a self-supervised CT reconstruction method based on adaptive domain transformation according to the present invention;
[0047] Figure 7 This is a comparison chart of the reconstructed images of the FBP, SART-TV, N2-Learned, NeRP, DuDoTrans, and ADT algorithms in a self-supervised CT reconstruction method based on adaptive domain transformation according to the present invention when N = 30, N = 60, and N = 180;
[0048] Figure 8 This is another comparison chart of the reconstructed images of the FBP, SART-TV, N2-Learned, NeRP, DuDoTrans, and ADT algorithms in a self-supervised CT reconstruction method based on adaptive domain transformation according to the present invention when N = 30, N = 60, and N = 180;
[0049] Figure 9 This is a comparison chart of PSNR, SSIM, and time in the LoDoPaB-test and AAPM-test of a self-supervised CT reconstruction method based on adaptive domain transformation according to the present invention;
[0050] Figure 10 This is a comparison chart of the performance of each algorithm in a self-supervised CT reconstruction method based on adaptive domain transformation according to the present invention under different evaluation metrics. Detailed implementation manners
[0051] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0052] Please refer to Figures 1 - 10 , the present invention provides a technical solution:
[0053] Refer to Figures 1 - 10 shown, an embodiment of a self-supervised CT reconstruction method based on adaptive domain transformation:
[0055] This embodiment will introduce in detail the specific implementation process of a self-supervised CT reconstruction method based on adaptive domain transformation. This method aims to solve the problem of image reconstruction artifacts introduced in the sparse CT imaging process, while reducing the signal acquisition density of CT, reducing the overall radiation dose and exposure time of X-rays to the human body. This embodiment will be described in conjunction with the accompanying drawings to enable those skilled in the art to better understand and implement this method.
[0056] As shown in the attached Figure 1 figures, the technologies adopted by the self-supervised CT reconstruction method based on adaptive domain transformation mainly include three aspects: adaptive domain transformation, image optimization network, and self-supervised learning.
[0057] (1) Adaptive domain transformation
[0058] Step 1: Sinogram-CT image cross-domain transformation module. When performing CT projection on any point m of the body tissue (x,y ), as the projection angle θ of the scanner rotates, the position of this point on the original signal detector will also change accordingly. This position is reflected in the distance t from the target point to the center point on the y-axis of the sinogram, and the projection angle θ is reflected in the x-axis of the sinogram. Given the reconstruction target position coordinates (x, y) and the projection angle θ, the position t of the target on the detector can be calculated by the following formula:
[0059]
[0060] Here, W and H are the width and height of the reconstruction target image, and D represents the width of the detector in the sinogram. Combining with the Radon transform theory in the projection process, the signal intensity p of any point (t, θ) in the sinogram is:
[0061] p(t, θ) = ∫∫m (x,y) δ(xcosθ + ysinθ - t)dxdy
[0062] For N projection angles, using the above two formulas to calculate the coordinates (t (x,y) in the sinogram of the reconstruction target point m j , θ j ), j ∈ {1, 2, 3,..., N} and its corresponding signal intensity p(t j , θ j ). This signal forms a sine curve highly correlated with the reconstruction target pixel (x, y). This method introduces a tomographic information matrix M(x, y, θ) ∈ R 3 to represent and calculate the intensity of the sine curve of each reconstruction target pixel. The first two dimensions of M are consistent with the reconstruction image dimension, and its third-dimensional information represents the intensity information on the sine curve, providing a rich and interpretable information and data bridge for the cross-domain reconstruction from the sinogram to the CT image. Attached Figure 2Shows the calculation process of the tomographic information matrix for three reconstructed target pixels, represented in blue, yellow, and pink respectively.
[0063] Step 2: Tomographic attention reconstruction module. To simulate the scattering effect in the X-ray imaging process and apply it to more accurate medical image reconstruction, this method defines a tomographic attention matrix G(x, y, θ) ∈ R H×W×N , which is composed of a spatial attention α(x, y) ∈ R H×W and a perspective attention ω(θ) ∈ R N . The spatial attention reflects the scattering influence weight coefficient between different pixel points in space, making P represents the set of reconstructed target pixels, is the theoretical reconstruction result without scattering phenomenon; while the perspective attention reflects the contribution weight of the original signals from different angles to the theoretical reconstruction result when reconstructing the target pixel points, that is In practical applications, this attention matrix G(x, y, θ) is obtained by the learning method of the attention mechanism, and its implementation details are as attached Figure 3 . Furthermore, the intensity of the reconstructed target pixel is obtained by the following formula:
[0064] m (x,y) = σ(OutConv(G(x, y, θ) × M(x, y, θ)))
[0065] where OutConv(·) represents the 1×1 convolution function for dimension optimization, and σ(·) is the non-linear activation function Sigmoid, which is used to control the numerical range of the output layer to be [0, 1], so that the tomographic attention module can obtain an initial reconstruction result m with clear contours and structural information and reduce the dimension.
[0066] (2) Image optimization network
[0067] Assume that the reconstruction result obtained by the above adaptive domain transformation method is R' ∈ R H×W , and we use a U-Net encoder-decoder network to optimize the reconstruction result to obtain the final reconstruction result R ∈ R H×W , as shown in the attachment Figure 4 . The encoder consists of 4 groups of convolutional layers. Each group of convolutional layers uses a 3×3 convolutional kernel to extract image features, and uses the activation function Relu to add non-linearity to the feature map, combined with the max-pooling operation to reduce the spatial size of the feature map and increase the number of channels of the feature map at the same time, so as to reduce the input data volume and improve the high-level feature expression ability. After adopting the above encoder structure, a deep feature map f ∈ R with high-level semantic expression ability can be obtained H’×W’×1024Subsequently, 4 sets of convolutional layers are also used as the decoder to map the high-level semantic features f to the CT image space at high resolution. Through layer-by-layer upsampling and convolutional operations, the details and dimensions of the image are gradually restored to achieve high-precision CT image reconstruction. At the output end of the decoder, the image optimization network uses an output layer and a Sigmoid activation function to normalize the result between 0 and 1, obtaining the optimized reconstructed image R∈R H×W In addition, to reduce information loss during the encoding process and improve the reconstruction quality of the detailed information in the image, we introduce a skip connection layer between the encoder and the decoder.
[0068] (3) Self-supervised learning
[0069] To avoid a large number of reconstructed images participating in model training, we convert the reconstructed image to the sinogram space to calculate the loss function instead of calculating it in the image space, so as to achieve self-supervised learning. For this purpose, the Radon transform needs to be used to project the reconstructed result R output by the network into the sinogram space, and its calculation formula is:
[0070] S'(t,θ)=∫ L f(x,y)dl,L:t=xcosθ+ysinθ
[0071] Furthermore, the loss function in the sinogram space can be calculated by the following formula:
[0072]
[0073] where ‖·‖ 2 represents the Euclidean distance between the predicted sinogram and the observed sinogram, which directly reflects the global difference between the two, ensuring the correctness and accuracy of the network in image reconstruction. ε is a decimal number approaching 0, which can avoid the abnormal situation of the loss function gradient or derivative when the local error value approaches zero.
[0074] Summary:
[0075] By continuously adjusting the model parameters, the loss function is minimized, thereby optimizing the entire CT reconstruction model. In practical applications, relevant parameters can be adjusted according to different scanning requirements and device conditions to achieve the best reconstruction effect. For example, when facing scans of different parts, the number N of projection angles can be appropriately adjusted, or the parameters of the convolutional layers in the network can be adjusted according to the performance of the device.
[0076] This embodiment demonstrates the specific implementation process of the self-supervised CT reconstruction method based on adaptive domain transformation through detailed steps. This method can effectively avoid the artifact phenomenon that appears in traditional methods in under-sampled reconstructions with low radiation such as sparse and limited-angle, and achieve fast reconstruction while ensuring high-quality image reconstruction, with significant advantages.
Claims
1. A self-supervised CT reconstruction method based on adaptive domain transformation, characterized in that Including the following steps: S1. Adaptive domain transformation; S1.
1. Sinogram-CT Image Cross-Domain Transformation: For any point m on the body tissue (x,y) CT projection, according to the scanner rotation projection angle θ and the reconstructed target position coordinates (x, y), through the formula: Calculate the position \(t\) of the target on the detector, where \(w\) is the width of the reconstructed target image, \(H\) is the height, \(D\) is the width of the detector, and determine the signal intensity \(p\) of any point \((t,\theta)\) in the sinogram in combination with the Radon transform theory; for \(N\) projection angles, calculate the reconstructed target point \(m\). (x,y) The coordinates \((t\) j ,\theta\) j ) and its corresponding signal intensity \(p(t\) j ,\theta\) j ), introduce the tomographic information matrix \(M(x,y,\theta)\in R\) 3 to represent and calculate the intensity of the sine curve of each reconstructed target pixel. S1.2, Tomographic attention reconstruction: Define the tomographic attention matrix \(G(x, y, \theta)\in\mathbb{R}\) H×W×N , which is respectively composed of the spatial attention \(\alpha(x, y)\in\mathbb{R}\) H×W and the perspective attention \(\omega(\theta)\in\mathbb{R}\) N ; Through: Calculate the relevant weights; Use the attention mechanism to learn and obtain the attention matrix G(x, y, θ), and use the formula: m (x,y) = σ(OutConv(G(x, y, θ) × M(x, y, θ))); Calculate the intensity of the reconstructed target pixel to obtain the initial reconstruction result m; S2. Image Optimization Network: The U-Net encoder-decoder network is used to optimize the reconstruction result R' ∈ R obtained by the adaptive domain transformation. H×W The encoder consists of 4 sets of convolutional layers. Each set of convolutional layers uses a 3×3 convolutional kernel to extract image features, and the activation function ReLU is used in combination with the max pooling operation to obtain a deep feature map with high-level semantic expression ability. f ∈ R H′×W′×1024 ; The decoder adopts 4 groups of convolutional layers, and maps the high-level semantic feature f to the CT image space at high resolution through layer-by-layer upsampling and convolution operations; At the output end of the decoder, an output layer and a Sigmoid activation function are used to normalize the result between 0 and 1 to obtain the optimized reconstructed image: And a skip connection layer is introduced between the encoder and the decoder; R ∈ R H×W , S3. Self-supervised learning: Use the Radon transform, according to the formula: Project the reconstructed result R output by the network into the sinogram space; In the sinogram space, use the formula: S'(t,θ) = ∫ L f(x,y)dl, L: t = xcosθ + ysinθ; Indicates the Euclidean distance between the predicted sinogram and the observed sinogram, ε Calculate the loss function, where ‖·‖ 2 Is a decimal number approaching 0, and the CT reconstruction model is optimized by minimizing the loss function. In the sinogram- 2. The self-supervised CT reconstruction method based on adaptive domain transformation according to claim 1, wherein: CT image cross-domain transformation module, W is the width of the reconstructed target image, H is the height of the reconstructed target image, and D is the width of the detector in the sinogram. In the tomographic attention reconstruction module, the calculation of the spatial attention α(x, y) involves the set P of reconstructed target pixels, which is used to reflect the scattering influence weight coefficients between different pixel points in space.
3. The self-supervised CT reconstruction method based on adaptive domain transformation according to claim 1, wherein: In the tomographic attention reconstruction module, the view angle attention ω(θ) is for N projection angles, which is used to reflect the contribution weights of the original signals at different angles to the theoretical reconstruction result.
4. The self-supervised CT reconstruction method based on adaptive domain transformation according to claim 1, characterized in that: In the encoder of the image optimization network, the maximum pooling operation is combined to reduce the spatial size of the feature map and increase the number of channels of the feature map at the same time, which is used to reduce the input data volume.
5. A self-supervised CT reconstruction method based on adaptive domain transformation according to claim 1, characterized in that: In the decoder of the image optimization network, the details and size of the image are gradually restored through layer-by-layer upsampling and convolution operations.
6. The self-supervised CT reconstruction method based on adaptive domain transformation according to claim 1, wherein: In self-supervised learning, when calculating the Radon transform, the projection relationship is determined according to: L: t = xcosθ + ysinθ.
7. A self-supervised CT reconstruction method based on adaptive domain transformation according to claim 1, characterized in that: In adaptive domain transformation, when the tomographic attention reconstruction module obtains the attention matrix G(x, y, θ), a deep learning-based attention mechanism is adopted.
8. The self-supervised CT reconstruction method based on adaptive domain transformation according to claim 1, wherein: