Method and system for sparse-angle ct image artifact removal
Patent Information
- Application Number
- CN202411161056.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-22
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-08-22
AI Technical Summary
这种方法虽然在某些情况下能够取得不错的效果,仍然存在一些问题,例如:CT图像本质上是三维数据,仅处理二维断层会忽略相邻层之间的空间关联性,导致重要的三维结构信息丢失;又例如,由于每个断层是独立处理的,重建后的三维图像在不同切面(如矢状面和冠状面)上可能出现不连续或不一致的情况,影响诊断的准确性;又例如,对每个断层单独进行处理需要频繁的数据输入输出,在处理大量CT图像时效率较低;二维处理方法难以有效捕捉复杂的三维伪影模式,导致处理效果在不同部位或不同扫描参数下的CT图像上表现不一致;又例如,传统的损失函数(如均方误差)往往导致生成的图像过于平滑,缺乏真实感,不符合人类视觉系统的感知特性
[0073]本申请提出了一种用于稀疏角度CT图像伪影去除的创新方法,其核心在于充分利用CT图像的三维信息和先进的深度学习技术。该方法主要包含三个关键点,每个关键点都带来了显著的技术效果。
Smart Images

Figure CN119006315B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of medical image processing, and in particular to artifact removal techniques for sparse angle computed tomography (CT) images. Background Technology
[0002] Computed tomography (CT) is a widely used medical imaging technology that plays a crucial role in disease diagnosis, surgical planning, and treatment evaluation. However, the high-dose radiation used in traditional CT scans poses potential risks to patients. To reduce radiation dose, sparse-angle CT technology has gained increasing attention in recent years. This technology reduces radiation dose by decreasing the projection angle, but it also introduces the problem of degraded image quality, primarily manifested as artifacts.
[0003] Artifacts in sparse-angle CT images not only affect the visual quality of the images, but more importantly, may lead to incorrect medical diagnoses. Traditional artifact removal methods are mainly based on image processing techniques, such as filtering and iterative reconstruction algorithms. However, these methods often struggle to achieve a good balance between reducing artifacts and preserving image details. With the development of deep learning technology, neural network-based methods have shown great potential in the field of image processing and are increasingly being applied to artifact removal tasks in CT images.
[0004] Currently, most deep learning-based CT image artifact removal methods focus on processing two-dimensional tomographic images. While this approach can achieve good results in some cases, it still has several drawbacks. For example, CT images are inherently three-dimensional data, and processing only two-dimensional slices ignores the spatial relationships between adjacent slices, leading to the loss of important three-dimensional structural information. Furthermore, because each slice is processed independently, the reconstructed three-dimensional image may exhibit discontinuities or inconsistencies across different planes (such as sagittal and coronal planes), affecting diagnostic accuracy. Additionally, processing each slice individually requires frequent data input and output, resulting in low efficiency when processing large numbers of CT images. Two-dimensional processing methods struggle to effectively capture complex three-dimensional artifact patterns, leading to inconsistent processing results across different locations or under different scanning parameters. Finally, traditional loss functions (such as mean squared error) often result in overly smooth images that lack realism and do not conform to the perceptual characteristics of the human visual system.
[0005] These problems severely restrict the widespread application of sparse-angle CT technology in clinical practice. Therefore, there is an urgent need for a novel artifact removal method that can fully utilize the three-dimensional information of CT images, improve spatial consistency, enhance processing efficiency, stabilize image quality, and improve perceived quality. This method should not only effectively remove artifacts but also generate visually more natural images that conform to doctors' reading habits while preserving key diagnostic information. Furthermore, considering the real-time requirements of the medical field, this method also needs to be capable of processing large-scale CT data to meet the growing clinical needs. Summary of the Invention
[0006] The purpose of this application is to provide a method and system for removing artifacts in sparse angle CT images, so as to solve the problems mentioned in the background art.
[0007] This application discloses a method for artifact removal in sparse angle CT images, comprising the following steps:
[0008] Acquire sparse angle CT images and full angle CT images, and preprocess them to obtain preprocessed sparse angle CT images and full angle CT images;
[0009] The preprocessed sparse angle CT image and full angle CT image are divided into several training 3D sub-blocks; and a block-based sparse angle CT image artifact removal network is constructed, the network including two pairs of generators and discriminators, one pair is used to transform the training 3D sub-blocks from the sparse angle image domain to the full angle image domain, and the other pair is used to perform the reverse process.
[0010] The network is trained using the three-dimensional sub-blocks used for training. During the training process, a loss function including adversarial loss, cycle consistency loss, and reconstruction loss is used to optimize the network, resulting in a trained network.
[0011] After preprocessing the sparse angle CT image to be processed, it is divided into several three-dimensional sub-blocks of the same size as the three-dimensional sub-blocks according to spatial position. The three-dimensional sub-blocks to be processed are then input into the trained network for processing to obtain three-dimensional sub-blocks after artifact removal. The artifact-removed three-dimensional sub-blocks are then merged to form a CT image after artifact removal.
[0012] In a preferred embodiment, the generator of the volume-based sparse angular CT image artifact removal network includes, in sequence, an encoder module, a dilated spatial convolutional pooling pyramid module, multiple volume attention-guided residual modules, and a decoder module; wherein:
[0013] The encoder module consists of multiple three-dimensional convolutional layers connected in series. Each three-dimensional convolutional layer is followed by an instance normalization layer and a ReLU activation function to extract features of the training three-dimensional sub-blocks.
[0014] The dilated spatial convolutional pooling pyramid module receives the output of the encoder module, including multiple dilated convolutional branches with different dilation rates and a global average pooling branch. The outputs of each branch are concatenated along the channel direction to capture multi-scale features.
[0015] The volume attention-guided residual module includes a volume attention weighted branch, a convolutional branch, and a residual connection branch. The volume attention weighted branch pools the input features in four dimensions: channel, depth, height, and width. After processing through convolutional layers, instance normalization layers, and a sigmoid activation function, it captures cross-dimensional correlations between different dimensions and multiplies them with the input features. The convolutional branch includes multiple three-dimensional convolutional layers. The residual connection branch directly passes the input features.
[0016] The decoder module consists of multiple 3D transposed convolutional layers connected in series. Each 3D transposed convolutional layer is followed by an instance normalization layer and an activation function to reconstruct the processed features into an output image with the same size as the input.
[0017] In a preferred embodiment, the encoder module consists of three cascaded three-dimensional convolutional layers, each followed by an instance normalization layer and a ReLU activation function. The first three-dimensional convolutional layer has a kernel size of 7×7×7 and a stride of 1, while the second and third three-dimensional convolutional layers have kernel sizes of 3×3×3 and a stride of 2.
[0018] In a preferred embodiment, the dilated spatial convolutional pooling pyramid module includes a 1×1×1 convolutional branch, three 3×3×3 dilated convolutional branches with different dilation rates, and a global average pooling branch, with the outputs of each branch spliced along the channel direction.
[0019] In a preferred embodiment, the volume attention-guided residual module includes:
[0020] - The volumetric attention weighted branch pools the input features in four dimensions: channel, depth, height, and width. Then, it captures the cross-dimensional correlation between each of the three dimensions through convolutional layers, instance normalization layers, and the Sigmoid activation function, and then multiplies it with the input features.
[0021] - Convolutional branch, consisting of two cascaded 3×3×3 three-dimensional convolutional layers with a stride of 1;
[0022] - Residual connection branches are used to directly pass input features.
[0023] In a preferred embodiment, the decoder module consists of three cascaded three-dimensional transposed convolutional layers, wherein the kernel size of the first and second three-dimensional transposed convolutional layers is 3×3×3 with a stride of 2, and each layer is followed by an instance normalization layer and a ReLU activation function. The kernel size of the third three-dimensional transposed convolutional layer is 7×7×7 with a stride of 1, and a Tanh activation function is added.
[0024] In a preferred embodiment, the cycle consistency loss in the loss function is calculated using a perceptual metric to better reflect the characteristics of the human visual system.
[0025] In a preferred embodiment, the perceptual metric is a learned perceptual patch similarity (LPIPS) metric, which is used to calculate the perceptual difference between the generated image and the original image, thereby avoiding over-smoothing of the generated image by capturing high-level features.
[0026] In a preferred embodiment, the cycle consistency loss includes:
[0027] After the sparse angle CT image is transformed to the full angle CT image domain by the first generator A, it is then transformed back to the sparse angle CT image domain by the second generator B. The LPIPS metric between the obtained result and the original sparse angle CT image is obtained.
[0028] After transforming the full-angle CT image to the sparse angle CT image domain through the second generator B, it is then transformed back to the full-angle CT image domain through the first generator A. The LPIPS metric between the obtained result and the original full-angle CT image is then calculated.
[0029] In a preferred embodiment, the formula for calculating the cycle consistency loss is:
[0030] L cycle =E[||B(A(s))-s|| LPIPS ]+E[||A(B(f))-f|| LPIPS (5)
[0031] L cycle This represents the loss of cycle consistency.
[0032] E[·] represents the expectation operation;
[0033] ||·|| LPIPS This indicates that the distance was calculated using LPIPS (Learned Perceptual Patch Similarity) metric.
[0034] s represents a sparse angle CT image;
[0035] f represents a full-angle CT image;
[0036] A(·) represents the first generator that transforms sparse angle CT images into the full angle CT image domain;
[0037] B(·) represents the second generator that transforms the full-angle CT image to the sparse-angle CT image domain;
[0038] B(A(s)) represents the result of transforming the sparse angle CT image s into the full angle CT image domain first through generator A, and then transforming it back into the sparse angle CT image domain through generator B.
[0039] A(B(f)) represents the result of transforming the full-angle CT image f into the sparse angle CT image domain through generator B, and then transforming it back into the full-angle CT image domain through generator A.
[0040] In a preferred embodiment, the loss function further includes a reconstruction loss for calculating the perceptual difference between the generated image and the ground truth image, guiding the network to optimize image details.
[0041] In a preferred embodiment, the formula for calculating the reconstruction loss is:
[0042] L rec =E[||A(s)-f|| LPIPS ]+E[||B(f)-s|| LPIPS (6)
[0043] in,
[0044] L rec Indicates the losses incurred during reconstruction;
[0045] E[·] represents the expectation operation;
[0046] ||·|| LPIPS This indicates that the distance was calculated using LPIPS (Learned Perceptual Patch Similarity) metric.
[0047] s represents a sparse angle CT image;
[0048] f represents a full-angle CT image;
[0049] A(·) represents a generator that transforms sparse angular CT images into the full-angle CT image domain;
[0050] B(·) represents a generator that transforms a full-angle CT image to a sparse-angle CT image domain;
[0051] A(s) represents the result of transforming the sparse angle CT image s into the full angle CT image domain through generator A;
[0052] B(f) represents the result of transforming the full-angle CT image f into the sparse-angle CT image domain through generator B.
[0053] In a preferred embodiment, the total loss of the loss function is a weighted sum of the adversarial loss, the cycle consistency loss, and the reconstruction loss, calculated as follows:
[0054] L total =L GAN (A,D F )+L GAN (B,D S ) +λ1L cycle +λ2L rec (7)
[0055] in,
[0056] L total Indicates the total loss;
[0057] L GAN (A,D F () represents generator A and discriminator D F Adversarial loss between angles is used to optimize image transformation from sparse angles to full angles;
[0058] L GAN (B,D S () represents generator B and discriminator D S The adversarial loss between them is used to optimize the image transformation from full angle to sparse angle;
[0059] L cycle This represents the loss of cycle consistency.
[0060] L rec Indicates the losses incurred during reconstruction;
[0061] λ1 and λ2 are weighting coefficients used to balance the contributions of different loss terms;
[0062] A represents a generator that transforms sparse angular CT images into the full-angle CT image domain;
[0063] B represents the generator that transforms full-angle CT images into the sparse-angle CT image domain;
[0064] D F A discriminator representing the full-angle image domain is used to distinguish between real full-angle CT images and images generated by generator A;
[0065] D S A discriminator representing the sparse angular image domain is used to distinguish between real sparse angular CT images and images generated by generator B.
[0066] This application also discloses an apparatus for removing artifacts from sparse angle CT images, comprising:
[0067] The data preprocessing module is used to acquire sparse angle CT images and full-angle CT images, and preprocess them to obtain preprocessed sparse angle CT images and full-angle CT images.
[0068] The network construction module is used to divide the preprocessed sparse angle CT image and full angle CT image into several training three-dimensional sub-blocks; and to construct a block-based sparse angle CT image artifact removal network, which includes two pairs of generators and discriminators, one pair of which is used to transform the training three-dimensional sub-blocks from the sparse angle image domain to the full angle image domain, and the other pair is used to perform the reverse process.
[0069] The network training module is used to train the network using the training 3D sub-blocks. During the training process, a loss function including adversarial loss, cycle consistency loss and reconstruction loss is used to optimize the network to obtain the trained network.
[0070] The image processing module is used to preprocess the sparse angle CT image to be processed, divide it into several three-dimensional sub-blocks of the same size as the three-dimensional sub-blocks according to the spatial position order, and input the three-dimensional sub-blocks to be processed into the trained network in sequence to obtain the three-dimensional sub-blocks after removing artifacts.
[0071] The image synthesis module is used to merge the artifact-removed 3D sub-blocks to form an artifact-removed CT image.
[0072] The embodiments of this application have the following technical effects:
[0073] This application proposes an innovative method for artifact removal in sparse-angle CT images, the core of which lies in fully utilizing the three-dimensional information of CT images and advanced deep learning techniques. This method mainly comprises three key points, each leading to significant technical improvements.
[0074] First, this application employs a three-dimensional sub-block approach, rather than a two-dimensional slice approach, to perform artifact removal on sparse angular CT images. This innovation fully utilizes the three-dimensional information of CT images. Traditional methods typically treat CT images as a series of independent two-dimensional slices, ignoring the spatial relationships between slices. This application, however, effectively captures the structural information of CT images in various directions by processing three-dimensional sub-blocks containing multiple adjacent slices. The technical effects of this method are significant: the processed images exhibit better structural coherence in the transverse, sagittal, and coronal planes, greatly improving the restoration of image details. Particularly in the sagittal and coronal planes, this method effectively avoids the discontinuities common in traditional methods, making the reconstructed three-dimensional images more realistic and reliable. Furthermore, by simultaneously considering the information of multiple adjacent slices, this method can more accurately distinguish between real structures and artifacts in the image, thereby achieving more precise artifact removal.
[0075] Secondly, this application specifically designs a volume-based network structure for the task of artifact removal in sparse angular CT images. The core of this network lies in its volume attention-guided residual module, which effectively captures cross-dimensional correlations between different dimensions, significantly enhancing the network's expressive power. Specifically, this module processes the channel, depth, height, and width dimensions through four branches, capturing the complex interactions between them. This design brings several technical benefits: First, it greatly improves the network's understanding and processing of three-dimensional structural information, enabling the network to more accurately identify and remove complex artifact patterns; second, the network's multi-branch structure and residual connections not only enhance feature extraction capabilities but also make the training process more stable and efficient by allowing direct gradient transfer; finally, this design makes the network more adaptable when processing CT images of different locations or with different scanning parameters.
[0076] Finally, this application employs the LPIPS metric in calculating the cycle consistency loss, proposing the concept of cycle consistency perceptual loss. This loss function is more in line with the characteristics of the human visual system, effectively avoiding the generation of overly smoothed images. Traditional pixel-level loss functions often lead to the loss of image details, while the LPIPS metric, by considering the similarity of high-level features, can better preserve the texture and detail information of the image. The technical effects of this innovation are significant: First, the processed images are visually more natural and realistic, with richer details, which is crucial for medical diagnosis; second, this method achieves a better balance between artifact removal and preservation of important anatomical structural details, improving the clinical applicability of the processing results; finally, because it is more in line with human visual perception, images processed by this method are easier for doctors to accept and interpret. In addition, this application significantly optimizes processing speed by processing in units of three-dimensional sub-blocks. Two-dimensional processing methods require frequent loading and processing of each tomographic image, while this application reduces the frequency of data loading by processing three-dimensional sub-blocks containing multiple adjacent slices. Especially in large-scale CT data processing scenarios, this improvement significantly enhances processing efficiency. Meanwhile, this application considers the pressure on video memory when processing 3D images, processing 3D sub-blocks instead of the entire 3D image. This makes it possible to process larger-sized 3D data under the same hardware conditions, effectively alleviating the video memory pressure in high-resolution CT image processing. Furthermore, the strategy for sampling 3D sub-blocks during training in this application is to sample sub-blocks at random locations within the entire 3D image each round, rather than manually cutting sub-blocks when creating the dataset. This increases the diversity of training samples, allowing the network to learn more information. In summary, especially in large-scale CT data processing scenarios, these improvements significantly enhance processing efficiency and capabilities, enabling efficient CT image artifact removal without increasing hardware costs.
[0077] In summary, these three key aspects not only individually bring significant technological advancements, but their synergistic effect also produces a comprehensive result that transcends the impact of any single technique. The 3D sub-block processing method provides the network with rich spatial information, and the specially designed network structure fully utilizes this information through its volume attention-guided residual modules, accurately capturing the correlations between different dimensions. This structure complements the cyclic consistency perceptual loss, jointly ensuring the preservation of subtle anatomical details while removing artifacts. In particular, the LPIPS metric, which considers human visual perception when evaluating image quality, combined with the 3D sub-block processing method, maintains a high degree of perceptual consistency throughout the entire 3D space.
[0078] The synergistic effect is significant: for a CT image containing 150 slices, the processing time was reduced from 18 seconds (using U-former as an example) to just 3 seconds, a six-fold increase in efficiency, while simultaneously achieving unprecedented image quality improvements. This combination of high efficiency and high quality makes this method highly valuable in clinical and research scenarios requiring rapid processing of large numbers of sparse-angle CT images. More notably, the advantages of this method increase exponentially with the amount of data to be processed and the number of slices. This is because the 3D processing method reduces the need for switching between different dimensions, the efficient design of the network structure reduces computational complexity, and the use of perceptual loss ensures excellent image quality even during high-speed processing. This scalability makes this method ideal for large-scale CT image data processing, bringing revolutionary progress to the field of medical imaging.
[0079] The specification of this application contains numerous technical features distributed across various technical solutions. Listing all possible combinations of these technical features (i.e., technical solutions) would make the specification excessively lengthy. To avoid this problem, the various technical features disclosed in the above-described invention, the various technical features disclosed in the following embodiments and examples, and the various technical features disclosed in the accompanying drawings can be freely combined to form various new technical solutions (all of which are considered to have been described in this specification), unless such a combination of technical features is technically infeasible. For example, one example discloses feature A+B+C, and another example discloses feature A+B+D+E. Features C and D are equivalent technical means that serve the same function, and technically only one needs to be used; they cannot be used simultaneously. Feature E can technically be combined with feature C. Therefore, the solution A+B+C+D should not be considered as described because it is technically infeasible, while the solution A+B+C+E should be considered as described. Attached Figure Description
[0080] Figure 1 This is a schematic diagram of the overall process of the method for removing artifacts in sparse angle CT images according to this application.
[0081] Figure 2 This is a generator structure diagram of a block-based sparse angle CT image artifact removal network in the method for sparse angle CT image artifact removal according to the first embodiment of this application.
[0082] Figure 3 This is a schematic diagram illustrating the training process and loss function of a block-based sparse angle CT image artifact removal network in the method for sparse angle CT image artifact removal according to the first embodiment of this application.
[0083] Figure 4 This is a schematic diagram illustrating the effect of a method for removing artifacts from sparse angle CT images according to the first embodiment of this application.
[0084] Figure 5 This is a flowchart illustrating a method for removing artifacts from sparse-angle CT images according to a first embodiment of this application.
[0085] Figure 6 This is a schematic diagram of the structure of a system for removing artifacts from sparse angle CT images according to the second embodiment of this application. Detailed Implementation
[0086] In the following description, many technical details are presented to help the reader better understand this application. However, those skilled in the art will understand that the technical solutions claimed in this application can be implemented even without these technical details and various variations and modifications based on the following embodiments.
[0087] Explanation of some concepts:
[0088] Sparse view: In computed tomography (CT) imaging, sparse view is a technique that reduces radiation dose by decreasing the number of projection angles. Compared to traditional full-angle scans, sparse view imaging acquires less projection data, making it more prone to artifacts or noise in the image. This technique is typically used in scenarios where it is necessary to reduce patient exposure to X-ray radiation, such as pediatric imaging or patients requiring frequent scans. The number of projection angles in a sparse view is usually less than one-third of the number of projection angles in a full-angle scan; a number of angles less than 60 is considered extremely sparse.
[0089] Full view: Full view refers to scanning using the entire range of projection angles in CT imaging. It typically covers a large number of projection angles to provide more comprehensive image data and higher image quality. Full view scans generate more complete data, thus reducing artifacts and providing clearer images, making them suitable for medical diagnoses that require high precision and detailed anatomical information.
[0090] Artifacts: In CT images, artifacts or erroneous information caused by incomplete data or problems with the reconstruction algorithm can interfere with medical diagnosis.
[0091] 3D sub-blocks: CT images are divided into 3D blocks containing multiple adjacent slices for processing, so as to make full use of the 3D information of the image and improve the effect of artifact removal and image quality.
[0092] Two-dimensional tomography: Traditional CT image processing methods treat CT images as independent two-dimensional slices, which may ignore the spatial correlation between slices.
[0093] Volume attention-guided residual modules: These modules introduce attention mechanisms and residual connections into the network structure, which can capture cross-dimensional correlations between different dimensions, thereby enhancing the network's expressive power and training stability.
[0094] LPIPS (Learned Perceptual Image Patch Similarity) is a deep learning-based image similarity metric that better aligns with the human visual system and avoids generating overly smooth images.
[0095] Cyclic consistency loss: In deep learning, this loss measures the difference between two inverse transformations of an image to ensure the reversibility of the transformation process and the consistency of the image content.
[0096] Video memory pressure: refers to the usage of the graphics card's memory (VRAM) during image processing. Dividing a 3D image into sub-blocks for processing can reduce video memory pressure, allowing the graphics card to process larger amounts of data.
[0097] Deep learning: an artificial intelligence technology that processes data and recognizes patterns by simulating the structure and function of the human brain's neural network, and is widely used in fields such as image processing and natural language processing.
[0098] CT (Computed Tomography): Computed tomography uses a computer to process X-ray projections from multiple angles to generate three-dimensional images of internal structures. It is widely used in medical diagnosis and industrial inspection.
[0099] Generative Adversarial Networks (GANs): A deep learning model architecture consisting of two adversarial neural networks, a generator and a discriminator. In this method, two pairs of GANs are used to achieve bidirectional conversion between sparse angular CT images and full-angle CT images.
[0100] The following is a brief summary of some of the innovative aspects of this application:
[0101] First, this application has a wide range of applications and scenarios in medical image processing, particularly in artifact removal and image quality improvement of CT images. For example:
[0102] In emergency medicine, this technology has particularly significant applications. Rapid assessment of trauma patients requires quick CT scans, and while sparse-angle CT can speed up the scan, it can introduce artifacts. This technology removes these artifacts, ensuring sufficiently high image quality to support emergency medical decisions, helping doctors quickly identify internal injuries and develop treatment plans, thus improving treatment efficiency and effectiveness.
[0103] Some medical institutions may use older CT scanners, which may produce images of lower quality than the latest equipment. This technology can upscale the relatively low-quality images from older equipment to near-newer levels, thereby extending the lifespan of older equipment and saving medical resources. This is especially important for hospitals and clinics with limited budgets, enabling them to improve the quality of healthcare services without replacing the equipment.
[0104] For certain patient groups, such as children, pregnant women, or patients with chronic diseases who require frequent CT scans, reducing radiation dose is particularly important. This technology can provide high-quality images with low-dose scans, making it especially suitable for these patient groups. By reducing radiation exposure, patient health is protected while ensuring diagnostic accuracy.
[0105] CT scans are a crucial tool in tumor detection and monitoring. Using the technical solution presented in this application, high-quality CT images can be obtained while reducing radiation dose, facilitating early tumor detection and follow-up monitoring, and ensuring timely and accurate treatment.
[0106] CT imaging plays a crucial role in the diagnosis and assessment of cardiovascular diseases, such as coronary artery disease, aneurysms, and cardiac structural abnormalities. The technical solution presented in this application can provide clear images of the heart and blood vessels, helping physicians assess the condition and develop treatment plans, thereby improving diagnostic accuracy and treatment outcomes.
[0107] Low-dose CT scans are widely used for early screening of lung cancer. By applying the technical solution of this application, image quality can be improved while reducing radiation dose, thereby screening for lung diseases more effectively, which is of great significance, especially in the early detection of lung cancer.
[0108] In complex surgical procedures, precise preoperative planning and intraoperative navigation are crucial. High-quality CT images help surgeons understand the patient's anatomy in detail and develop the best surgical plan. Using the technology described in this application, clearer and more accurate CT images can be obtained, assisting surgeons in preoperative planning and intraoperative navigation, reducing surgical risks, and improving surgical success rates.
[0109] The rise of telemedicine has enabled patients in remote areas to enjoy high-quality medical services. The transmission and remote diagnosis of CT images are crucial components of telemedicine. The technical solution presented in this application effectively removes artifacts from sparse angle CT images, improving image quality and ensuring the accuracy and reliability of remote diagnosis.
[0110] Medical researchers and educators require a large number of high-quality CT images for research and teaching. The technical solution of this application can provide clear and accurate imaging data to support the work of researchers and educators, thereby promoting the development of medical research and improving the quality of teaching.
[0111] Pediatric patients are highly sensitive to radiation, therefore it is necessary to minimize radiation dose. The technical solution of this application can provide high-quality CT images with low-dose scanning, protecting children's health while ensuring diagnostic accuracy.
[0112] In emergency trauma assessment, CT scans are a key tool for rapidly evaluating internal damage. High-quality images help physicians make quick diagnostic and treatment decisions, improving the success rate of trauma patients' treatment.
[0113] The above-mentioned application areas and scenarios exemplify the wide and important application of the technical solution of this application in actual medical image processing, providing important support for improving the quality of diagnosis and treatment.
[0114] Based on the technical concept of this application, a block-based sparse angular CT image artifact removal network is proposed. The overall process is described in [link to application]. Figure 1 The network is trained using sparse angled CT images and full-angle CT images, employing a composite loss function that includes adversarial loss, cyclic consistency sensing loss, and reconstruction loss.
[0115] Considering that CT images are inherently three-dimensional, with content similarity and structural continuity between adjacent layers, these characteristics help the network distinguish structural details and artifacts in the image. Therefore, this application no longer limits itself to processing individual slices, but instead divides the CT image into three-dimensional sub-blocks for input into the network for processing. This method can effectively preserve information in the three-dimensional structure of the CT image, thereby helping the network to more accurately distinguish image details and artifacts.
[0116] This artifact removal network consists of two pairs of generators and discriminators with identical structures but non-shared weights. This design helps ensure the reversibility of the image transformation process and the consistency of the content before and after the transformation, while also guaranteeing image fidelity and the stability of the training process. One pair of generators and discriminators is responsible for transforming the image from a sparse angular image domain (S) to a full-angle image domain (F), while the other pair performs the reverse process.
[0117] The structure of the generator is as follows Figure 2 As shown, it comprises four key modules: an encoder module, a dilated spatial convolutional pooling pyramid module, a volume attention-guided residual module, and a decoder module. The dilated spatial convolutional pooling pyramid module (e.g.) Figure 2 (c) shows that dilated convolutions with different dilation rates are used to capture multi-scale image features, giving the network receptive fields of different sizes, thereby improving feature extraction capabilities and promoting the network's comprehensive understanding of images. Volume attention-guided residual modules (such as...) Figure 2 (B) contains four branches, each capturing the dimensional interaction weights between combinations of three dimensions: channel, depth, height, and width, and performing an averaging operation. In addition to the attention-weighted operation, the input tensor undergoes a convolutional branch for feature extraction, which is then directly passed through residual connections. This multi-branch structure not only enhances the network's expressive power but also improves training stability through direct gradient propagation.
[0118] The network optimization employs a composite loss function consisting of adversarial loss, cyclic consistency-aware loss, and reconstruction loss. Theoretically, a CT image should be restored to its original state after two inverse transformations. Taking a sparse angular CT image sub-block *s* as an example, generator A first transforms *s* from the S-domain to the F-domain to obtain A(s), and then generator B transforms A(s) back to the S-domain to obtain B(A(s)). Cyclic consistency-aware loss is used to measure the difference between B(A(s)) and the original image *s*, ensuring the reversibility of the transformation process and the consistency of image content. Furthermore, the intermediate result A(s) is compared with the corresponding full-angle CT image label to calculate the reconstruction loss. All these loss calculations use learned perceptual image patch similarity (LPIPS) to measure the differences between images.
[0119] The formulas for calculating each part of the loss function are as follows:
[0120] L GAN (A,D F )=E[log D F (f)]+E[log(1-D F (A(s)))] (3)
[0121] L GAN (B,D S )=E[log D S (s)]+E[log(1-D S (B(f)))] (4)
[0122] L cycle =E[||B(A(s))-s|| LPIPS ]+E[||A(B(f))-f|| LPIPS (5)
[0123] L rec =E[||A(s)-f|| LPIPS ]+E[||B(f)-s|| LPIPS (6)
[0124] In these formulas, L GAN (A,D F ) and L GAN (B,D S This constitutes an adversarial loss, where A and B are the generators for S→F and F→S, respectively, and D... F and D S This is the corresponding discriminator. s and f are samples from the S and F domains, respectively. L cycle Calculate the cycle consistency-aware loss, L rec Calculate the reconstruction loss, ||·|| LPIPS This represents the LPIPS metric.
[0125] L total =L GAN (A,D F )+L GAN (B,D S ) +λ1L cycle +λ2L rec (7)
[0126] During network training, each iteration samples a fixed-size image sub-block at a random location as input. After training, the sparse angular CT image to be processed is divided into sub-blocks of the same size as those sampled during training, according to spatial position. These sub-blocks are then sequentially input into generator A for processing. Finally, the processed sub-blocks are merged to form the final optimized result.
[0127] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0128] The first embodiment of this application relates to a method for artifact removal in sparse angle CT images, the process of which is as follows: Figure 5 As shown, the method includes the following steps:
[0129] Step 10: Acquire sparse angle CT images and full-angle CT images, and preprocess them to obtain preprocessed sparse angle CT images and full-angle CT images;
[0130] Step 20: Divide the preprocessed sparse angle CT image and full angle CT image into several training 3D sub-blocks; and construct a block-based sparse angle CT image artifact removal network, the network including two pairs of generators and discriminators, one pair for transforming the training 3D sub-blocks from the sparse angle image domain to the full angle image domain, and the other pair for performing the reverse process;
[0131] Step 30: Train the network using the training 3D sub-blocks. During the training process, optimize the network by using adversarial loss, cycle consistency loss, and reconstruction loss to obtain the trained network.
[0132] Step 40: After preprocessing the sparse angle CT image to be processed, divide it into several three-dimensional sub-blocks of the same size as the three-dimensional sub-blocks according to spatial position. Input the three-dimensional sub-blocks to be processed into the trained network in sequence to obtain three-dimensional sub-blocks after artifact removal. Then, merge the three-dimensional sub-blocks after artifact removal to form a CT image after artifact removal.
[0133] The following provides further explanation of each step.
[0134] In step 10, firstly, sparse-angle CT images and full-angle CT images are acquired using a CT scanning device. Sparse-angle CT images reduce radiation dose by decreasing the projection angle, thereby reducing potential harm to the patient. However, this method may result in artifacts in the images. Full-angle CT images, on the other hand, are acquired using a conventional full-angle scanning method, offering higher image quality and integrity.
[0135] Next, the acquired sparse angle CT images and full-angle CT images are preprocessed. Preprocessing steps include resampling, cropping, and grayscale normalization. Resampling adjusts the spatial resolution of the CT images to ensure different images have the same resolution. This step is crucial because it ensures that the spatial features of the images can be consistently processed and analyzed in subsequent processing. Cropping cropps the CT images to a specific size to fit the network's input requirements. Typically, CT images may be too large or irregular in size; cropping adjusts the images to a uniform size, improving processing efficiency and effectiveness. Grayscale normalization standardizes the grayscale value range of the image, usually adjusting it to between [-1,1] or [0,1]. The grayscale values of CT images represent the density of different tissues; normalization adjusts the grayscale values of different images to the same range. This step improves the operability and stability of the image data in the network, avoiding training instability caused by different grayscale value ranges.
[0136] For example, suppose the CT image processed in this step is a scan of a walnut. A sparse-angle CT image might be acquired using 40 projection angles, while a full-angle CT image is acquired using 1200 projection angles. The original image size is (512, 512, 512), and this resolution is maintained during resampling. Then, the image is cropped to (360, 360, 150) to ensure all images are the same size. Next, grayscale normalization is performed, adjusting the grayscale values to the range [-1, 1].
[0137] These preprocessing steps ensure consistency and adaptability of sparse-angle and full-angle CT images in subsequent processing and analysis, providing a solid foundation for subsequent artifact removal and image reconstruction. These techniques, combined, form a complete preprocessing workflow aimed at improving the efficiency and quality of CT image processing.
[0138] In step 20, the preprocessed sparse angle CT images and full-angle CT images are first divided into several training 3D sub-blocks. This process involves dividing the CT images into small blocks according to their spatial location, with each block containing multiple adjacent slices, thus forming 3D data blocks. This division method ensures that each sub-block not only contains information from a single slice but also captures the spatial relationships between adjacent slices, providing rich 3D structural information for subsequent network training.
[0139] Next, a block-based sparse angular CT image artifact removal network is constructed. The core of this network consists of two pairs of generators and discriminators, a design intended to achieve bidirectional transformation between image domains. Specifically, one pair of generators and discriminators is responsible for transforming the training 3D sub-blocks from the sparse angular image domain to the full angular image domain, while the other pair performs the reverse process, transforming the 3D sub-blocks from the full angular image domain back to the sparse angular image domain. The generator's role is to generate images in the target domain, while the discriminator is used to distinguish the generated images from the real images.
[0140] This bidirectional conversion design enables the network to learn and understand the fundamental differences between sparse-angle and full-angle CT images, thereby progressively optimizing the image conversion effect during training. By simultaneously learning features from both image domains, the network can more effectively remove artifacts while preserving important details in the images.
[0141] In summary, this step, by dividing CT images into three-dimensional sub-blocks and utilizing bidirectional generative adversarial networks to achieve the conversion between sparse angular and full-angle image domains, provides a technical basis for removing artifacts and improving image quality.
[0142] In summary, step 20 demonstrates the innovative thinking behind this method in data processing and network architecture design. By combining 3D data processing with an advanced GAN architecture, this method not only improves artifact removal but also enhances the network's learning ability and generalization, providing a new paradigm for CT image reconstruction.
[0143] In step 30, the network is first trained using the 3D training sub-blocks defined in step 20. During training, these 3D sub-blocks are sequentially input into the network. To optimize network performance and ensure effective artifact removal and image detail preservation, three main loss functions are employed during training: adversarial loss, cycle consistency loss, and reconstruction loss.
[0144] Adversarial loss stems from the training concept of Generative Adversarial Networks (GANs), which improves the quality of generated images through adversarial training between the generator and the discriminator. Specifically, the generator attempts to generate realistic CT images, while the discriminator tries to distinguish between the generated images and real images. This process iterates continuously, making the images generated by the generator increasingly realistic, thereby improving the network's artifact removal capabilities.
[0145] Cyclic consistency loss ensures that the transformation between the sparse angle and full-angle image domains is reversible. That is, after a sparse angle CT image is converted to a full-angle CT image by a generator, and then converted back to a sparse angle CT image by another generator, the resulting image should be as consistent as possible with the original sparse angle CT image. Similarly, after a full-angle CT image is converted to a sparse angle CT image by a generator, and then converted back to a full-angle CT image, the resulting image should be as consistent as possible with the original full-angle CT image. This loss function ensures the stability and consistency of the transformation process by measuring the difference in image transformation between the two domains.
[0146] The reconstruction loss measures the overall structural similarity between the generated and target images. It ensures that the network generates high-quality images while preserving as much detail and structural information as possible from the original images. Reconstruction loss typically employs advanced similarity metrics rather than simple pixel-level differences, thus better preserving the image's detailed features.
[0147] By comprehensively applying these three loss functions, the network is optimized, gradually improving its artifact removal and image reconstruction capabilities. After multiple iterations of training, a fully trained network is finally obtained, capable of generating high-quality artifact-free CT images when given new sparse angle CT images as input.
[0148] It should be noted that using 3D sub-blocks for network training in this step is an innovative approach. This method not only utilizes the 3D characteristics of CT images but also cleverly solves the computational resource problem of training with large-scale 3D data. By randomly sampling 3D sub-blocks instead of using the entire CT volume, this method achieves the following goals: Increased diversity of training samples: Random sampling generates a large number of different 3D sub-blocks, greatly increasing the diversity of training data and helping to improve the network's generalization ability. Balance between computational efficiency and learning effectiveness: Using sub-blocks instead of the entire CT volume significantly reduces the computational load per iteration, enabling effective training even with limited computational resources. Implicit data augmentation: Different sub-blocks may contain varying degrees of artifacts and anatomical structures; this diversity itself is a form of implicit data augmentation, helping the network learn more robust features.
[0149] Secondly, this method employs a multi-component composite loss function, another key innovation in the training process. This loss function comprises three main parts: adversarial loss, cycle consistency loss, and reconstruction loss. Specifically, adversarial loss: This originates from the idea of Generative Adversarial Networks (GANs). Through adversarial training between the generator and discriminator, the network can generate images that are more realistic and better conform to the distribution of the target domain. In CT image reconstruction, this means that the generated full-angle CT images should not only be free of artifacts but also retain the texture and detail features of the real CT image. Cycle consistency loss: This is a unique innovation of this method. By ensuring that the transformation between the two domains is reversible (sparse angle → full angle → sparse angle, and the reverse process), the network is forced to learn to preserve the essential features of the image. This not only improves the accuracy of reconstruction but also helps to preserve subtle but important structural details that may be mistakenly deleted during artifact removal. Reconstruction loss: This ensures the similarity of the generated image to the target image in terms of overall structure. It is worth noting that this method may employ advanced similarity metrics (such as perceptual loss) rather than simple pixel-level differences, which is more in line with the characteristics of the human visual system.
[0150] The combination of these three losses forms a multi-objective optimization problem, with each loss targeting a different aspect of the reconstruction process. This complex optimization strategy requires the network to find a balance under multiple constraints to produce high-quality, realistic, and detail-preserving reconstruction results.
[0151] In step 40, the sparse angular CT image to be processed is first preprocessed. This includes adjusting the spatial resolution, cropping the image, and normalizing the grayscale values to ensure image consistency and adaptability. The preprocessed image is then segmented into several 3D sub-blocks of the same size as the training sub-blocks according to their spatial location. This segmentation process divides the entire CT image into multiple small 3D blocks, each containing multiple adjacent slices, thus preserving the 3D structural information of the image.
[0152] Next, these unprocessed 3D sub-blocks are sequentially fed into the pre-trained network for processing. The network performs artifact removal on each input 3D sub-block, generating artifact-free 3D sub-blocks. This step utilizes the previously trained network model to ensure that each sub-block retains its important structural information while removing artifacts.
[0153] Finally, all the artifact-removed 3D sub-blocks are merged to form a complete artifact-removed CT image. This merging process requires reassembling the processed sub-blocks according to their original spatial positions to reconstruct a complete 3D CT image. This method ensures the coherence and integrity of the entire CT image after processing, providing technical support for the final image reconstruction.
[0154] Step 40 addresses the challenges of large-scale 3D data processing while ensuring high quality and consistency of the results. The innovation of this step lies in the following aspects:
[0155] First, the sparse angular CT images to be processed are segmented into 3D sub-blocks according to their spatial location, demonstrating a deep understanding of 3D data structures. Unlike random sampling during training, ordered segmentation is used here to ensure that the entire CT volume is processed completely and systematically. This method has several significant advantages: Preserving spatial continuity: Ordered segmentation ensures that the spatial relationships between adjacent sub-blocks are preserved, which is crucial for maintaining the coherence of the overall image. Adapting to large-scale data: By processing in blocks, this method cleverly overcomes GPU memory limitations, enabling efficient processing even of large CT datasets. Parallel processing potential: This block strategy enables parallel computing, significantly improving processing speed.
[0156] Secondly, the 3D sub-blocks to be processed are sequentially input into the trained network for processing. This process demonstrates the flexibility and adaptability of the method: Consistent processing: Each sub-block undergoes the same network processing, ensuring consistency in the processing of the entire CT volume. Local-global balancing: When processing each local sub-block, the network actually operates based on knowledge learned within a larger spatial context, which helps maintain the consistency of the global structure while removing local artifacts. Dynamic adaptation: Although the network is trained on sub-blocks of fixed size, this processing method allows the network to dynamically adapt to CT data of different sizes and shapes.
[0157] Finally, the artifact-free 3D sub-blocks are merged to form the final artifact-free CT image.
[0158] In summary, according to the above-described process of this embodiment, firstly, this method innovatively employs three-dimensional sub-blocks as processing units, rather than traditional two-dimensional slices. This method deeply recognizes that CT images are essentially three-dimensional data, with rich spatial correlation information existing between adjacent slices. By processing multiple adjacent slices simultaneously, the network can capture complex three-dimensional structural features, which is crucial for accurately distinguishing between real anatomical structures and artifacts. This three-dimensional processing strategy not only improves the accuracy of artifact removal but also significantly improves the coherence of reconstructed images across different sections, solving the common problem of tomographic discontinuity in traditional methods.
[0159] Secondly, this method employs a unique bidirectional transformation mechanism comprising two pairs of generators and discriminators. Inspired by the field of image style transfer, this design is innovative in its application to CT image processing. By simultaneously learning the forward mapping from sparse angles to full angles and the reverse mapping from full angles to sparse angles, the network gains a more comprehensive understanding of the relationship between the two image domains. This bidirectional learning not only enhances the network's robustness but also lays the foundation for subsequent cycle consistency constraints.
[0160] Another highlight of the method is the use of a composite loss function. Specifically, it introduces a perceptual-based cycle consistency loss, employing the LPIPS (Learned Perceptual Image Patch Similarity) metric to evaluate image similarity. This approach goes beyond traditional pixel-level comparison, better aligning with the perceptual characteristics of the human visual system. By considering the similarity of high-level features, this method can remove artifacts while preserving subtle texture details, avoiding over-smoothing, which is crucial for medical diagnosis.
[0161] Finally, this method employs a clever sliding window strategy in practical applications. By segmenting large CT datasets into overlapping 3D sub-blocks, processing and reassembling them block by block, the method effectively balances computational efficiency and processing quality. This strategy not only addresses GPU memory limitations but also ensures the continuity of processing results through smooth transitions in overlapping regions, providing a feasible solution for the efficient processing of large-scale CT data.
[0162] This approach, by integrating 3D processing, bidirectional learning, perceptual loss, and efficient implementation strategies, forms a unique and powerful framework that opens up new avenues for high-quality reconstruction of sparse angular CT images.
[0163] A block-based sparse angular CT image artifact removal network
[0164] The following section provides a more detailed explanation of the specific structure of the block-based sparse angular CT image artifact removal network.
[0165] Optionally, the generator of the block-based sparse angular CT image artifact removal network includes an encoder module, a dilated spatial convolutional pooling pyramid module, multiple volume attention-guided residual modules, and a decoder module connected in series. The encoder module consists of multiple 3D convolutional layers connected in series, with an instance normalization layer and a ReLU activation function added sequentially after each 3D convolutional layer to extract features from the training 3D sub-blocks. The dilated spatial convolutional pooling pyramid module receives the output of the encoder module and includes multiple dilated convolutional branches with different dilation rates and a global average pooling branch. The outputs of each branch are concatenated along the channel direction to capture multi-scale features. The volume attention... The force-guided residual module includes a volume attention weighted branch, a convolutional branch, and a residual connection branch. The volume attention weighted branch pools the input features in four dimensions: channel, depth, height, and width. After processing by convolutional layers, instance normalization layers, and a sigmoid activation function, it captures cross-dimensional correlations between different dimensions and multiplies them with the input features. The convolutional branch includes multiple three-dimensional convolutional layers. The residual connection branch directly passes the input features. The decoder module consists of multiple three-dimensional transposed convolutional layers connected in series. Each three-dimensional transposed convolutional layer is followed by an instance normalization layer and an activation function to reconstruct the processed features into an output image with the same size as the input.
[0166] Optionally, the encoder module consists of three cascaded three-dimensional convolutional layers, with each three-dimensional convolutional layer followed by an instance normalization layer and a ReLU activation function. The first three-dimensional convolutional layer has a kernel size of 7×7×7 and a stride of 1, while the second and third three-dimensional convolutional layers have kernel sizes of 3×3×3 and a stride of 2.
[0167] Optionally, the dilated spatial convolutional pooling pyramid module includes a 1×1×1 convolutional branch, three 3×3×3 dilated convolutional branches with different dilation rates, and a global average pooling branch, with the outputs of each branch spliced along the channel direction.
[0168] Optionally, the volume attention-guided residual module includes: a volume attention weighted branch, which pools the input features in four dimensions (channel, depth, height, and width), then captures the cross-dimensional correlation between each of the three dimensions through convolutional layers, instance normalization layers, and a sigmoid activation function, and then multiplies it with the input features; a convolutional branch, which includes two cascaded 3×3×3 three-dimensional convolutional layers with a stride of 1; and a residual connection branch, which is used to directly pass the input features.
[0169] Optionally, the decoder module consists of three cascaded 3D transposed convolutional layers, wherein the kernel size of the first and second 3D transposed convolutional layers is 3×3×3 with a stride of 2, and each layer is followed by an instance normalization layer and a ReLU activation function. The kernel size of the third 3D transposed convolutional layer is 7×7×7 with a stride of 1, and a Tanh activation function is added.
[0170] In other words, the generator of the block-based sparse angular CT image artifact removal network comprises an encoder module, a dilated spatial convolutional pooling pyramid module, multiple volume attention-guided residual modules, and a decoder module, all connected in series. The encoder module consists of multiple cascaded 3D convolutional layers, each followed by an instance normalization layer and a ReLU activation function to extract features from the training 3D sub-blocks. This design enables the encoder module to effectively capture and extract key feature information from CT images.
[0171] The dilated spatial convolutional pooling pyramid module receives the output from the encoder module and contains multiple dilated convolutional branches with different dilation rates and a global average pooling branch. The outputs of each branch are concatenated along the channel direction to capture multi-scale features. The dilated convolutional branches with different dilation rates can capture feature information at different scales, while the global average pooling branch provides global contextual information.
[0172] The volume attention-guided residual module comprises a volume attention weighted branch, a convolutional branch, and a residual connection branch. The volume attention weighted branch pools the input features along four dimensions (channel, depth, height, and width) separately, then processes them through convolutional layers, instance normalization layers, and a sigmoid activation function to capture cross-dimensional correlations before multiplying them with the input features. The convolutional branch includes multiple 3D convolutional layers for further processing of feature information. The residual connection branch directly passes the input features; this design preserves the original feature information and enhances the network's stability.
[0173] The decoder module consists of multiple cascaded 3D transposed convolutional layers. Each 3D transposed convolutional layer is followed by an instance normalization layer and an activation function to reconstruct the processed features into an output image of the same size as the input. The design of the decoder module enables the network to effectively reconstruct the extracted feature information back to the original size of the CT image.
[0174] Furthermore, the encoder module consists of three cascaded 3D convolutional layers, each followed by an instance normalization layer and a ReLU activation function. The first 3D convolutional layer has a 7×7×7 kernel size and a stride of 1, while the second and third 3D convolutional layers have 3×3×3 kernel sizes and a stride of 2. The dilated spatial convolutional pooling pyramid module includes a 1×1×1 convolutional branch, three 3×3×3 dilated convolutional branches with different dilation rates, and a global average pooling branch. The outputs of each branch are concatenated along the channel direction. The volume attention-guided residual module includes: a volume attention weighted branch, which pools the input features along the four dimensions of channel, depth, height, and width, then captures the cross-dimensional correlations between each of the three dimensions through convolutional layers, instance normalization layers, and a sigmoid activation function, before multiplying it with the input features; a convolutional branch, which includes two cascaded 3×3×3 3D convolutional layers with a stride of 1; and a residual connection branch, used to directly pass the input features. The decoder module consists of three cascaded 3D transposed convolutional layers. The first and second layers each have a 3×3×3 kernel size and a stride of 2, followed by an instance normalization layer and a ReLU activation function. The third layer has a 7×7×7 kernel size and a stride of 1, and incorporates a Tanh activation function. Through this detailed module design and configuration, the entire network can efficiently process and reconstruct CT images.
[0175] The aforementioned network architecture is specifically designed to address artifact removal in sparse-angle CT images, adapting to the unique requirements of CT image processing. This network architecture demonstrates a profound understanding of CT image characteristics and cutting-edge deep learning technologies. It not only effectively processes 3D data but also enhances the network's ability to capture multi-scale features and complex spatial relationships through innovative module design. This carefully designed structure is likely to better preserve important anatomical details while removing CT image artifacts, providing a powerful and flexible solution for high-quality CT image reconstruction.
[0176] In this embodiment, the encoder module employs three cascaded three-dimensional convolutional layers, each with its own unique purpose and function.
[0177] The first 3D convolutional layer uses a large 7×7×7 convolutional kernel, a key innovation. This large kernel captures a wider range of spatial information in the initial stage, crucial for understanding large-scale structures in CT images. A stride of 1 ensures that not too much spatial detail is lost in the first layer. This design allows the network to acquire rich contextual information from the outset, laying a solid foundation for subsequent feature extraction.
[0178] The second and third 3D convolutional layers employ smaller 3×3×3 convolutional kernels. This is based on a key finding in deep learning: stacking multiple small convolutional kernels can achieve the same receptive field as large kernels while significantly reducing the number of parameters and improving computational efficiency. The stride of these two layers is set to 2, achieving downsampling of the feature maps. This progressive downsampling strategy not only reduces the spatial dimensionality of the feature maps but also gradually increases the number of channels, enabling the network to learn more abstract and higher-level feature representations.
[0179] Each convolutional layer is followed by an instance normalization layer and a ReLU activation function. This combination is widely used in modern deep learning networks for several reasons: the instance normalization layer normalizes the features of each instance (in this case, each CT image). This operation is particularly effective for processing medical images, as different CT scans may have different brightness and contrast. Instance normalization reduces the impact of such variations, allowing the network to focus more on learning the structural features of the image rather than intensity changes. The ReLU activation function introduces non-linearity, which is crucial for enhancing the network's expressive power. It allows the network to learn complex non-linear mappings, which is especially important when dealing with complex artifact patterns in CT images. Furthermore, the simplicity of ReLU makes gradient calculation efficient, helping to accelerate the network training process. This three-layer structure embodies a balanced approach: the large convolutional kernel in the first layer captures broad spatial information, while the smaller convolutional kernels and downsampling in the latter two layers progressively extract more abstract features. This progressive feature extraction strategy is well-suited for processing multi-scale information in CT images, from large-scale anatomical structures to small-scale texture details.
[0180] In summary, this encoder design demonstrates a profound understanding of 3D medical image processing. It not only effectively extracts key features from CT images but also achieves a good balance between computational efficiency and feature extraction capabilities through reasonable parameter settings and structural arrangement. This design lays a solid foundation for subsequent artifact removal tasks and is expected to significantly improve the quality of reconstructed CT images.
[0181] The hollow spatial convolutional pooling pyramid module contains five parallel branches, each with its specific function and advantages. This multi-branch structure allows the network to capture information from different spatial scales and receptive fields simultaneously, which is crucial for processing complex structures and artifact patterns in CT images.
[0182] The 1×1×1 convolution branch is a clever design element in this module. This small-sized convolution is primarily used for cross-channel information integration and dimensionality reduction. In CT image processing, it can effectively fuse features from different channels, extracting more refined representations while reducing computational cost. This branch can capture pixel-level local information, providing crucial details for subsequent feature fusion.
[0183] The core innovation of this module is the three 3×3×3 dilated convolution branches with different dilation rates. The unique feature of dilated convolution (also known as dilated convolution) is its ability to significantly expand the receptive field without increasing the number of parameters or computational complexity. By using different dilation rates, the network can simultaneously focus on features at different scales. For example, a smaller dilation rate might focus on capturing detailed textures, while a larger dilation rate helps in understanding large-scale structures. This design is particularly suitable for CT images, as CT images often contain both large-scale anatomical structures and subtle lesion features.
[0184] The introduction of the global average pooling branch reflects the consideration of global contextual information. By performing average pooling on the entire feature map, this branch can capture global statistical information. In CT image processing, this global information can help the network understand the overall anatomical structure and large-scale artifact patterns, thus providing important contextual guidance during local feature extraction.
[0185] Finally, the outputs of each branch are concatenated along the channel direction. This operation preserves the unique features extracted by each branch, allowing subsequent network layers to learn how to best utilize this multi-scale information. In the task of artifact removal in CT images, this fusion can help the network simultaneously consider local details, intermediate-scale structures, and global context, thereby making more accurate judgments.
[0186] In summary, this dilated spatial convolutional pooling pyramid module not only effectively captures multi-scale features but also achieves a balance between the breadth of feature extraction and computational efficiency through its ingenious design. This structure excels in handling complex CT image artifacts, simultaneously focusing on large-scale structures and minute details, providing strong support for high-quality CT image reconstruction.
[0187] The volume attention-guided residual module is designed to process the three-dimensional characteristics of CT images. This module comprises three key branches, each with its own specific functions and technical considerations.
[0188] Regarding the volume attention weighted branch, pooling operations are first performed on the four dimensions of the input features (channels, depth, height, and width). This dimensional pooling allows the network to independently consider the feature distribution in each dimension. Subsequently, through a cascaded process of convolutional layers, instance normalization layers, and a sigmoid activation function, the network learns to capture cross-dimensional correlations between each of the three dimensions. The technical significance of this design is that it enables the network to understand and utilize the complex three-dimensional spatial relationships in CT images. For example, the network can learn that the correlation of certain features in the depth and height dimensions may be more important than in the width dimension. Finally, the resulting attention weights are multiplied by the input features, achieving adaptive adjustment of the input features.
[0189] The convolutional branch employs two cascaded 3×3×3 three-dimensional convolutional layers with a stride of 1. This design allows the network to perform local feature extraction and transformation. Using two convolutional layers instead of a single larger convolutional layer aligns with an important finding in deep learning: stacking multiple small convolutional kernels can achieve the same receptive field as a large kernel while reducing the number of parameters and improving computational efficiency. The stride of 1 ensures preservation of spatial resolution, which is crucial for retaining details in CT images.
[0190] Residual connections directly pass input features. From a technical perspective, residual connections have several important functions: First, they alleviate the vanishing gradient problem in deep networks, making training deep networks easier; second, they allow the network to learn identity mappings, which is advantageous in certain situations (e.g., when the input is already high-quality features); and finally, they provide a shortcut for the network, making it easier for information to flow between different layers.
[0191] The outputs of these three branches are eventually merged to form the module's output. This design allows the network to flexibly determine how best to utilize the attention mechanism, convolutional transformation, and original features. In CT image processing, this structure helps the network better handle complex 3D structures while maintaining sensitivity to detail, which is crucial for high-quality artifact removal and image reconstruction.
[0192] The decoder module's structure has been carefully tailored to the characteristics of 3D data. The module consists of three cascaded 3D transposed convolutional layers, each designed with specific technical considerations.
[0193] The first two 3D transposed convolutional layers use 3×3×3 kernels with a stride of 2. This configuration offers several technical advantages: First, the 3×3×3 kernel size effectively captures local spatial information while maintaining a small number of parameters. This is based on an important finding in deep learning: stacking multiple small kernels can achieve the same receptive field as a large kernel while reducing computational complexity. Second, a stride of 2 means that the spatial dimension of the feature map is doubled with each operation. This gradual upsampling strategy helps to progressively recover spatial details and avoids artifacts that may result from direct, large upsampling.
[0194] Each instance normalization layer followed by the ReLU activation function after the transposed convolutional layer also has a specific role. Instance normalization helps stabilize the feature distribution and reduce internal covariate shifts, which is particularly effective for processing CT images that may have intensity variations. The ReLU activation function introduces non-linearity, enhancing the network's expressive power and facilitating gradient propagation.
[0195] The third transposed convolutional layer employs a large 7×7×7 kernel and a stride of 1. This design has its unique technical considerations: the large kernel captures a wider range of spatial information in the final stage, helping to recover more detailed image information. The stride of 1 ensures that the last layer does not change the spatial size of the feature map, but focuses on fine-tuning the features.
[0196] The final layer uses the Tanh activation function instead of ReLU, a choice made for technical reasons: the Tanh function's output range is [-1, 1], which matches the CT image data normalized to the [-1, 1] range during preprocessing. Tanh has a large gradient near zero, which is beneficial for handling subtle changes in the image and helps restore fine image details. Furthermore, Tanh is a bounded function, which helps control the output range and avoids producing excessively large values.
[0197] This three-layer structure embodies a balanced approach: the first two layers progressively increase the spatial dimension of the feature maps using smaller convolutional kernels and a stride of 2, while the final layer finely adjusts the features using a large convolutional kernel and a stride of 1. This progressive reconstruction strategy is well-suited for processing multi-scale information in CT images, effectively reconstructing everything from large-scale structures to small-scale texture details.
[0198] It should be noted that the above descriptions are preferred embodiments of this application and are exemplary, not limiting. Besides the specific structures and steps described, this application can also employ various other implementation methods. For example, the encoder module can use convolutional layers of different numbers and sizes, and is not limited to three cascaded three-dimensional convolutional layers; the number of volume attention-guided residual modules can be adjusted according to actual needs, and the branching structure in the module can also be adjusted according to actual needs, such as increasing or decreasing the number of convolutional layers, or using other types of activation functions. Furthermore, the configuration of the decoder module can also be optimized according to different application scenarios, such as using transposed convolutional kernels of different sizes. Therefore, this application is not limited to the specific embodiments described above, but also includes all equivalent transformations and improvements made based on the technical concept of this application.
[0199] loss function
[0200] The loss function will be explained in further detail below.
[0201] Optionally, the cycle consistency loss in the loss function is calculated using a perceptual metric to better reflect the characteristics of the human visual system.
[0202] Optionally, the perceptual metric method is the learned perceptual patch similarity (LPIPS) metric, which is used to calculate the perceptual difference between the generated image and the original image, and avoids over-smoothing of the generated image by capturing high-level features.
[0203] Optionally, the cycle consistency loss includes:
[0204] After the sparse angle CT image is transformed to the full angle CT image domain by the first generator A, it is then transformed back to the sparse angle CT image domain by the second generator B. The LPIPS metric between the obtained result and the original sparse angle CT image is obtained.
[0205] After transforming the full-angle CT image to the sparse angle CT image domain through the second generator B, it is then transformed back to the full-angle CT image domain through the first generator A. The LPIPS metric between the obtained result and the original full-angle CT image is then calculated.
[0206] Specifically, the cycle consistency loss in the loss function is calculated using a perceptual metric to better align with the characteristics of the human visual system. Specifically, this perceptual metric is the Learned Perceptual Patch Similarity (LPIPS) metric. The LPIPS metric is used to calculate the perceptual difference between the generated image and the original image, avoiding over-smoothing of the generated image by capturing high-level features.
[0207] The calculation process for cycle consistency loss is as follows: First, the sparse angle CT image is transformed to the full-angle CT image domain by a first generator A, and then transformed back to the sparse angle CT image domain by a second generator B. The resulting image is then compared with the original sparse angle CT image using LPIPS to calculate the perceptual difference between them. Similarly, the full-angle CT image is transformed to the sparse angle CT image domain by the second generator B, and then transformed back to the full-angle CT image domain by the first generator A. The final resulting image is then compared with the original full-angle CT image using LPIPS to calculate the perceptual difference between them.
[0208] In this way, cycle consistency loss can ensure that the image maintains high-quality perceptual consistency during the transformation between the two domains, thereby improving the overall effect of image processing and reconstruction.
[0209] It is understood that in this embodiment, a perceptual measurement method is adopted, specifically using the learned perceptual patch similarity (LPIPS) metric, in order to better conform to the characteristics of the human visual system.
[0210] Traditional loss functions, such as mean squared error (MSE) or structural similarity (SSIM), tend to perform comparisons at the pixel level. However, these methods may not adequately capture the differences in image quality perceived by the human visual system. The LPIPS metric was introduced to address this issue.
[0211] LPIPS is a deep learning-based perceptual similarity measure. It uses pre-trained neural networks (typically trained on large-scale image classification tasks) to extract high-level features from images. These high-level features are closer to the human visual system's understanding of image content, including texture, shape, and semantic information. By comparing these high-level features rather than raw pixel values, LPIPS is able to better assess the perceptual similarity between images.
[0212] In CT image processing, using LPIPS offers several important technical advantages, such as:
[0213] Avoiding over-smoothing: Traditional pixel-level loss functions often result in overly smooth images, losing important texture details. LPIPS, by focusing on high-level features, better preserves these details. Preserving structural information: LPIPS captures the structural information of images, which is particularly important for preserving complex anatomical structures in CT images. Aligning with medical diagnostic needs: In medical image processing, preserving key diagnostic information is more important than maintaining pixel-level precision. LPIPS' characteristics better meet this requirement.
[0214] The calculation process of the cycle consistency loss demonstrates the innovation of this method: 1) Cyclic conversion from sparse angle to full angle: The sparse angle CT image is converted into a full angle image by generator A, and then converted back into a sparse angle image by generator B. Finally, the LPIPS metric between this result and the original sparse angle image is calculated. 2) Cyclic conversion from full angle to sparse angle: Similarly, the full angle CT image is converted into a sparse angle image by generator B, and then converted back into a full angle image by generator A. Then the LPIPS metric between this result and the original full angle image is calculated.
[0215] This bidirectional recurrent design ensures the consistency and reversibility of the transformation process. By using the LPIPS metric, the network is trained to maintain consistency not only at the pixel level but also at high-level features and perceptual quality. This is particularly important for preserving key diagnostic information in CT images.
[0216] In summary, this LPIPS-based cyclic consistency loss design reflects a deep understanding of the specific needs of medical image processing, providing an optimization target that is more in line with human visual perception and medical diagnostic needs for high-quality CT image reconstruction.
[0217] It should be noted that the above descriptions are preferred embodiments of this application and are exemplary, not limiting. Besides the specific structures and steps described, this application can also employ various other implementation methods. For example, the calculation method for the cycle consistency loss is not limited to Learned Perceptual Image Patch Similarity (LPIPS), but can also use other perceptual metrics, such as Structural Similarity Index (SSIM) or other deep learning-based image similarity metrics. Furthermore, the specific implementation of the cycle consistency loss can also differ; for example, network performance can be optimized by adding additional loss terms. Therefore, this application is not limited to the specific embodiments described above, but also includes all equivalent transformations and improvements made according to the technical concept of this application.
[0218] Optionally, the formula for calculating the cycle consistency loss is:
[0219] L cycle =E[||B(A(s))-s|| LPIPS ]+E[||A(B(f))-f|| LPIPS (5)
[0220] L cycle This represents the loss of cycle consistency.
[0221] E[·] represents the expectation operation;
[0222] ||·|| LPlPS This indicates that the distance was calculated using LPIPS (Learned Perceptual Patch Similarity) metric.
[0223] s represents a sparse angle CT image;
[0224] f represents a full-angle CT image;
[0225] A(·) represents the first generator that transforms sparse angle CT images into the full angle CT image domain;
[0226] B(·) represents the second generator that transforms the full-angle CT image to the sparse-angle CT image domain;
[0227] B(A(s)) represents the result of transforming the sparse angle CT image s into the full angle CT image domain first through generator A, and then transforming it back into the sparse angle CT image domain through generator B.
[0228] A(B(f)) represents the result of transforming the full-angle CT image f into the sparse angle CT image domain through generator B, and then transforming it back into the full-angle CT image domain through generator A.
[0229] The formula above describes the calculation method for cycle consistency loss, which is an innovative application based on the Learned Perceptual Image Patch Similarity (LPIPS) metric. This loss function is designed to ensure consistency of images when translating between two domains, while taking into account the characteristics of human visual perception.
[0230] Using LPIPS metrics instead of traditional pixel-level metrics (such as mean squared error) offers several key advantages: 1) Perceptual similarity: LPIPS captures image similarities that more closely resemble human visual perception, which is particularly important for medical images. 2) Preservation of structural information: It better preserves the structural and texture information of images, avoiding over-smoothing. 3) High-level feature comparison: LPIPS extracts high-level features through pre-trained neural networks, which better reflect the semantic content of images.
[0231] This cycle consistency loss ensures that the image transformation between the two domains is consistent and reversible. It encourages the network to learn to preserve the essential features of the image, without losing important diagnostic information even during artifact removal. This is particularly important for CT image reconstruction, as it requires striking a balance between artifact removal and detail preservation.
[0232] In summary, the design of this loss function reflects a deep understanding of the specific needs of CT image processing, combining advanced deep learning technology and expertise in medical image processing to provide a powerful optimization target for high-quality CT image reconstruction.
[0233] Optionally, the loss function may further include a reconstruction loss, used to calculate the perceptual difference between the generated image and the ground truth image, guiding the network to optimize image details.
[0234] It should be noted that the loss function also includes a reconstruction loss, used to calculate the perceptual difference between the generated image and the ground truth image, thereby guiding the network to optimize image details. Specifically, the reconstruction loss evaluates image quality by comparing the differences between the generated image and the ground truth image. The calculation of the reconstruction loss is based on a perceptual metric, which measures the similarity between the generated image and the ground truth image in high-level features, rather than just pixel-level differences. Perceptual metrics can capture the texture, shape, and other complex feature information of an image, allowing the network to retain and recover more image details during optimization. By introducing the reconstruction loss, the network can be continuously adjusted and improved during training to generate higher-quality, more realistic images.
[0235] Specifically, the reconstruction loss is introduced to further improve the quality of the generated image, especially in terms of accuracy in detail.
[0236] The core idea of reconstruction loss is to directly compare the differences between the generated image and the ground truth image (i.e., the target image). In the context of CT image reconstruction, this means comparing the differences between the image after artifact removal and the ideal full-angle CT image. This comparison is not just a simple pixel-level comparison, but employs a perceptual difference metric.
[0237] Using perceptual discriminant analysis instead of traditional pixel-level discriminant analysis (such as mean squared error) has several important technical considerations: 1) Preserving structural information: Perceptual metrics focus more on the structural and textural information of an image, rather than just the precise matching of pixel values. This is crucial for preserving key anatomical structures and lesion features in CT images. 2) Avoiding over-smoothing: Traditional pixel-level loss functions often result in overly smooth images, losing important details. Perceptual metrics can better preserve these diagnostically valuable details. 3) Aligning with the human visual system: Perceptual metrics are closer to how the human visual system judges image quality, which is especially important for medical images, as these images are ultimately interpreted by doctors. 4) High-level feature comparison: Many perceptual metrics (such as LPIPS mentioned above) use pre-trained neural networks to extract high-level features. These features better reflect the semantic content of the image, rather than just low-level pixel information.
[0238] By introducing a reconstruction loss, the network is trained not only to remove artifacts but also to reconstruct a result as similar as possible to the ground truth image. This process can be viewed as a form of supervised learning, where the ground truth image provides a clear objective.
[0239] The inclusion of reconstruction loss has several important impacts on network training: Detail optimization: It guides the network to pay more attention to image details, ensuring accurate reconstruction of important anatomical structures and potential lesion features. Balancing denoising and fidelity: Reconstruction loss helps find a balance between removing artifacts (primarily handled by cycle consistency loss) and preserving realistic information. Accelerated convergence: By providing a more direct optimization objective, reconstruction loss can help the network converge to the ideal solution faster. Improved generalization ability: By comparing multiple real samples, the network can learn more generalized image reconstruction rules, rather than relying solely on cycle consistency.
[0240] In summary, the introduction of reconstruction loss reflects a comprehensive consideration of the CT image reconstruction task. It not only focuses on artifact removal but also emphasizes the overall quality and detail accuracy of the reconstructed image. This multi-objective optimization strategy helps generate high-quality CT images that are both clear and artifact-free while preserving crucial diagnostic information.
[0241] Optionally, the formula for calculating the reconstruction loss is:
[0242] L rec =E[||A(s)-f|| LPIPS ]+E[||B(f)-s|| LPIPS (6)
[0243] in,
[0244] L rec Indicates the losses incurred during reconstruction;
[0245] E[·] represents the expectation operation;
[0246] ||·|| LPIPS This indicates that the distance was calculated using LPIPS (Learned Perceptual Patch Similarity) metric.
[0247] s represents a sparse angle CT image;
[0248] f represents a full-angle CT image;
[0249] A(·) represents a generator that transforms sparse angular CT images into the full-angle CT image domain;
[0250] B(·) represents a generator that transforms a full-angle CT image to a sparse-angle CT image domain;
[0251] A(s) represents the result of transforming the sparse angle CT image s into the full angle CT image domain through generator A;
[0252] B(f) represents the result of transforming the full-angle CT image f into the sparse-angle CT image domain through generator B.
[0253] The above formula describes the method for calculating the reconstruction loss, which is an innovative application based on the Learned Perceptual Image Patch Similarity (LPIPS) metric. The reconstruction loss is designed to directly evaluate the perceptual similarity between the generated image and the target image, thereby guiding the network to generate high-quality reconstruction results.
[0254] Using LPIPS metrics instead of traditional pixel-level metrics (such as mean squared error) has several important advantages: Perceptual similarity: LPIPS captures image similarity that more closely resembles human visual perception, which is particularly important for the quality assessment of medical images. Preservation of structural information: It better preserves the structural and textural information of images, avoiding the generation of overly smoothed results. High-level feature comparison: LPIPS extracts high-level features through pre-trained neural networks, which better reflect the semantic content of images. High adaptability: LPIPS can capture subtle differences that are difficult for the human eye to quantify but can perceive.
[0255] This reconstruction loss function serves several key purposes: Direct optimization objective: It provides the network with a direct optimization objective—the generated image should be as close as possible to the real image. Bidirectional consistency: By simultaneously considering the transformation from sparse angles to full angles and from full angles to sparse angles, it ensures bidirectional consistency in the transformation process. Detail preservation: The use of LPIPS helps preserve important image details, which is crucial for identifying minute structures and potential lesions in CT images. Balancing denoising and fidelity: This loss function helps the network find a balance between artifact removal and preserving realistic information.
[0256] By calculating the expected values of these two parts, the reconstruction loss L is obtained. rec This loss function guides the network to more closely resemble real images during the image generation process, thereby optimizing image details and improving the quality of the generated images.
[0257] In summary, this reconstruction loss design reflects a deep understanding of the CT image reconstruction task. It not only focuses on artifact removal but also emphasizes the overall quality and perceptual similarity of the reconstructed image. This approach helps generate high-quality CT images that are both clear and artifact-free while preserving key diagnostic information, providing a powerful optimization objective for medical image processing.
[0258] Optionally, the overall loss of the loss function is a weighted sum of the adversarial loss, the cycle consistency loss, and the reconstruction loss, calculated as follows:
[0259] L total =L GAN (A,D F )+L GAN (B,D S ) +λ1L cycle +λ2L rec(7)
[0260] in,
[0261] L total Indicates the total loss;
[0262] L GAN (A,D F () represents generator A and discriminator D F Adversarial loss between angles is used to optimize image transformation from sparse angles to full angles;
[0263] L GAN (B,D S () represents generator B and discriminator D S The adversarial loss between them is used to optimize the image transformation from full angle to sparse angle;
[0264] L cycle This represents the loss of cycle consistency.
[0265] L rec Indicates the losses incurred during reconstruction;
[0266] λ1 and λ2 are weighting coefficients used to balance the contributions of different loss terms;
[0267] A represents a generator that transforms sparse angular CT images into the full-angle CT image domain;
[0268] B represents the generator that transforms full-angle CT images into the sparse-angle CT image domain;
[0269] D F A discriminator representing the full-angle image domain is used to distinguish between real full-angle CT images and images generated by generator A;
[0270] D S A discriminator representing the sparse angular image domain is used to distinguish between real sparse angular CT images and images generated by generator B.
[0271] The formula above describes the overall loss function used in network training, which is a weighted combination of multiple loss terms. This composite loss function design reflects a comprehensive consideration of the CT image reconstruction task, aiming to optimize network performance from multiple perspectives.
[0272] Among them, L total Indicates the total loss. L GAN (A,D F () represents generator A and discriminator D F The adversarial loss between the two is used to optimize the image transformation from sparse angles to full angles. Specifically, generator A generates full-angle CT images, while discriminator D... FThis is used to distinguish images generated by generator A from real full-angle CT images. Adversarial loss improves the generator's performance through this adversarial mechanism, making the generated images more realistic.
[0273] Similarly, L GAN (B,D S () represents generator B and discriminator D S An adversarial loss is used to optimize the image transformation from full-angle to sparse-angle. Generator B generates sparse-angle CT images, and discriminator D... S This is used to distinguish images generated by generator B from real sparse angle CT images. This adversarial mechanism also improves the performance of generator B, making its generated sparse angle CT images more realistic.
[0274] In addition, L cycle Representing the cycle consistency loss, it is used to ensure that an image remains consistent during transformations between two domains. Cycle consistency loss ensures the reversibility and stability of the transformation process by measuring the differences between the image in the two transformations.
[0275] L rec Representing the reconstruction loss, it measures the perceptual difference between the generated image and the ground truth image. The reconstruction loss guides the network to retain more detail during image generation by calculating the high-level feature similarity between the generated and ground truth images.
[0276] Weight coefficients λ1 and λ2 are used to balance the contributions of different loss terms. These coefficients adjust the relative importance of adversarial loss, cycle consistency loss, and reconstruction loss in the overall loss. By setting these weight coefficients appropriately, it can be ensured that the network simultaneously optimizes the realism, consistency, and detail preservation of the generated images during training.
[0277] Generators A and B are used to convert sparse angle CT images to the full-angle CT image domain and to convert full-angle CT images to the sparse angle CT image domain, respectively. Discriminator D... F Discriminator D is used to distinguish between real full-angle CT images and images generated by generator A. S This is used to distinguish between real sparse angle CT images and images generated by Generator B. By integrating these different loss functions and network components, the overall network can more effectively remove artifacts and generate high-quality CT images.
[0278] This composite loss function design offers several key advantages: Multi-faceted optimization: It simultaneously considers image realism, transformation consistency, and reconstruction accuracy. Balancing different objectives: By adjusting weights, a balance is found between artifact removal, detail preservation, and generating realistic images. Enhanced perceptual quality: The LPIPS metric focuses on image quality perceived by the human visual system. Bidirectional learning: Optimization in two directions ensures bidirectional consistency and reliability of the transformation. High adaptability: Weights can be adjusted to meet different CT image reconstruction needs.
[0279] This overall loss function reflects a deep understanding and comprehensive consideration of the CT image reconstruction task, integrating multiple advanced deep learning concepts to generate high-quality, artifact-free CT images that retain key diagnostic information.
[0280] It should be noted that the above descriptions are preferred embodiments of this application and are exemplary, not limiting. Besides the specific structures and steps described, this application can also employ various other implementation methods. For example, the calculation method for cycle consistency loss is not limited to Learned Perceptual Image Patch Similarity (LPIPS), but can also use Structural Similarity Index (SSIM) or other deep learning-based image similarity metrics. The calculation of reconstruction loss can also use other perceptual metrics to replace LPIPS. Furthermore, the generator and discriminator structures for adversarial loss can be adjusted according to specific applications, such as employing different network architectures or optimization algorithms. The weight coefficients (λ1 and λ2) of the overall loss can also be adjusted according to different training requirements to balance the contributions of each loss term. Therefore, this application is not limited to the specific embodiments described above, but also includes all equivalent transformations and improvements made based on the technical concept of this application.
[0281] The following is combined with Figure 2 , Figure 3 and Figure 4 The following is a further explanation of this embodiment.
[0282] Figure 2 This paper presents the generator architecture of a volume-based sparse angular CT image artifact removal network. The architecture, from left to right, illustrates the entire processing flow from the input CT image with artifacts to the output optimized result. The generator architecture mainly consists of four parts: an encoder, a dilated spatial convolutional pooling pyramid module, a volume attention-guided residual module, and a decoder.
[0283] The encoder, located on the left side of the structure, contains multiple 3D convolutional layers. Each convolutional layer is followed by instance normalization and a ReLU activation function, a design that helps extract multi-level features from the input image. Following the encoder, the dilated spatial convolutional pooling pyramid module (labeled c in the figure) plays a crucial role. This module contains multiple parallel convolutional branches, each using a different dilation rate, enabling the network to effectively capture multi-scale features and enhancing its ability to recognize artifacts of varying sizes.
[0284] The core of the network consists of six cascaded volume attention-guided residual modules (labeled B in the diagram). Each module contains a sophisticated volume attention mechanism that captures the correlations between four dimensions: channel (C), depth (D), height (H), and width (W). Furthermore, each module includes convolutional branches for further feature extraction, and residual connections that directly pass input features. This design enhances the network's expressive power and facilitates stable gradient propagation.
[0285] The decoder, located on the right side of the structure, consists of multiple 3D transposed convolutional layers. These layers reconstruct the processed features into the final output image. Notably, several skip connections are also present between the encoder and decoder; this design helps preserve the detail information of the original image, ensuring high-quality output.
[0286] This network architecture design embodies several key innovations: First, by using 3D manipulation to process volumes, it fully utilizes the three-dimensional information of CT images; second, the volume attention mechanism can capture complex correlations between different dimensions, significantly enhancing the network's expressive power; third, the introduction of dilated convolutional pyramid modules enables the network to effectively capture multi-scale features; and finally, the use of residual structures and skip connections not only facilitates the network training process but also better preserves image details.
[0287] Overall, this network architecture aims to process 3D CT images more effectively, improve artifact removal, and ensure high-quality output images. This method fully considers the three-dimensional characteristics of CT images and the complexity of artifacts, providing a powerful solution for artifact removal in sparse angular CT images.
[0288] Figure 3 This paper demonstrates the training process and loss function composition of a block-based sparse-angle CT image artifact removal network. The figure clearly shows how the network processes both sparse-angle and full-angle CT images simultaneously, and the various loss functions applied during training.
[0289] In the upper part of the figure, we can see the processing flow of the sparse angle CT image s. First, s is processed by generator A, transforming it from the sparse angle image domain (S) to the full angle image domain (F), resulting in A(s). Then, A(s) is transformed back to the sparse angle image domain by generator B, resulting in B(A(s)). This process forms a complete loop used to calculate the cycle consistency loss L. cycle_s Simultaneously, A(s) is compared with the actual full-angle CT image f to calculate the reconstruction loss L. rec_s In addition, A(s) is also fed into the discriminator D. F Used to calculate adversarial loss L GAN (A,D F ).
[0290] The lower half of the figure illustrates a similar processing procedure for a full-angle CT image f. f is first converted to the sparse angular image domain by generator B, yielding B(f), and then converted back to the full-angle image domain by generator A, yielding A(B(f)). This process is used to calculate another cycle consistency loss L. rec_f B(f) is compared with the true sparse angular CT image s to calculate the reconstruction loss L. rec_f Similarly, B(f) is fed into the discriminator D. s Used to calculate adversarial loss L GAN (B,D S )).
[0291] This complex training process and diverse loss function design reflect several key characteristics of the network: First, bidirectional transformation (S→F and F→S) ensures that the network can learn the essential relationship between sparse angle and full-angle CT images; second, cycle consistency loss guarantees the reversibility of the transformation process, which helps to preserve important image features; third, reconstruction loss ensures the similarity between the generated image and the real image; finally, adversarial loss prompts the generator to produce more realistic images, improving the quality of artifact removal.
[0292] In summary, this carefully designed training process and loss function combination enables the network to comprehensively learn the features of sparse angular CT images, effectively removing artifacts while preserving image realism and detail. This method not only considers image transformation in a single direction but also ensures the consistency and reliability of the transformation through bidirectional loops, providing strong support for high-quality reconstruction of sparse angular CT images.
[0293] Figure 4This paper compares the performance of our method with several other state-of-the-art methods in artifact removal from sparse-angle CT images. The image comprises six columns of images, from left to right: a full-angle CT image (as a reference standard), the original sparse-angle CT image, the result processed by the U-former method, the result processed by the PTNet3D method, the result processed by the Pix2Pix3D method, and the result processed by our method. The images are arranged in two rows: the upper row shows the transverse images, and the lower row shows the sagittal images. This multi-angle presentation helps to comprehensively evaluate the performance of each method.
[0294] In the cross-sectional images, obvious striped artifacts are clearly visible in the original sparse-angle CT images, severely affecting image quality and detail recognition. In contrast, the images processed by the method in this application almost completely remove these artifacts, and the image quality is close to that of full-angle CT images. Especially in the area marked by the yellow box in the figure, the method in this application demonstrates excellent detail preservation capabilities, and the structural clarity is significantly better than other methods.
[0295] The comparison of sagittal images further highlights the advantages of the proposed method. In the original sparse angular CT images, severe blurring and streaking artifacts are observed, causing the image to almost lose its continuity in the vertical direction. Other methods have improved this situation to some extent, but varying degrees of blurring and structural breaks still exist. In contrast, the sagittal images processed by the proposed method exhibit excellent continuity and clarity, especially in the area marked by the red box, where the restoration of fine structures is very close to that of full-angle CT images.
[0296] Through this visual comparison, we can intuitively see that the method proposed in this application performs excellently in removing artifacts, preserving details, and maintaining image continuity. Especially when dealing with complex three-dimensional structures, this method can maintain high-quality image reconstruction across different cross-sections, which is of great significance for improving the application value of CT images in clinical diagnosis. In summary, Figure 4 This strongly demonstrates the superiority of the proposed method in the task of removing artifacts from sparse angle CT images. It not only shows significant results on the cross-section but also demonstrates a clear advantage in maintaining the three-dimensional consistency of the image.
[0297] The above embodiments have the following technical effects:
[0298] The above embodiments propose an innovative method for artifact removal in sparse-angle CT images, the core of which lies in fully utilizing the three-dimensional information of CT images and advanced deep learning technology. This method mainly comprises three key points, each of which brings significant technical improvements.
[0299] First, the above embodiments employ a three-dimensional sub-block approach, rather than a two-dimensional slice approach, to perform artifact removal on sparse angular CT images. This innovation fully utilizes the three-dimensional information of CT images. Traditional methods typically treat CT images as a series of independent two-dimensional slices, ignoring the spatial relationships between slices. The above embodiments, however, effectively capture the structural information of CT images in various directions by processing three-dimensional sub-blocks containing multiple adjacent slices. The technical effects of this method are significant: the processed images exhibit better structural coherence in the transverse, sagittal, and coronal planes, greatly improving the restoration of image details. Particularly in the sagittal and coronal planes, this method effectively avoids the discontinuities common in traditional methods, making the reconstructed three-dimensional images more realistic and reliable. Furthermore, by simultaneously considering the information of multiple adjacent slices, this method can more accurately distinguish between real structures and artifacts in the image, thereby achieving more precise artifact removal.
[0300] Secondly, the above embodiment specifically designed a volume-based network structure for the task of artifact removal in sparse angular CT images. The core of this network lies in its volume attention-guided residual module, which effectively captures cross-dimensional correlations between different dimensions, significantly enhancing the network's expressive power. Specifically, this module processes the channel, depth, height, and width dimensions through four branches, capturing the complex interactive relationships between them. This design brings several technical benefits: First, it greatly improves the network's understanding and processing of three-dimensional structural information, enabling the network to more accurately identify and remove complex artifact patterns; second, the network's multi-branch structure and residual connections not only enhance feature extraction capabilities but also make the training process more stable and efficient by allowing direct gradient transfer; finally, this design makes the network more adaptable when processing CT images of different locations or with different scanning parameters.
[0301] Finally, the above embodiments employ the LPIPS metric when calculating the cycle consistency loss, introducing the concept of cycle consistency perceptual loss. This loss function is more in line with the characteristics of the human visual system, effectively avoiding the generation of overly smoothed images. Traditional pixel-level loss functions often lead to the loss of image details, while the LPIPS metric, by considering the similarity of high-level features, can better preserve the texture and detail information of the image. The technical effects of this innovation are significant: First, the processed images are visually more natural and realistic, with richer details, which is crucial for medical diagnosis; second, this method achieves a better balance between artifact removal and preservation of important anatomical structural details, improving the clinical applicability of the processing results; finally, because it is more in line with human visual perception, images processed by this method are more easily accepted and interpreted by doctors.
[0302] Furthermore, this application significantly optimizes processing speed by processing in units of three-dimensional sub-blocks. Two-dimensional processing methods require frequent loading and processing of each tomographic image, while this application reduces the frequency of data loading by processing three-dimensional sub-blocks containing multiple adjacent slices. This improvement significantly enhances processing efficiency, especially in large-scale CT data processing scenarios. Simultaneously, this application considers the memory pressure of processing three-dimensional images, processing three-dimensional sub-blocks instead of the entire three-dimensional image. This makes it possible to process larger-sized three-dimensional data under the same hardware conditions, effectively alleviating memory pressure in high-resolution CT image processing. Moreover, the strategy of sampling three-dimensional sub-blocks during training in this application involves sampling sub-blocks at random locations within the entire three-dimensional image each round, rather than manually cutting sub-blocks when creating the dataset. This increases the diversity of training samples, allowing the network to learn more information. In summary, especially in large-scale CT data processing scenarios, these improvements significantly enhance processing efficiency and capacity, enabling efficient CT image artifact removal without increasing hardware costs.
[0303] In summary, these three key aspects not only individually bring significant technological advancements, but their synergistic effect also produces a comprehensive result that transcends the impact of any single technique. The 3D sub-block processing method provides the network with rich spatial information, and the specially designed network structure fully utilizes this information through its volume attention-guided residual modules, accurately capturing the correlations between different dimensions. This structure complements the cyclic consistency perceptual loss, jointly ensuring the preservation of subtle anatomical details while removing artifacts. In particular, the LPIPS metric, which considers human visual perception when evaluating image quality, combined with the 3D sub-block processing method, maintains a high degree of perceptual consistency throughout the entire 3D space.
[0304] The synergistic effect is significant: for a CT image containing 150 slices, the processing time was reduced from 18 seconds (using U-former as an example) to just 3 seconds, a six-fold increase in efficiency, while simultaneously achieving unprecedented image quality improvements. This combination of high efficiency and high quality makes this method highly valuable in clinical and research scenarios requiring rapid processing of large numbers of sparse-angle CT images. More notably, the advantages of this method increase exponentially with the amount of data to be processed and the number of slices. This is because the 3D processing method reduces the need for switching between different dimensions, the efficient design of the network structure reduces computational complexity, and the use of perceptual loss ensures excellent image quality even during high-speed processing. This scalability makes this method ideal for large-scale CT image data processing, bringing revolutionary progress to the field of medical imaging.
[0305] To better understand the technical solution of this application, a specific example is provided below. The details listed in this example are mainly for ease of understanding and are not intended to limit the scope of protection of this application.
[0306] This example presents a method for artifact removal in sparse angular CT images, including the following steps:
[0307] Step 110: Obtain sparse angle CT images and full-angle CT images, and preprocess the data, including resampling, size cropping and grayscale normalization.
[0308] Preferably, in step 110, the object in the CT image is a walnut. In this embodiment, the full angle is preferably 1200 projection angles. To demonstrate the effectiveness of the method in the case of extremely sparse angles, the sparse angle is preferably one-fortieth to one-twentieth of the full angle, that is, 30 to 60 projection angles, specifically 40 projection angles, which is one-thirtieth of the number of projection angles in the full angle. The original image size is (512, 512, 512), and without resampling, its size is cropped to (360, 360, 150), and the grayscale values are normalized to the range [-1, 1].
[0309] Step 120: Construct a volume-based sparse angular CT image artifact removal network. The model includes two pairs of generators and discriminators. One pair of generators and discriminators transforms the image from the sparse angular image domain (S) to the full angular image domain (F), while the other pair performs the reverse process. The generator includes an encoder module, a dilated spatial convolutional pooling pyramid module, a volume attention-guided residual module, and a decoder module; the discriminator consists of 5 downsampled convolutional blocks.
[0310] More specifically, step 120 may further include the following sub-steps:
[0311] Step 121: The generator described in step 120 consists of an encoder module, a dilated spatial convolutional pooling pyramid module, six attention-guided residual modules, and a decoder module connected in series. The final output image has the same size as the original image. The structure of each module is as follows:
[0312] The encoder module consists of three 3D convolutional layers connected in series. The first convolutional layer has a kernel size of 7×7×7 and a stride of 1. The second and third convolutional layers have kernel sizes of 3×3×3 and a stride of 2. An instance regularization layer and a ReLU activation function are added after each convolutional layer.
[0313] The dilated spatial convolutional pooling pyramid module consists of a multi-branch structure with a 1×1×1 convolution, three dilated convolutions with different dilation rates, and a global average pooling layer. The kernel size of the three dilated convolutions is 3×3×3, the stride is 1, and the dilation rates are 1, 2, and 4, respectively. The feature layers output by the previous encoder module pass through the above five branches and are spliced along the channels.
[0314] The volume attention-guided residual block consists of three branches: a volume attention weighted branch, a convolutional branch, and a residual connection branch. The volume attention weighted branch pools the four dimensions of channel, volume length, volume width, and volume height, and then passes them through a convolutional layer, an instance regularization layer, and a sigmoid activation function to capture the cross-dimensional correlation between each of the three dimensions. The average of these correlations is then multiplied by the input features. The convolutional branch consists of two convolutional layers with a kernel size of 3×3×3 and a stride of 1.
[0315] The decoder module consists of three transposed convolutional layers connected in series. The first and second transposed convolutional layers have a kernel size of 3×3×3 and a stride of 2. Each layer is followed by an instance regularization layer and a ReLU activation function. The third transposed convolutional layer has a kernel size of 7×7×7 and a stride of 1. It is followed by Tanh activation.
[0316] Step 122: The discriminator described in step 120 consists of five downsampled feature extraction layers connected in series. The feature extraction layers include convolutional layers, instance regularization layers, and LeakyReLu activation functions.
[0317] Step 130: Train the network using the preprocessed data. In each round, sample image sub-blocks of a certain size at random locations and input them into the network for training. Preferably, for step 130, the size of the sampled image sub-blocks is (180, 180, 28), and the training process is as follows: Figure 3 As shown, taking a sparse angle CT image sub-block s as an example, the generator A that executes S→F transforms s into A(s), and then the generator B that executes F→S transforms it back into S and obtains the result B(A(s)).
[0318] Step 140: Optimize the model using a loss function consisting of adversarial loss, cycle consistency-aware loss, and reconstruction loss. The adversarial loss, cycle consistency-aware loss, and reconstruction loss may further include the following sub-steps:
[0319] Step 141: Define adversarial loss: Adversarial loss is used to improve the image generation ability of the generator and the classification ability of the discriminator;
[0320] L GAN (A,D F )=E[log D F (f)]+E[log(1-DF (A(s)))] (3)
[0321] L GAN (B,D S )=E[log D S (s)]+E[log(1-D S (B(f)))] (4)
[0322] Where S is the sparse angular CT image domain, s is a sample in S, F is the full-angle CT image domain, and f is a sample in F. A is a generator that transforms the sparse angular CT image domain to the full-angle CT image domain, and B is a generator that transforms the full-angle CT image domain to the sparse angular CT image domain. D F To determine whether an image belongs to the full-angle CT image domain, D S A discriminator for determining whether an image belongs to the sparse angular CT image domain.
[0323] E[logD F (f)]: This part is the expectation operation, representing the discriminator D. F The probability of identifying a sample in the full-angle CT image domain as a full-angle CT image. E[log D] S The same logic applies to (s).
[0324] E[log(1-D F [(A(s)))]: This is also the expectation operation, representing the discriminator D. F The probability that the result of transforming samples from the sparse angle CT image domain to the full angle CT image domain is a sparse angle CT image is determined by E[log(1-D s The same logic applies to (B(f)))].
[0325] Step 142: Define the cyclic consistency perceptual loss. After the sparse angle CT image is transformed to the full angle CT image domain by generator A, it is then transformed back to the sparse angle CT image domain by another generator B. The result should theoretically be the same as the original image. The difference between the transformed result and the original image is measured using LPIPS, which is the cyclic consistency perceptual loss. The cyclic transformation process of the full angle image is the same.
[0326] L cycle =E[||B(A(s))-s|| LPIPS ]+E[||A(B(f))-f|| LPIPS (5)
[0327] The first term is the formula for calculating the cycle consistency loss for sparse angle CT images. For a sparse angle CT image s, it is first transformed to the full-angle CT image domain by generator sub-network A, and then transformed back to the sparse angle CT image domain by generator sub-network B. Theoretically, it should be the same as the original image s. The perceptual loss between it and s is calculated, and the expected value should be as small as possible. The second term is calculated similarly.
[0328] Step 143: Define the reconstruction loss, calculate the perceptual loss between the generated image and the ground truth image, and guide the network to optimize image details;
[0329] L rec =E[||A(s)-f|| LPIPS ]+E[||B(f)-s|| LPIPS (6)
[0330] E[||A(s)-f|| PPIPS A(s) represents the result of transforming the sparse angular CT image s to the full-angle CT image domain, and f is the ground truth image of s in the full-angle CT image domain. The goal is to calculate the perceptual loss between the transformed image and the ground truth image, aiming to minimize this loss. E[||B(f)-s||] LPIPS The same principle applies.
[0331] Step 144: Define the overall loss by adding the adversarial loss, the cyclic consistency perception loss, and the reconstruction loss according to certain weights to obtain the overall loss.
[0332] L total =L GAN (A,D F )+L GAN (B,D s )+λ1L cycle +λ2L rec (7)
[0333] Preferably, in step 140, the weight coefficients λ1, λ2, and λ3 are all set to 10, the model is optimized using the Adam optimizer after 300 iterations.
[0334] Step 150: Divide the sparse angle CT image into several sub-blocks of the same size as the sampling size during training according to the spatial position order, input them into generator A for processing in sequence, and finally merge them to form the optimized result.
[0335] Preferably, in step 150, the sparse angle CT image is divided into several sub-blocks of size (180, 180, 28) according to spatial position, and then the sub-blocks are sequentially input into generator A for processing. The output sub-blocks are then filled into the corresponding positions. For cases where the sub-blocks are not divisible, the last extra incomplete sub-block is also taken from the end to the front and a sub-block of size (180, 180, 28) is input into generator A. The corresponding part of the output result is then filled to obtain the final result after artifact removal.
[0336] The second embodiment of this application relates to a system for artifact removal in sparse angle CT images, the structure of which is as follows: Figure 6 As shown, the system for artifact removal in sparse-angle CT images includes:
[0337] The data preprocessing module is used to acquire sparse angle CT images and full-angle CT images, and preprocess them to obtain preprocessed sparse angle CT images and full-angle CT images.
[0338] The network construction module is used to divide the preprocessed sparse angle CT image and full angle CT image into several training three-dimensional sub-blocks; and to construct a block-based sparse angle CT image artifact removal network, which includes two pairs of generators and discriminators, one pair of which is used to transform the training three-dimensional sub-blocks from the sparse angle image domain to the full angle image domain, and the other pair is used to perform the reverse process.
[0339] The network training module is used to train the network using the training 3D sub-blocks. During the training process, a loss function including adversarial loss, cycle consistency loss and reconstruction loss is used to optimize the network to obtain the trained network.
[0340] The image processing module is used to preprocess the sparse angle CT image to be processed, divide it into several three-dimensional sub-blocks of the same size as the three-dimensional sub-blocks according to the spatial position order, and input the three-dimensional sub-blocks to be processed into the trained network in sequence to obtain the three-dimensional sub-blocks after removing artifacts.
[0341] The image synthesis module is used to merge the artifact-removed 3D sub-blocks to form an artifact-removed CT image.
[0342] The first embodiment is a method embodiment corresponding to this embodiment. The technical details in the first embodiment can be applied to this embodiment, and the technical details in this embodiment can also be applied to the first embodiment.
[0343] It should be noted that those skilled in the art should understand that the implementation functions of each module shown in the above-described embodiments of the system for removing artifacts from sparse-angle CT images can be understood with reference to the relevant descriptions of the methods for removing artifacts from sparse-angle CT images described above. The functions of each module shown in the above-described embodiments of the system for removing artifacts from sparse-angle CT images can be implemented by a program (executable instructions) running on a processor, or by specific logic circuits. If the system for removing artifacts from sparse-angle CT images described in this application is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any particular combination of hardware and software.
[0344] Accordingly, this application also provides a computer storage medium storing computer-executable instructions, which, when executed by a processor, implement the various method implementations of this application.
[0345] Furthermore, this application also provides a system for artifact removal in sparse angle CT images, including a memory for storing computer-executable instructions and a processor; the processor is used to implement the steps in the above-described method embodiments when executing the computer-executable instructions in the memory. The processor may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The aforementioned memory may be read-only memory (ROM), random access memory (RAM), flash memory, hard disk, or solid-state drive, etc. The steps of the methods disclosed in the embodiments of this application can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules in the processor.
[0346] All documents mentioned in this application are considered to be incorporated in their entirety into the disclosure of this application so that they can serve as a basis for modifications if necessary. Furthermore, it should be understood that after reading the foregoing disclosure of this application, those skilled in the art can make various alterations or modifications to this application, and these equivalent forms also fall within the scope of protection claimed in this application.
Claims
1. A method for artifact removal in sparse-angle CT images, characterized in that, Includes the following steps: Acquire sparse angle CT images and full angle CT images, and preprocess them to obtain preprocessed sparse angle CT images and full angle CT images; The preprocessed sparse angle CT image and full angle CT image are divided into several training 3D sub-blocks; and a block-based sparse angle CT image artifact removal network is constructed, the network including two pairs of generators and discriminators, one pair is used to transform the training 3D sub-blocks from the sparse angle image domain to the full angle image domain, and the other pair is used to perform the reverse process. The network is trained using the three-dimensional sub-blocks used for training. During the training process, a loss function including adversarial loss, cycle consistency loss, and reconstruction loss is used to optimize the network, resulting in a trained network. After preprocessing the sparse angle CT image to be processed, it is divided into several three-dimensional sub-blocks of the same size as the three-dimensional sub-blocks according to spatial position. The three-dimensional sub-blocks to be processed are then input into the trained network for processing to obtain three-dimensional sub-blocks after artifact removal. The artifact-removed three-dimensional sub-blocks are then merged to form a CT image after artifact removal.
2. The method according to claim 1, characterized in that, The generator of the block-based sparse angular CT image artifact removal network includes an encoder module, a dilated spatial convolutional pooling pyramid module, multiple volume attention-guided residual modules, and a decoder module connected in series; wherein: The encoder module consists of multiple three-dimensional convolutional layers connected in series. Each three-dimensional convolutional layer is followed by an instance normalization layer and a ReLU activation function to extract features of the training three-dimensional sub-blocks. The dilated spatial convolutional pooling pyramid module receives the output of the encoder module, including multiple dilated convolutional branches with different dilation rates and a global average pooling branch. The outputs of each branch are concatenated along the channel direction to capture multi-scale features. The volume attention-guided residual module includes a volume attention weighted branch, a convolutional branch, and a residual connection branch. The volume attention weighted branch pools the input features in four dimensions: channel, depth, height, and width. After processing through convolutional layers, instance normalization layers, and a sigmoid activation function, it captures cross-dimensional correlations between different dimensions and multiplies them with the input features. The convolutional branch includes multiple three-dimensional convolutional layers. The residual connection branch directly passes the input features. The decoder module consists of multiple 3D transposed convolutional layers connected in series. Each 3D transposed convolutional layer is followed by an instance normalization layer and an activation function to reconstruct the processed features into an output image with the same size as the input.
3. The method according to claim 2, characterized in that, The dilated spatial convolution pooling pyramid module includes a 1×1×1 convolution branch, three 3×3×3 dilated convolution branches with different dilation rates, and a global average pooling branch. The outputs of each branch are spliced along the channel direction.
4. The method according to claim 2, characterized in that, The volume attention-guided residual module includes: - The volumetric attention weighted branch pools the input features in four dimensions: channel, depth, height, and width. Then, it captures the cross-dimensional correlation between each of the three dimensions through convolutional layers, instance normalization layers, and the Sigmoid activation function, and then multiplies it with the input features. - Convolutional branch, consisting of two cascaded 3×3×3 three-dimensional convolutional layers with a stride of 1; - Residual connection branches are used to directly pass input features.
5. The method according to claim 1, characterized in that, The cycle consistency loss in the loss function is calculated using a perceptual metric to better align with the characteristics of the human visual system. Furthermore, the perceptual metric is the Learned Perceptual Patch Similarity (LPIPS) metric, which is used to calculate the perceptual difference between the generated image and the original image. By capturing high-level features, it avoids over-smoothing of the generated image.
6. The method according to claim 5, characterized in that, The cycle consistency loss includes: After the sparse angle CT image is transformed to the full angle CT image domain by the first generator A, it is then transformed back to the sparse angle CT image domain by the second generator B. The LPIPS metric between the obtained result and the original sparse angle CT image is obtained. After transforming the full-angle CT image to the sparse angle CT image domain through the second generator B, and then transforming it back to the full-angle CT image domain through the first generator A, the LPIPS metric between the obtained result and the original full-angle CT image is calculated. The formula for calculating the cycle consistency loss is as follows: L cycle =E[||B(A(s))-s|| LPIPS ]+E[||A(B(f))-f|| LPIPS ] (5) L cycle This represents the loss of cycle consistency. E[·] represents the expectation operation; ||·|| LPIPS This indicates that the distance was calculated using LPIPS (Learned Perceptual Patch Similarity) metric. s represents a sparse angle CT image; f represents a full-angle CT image; A(·) represents the first generator that transforms sparse angle CT images into the full angle CT image domain; B(·) represents the second generator that transforms the full-angle CT image to the sparse-angle CT image domain; B(A(s)) represents the result of transforming the sparse angle CT image s into the full angle CT image domain first through generator A, and then transforming it back into the sparse angle CT image domain through generator B. A(B(f)) represents the result of transforming the full-angle CT image f into the sparse angle CT image domain through generator B, and then transforming it back into the full-angle CT image domain through generator A.
7. The method according to claim 6, characterized in that, The loss function also includes a reconstruction loss, used to calculate the perceptual difference between the generated image and the ground truth image, guiding the network to optimize image details. The formula for calculating the reconstruction loss is as follows: L rec =E[||A(s)-f|| LPIPS ]+E[||B(f)-s|| LPIPS ] (6) in, L rec Indicates the losses incurred during reconstruction; E[·] represents the expectation operation; ||·|| LPIPS This indicates that the distance was calculated using LPIPS (Learned Perceptual Patch Similarity) metric. s represents a sparse angle CT image; f represents a full-angle CT image; A(·) represents a generator that transforms sparse angular CT images into the full-angle CT image domain; B(·) represents a generator that transforms a full-angle CT image to a sparse-angle CT image domain; A(s) represents the result of transforming the sparse angle CT image s into the full angle CT image domain through generator A; B(f) represents the result of transforming the full-angle CT image f into the sparse-angle CT image domain through generator B.
8. The method according to claim 7, characterized in that, The overall loss of the loss function is a weighted sum of the adversarial loss, the cycle consistency loss, and the reconstruction loss, calculated as follows: L total =L GAN (A,D F )+L GAN (B,D S )+λ1L cycle +λ2L rec (7) in, L total Indicates the total loss; L GAN (A, D) F () represents generator A and discriminator D F Adversarial loss between angles is used to optimize image transformation from sparse angles to full angles; L GAN (B, D) S () represents generator B and discriminator D S The adversarial loss between them is used to optimize the image transformation from full angle to sparse angle; L cycle This represents the loss of cycle consistency. L rec Indicates the losses incurred during reconstruction; λ1 and λ2 are weighting coefficients used to balance the contributions of different loss terms; A represents a generator that transforms sparse angular CT images into the full-angle CT image domain; B represents the generator that transforms full-angle CT images into the sparse-angle CT image domain; D F A discriminator representing the full-angle image domain is used to distinguish between real full-angle CT images and images generated by generator A; D S A discriminator representing the sparse angular image domain is used to distinguish between real sparse angular CT images and images generated by generator B.
9. An apparatus for removing artifacts from sparse angle CT images, characterized in that, include: The data preprocessing module is used to acquire sparse angle CT images and full-angle CT images, and preprocess them to obtain preprocessed sparse angle CT images and full-angle CT images. The network construction module is used to divide the preprocessed sparse angle CT image and full angle CT image into several training three-dimensional sub-blocks; and to construct a block-based sparse angle CT image artifact removal network, which includes two pairs of generators and discriminators, one pair of which is used to transform the training three-dimensional sub-blocks from the sparse angle image domain to the full angle image domain, and the other pair is used to perform the reverse process. The network training module is used to train the network using the training 3D sub-blocks. During the training process, a loss function including adversarial loss, cycle consistency loss and reconstruction loss is used to optimize the network to obtain the trained network. The image processing module is used to preprocess the sparse angle CT image to be processed, divide it into several three-dimensional sub-blocks of the same size as the three-dimensional sub-blocks according to the spatial position order, and input the three-dimensional sub-blocks to be processed into the trained network in sequence to obtain the three-dimensional sub-blocks after removing artifacts. The image synthesis module is used to merge the artifact-removed 3D sub-blocks to form an artifact-removed CT image.
Citation Information
Patent Citations
Thoracic cavity CT image artifact removal method and system based on three-dimensional generative adversarial network
CN114897726A
Generative adversarial network-based artifact suppression method for extreme sparse view CT (Computed Tomography) reconstruction
CN115239588A