CT metal artifact removal method based on geometric perception graph learning
By using geometric perceptual map learning and geometric contrast loss function, the problems of inaccurate artifact localization and insufficient detail preservation in CT images are solved, achieving high-precision artifact removal and image detail restoration, which is suitable for the removal of CT metal artifacts.
Patent Information
- Application Number
- CN202510907166.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-11-25
AI Technical Summary
Existing CT metal artifact removal methods suffer from inaccurate artifact localization in the image domain, insufficient preservation of image details, and lack of geometric priors, resulting in poor diagnostic interference and artifact suppression.
A geometry-aware map-based learning approach is adopted, which constructs point-to-point geometric artifact maps and geometry-aware map neural networks, and combines them with a geometric contrast loss function to accurately locate artifacts and preserve image details.
It improves the accuracy of artifact localization, enhances the ability to preserve image details, has strong applicability, reduces the interference of metal artifacts on clinical diagnosis, and has a visualized artifact attention map to enhance the interpretability of the algorithm.
Smart Images

Figure CN121010657A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of medical image analysis technology, specifically relating to a method for removing CT metal artifacts. Background Technology
[0002] Computed tomography (CT) has become a widely used imaging tool in clinical diagnosis and screening. However, when a patient has high-density metal implants (such as dental fillings, metal prostheses, etc.), the metal will excessively absorb X-rays during the projection process, resulting in severe artifacts in the original projection data. In the reconstructed CT image, these artifacts appear as radial or strip-shaped patterns, which not only obscure the clinically interesting region but may also mislead diagnostic conclusions [1,2]. To address the problem of metal artifacts, metal artifact removal algorithms have been proposed, and two main approaches have been developed: the first is projection domain artifact removal. This method identifies and locates the "metal projection" region in the projection domain and interpolates or reconstructs this region to reduce artifact propagation; although it has achieved significant performance improvements, this type of method has high requirements for the availability of the original projection data, and new global "secondary artifacts" are easily introduced during the interpolation process [1,2,4,7]; the other approach is image domain artifact removal. This method directly uses deep neural networks or filtering methods to suppress artifacts on the reconstructed CT images [3, 4, 6]. Since image domain artifacts lack obvious spatial local features and are superimposed on real tissue structures, the network must distinguish between artifacts and clinical details while ensuring that details are not overly smoothed. Currently, most image domain methods treat all pixels equally, making it difficult to provide stronger suppression for areas with severe artifacts. Furthermore, there is a lack of effective geometric priors to guide network learning. In summary, existing metal artifact removal methods have advantages and disadvantages in both the projection and image domains: projection domain methods rely on projection data and may introduce new artifacts, while image domain methods face three major challenges: difficulty in artifact localization, insufficient detail preservation and suppression capabilities, and a lack of geometric priors. Summary of the Invention
[0003] The purpose of this invention is to propose a CT metal artifact removal method based on geometric perception map learning, which features high artifact localization accuracy, good image detail preservation, and wide applicability. This addresses the problems of existing technologies, such as the difficulty in accurately locating artifacts in image domain MAR, the weakening of clinical details, and the underutilization of artifact geometric features.
[0004] The proposed CT metal artifact removal method based on geometric perception map learning first generates an artifact map by constructing a point-to-point geometric artifact map based on the position and shape of the metal implant in the CT image to approximate the distribution of severe artifacts between metal implant areas. Then, a geometric perception map neural network is used for modeling. Based on the artifact map, a graph neural network module including full-image convolution and local aggregation is employed to learn the geometric relationships of artifacts layer by layer from global to local, outputting an artifact attention weight map. Finally, a geometric contrast loss-guided training method is designed. During network training, a geometric contrast loss function is introduced to enhance the network's ability to identify the location and intensity of artifacts through the comparison and optimization of positive and negative samples in the geometric structure space.
[0005] The specific steps are as follows:
[0006] (I) Construction of artifact images and generation of spatial masks based on geometric priors;
[0007] Based on the location and shape of the metallic implant in the CT image, a point-to-point geometric artifact map is constructed to approximately identify the areas where heavy metal artifacts may occur, and this map is converted into a spatial artifact mask as a stable geometric prior. The specific process is as follows:
[0008] (1) Metal Region Segmentation and Boundary Extraction: First, threshold segmentation of the metal region is performed on the reconstructed CT image. Specifically, based on the Hounsfield Unit (HU), pixels with a HU value greater than 2800 are extracted as the initial metal mask. Subsequently, connected component analysis is performed on the metal mask to identify and label each independent metal implant region. To reduce subsequent computational complexity, only the boundary pixels of each metal implant region are extracted, forming multiple independent boundary pixel sets Pi, where each set corresponds to the boundary of an implant.
[0009] (2) Artifact Graph Construction: After obtaining the boundary pixel sets of each metal implant, the pixels from the boundaries of different metal implants are paired up and connected to form edges, thus constructing a preliminary artifact graph G. To achieve a balance between computational cost and performance, a preset sampling rate is introduced to control the total number of edges. Specifically, the sampling rate is set to 0.1, that is, 10% of the possible pixel pair connections are randomly sampled to form the edges of the graph. At the same time, to ensure that the connectivity information of the artifacts is not lost, it is mandatory that at least one connecting edge is retained between any two different metal implants.
[0010] (3) Spatial Artifact Mask Generation: To ensure the discrete graph structure is compatible with the regular grid input requirements of mainstream MAR networks (such as CNNs), the constructed artifact map G is converted into a spatial artifact mask MG. Specifically, a zero-based mask of the same size as the original CT image is created. Then, each edge in the artifact map G is traversed, and the values of the pixels traversed by that edge are incremented by 1. After traversal, the entire mask is subjected to min-max normalization to ensure that its pixel values fall within the [0,1] interval. This generated artifact mask MG can roughly indicate the distribution area of severe artifacts and will serve as a geometric prior for subsequent network training.
[0011] (II) Construction of GeoGNN (Geometric Perceptual Graph Neural Network);
[0012] Based on the artifact prior, a GeoGNN module is designed to learn the geometric relationships of artifacts layer by layer from global to local, and to dynamically generate more accurate artifact localization for each CT image. The specific process is as follows:
[0013] (1) Feature preparation: The GeoGNN module is inserted as a plug-in into any existing MAR backbone network. For a given input feature map Z (from the backbone network), the last c channels are first extracted and reshaped into a two-dimensional feature matrix F. Then, through two independent lightweight linear mappings, global features H for global modeling and local features V for local modeling are generated from F respectively;
[0014] (2) Global artifact modeling and feature update:
[0015] (a) Global Graph Construction: To capture the overall distribution of artifacts, the global feature H is first divided into k spatial patches and their corresponding central features C using average pooling. Then, a hybrid adjacency matrix A is constructed on top of these centers. This matrix is a weighted fusion of two parts: ① Local adjacency matrix A_local: connecting the eight neighborhoods of each patch center to preserve the local spatial structure. ② Global adjacency matrix A_global: calculated using a sparsified self-attention mechanism to capture long-range dependencies. The sparsity operation is implemented using a quantization operator q(·), which retains only elements greater than or equal to the 95th percentile in the self-attention score matrix, thus focusing on the most relevant global relationships. The final adjacency matrix A is obtained by fusing A_local and A_global with weight β and then normalizing, and serves as the edge matrix of the global artifact graph.
[0016] (b) Global feature update: Apply a graph convolutional network (GCN) to the constructed global artifact map to update the patch center features and obtain the updated center features Cb.
[0017] (3) Local artifact modeling and feature updating:
[0018] (a) Local graph construction: Using the updated global central feature Cb as the cluster center, all pixels are dynamically clustered into different clusters based on the cosine similarity between each pixel feature in the local feature V and these central features.
[0019] (b) Local Feature Update: Within each cluster, a local subgraph is constructed, where the cluster center and its contained pixels constitute nodes. Then, weighted message passing and feature aggregation are performed on this subgraph. Finally, the features of all pixels are updated, outputting artifact-aware features Za containing accurate geometric information.
[0020] (III) Feature Fusion and Reconstruction
[0021] The geometric perception features learned by the GeoGNN module are fused with the original features to enhance the artifact removal capability of the backbone network. The specific process is as follows:
[0022] (1) The artifact-aware feature Za output by the GeoGNN module is fused with the original feature map Z of the backbone network. Specifically, a 1×1 convolutional layer is used to fuse the two to generate an enhanced feature map Zb.
[0023] (2) Replace the feature map at the corresponding position in the original artifact removal backbone network with the enhanced feature Zb, and then the network continues to perform normal forward propagation to complete the subsequent artifact suppression and image detail reconstruction tasks.
[0024] Step 4: Artifact Attention Map Generation;
[0025] After decoding the artifact-aware features learned by GeoGNN, an artifact attention map is generated that can be used for visualization and supervised learning. The specific process is as follows:
[0026] (1) First, average pooling is performed on the artifact-aware feature maps Za output by each layer of the GeoGNN module in the network according to the channel dimension.
[0027] (2) Next, all feature maps that have been averaged through the channels are scaled up or down to the same spatial size as the spatial artifact mask MG generated in the first step.
[0028] (3) Finally, these uniformly sized feature maps are concatenated along the channel dimension and passed through a 1×1 convolutional decoder to generate the final artifact attention map EG. This map visually shows the artifact regions that the network is interested in.
[0029] (v) Constructing geometric contrastive loss to guide network training;
[0030] A novel geometric contrast loss function is designed to enhance the network's ability to identify the location and intensity of artifacts by comparing positive and negative samples in the geometric structure space. The specific process is as follows:
[0031] (1) Queue Construction: During network training, the artifact attention map EG generated in step (iv) and the spatial artifact mask MG generated in step 1 are flattened into one-dimensional feature vectors zE and zM, respectively. These vectors are then stored in their respective "first-in, first-out" (FIFO) queues. In this invention, the length of the queue is set to 600. Geometric contrast loss is only calculated after the queue is full.
[0032] (2) Loss Calculation: A selective supervised contrastive loss function is used to guide the training of the GeoGNN module. The specific form of this loss function is as follows:
[0033]
[0034] Among them, z E z M Here, EG and MG are the flattened feature representations, respectively, and τ is the temperature hyperparameter. This is the memory queue for MG. The Select operation selects the top k data pairs most similar to the current metal mask M based on cosine similarity. The goal of the loss function is to maximize the positive sample pairs (z) in geometric space. E , z M The similarity between zE and other anatomical contexts (i.e., aligning the attention map EG generated by the network with the geometric prior MG) is minimized, while simultaneously minimizing the similarity between the current zE and artifact patterns (negative samples) in other different anatomical contexts. This leverages the stability of the geometric prior while preserving the individual differences in artifact patterns across different CT images, thereby enhancing the network's ability to suppress severe artifacts and effectively avoiding over-smoothing of normal tissue details.
[0035] The present invention also includes a plug-and-play CT metal artifact removal system based on the above method, specifically comprising:
[0036] (1) Graph construction module: used to perform step (i) to perform metal segmentation on CT images and construct artifact map G and spatial artifact mask MG.
[0037] (2) Geometric perception graph neural network module: used to execute step (ii), accept input feature Z and output artifact perception feature Za.
[0038] (3) Feature fusion module: used to perform step (iii) to fuse Za and Z to generate enhanced feature Zb.
[0039] (4) Attention Decoding Module: Used to execute step (iv) and generate artifact attention map EG.
[0040] (5) Loss Calculation Module: Used to execute step (5), calculate geometric contrast loss based on EG and MG to guide learning.
[0041] (6) Interface unit: responsible for outputting the enhanced feature Zb to any existing metal artifact removal backbone network (such as U-Net, CNN, Transformer, etc.) for subsequent image reconstruction.
[0042] This invention, by organically combining geometric modeling, graph neural networks, and contrastive learning, introduces a stable and effective geometric prior for artifacts in image-domain MAR for the first time. As a plug-and-play framework, it demonstrates excellent performance on both simulated and real clinical datasets, offering advantages such as high artifact localization accuracy, good image detail preservation, and strong applicability. Furthermore, it enhances the interpretability of the algorithm through visualized attention maps, effectively mitigating the interference of metal artifacts on clinical diagnosis. Attached Figure Description
[0043] Figure 1 This is a flowchart illustrating the CT metal artifact removal method based on geometric perception map learning according to the present invention.
[0044] Figure 2 This describes the process of constructing the artifact map based on geometric priors in this invention.
[0045] Figure 3 This paper distinguishes between metal artifact removal guided by prior knowledge of chord map masking and metal artifact removal guided by geometry map proposed in this paper.
[0046] Figure 4 The images show the effects of different metal artifact methods on a dental simulation dataset. Results prefixed with GraphMAR indicate the use of this invention to enhance the corresponding pedestal model. The displayed CT window ranges from [-200, 300] HU. The red mask represents the shape and location of the metal artifacts.
[0047] Figure 5 This image shows the effects of different metal artifact methods on real clinical dental data. Results prefixed with GraphMAR are those using the enhancement method of this invention to the corresponding pedestal model. The displayed CT window range is [-200, 300] HU. The red mask represents the shape and location of the metal.
[0048] Figure 6 This is a visualization of the attention map of metal artifacts output by the present invention. Detailed Implementation
[0049] The present invention will be further introduced through specific simulation examples below, and the effectiveness of the present invention in removing metal artifacts from metal artifact data will be demonstrated, as well as the comparison with other methods, including image quality and quantization indicators.
[0050] On the DDMAR dataset, 400 CT images were first selected from the DeepLesion dataset. Following the QuadNet method, 90 metal blocks were randomly added to each image, generating 36,000 training samples containing metal artifacts. Then, 10 metal blocks were randomly added to each of another 200 CT images, generating 2,000 test samples. The reconstructed image size was 512×512, with a corresponding sinogram size of 640×640. On the Dental real-world oral CT dataset, data from 104 de-identified patients were collected, including 1,125 images with metal artifacts and 2,233 clean images unaffected by artifacts. To construct a paired training set, 7,200 pairs of images with / without artifacts were randomly generated from 45 clean images as training data, while an additional 800 pairs were reserved as test data. Quantitative metrics were evaluated on the simulated paired data, and a subjective visual comparison was performed on the 1,125 real metal artifact images. Since the dataset does not provide sinograms or traditional interpolation prior maps, all methods that rely on the projection domain are excluded in this comparison.
[0051] This invention compares various types of metal artifact removal methods, including both classic algorithms and the current best-performing methods. The proposed model is compared with the following state-of-the-art methods: LI[1], NMAR[2], FBPConvNet[3], FredNet[4], DuDoNet++[7], DANNet[8], ACDNet[5], QuadNet[9], and RiseMAR[6]. Among them, LI and NMAR are classic baseline methods widely used in the MAR field. FBPConvNet and FredNet use U-Net and Fourier networks for image post-processing, respectively. ACDNet is an image domain network that uses a convolutional dictionary network to encode priors related to metal artifacts. RiseMAR is an advanced image domain framework that improves performance by introducing radiologist feedback and uses ProCT
[41] as its network implementation. DuDoNet++, DANNet, and QuadNet are dual-domain methods that use independent networks in the sinusoidal domain and the image domain, respectively, to improve reconstruction quality. All models were trained and tested on the same dataset to ensure fairness in the comparison. The code implementation in CNNMAR was used for LI and NMAR
[15] , while the official code was used for ACDNet, FredNet, RiseMAR and QuadNet. For FBPConvNet and DuDoNet++, the original paper was strictly followed for reproduction.
[0052] GraphMAR is implemented in PyTorch and can be seamlessly integrated into four existing image domain MAR backbone networks: FBPConvNet, FredNet, RiseMAR, and DuDoNet++. In the UNet backbone of FBPConvNet and DuDoNet++, a GeoGNN is inserted after each of the four downsampling modules; in the RiseMAR encoder, a GeoGNN is inserted after each of the two downsampling steps; and in FredNet's seven residual blocks, a GeoGNN is inserted after every two residual blocks, for a total of four. To balance performance and computational cost, the number of feature channels in the input GeoGNN is set to 8% of the total number of channels in the original network, and the feature map is divided into 64 patches for parallel processing in the graph neural network.
[0053] The model was trained using the Adam optimizer with β1 = 0.5 and β2 = 0.999 for a total of 120 epochs. The initial learning rate was set to 4 × 10⁻⁶. -4The learning rate was halved every 35 epochs. Training was performed in parallel on eight NVIDIA 4090 GPUs, with a total batch size of 16. Peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) were used to measure the experimental results. PSNR was defined as follows:
[0054]
[0055] SSIM is defined as follows:
[0056]
[0057] PSNR represents the pixel-level match between the algorithm's artifact removal results and normal-dose CT, while SSIM represents the structural similarity between the two. We use the improvement relative to the input and the corresponding baseline model as the relative improvement ratio. The definition is as follows:
[0058]
[0059] Where v i and This represents a model enhanced using the method proposed in this invention, and a corresponding baseline model, v max This represents the highest possible index under ideal conditions, with ideal indices of 50 for PSNR and 100 for SSIM. Additionally, v input The metric representing the input artifact map serves as the minimum performance benchmark.
[0060] Experiment Example 1: Comparison of different metal artifact removal methods on the DDMAR dataset
[0061] Table 1 Comparison of different metal artifact removal methods on the DDMAR dataset
[0062]
[0063] Quantitative evaluations of different methods were conducted on the DDMAR dataset. The results show that deep learning methods generally outperform traditional methods, with dual-domain methods outperforming pure image-domain methods due to the utilization of projection domain sinogram information. Integrating GraphMAR into FBPConvNet, FredNet, and DuDoNet++ improved the average PSNR by 0.66dB, 0.21dB, and 0.57dB, respectively, achieving performance improvements across five sets of images with different metal sizes. Notably, FBPConvNet showed a greater gain than FredNet, indicating that GeoGNN's enhanced global modeling capabilities through graph structure are more beneficial to convolutional networks with limited receptive fields. Furthermore, even when integrated only in the image domain, GraphMAR remains compatible with dual-domain networks, effectively suppressing secondary artifacts and further improving overall performance.
[0064] Experiment Example 2: Comparison of different metal artifact removal methods on the Dental dataset
[0065] Table 2 Comparison of different metal artifact removal methods on the oral simulation dataset
[0066]
[0067]
[0068] Quantitative evaluation results on the Dental dataset show that, due to the difficulty in obtaining sinogram data and accurate prior images in clinical scenarios, this experiment only compared image-domain methods. While LI and NMAR can improve reconstruction quality to some extent, LI's performance significantly declines in the group with the smallest metal implant size. This is mainly because the sinogram obtained directly from image projection is not accurate and may introduce new secondary artifacts, leading to performance degradation and reduced reliability of clinical decisions. Integrating GraphMAR into the existing backbone network resulted in a continuous improvement in the performance of all models. The gain of GraphMAR is particularly significant when the metal implant is small, indicating that its geometric prior and graph neural network can more effectively recover details in terms of weak artifact suppression. For images with larger metal implants, although the backbone network still has difficulty in recovering in areas with severe artifacts, the assistance of GraphMAR still brings observable improvements. Overall, GraphMAR requires only a few network layers to modify and is completely independent of sinogram data, demonstrating good versatility and clinical applicability. GraphMAR demonstrates stable and significant performance improvements in both simulated data and real-world oral CT artifact removal tasks. Figure 4 The visual results demonstrate the competitive approach. It can be seen that the present invention effectively enhances the performance of the baseline.
[0069] Experiment Example 3: Comparison of different metal artifact removal methods on real clinical datasets
[0070] The model weights given in Table 2 are used to test real artifact images (e.g., Figure 5 As shown in the image, the original backbone network still plays a crucial role in overall performance, with RiseMAR demonstrating excellent performance in all three cases through its spatial-frequency domain Transformer modeling of long-range dependencies. Integrating GraphMAR into the three backbone networks further improved the MAR results for all methods. Specifically, in the first two rows of scans affected by metal braces and orthodontic brackets, GraphMAR effectively guided the model to restore accurate anatomical structures and significantly suppressed strip artifacts. In the third row of images containing small metal implants, GraphMAR's enhancement of details such as the medullary cavity was particularly noticeable. Furthermore, comparing the performance of GraphMAR-enhanced FBPConvNet with the original FBPConvNet in the first two rows shows that GraphMAR is more advantageous in reducing long-range artifacts when the metal implant span is large. Overall, GraphMAR's output is more regular and conforms to anatomical features, fully validating its effectiveness in removing metal artifacts in real-world clinical scenarios. Figure 6 The attention map generated by the method proposed in this invention is illustrated. This attention map effectively outlines artifact regions, providing interpretability for downstream applications and helping to improve the reliability of artifact removal algorithms.
[0071] References:
[0072] [1] WA Kalender, R. Hebel, and J. Ebersberger, "Reduction of CT artifacts caused by metallic implants." Radiology, vol. 164, no. 2, pp. 576–577, 1987.
[0073] [2]E.Meyer, R.Raupach, M.Lell, B.Schmidt, and M.Kachelrieβ,
[0074] “Normalized metal artifact reduction NMAR in computed tomography,”
[0075] Med.Phys.,vol.37,no.10,pp.5482–5493,2010.
[0076] [3]K.H.Jin,M.T.McCann,E.Froustey,and M.Unser,“Deep convolutionalneural network for inverse problems in imaging,”IEEE Trans.
[0077] Med.Imaging,vol.26,no.9,pp.4509–4522,2017.
[0078] [4]Z.Li,C.Ma,J.Chen,J.Zhang,and H.Shan,“Learning to distill globalrepresentation for sparse-view CT,”in Proceedings of the Proc.
[0079] IEEE / CVF Int.Conf.Comput.Vis.,2023,pp.21 196–21 207.
[0080] [5]H.Wang,Y.Li,D.Meng,and Y.Zheng,“Adaptive convolutional dictionarynetwork for CT metal artifact reduction,”arXiv preprint arXiv:2205.07471,2022.
[0081] [6]C.Ma,Z.Li,Y.Li,J.Han,J.Zhang,Y.Zhang,J.Liu,and H.Shan,
[0082] “Radiologist-in-the-loop self-training for generalizable CT metalartifactreduction,”IEEE Trans.Med.Imaging,2025.
[0083] [7]Y.Lyu,W.-A.Lin,H.Liao,J.Lu,and S.K.Zhou,“Encoding metalmaskprojection for metal artifact reduction in computed tomography,”inProc.Int.Conf.Med.Image Comput.Comput.-Assisted Intervention,2020,pp.147–157.
[0084] [8]T.Wang et al.,“Dual-domain adaptive-scaling non-local networkforCT metal artifact reduction,”in Proc.Int.Conf.Med.Image Comput.Comput.-Assisted Intervention,2021,pp.243–253.
[0085] [9]Z.Li,Q.Gao,Y.Wu,C.Niu,J.Zhang,M.Wang,G.Wang,and H.Shan,“Quad-Net:Quad-domain network for CT metal artifact reduction,”IEEE Trans.Med.Imaging,pp.1–1,2024。
Claims
1. A method for removing CT metal artifacts based on geometric perception map learning, characterized in that, include: Artifact map generation involves constructing point-to-point geometric artifact maps based on the location and shape of metal implants in CT images to approximate the distribution of severe artifacts between metal implant areas. Geometric perceptual graph neural network modeling is then performed. Building upon the artifact maps, a graph neural network module incorporating full-image convolution and local aggregation is used to learn the geometric relationships of artifacts layer by layer from global to local, outputting an artifact attention weight map. Geometric contrast loss-guided training is designed, introducing a geometric contrast loss function during network training. This function optimizes the network's ability to identify artifact location and intensity through comparison of positive and negative samples in the geometric structure space. The specific steps are as follows: (I) Construction of artifact images and generation of spatial masks based on geometric priors; Based on the location and shape of the metal implant in the CT image, a point-to-point geometric artifact map is constructed to approximately identify the area where heavy metal artifacts may occur, and it is converted into a spatial artifact mask as a stable geometric prior. (II) Construction of a geometric perception graph neural network; Based on the artifact prior, a GeoGNN module is designed to learn the geometric relationships of artifacts layer by layer from global to local, and to dynamically generate more accurate artifact localization for each CT image. (III) Feature fusion and reconstruction; The geometric perception features learned by the GeoGNN module are fused with the original features to enhance the artifact removal capability of the backbone network. (iv) Artifact attention map generation; Decode the artifact perception features learned by GeoGNN to generate artifact attention maps that can be used for visualization and supervised learning; (v) Constructing geometric contrastive loss to guide network training; We designed a geometric contrast loss function to enhance the network's ability to identify the location and intensity of artifacts by comparing positive and negative samples in the geometric structure space.
2. The CT metal artifact removal method according to claim 1, characterized in that, The specific process for step (one) is as follows: (1) Metal region segmentation and boundary extraction: First, threshold segmentation of the metal region is performed on the reconstructed CT image; specifically, based on the Henle unit (HU), pixels with a HU value greater than 2800 in the image are extracted as the initial metal mask; Subsequently, connected component analysis is performed on the metal mask to identify and mark each independent metal implant region. To reduce the computational complexity of the subsequent steps, only the boundary pixels of each metal implant region are extracted to form multiple independent boundary pixel sets Pi, where each set corresponds to the boundary of an implant. (2) Artifact map construction: After obtaining the boundary pixel set of each metal implant, the pixels from the boundaries of different metal implants are paired up and these paired pixels are connected to form edges to construct a preliminary artifact map G; In order to achieve a balance between computation and performance, a preset sampling rate is introduced to control the total number of edges; To ensure that the connectivity information of the artifacts is not lost, at least one connecting edge is required between any two different metal implants. (3) Spatial artifact mask generation: In order to make the discrete graph structure compatible with the requirements of the mainstream MAR network for regular grid input, the constructed artifact map G is converted into a spatial artifact mask MG. The specific operation is as follows: create an all-zero mask with the same size as the original CT image, and then traverse each edge in the artifact map G, and increment the value of the pixel position passed by the edge by 1. After the traversal is completed, the entire mask is subjected to min-max normalization so that its pixel value falls within the interval [0,1]. The artifact mask MG is used to indicate the distribution area of heavy artifacts, as a geometric prior for subsequent network training.
3. The CT metal artifact removal method according to claim 2, characterized in that, The specific process for step (two) is as follows: (1) Feature preparation: The GeoGNN module is inserted into the MAR backbone network as a plug-in; for a given input feature map Z, the last c channels are first extracted and reshaped into a two-dimensional feature matrix F; then, through two independent lightweight linear mappings, global features H for global modeling and local features V for local modeling are generated from F respectively. (2) Global artifact modeling and feature update: (a) Global graph construction: In order to capture the overall distribution of artifacts, firstly, the global feature H is divided into k spatial patches and their corresponding central features C through average pooling; then, a hybrid adjacency matrix A is constructed on these centers. (b) Global feature update: Apply a graph convolutional network (GCN) to the constructed global artifact map to update the patch center features and obtain the updated center features Cb; (3) Local artifact modeling and feature updating: (a) Local graph construction: Using the updated global central feature Cb as the cluster center, all pixels are dynamically clustered into different clusters based on the cosine similarity between each pixel feature in the local feature V and these central features. (b) Local feature update: Within each cluster, a local subgraph is constructed, in which the cluster center and its contained pixels constitute nodes; then, weighted message passing and feature aggregation are performed on this subgraph; finally, the features of all pixels are updated, and the artifact-aware features Za containing accurate geometric information are output.
4. The CT metal artifact removal method according to claim 3, characterized in that, The adjacency matrix A mentioned in step (II) is formed by weighted fusion of two parts: ① Local adjacency matrix A_local: connects the eight neighborhoods of each patch center to preserve the local structure of the space; ② Global adjacency matrix A_global: calculated through a sparsified self-attention mechanism to capture long-range dependencies; wherein, the sparsification operation is implemented through a quantization operator q(·), which only retains elements greater than or equal to the 95th percentile in the self-attention score matrix, thereby focusing on the most relevant global relationships; the final adjacency matrix A is obtained by fusing Alocal and Aglobal with weight β and then normalizing, and serves as the edge matrix of the global artifact graph.
5. The CT metal artifact removal method according to claim 4, characterized in that, The specific process for step (three) is as follows: (1) The artifact-aware feature Za output by the GeoGNN module is fused with the original feature map Z of the backbone network. Specifically, a 1×1 convolutional layer is used to fuse the two to generate an enhanced feature map Zb. (2) Replace the feature map at the corresponding position in the original artifact removal backbone network with the enhanced feature Zb, and then the network continues to perform normal forward propagation to complete the subsequent artifact suppression and image detail reconstruction tasks.
6. The CT metal artifact removal method according to claim 5, characterized in that, The specific process for step (four) is as follows: (1) Perform channel-dimensional average pooling on the artifact-aware feature maps Za output by each layer of the GeoGNN module in the network; (2) All feature maps that have been averaged through the channels are scaled up or down to the same spatial size as the spatial artifact mask MG generated in the first step. (3) These feature maps of uniform size are spliced together along the channel dimension and then passed through a 1×1 convolutional decoder to generate the final artifact attention map EG; this map visually shows the artifact regions that the network is interested in.
7. The CT metal artifact removal method according to claim 6, characterized in that, The specific process for step (five) is as follows: (1) Queue construction: During network training, the artifact attention map EG generated in step (IV) and the spatial artifact mask MG generated in step (I) are flattened into one-dimensional feature vectors zE and zM respectively; then, these vectors are stored in their respective first-in-first-out (FIFO) queues. (2) Loss Calculation: A selectively supervised contrastive loss function is used to guide the training of the GeoGNN module. The specific form of this loss function is as follows: Among them, z E z M Here, EG and MG are the flattened feature representations, respectively, and τ is the temperature hyperparameter. The memory queue for MG; the Select operation selects the top k data pairs most similar to the current metal mask M based on cosine similarity; the objective of this loss function is to maximize the positive sample pairs (z) in geometric space. E , z M The similarity between zE and other anatomical contexts is calculated by aligning the attention map EG generated by the network with the geometric prior MG, while minimizing the similarity between the current zE and artifact patterns in other different anatomical contexts.
8. A plug-and-play CT metal artifact removal system based on the CT metal artifact removal method according to any one of claims 1-7, specifically comprising: (1) Graph construction module: used to perform the operation in step (i); (2) Geometric Perceptual Graph Neural Network Module: Used to perform the operations in step (ii); (3) Feature fusion module: used to perform the operation in step (iii); (4) Attention Decoding Module: Used to perform the operation in step (iv); (5) Loss Calculation Module: Used to perform the operations in step (five); (6) Interface unit: responsible for outputting the enhanced feature Zb to any existing metal artifact removal backbone network for subsequent image reconstruction.