A dual-domain artifact correction method based on the fusion of multi-level data and physical priors

A dual-domain artifact correction method that integrates multi-level data with physical priors solves the artifact problem in industrial CT inspection of high-density metal components, achieving high-precision and robust artifact correction applicable to imaging needs in multiple fields and improving detection efficiency and accuracy.

CN121329834BActive Publication Date: 2026-04-07ZHEJIANG UNIV +1
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-16
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing industrial CT inspection technologies suffer from severe metal artifacts, high data dependence, poor physical consistency, and low computational efficiency when dealing with high-density metal components, making it difficult to meet the needs of high-end manufacturing and safety inspection.

Method used

A dual-domain artifact correction method is adopted, which integrates multi-level data with physical priors. Projection data is acquired through multi-level slow-switching scanning to construct a multi-channel virtual image. Deep learning networks in the projection domain and image domain are combined, and metal trajectory masks and the Swing Transformer architecture are used for collaborative correction. Combined with a composite loss function and a dynamic training optimization strategy, high-precision and high-robust artifact correction is achieved.

Benefits of technology

It effectively removes diffuse metal artifacts, improves imaging clarity and feature recognition, and is suitable for fields such as medical, industrial, security and aerospace. It has high adaptability and real-time performance, outputs accurate images that conform to physical laws, reduces detection risks and improves efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121329834B_ABST
    Figure CN121329834B_ABST
Patent Text Reader

Abstract

This invention discloses a dual-domain artifact correction method based on the fusion of multi-level data and physical priors. By constructing a complete technical solution combining multi-level data acquisition, dual-domain collaborative correction, and physical model constraints, it achieves efficient and accurate removal of diffuse metal artifacts. Its beneficial effects cover multiple fields such as medical, industrial, security, and aerospace, and it has strong versatility and adaptability. Based on a multi-level slow-switching scanning protocol, it completes full-angle scanning at least two energy levels by dynamically adjusting the X-ray source parameters. The acquired complete multi-level projection data is transformed into a high-dimensional tensor through an image fusion method that stitches together channels. The constructed multi-channel virtual image completely preserves the attenuation characteristics and structural information of the target object at each energy level, providing a more comprehensive and discriminative input source for deep learning networks. This fundamentally strengthens the data foundation for accurate correction in different fields and adapts to various imaging scenarios containing metal targets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of industrial CT imaging and non-destructive testing technology, and discloses a dual-domain artifact correction method based on the fusion of multi-level data and physical priors. Background Technology

[0002] Industrial computed tomography (ICT) is a core technology for non-destructive testing of key components in critical fields such as nuclear industry, aerospace, and automobile manufacturing. Its unique advantage of non-destructively revealing the internal three-dimensional structure and defects of objects is directly related to the safety of the entire system. However, when dealing with complex industrial workpieces containing high atomic number, high density metals such as tungsten, platinum, and lead, or metal components such as titanium alloys and cobalt-chromium alloys, ICT technology faces severe challenges: X-ray photons encounter strong photoelectric absorption and Compton scattering when passing through these regions, leading to two core problems: "photon starvation" and "beam hardening." "Photon starvation" results in insufficient X-ray photons penetrating the metal region, drastically deteriorating the detector signal-to-noise ratio and causing bright striped artifacts in the reconstructed image; "beam hardening," due to the contradiction between the "hardening" of the X-ray energy spectrum and the single-energy assumption of the reconstruction algorithm, produces problems such as dark bands, cup-shaped artifacts, and boundary "overflow." These two effects, combined with scattering noise, form complex diffuse and mottled noise. This not only masks the fine structures such as microcracks and pores in the adjacent areas of the metal, but also blurs the precise boundaries of the metal itself. This greatly interferes with defect identification, size measurement, and material interface judgment, seriously affecting the accuracy and reliability of the detection.

[0003] To address the long-standing technical challenge of metal artifacts, academia and industry have developed three main correction methods. Algorithms based on projection data interpolation (such as LI-MAR) are a classic approach. They identify contaminated ray paths in a sine wave, replace the "bad" data with surrounding uncontaminated data, and then reconstruct the image. While computationally efficient and easy to implement, this approach suffers from drawbacks such as compromising the integrity of projection data, introducing new errors, difficulty handling large-scale complex multi-metal artifacts, and sacrificing details in non-metallic areas. Algorithms based on iterative reconstruction (such as the MAP algorithm) integrate correction into the reconstruction process. By constructing a system model incorporating physical effects and prior constraints, they iteratively optimize the matching degree between the image and measured data. Theoretically, this results in better image quality and stronger robustness to noise and missing data. However, it requires multiple orthographic and backprojection operations, leading to extremely high computational costs and reconstruction time far exceeding that of analytical methods, making it difficult to meet the efficiency requirements of industrial inspection. The emergence of deep learning methods in recent years has opened up a new paradigm. Representative works such as CNNMAR, ADNNet, and DuDoNet achieve correction by learning data mapping relationships. In particular, the dual-domain learning framework has significantly improved performance. However, it still faces bottlenecks in real-world industrial scenarios: it relies heavily on high-quality paired training data; the "domain difference" between simulation data and real-world scenarios affects generalization ability; the model lacks constraints from physical laws, which may lead to structural distortions that violate physical laws; and its artifact suppression ability and robustness are insufficient when facing extreme conditions such as high metal filling rates and multi-metal mixtures. The high complexity of advanced models contradicts the real-time requirements of industrial inspection. Therefore, existing technologies are insufficient in terms of data dependence, physical consistency, adaptability to extreme conditions, and computational efficiency. There is an urgent need to develop a new generation of metal artifact correction methods that deeply integrate physical prior knowledge, effectively reduce the dependence on paired data, and balance high accuracy and high reliability, in order to overcome the application bottlenecks of ICT technology in high-end manufacturing and safety inspection fields. Summary of the Invention

[0004] This invention aims to solve the problem of severe metal artifacts caused by high-density metal components in industrial CT inspection. Existing deep learning methods suffer from three main drawbacks: over-reliance on perfectly paired training data, lack of physical model constraints leading to distorted results, and insufficient robustness under extreme conditions. To address these shortcomings, this invention proposes a dual-domain artifact correction method based on the fusion of multi-level data and physical priors. By organically combining multi-energy spectral scanning with physical priors, high-precision and highly robust metal artifact correction is achieved without relying on perfectly paired data.

[0005] The core of this invention lies in constructing a dual-domain deep learning network model called DudoNet. This model achieves end-to-end optimization from data acquisition to image reconstruction through collaborative optimization of the projection domain and image domain, combined with prior knowledge of X-ray energy spectrum physics. The overall architecture of this invention comprises four key components: a multi-level data acquisition module, a dual-domain deep learning network, a physically constrained loss function, and a dynamic training optimization strategy. This invention is achieved through the following technical solutions:

[0006] This invention discloses a dual-domain artifact correction method based on the fusion of multi-level data and physical priors, comprising the following steps:

[0007] S1.1 Multi-level data acquisition: Control the CT scanning equipment and use the multi-level slow switching scanning protocol to acquire multi-level projection data of the object under test at at least two different X-ray energy levels;

[0008] S1.2 Virtual Image Construction: Based on multi-level projection data, a multi-channel multi-level virtual image is constructed using an image fusion method, which serves as the input to the deep learning network;

[0009] S1.3 Dual-domain deep learning network construction: A composite loss function containing the physical model of X-ray energy spectrum attenuation is adopted, and the model is trained and optimized using multi-level virtual images and corresponding artifact-free reference images to obtain a dual-domain deep learning network model.

[0010] S1.4 Dual-domain Cooperative Correction and Image Output: The multi-level virtual image is input into the trained dual-domain deep learning network model. Through the cooperative processing of the projection domain sub-network and the image domain sub-network, a CT image with suppressed metal artifacts is output.

[0011] As a further improvement, the multi-level slow switching scanning protocol of the present invention specifically means: by dynamically adjusting the tube voltage of the X-ray source, full-angle scanning is completed sequentially at multiple energy levels;

[0012] To acquire multi-level projection data of the object under test at at least two different X-ray energy levels, the following steps are taken: at each energy level, complete projection data of the object under test is collected to ensure that the data covers the 360-degree scanning range.

[0013] As a further improvement, the image fusion method in S1.2 of this invention specifically involves: stitching together projection data or reconstructed images of different energy levels along the channel dimension to generate a two-dimensional tensor. , where C represents the number of energy levels.

[0014] As a further improvement, the dual-domain deep learning network model of the present invention includes:

[0015] Projection domain subnetwork: Using a metal trajectory mask as a spatial guiding clue, it performs targeted repair on projection data containing metal artifacts; the metal trajectory mask is generated by identifying abrupt regions caused by metal in the projection data, and is used to locate the root cause of artifacts and constrain the repair range; Image domain subnetwork: Receives the preliminary corrected image output by the projection domain subnetwork, removes residual artifacts and restores image details through fine-grained feature learning; The subnetwork adopts a cross-scale feature interaction mechanism to fuse high-frequency detail information with low-frequency structural information to improve repair accuracy.

[0016] As a further improvement, the projection domain sub-network described in this invention specifically includes:

[0017] Dual-input collaborative mechanism: The metal trajectory mask and the projection data containing artifacts are fed into the network in parallel as dual inputs, and the mask information participates in the feature extraction and repair process throughout the entire process;

[0018] Structured repair module: It adopts an encoder-decoder structure. In the encoder stage, it uses the feature weights of the metal trajectory region through the attention mechanism. In the decoder stage, it introduces reference features of the non-artifact region through skip connections.

[0019] As a further improvement, the image domain sub-network described in this invention specifically includes:

[0020] Feature extraction module: Employs a Swing Transformer-based encoder to extract image features through window multi-head self-attention and shifted window multi-head self-attention mechanisms;

[0021] Image reconstruction module: A decoder based on Swing Transformer is used to upsample and reconstruct the image features obtained by the encoder.

[0022] As a further improvement, the feature extraction module of the present invention specifically includes: an encoder structure: an encoder composed of 4 Swing Transformer Layers, each Layer including layer normalization, multi-head self-attention of windows, multi-head self-attention of shifted windows and multi-layer perceptron; and a decoder structure: a decoder composed of 4 Swing Transformer Layers, which upsamples and reconstructs the encoded features.

[0023] As a further improvement, the composite loss function in S1.3 of this invention is specifically as follows:

[0024] The loss function is composed of: ;

[0025] The meanings of each parameter are as follows:

[0026] Pixel-level loss: Calculate the pixel-level differences between the output image and the target image;

[0027] Perceived loss: Calculate feature map differences based on pre-trained VGG networks;

[0028] Physical loss: Consistency loss is calculated based on a physical model of energy spectrum decay.

[0029] As a further improvement, the physical loss described in this invention specifically includes: theoretical attenuation coefficient calculation: based on the known X-ray energy spectrum distribution, detector response matrix and multi-level projection data, the theoretical attenuation coefficient distribution of the object under test is obtained through analytical calculation; consistency constraint: the difference between the network-predicted attenuation coefficient and the theoretical attenuation coefficient is calculated as the physical consistency loss.

[0030] As a further improvement, the S1.3 model training optimization described in this invention specifically includes: adopting a dynamic training ratio adjustment strategy: dual-domain loss monitoring: calculating the average percentage loss of the projection domain sub-network and the image domain sub-network within a preset training period; gradient difference calculation: calculating the loss gradient difference between the two domains based on the average percentage loss; dynamic adaptation of iteration ratio: dynamically adjusting the ratio of the number of training iterations of the two sub-networks according to the loss gradient difference. When the loss gradient of a certain sub-network is large, its number of training iterations is increased to accelerate convergence, thereby achieving dual-domain collaborative optimization.

[0031] The beneficial effects of this invention are as follows:

[0032] This invention achieves efficient and accurate removal of diffuse metal artifacts by constructing a complete technical solution combining multi-level data acquisition, dual-domain collaborative correction, and physical model constraints. Its beneficial effects cover multiple fields such as medical, industrial, security, and aerospace, demonstrating strong versatility and adaptability. Based on a multi-level slow-switching scanning protocol, it completes full-angle scanning at at least two energy levels by dynamically adjusting X-ray source parameters. The acquired complete multi-level projection data is transformed into a high-dimensional tensor through an image fusion method that stitches together channels. The constructed multi-channel virtual image fully preserves the attenuation characteristics and structural information of the target object at each energy level, providing a more comprehensive and discriminative input source for deep learning networks. This fundamentally strengthens the data foundation for accurate correction in different fields and adapts to various imaging scenarios containing metal targets.

[0033] This invention employs a dual-domain deep learning network that, through the collaborative processing of the projection domain and the image domain, constructs a complete correction chain from artifact root cause suppression to detail restoration. The projection domain sub-network, guided by a metal trajectory mask, specifically repairs artifact-laden projection data, blocking artifact propagation. The image domain sub-network, based on the Swing Transformer architecture, achieves cross-scale feature interaction, thoroughly removing residual artifacts while accurately preserving target details. This advantage is not only applicable to medical CT images but also meets the needs of internal defect identification in metal components in industrial non-destructive testing, item identification under metal interference in luggage / cargo during security checks, and imaging inspection of metal parts in the aerospace field. It significantly improves the structural clarity and feature recognition of images in different scenarios, resolving the core contradiction between "artifact interference" and "detail loss" in imaging various metal-containing targets.

[0034] During model training, a composite loss function incorporating a physical model of ray energy spectrum attenuation provides strict physical constraints for network optimization. Through the organic combination of pixel-level loss, perceptual loss, and physical loss, it ensures both the visual naturalness of the image and avoids false structures that violate imaging physics. This characteristic enables it to output accurate images that conform to actual physical characteristics in fields with extremely high requirements for physical consistency, such as geological exploration (imaging analysis of areas containing metallic minerals) and industrial material testing (internal structural characterization of metal matrix composites). This provides a reliable basis for subsequent data analysis, defect identification, and resource exploration, eliminating decision-making errors caused by image distortion.

[0035] This technology employs a dynamic training ratio adjustment strategy to achieve collaborative optimization of dual-domain sub-networks. By monitoring the difference in loss gradients, it adaptively adjusts the training iteration ratio, accelerating model convergence and improving overall performance. This solution boasts strong compatibility with existing X-ray imaging equipment (CT, DR, industrial inspection X-ray equipment, etc.), requiring no large-scale hardware modifications and quickly adapting to the imaging needs of different fields. Whether in medical clinical diagnosis and treatment, industrial product quality control, aerospace component inspection, or geological resource exploration, it provides clear, accurate, and reliable image support, helping various fields improve detection efficiency and reduce decision-making risks. It possesses broad cross-domain application value and promising prospects for widespread adoption. Attached Figure Description

[0036] Figure 1 Flowchart for using dual-domain deep learning to remove artifacts from CT images;

[0037] Figure 2 The experimental results and metal artifact removal results of 2DectCT data are presented publicly. Detailed Implementation

[0038] To provide a detailed description of the present invention, the technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. Furthermore, for clarity, examples and corresponding explanations will be given.

[0039] This invention designs a deep learning method for CT metal artifacts based on prior physical information fusion, the main process of which is as follows: Figure 1 As shown, the specific steps are as follows:

[0040] S1: Select a suitable X-ray source based on the specific test piece, determine the material of the anode target of the X-ray source, consult the manual to determine the characteristic peak energy value of the anode material, and at the same time determine the approximate proportion of the material filling the test piece, the structural complexity, and the size and shape as prior information for subsequent steps.

[0041] The data used includes two core types of data:

[0042] 1) Real industrial workpiece scanning data (such as high-temperature alloy turbine blades in the aerospace field, aluminum alloy cylinder blocks in the automotive industry, and stainless steel pipe fittings in the nuclear power field), a total of 120 different structural workpieces were collected, and multi-level projection data were obtained for each workpiece under 3-5 different CT equipment parameters (tube voltage 40-300Kev, tube current 50-200mA).

[0043] 2) Simulated data generation (projection data with different metal contents and different artifact intensities simulated based on the Monte Carlo method, used to supplement extreme working condition samples, such as dense workpieces with a metal content of more than 40% and thin metal sandwich structures with severe photon starvation).

[0044] All data is precisely labeled, including:

[0045] ① Metal region mask (using the "threshold segmentation + manual correction" method to ensure that the metal boundary positioning error is ≤1 pixel);

[0046] ② Artifact-free reference images (for non-metallic workpieces or low-artifact workpieces with metal content <5%, obtained by averaging multiple scans; for workpieces with high metal content, synchrotron radiation CT scan data are used as the gold standard).

[0047] ③ Label key quantitative indicators (such as the size of internal defects in the workpiece and the CT value range of density uniformity areas). Perform consistency correction on the original projection data (including detector dark current correction, gain correction, and scattering correction) to ensure that the data from different devices and different scanning batches are comparable, and control the grayscale deviation of the preprocessed projection data within ±2%.

[0048] S1.1 Multi-level data acquisition: Control the CT scanning equipment and use a multi-level slow switching scanning protocol to acquire projection data of the object under test at at least two different X-ray energy levels.

[0049] Prior Information Acquisition and Customized Scanning Scheme Design: Anode Target Characteristic Energy Spectrum Analysis: Before initiating the scan, the operator must first select the optimal X-ray source anode target based on the possible metal composition of the workpiece being tested. For example, for heavy metals (such as tungsten and lead), a tungsten target is recommended to effectively penetrate using its high-energy characteristic X-rays (such as Kα1 approximately 59.3 keV and Kβ1 approximately 67.2 keV); for light metals (such as aluminum and titanium), a molybdenum target (Kα1 approximately 17.5 keV) or a silver target (Kα1 approximately 22.1 keV) can be selected. By consulting the X-ray Data Handbook or the NIST database, the characteristic X-ray energy values ​​and relative intensities of the selected target are accurately obtained as the basis for subsequent physical models.

[0050] Preliminary assessment of the material composition and structure of the test component: Using design drawings, process documents, or results from previous non-destructive testing (such as ultrasonic testing), obtain the approximate mass density, equivalent atomic number, structural complexity (e.g., presence of thin walls or cavities), and overall dimensions of the internal filling material of the test component. This step aims to preliminarily assess the severity of beam hardening and scattering. For example, for an aluminum alloy casting containing internal steel fasteners, the volume ratio of steel to aluminum, their spatial distribution, and the maximum wall thickness of the casting need to be estimated in advance.

[0051] Energy Level and Sequence Optimization: Based on the prior information mentioned above, the energy spectrum distribution and signal-to-noise ratio received by the detector after X-ray penetration of the workpiece are calculated using Monte Carlo simulation (e.g., using Geant4 software) or empirical formulas at different tube voltages (kVp). The key implementation point is that at least two selected energy levels must be able to form an effective "energy contrast." Typically, a low energy level (e.g., 80-120 kVp) is used to obtain high-contrast images but is susceptible to artifacts, and a high energy level (e.g., 140-220 kVp) is used to ensure penetration capability into high-density areas. For complex workpieces, a three-level protocol (e.g., 80 kVp, 150 kVp, 220 kVp) can be used. The scanning sequence adopts a "slow switching" mode, that is, after exposing all preset energy levels sequentially at the same projection angle, the turntable rotates to the next angle to ensure strict spatial alignment of different energy level data.

[0052] Precise control of tube voltage and current: Dynamic switching of tube voltage (kVp) is achieved through programming via the CT equipment's API interface or control scripts. To ensure the stability of the X-ray output after each switch, a sufficient settling time (typically 100-500 milliseconds) is set, and the tube current (mA) is monitored in real time to ensure it meets the preset value. Implementation example: For aluminum alloy castings, the sequence is set as follows: [Angle θ: Exposure @ 100kVp / 2mA -> Stabilization wait 200ms -> Exposure @ 180kVp / 1.5mA -> Turntable rotation Δθ].

[0053] Full-angle data acquisition and geometric calibration: At each energy level, it is essential to ensure that the projected data covers the entire 360-degree range. The number of views must satisfy the Shannon-Nyquist sampling theorem, typically no less than 1000. Immediately before and after scanning, perform system geometric calibration using standard components (such as steel ball phantoms) to precisely calibrate parameters such as source-to-detector distance (SDD), source-to-object distance (SOD), detector pixel size, and rotation center offset. These calibration parameters are then embedded in the original data header file for subsequent reconstruction.

[0054] Dark-field and bright-field correction: Before and after each energy level scan, dark-field images (X-ray off) and bright-field images (no sample) at the corresponding energies are acquired to normalize and correct the raw projection data, eliminating the effects of detector dark current and response inconsistencies. The correction formula is: P_corrected = -log( (P_raw - Dark) / (Flat - Dark)), where P is the projection value.

[0055] S2: Virtual Image Construction: Based on the multi-level projection data, a multi-channel multi-level virtual image is constructed using image fusion technology, which serves as the input to the deep learning network.

[0056] Preliminary image reconstruction and energy spectrum feature extraction:

[0057] For the corrected projection data (P_corrected) at each energy level, an analytical reconstruction algorithm (such as filtered back projection FBP or FDK algorithm) is first used to quickly reconstruct the data, resulting in a set of initial CT images I_i(x, y) at different energy levels, where i represents the energy level index.

[0058] Due to the presence of metal artifacts, these I_i images are of poor quality, but they contain differentiated information about the interaction between X-rays of different energies and matter. Key to implementation: All I_i images must undergo rigorous image registration to ensure a one-to-one spatial correspondence. Utilizing the inherent geometric alignment during scanning, typically only sub-pixel-level fine-tuning is required.

[0059] This method does not directly use I_i, which suffers from severe artifacts, but instead uses it as a foundation to construct a more information-rich, multi-channel virtual image that is more conducive to network learning. Assuming we have N energy levels, we construct a virtual image V with M channels (typically M >= N).

[0060] The channel construction strategy is as follows:

[0061] Channels 1-N: Directly store the original reconstructed images I_1, I_2, ..., I_N for each energy level. Channel N+1: Effective atomic number (Z_eff) mapping. Utilizing the principle of dual-energy or tri-energy CT, the equivalent atomic number of each voxel is calculated based on the CT value ratio of the images at each energy level, using a pre-calibrated matrix material decomposition model (such as the water-iodine model or the photoelectric-Compton scattering model). This channel significantly enhances the network's ability to distinguish material composition.

[0062] Channel N+2: Electron density (ρ_e) mapping. Also obtained through base matter decomposition, it provides information on matter density and is an important input for physical constraints. Channel N+3: Primary metal artifact estimation map. By performing simple global thresholding on the high-energy image I_high, a binary mask of the metal region is extracted. This mask is then used to perform morphological operations and radial diffusion on the low-energy image I_low, generating a coarse artifact distribution map to guide the network to focus on the problem region.

[0063] Ultimately, the virtual image V is an H x W x M tensor (H and W are the image height and width), which integrates the original information of the multi-energy spectrum and the derived physical features, providing an information dimension far exceeding that of the single-energy image for subsequent deep learning.

[0064] S3: Dual-domain collaborative correction: The multi-level virtual image is input into a pre-trained dual-domain deep learning network model. Through the collaborative processing of the projection domain sub-network and the image domain sub-network, a CT image with suppressed metal artifacts is output.

[0065] The constructed multi-channel virtual image V is input into Dudo-Net. The network first passes through a shared feature extraction module, consisting of three 3x3 convolutional layers, to extract basic, domain-independent fusion features from V. Subsequently, the feature stream is copied and fed into the projection domain sub-network (P-Net) and the image domain sub-network (I-Net), respectively.

[0066] Deep analysis of the Projection Domain Subnetwork (P-Net): Input / Output Representation: The input to P-Net is not the original sine map, but a virtual projection defect map generated by mapping shared features through a 1x1 convolutional layer. Its output is the corrected virtual projection map.

[0067] An improved U-Net encoder architecture: The encoder consists of 5 layers, each containing two densely connected blocks. Each dense block comprises three 3x3 convolutions, and the output of each convolutional layer is directly connected to the input of all subsequent layers within the same dense block (channel concatenation). This design greatly enhances gradient flow and feature reuse. Each dense block is followed by a 2x2 max pooling operation for downsampling.

[0068] Attention-Gated Decoder: The decoder also has 5 layers. Each layer first performs 2x2 transposed convolutional upsampling, and then concatenates the upsampled result with the corresponding skip connection features from the encoder. The key innovation is that before concatenation, the skip connection features pass through a spatial attention-gated module. This module uses the current features of the decoder as the gating signal to generate a spatial weight map, highlighting areas that need to be repaired, such as metal trajectories. Formulaically, this is expressed as: α = σ(ψ_enc(X_enc) + ψ_dec(X_dec)), where σ is the Sigmoid function and ψ is the convolutional layer. The weighted feature is X_enc_att = X_enc ⊙ α.

[0069] Injection of metal mask prior: In skip connections, the binary metal mask generated in step S2 is upsampled / downsampled to the same size as the current feature map, and concatenated with X_enc_att as an extra channel before being fed into the decoder convolutional block. This provides the network with an explicit geometric prior.

[0070] Input / Output Representation: The input to I-Net is shared features, and the output is the final corrected CT image. Swin Transformer-based Backbone Network: Layered Design: The network is divided into 4 stages. Stage 1: The input image is segmented into 4x4 patches, passed through a linear embedding layer, and then fed into a Swin Transformer Block. This stage uses a window size M=4 and focuses on extracting local features such as edges and textures.

[0071] Stages 2 / 3 / 4: Each stage begins with a Patch Merging layer (similar to convolutional downsampling) to reduce resolution and increase the number of channels, followed by Swin TransformerBlocks with window sizes of M=8, 16, and 16 respectively. The deep, large windows enable the network to establish long-range dependencies and understand the global distribution patterns of artifacts throughout the image.

[0072] Adaptive Attention Span: In the window self-attention computation of each Swin Transformer Block, a learnable relative position bias B is introduced. Attention = Softmax(QK^T / √d + B). The span of B (i.e., the range of relative distances it covers) is not fixed, but dynamically adjusted according to the average activation of the current feature map, automatically expanding the receptive field in regions with complex features.

[0073] Cross-Scale Feature Interaction Module (CFIM): A bidirectional feature pyramid is introduced at the decoding path (if needed) or at the end of the network. It receives multi-scale features from Stages 2, 3, and 4, and performs feature fusion through top-down and bottom-up pathways combined with lateral connections. The final output is a high-quality feature map that integrates global context and local details, which is then passed through a 1x1 convolution to generate the final image.

[0074] The corrected virtual projection map output by P-Net can be reconstructed into an image using a differentiable projection operator (such as a neural network approximation of Radon transform).

[0075] The reconstructed image is weighted and fused with the output image of I-Net at the feature level, or a residual connection is used. Simultaneously, a dual-domain consistency loss is designed to enforce content consistency between the P-Net-reconstructed image and the direct output of I-Net. This design allows the two sub-networks to supervise each other and collaboratively optimize during training, rather than working independently.

[0076] S4: Model Training Optimization: The dual-domain deep learning network model is trained using a composite loss function, which includes a physical consistency loss term based on the physical model of X-ray energy spectrum decay.

[0077] The construction and calculation of the composite loss function:

[0078] The total loss function L_total is a weighted sum of multiple loss terms:

[0079] L_total = λ_phy * L_physical + λ_adv * L_adversarial + λ_perc * L_perceptual + λ_rec * L_rec

[0080] Physical consistency loss (L_physical): This is the core of this invention. It is based on discretized X-ray energy spectra and Beer-Lambert's law. Let μ(E, x) be the linear decay coefficient of energy E at spatial position x, S(E) be the normalized X-ray source energy spectrum, and P be the projection value. Theoretically, we have P = -log(∫ S(E) exp(-∫ μ(E, x) dl ) dE ).

[0081] Implementation: The corrected image output by the network can be considered as an estimate of μ, μ_est. We map μ_est to the basis material coefficients a_1, a_2 through basis material decomposition (e.g., μ_est(E, x) ≈ a_1(x) * f_p(E) + a_2(x) * f_C(E), where f_p and f_C are known energy dependence functions of the photoelectric effect and Compton scattering, respectively). Then, using these coefficients and the known S(E), we forward project to calculate the synthesized projection data P_synth.

[0082] Loss calculation: L_physical = || P_measured - P_synth ||_2^2 + TV(a_1) + TV(a_2). Where P_measured is the relatively reliable projection data acquired at high energy levels. This loss term forces the network to output a physically plausible image whose projection should match the actual measurements. The total variational (TV) regularization term is used to promote the smoothness of the results.

[0083] Adversarial Loss (L_adversarial): A discriminator network D is used, whose goal is to distinguish the network's output image from a "clean" real CT image (a reference image without metal artifacts, obtained through high-dose scanning or numerical simulation). The generator (i.e., Dudo-Net) aims to fool D. L_adversarial = E[log(1 - D(G(V)))], where G is Dudo-Net. This helps generate visually more realistic and naturally textured images.

[0084] Perceptual Loss (L_perceptual): Utilizes a VGG network pre-trained on a large image dataset (such as ImageNet). It calculates the difference between the feature maps of the network's output image and the target image at a specific intermediate layer of the VGG network (e.g., ReLU3_3): L_perceptual = || Φ(G(V)) - Φ(I_target) ||_2^2. This loss guides the network to reconstruct results consistent with the real image at the semantic feature level, helping to recover high-frequency details.

[0085] Pixel-level reconstruction loss (L_rec): The most basic loss, such as the L1 or L2 norm, directly constrains how close the output image is to the target image in terms of pixel values. L_rec = || G(V) - I_target ||_1.

[0086] The operation of the dynamic training optimization system:

[0087] Multi-objective weight optimization: The initial weights λ_i are determined on a small validation set through grid search or Bayesian optimization to find an equilibrium point on the Pareto front. During training, a dynamic weight adjustment strategy is employed. For example, every K epochs, the rate of change r_i = L_i / L_i^0 of each loss term L_i relative to its initial value L_i^0 is calculated. Then, the weights are adjusted according to λ_i' = λ_i * (r_i / (∏ r_j)^{1 / N}), giving higher weights to loss terms that decrease more slowly, thereby balancing the learning progress of each task.

[0088] Reinforcement learning-driven hyperparameter tuning: Modeling the entire training process as a Markov decision process (MDP).

[0089] State: Includes the current epoch, each loss value, gradient norm, validation set PSNR / SSIM, etc.

[0090] Action: Make small increases or decreases to the learning rate, batch size, loss weight λ_phy, etc.

[0091] Reward: Defined as the improvement in validation set performance metrics (such as ΔPSNR) minus the penalty for training time.

[0092] A lightweight agent (such as using the PPO algorithm) interacts with the environment (during the training process) and learns how to dynamically adjust hyperparameters to achieve optimal performance as quickly as possible. S5: System Deployment and Inference

[0093] After training is complete, save the optimal Dudo-Net model weights.

[0094] In actual industrial CT inspection, for a new test piece, only steps S1 and S2 need to be performed to obtain its multi-level projection data and construct a virtual image.

[0095] By inputting the virtual image into the pre-trained Dudo-Net and performing a single forward propagation, a high-quality CT image with significantly suppressed metal artifacts can be output within seconds, which can then be used for subsequent tasks such as defect analysis and dimensional measurement.

[0096] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

[0097] Summary of the advantages of the S8 technical solution and comparison with existing methods

[0098] Experimental results and metal artifact elimination using publicly available 2DectCT data are as follows: Figure 2 As shown, Figure 2 The image reconstruction effects of filtered backprojection, GAN network, U network, and dual-domain network were compared from three dimensions: reconstruction results, difference quantification, and region of interest (ROI) details. Intuitively, the dual-domain network reconstruction results showed the highest degree of fit to the ground truth in terms of overall shape and detail. In the difference map, the dual-domain network only displayed dark blue (bias <10%), representing low differences, which was far superior to the high proportion of red and yellow areas in filtered backprojection (bias of 20%–30%) and the obvious yellow and orange areas in GAN and U networks (bias of approximately 10%–20%). The magnified ROI also showed that the dual-domain network's restoration of texture and structure far exceeded other methods, fully demonstrating that the improved dual-domain network performed better in image reconstruction tasks. By comparing the core characteristics with traditional methods and existing deep learning methods, the technical breakthroughs of this invention were clarified, as shown in Table 1.

[0099] Table 1 compares the core characteristics of different CT image artifact correction methods.

[0100] Comparison Dimensions Traditional iterative reconstruction methods (such as ART) Single-domain deep learning methods (such as U-Net) Physically Constrained Two-Domain Method This invention (physical information fusion dual-domain method) Data dependency No labeled data is required, but a large number of prior parameters (such as attenuation coefficients and regularization parameters) are needed. Relying on a large amount of "artifact-without-artifact" pairing data (≥100,000 pairs) It relies on multi-level data, but requires no physical parameters. Only multi-level projection data and a few physical parameters (such as metal material and energy spectrum range) are required; no pairing annotation is needed. Physical consistency There is a fundamental physical model (Beer-Lambert law), but the details of energy spectrum hardening are ignored. Without physical constraints, CT values ​​are prone to distortion (deviation of 15-20 HU). Without physical constraints, the quantitative accuracy is low (ΔHU±12HU). The integrated multi-spectral attenuation model shows a CT value deviation of ≤±5HU, meeting industrial quantitative standards. Robustness under extreme conditions Failure under photon starvation scenarios (artifact suppression rate <50%) When the metal content is greater than 30%, the performance drops sharply (PSNR < 28dB). Poor generalization ability in extreme scenarios (SSIM fluctuation > 0.1) In extreme scenarios, artifact suppression rate is ≥80%, PSNR is ≥32dB, and robustness is improved by more than 20%. Industrial applicability Slow processing speed (single image > 5s), does not support real-time detection. It is fast (<0.5s), but requires retraining for different devices. It's fast, but has poor device compatibility. Fast speed (<0.3s), supports multi-device adaptation, no retraining required. Preservation of complex structural details Details that are easily blurred (e.g., the rate of missed detection of 0.2mm pores > 30%) Detail is well preserved, but artifacts are prone to appear at the edges. Detail retention is average, but metal edges are distorted. Defect miss rate <5%, metal boundary error ≤1 pixel, optimal detail preservation.

[0101] S9 technical solution expansion and advanced optimization

[0102] To further expand the application scenarios and performance limits of the method, the following three advanced directions are designed:

[0103] S9.1 Multimodal Data Fusion Extension

[0104] To address the difficulty of distinguishing between artifacts and real defects using single CT data, multimodal data such as ultrasound and infrared are integrated to improve detection accuracy.

[0105] Multimodal data acquisition: While performing CT scanning, the ultrasonic A-scan data of the workpiece is acquired simultaneously (to obtain the acoustic characteristics of the internal structure and distinguish between artifacts and cracks) and infrared thermal imaging data (to obtain the differences in thermal conductivity of materials and distinguish between artifacts and pores).

[0106] Cross-modal feature fusion module: A "multimodal attention fusion layer" is added after the image domain sub-network of DudoNet to adaptively allocate the weights of CT image features, ultrasound features, and infrared features (e.g., the weight of ultrasound features in the crack area is increased to 0.6) to achieve cross-modal information complementarity.

[0107] Application scenarios: Primarily used for high-precision component inspection in aerospace (such as defect detection of thermal barrier coating on turbine blades), improving defect identification accuracy from 92% to over 98%.

[0108] S9.2 Lightweight Deployment Optimization

[0109] To address the needs of portable CT equipment in industrial settings (limited computing power and insufficient memory), the model was modified to be lightweight.

[0110] Model pruning: A combination of "channel pruning + layer pruning" strategy was adopted to remove redundant convolutional channels in the projection domain subnetwork (retaining 60% of the core channels) and redundant Transformer blocks in the image domain subnetwork (retaining 70% of the core blocks), reducing the model parameters from 120M to 35M;

[0111] Quantization and distillation: The model weights are quantized from 32-bit floating-point to 8-bit integers, while retaining more than 95% of the performance through knowledge distillation (using the original DudoNet as the teacher model and the lightweight model as the student model);

[0112] Deployment results: The lightweight model can run on edge computing devices (such as NVIDIA Jetson AGX), with a single image processing time of <0.8s, meeting the needs of real-time on-site detection, and memory usage of <2GB.

[0113] S9.3 Semi-supervised / Unsupervised Learning Extension

[0114] To address the scarcity of labeled data in industrial settings, this paper expands upon semi-supervised and unsupervised learning models:

[0115] Semi-supervised training strategy: The model is trained using "a small amount of labeled data (10%) + a large amount of unlabeled data (90%)", and the information from the unlabeled data is utilized through consistency regularization (requiring that the correction results of the same workpiece are consistent under different noise disturbances);

[0116] Unsupervised training strategy: Based on the "physical consistency of multi-level data itself", an unsupervised loss is designed (e.g., the decay coefficient of the same region under different energies should meet the theoretical proportional relationship), which can be trained without any labeled data;

[0117] Performance verification: In scenarios where only 5% of the data is labeled, the artifact suppression rate of the semi-supervised mode is still ≥80%, and the performance difference with the fully supervised mode (100% labeling) is <5%, which significantly reduces the cost of data labeling.

[0118] S10 Application Examples and Effect Verification

[0119] Taking "CT inspection of high-temperature alloy turbine blades for aero-engines" as an example, the practical application effect of this invention is verified:

[0120] Workpiece background: The turbine blade is made of GH4169 high-temperature alloy (metal content is about 35%), which contains complex cooling channels (thickness 0.3-0.5mm, which are prone to photon starvation). The defects to be detected include microcracks (0.1-0.3mm) in the cooling channel wall and pores (0.2-0.8mm) in the matrix.

[0121] Testing equipment: Industrial CT equipment (tube voltage range 80-220Kev, detector pixel size 0.1mm), using the multi-level scanning protocol of this invention (energy levels: 80, 120, 160, 220Kev).

[0122] Processing result:

[0123] Artifact suppression effect: Before correction, there were obvious strip-shaped artifacts at the metal boundary of the blade (standard deviation of artifact area σ=45HU), after correction σ=6HU, the artifact suppression rate reached 86.7%;

[0124] Quantitative detection accuracy: The deviation of the CT value of the cooling channel wall was reduced from ±18HU before correction to ±4HU, and the measurement error of 0.2mm microcracks was ±0.03mm, which meets the aviation inspection standards;

[0125] Detection efficiency: The total scanning and correction time for a single leaf (1024×1024 pixels) is 2.5 minutes, which is 50% more efficient than the traditional method (5 minutes).

[0126] Conclusion: This invention can effectively solve the problem of metal artifacts in high-temperature alloy blades, meeting the industrial demand for high-precision and high-efficiency detection.

[0127] This invention overcomes the bottlenecks of existing methods, such as dependence on labeled data, poor physical consistency, and insufficient robustness, by deeply integrating physical information and deep learning. It provides a complete and feasible technical solution for metal artifact removal in industrial CT, which can be widely used in high-precision detection scenarios in aerospace, automotive, nuclear power and other fields, and has important engineering application value.

[0128] It will be understood by those skilled in the art that the above description is merely a single example of the invention and is not intended to limit the invention. Although the invention has been described in detail with reference to the foregoing examples, those skilled in the art can still modify the technical solutions described in the foregoing examples or make equivalent substitutions for some of the technical features. All modifications and equivalent substitutions made within the spirit and principles of the invention should be included within the scope of protection of the invention.

Claims

1. A dual-domain artifact correction method based on the fusion of multi-level data and physical priors, characterized in that, Includes the following steps: S1.1 Multi-level data acquisition: Control the CT scanning equipment and use the multi-level slow switching scanning protocol to acquire multi-level projection data of the object under test at at least two different X-ray energy levels; S1.2 Virtual Image Construction: Based on the multi-level projection data, a multi-channel multi-level virtual image is constructed using an image fusion method, which serves as the input to the deep learning network; S1.3 Dual-domain deep learning network construction: A composite loss function containing the physical model of X-ray energy spectrum attenuation is adopted, and the model is trained and optimized using multi-level virtual images and corresponding artifact-free reference images to obtain a dual-domain deep learning network model. S1.4 Dual-domain collaborative correction and image output: The multi-level virtual image is input into the trained dual-domain deep learning network model. Through the collaborative processing of the projection domain sub-network and the image domain sub-network, a CT image with suppressed metal artifacts is output. The aforementioned multi-level slow switching scanning protocol specifically involves: dynamically adjusting the tube voltage of the X-ray source to sequentially complete a full-angle scan at multiple energy levels; The acquisition of multi-level projection data of the object under test at at least two different X-ray energy levels specifically involves: acquiring complete projection data of the object under test at each energy level to ensure that the data covers the 360-degree scanning range; The image fusion method in S1.2 specifically involves stitching together projection data or reconstructed images of different energy levels along the channel dimension to generate a two-dimensional tensor. , where C represents the number of energy levels; The dual-domain deep learning network model includes: Projection domain subnetwork: Using a metal trajectory mask as a spatial guiding clue, it performs targeted repair on multi-level projection data containing metal artifacts; The metal trajectory mask is generated by identifying abrupt regions caused by metal in the projection data, and is used to locate the root cause of artifacts and constrain the repair range; the image domain sub-network receives the preliminary corrected image output by the projection domain sub-network, removes residual artifacts and restores image details through fine-grained feature learning; the sub-network adopts a cross-scale feature interaction mechanism to fuse high-frequency detail information with low-frequency structural information to improve repair accuracy.

2. The method according to claim 1, characterized in that, The projection domain subnetwork specifically includes: Collaborative dual-input module: The metal trajectory mask and multi-level projection data containing artifacts are fed into the network in parallel as dual inputs, and the mask information participates in the feature extraction and repair process throughout the entire process. Structured repair module: It adopts an encoder-decoder structure. In the encoder stage, it uses the feature weights of the metal trajectory region through the attention mechanism. In the decoder stage, it introduces reference features of the non-artifact region through skip connections.

3. The method according to claim 2, characterized in that, The image domain subnetwork specifically includes: Feature extraction module: Employs a Swing Transformer-based encoder to extract image features through window multi-head self-attention and shifted window multi-head self-attention mechanisms; Image reconstruction module: It uses a Swing Transformer-based decoder to upsample and reconstruct the image features obtained by the Swing Transformer encoder.

4. The method according to claim 3, characterized in that, The feature extraction module specifically includes: an encoder structure: an encoder composed of 4 Swing Transformer Layers, each Layer containing layer normalization, multi-head self-attention for windows, multi-head self-attention for shifted windows, and a multilayer perceptron; and a decoder structure: a decoder composed of 4 Swing Transformer Layers, which upsamples and reconstructs the encoded features.

5. The method according to claim 1, 2, 3, or 4, characterized in that, The composite loss function in S1.3 is specifically as follows: The loss function is composed of: ; The meanings of each parameter are as follows: Pixel-level loss: Calculate the pixel-level differences between the output image and the target image; Perceived loss: Calculate feature map differences based on pre-trained VGG networks; Physical loss: Consistency loss is calculated based on a physical model of energy spectrum decay.

6. The method according to claim 5, characterized in that, The physical losses specifically include: theoretical attenuation coefficient calculation: based on the known X-ray energy spectrum distribution, detector response matrix and multi-level projection data, the theoretical attenuation coefficient distribution of the measured object is obtained through analytical calculation; consistency constraint: the difference between the network-predicted attenuation coefficient and the theoretical attenuation coefficient is calculated as the physical consistency loss.

7. The method according to claim 1, characterized in that, The S1.3 model training optimization adopts a dynamic training ratio adjustment strategy, which specifically includes: dual-domain loss monitoring: calculating the average percentage loss of the projection domain sub-network and the image domain sub-network within a preset training period; gradient difference calculation: calculating the loss gradient difference between the two domains based on the average percentage loss; dynamic adaptation of iteration ratio: dynamically adjusting the ratio of the number of training iterations of the two sub-networks according to the loss gradient difference. When the loss gradient of a certain sub-network is large, its number of training iterations is increased to accelerate convergence, thereby achieving dual-domain collaborative optimization.

Citation Information

Patent Citations

  • A metal artifact correction method for multi-energy spectrum X-ray CT imaging

    CN109146994A

  • Novel method for removing metal artifacts of CT (Computed Tomography) image

    CN118628599A

  • Sparse finite angle CBCT reconstruction method and system based on residual diffusion and storage medium

    CN120510295A