Metal artifact reduction method based on time-varying constraint adversarial generative network model
By constructing a GAN network model with time-domain variable constraints, the problems of poor metal artifact removal and insufficient detail fidelity in existing technologies are solved, achieving better metal artifact removal and detail fidelity, and making it suitable for a wide range of application scenarios.
Patent Information
- Application Number
- CN202310878651.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-18
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2043-07-18
AI Technical Summary
Existing methods for removing metal artifacts fail to meet practical needs in terms of removal effectiveness and detail fidelity. In particular, unsupervised methods perform poorly when dealing with complex artifacts and tissue details, while supervised methods are insufficient in terms of paired data requirements.
A GAN network model with temporally variable constraints (MARGANVAC) is constructed. By introducing time-varying constraint terms, adaptive fidelity constraints are applied to each part of the whole image. The model includes a generator G, a discriminator D, and a registration network R. By using residual learning and self-reconstruction loss functions, combined with time-varying loss and adversarial loss, the generator is optimized to generate more realistic artifact-free images.
It achieves better metal artifact removal and detail fidelity, and is suitable for a wider range of applications, especially performing exceptionally well under real metal artifact conditions.
Smart Images

Figure CN116894783B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to a metal artifact removal method, in particular to a metal artifact removal method based on a time-varying constraint adversarial generative network model, and belongs to the technical field of improving CT image quality. BACKGROUND
[0002] The presence of metal implants can cause serious metal artifacts in CT images, which seriously reduces the image quality. In the past few decades, many methods have been proposed to reduce metal artifacts. Although these traditional methods have produced certain effects in metal artifact removal, there is still a long way to go to meet the actual needs. In recent years, with the rise of deep learning, deep learning models have also been gradually used in the research of metal artifact removal and have produced good results. In reference [1, 2], Park et al. use U-net to learn the mapping from the geometric shape of the metal track area in the projection map to the corresponding beam hardening factor to correct the inaccurate projection data in the projection map. The CNNMAR method combines the uncorrected image, the LI image and the BHC image into a three-channel image input into the convolutional neural network to generate a prior image, and then uses its projection result to guide the interpolation process to produce a picture without artifacts. Based on the idea of conditional generative adversarial network (cGAN), Wang et al. proposed cGANMAR in reference [3], which learns the mapping from the image with artifacts to the image without artifacts, and uses PatchGAN as the discriminator. The above methods are supervised methods that require good paired data as the training set. Many scholars have devoted themselves to the research of unsupervised learning, which has liberated the demand for paired data in the problem of metal artifact removal. Based on the idea of CycleGAN, Lee et al. proposed Attention-guided beta-CycleGAN in reference [4], which uses an attention mechanism to focus on the unique features of metal artifacts in the spatial domain and the channel domain, and is trained in an unsupervised manner, and the results have good robustness. In reference [5], Liao et al. creatively proposed the ADN network, which uses the concept of hidden space, uses different encoders to encode the image with artifacts into the content space containing only image information features, and uses another encoder to encode the image with artifacts into the artifact space containing only artifacts, realizing the decoupling of artifacts and tissue details. The unsupervised method eliminates the demand for paired data, but it does not have good processing capability for complex artifacts and tissue details produced under clinical conditions. The current mainstream supervised method is more superior in performance than the unsupervised method, but it still cannot meet the demand of actual application in metal artifact removal and image detail fidelity.
[0003] REFERENCES
[0004] [1] Park H S, Chung Y E, Lee S M, et al. Sinogram-consistency learning in CT for metal artifact reduction [J]. arXiv preprint arXiv:00607, 2017, 1.
[0005] [2] Park H S, Lee S M, Kim H P, et al. CT sinogram-consistency learning for metal-induced beam hardening correction [J]. Medical physics, 2018, 45(12): 5376-84.
[0006] [3] Wang J, Zhao Y, Noble J H, et al. Conditional generative adversarial networks for metal artifact reduction in CT images of the ear [C]. International Conference on Medical Image Computing and Computer-Assisted Intervention. 2018: 3-11.
[0007] [4] Lee J, Gu J, Ye J C. Unsupervised CT Metal Artifact Learning Using Attention-Guided β-CycleGAN [J]. IEEE Transactions on Medical Imaging, 2021, 40(12): 3932-44.
[0008] [5] Liao H, Lin W-A, Zhou S K, et al. ADN: artifact disentanglement network for unsupervised metal artifact reduction [J]. IEEE Transactions on Medical Imaging, 2019, 39(3): 634-43. SUMMARY
[0009] Objective: To address the problems and shortcomings of existing technologies, this invention provides a metal artifact removal method based on a time-varying constraint-based generative adversarial network (GAN) model. A metal artifact removal model with time-varying constraints (MARGANVAC) is constructed. This model introduces time-varying constraint terms to adaptively constrain the fidelity of various parts of the entire image, thereby more effectively training the generator to produce images with better detail fidelity and metal artifact removal. Compared with current mainstream models, this method not only achieves better artifact removal results but also has a wider range of applications.
[0010] Technical solution: A metal artifact removal method based on a time-varying constraint adversarial generative network model; a metal artifact removal model with time-domain variable constraints (MARGANVAC) is constructed. This model introduces time-varying constraint terms on the GAN network to adaptively constrain the fidelity of each part of the whole image.
[0011] The GAN network model for removing metal artifacts with temporally variable constraints comprises three modules: a generator G, a discriminator D, and a registration network R. The CT image to be processed undergoes random affine transformation and is then input into the generator G, which outputs to the registration network and the discriminator D. The registration network R is used for random sampling, and the neighborhood of the sampled pixels gradually decreases with the increase of the number of iterations. The registration network R can adaptively adjust its parameters without human intervention, thereby enabling the generator to produce more realistic images.
[0012] Furthermore, the training process of the GAN network model for removing metal artifacts with time-domain variable constraints is as follows:
[0013] First, for each input image x with metal artifacts... a Perform a random affine transformation on the image x. a A random affine transformation is performed on the reference image x, which is unaffected by artifacts, to obtain images xi and xj respectively. a Image after x-transformation of reference image and x T When the image After passing through the generator G, we obtain an image with artifacts removed. Will and x T Input the registration network R; the physical meaning of the registration network can be expressed by a function. Let G(x) represent the first part of the function. a ;θ G ) is the output of the generator subnetwork, θ G The parameters of the generator subnetwork are represented by j, where j is the pixel index, and φ in the second part represents an abstract sampling function used to sample pixels x. jSampling is performed on pixels within their neighborhood, σ t σ is a parameter used to control the size of the neighborhood. As training iterates, the parameter σ... t It gradually converges to zero. At the beginning of training, the registration network performs poorly, therefore the output... Compared with the affine transformed ground truth image x t There will be a certain positional deviation between them, and these deviations are variable, meaning that each pixel in each input image has an independent and different deviation in each iteration. This characteristic can help simulate the function φ(x) j ,σ t The random sampling process involves the registration network. Furthermore, as training progresses, the registration performance of the network improves, meaning that the parameter σ... t It will decrease if the entire model converges well, σ t It will drop to zero.
[0014] A pre-trained GAN network with temporally variable constraints is used to generate metal artifact-free images from CT images with metal artifacts.
[0015] Furthermore, residual learning is introduced into the generator G; self-reconstruction is introduced as a constraint to regularize the generator; the generator contains an encoder and a decoder; the encoder is used to map image samples from the image domain to the latent space, where content information and features of metal artifacts are separated; the decoder reconstructs the separated content information into an artifact-free image; between the encoder and decoder, a deep subnetwork consisting of 21 Inception-ResNet modules is introduced to improve the separation ability of artifact features and content information in the latent space.
[0016] Furthermore, in the training process of the GAN network with time-domain variable constraints for removing metal artifacts, self-reconstruction is introduced as a constraint to regularize the generator in order to better distinguish between metal artifacts and content information. During self-reconstruction, the artifact-free image y is stitched together into [y, y], and then input into the generator. Before stitching, a random affine transformation is performed on the artifact-free image y to obtain the result A3(y), which is then stitched together to obtain [A3(y), A3(y)]. Accordingly, in self-reconstruction, the generator output becomes: G([A3(y), A3(y)]; θ G The reference image y that is unaffected by artifacts is also y;
[0017] Furthermore, for image x a The linear interpolation (LI) method is used to obtain the estimated value of the metal trajectory portion in the projected image in the sinusoidal domain, and the LI-corrected reconstructed image x is obtained by the FBP or FDK method. [LI]a The image x corrected by the LI method[LI]a and the image x affected by artifacts a respectively, and the affine transformation results A1(x [LI]a ) and A1(x a ), A1(x [LI]a ) and A1(x a ) are connected in the channel dimension as the input of the generator G.
[0018] Further, the residual learning is introduced in the generator G, and the image after removing the metal artifacts is represented as: [x a , x [LI]a ] represents the connection operation of x a and x [LI]a ; in the self-reconstruction, the generator output becomes and the concatenation of A1(x ) and A2(x) is taken as the input of the registration network R, A2(x) represents the random affine transformation of the reference image x, and the output of the registration network R is a deformation vector field T x , θ R is the registration network parameter, after obtaining the deformation vector field T x , the resampled image x is obtained by applying T x to the input image x According to the same steps as generating T x , the deformation vector field T y is obtained. A4(y) represents the random affine transformation of the reference image y; and in the case of self-reconstruction, the corresponding resampled image x The generator output and A2(x) are taken as the input of the discriminator D, and in the case of self-reconstruction, the generator output and A4(y) are taken as the input of the discriminator D.
[0019] Further, the GAN network with time domain variable constraint metal artifact reduction model contains four loss functions, which are artifact correction loss self-reconstruction loss adversarial loss and smooth loss for constraining the deformation vector field The total loss is represented as the weighted sum of these losses Each λ represents the weight coefficient of the corresponding loss.
[0020] Correction loss The network R and the generator G are trained simultaneously, and the correction loss is designed as:
[0021]
[0022] The L1 norm is used to make the generator retain more details while reducing metal artifacts, represents the expectation, represents that x comes from the dataset The domain of the CT image affected by metal artifacts is
[0023] The self-reconstruction loss The training goal is to make the generator learn to remove metal artifacts, while more constraints need to be introduced to make the generator retain more content information; that is, to reduce metal artifacts as much as possible in the presence of metal and artifacts, and to retain all image content information as much as possible in the absence of metal artifacts.
[0024]
[0025] wherein, represents the expectation, represents that y comes from the dataset The domain of the image not affected by artifacts is
[0026] The adversarial loss Adversarial learning encourages the generator G to generate more realistic artifact-free images. To achieve this goal, the generator should have the ability to distinguish between artifact information and content information in the latent space. To enhance this ability, two adversarial learning strategies are introduced. One strategy is to improve the ability of the latent space to recognize metal artifacts, and the other strategy is to improve the ability of the latent space to preserve content information. The input data of the two strategies are images affected by metal artifacts and images without metal artifacts, respectively. The two strategies are executed simultaneously in adversarial learning, and their respective losses can be written as:
[0027]
[0028]
[0029] The base of the logarithm is 2, and D(·) represents the output of the discriminator D; therefore, the total adversarial loss is:
[0030]
[0031] Smooth loss In minimizing the combined loss of the correction loss and the self-reconstruction loss, the resulting deformation vector field can become non-smooth, which is physically unrealistic. To solve this problem, a diffusion regularization operator in the direction of the gradient of the deformation vector field is introduced in the original registration network to constrain the deformation vector field to be smooth. Because the generator has two outputs and There are two corresponding regularization terms for the deformation vector field. Therefore, the total smoothing loss can be expressed as:
[0032]
[0033] where is the gradient of the deformation vector field T. In practice, the difference between pixels is used to approximate the gradient of each pixel in the deformation vector field.
[0034] A computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, the processor implementing the metal artifact removal method based on the time-varying constraint adversarial generative network model as described above when executing the computer program.
[0035] A computer-readable storage medium storing a computer program for executing the metal artifact removal method based on the time-varying constraint adversarial generative network model as described above.
[0036] Advantages: Compared with the prior art, the present application constructs a GAN network metal artifact removal model (MARGANVAC) with time-varying constraints. The model performs adaptive fidelity constraints on each part of the whole image by introducing a time-varying constraint term, thereby more effectively training the generator to generate images with better detail fidelity and metal artifact removal effect. Compared with the current mainstream model, the method not only has better artifact removal effect, but also has a wider range of application scenarios. BRIEF DESCRIPTION OF DRAWINGS
[0037] Figure 1 is a network diagram of a conventional GAN model;
[0038] Figure 2 is a result display diagram of the MAR method based on the conventional GAN architecture;
[0039] Figure 3 is a GAN network diagram introducing a registration network as a sampling function;
[0040] Figure 4 is an affine transformation diagram, where (a) is the original image, (b) is the image after affine transformation of (a), and (c) is the difference image between (a) and (b);
[0041] Figure 5This is a schematic diagram of the overall architecture of the MARGANVAC model according to an embodiment of the present invention;
[0042] Figure 6 This is a schematic diagram of the generator structure;
[0043] Figure 7 A schematic diagram of the basic components of a generator;
[0044] Figure 8 This is a schematic diagram of the discriminator structure;
[0045] Figure 9 A schematic diagram illustrating the qualitative comparison of different methods on the deepplesion dataset;
[0046] Figure 10 A schematic diagram illustrating the qualitative comparison of different methods on a MicroCT synthetic artifact dataset;
[0047] Figure 11 This is a schematic diagram showing the results of a qualitative comparison of different methods on a cone-beam MicroCT dataset containing real metal artifacts. Detailed Implementation
[0048] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. After reading the present invention, any modifications of the present invention in various equivalent forms by those skilled in the art will fall within the scope defined by the appended claims.
[0049] Let the domain of a CT image affected by metal artifacts be... The domain of an image unaffected by artifacts is The goal of a metal artifact removal network is to find a mapping function f(x) a ) = x, where When there is a paired dataset At that time, a supervised network is trained using a paired dataset to learn f.
[0050] GAN networks have better fitting capabilities than common convolutional neural networks (CNNs) and can generate images with more anatomical details. Therefore, GAN models are a very good foundational network model for MAR (Magnetic Image Processing). A drawback of traditional GAN models is that the generated details may not be consistent with the ground truth. Conventional GAN networks, such as... Figure 1 As shown, it consists of three parts. The first part is the generator subnetwork G(x) a ;θ G The second part is the discriminant subnetwork D(x). a ||x;θ DThe third part is the fidelity loss based on high-level features extracted using the L1 or L2 norm or certain specific networks. Generally, designing a powerful generator G(x) involves... a ;θ G ) and the matching discriminator D(x) a ||x;θ D Removing metal artifacts is very simple and convenient because the characteristics of metal artifacts are very different from the characteristics of image content information. Figure 2 The image shows an artifact-removed image obtained using a traditional GAN-based MAR model. It can be seen that the texture features of specific metallic artifacts (such as stripes and shadows) are eliminated, and the generator produces an image that appears to be free of metallic artifacts. In other words, the powerful generator successfully removes the artifacts. x in the domain a Image converted In the domain Images. However, it is clear that the images generated by the generator... The content information is inconsistent with the ground truth image x. This preliminary experiment shows that the generative adversarial network has been trained well enough to achieve translation between images from different domains, but the fidelity loss is insufficient to guide the generator to optimize and generate images highly similar to the ground truth image. Common fidelity losses include...
[0051] L p =||C(x) a ;θ G )-x|| p (1.1)
[0052] Where p is 1 or 2, representing the L1 or L2 norm. Equation (1.1) is a pixel-level fidelity constraint, so it is a strong constraint. From many previous studies, such as the well-known ResNet model, encoding residual vectors is more efficient than encoding original vectors. This is because residual vectors contain less encoded information, thus reducing the learning burden on the network. This residual encoding can be called spatial encoding. This invention attempts to extend the idea of residual encoding to the time domain, that is, to allow the network to learn content information features step by step. More specifically, the network first learns low-frequency features of the image (i.e., relatively smooth parts in CT images), and then captures high-frequency features (i.e., edge information and other structures in the image). Therefore, the learning process is gradual, and the network can evolve gradually. To achieve this goal, this invention needs to design a time-varying loss function to replace the time-invariant loss function.
[0053] L var =F t (G(x a ;θ G), x) (1.2)
[0054] where F t is a time-varying function with respect to generating image G(x a ; θ G ) and ground truth image x. The image is composed of low-frequency components and high-frequency components, which means that the image x can be represented as a combination of low-frequency components and high-frequency components. The image low-frequency component, due to its relatively smooth characteristics, can be regarded as a piecewise constant function, so the pixel x j has a greater probability of being similar to the adjacent pixel. If the GAN model is allowed to learn to generate the low-frequency part at the beginning, the pixel fidelity constraint is relaxed by introducing the neighborhood similarity constraint as follows:
[0055]
[0056] where j is the pixel index, where φ is a sampling function for sampling pixels within the neighborhood of pixel x j , and σ is a parameter for controlling the size of the neighborhood. The physical meaning of this loss function L p is that the pixel x j is similar to its adjacent pixel. G(x a ; θ G ) learns the low-frequency component of x. If the parameter σ can be dynamically changed to become a time-varying function, a loss function L var is obtained:
[0057]
[0058] The function φ(x j , σ t ) is a sampling function of the pixels in the neighborhood of x j , and the parameter σ t gradually converges to zero as the iteration proceeds, that is, the loss function eventually degenerates to:
[0059]
[0060] is a common fidelity loss, which means that the constraint is being strengthened in order to prompt the GAN to learn the high-frequency component. Next, a function φ(x j , σ t ) is instantiated to meet the above conditions, that is, the function has a sampling function and the parameter σ decreases over time; in theory, there should be many options to construct such a function, and similar function extensions are within the scope of the invention. The deformable registration network designed in the invention is a relatively good solution, which has the advantage of being able to adaptively adjust the network parameters without human intervention, so that the generator can produce more realistic artifact-free pictures. Figure 3The diagram shows a GAN network incorporating a registration network. First, for each input image x... a Perform a random affine transformation, and then perform another random affine transformation on the reference image x. This will give you the transformed image. and x T .when After passing through the generator, we obtain an image with artifacts removed. because and x T Images that are not pixel-level corresponding are used to compare their differences. and x T The input registration network partially corrected the distortion caused by the random affine transformation. For example... Figure 4 As shown, the magnitude of the random affine transformations adopted is not large and is within a controllable range, so the registration network can correct these random deviations.
[0061] At the start of training, the registration network performed poorly, therefore its output... Compared with the affine transformed ground truth image x T There will be a certain positional deviation between them, and these deviations are variable; that is, each pixel in each input image of each epoch has an independent and different deviation. This characteristic can help the network simulate the function φ(x). j , σ t The random sampling process involves the registration network. Furthermore, as training progresses, the registration performance of the network improves, i.e., the parameter σ... t It will decrease. If the entire model converges well, σ t It will drop to zero.
[0062] The overall architecture of the MARGANVAC model proposed in this invention is as follows: Figure 5 As shown, MARGANVAC consists of three modules: a generator G, a discriminator D, and a registration network R. In addition to the novel training mechanism with temporally variable constraints mentioned above, some effective techniques are introduced to make the generator more efficient to train. First, for the input image x... a The metal trajectory portion is estimated in the projected image in the sinusoidal domain using linear interpolation, and the LI-corrected reconstructed image x is obtained using FBP or FDK methods. [LI]a The image x corrected by the LI method [LI]a and images affected by artifacts x a The channels are concatenated along the channel dimension and used as input to the generator G. Secondly, to further reduce the learning burden on the generator, residual learning is introduced into the generator G. Therefore, the image after removing metal artifacts is represented as:
[0063]
[0064] where [m, n] denotes the concatenation operation of images m and n. To better distinguish the metal artifact and the content information, self-reconstruction is introduced as a constraint to regularize the generator. In the self-reconstruction process, the artifact-free image y is concatenated into [y, y] and then input into the generator, thus the self-reconstructed image is denoted as:
[0065]
[0066] To make the registration network R work properly, a random affine transformation is applied to each input artifact image x a and ground-truth image x. Suppose the corresponding affine transformed images are A1(x a ) and A2(x), respectively. In this way, the image without metal artifact is denoted as:
[0067]
[0068] Correspondingly, in the self-reconstruction, the output becomes:
[0069]
[0070] Then take the concatenation of A1(x ) and A2(x) as the input of the registration network R, and the output of the registration network R is a deformation vector field T x ,
[0071]
[0072] After obtaining the deformation vector field T x , the resampled image is obtained by applying T x to the input image
[0073]
[0074] In the same way as generating T x , the deformation vector field T y is obtained,
[0075]
[0076] and in the case of self-reconstruction, the corresponding resampled image
[0077]
[0078] To make the generator able to well separate the content information and the metal artifact, a specially designed generator is introduced, such as Figure 6The generator contains an encoder and a decoder. The encoder is used to map the image samples from the image domain to the latent space, in which the features of the content information and the metal artifacts are separated. The decoder reconstructs the separated content information into the artifact-free image. Between the encoder and the decoder, a deep subnetwork consisting of 21 Inception-ResNet modules is introduced to improve the separation ability of the artifact features and the content information in the latent space. The Inception structure proposed in GoogleNet makes full use of different sizes of convolution kernels, so the Inception structure has better performance in feature extraction. The Inception-ResNet module is formed by introducing the residual network structure into the Inception structure, so it is easy to build a deep network consisting of multiple Inception-ResNet modules to strengthen the separation ability of the content information and the artifact features without worrying about whether the network will converge. The structure of the Inception-ResNet module used is as shown in Figure 7 (c), in which the dimension of the input feature map is reduced by adding a 1x1 convolution operator, thereby reducing the amount of calculation. In the encoder, the core module is the down-sampling module as shown in Figure 7 (b), which consists of a convolution layer, an instance normalization (Instance Normalization) layer and a ReLU activation function. The reason for choosing instance normalization instead of batch normalization (Batch Normalization) is that instance normalization has been proven to have better performance for small-batch image generation tasks. In the decoder, the core module is the up-sampling module as shown in Figure 7 (e), which is similar to the down-sampling module, the only modification is that the transposed convolution replaces the convolution operation. The discriminator D is built based on the PatchGAN discriminator. The structure of the discriminator is as shown in Figure 8 The whole structure of the registration network R is based on the U-Net structure, but it is deeper than the traditional U-Net structure. In this very deep U-Net, all levels of features can be obtained to effectively promote the registration process.
[0079] Loss function design
[0080] The learning process of the model is also a process of encouraging the generator to reduce the metal artifacts while maintaining the content information. In each adversarial learning iteration, the generator G outputs the image with reduced metal artifacts; when the performance of G and D is improved, the output image looks more like an artifact-free image. The role of the introduced registration network is to help the generator G to gradually learn the content features, so the network weights θ R and θ G need to be updated simultaneously in the training process. Figure 5As can be seen, MARGANVAC contains four forms of loss, namely artifact correction loss self-reconstruction loss adversarial loss and a smoothing loss for constraining the deformation vector field The total loss can be represented as a weighted sum of these losses as follows:
[0081]
[0082] where λ are hyperparameters that balance the importance of each loss during the training process.
[0083] correction loss Since the whole model is based on supervised learning, the most direct and effective constraint is the correction loss, which minimizes the difference between the image without metal artifacts and the ground truth image. However, after random affine transformation of the input image, there is no longer a strict pixel-level corresponding ground truth image. To solve this problem, the registration network R and the generator G are trained simultaneously, and the correction loss is designed as:
[0084]
[0085] where is the resampled image obtained by equation (1.11), and A2(x) is the ground truth image after affine transformation as shown in equation (1.10). In many image-to-image conversion tasks, the L1 norm loss function has been proven to be more effective in restoring image details, so L1 norm is used here to encourage the generator to maintain more details while reducing metal artifacts.
[0086] self-reconstruction loss The training goal is to make the generator learn to remove metal artifacts. At the same time, more constraints need to be introduced to encourage the generator to keep more existing content information unchanged. The core goal is to let the network learn to identify metal artifacts and content information, i.e. reduce metal artifacts in the presence of metal and artifacts, and preserve all image content information in the absence of metal artifacts. Therefore, the self-reconstruction loss is introduced as follows:
[0087]
[0088] where, is the resampled image obtained by equation (1.13), and A4(y) is the affine transformation of the artifact-free image y.
[0089] adversarial loss The adversarial learning can encourage the generator G to generate more realistic artifact-free images. To achieve this goal, the generator should have the ability to distinguish artifact information and content information in the latent space. To enhance this ability, two adversarial learning strategies are introduced. One strategy is to improve the ability of the latent space to recognize metal artifacts, and the other way is to improve the ability of the latent space to preserve content information. The input data of the two strategies are the images affected by metal artifacts and the images without metal artifacts, respectively. The two strategies are executed simultaneously in adversarial learning, and their respective losses can be written as:
[0090]
[0091]
[0092] So the total adversarial loss is:
[0093]
[0094] Smooth loss When minimizing the combined loss of the correction loss and the self-reconstruction loss, the deformation vector field generated may become non-smooth, which is physically unrealistic. To solve this problem, a diffusion regularization operator in the gradient direction of the deformation vector field is introduced in the original registration network to constrain the deformation vector field to be smooth. Because the generator has two outputs and There are two corresponding deformation vector field regularization terms. Therefore, the total smooth loss can be expressed as:
[0095]
[0096] where is the gradient of the deformation vector field T. In practice, the difference between pixels is used to approximate the gradient of each pixel in the deformation vector field.
[0097] Obviously, those skilled in the art should understand that the steps of the metal artifact removal method of the time-varying constraint-based adversarial generation network model of the embodiments of the present application described above can be realized by general computing devices, which can be concentrated on a single computing device or distributed on a network composed of multiple computing devices, and optionally, they can be realized by program codes executable by computing devices, so that they can be stored in storage devices and executed by computing devices, and in some cases, the steps shown or described can be executed in different order, or they can be made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module. Thus, the embodiments of the present application are not limited to any specific combination of hardware and software.
[0098] Preparation before experiment
[0099] (1) Data set construction
[0100] Data set 1: deeplesion simulated metal artifact data set
[0101] From the deeplesion data set, 4000 non-artifact CT images were selected, and 90 metal masks were selected from the 100 binary metal masks of different shapes and sizes provided by CNNMAR to synthesize metal artifacts. A total of 360000 paired CT images were synthesized as training data sets. In addition, another 200 non-artifact CT images and the remaining 10 binary metal masks were selected from the deeplesion data set to generate a test data set containing 2000 paired CT images. Fan-beam projection was used, and the reconstruction method was FBP, and the image size was 256x256.
[0102] Data set 2: real artifact data set
[0103] The real artifact data set is a Micro CT data set, which is a collection of real CT projection images of bone samples with implanted metal and without implanted metal. A synthetic method based on real metal artifacts was used to synthesize paired training data sets. After obtaining the synthetic projection data of the inserted metal track, the corresponding reconstructed image was obtained by the FDK method. The training set contains 3868 images, and the test set contains 537 images. The image size is 364x364.
[0104] (2) Implementation and training details
[0105] The network is based on the Pytorch deep learning framework and runs on a computer equipped with an Nvidia 2080Ti GPU. In the training process, the Adam optimizer (optimizer parameters (β1, β2) = (0.5, 0.999)) is used to optimize the loss function. For the deeplesion data set, the batch size is set to 2, the learning rate is set to 0.0001, and the number of training epochs is 5. For the cone-beam Micro CT data set, the batch size is set to 1, the learning rate is set to 0.0001, and a total of 70 epochs are trained. The weights of the loss function are λ smooth = 10, λ Adv = 1, λ Corr = 20, λ Rec = 20.
[0106] (3) Evaluation index
[0107] On the synthetic artifacted paired dataset, the performance of all MAR methods, including the proposed method and other classical MAR methods, is quantitatively evaluated by structural similarity (SSIM), peak signal-to-noise ratio (PSNR) and standard deviation. Specifically, the higher the SSIM and PSNR values, the better the performance of metal artifact removal and content information preservation. The introduction of standard deviation is to evaluate different methods from the dimension of result stability, because in addition to SSIM and PSNR, the stability of the results of different models when encountering different data also needs to be concerned. On the real cone-beam Micro CT dataset with real metal artifacts, due to the lack of paired artifact-free CT images, only qualitative evaluation based on image visual evaluation is carried out.
[0108] MARGANVAC model effectiveness verification
[0109] In order to prove the excellent performance of the proposed method in metal artifact removal, several representative MAR methods are experimented, including the traditional LI algorithm, FSMAR algorithm, supervised CNNMAR, cGANMAR, U-Net, dual-domain method InDuDoNet, unsupervised ADN algorithm, CycleGAN and attention-guided β-CycleGAN algorithm. All methods are carried out on the public dataset deeplesion dataset and private Micro CT dataset. In addition to the proposed model, all methods are based on the publicly available code or the model proposed in the published paper.
[0110] Because the projection and reconstruction of CT images are involved in the training process of the dual-domain network, the FDK of cone-beam CT requires a large amount of computing resources, therefore in the training of the previous dual-domain network, fan-beam CT is used, and Micro CT data is cone-beam CT, therefore only the dual-domain network experiment is carried out on the deeplesion dataset, and no dual-domain experiment is carried out on the Micro CT dataset.
[0111] Evaluation on deeplesion dataset
[0112] Table 1.1 Comparison of PSNR (dB) and SSIM of different methods on deeplesion dataset
[0113]
[0114] In the experiments of this section, various MAR methods are qualitatively and quantitatively analyzed on the paired deeplesion dataset. The results of quantitative analysis are obtained based on 2000 test set images. As shown in Table 1.1, it can be intuitively seen that the deep learning methods are generally superior to the traditional methods, and the supervised methods are generally superior to the unsupervised methods. Among all the reproduced supervised methods, the performance of the dual-domain model InDuDoNet is the best. The PSNR and SSIM scores of the model proposed in the present application are only slightly inferior to InDuDoNet, while InDuDoNet is only suitable for fan-beam CT, but the model of the present application can not only be suitable for fan-beam, but also be suitable for cone-beam CT.
[0115] The images affected by artifacts, and their corresponding artifact-free images, and the images after removing artifacts by various MAR methods are shown in Figure 9 Figure 9 Examples of chest slices are given for demonstration. In the image affected by metal artifacts in Figure 9 , almost no tissue / organ structure in the vicinity of the metal implant can be seen, and the strip-shaped artifact runs through the entire image, and the image content information in the vicinity of the metal implant is severely damaged by the metal artifact. From the images after removing artifacts by various MAR methods, it can be seen that most of the strip-shaped artifacts far from the metal are removed, and the performance of the deep learning method is superior to that of the traditional LI method and the FSMAR method. After processing by the traditional methods LI and FSMAR, the tissue structure is still blurred, and there are secondary artifacts. The performance of the supervised methods CNNMAR, cGANMAR, InDuDoNet and U-Net is superior to that of the traditional methods, but they cannot well restore the missing content information. In contrast, the unsupervised methods tend to automatically generate rich details in the image, but many of these details are inconsistent with the actual missing structure. Therefore, it can be seen that some structures similar to tissues, organs or lesions actually do not exist. β-CycleGAN and ADN have improved in this regard, but cannot be completely avoided, because the unsupervised method lacks fidelity constraints. By analyzing the image after removing artifacts by the MARGANVAC method of the present application, it can be clearly seen that several blood vessel lumens are restored, and the vertebral body boundary is complete. These clear details cannot be seen in other images. These results show that the method of the present application has better performance in reducing metal artifacts and preserving content information. Although the InDuDoNet method is slightly superior to the method of the present application in evaluation indicators ssim and psnr, it can be seen from Figure 9 that the texture of the image restored by the method of the present application is closer to the real image than that of the InDuDoNet method.
[0116] The standard deviation (Std) of ssim and psnr of different MAR methods can be used to compare the stability of different methods. From Table 1.1, the Std of PSNR index of the present method ranks third, and the Std of SSIM index ranks first. Therefore, compared with other MAR methods, the present method can obtain more stable results while ensuring good image quality.
[0117] The proposed MAR methods were evaluated on cone-beam Micro CT artifact migration dataset
[0118] To verify the performance of the proposed MAR methods in real applications, a dataset acquired from cone-beam Micro CT was prepared. The paired data was generated by migrating the metal artifact-affected metal trajectories to the projections without metal implants, and the results of quantitative analysis of all MAR methods are shown in Table 1.2. It can be seen that the quality of the original images contaminated by metal artifacts is much worse than that of the simulation experiment in terms of SSIM and PSNR index. It is indicated that introducing real metal artifacts to synthesize artifact-affected images can cause more severe image degradation. From the quantitative analysis results of MAR methods in Table 1.2, it can be seen that the performance of all MAR methods has decreased, and the present method is superior to other MAR methods. The reduced value can be due to the lower quality of the original image affected by the artifact than that of the original image in the simulation experiment. The results of qualitative comparison of different MAR methods are shown in Figure 10 The cross-sectional images of the bone affected by artifacts are shown in Figure 10In the middle, the metal implant is inserted into the bone, resulting in severe image degradation and loss of content information. Stronger shadow artifacts can be seen than in the simulation experiment, and part of the bone tissue detail is completely lost. Part of the trabecular bone disappears, and the complete cortical bone is cut into several parts. From the results of the CycleGAN and ADN de-artifacting methods, it can be seen that the gap between the segmented parts of the cortical bone is not well filled, and the missing trabecular bone is not well recovered. The β-CycleGAN performs slightly better than the CycleGAN and ADN in both aspects. Compared with the unsupervised method, the supervised method, especially the CNNMAR, has greater improvement in both aspects, while the U-Net method tends to make the picture smooth and fuzzy. By comparing the images obtained by the method proposed in the application and the CNNMAR, it can be seen that the cortical bone boundary recovered by the method of the application is clearer and closer to the ground truth image. The soft tissue area close to the metal implant is more easily contaminated by metal artifacts because the CT value is much smaller than that of bone tissue. However, in clinical applications, correct soft tissue imaging is crucial for diagnosis. Therefore, soft tissue recovery performance should be a key indicator for reducing metal artifacts. By comparing the soft tissue, it can be found that the method of the application has better soft tissue recovery performance while removing metal artifacts. In summary, all the above deep learning-based MAR methods can remove most of the metal artifacts, but there are large differences in content information preservation, that is, the supervised method is better than the unsupervised method, and the method proposed in the application performs best.
[0119] Table 1.2 Comparison of PSNR (dB) and SSIM of different methods on the Micro CT synthetic artifact dataset
[0120]
[0121] Performance evaluation on cone beam Micro CT real artifact images
[0122] Further experiments on the Micro CT dataset affected by real metal artifacts were conducted on the method proposed in the application to test its performance in actual applications. Since there is no ground truth picture, only a qualitative evaluation method can be used to evaluate its performance. All MAR models are first trained on the Micro CT paired training dataset generated using the artifact migration method, and then tested on the real artifact images of the cone beam Micro CT. Figure 11 Real artifact-affected images and corresponding de-artifact images obtained by different MAR methods are shown. From the real artifact image, it can be seen that the artifact is visually different from the ground truth image, and the soft tissue near the metal implant is also affected. The image obtained by the CycleGAN method is still affected by the artifact, and the soft tissue near the metal implant is also affected. The image obtained by the ADN method is better than the CycleGAN method, but the artifact is still visible, and the soft tissue near the metal implant is also affected. The image obtained by the β-CycleGAN method is better than the CycleGAN and ADN methods, but the artifact is still visible, and the soft tissue near the metal implant is also affected. The image obtained by the U-Net method is better than the CycleGAN, ADN and β-CycleGAN methods, but the artifact is still visible, and the soft tissue near the metal implant is also affected. The image obtained by the CNNMAR method is better than the CycleGAN, ADN, β-CycleGAN and U-Net methods, and the artifact is not visible, but the soft tissue near the metal implant is still affected. The image obtained by the method proposed in the application is better than the CycleGAN, ADN, β-CycleGAN, U-Net and CNNMAR methods, and the artifact is not visible, and the soft tissue near the metal implant is also not affected. Figure 10The artifacts in the synthetic images are very similar. By comparing the local zoom-in images of the results from different MAR methods, it can be seen that the supervised method has better artifact removal performance than the unsupervised method. It can be shown that the training of the supervised method is successful. In other words, with the help of the paired dataset synthesized by the artifact migration method, the supervised method is successfully extended to practical applications. In all the comparison methods, it is obvious that CNNMAR and the method of the present application perform better in removing metal artifacts and restoring image content information. Further comparison of the details in the local zoom-in images shows that the trabecular bone anatomy obtained by the method of the present application is clearer, the soft tissue area around the metal has almost no residual artifacts, and the soft tissue boundary is smoother. This result is consistent with the above simulation experiment.
Claims
1. A method for metal artifact reduction based on a time-varying constraint adversarial generative network model, characterized in that, A GAN network with time-varying constraint, referred to as MARGANVAC, is constructed to remove metal artifacts from CT images. The GAN network with time-varying constraint comprises three modules, namely, a generator G, a discriminator D, and a registration network R. The CT image with metal artifacts is input into the trained GAN network with time-varying constraint to generate a metal artifact removed image.
2. The metal artifact reduction method of the time-varying constraint-based generative adversarial network model according to claim 1, characterized in that, The training process of the GAN network with time-varying constraint is as follows: First, each input image x a with metal artifacts is subjected to a random affine transformation, and then a reference image x a of the image x a which is not affected by artifacts is subjected to a random affine transformation, respectively obtaining the transformed image x and the reference image x T ; when the image x passes through the generator G, the image x without artifacts is obtained and x T are input into the registration network R; the loss function of the registration network is where the first part of the loss function is the generator sub-network G(x a ; θ G ), θ G are parameters of the generator subnetwork, j is a pixel index, where φ is a sampling function that samples pixels within a neighborhood of pixel x j , σ T is a parameter that controls the size of the neighborhood, and σ t gradually converges to zero as iterations proceed during training. 3.The metal artifact reduction method of the time-varying constraint-based generative adversarial network model according to claim 1, wherein, The generator G is introduced with residual learning, and self-reconstruction is introduced as a constraint to regularize the generator. The generator comprises an encoder and a decoder.
4. The metal artifact reduction method of the time-varying constraint-based generative adversarial network model according to claim 1, characterized in that, The training process of the GAN network with time domain variable constraint de-metal artifact model introduces self-reconstruction as a constraint to regularize the generator, in the self-reconstruction process, the artifact-free image y is spliced into [y, y], and then input into the generator; before splicing, the artifact-free image y is randomly subjected to affine transformation to obtain a result A3(y), and splicing obtains [A3(y), A3(y)], accordingly, in the self-reconstruction, the generator output becomes: G([A3(y), A3(y)]; theta G ); the reference image corresponding to the artifact-free image y which is not affected by the artifact is also y, and theta G is the parameter of the generator subnetwork.
5. The metal artifact reduction method of the time-varying constraint-based generative adversarial network model according to claim 2, characterized in that, For image x a , the estimated value of the metal track part in the projection image of the sinogram domain is obtained by using the linear interpolation method, and the LI correction reconstruction image x [LI]a is obtained by the FBP or FDK method; the image x [LI]a corrected by the LI method and the image x a affected by the artifact are respectively subjected to random affine transformation to obtain affine transformation results A1(x [LI]a ) and A1(x a ), and A1(x [LI]a ) and A1(x a ) are connected in the channel dimension as the input of the generator G.
6. The metal artifact reduction method of the time-varying constraint-based generative adversarial network model according to claim 5, characterized in that, Residual learning is introduced in the generator G, the image after removing metal artifacts is denoted as: [x a ,x [LI]a ] represents the concatenation operation of x a and x [LI]a ; in the self-reconstruction, the generator output becomes A3(y) is obtained by randomly affine transforming the artifact-free image y; then the concatenation of and A2(x) is taken as the input of the registration network R, A2(x) represents the random affine transformation of the reference image x, the output of the registration network R is a deformation vector field T x , θ R is the registration network parameter, θ G is the parameter of the generator sub-network, after obtaining the deformation vector field T x , the resampled image is obtained by applying T x to the input image According to the same steps as obtaining T x , the deformation vector field T y , A4(y) represents the random affine transformation of the reference image y; and in the case of self-reconstruction, the corresponding resampled image The concatenation of the input image and A2(x) is taken as the input of the discriminator D, and in the case of self-reconstruction, the concatenation of the generator output and A4(y) is taken as the input of the discriminator D.
7. The metal artifact reduction method of claim 6, wherein, The GAN network with time domain variable constraint de-metal artifact model contains four forms of loss, which are artifact correction loss self-reconstruction loss adversarial loss and smoothing loss for constraining deformation vector field The total loss is represented as the weighted sum of these losses Each λ represents the weight coefficient of the corresponding loss; correction loss The registration network R and the generator G are trained simultaneously, and the correction loss is designed as: The LI norm is used to encourage the generator to preserve more details while reducing metal artifacts, represents a desire, represents that x comes from the dataset The domain of the CT images affected by metal artifacts is Self-reconstruction loss The training objective is to make the generator learn to remove the metal artifacts, while more constraints need to be introduced to encourage the generator to keep more of the existing content information unchanged; The encoder is used to map image samples from the image domain to the latent space, in which the content information and the features of metal artifacts are separated. wherein, represents that, represents that y comes from the dataset The domain of the image that is not affected by the artifact is adversarial loss : Adversarial learning encourages the generator G to generate more realistic artifact-free images; to achieve this goal, the generator has the ability to distinguish artifact information and content information in the latent space; to enhance this ability, two adversarial learning strategies are introduced; The decoder reconstructs the separated content information into an artifact-free image as a reference image. A deep subnetwork composed of 21 Inception-ResNet modules is introduced between the encoder and the decoder. Smooth loss is: wherein is the gradient of the deformation vector field T, is the gradient of the deformation vector field T x is the gradient of the deformation vector field T, is the gradient of the deformation vector field T y is the gradient of the deformation vector field T.
8. A computer device, comprising: The metal artifacts are reduced in the presence of metal and artifacts, and all image content information is preserved in the absence of metal artifacts.
9. A computer-readable storage medium, characterized in that: One strategy is to improve the ability of the latent space to identify metal artifacts, and the other is to improve the ability of the latent space to preserve content information. The input data of the two strategies are images affected by metal artifacts and images without metal artifacts, respectively. The two strategies are executed simultaneously in adversarial learning, and their respective losses are as follows: The base of log is 2, and D(.) represents the output of the discriminator D. Therefore, the total adversarial loss is as follows: The computer device comprises a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the metal artifact removal method based on the time-varying constraint adversarial generation network model is realized. The computer readable storage medium stores a computer program for executing the metal artifact removal method based on the time-varying constraint adversarial generation network model.
Citation Information
Patent Citations
CT image recovery method based on unsupervised learning
CN110675461A
CT double-domain joint metal artifact correction method based on generative adversarial network
CN112508808A