Calcification artifact removal method and system for CT image, terminal and medium
Through an artifact removal model based on adversarial networks, combined with preprocessing, semantic feature extraction and feature fusion technology, the problem of difficult removal of calcified artifacts in CT images is solved, and more accurate discrimination of stenosis and the reduction of misdiagnosis rate is achieved.
Patent Information
- Application Number
- CN202311581005.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-24
- Publication Date
- 2025-05-27
AI Technical Summary
The presence of calcified artifacts in CT images leads to an increase in the diagnostic false positive rate, and the prior art is difficult to effectively remove artifacts, especially when the ratio of the calcified region to the entire section is dysregulated.
Using an artifact removal model based on adversarial network, the calcified artifacts in CT images are removed through preprocessing, semantic feature information extraction, feature fusion and perceived loss function combination, the generator and discriminator cooperate with the perceived extraction module and the semantic judgment module.
It significantly improves the artifact removal effect of CT images, which can effectively assist doctors in improving the accuracy of discrimination of stenosis and reduce the rate of misdiagnosis.
Smart Images

Figure CN120047550A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of CT image processing, and in particular, to a method, system, terminal and medium for removing calcification artifacts from CT images. Background Art
[0002] CT (Computed Tomography), namely computed tomography, has been widely used in the clinical diagnosis of various diseases due to its characteristics of fast scanning time, high clarity, non-invasiveness, etc. However, due to reasons such as different CT scanner models, different parameter settings, and differences in patients themselves, CT images are easily affected by noise and artifacts. For patients with calcified plaques, there will be calcification artifacts in CT images, also known as calcium blooming or blooming artifacts. Specifically, when an image voxel contains a high-density calcified plaque and a much lower-density blood vessel, the imaging result is that the attenuation average value of the two voxels is close to the attenuation average value of the high-density material, making the calcified plaque look larger than its actual size. The blooming artifact will lead to an overestimation of the severity of vascular stenosis, further resulting in an increased problem of false positive diagnosis rate. A study including 291 patient samples showed that when ≥50% arterial stenosis was used as the critical value, calcification would reduce the accuracy of adjacent lumen evaluation, resulting in increased inaccuracy of CT images.
[0003] To solve this thorny problem, researchers have developed many methods. These methods can be divided into three categories: 1) high-resolution CT hardware; 2) high-quality iterative reconstruction methods; 3) image post-processing methods. However, the first two methods are either limited by hardware requirements or require a large number of prior assumptions and high computational costs, which leads to bottlenecks in their practical applications. However, the post-processing method can be directly processed in the image domain, and its computational efficiency is higher than that of the high-quality iterative reconstruction method. Some traditional methods include non-local mean method, improved K-SVD method, three-dimensional block matching algorithm (BM3D), etc. Using these methods can effectively improve the image quality, but also make the image overly smoothed. At the same time, since CT noise is often non-uniformly distributed, these methods are difficult to work effectively.
[0004] In recent years, the development of deep learning has provided new ideas for this problem. Compared with traditional algorithms, deep learning has shown great potential in the field of medical image processing with its excellent feature extraction ability and powerful nonlinear fitting ability. Convolutional neural networks (CNNs) have been proven to perform well in tasks such as image detection, classification, and segmentation of medical images. Among them, the U-type network is widely used in the field of medical images because of its special network structure that combines low-level detail information and high-level semantic information, which is very consistent with the characteristics of medical images. In addition, unlike general CNNs, generative adversarial networks (GANs) have also attracted widespread attention from researchers in solving clinical problems because of their compatibility with unpaired data. GAN consists of a generator G and a discriminator D. The former is used to generate simulated real images, and the latter is used to distinguish the authenticity of images.
[0005] In the past, many deep learning networks based on CNN and GAN have been shown to greatly reduce noise and improve image clarity. Among them, RED-CNN is one of the earliest denoising networks used for low-dose CT (LDCT). In addition, models such as W-GAN, Cyc-GAN, and SGD-NET have improved the denoising ability of the network from different angles. Although the above methods have made significant progress in CT denoising, the goal of these methods is to remove general noise, so the effect on removing artifacts is very limited. Due to the particularity of calcification artifacts, some problems remain to be solved. For example, the generation mechanism of calcification artifacts is partial volume effect, while the generation of general noise is caused by beam hardening and hardware noise, which leads to different emphasis on image processing. In addition, due to the imbalance between the calcification area and the entire slice, it is difficult for general denoising models to focus on high semantic parts. Summary of the invention
[0006] In view of the above-mentioned deficiencies in the prior art, the present invention provides a method, system, terminal and medium for removing calcification artifacts from CT images.
[0007] According to one aspect of the present invention, a method for removing calcification artifacts from a CT image is provided, comprising:
[0008] Preprocessing the CT image to obtain a preprocessed CT image;
[0009] Extracting semantic feature information from the preprocessed CT image;
[0010] Provide an artifact removal model based on a generative adversarial network. The artifact removal model includes: a perception extraction module, a generator, and a discriminator. Among them, after fusing the features of the preprocessed original CT image and the semantic feature information, it is used as the input of the artifact removal model. The generator is used to generate a reconstructed CT image after artifact removal, and a perception extraction module is introduced to extract the feature maps of each layer of the neural network and calculate the perception loss. The discriminator is used to determine the probability that the generated image is real, and a semantic determination module is added to the discriminator to determine the authenticity of the generated CT image after artifact removal through semantic information.
[0011] Preferably, the preprocessing of the CT image includes:
[0012] Convert the format of the CT image;
[0013] Perform image pixel value normalization and image center region cropping on the CT image after format conversion.
[0014] Preferably, the extraction of semantic feature information from the preprocessed CT image includes:
[0015] Provide a feature extraction network model based on a U-shaped network (U-shaped network A), train the feature extraction network model, and use the preprocessed CT image as the input of the feature extraction network model to output the calcified plaque and the detection branch region, and use the calcified plaque and the detection branch region as semantic feature information.
[0016] Preferably, the feature extraction network model based on the U-shaped network includes: four encoding modules and corresponding decoding modules. Among them, the encoding module performs convolution and downsampling operations to gradually reduce the resolution of the input image and extract image features. The decoding module performs deconvolution operations to gradually restore the image resolution and perform feature fusion through skip connections.
[0017] Preferably, the feature fusion of the preprocessed CT image and the semantic feature information includes:
[0018] Provide a global-local feature fusion network model;
[0019] Input the preprocessed CT image into the global branch of the global-local feature fusion network model to output a global feature map;
[0020] Input the semantic feature information into the local semantic branch of the global-local feature fusion network model to output a local feature map;
[0021] Merge the global feature map and the local feature map to obtain three-dimensional feature information after feature fusion.
[0022] Preferably, the generator is constructed based on an improved U-shaped network (U-shaped network B) and a perceptual extraction module is introduced; where:
[0023] The improved U-shaped network includes: five encoding sub-modules and five decoding sub-modules; among them, the first encoding module is composed of a global-local feature fusion network model for feature fusion, and a Squeeze-and-Excitation network is inserted into each skip layer to realize adaptive allocation of feature weights of different channels;
[0024] The perceptual extraction module is used to extract the feature map of each layer of the improved U-shaped network and calculate the perceptual loss; among them, the perceptual extraction module is migrated from the encoding module part of the feature extraction network model for extracting semantic feature information for self-supervised learning, and the network structure and parameter settings before and after migration remain consistent;
[0025] The generator adopts a semantic similarity loss L SSL to measure the semantic distance between the generated image and the gold standard, and assigns higher weights to the calcified plaque and blood vessel branch regions in the image so that the network pays more attention to key information:
[0026]
[0027] In the formula, is the expectation function, x and y are images with and without calcification artifacts respectively, φ(·) is the network migrated from the encoding part of the feature extraction network model, G(·) is the generator, x s , y s are semantic rich regions, γ 1 (γ 1 > 1) is the weight parameter, and W, H, and C are the width, height, and number of channels of the feature map respectively.
[0028] Preferably, the discriminator is used to judge whether the input image is a real sample or a fake sample generated by the generator, and includes: a local discriminator and a global discriminator; among them, the local discriminator is used to map the input into a matrix of size N×N and discriminate the probability of each matrix element being a real sample. A semantic determination module is added to the local discriminator. The added semantic determination module assigns corresponding weights to different regions of the image according to the importance of semantic information, so that the calcified plaque and the detected branch region in the image obtain higher weights than the background region, so that the overall discrimination process further integrates semantic information; the global discriminator is used to map the input into the probability that the entire generated image is a real sample; the loss function L of the discriminator D is:
[0029] LD = γ 2 L g +(L ps + γ 3 L pl )
[0030] wherein, L g is the global loss, and L ps , L pl respectively represent the local losses of high semantic information and low semantic information, and γ 2 , γ 3 ∈ (0, 1) are weight parameters.
[0031] According to another aspect of the present invention, there is provided a system for removing calcification artifacts from CT images, comprising:
[0032] A data processing module, which is used to preprocess the CT image to obtain a preprocessed CT image;
[0033] A semantic extraction module, which is used to extract semantic feature information from the preprocessed CT image;
[0034] An artifact removal model module, which is used to provide an artifact removal model based on an adversarial network. The artifact removal model includes: a perception extraction module, a generator, and a discriminator; wherein, the preprocessed CT image and the semantic feature information are fused as the input of the artifact removal model. The generator is used to generate an artifact-removed CT image, and a perception extraction module is introduced to extract the feature map of each layer of the neural network and calculate the perception loss; the discriminator is used to determine the probability that the generated image is real, and a semantic determination module is added to the discriminator to determine the authenticity of the generated artifact-removed CT image through semantic information.
[0035] According to a third aspect of the present invention, there is provided a computer terminal, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it can be used to execute the method described in any one of the above of the present invention, or, run the system described in any one of the above of the present invention.
[0036] According to a fourth aspect of the present invention, there is provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it can be used to execute the method described in any one of the above of the present invention, or, run the system described in any one of the above of the present invention.
[0037] Due to the adoption of the above technical solutions, compared with the prior art, the present invention has at least one of the following beneficial effects:
[0038] The method, system, terminal and medium for removing calcification artifacts from CT images provided by the present invention propose a brand-new artifact removal network model based on the adversarial network. By feature fusion, perceptual loss function and weighted matrix discrimination results, the semantic information obtained is fully utilized to guide the network to focus on the semantically concentrated regions.
[0039] The method, system, terminal and medium for removing calcification artifacts from CT images provided by the present invention are verified by real clinical data. The results show that it has a better artifact removal effect, can effectively assist doctors in improving the discrimination accuracy of the stenosis degree, and thus reduce the misdiagnosis rate. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] By reading the following detailed description of the non-restrictive embodiments with reference to the accompanying drawings, other features, objects and advantages of the present invention will become more apparent:
[0041] Figure 1 It is a flowchart of the method for removing calcification artifacts from CT images in an embodiment of the present invention.
[0042] Figure 2 It is a working schematic diagram of the method for removing calcification artifacts from coronary CT images in a preferred embodiment of the present invention.
[0043] Figure 3 It is a working schematic diagram of the global-local feature fusion network model in a preferred embodiment of the present invention;
[0044] Figure 4 It is a working schematic diagram of the perceptual extraction module in a preferred embodiment of the present invention;
[0045] Figure 5 It is a working schematic diagram of the global-local semantic determination module in a preferred embodiment of the present invention.
[0046] Figure 6 It is a schematic diagram of the artifact removal performance of the method for removing calcification artifacts from CT images in a preferred embodiment of the present invention.
[0047] Figure 7 It is an artifact removal effect diagram of the method for removing calcification artifacts from CT images in a preferred embodiment of the present invention on the test set.
[0048] Figure 8 It is a schematic diagram of the composition modules of the system for removing calcification artifacts from CT images in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0049] The embodiments of the present invention will be described in detail below: These embodiments are implemented on the premise of the technical solution of the present invention, and detailed implementation manners and specific operation processes are given. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several deformations and improvements can be made, and these all belong to the protection scope of the present invention.
[0050] The generation mechanism of calcification artifacts is partial volume effect, while the generation of general noise is caused by beam hardening and hardware noise, which leads to different emphases on image processing; in addition, due to the imbalance between the calcification area and the proportion of the entire slice, it is difficult for general denoising models to focus on the high-semantic part. In order to remove the artifacts caused by the above two special differences, an embodiment of the present invention provides a method for removing calcification artifacts in CT images. Through a neural network model, this method can effectively remove calcification artifacts and reduce general noise in CT images, so as to improve the quality of CT images, and further improve the doctor's discrimination accuracy of the stenosis degree of the detection area in CT images, thereby reducing the misdiagnosis rate of related diseases.
[0051] Specifically, as Figure 1 shown, the method for removing calcification artifacts in CT images provided by this embodiment may include the following operations:
[0052] S1, preprocess the CT image to obtain the preprocessed CT image;
[0053] S2, extract semantic feature information from the preprocessed CT image;
[0054] S3, provide an artifact removal model based on a confrontation network. The artifact removal model includes: a perceptual extraction module, a generator, and a discriminator; among them, the preprocessed CT image and the semantic feature information are fused as the input of the artifact removal model. The generator is used to generate the CT image after artifact removal, and a perceptual extraction module is introduced to extract the feature map of each layer of the neural network and calculate the perceptual loss, which measures the distance between the generated image and the gold standard in the feature space; the discriminator is used to determine the probability that the generated image is real, and a semantic determination module is added to the discriminator to determine the authenticity of the generated CT image after artifact removal through semantic information.
[0055] In some preferred implementation manners of S1, preprocessing the CT image may further include the following operations:
[0056] S11, convert the format of the CT image;
[0057] S12, perform image pixel value normalization and image central region cropping processing on the CT image after format conversion.
[0058] In some preferred embodiments of S2, for extracting semantic feature information from the preprocessed CT images, the following operations may further be included:
[0059] S21, provide a feature extraction network model based on a U-shaped network and train the feature extraction network model;
[0060] S22, use the preprocessed CT images as the input of the trained feature extraction network model, output calcified plaques and detection branch regions, and use the calcified plaques and detection branch regions as semantic feature information.
[0061] In some preferred embodiments of S21, the feature extraction network model based on a U-shaped network may further include: four encoding modules and corresponding decoding modules; wherein, the encoding modules perform convolution and downsampling operations to gradually reduce the resolution of the input images and extract image features; the decoding modules perform deconvolution operations to gradually restore the image resolution and achieve feature fusion.
[0062] In some preferred embodiments of S3, for feature fusion of the preprocessed CT images and semantic feature information, the following operations may further be included:
[0063] S31, provide a global-local feature fusion network model;
[0064] S32, input the preprocessed CT images into the global branch of the global-local feature fusion network model and output a global feature map;
[0065] S33, input the semantic feature information into the local semantic branch of the global-local feature fusion network model and output a local feature map;
[0066] S34, merge the global feature map and the local feature map to obtain three-dimensional feature information after feature fusion.
[0067] In some preferred embodiments of S3, the generator is constructed based on an improved U-shaped network and a perception extraction module is introduced. Wherein:
[0068] The improved U-shaped network structure includes: five encoding sub-modules and five decoding sub-modules; wherein, the first encoding module is composed of a global-local feature fusion network model for performing feature fusion; a Squeeze-and-Excitation network is inserted in each skip layer to realize adaptive allocation of feature weights of different channels;
[0069] The perception extraction module is migrated from the encoding module part of the feature extraction network model for extracting semantic feature information to self-supervised learning, and is used to extract the feature maps of each layer of the improved U-shaped network and calculate the perception loss. The network structure and parameter settings before and after migration remain consistent.
[0070] The generator adopts the structured semantic loss L SSL to measure the semantic distance between the generated image and the gold standard, and assigns higher weights to the calcified plaque and vascular branch regions in the image, so that the network pays more attention to key information:
[0071]
[0072] In the formula, x and y are the images with and without calcification artifacts respectively, φ is the network migrated from the encoding part of the feature extraction network model, x s , y s are the semantically rich regions, γ 1 (γ 1 > 1) is the weight parameter, W, H, and C are the width, height, and number of channels of the feature map respectively, is the expectation function, and G(·) is the generator.
[0073] In some preferred embodiments of S3, the discriminator is used to determine whether the input image is a real sample or a fake sample generated by the generator, including: a local discriminator and a global discriminator; wherein, the local discriminator is used to map the input into a matrix of size N×N and discriminate the probability of each matrix element being a real sample. A semantic determination module is added to the local discriminator, and this added semantic determination module assigns corresponding weights to different regions of the image according to the importance of semantic information, so that the calcified plaque and the detected branch region in the image obtain higher weights than the background region, thereby making the overall discrimination process further integrate semantic information; the global discriminator is used to map the input into the probability that the entire generated image is a real sample; the loss function of the discriminator is:
[0074] L D = γ 2 L g +(L ps + γ 3 L pl )
[0075] In the formula, L g is the global loss, L ps , L pl represent the local losses of high semantic information and low semantic information respectively, γ 2 , γ 3 ∈(0, 1) are the weight parameters.
[0076] The following further elaborates in detail on the technical solution provided in the above embodiments of the present invention in conjunction with a preferred embodiment.
[0077] The method for removing calcification artifacts from CT images provided by this preferred embodiment includes the following steps:
[0078] Step 1, preprocess the CT image;
[0079] Step 2, extract key information (semantic feature information) from the preprocessed CT image;
[0080] Step 3, after fusing the CT image preprocessed in Step 1 with the key information extracted in Step 2, input it into the artifact removal model based on the adversarial network; among them, in the artifact removal model, the perceptual extraction module is introduced in the generator and the semantic determination module is added to the matrix discriminator to improve the artifact removal and noise reduction effect.
[0081] In a preferred embodiment of Step 1, the preprocessing operations on the CT image specifically include: image format conversion, image pixel value normalization, and image central cropping, etc. Among them, image format conversion involves re-encoding the pixel data of the image, modifying the resolution or color space of the image to ensure its correct representation in the new format. Image pixel value normalization is the process of mapping the pixel values of the image to a standard range or distribution, which usually involves linearly transforming the pixel values so that their range is between 0 and 1 or between -1 and 1. The purpose of normalization here is to eliminate the intensity differences between images so that the images are easier to process in the subsequent network, improving the stability and convergence speed of training. Image central cropping is the process of intercepting a specific-sized area from the center position of the image, which can ensure that important image content is located in the cropped image for training or inference, and is also used here to remove the noise around the image or exclude the influence of irrelevant background information.
[0082] In a preferred embodiment of Step 2, the extraction of key information from the CT image specifically includes: using a U-shaped network to construct a feature extraction network model for extracting key information. In this preferred embodiment, the key information to be extracted is the calcified plaque and coronary artery in CCTA. Specifically, there are three coronary arteries, namely the left anterior descending branch (LAD), the left circumflex branch (LCx), and the right coronary artery (RCA). The input of the feature extraction network model is H×W, and the output is H×W×C, where H and W respectively refer to the image height and width, and C refers to the number of channels. When constructing the feature extraction network, four encoding modules and corresponding decoding modules are used. Each encoding module consists of two convolutional layers and one max-pooling layer, and each decoding module consists of one transposed convolutional layer and two convolutional layers, and uses skip connections with the encoding module of the same feature dimension. Through the feature extraction network, the segmentation results of the calcified plaque and coronary artery can be obtained.
[0083] In a preferred embodiment of Step 3, the specific method of feature fusion includes: using the strategy of a global-local feature fusion network model, where global features help retain necessary context information, while local features can effectively filter information and focus more on semantically meaningful regions. The purpose of this global-local feature fusion network model is to better fuse the obtained semantic feature information with global information. The specific working principle is as Figure 3 shown. First, the first encoding module of the feature extraction network model is migrated to the global branch, and the output has the same width and height as the input, but the number of channels is correspondingly modified; then the same encoding module structure is placed in the local semantic branch. The training of the local semantic branch is cascaded and matches the input size of the subsequent artifact removal network model; finally, the output of the local semantic branch is merged with the output of the global branch to obtain three-dimensional feature information for input to the subsequent artifact removal network model. This can not only fully utilize the original global semantic information but also enable semantic factors focusing on local regions of interest to be added to the network output result.
[0084] In a preferred embodiment of Step 3, the specific method of using the artifact removal network model for noise reduction includes: first, the result of feature fusion of local features and global features is input into the generator to obtain a prediction result, and then this result is sent into the discriminator.
[0085] In a preferred embodiment of Step 3, the loss function used by the perception extraction module is a semantic similarity loss, which is improved on the basis of the perception loss. Its working principle is as Figure 4 shown. After obtaining the predicted artifact removal image through the generator, it is further input into the perception extraction module, and a semantic similarity loss is obtained as the loss function, and weight differentiation is carried out according to the importance of different semantics in the image domain. The specific formula is:
[0086]
[0087] In a preferred embodiment of Step 3, the specific method of constructing the generator includes: using an improved U-shaped network to construct the generator network, which is specifically composed of five encoding sub-modules and five decoding sub-modules. Except for the upsampling and downsampling layers in each module, each module is composed of two groups of "convolution, normalization, and Relu activation functions". Among them, the first encoding sub-module is replaced by a global-local feature fusion network model. At the same time, to further use the attention mechanism to help the network better combine global and local features, the 'Squeeze&Excitation' [scSE] model is inserted into each skip layer to enhance meaningful features and suppress redundant features.
[0088] In a preferred embodiment of Step 3, the specific method for constructing the matrix discriminator includes: The matrix discriminator consists of two parts, a local discriminator and a global discriminator, and its working principle is as Figure 5 shown. Each local discriminator focuses on different positions and semantics respectively, and is composed of four sub-modules, and each sub-module is composed of convolutional, normalization, LeakyReLU, and Dropout layers; at the same time, a global discriminator is added on the basis of the original concept. Patches with high semantic information among these discriminators will obtain high weight values, so that the overall discrimination process further integrates semantic information. The specific formula of the discriminator loss function is:
[0089] L D =γ 2 L g +(L ps +γ 3 L pl )
[0090] Next, in combination with a specific application example, the technical solutions provided in the above embodiments of the present invention will be further described. In this specific application example, the coronary artery is used as the detection area to remove the calcification artifacts in the CT image.
[0091] As Figure 2 shown, the method for removing the calcification artifacts in the coronary CT image adopted in this specific application example specifically includes the following steps:
[0092] The first step is to preprocess the coronary CT image;
[0093] The second step is to extract the key information of the coronary CT image;
[0094] The third step is to input the original data preprocessed in the first step and the key data extracted in the second step into the generative adversarial network for image reconstruction after fusion; the perception extraction module in the generator and the semantic determination module in the discriminator are used to improve the artifact removal and noise reduction effects.
[0095] Next, this specific application example will be described in detail.
[0096] The preprocessing operation of the coronary CT image in the first step specifically includes: converting the obtained CT data from the original DICOM, NIFTI and other medical image formats into the NumPy file format in Python (the file extension is.npy). Use the normalization function to normalize the gray value of each slice to the interval [0-1]. Then, remove the surrounding background information through central cropping to obtain a CT image with a size of [320, 320].
[0097] The specific extraction of key information in the second step refers to the extraction of semantic feature information from coronary CT images. In this step, calcified plaques and coronary arteries in CCTA are regarded as specific semantic feature information. The specific coronary arteries involve three major branches, namely the left anterior descending artery (LAD), the left circumflex artery (LCx), and the right coronary artery (RCA). Among them, the coronary artery region and the calcified plaque region are marked by clinicians, and two semantic annotation mask images are obtained by setting different thresholds respectively.
[0098] In the second step, a U-shaped network is used as the feature extraction network model to facilitate the development of the artifact removal work in the third step. Since CT images are grayscale images, the input of the feature extraction network model is H×W, and the output is H×W×C, where H is the height of the feature map, W is the width, and C is the number of channels. During training, the input of the feature extraction network model is the preprocessed CT image data, and the output is the semantic mask corresponding to the image. If the network depth of the feature extraction network model is too deep, some detailed information will be lost. Since the proportion of calcified plaques in the whole slice is small, this specific application example moderately adopts four encoding modules and corresponding four decoding modules to form a feature extraction network model based on the encoder-decoder structure. Each encoding module consists of two 3×3 convolutional layers and one 2×2 max pooling layer, and the number of convolutional kernels is 32, 64, 128, and 256 respectively; each decoding module consists of one 2×2 transposed convolutional layer and two 3×3 convolutional layers, and the number of convolutional kernels is 256, 128, 64, and 32 respectively. The ReLU function is used as the activation function for the convolutional layers, and the Dropout algorithm is used after the layers to prevent overfitting problems. Finally, the segmentation results of calcified plaques and coronary arteries are obtained, and the extraction of key information is completed.
[0099] After passing through the feature extraction network model, key semantic information can be obtained from CT images. In this specific application example, it is the coronary artery region and the calcified plaque region. It should be noted that if the slice does not contain key semantic information, such as coronary arteries or calcified plaques, the feature extraction network model will not extract the key regions, but the method and system of the present invention still have the functions of noise reduction and image quality improvement.
[0100] The original information and key semantic information obtained in the second step are fused through the global-local feature fusion network model and then input into the artifact removal network model. The fusion process is actually an application of the global-local fusion model (GLFM). In the specific application example, the input of the artifact removal network model in the third step comes from the combined features provided by the global-local feature fusion network model. The purpose of this feature fusion network model is to better fuse the obtained semantic feature information with the global information, such as Figure 3As shown. Global features help preserve necessary context information, while local features can effectively filter information and focus more on regions with semantic information. To fuse data from the two branches with the same weight, the two branches need to have the same size and the same number of channels. The global branch F Global has its architecture and hyperparameters migrated from the feature extraction network model, that is, the first encoding module of the feature extraction network model is migrated to the global branch. The output has the same width and height as the input, while the number of channels is changed from 1 to 16. The same encoder structure as the feature extraction network model is used for the local semantic branch F Local . Cascade training is carried out and matched with subsequent modules. Finally, the output of the local semantic branch is merged with the output of the global branch to obtain 3D feature information with 32 channels, which is input into the subsequent artifact removal network.
[0101] In the third step, the artifact removal network model after feature fusion is specifically composed of a generator and a discriminator module. First, the composition of the generator will be detailed below.
[0102] In this specific application example, an improved U-shaped network is used as the generator network. The generator consists of five encoding sub-modules and five corresponding decoding sub-modules. The function of the first encoding sub-module is replaced by a global-local fusion model. The input of the generator is the comprehensive features extracted from the original image and the semantic information provided by the fusion model, with a size of H×W×C, and the output is the prediction result with a size of H×W. Each encoding sub-module (except the first one) consists of two 3×3 convolutional layers, a batch normalization layer, and a 2×2 max pooling layer, where the number of convolutional kernels is 32, 64, 128, and 256 respectively; each decoding sub-module consists of a 2×2 transposed convolutional layer and two 3×3 convolutional layers, the transposed convolution stride is 2, and the number of convolutional kernels is 256, 128, 64, 32, and 16 respectively. The convolutional layer uses the ReLU activation function. At the same time, the attention mechanism is further used to help the network better combine global and local features, and the [scSE] model is inserted into each skip layer to enhance meaningful features and suppress useless features.
[0103] When measuring the loss of the network, this specific application example quotes the perceptual loss to replace the mean square error, so as to pay more attention to the evaluation at the semantic level. The perceptual loss measures the similarity between the feature maps of the original image and the predicted image after passing through the intermediate layer of the network, and punishes the output of the feature level that is perceptually unreasonable. The traditional perceptual loss is shown in formula (1), where VGG is the perceptual network, h, w, and c represent the height, width, and number of channels of the feature map respectively, and ||·|| is the Frobenius norm.
[0104]
[0105] The semantic similarity loss in this specific application example is improved based on the perceptual loss and is obtained through the perceptual extraction module, as shown in Equation (2).
[0106] First, the structure of the semantic similarity loss is migrated from the encoding part of the feature extraction network model to self-supervised learning. The parameters obtained by the feature extraction network model can better reflect the true features of coronary CT images, rather than using the parameters pre-trained by VGG or autoencoders. Since this network has been trained in previous tasks, its parameters can be directly used for semantic similarity loss extraction without increasing the overall training time of the network. In this specific application example, the semantic feature map is further used as a weight reference for the semantic similarity loss, so that regions with rich semantics, namely the calcified plaque region and the coronary artery branch region in this specific application example, have higher weights. The improved semantic similarity loss formula is as shown in Equation (2), where x and y are images with and without calcification artifacts respectively, φ is the network migrated from the encoding part of the feature extraction network model, x s , y s are regions with rich semantics, γ 1 is the weight parameter, and W, H, and C are the width, height, and number of channels of the feature map respectively.
[0107]
[0108] The matrix discriminator network architecture in the third step is described in detail below, as Figure 5 shown. The discriminator consists of four sub-modules, and each sub-module consists of a convolutional layer, a batch normalization layer, a LeakyReLU, and a dropout layer. Each sub-module compresses the height and width of the feature map to half of the original size, and the number of channels is successively doubled from 32 to 256, and the final size is Then, the 256 channels are compressed to form a local discriminator. Among them, each output X i,j judges whether the corresponding (i, j) block region in the image is true. Each local discriminator focuses on each different position and semantics respectively, and guides the generation network to adapt to more realistic images. On this basis, a global discriminator is used to make up for the lack of overall consistency. To avoid the problem that too many discriminators increase the number of trainable parameters and cause training instability, semantic feature information is further introduced, so that regions with high semantic information content obtain high weights, while the remaining regions are given low weights. The loss function of the specific discriminator network is as shown in Equation (3), where L g is the global loss, L ps , L pl represent the local losses of high semantic information and low semantic information respectively, γ 2 , γ 3is a weight parameter.
[0109] L D = γ 2 L g +(L ps + γ 3 L pl ) (3)
[0110] It should be noted that the extraction and application of semantic feature information can contribute to the overall framework of the present invention in three aspects. First, after the global-local feature fusion network model, the extracted semantic features can be combined with the original data into the network, and the features therein can be regarded as the prior knowledge of the original data, so as to better guide the network to focus on the feature regions. Second, after the generator obtains the prediction result, the original perceptual loss can be further inspired by the semantic features, making the weights of the image regions corresponding to the semantic features higher. Finally, in the discriminator, different from the traditional method of only using a "true / false" boolean value as the discrimination result, the predicted image can be further converted through convolution to obtain multiple information blocks. Each block is used to constitute a local loss function, and the importance of each local loss is determined by the semantic features. Combining the results of the local losses with the global discrimination loss weighted can help the discriminator make a better judgment.
[0111] Through the method for removing calcification artifacts from coronary CT images provided by the above embodiments of the present invention, the result performance diagrams of the artifact diameter removal ratios of different models in three blood vessels are as Figure 6 shown; the effect diagrams of different models in actual clinical data (test set) and the comparison diagrams of artifact removal with the gold standard are as Figure 7 shown.
[0112] An embodiment of the present invention provides a system for removing calcification artifacts from CT images, as Figure 8 shown, the system may include:
[0113] A data processing module, which is used to preprocess the CT image to obtain a preprocessed CT image;
[0114] A semantic extraction module, which is used to extract semantic feature information from the preprocessed CT image;
[0115] An artifact removal model module, which is used to provide an artifact removal model based on a confrontation network. The artifact removal model includes: a perceptual extraction module, a generator, and a discriminator. Among them, the preprocessed CT image is fused with semantic feature information and used as the input of the artifact removal model. The generator is used to generate the artifact-removed CT image, and a perceptual extraction module is introduced to extract the feature map of each layer of the neural network and calculate the perceptual loss. The discriminator is used to determine the probability that the generated image is real. A semantic determination module is added to the discriminator to determine the authenticity of the generated artifact-removed CT image through semantic information.
[0116] The following further describes each functional module provided in the above embodiments of the present invention.
[0117] The functional modules adopted by the calcification artifact removal system for CT images provided in the above embodiments of the present invention include:
[0118] Data processing module: preprocessing of CT images;
[0119] Semantic extraction module: a feature extraction network model;
[0120] Feature fusion module: a module that fuses global information and key information;
[0121] Artifact removal model module: includes a generator module and a discriminator module. Among them, the generator module introduces a perceptual extraction module to generate the artifact-removed CT image; the discriminator module adds a semantic determination module to determine the authenticity of the generated artifact-removed CT image through semantic information.
[0122] The data processing module specifically includes: image format conversion, image pixel value normalization, and image center cropping, etc. Among them, image format conversion involves re-encoding the pixel data of the image, modifying the resolution or color space of the image to ensure its correct representation in the new format. Image pixel value normalization is the process of mapping the pixel values of the image to a standard range or distribution, which usually involves linearly transforming the pixel values to eliminate the intensity differences between images, so that the images are easier to process in the subsequent network, improving the stability and convergence speed of training. Image center cropping is the process of intercepting a specific size area from the center position of the image, which can ensure that important image content is located in the cropped image for training or inference.
[0123] The semantic extraction module specifically includes: using a U-shaped network as the feature extraction network model and training it to extract calcified plaques and coronary arteries as key information.
[0124] The feature fusion module specifically includes: using the idea of global-local fusion, first migrating the first encoding module of the feature extraction network to the global branch, with the output having the same width and length as the input, but correspondingly modifying the number of channels; then placing the same encoding model structure in the local-semantic branch, where the training of the local-semantic branch is cascaded and matches the following network. A global-local fusion module is formed. Finally, the output of the local-semantic branch is merged with the output of the global branch to obtain three-dimensional feature information for input into the subsequent artifact removal network.
[0125] The generator module specifically includes: constructing the generator module using an improved U-shaped network, which is specifically composed of five encoding sub-modules and five decoding sub-modules. Among them, the first encoding sub-module is replaced by the global-local fusion module. Except for the necessary upsampling layer and downsampling layer in each module, each module is composed of two groups of "convolution, normalization, and ReLU activation function". At the same time, to further use the attention mechanism to help the network better combine global and local features, the 'Squeeze&Excitation' [scSE] model is inserted in each skip layer to enhance meaningful features.
[0126] The perceptual extraction module specifically includes: on the basis of the general perceptual loss, migrating from the encoding part of the semantic extraction network to self-supervised learning, proposing the semantic similarity loss, and at the same time further using the semantic feature map as the weight reference for the semantic similarity loss to build the perceptual extraction module.
[0127] The discriminator module specifically includes: this module consists of two parts, a local discriminator and a global discriminator. Each local discriminator focuses on different positions and semantics respectively and is composed of four sub-modules, and each sub-module is composed of convolutional, normalization, LeakyReLU, and Dropout layers; at the same time, the global discriminator is retained. Patches with high semantic information between local discriminators will obtain high weight values, so that the overall discrimination process further integrates semantic information.
[0128] It should be noted that the steps in the method provided by the present invention can be implemented by corresponding modules, devices, units, etc. in the system. Those skilled in the art can refer to the technical solution of the method to implement the composition of the system, that is, the embodiments in the method can be understood as preferred examples for constructing the system, which will not be elaborated here.
[0129] An embodiment of the present invention provides a computer terminal, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it can be used to execute the method of any one of the above embodiments of the present invention, or run the system of any one of the above embodiments of the present invention.
[0130] Optionally, a memory for storing programs; the memory may include volatile memory (e.g., random-access memory, such as static random-access memory (SRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDR SDRAM), etc.); the memory may also include non-volatile memory, such as flash memory. The memory is used to store computer programs (such as application programs and functional modules for implementing the above methods), computer instructions, etc. The above computer programs, computer instructions, etc. can be partitioned and stored in one or more memories. And the above computer programs, computer instructions, data, etc. can be called by the processor.
[0131] The above computer programs, computer instructions, etc. can be partitioned and stored in one or more memories. And the above computer programs, computer instructions, data, etc. can be called by the processor.
[0132] A processor for executing the computer programs stored in the memory to implement each step in the method or each module in the system according to the above embodiments. For specific details, reference can be made to the relevant descriptions in the previous method and system embodiments.
[0133] The processor and the memory can be of an independent structure or an integrated structure integrated together. When the processor and the memory are of an independent structure, the memory and the processor can be coupled and connected through a bus.
[0134] An embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it can be used to execute the method according to any one of the above embodiments of the present invention, or to run the system according to any one of the above embodiments of the present invention.
[0135] The method, system, terminal and medium for removing calcification artifacts in CT images provided by the above embodiments of the present invention propose a brand-new adversarial network model. By feature fusion, perceptual loss function and weighted matrix discrimination results, the semantic information obtained is fully utilized to guide the network to focus on the semantically concentrated areas; verified by real clinical data, the results show that it has a better artifact removal effect, can effectively assist doctors in improving the discrimination accuracy of stenosis degree, and thus reduce the misdiagnosis rate.
[0136] Those skilled in the art know that, in addition to implementing the system and its various devices provided by the present invention in the form of pure computer-readable program code, it is entirely possible to make the system and its various devices provided by the present invention implement the same functions in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. by logically programming the method steps. Therefore, the system and its various devices provided by the present invention can be regarded as a kind of hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component; the devices for implementing various functions can also be regarded as either software modules for implementing the method or the structures within the hardware component.
[0137] Matters not described in detail in the above embodiments of the present invention are all well-known technologies in the art.
[0138] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various deformations or modifications within the scope of the claims, which does not affect the essence of the present invention.
Claims
1. A method for removing calcification artifacts from CT images, characterized in that, it includes: Preprocessing the CT image to obtain a preprocessed CT image; Extracting semantic feature information from the preprocessed CT image; Providing an artifact removal model based on an adversarial network, the artifact removal model includes: a perceptual extraction module, a generator, and a discriminator; wherein, the preprocessed original CT image and the semantic feature information are fused as the input of the artifact removal model, the generator is used to generate a reconstructed CT image after artifact removal, and a perceptual extraction module is introduced to extract the feature map of each layer of the neural network and calculate the perceptual loss; the discriminator is used to determine the probability that the generated image is real, and a semantic determination module is added to the discriminator to determine the authenticity of the generated CT image after artifact removal through semantic information.
2. The method for removing calcification artifacts from CT images according to claim 1, characterized in that, The preprocessing of the CT image includes: Converting the format of the CT image; Performing image pixel value normalization and image center region cropping on the CT image after format conversion.
3. The method for removing calcification artifacts from CT images according to claim 1, characterized in that, The extraction of semantic feature information from the preprocessed CT image includes: Providing a feature extraction network model based on U-shaped network A, training the feature extraction network model, and using the preprocessed CT image as the input of the feature extraction network model, outputting calcified plaques and detection branch regions, and using the calcified plaques and detection branch regions as semantic feature information.
4. The method for removing calcification artifacts from CT images according to claim 3, characterized in that, The feature extraction network model based on U-shaped network A includes: four encoding modules and corresponding decoding modules; wherein, the encoding modules perform convolution and downsampling operations to gradually reduce the resolution of the input image and extract image features; the decoding modules perform deconvolution operations to gradually restore the image resolution and perform feature fusion through skip connections.
5. The method for removing calcification artifacts from CT images according to claim 1, characterized in that, The feature fusion of the preprocessed CT image and the semantic feature information includes: Providing a global-local feature fusion network model; Inputting the preprocessed CT image into the global branch of the global-local feature fusion network model to output a global feature map; Inputting the semantic feature information into the local semantic branch of the global-local feature fusion network model to output a local feature map; Combining the global feature map and the local feature map to obtain three-dimensional feature information after feature fusion.
6. The method for removing calcification artifacts from CT images according to claim 1, characterized in that, The generator is constructed based on a U-shaped network B and a perceptual extraction module is introduced; wherein: The U-shaped network B includes: five encoding sub-modules and five decoding sub-modules; among them, the first encoding module is composed of a global-local feature fusion network model for feature fusion, and a Squeeze-and-Excitation network is inserted into each skip layer to realize adaptive allocation of feature weights for different channels; The perception extraction module is used to extract the feature map of each layer of the U-shaped network B and calculate the perception loss; among them, the perception extraction module is migrated from the encoding module part of the feature extraction network model for extracting semantic feature information for self-supervised learning, and the network structure and parameter settings before and after migration remain consistent; The generator adopts a semantic similarity loss L SSL to measure the semantic distance between the generated image and the gold standard, and assigns higher weights to the calcified plaque and vascular branch regions in the image, so that the network pays more attention to key information: In the formula, is the desired function, x and y are the images with and without calcification artifacts respectively, φ(·) is the network migrated from the encoding part of the feature extraction network model, G(·) is the generator, x s , y s are the semantically rich regions, and γ 1 (γ 1 > 1) is the weight parameter, and W, H, and c are the width, height, and number of channels of the feature map respectively.
7. The method for removing calcification artifacts from CT images according to claim 1, characterized in that, The discriminator is used to determine whether the input image is a real sample or a fake sample generated by the generator, and includes: a local discriminator and a global discriminator; wherein, the local discriminator is used to map the input into a matrix of size n×n and discriminate the probability of each matrix element being a real sample. A semantic determination module is added to the local discriminator. The semantic determination module assigns corresponding weights to different regions of the image according to the importance of semantic information, so that the calcified plaque and the detection branch region in the image obtain higher weights than the background region, thereby making the overall discrimination process further integrate semantic information; the global discriminator is used to map the input into the probability that the entire generated image is a real sample; the loss function L of the discriminator D is as follows: L D = γ 2 L g +(L ps + γ 3 L pl ) where, L g is the global loss, and L ps , L pl represent the local losses of high semantic information and low semantic information respectively, and γ 2 , γ 3 ∈(0, 1) are weight parameters.
8. A system for removing calcification artifacts from CT images, characterized in that, comprising: A data processing module, which is used to preprocess the CT image to obtain a preprocessed CT image; A semantic extraction module, which is used to extract semantic feature information from the preprocessed CT image; An artifact removal model module, which is used to provide an artifact removal model based on an adversarial network. The artifact removal model includes: a perception extraction module, a generator, and a discriminator; among them, the preprocessed CT image and the semantic feature information are fused as the input of the artifact removal model. The generator is used to generate a CT image after artifact removal, and a perception extraction module is introduced to extract the feature map of each layer of the neural network and calculate the perception loss; the discriminator is used to determine the probability that the generated image is real, and a semantic determination module is added to the discriminator to determine the authenticity of the generated CT image after artifact removal through semantic information.
9. A computer terminal, including a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it can be used to execute the method described in any one of claims 1-7, or run the system described in claim 8.
10. A computer-readable storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by the processor, it can be used to execute the method described in any one of claims 1-7, or run the system described in claim 8.
Citation Information
Cited By
Image processing method based on deep learning
CN121903861A