CT image metal artifact correction method and device based on U-Net model, electronic equipment and storage medium
By introducing a multi-scale attention mechanism and an adaptive fusion module into the U-Net network model, the problem that the U-Net network model has difficulty capturing changes in metal traces in CT images is solved, and higher-precision metal artifact correction is achieved, reducing artifact residues and improving image quality.
Patent Information
- Application Number
- CN202510604408.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-09-16
AI Technical Summary
When processing metal artifacts in CT images, the existing U-Net network model has difficulty in effectively capturing the changes in metal traces at different scales of the projected sinusoidal graph, resulting in low correction accuracy and efficiency.
A multi-scale attention mechanism and adaptive fusion module are introduced to highlight the metal artifact area and tissue structure characteristics through multi-scale feature extraction, and the correction strategy is dynamically adjusted to avoid overfitting or underfitting.
Improved the accuracy of metal artifact correction, reduced residual artifacts, and improved image quality.
Smart Images

Figure CN120655748A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a method, device, electronic device, and storage medium for correcting metal artifacts in CT images based on a U-Net model. Background Art
[0002] Currently, computed tomography (CT) is widely used in medicine, industry, and other fields. However, when metal artifacts appear in CT images, the image quality will be severely degraded.
[0003] In related technologies, U-Net networks are commonly used for end-to-end metal artifact correction in CT images. However, in practical applications, it has been found that due to the large variations in the distribution and intensity of metal traces at different scales in the projected sinusoidal image, the skip connection method of the U-Net network model is often unable to effectively capture these variations, affecting the accuracy and efficiency of the network model in correcting metal artifacts.
[0004] In summary, the technical problems existing in the relevant technologies need to be improved. Summary of the Invention
[0005] The embodiments of the present application provide a method, device, electronic device and storage medium for correcting metal artifacts in CT images based on a U-Net model, which can improve the correction accuracy of metal artifacts, avoid overfitting or underfitting, effectively reduce artifact residues and improve the quality of metal artifact correction.
[0006] On the one hand, an embodiment of the present application provides a method for correcting metal artifacts in CT images based on a U-Net model, the method comprising the following steps:
[0007] Acquiring a CT image to be corrected;
[0008] Inputting the CT image to be corrected into a U-Net network model to obtain a CT image output by the U-Net network model that has completed metal artifact correction;
[0009] The U-Net network model includes a multi-scale attention mechanism and an adaptive fusion module.
[0010] Optionally, the multi-scale attention mechanism includes:
[0011] Obtaining encoder features and decoder features; wherein the encoder features and the decoder features are located at the same layer depth of the U-Net network model;
[0012] The encoder features are extracted using the first convolution kernel, the second convolution kernel, and the third convolution kernel, respectively, and the three feature extraction results are fused to obtain encoder multi-scale features;
[0013] Extracting features of the decoder using the first convolution kernel, the second convolution kernel, and the third convolution kernel respectively, and fusing the three feature extraction results to obtain multi-scale features of the decoder;
[0014] Obtaining a metal projection trace map, performing channel splicing on the encoder multi-scale features, the decoder multi-scale features, and the metal projection trace map, and generating attention weights through an adaptive attention mechanism;
[0015] Feature fusion is performed based on the generated attention weights, the encoder features, and the decoder features to obtain and output multi-scale attention features.
[0016] Optionally, the metal projection trace diagram is a metal projection trace diagram that has been binarized.
[0017] Optionally, the adaptive fusion module includes:
[0018] Obtain a binary metal projection trace image as conditional input;
[0019] Obtaining a pre-processed projection domain linear interpolation correction image;
[0020] Performing feature extraction on the binarized metal projection trace image and the projection domain linear interpolation correction image respectively, performing feature fusion on the feature extraction results, and generating adaptive attention weights through an adaptive attention mechanism;
[0021] Feature fusion is performed according to the adaptive attention weight, the binarized metal projection trace map and the projection domain linear interpolation correction image to obtain and output an adaptive fusion feature.
[0022] Optionally, the loss function of the U-Net network model includes reconstruction loss, consistency loss and mask guidance loss.
[0023] Optionally, the method further includes:
[0024] Constructing the reconstruction loss according to the mean square error;
[0025] constructing the consistency loss according to a structural similarity index;
[0026] The mask-guided loss is constructed according to the binarized metal projection trace map.
[0027] Optionally, the U-Net network model is trained based on the following steps:
[0028] Acquire multiple historical CT images, and obtain metal artifact correction results corresponding to the multiple historical CT images;
[0029] Taking each of the historical CT images as a sample and the metal artifact correction result corresponding to each of the historical CT images as a sample label corresponding to the sample, to construct a training data set;
[0030] The U-Net network model is pre-trained using the training data set.
[0031] On the other hand, an embodiment of the present application provides a CT image metal artifact correction device based on a U-Net model, the device comprising:
[0032] An image acquisition module, used for acquiring a CT image to be corrected;
[0033] An image correction module is used to input the CT image to be corrected into a U-Net network model to obtain a CT image with metal artifact correction completed, which is output by the U-Net network model;
[0034] The U-Net network model includes a multi-scale attention mechanism and an adaptive fusion module.
[0035] On the other hand, an embodiment of the present application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above-mentioned CT image metal artifact correction method based on the U-Net model.
[0036] On the other hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned CT image metal artifact correction method based on the U-Net model.
[0037] The embodiment of the present application introduces a multi-scale attention mechanism into the U-Net network model, which can perform multi-scale feature extraction in the skip connection, highlight the metal artifact area and tissue structure characteristics, improve the correction accuracy of the metal artifact, and dynamically adjust the correction strategy through the adaptive fusion module to avoid overfitting or underfitting, effectively reduce the artifact residue and improve the quality of metal artifact correction. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 Schematic diagram of an implementation environment of a method for correcting metal artifacts in CT images based on a U-Net model provided in an embodiment of the present application;
[0039] Figure 2 1 is a flow chart of a method for correcting metal artifacts in CT images based on a U-Net model provided in an embodiment of the present application;
[0040] Figure 3This is a schematic diagram of the structure of a multi-scale attention mechanism provided in an embodiment of the present application;
[0041] Figure 4 This is a schematic diagram of the structure of an adaptive fusion module provided in an embodiment of the present application;
[0042] Figure 5 1 is a structural diagram of a CT image metal artifact correction device based on a U-Net model provided in an embodiment of the present application;
[0043] Figure 6 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0044] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are merely examples of devices and methods consistent with some aspects of the embodiments of the present application as detailed in the appended claims.
[0045] It will be understood that the terms "first", "second", etc. used in this application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0046] The terms "at least one", "plurality", "each", "any", etc. used in this application include "at least one", "two" or more, "plurality" or "each", "any" or "any one", "each" or "any one" as used herein.
[0047] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0048] Currently, computed tomography (CT) is widely used in medicine, industry, and other fields. However, when metal artifacts appear in CT images, the image quality will be severely degraded.
[0049] In related technologies, U-Net networks are commonly used for end-to-end metal artifact correction in CT images. However, in practical applications, it has been found that due to the large variations in the distribution and intensity of metal traces at different scales in the projected sinusoidal image, the skip connection method of the U-Net network model is often unable to effectively capture these variations, affecting the accuracy and efficiency of the network model in correcting metal artifacts.
[0050] In view of this, the embodiments of the present application provide a method, device, electronic device and storage medium for metal artifact correction of CT images based on the U-Net model. By introducing a multi-scale attention mechanism into the U-Net network model, multi-scale feature extraction can be performed in the jump connection, highlighting the metal artifact area and tissue structure characteristics, improving the correction accuracy of the metal artifact, and dynamically adjusting the correction strategy through the adaptive fusion module to avoid overfitting or underfitting, effectively reducing artifact residue and improving the quality of metal artifact correction.
[0051] It should be noted that in each specific embodiment of the present application, when it comes to the need to perform relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of such data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present application needs to obtain the user's sensitive personal information, the user's separate permission or consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the user's separate permission or consent, the necessary user-related data for the normal operation of the embodiment of the present application will be obtained.
[0052] The following describes the specific implementation of the embodiment of the present application in detail with reference to the accompanying drawings. First, a method for correcting metal artifacts in CT images based on a U-Net model provided in the embodiment of the present application is described with reference to the accompanying drawings.
[0053] Please refer to Figure 1 , Figure 1 1 is a schematic diagram of an implementation environment for a method for correcting metal artifacts in CT images based on a U-Net model provided in an embodiment of the present application. In this implementation environment, the main hardware and software components involved include a terminal processor 110 and a server 120.
[0054] Specifically, the terminal processor 110 may be installed with a control program for a method for correcting metal artifacts in CT images based on a U-Net model, and the server 120 serves as the backend server for the control program. The terminal processor 110 and the backend server 120 are communicatively connected. The method for correcting metal artifacts in CT images based on a U-Net model provided in the embodiments of the present application may be executed on the terminal processor 110.
[0055] Server 120 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), as well as big data and artificial intelligence platforms.
[0056] In addition, the server 120 may also be a node server in the blockchain network.
[0057] A communication connection can be established between the terminal processor 110 and the server 120 via a wireless network. The wireless network uses standard communication technologies and / or protocols. The network can be the Internet or any other network, including, but not limited to, a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile network, or any combination of a wireless network, a private network, or a virtual private network. Furthermore, the aforementioned software and hardware entities can use the same or different communication connection methods, and this application does not impose any specific restrictions on this.
[0058] Of course, it is understandable that Figure 1 The implementation environment in the embodiment of the present application is only some optional application scenarios of the CT image metal artifact correction method based on the U-Net model. The actual application is not fixed. Figure 1 The software and hardware environment shown is not specifically limited in this application.
[0059] like Figure 2 As shown, Figure 2 This is a flow chart of a method for correcting metal artifacts in CT images based on a U-Net model provided in an embodiment of the present application, specifically including but not limited to steps 100 to 200.
[0060] Step 100: Acquire a CT image to be corrected.
[0061] In an embodiment of the present application, the CT image to be corrected can be obtained by responding to the CT image uploaded by the user, or according to a preset trigger mechanism, when the number of CT images to be corrected in the database meets a certain threshold, the step of obtaining the data of the CT image to be corrected is triggered, and all the CT images to be corrected are read from the database.
[0062] Step 200: Input the CT image to be corrected into a U-Net network model to obtain a CT image with metal artifact correction output by the U-Net network model;
[0063] The U-Net network model includes a multi-scale attention mechanism and an adaptive fusion module.
[0064] In an embodiment of the present application, the acquired CT image to be corrected is input into a U-Net network model, which performs feature extraction and metal artifact recognition, and finally outputs a CT image after metal artifact correction.
[0065] In practical applications, the original U-Net model performs end-to-end metal artifact correction on CT images, using convolution operations to extract metal artifact features from CT images and then learning to correct them through supervised training. In the downsampling encoding and upsampling decoding stages of the U-Net model, the U-Net network uses skip connections between the encoder and decoder at corresponding depths to fuse features extracted from the encoder's feature map into the corresponding decoder. This better preserves some image information and allows for better image detail recovery during the decoder's upsampling phase. However, the original U-Net model performs poorly on metal artifact correction in the projection domain because metal artifact information is primarily contained in the projection data derived from metal regions. This data appears as sinusoidal curves on the projection sinusoidal graph, known as metal traces. Because the distribution and intensity of metal traces vary significantly across different scales in the projection sinusoidal graph, the original U-Net's skip connection approach often fails to effectively capture these variations.
[0066] Therefore, this application introduces a multi-scale attention mechanism into the U-Net network model, which can perform multi-scale feature extraction in the jump connection, highlight the metal artifact area and tissue structure characteristics, improve the correction accuracy of metal artifacts, and dynamically adjust the correction strategy through the adaptive fusion module to avoid overfitting or underfitting, effectively reduce artifact residues and improve the quality of metal artifact correction.
[0067] Specifically, as an optional implementation, the multi-scale attention mechanism includes:
[0068] Obtaining encoder features and decoder features; wherein the encoder features and the decoder features are located at the same layer depth of the U-Net network model;
[0069] The encoder features are extracted using the first convolution kernel, the second convolution kernel, and the third convolution kernel, respectively, and the three feature extraction results are fused to obtain encoder multi-scale features;
[0070] Extracting features of the decoder using the first convolution kernel, the second convolution kernel, and the third convolution kernel respectively, and fusing the three feature extraction results to obtain multi-scale features of the decoder;
[0071] Obtaining a metal projection trace map, performing channel splicing on the encoder multi-scale features, the decoder multi-scale features, and the metal projection trace map, and generating attention weights through an adaptive attention mechanism;
[0072] Feature fusion is performed based on the generated attention weights, the encoder features, and the decoder features to obtain and output multi-scale attention features.
[0073] In the examples of this application, please refer to Figure 3 , Figure 3 This is a structural diagram of a multi-scale attention mechanism provided in an embodiment of the present application. The multi-scale attention mechanism introduced by the U-Net network model of the present application mainly obtains the encoder features and decoder features at the same layer depth (the encoder features and decoder features of each layer of the U-Net network model can be used to highlight the metal artifact area and tissue structure features through the multi-scale attention mechanism), and extracts feature maps of different scales through convolution kernels of different sizes, thereby helping the network model capture metal trace features at different scales during the jump connection process.
[0074] In practical applications, such as Figure 3 As shown in the figure, after obtaining the encoder features and decoder features at the same depth of the U-Net network model, a 3x3 convolution kernel can be used as the first convolution kernel, a 5x5 convolution kernel can be used as the second convolution kernel, and a 7x7 convolution kernel can be used as the third convolution kernel. The encoder features are extracted by the first convolution kernel, the second convolution kernel, and the third convolution kernel to obtain feature extraction results at three scales, and feature fusion is performed to obtain encoder multi-scale features. Correspondingly, multi-scale feature extraction can also be performed on the decoder features by the first convolution kernel, the second convolution kernel, and the third convolution kernel, and the feature extraction results are fused to obtain decoder multi-scale features.
[0075] Furthermore, after obtaining the encoder and decoder multi-scale features, the metal projection trace map can be further introduced. Encoder features are often closer to the original input features and contain rich spatial details, but may contain strong artifacts. Decoder features tend to be more like corrected features, but have relatively low resolution. Traditional U-Nets simply add or concatenate the two, failing to fully distinguish which regions should rely more on encoder features and which should utilize decoder features. Therefore, the encoder and decoder multi-scale features, along with the metal projection trace map, are channel-wise concatenated, and attention weights are generated using an adaptive attention mechanism. The adaptive attention mechanism generates attention weights using the following process: 1×1 convolution → batch normalization → ReLU rectified linear unit → second 1×1 convolution → sigmoid activation function. Finally, the attention weights are mapped to the range (0, 1). This allows the weights of feature maps at different scales and locations to be calculated, highlighting the features of metal trace regions. The introduction of the metal projection trace map helps the network model accurately locate metal trace regions and embed them into the feature space.
[0076] Finally, the network model is taught to favor the weights of the decoder features in metal artifact areas (mask = 1), and to favor the weights of the encoder features in non-metal areas (mask = 0). Specifically, the attention weights are multiplied by the decoder features, and the difference between the attention weights and 1 is multiplied by the encoder features. The two operation results are then channel-concatenated to obtain and output multi-scale attention features through feature fusion.
[0077] In practical applications, the attention weight is mapped to (0, 1). When the attention weight is close to 1, the output multi-scale attention feature is closer to the decoder feature. When the attention weight is close to 0, the output multi-scale attention feature is closer to the encoder feature.
[0078] In practical applications, the introduced metal projection trace map may be a metal projection trace map that has been binarized, which can better help the network accurately locate the metal trace area.
[0079] Therefore, this application introduces a multi-scale attention mechanism into the U-Net network model, effectively focusing on helping the network capture the metal trace features at different scales during the jump connection process. First, by obtaining feature maps of different scales, the characteristics of the metal traces in the projected sinusoidal graph are captured at different scales; secondly, for the metal traces at different scales and regions in the projected sinusoidal graph, their distribution and intensity vary greatly. Through the adaptive attention mechanism, the feature maps obtained at different scales and positions are weighted to highlight the metal trace areas and suppress irrelevant areas, so that the network model can dynamically adjust the correction strategy according to the input metal trace distribution, so that the network model can automatically learn to increase the correction intensity in the metal artifact area, while retaining the original information as much as possible in the non-artifact area.
[0080] Specifically, as an optional implementation, the adaptive fusion module includes:
[0081] Obtain a binary metal projection trace image as conditional input;
[0082] Obtaining a pre-processed projection domain linear interpolation correction image;
[0083] Performing feature extraction on the binarized metal projection trace image and the projection domain linear interpolation correction image respectively, performing feature fusion on the feature extraction results, and generating adaptive attention weights through an adaptive attention mechanism;
[0084] Feature fusion is performed according to the adaptive attention weight, the binarized metal projection trace map and the projection domain linear interpolation correction image to obtain and output an adaptive fusion feature.
[0085] In the examples of this application, please refer to Figure 4 , Figure 4 This is a structural diagram of an adaptive fusion module provided in an embodiment of the present application. The adaptive fusion module introduced by the U-Net network model of the present application calculates the fusion weight of the metal trace and projection sinusoidal graph features by taking the binarized metal projection trace map as an additional conditional input and combining it with the preprocessed linear interpolation corrected projection sinusoidal graph through adaptive weight generation, so that the network can dynamically adjust the correction strategy according to the input metal trace distribution to avoid overfitting or underfitting.
[0086] In practical applications, such as Figure 4As shown in the figure, after obtaining the binarized metal projection trace image and the projection domain linear interpolation correction image, feature extraction is performed on each of the two images through convolution operations. For example, a convolution operation with a kernel size of 3×3 can be used. Furthermore, the feature extraction results are concatenated using Concat channel concatenation to form feature maps. Adaptive attention weights are then generated using an adaptive attention mechanism. The features are then mapped to a single channel through a 1×1 convolution layer, a BN layer, and a RELU layer. The features are then normalized using a Sigmoid activation function for adaptive fusion.
[0087] Finally, the adaptive attention weight, the binarized metal projection trace map and the projection domain linear interpolation correction image are feature fused to obtain and output the adaptive fusion feature. Specifically, the adaptive attention weight is multiplied by the projection domain linear interpolation correction image, and the difference between the attention weight and 1 is multiplied by the binarized metal projection trace map, and the two operation results are channel-joined to obtain and output the adaptive fusion feature through feature fusion.
[0088] Therefore, this application effectively focuses on preprocessing the input projection sinusoidal graph by introducing an adaptive fusion module into the U-Net network model, that is, performing linear interpolation on the projection sinusoidal graph after removing the metal traces. However, linear interpolation cannot completely remove metal artifacts, and will also introduce new smoothing artifact problems. Through the adaptive fusion module, the network can dynamically adjust the correction strategy according to the input metal trace distribution, reduce the residual artifacts after correction and better retain the real detailed structure, thereby improving the network's ability to correct complex metal artifacts.
[0089] In practical applications, by introducing a multi-scale attention mechanism and an adaptive fusion module into the U-Net network model, the metal artifact correction quality of the network model can be effectively improved, as shown in Table 1. Table 1 shows the test results of metal artifact correction of CT images of different algorithm models. By selecting three different test samples and different algorithm models for CT image metal artifact correction tests, and selecting image quality evaluation indicators PSNR (peak signal-to-noise ratio) and SSIM (structural similarity index), the LIMAR algorithm, as a traditional MAR algorithm, has the lowest evaluation index result, while cGANMAR, as a MAR algorithm based on deep learning, has a slight improvement in the correction result in terms of image quality evaluation index. The original U-Net network model, thanks to its skip connection structure, can better preserve the detailed structure in the image. After introducing a multi-scale attention mechanism and an adaptive fusion module into the U-Net network model, the results obtained after correction on the three groups of test samples are better than the original U-Net model in terms of PSNR and SSIM, and the best evaluation indicators of PSNR and SSIM are achieved. Therefore, the U-Net model of this application can effectively reduce artifact residue and improve the quality of metal artifact correction.
[0090] Table 1 Test results of CT image metal artifact correction with different algorithm models
[0091]
[0092] Specifically, as an optional implementation, the loss function of the U-Net network model includes reconstruction loss, consistency loss and mask guidance loss.
[0093] In an embodiment of the present application, the reconstruction loss is used to measure the difference between the final output image of the network and the label image at the pixel level, which can directly constrain the accuracy of the network reconstruction in the image domain and reduce global brightness deviation and local color difference; the consistency loss is used to evaluate the consistency of the reconstructed image and the true image in structure and contrast, which can improve the perceptual quality of the network model (contrast, structural similarity), and the restored image is more consistent with the true value in texture and edges; the mask-guided loss is used to impose a stronger pixel-level penalty in the metal projection trace area, guiding the network to pay special attention to the reconstruction accuracy of these key areas.
[0094] As an optional implementation, the method further includes:
[0095] Constructing the reconstruction loss according to the mean square error;
[0096] constructing the consistency loss according to a structural similarity index;
[0097] The mask-guided loss is constructed according to the binarized metal projection trace map.
[0098] In an embodiment of the present application, the mean square error (MSE) can be selected to construct the reconstruction loss, and the structural similarity index (SSIM) can be selected to construct the consistency loss. The final loss function of the network model can also be constructed by introducing a binary metal projection trace map into the mask-guided loss and assigning weights to the three losses.
[0099] In practical applications, the network model's loss function is expressed as L = L1 + 0.1L2 + 0.5L3, where L is the U-Net network model's loss function, L1 is the reconstruction loss, L2 is the consistency loss, and L3 is the mask-guided loss. By combining these three loss functions, global pixel differences, local structure restoration, and focused correction of metal projection traces are taken into account. This not only helps the network learn more realistic and clearly structured images, but also ensures a higher degree of restoration in key metal projection trace areas, resulting in superior image reconstruction and artifact removal.
[0100] Specifically, as an optional implementation, the U-Net network model is trained based on the following steps:
[0101] Acquire multiple historical CT images, and obtain metal artifact correction results corresponding to the multiple historical CT images;
[0102] Taking each of the historical CT images as a sample and the metal artifact correction result corresponding to each of the historical CT images as a sample label corresponding to the sample, to construct a training data set;
[0103] The U-Net network model is pre-trained using the training data set.
[0104] In an embodiment of the present application, multiple historical CT images and corresponding metal artifact correction results of the multiple historical CT images can be obtained during the metal artifact correction processing of the historical CT images, and then any historical CT image is used as a sample and the metal artifact correction result of the any historical CT image is used as a sample label corresponding to the sample to form a group of training samples. A training data set is constructed by obtaining multiple training samples, and then the training data set is input into the U-Net model. The model parameters in the U-Net model are adjusted according to each output result of the U-Net model, and finally the pre-training process of the U-Net model is completed.
[0105] The pre-training process of the U-Net model can be considered completed after reaching a preset number of pre-training times; or the pre-training process of the U-Net model can be considered completed when the training output results of the U-Net model converge.
[0106] In practical applications, after constructing the training data set, it can also be divided into training set, validation set and test set. The training set accounts for 80% of the training data set, the validation set accounts for 10% of the training data set, and the test set accounts for 10% of the training data set. The training set is used for model training, and the validation set and test set are used for model verification.
[0107] See also Figure 5 , Figure 5 This is a schematic diagram of the structure of a CT image metal artifact correction device based on a U-Net model provided in an embodiment of the present application. This embodiment of the present application also provides a CT image metal artifact correction device based on a U-Net model, which can implement the above-mentioned CT image metal artifact correction method based on the U-Net model. The device includes:
[0108] An image acquisition module 510 is used to acquire a CT image to be corrected;
[0109] An image correction module 520 is configured to input the CT image to be corrected into a U-Net network model to obtain a CT image output by the U-Net network model that has completed metal artifact correction;
[0110] The U-Net network model includes a multi-scale attention mechanism and an adaptive fusion module.
[0111] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0112] See also Figure 6 , Figure 6 : This is a hardware structure diagram of an electronic device provided in an embodiment of the present application. The electronic device includes:
[0113] The processor 601 may be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.
[0114] The memory 602 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 602 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 602 and is called by the processor 601 to execute the CT image metal artifact correction method based on the U-Net model according to the embodiments of this application.
[0115] Input / output interface 603, used to implement information input and output;
[0116] Communication interface 604, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);
[0117] Bus 605 , which transmits information between various components of the device (e.g., processor 601 , memory 602 , input / output interface 603 , and communication interface 604 );
[0118] The processor 601 , the memory 602 , the input / output interface 603 and the communication interface 604 are connected to each other in communication within the device via a bus 605 .
[0119] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned CT image metal artifact correction method based on the U-Net model.
[0120] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiment, the functions specifically implemented by the present storage medium embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0121] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0122] The embodiments of the present application provide a method, device, electronic device, and storage medium for correcting metal artifacts in CT images based on a U-Net model. By introducing a multi-scale attention mechanism into the U-Net network model, the method can perform multi-scale feature extraction in skip connections, highlight metal artifact areas and tissue structure features, and improve the correction accuracy of metal artifacts. The correction strategy is dynamically adjusted through an adaptive fusion module to avoid overfitting or underfitting, effectively reducing artifact residues and improving the quality of metal artifact correction.
[0123] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0124] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0125] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0126] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.
[0127] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0128] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0129] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0130] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0131] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0132] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0133] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.
Claims
1. A method for correcting metal artifacts in CT images based on a U-Net model, characterized in that: The method comprises the following steps: Acquiring a CT image to be corrected; Inputting the CT image to be corrected into a U-Net network model to obtain a CT image output by the U-Net network model that has completed metal artifact correction; The U-Net network model includes a multi-scale attention mechanism and an adaptive fusion module.
2. The method for correcting metal artifacts in CT images based on the U-Net model according to claim 1, characterized in that: The multi-scale attention mechanism includes: Obtaining encoder features and decoder features; wherein the encoder features and the decoder features are located at the same layer depth of the U-Net network model; The encoder features are extracted using the first convolution kernel, the second convolution kernel, and the third convolution kernel, respectively, and the three feature extraction results are fused to obtain encoder multi-scale features; Using the first convolution kernel, the second convolution kernel, and the third convolution kernel to extract features of the decoder respectively, and fusing the three feature extraction results to obtain multi-scale features of the decoder; Obtaining a metal projection trace map, performing channel splicing on the encoder multi-scale features, the decoder multi-scale features, and the metal projection trace map, and generating attention weights through an adaptive attention mechanism; Feature fusion is performed based on the generated attention weights, the encoder features, and the decoder features to obtain and output multi-scale attention features.
3. The method for correcting metal artifacts in CT images based on the U-Net model according to claim 2, characterized in that: The metal projection trace diagram is a metal projection trace diagram that has been binarized.
4. The method for correcting metal artifacts in CT images based on the U-Net model according to claim 1, characterized in that: The adaptive fusion module includes: Obtain a binary metal projection trace image as conditional input; Obtaining a pre-processed projection domain linear interpolation correction image; Performing feature extraction on the binarized metal projection trace image and the projection domain linear interpolation correction image respectively, performing feature fusion on the feature extraction results, and generating adaptive attention weights through an adaptive attention mechanism; Feature fusion is performed according to the adaptive attention weight, the binarized metal projection trace map and the projection domain linear interpolation correction image to obtain and output an adaptive fusion feature.
5. The method for correcting metal artifacts in CT images based on the U-Net model according to claim 1, characterized in that: The loss function of the U-Net network model includes reconstruction loss, consistency loss and mask guidance loss.
6. The method for correcting metal artifacts in CT images based on the U-Net model according to claim 5, characterized in that: The method further comprises: Constructing the reconstruction loss according to the mean square error; constructing the consistency loss according to a structural similarity index; The mask-guided loss is constructed according to the binarized metal projection trace map.
7. The method for correcting metal artifacts in CT images based on the U-Net model according to claim 1, characterized in that: The U-Net network model is trained based on the following steps: Acquire multiple historical CT images, and obtain metal artifact correction results corresponding to the multiple historical CT images; Taking each of the historical CT images as a sample and the metal artifact correction result corresponding to each of the historical CT images as a sample label corresponding to the sample, to construct a training data set; The U-Net network model is pre-trained using the training data set.
8. A CT image metal artifact correction device based on a U-Net model, characterized in that: The device comprises: An image acquisition module, used for acquiring a CT image to be corrected; An image correction module is used to input the CT image to be corrected into a U-Net network model to obtain a CT image with metal artifact correction completed, which is output by the U-Net network model; The U-Net network model includes a multi-scale attention mechanism and an adaptive fusion module.
9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the CT image metal artifact correction method based on the U-Net model according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for correcting metal artifacts in CT images based on a U-Net model according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
PAM motion artifact removal method based on CBAM-U-Net
CN121788402A