Image artifact restoration method and system based on anti-fact diffusion model
Through the image artifact repair method based on the counterfactual diffusion model, the texture and structural denoising network are trained using artifact-free data, the problem of lack of paired data in the existing technology is solved, and efficient repair of real artifact data is achieved, image quality is improved and cost is reduced.
Patent Information
- Application Number
- CN202510027888.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-08
AI Technical Summary
The lack of paired data in image artifact repair in the prior artifact repair makes it difficult to effectively repair real artifact data. The traditional method has limited effect when dealing with artifacts of nonlinear and complex structures.
The image artifact repair method based on the counterfactual diffusion model is adopted. In the training stage, only artifact-free data is used, and the image artifact repair model is constructed through the texture denoising network, the structure denoising network and the discriminator network to achieve automatic artifact repair.
This method can effectively improve image quality and is suitable for a variety of imaging modes and artifact types, significantly reducing the repeated operations and repair steps in imaging examinations, and reducing the time and cost of hospitals during image repair.
Smart Images

Figure CN119941579A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and more particularly to an image artifact restoration method and system based on a counterfactual diffusion model. Background Art
[0002] Imaging images are an important part of modern medicine and play an important role in assisting clinical diagnosis and treatment. Imaging images can provide a clear and intuitive understanding of the patient's internal tissues and structures, helping doctors quickly locate and confirm the lesion area and develop a safer and more effective surgical plan. Therefore, the imaging quality of imaging images is extremely important. Although the acquisition speed and accuracy of imaging images are constantly being optimized with the upgrade and iteration of hardware and technology, due to the complexity of the imaging process, the imaging process is easily affected by various factors, resulting in the generation of noise and artifacts, which affects the subsequent use of the image.
[0003] Traditional methods for repairing image artifacts usually rely on image processing techniques, such as filtering technology, iterative reconstruction algorithms, projection interpolation methods, hybrid methods, and model correction methods. These traditional methods work well in certain scenarios and can improve image quality to a certain extent, but they often rely on manually designed features and parameters and lack the ability to automatically adapt to complex image content. Therefore, traditional methods have limited effects when dealing with artifacts with nonlinear and complex structures, and are difficult to generalize to multiple imaging modes and artifact types.
[0004] In the existing related technologies, the research on image artifact removal is based on deep learning methods. Compared with previous traditional methods, deep learning methods have shown improved performance and reduced running time. However, most deep learning methods are based on supervised learning methods when repairing image artifacts. Due to the difficulty in obtaining paired artifact-free and damaged artifact data, these methods usually use simulated image artifact images to train the network. Therefore, it is difficult to apply them to the repair of real artifact data. In order to overcome the limitations of deep learning methods based on synthetic data, deep learning methods using unpaired data have also been explored. Some methods regard the image artifact removal problem as a conversion between two image domains (artifact-free domain and artifact domain) and solve it based on a cyclic generative adversarial network. Although they utilize real image artifact data, the performance of these algorithms is often limited due to the lack of an explicit image artifact removal mechanism.
[0005] Therefore, how to propose an image artifact restoration method and system based on the counterfactual diffusion model, use only artifact-free data in the training stage, solve the limitation of the current artifact restoration technology due to the lack of paired data, and effectively improve the image quality is an urgent problem that technicians in this field need to solve. Summary of the invention
[0006] In view of this, the present invention provides an image artifact restoration method and system based on a counterfactual diffusion model, which only uses artifact-free data in the training stage to solve the limitation of the current artifact restoration technology due to the lack of paired data and effectively improve the image quality. In order to achieve the above purpose, the present invention adopts the following technical solutions:
[0007] An image artifact restoration method based on a counterfactual diffusion model, comprising:
[0008] Acquiring multimodal data, and performing data screening on the multimodal data;
[0009] Preprocess the screened data and construct a standardized data set;
[0010] Construct an image artifact restoration model based on the counterfactual diffusion model;
[0011] Training the image artifact restoration model using a standardized data set;
[0012] The artifacts of the image contaminated by artifacts are repaired using the trained image artifact repair model.
[0013] Optionally, the data screening includes: based on an imaging annotation method, making the screened image data an artifact-free image.
[0014] Optionally, the preprocessing of the filtered data includes:
[0015] Perform data normalization on data with different dimensions, modes and intensity values;
[0016] Split the 3D image along three axes and discard images with extreme aspect ratios;
[0017] The distribution of input data is unified through normalization and resizing operations.
[0018] Optionally, the image artifact restoration model includes: a texture denoising network, a structural denoising network and a discriminator network. The original image is gradually denoised through a forward diffusion process, and the image is gradually restored through a reverse denoising process. The texture denoising network adopts a UNet-based denoising network to repair the input image by denoising and denoising. The structural denoising network is used to guide the texture denoising process, and the discriminator network is used to measure the semantic correlation between the results of the structural denoising and the texture denoising results.
[0019] Optionally, the training of the image artifact restoration model using a standardized data set includes: learning the anatomical structure and context information of the artifact-free image by training only the artifact-free image.
[0020] Optionally, the method of learning the anatomical structure and contextual information of artifact-free images by training only artifact-free images includes: inputting the screened artifact-free images in the training phase to train the texture denoising network, the structure denoising network and the discriminator network, respectively training the texture denoising network and the structure denoising network through a joint optimization strategy, and in the forward process of the texture denoising network, gradually injecting Gaussian noise into the artifact-free image, and in the reverse process, using the denoising result of the structure denoising network to guide the texture denoising network to predict noise, and gradually reconstructing the image, and the discriminator network is used as a supervision mechanism to optimize the semantic consistency between the two.
[0021] Optionally, the training process of the discriminator network is:
[0022] The discriminator network D is trained to calculate the result y of the structure denoising network t-1 And the result of texture denoising network x t-1 The semantic relevance score between them is calculated using the discriminator loss L dis and triplet loss L tri To optimize the discriminator network;
[0023] An adaptive resampling strategy is used during training. t-1 and x t-1 When the semantic relevance score between t-1 Adding noise generation Then through the After denoising, the updated Afterwards use To guide the texture denoising network to generate updated Then, evaluate and Repeat the above steps to adaptively adjust the semantic correlation between the denoised original image and the structure.
[0024] Optionally, performing artifact restoration on an image contaminated by artifacts by using a trained image artifact restoration model includes: selectively performing denoising and resampling only in the artifact area to maximize retention of the original texture and anatomical structure of the artifact-free area.
[0025] Optionally, the performing artifact repair includes:
[0026] Determine the counterfactual probability according to the goal of artifact restoration, and redetermine the counterfactual probability according to the fact that the influence of the artifact is limited to a specific area, while other areas remain unchanged;
[0027] In the inference stage, the reverse diffusion resampling operation is performed only on the artifact area, while the original texture of the artifact-free area remains unchanged. The artifact image to be repaired and the corresponding artifact labeled image are input, and the sum of the diffused artifact-free area at time t-1 and the artifact area of the denoising resampling result at time t is taken as the repair result at time t-1; the repair result at time t-1 is used as the input of the next denoising step, that is, from t-1 to t-2, and the repair result obtained at each time step is used as the input of the next time step.
[0028] Optionally, an image artifact restoration system based on a counterfactual diffusion model comprises:
[0029] Data acquisition module: used to acquire multimodal data and perform data screening on the multimodal data;
[0030] Preprocessing module: used to preprocess the filtered data and build a standardized data set;
[0031] Model building module: used to build an image artifact restoration model based on the counterfactual diffusion model;
[0032] Model training module: used to train the image artifact restoration model using a standardized data set;
[0033] Artifact repair module: used to repair artifacts of images contaminated by artifacts through the trained image artifact repair model.
[0034] It can be seen from the above technical solutions that, compared with the prior art, the present invention discloses a method and system for image artifact restoration based on a counterfactual diffusion model, which has the following beneficial effects:
[0035] The present invention proposes a method for repairing image artifacts based on a counterfactual diffusion model, comprising: obtaining multimodal data, screening the multimodal data; preprocessing the screened data to construct a standardized data set; constructing an image artifact repair model based on the counterfactual diffusion model; training the image artifact repair model through a standardized data set; and performing artifact repair on an image contaminated by artifacts through the trained image artifact repair model. (1) The present invention solves the repair of most imaging image artifacts and is not limited to a single part of a single modality. At the same time, when processing images of different modalities, the present invention can still ensure the quality of the repaired image and maintain the high resolution and details of the image; (2) The present invention not only improves the technical performance of image repair, but also brings significant economic benefits, significantly reduces repeated operations and repair steps in imaging examinations, reduces the time and cost of hospitals in the image repair process, and can save costs for hospitals. At the same time, the technical application of the present invention drives the industrialization process of a new generation of medical image processing systems and provides strong technical support for the commercialization of related technologies. Especially in the areas of intelligent image processing equipment and remote medical diagnosis systems, it is expected to drive the continuous expansion of the market size; (3) Artifact repair can effectively improve image quality, especially in medical images. After artifact repair, the key details in the image are preserved, which enhances the doctor's judgment accuracy of the disease, enables patients to discover the disease in time, and provides strong support for the diagnosis and treatment of the patient's disease; (4) The technology of the present invention can promote the application and popularization of imaging artificial intelligence technology, promote the development of intelligent medicine, and reduce the demand for medical resources, especially in resource-constrained or remote areas, which has important promotion significance. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0037] Figure 1 A flow chart of an image artifact restoration method based on a counterfactual diffusion model provided by the present invention.
[0038] Figure 2 A schematic diagram of the training phase based on the counterfactual diffusion model provided by the present invention.
[0039] Figure 3 A schematic diagram of the reasoning stages based on the counterfactual diffusion model provided by the present invention.
[0040] Figure 4This is a schematic diagram of the comparison before and after the artifact restoration provided by the present invention. DETAILED DESCRIPTION
[0041] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0042] The embodiment of the present invention discloses a method for repairing image artifacts based on a counterfactual diffusion model, comprising:
[0043] Acquiring multimodal data, and performing data screening on the multimodal data;
[0044] Preprocess the screened data and construct a standardized data set;
[0045] Construct an image artifact restoration model based on the counterfactual diffusion model;
[0046] Training the image artifact restoration model using a standardized data set;
[0047] The artifacts of the image contaminated by artifacts are repaired using the trained image artifact repair model.
[0048] Furthermore, the data screening includes: based on the imaging annotation method, making the screened image data an artifact-free image.
[0049] Furthermore, the preprocessing of the filtered data includes:
[0050] Perform data normalization on data with different dimensions, modes and intensity values;
[0051] Split the 3D image along three axes and discard images with extreme aspect ratios;
[0052] The distribution of input data is unified through normalization and resizing operations.
[0053] Furthermore, the image artifact restoration model includes: a texture denoising network, a structure denoising network and a discriminator network. The original image is gradually denoised through a forward diffusion process, and the image is gradually restored through a reverse denoising process. The texture denoising network adopts a UNet-based denoising network to repair the input image by denoising and denoising. The structure denoising network is used to guide the texture denoising process, and the discriminator network is used to measure the semantic correlation between the structure denoising result and the texture denoising result.
[0054] Furthermore, the training of the image artifact restoration model using a standardized data set includes: learning the anatomical structure and context information of the artifact-free image by training only the artifact-free image.
[0055] Furthermore, the method of learning the anatomical structure and contextual information of artifact-free images by training only artifact-free images includes: inputting the screened artifact-free images in the training stage to train the texture denoising network, the structure denoising network and the discriminator network, respectively training the texture denoising network and the structure denoising network through a joint optimization strategy, and in the forward process of the texture denoising network, gradually injecting Gaussian noise into the artifact-free image, and in the reverse process, using the denoising result of the structure denoising network to guide the texture denoising network to predict noise, and gradually reconstructing the image, and the discriminator network is used as a supervision mechanism to optimize the semantic consistency between the two.
[0056] Furthermore, the training process of the discriminator network is:
[0057] The discriminator network D is trained to calculate the result y of the structure denoising network t-1 And the result of texture denoising network x t-1 The semantic relevance score between them is calculated using the discriminator loss L dis and triplet loss L tri To optimize the discriminator network;
[0058] An adaptive resampling strategy is used during training. t-1 and x t-1 When the semantic relevance score between t-1 Adding noise generation Then through the After denoising, the updated Afterwards use To guide the texture denoising network to generate updated Then, evaluate and Repeat the above steps to adaptively adjust the semantic correlation between the denoised original image and the structure.
[0059] Furthermore, the artifact restoration of the image contaminated by artifacts by using the trained image artifact restoration model includes: selectively performing denoising and resampling only in the artifact area to maximize the retention of the original texture and anatomical structure of the artifact-free area.
[0060] Further, the performing of artifact repair includes:
[0061] Determine the counterfactual probability according to the goal of artifact restoration, and redetermine the counterfactual probability according to the fact that the influence of the artifact is limited to a specific area, while other areas remain unchanged;
[0062] In the inference stage, the reverse diffusion resampling operation is performed only on the artifact area, while the original texture of the artifact-free area remains unchanged. The artifact image to be repaired and the corresponding artifact labeled image are input, and the sum of the diffused artifact-free area at time t-1 and the artifact area of the denoising resampling result at time t is taken as the repair result at time t-1; the repair result at time t-1 is used as the input of the next denoising step, that is, from t-1 to t-2, and the repair result obtained at each time step is used as the input of the next time step.
[0063] In a specific implementation, an image artifact restoration system based on a counterfactual diffusion model includes:
[0064] Data acquisition module: used to acquire multimodal data and perform data screening on the multimodal data;
[0065] Preprocessing module: used to preprocess the filtered data and build a standardized data set;
[0066] Model building module: used to build an image artifact restoration model based on the counterfactual diffusion model;
[0067] Model training module: used to train the image artifact restoration model using a standardized data set;
[0068] Artifact repair module: used to repair artifacts of images contaminated by artifacts through the trained image artifact repair model.
[0069] In a specific implementation, a method for repairing image artifacts based on a counterfactual diffusion model comprises the following specific steps:
[0070] S1. Data acquisition step: acquiring a large amount of medical imaging data and screening to obtain high-quality artifact-free data;
[0071] S2, data preprocessing step: normalize the collected various modal data;
[0072] S3. Construct an image artifact restoration model based on the counterfactual diffusion model, wherein the model includes three parts: a texture denoising network, a structure denoising network to guide the texture denoising process, and a discriminator network to measure the semantic correlation between the structure denoising results and the texture denoising results;
[0073] S4, training step: using artifact-free images to learn the ability to generate local anatomical result representations from contextual information;
[0074] S5, reasoning step: selectively perform denoising resampling only in artifact regions to maximize the preservation of the original texture and anatomical structure of artifact-free regions;
[0075] Furthermore, the S1 specifically includes the following steps:
[0076] S101. The first stage of building a dataset is to collect candidate images. We searched for images on various websites, such as TCIA, OpenNeuro, Grand Challenge, Synapse, GitHub, etc., and obtained a huge dataset of images after searching open resources related to medicine.
[0077] S102, screening the data collected in S101. In the data screening process, based on the labeling method of imaging experts, it is ensured that the screened image data only contain artifact-free images, so as to provide reliable training samples for subsequent artifact restoration tasks.
[0078] Furthermore, the S2 specifically includes the following steps:
[0079] S201, normalizing data of different data dimensions, modalities, and intensity values. Normalizing voxel or pixel values that vary greatly due to different modalities and acquisition methods into a unified format, including first normalizing the original image to [0, 1], and then multiplying by 255 to obtain an upper limit value.
[0080] S202, split all 3D images along three axes and discard images with extreme aspect ratios. Specifically, slice images whose shortest side length is less than half of the longest side length are discarded to prevent the target area from becoming extremely blurred after resizing.
[0081] S203. Through normalization and resizing operations, the distribution of input data is unified, reducing noise interference that may occur during model learning.
[0082] Furthermore, S3 specifically includes the following steps:
[0083] S301. An image artifact restoration model based on a counterfactual diffusion model is constructed. The counterfactual diffusion model proposed in the present invention is based on the theoretical framework of the diffusion model.
[0084] The diffusion model gradually adds noise to the original image through the forward diffusion process, and gradually restores the image through the reverse denoising process. In the forward diffusion process, the medical image x0 is perturbed by noise for T time steps, and gradually generates a Gaussian distributed image, whose conditional probability distribution is described by the following formula:
[0085]
[0086] Where N represents the normal distribution, I is the unit matrix, α t is the hyperparameter in the diffusion process, is a quantitative parameter that evolves over time step t, often referred to as the “cumulative noise coefficient”. In the reverse denoising process, the denoising network f θ (x t |t) Learn the conditional probability distribution P(x) t-1 |x t ), by gradually optimizing P(x t-1 |x t ) to achieve gradual repair of artifact areas.
[0087] 302. The counterfactual diffusion model constructed by the present invention includes three parts and two stages. The three main parts of the model are texture denoising network, structure denoising network and discriminator network. The texture denoising network adopts a denoising network based on UNet to achieve the purpose of restoration by denoising the input image. The structure denoising network is used to guide the texture denoising process, and the discriminator network is used to measure the semantic correlation between the results of structure denoising and texture denoising. The proposed structure-guided denoising network shows more outstanding ability in learning the anatomical structure representation of artifact-free images. The two main stages of the model are training and inference. In the training stage, the model learns the anatomical structure and contextual information of artifact-free images by training only artifact-free images. In the inference stage, denoising resampling is selectively performed only in the artifact area to maximize the retention of the original texture and anatomical structure of the artifact-free area.
[0088] Furthermore, the S4 specifically includes the following steps:
[0089] S401. The present invention does not require any paired artifact images and artifact-free images during the training phase, and can complete the model training by relying only on artifact-free images. Compared with the strong dependence on paired images in traditional methods, the present invention significantly reduces the requirements for data set construction and improves the model's generalization ability for complex artifacts in real scenes.
[0090] S402, in the training phase, the artifact-free image screened in step S1 is input, and then the three networks mentioned in S3 are trained respectively. The texture denoising network and the structure denoising network are trained respectively through a joint optimization strategy, in which the discriminator network is used as a supervision mechanism to optimize the semantic consistency between the two.
[0091] S403, firstly, the texture denoising network is trained. In the forward process, Gaussian noise is gradually injected into the artifact-free image. In the reverse process, the denoising result of the structural denoising network is used to guide the texture denoising network to predict the noise and gradually reconstruct the image. The structural denoising network introduced in the present invention uses the structure as a guide to improve the global semantic integrity and reliability of the restoration result.
[0092] S404. The specific contents of the structural denoising network training mentioned in S403 are as follows:
[0093] The input of the structural denoising network is the artifact-free image input in S402, which is then gradually denoised. During the denoising process, the semantics of the structure becomes increasingly sparse over time, and the input image is gradually degraded into a combination of a sparse edge map and Gaussian noise through noise scheduling. Next, the denoised result is denoised using a denoising network, and the denoised result is used to guide the texture denoising process at this time step. The present invention solves the shortcomings of the traditional diffusion model in maintaining the global structure by introducing a structural denoising network.
[0094] S405: training the discriminator D. The specific contents are as follows:
[0095] According to S404, in order to ensure the effectiveness of the results of the structure denoising network in guiding the texture denoising network to perform denoising, a discriminator network D is trained to calculate the result y of the structure denoising network t-1 And the result of texture denoising network x t-1 The semantic relevance score D(y t-1 ,x t-1 ,t-1)(abbreviated as D(y t-1 ,x t-1 )) to ensure the semantic relevance between structure and texture. At the same time, the discriminator loss L dis and triplet loss L tri To optimize D.
[0096] S406, an adaptive resampling strategy is used during training to adjust the semantic relevance between the original image and the structure according to the score value provided by the discriminator mentioned in S405. Specifically, when y t-1 and x t-1 When the semantic relevance score between t-1 Adding noise generation Then through the After denoising, the updated Afterwards use To guide the texture denoising network to generate updated Then, evaluate and The semantic correlation between the original image and the structure is adjusted adaptively by repeating the above steps.
[0097] S407: The computational cost of the diffusion model is high. In order to reduce the computational cost of the diffusion model during the reasoning process, the present invention adopts a progressive distillation method during training, aiming to accelerate the reasoning process by reducing the denoising steps of the diffusion model. The specific implementation is as follows:
[0098] First, the denoising network f trained after T denoising steps is θ (x t |t,m) as the teacher model, and build a student model based on it The student model inherits the parameters of the teacher model as initialization, but its denoising step is reduced to This reduces the amount of computation for each inference to half of the original amount. By minimizing the distillation loss function, the output of the student model is made consistent with the output of the teacher model at time step t-1, thus completing the distillation training. The objective function of the training is as follows:
[0099]
[0100] Among them, x t represents a noisy image with t time steps, L Distill is the knowledge distillation loss, m is the label indicating the artifact area, and the student model is distilled n times (i.e., each time the denoising step is reduced to the original ), and finally obtain a student model with significantly improved computational efficiency.
[0101] Furthermore, the S5 specifically includes the following steps:
[0102] S501, the goal of artifact repair is to restore the image x that is contaminated by artifacts a Restored to its ideal artifact-free state This is a counterfactual state that is unobservable in reality. Its counterfactual probability can be expressed as:
[0103]
[0104] Where do(C=c f ) is the intervention operator in causal inference, which means setting condition C to the artifact-free state c through external intervention f , x a represents an image contaminated by artifacts, c f Represents the ideal state without artifacts, Represents an ideal image without artifacts.
[0105] Specifically, in the artifact restoration task, the impact of the artifact is limited to a specific area, while other areas remain unchanged. Therefore, the label m is introduced to indicate the artifact area, and the label is used to distinguish the artifact area from the unaffected area. Specifically, the known image information can be expressed as:
[0106] X known =x a ⊙(1-m);
[0107] Among them, m is a binary image containing only 0 and 1, 1 represents the current pixel is an artifact area, that is, the area needs to be repaired, 0 represents the current pixel is an artifact-free area, that is, the normal anatomical structure area needs to be preserved, and ⊙ is a pixel-level multiplication operator. The counterfactual probability of repair in this case can be further expressed as:
[0108]
[0109] Among them, M = m represents the artifact area mark, which specifies the application area of the do operator. Through counterfactual reasoning, the image can be modeled with the assumption of an artifact-free state (i.e., an ideal state), thereby significantly reducing the impact of artifacts on subsequent image analysis tasks.
[0110] S502: In the inference stage, only the artifact area is subjected to the reverse diffusion resampling operation, while the artifact-free area keeps the original texture unchanged, so as to achieve maximum detail preservation and improve the restoration efficiency. The input is the artifact image to be restored and the corresponding artifact marker image. Finally, the diffusion artifact-free area at time t-1 is The artifact area of the denoising resampling result at time t The sum of the repair results y at time t-1 t-1 .
[0111] S503: According to S502, in the inference stage, the input image y is firstly denoised to obtain the denoised result at time t-1. Next, the noise-added results are Operation is performed to obtain the diffusion artifact-free area at time t-1.
[0112] S504: Next, the repair result at time t is Recorded as Denoising is performed under the guidance of the structure-guided network denoising results to obtain The denoising and resampling results from t to t-1 The artifact area of the denoising resampling result at time t is obtained, and finally the result is added to the result of the diffusion artifact-free area at time t-1 obtained in S502 to obtain the restoration result at time t-1
[0113] S505: Based on the repair result obtained in S504 It will be used as the input for the next denoising step, from t-1 to t-2 That is, the repair result obtained at each time step will be used as the input of the next time step.
[0114] In a specific implementation, a method for repairing image artifacts based on a counterfactual diffusion model, such as Figure 1 As shown, the specific steps are as follows:
[0115] Step S1, obtaining data: obtaining a large amount of medical imaging data, and screening to obtain high-quality artifact-free data;
[0116] Step S2, data preprocessing: normalizing the collected various modal data;
[0117] Step S3, constructing a model: constructing an image artifact restoration model based on a counterfactual diffusion model;
[0118] Step S4, training: no paired data is required, only artifact-free images are used to learn the ability to generate local anatomical result representations from contextual information;
[0119] Step S5, reasoning: selectively perform denoising resampling only in the artifact area to maximize the preservation of the original texture and image style of the artifact-free area.
[0120] In a specific embodiment, step S2 specifically includes the following steps:
[0121] Step S201: For a 2D image, normalize the voxel or pixel values to a consistent image format. Specifically, first convert the original image into Normalize to [0,1], then multiply by 255 to get the upper limit. min and x max are the minimum and maximum values in the image, respectively.
[0122] Step S202: For 3D images, all 3D images are split along three axes. For the slices of each axis of the 3D image, only 30 slices in the middle are retained. By selecting a portion of the slices in the middle, the core area of the anatomical structure can usually be captured, avoiding noise interference in the edge area, and achieving a more efficient processing effect. At the same time, images with extreme aspect ratios are discarded. Specifically, slice images whose shortest side length is less than half of the longest side length are discarded to prevent the target area from becoming extremely blurred during subsequent scaling. Afterwards, all processed images are scaled to 256 images and saved.
[0123] Step S203: For the artifact image used in the inference stage, the corresponding artifact area needs to be marked. The marked image is a binary image containing only 0 and 1, 1 represents that the current pixel is an artifact area, and 0 represents that the current pixel is an artifact-free area. The marking of artifacts is completed by radiologists with rich experience.
[0124] In a specific embodiment, step S3 specifically includes the following steps:
[0125] Step S301: construct an image artifact restoration model based on a counterfactual diffusion model. The counterfactual diffusion model proposed in the present invention is based on the theoretical framework of a diffusion model. The diffusion model gradually adds noise to the original image through a forward diffusion process, and gradually restores the image through a reverse denoising process.
[0126] Step S302, the image artifact restoration model based on the counterfactual diffusion model constructed in step S301 includes three parts and two stages. The three main parts of the model are texture denoising network, structure denoising network and discriminator network. The two main stages of the model are training and inference. In the training stage, the model learns the anatomical structure and contextual information of the artifact-free image by training only the artifact-free image. In the inference stage, denoising resampling is performed only in the artifact area to maximize the retention of the original texture and anatomical structure of the artifact-free area.
[0127] Step S303, the texture denoising network is based on the UNet architecture, and the basic structure is an encoder-decoder. The input image is extracted through the convolution layer and downsampling of the encoder, and then the spatial resolution of the image is gradually restored through the upsampling layer in the decoder. The final output image size is consistent with the input image. In texture denoising, the number of basic channels of the network is set to 64, and the network depth is set to 4. The network extracts and restores multi-level features by inputting the texture information of the color image and combining the conditional input. The encoding stage extracts texture features layer by layer, and the decoding stage combines the conditional input to restore the texture details of the image layer by layer.
[0128] Step S304, the structural denoising network of this embodiment uses the same denoising network as the texture denoising. At each layer, the input feature map is subjected to multi-level convolution and pooling operations to extract high-dimensional features, and the network weights are adjusted through conditional input (edge map) so that the network can make full use of the edge features of the input image. The encoding stage extracts features layer by layer, and the decoding stage restores the spatial resolution of the image layer by layer, and finally outputs the original image after denoising that is consistent with the input.
[0129] Step S305: The discriminator network of this embodiment uses the UNet architecture that is consistent with the structure denoising network. The discriminator loss L is also used. dis and triplet loss L tri To optimize D.
[0130] In a specific implementation, step S4 specifically includes the following steps:
[0131] Step S401, all algorithms of this embodiment are implemented on an NVIDIA GeForce RTX4090 GPU with 24GB of memory using PyTorch (version 1.13.0) and Python (version 3.8.20).
[0132] Step S402: In the embodiment disclosed in the embodiment, when training the texture denoising network, the structure denoising network and the discriminator network, the Adam optimizer is used to accelerate the convergence process, beta1 is set to 0.9, beta2 is set to 0.99, and the batch size is 16. The learning rate size is 1×10 -4 , the decay factor of the learning rate update is 0.5, the diffusion model time step T = 400, and the input image size is 256 × 256. The maximum noise intensity in the denoising stage is set to 30 and the minimum noise intensity is set to 0.005, ensuring that the noise is almost completely eliminated in the later stage of the diffusion process, thus generating high-quality images.
[0133] Step S403, in the structural denoising stage of the embodiment, the input image is the image to be repaired, and the final state of the denoising is the edge map of the image to be repaired plus Gaussian noise. The edge map is obtained by edge detection based on the input image using the Canny algorithm. The Canny algorithm realizes edge extraction by calculating the gradient amplitude and connecting the double threshold edge. The formula is as follows:
[0134]
[0135] in, and Represent the gradients in the x and y directions respectively. Through non-maximum suppression and double threshold screening, the edge image E is finally generated. edge .
[0136] Step S404: The specific contents of the structural denoising in step S403 are as follows:
[0137] The input data includes an artifact-free image and its corresponding edge map. The edge map is extracted from the input image by the Canny algorithm and used as a conditional input to guide the network's repair. When training the structure denoising network, the input image size is 256×256. The input image passes through the initial convolution layer, the feature map size remains at 256×256, and the number of channels increases to 64. After the first downsampling operation, the feature map size becomes 128×128, and the number of channels increases to 128. After the second downsampling operation, the feature map size becomes 64×64, and the number of channels increases to 256. After the third downsampling operation, the feature map size becomes 32×32, and the number of channels increases to 512. The feature map maintains a size of 32×32 and a number of channels of 512 in the middle layer, and the global features are further extracted through the attention module. In the upsampling stage, the feature map size is restored layer by layer, and multi-scale features are fused through jump connections, and finally restored to the original resolution of 256×256. The mean square error (MSE) is used as the loss function to optimize the difference between the generated image and the target image. The weight parameter of the conditional input is set to 0.1 to dynamically adjust the influence of the conditional input in structure denoising. Figure 2 As shown in the figure, the input of the structure denoising is an artifact-free image. In the forward process, the input image is gradually denoised. In the process of denoising, the structure of the input image becomes sparser. Finally, the corresponding edge image plus noise x is obtained. t Next, we use the UNet denoising network to predict the noise and get the denoising result x t-1 .
[0138] Step S405: When training the texture denoising network, the input data are the artifact-free image and the restoration result image of the structure denoising stage. The restoration result of the structure denoising stage will provide guidance for texture denoising, so that the model can better learn the anatomical structure of the artifact-free image during the training stage. The weight parameter of the conditional input is set to 0.1 to dynamically adjust the influence of the conditional input in texture restoration. The mean square error (MSE) is used as the loss function to optimize the error between the generated image and the target texture image. The changes in the network structure and feature map of texture denoising during the training stage are consistent with the changes in the structure denoising in step S404. The specific details of the structure denoising are as follows. Figure 2 As shown, the input is an artifact-free image. Next, the image is denoised to obtain the denoised result y t Then the structure x of the structure denoising is used t-1 To guide the texture denoising process, the texture denoising result y is obtained. t-1 .
[0139] Step S406: In order to ensure the effectiveness of structure denoising guiding texture denoising, this embodiment introduces a discriminator network D for calculating the semantic correlation score D(y) between structure and texture. t-1 ,x t-1 ,t-1). During training, the discriminator uses the texture denoising result y t-1 and the current time step structure denoising result x t-1 As input to optimize the semantic relevance between the two.
[0140] Step S407: The discriminator network mentioned in step S406 adopts a convolutional neural network (CNN) structure, combined with time step embedding, and gradually extracts features through the encoding layer, and finally generates a semantic relevance score for the input image. Its input is the current time step t and the texture denoising result y of the current time step t-1 And the structure denoising result x t-1 The discriminator mainly consists of three parts: time step embedding layer, convolutional coding layer, and feature standard deviation module. The discriminator introduces the time step embedding module to convert the input time step t into a fixed-dimensional embedding vector t emded, as the conditional input. The embedding dimension is 128, and after being processed by the activation function, it participates in the convolution layer calculation. After that, the input image and the artifact-optimized image are subjected to feature extraction through multiple encoding layers. The starting convolution layer adjusts the number of channels of the input image to the basic number of channels of the convolution feature. Each convolution block of the multi-level downsampling convolution block contains multiple convolution layers, and combines time step embedding to gradually reduce the image resolution and extract deep features. The first downsampling layer increases the number of channels to 128, the second downsampling layer increases the number of channels to 256, the third downsampling layer increases the number of channels to 512, and the last layer keeps the number of channels at 512. In the final stage of the discriminator, the feature standard deviation module calculates the standard deviation of the feature map to enhance the perception of global features. This module improves the discriminator's discriminative ability by splicing the standard deviation feature into the final feature map.
[0141] Step S408, the loss function of the discriminator includes the following two parts: The first is the discriminator loss:
[0142]
[0143] in, Represents the variable y t The expected value of Represents the variable y t-1 The expected value of y t-1 Indicates the texture denoising result of the current time step, y t represents the image before texture denoising, x t-1 The current one represents the structure denoising result, and the second one is the triplet loss:
[0144]
[0145] in, represents the combination of the artifact-free area before denoising and the artifact area after denoising, M represents the artifact area, α is the boundary parameter used to control the convergence of the loss, and ||·||2 represents the L2 norm. The total loss of the discriminator is:
[0146] L total =L dis +λ tri L tri ;
[0147] Among them, λ tri Set to 1 to balance the two losses.
[0148] Step S409: This embodiment adopts an adaptive resampling strategy during training, and adjusts the semantic correlation between the original image and the structure according to the score value provided by the discriminator mentioned in step S407. Specifically, when y t-1 and x t-1When the semantic relevance score between t-1 Adding noise generation Then through the After denoising, the updated Afterwards use To guide the texture denoising network to generate updated Then, evaluate and Repeat the above steps to adaptively adjust the semantic correlation between the denoised original image and the structure.
[0149] Step S410: In order to reduce the computational cost of the diffusion model during the inference process, this embodiment adopts a progressive distillation method during training, aiming to accelerate the inference process by reducing the denoising steps of the diffusion model. The specific implementation is as follows:
[0150] First, the denoising network f trained after T denoising steps is θ (x t |t,c,m) as the teacher model, and build a student model based on it The student model inherits the parameters of the teacher model as initialization, but its denoising step is reduced to This reduces the amount of computation for each inference to half of the original amount. By minimizing the distillation loss function, the output of the student model is made consistent with the output of the teacher model at time step t-1, thus completing the distillation training. The objective function of the training is as follows:
[0151]
[0152] Among them, x t represents a noisy image with t time steps. By distilling the student model n times (i.e., reducing the denoising step to the original ), and finally obtain a student model with significantly improved computational efficiency.
[0153] Step S501, diffusion model time step T = 400, input image size is 256 × 256. The maximum noise intensity in the denoising stage is set to 30, and the minimum noise intensity is set to 0.005, ensuring that the noise is almost completely eliminated in the later stage of the diffusion process, thereby generating a high-quality image.
[0154] Step S502: In the inference phase, the inference environment must first be initialized. By configuring the distributed training framework, the computing devices required for inference are determined, and the GPU devices are set to ensure operation efficiency. The inference configuration file is read to obtain model parameters, file paths, and data set settings. Then a path is created for saving the results to ensure that the results generated by the inference can be stored according to the preset path.
[0155] Step S503, loading the test dataset and the artifact labeled image dataset is an important step in the inference phase. The test dataset is used to evaluate the performance of the model in the actual artifact removal task, while the artifact labeled image dataset is used to provide annotation information of the image artifact area. The dataset loading mechanism supports cyclic loading of artifact labeled images to ensure a stable supply of artifact labeled image data in the inference phase.
[0156] Step S504: In the inference phase, it is necessary to load pre-trained models, specifically texture denoising model, structure denoising model, and discriminator model. The weights of the trained texture denoising model and structure denoising model are loaded to ensure that the model has the ability to restore texture details and structural information. The discriminator D is loaded to evaluate the semantic consistency of the generated results.
[0157] Step S505: In the inference stage, only the artifact area is subjected to the reverse diffusion resampling operation, while the artifact-free area keeps the original texture unchanged, thereby achieving maximum detail preservation and improving the restoration efficiency. The input is the artifact image to be restored and the corresponding artifact marker image. Finally, the diffusion artifact-free area at time t-1 is The artifact area of the denoising resampling result at time t The sum of the repair results y at time t-1 t-1 Among them, m is a binary image containing only 0 and 1, 1 represents that the current pixel is an artifact area, that is, the area that needs to be repaired, 0 represents that the current pixel is an artifact-free area, that is, the normal anatomical structure area that needs to be preserved, and ⊙ is a pixel-level multiplication operator.
[0158] Step S506: Figure 3 As shown, in the inference stage, the input image y is first denoised to obtain the denoised result at time t-1 Next, the noise-added results are Then, the diffusion artifact-free area at time t-1 is obtained. Recorded as Denoising is performed under the guidance of the structure-guided network denoising results to obtain The denoising and resampling results from t to t-1 The artifact area of the denoising resampling result at time t is obtained, and finally the result is added to the result of the diffusion artifact-free area at time t obtained previously to obtain the repair result at time t-1
[0159] Step S507: finally, save the repaired image result to a specified file path.
[0160] In a specific embodiment, 100 patients' MRIs are cut along the axis to obtain 2D slices, and only the middle 30 slices are taken from each MRI, so a total of 3000 MRI slices are obtained as a training set. Then two patients' MRIs are cut along the axis after adding motion artifacts, and 30 slices are taken from each MRI, so a total of 60 slices are obtained as a test set.
[0161] The image artifact restoration model based on the counterfactual diffusion model constructed in step S3 is used for training to obtain a pre-trained model.
[0162] Use the pre-trained model to perform inference on the test set with motion artifacts using the steps in step S5 to obtain the restored artifact-free data. Figure 4 The figure shows a schematic diagram of the comparison before and after the artifact repair of the present invention. The left side is an MRI with motion artifacts, and the right side is the repaired image. This embodiment can effectively improve the image quality through artifact repair. After the artifact repair, the key details in the image are retained, which enhances the doctor's judgment accuracy of the disease, enables patients to discover the disease in time, and provides strong support for the patient's disease diagnosis and treatment.
[0163] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0164] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for image artifact restoration based on a counterfactual diffusion model, characterized in that: include: Acquiring multimodal data, and performing data screening on the multimodal data; Preprocess the screened data and construct a standardized data set; Construct an image artifact restoration model based on the counterfactual diffusion model; Training the image artifact restoration model using a standardized data set; The artifacts of the image contaminated by artifacts are repaired using the trained image artifact repair model.
2. The image artifact restoration method based on the counterfactual diffusion model according to claim 1, characterized in that: The data screening includes: based on the imaging annotation method, making the screened image data an artifact-free image.
3. The image artifact restoration method based on the counterfactual diffusion model according to claim 1, characterized in that: The preprocessing of the filtered data includes: Perform data normalization on data with different dimensions, modes and intensity values; Split the 3D image along three axes and discard images with extreme aspect ratios; The distribution of input data is unified through normalization and resizing operations.
4. The method for image artifact restoration based on a counterfactual diffusion model according to claim 1, characterized in that: The image artifact restoration model includes: a texture denoising network, a structure denoising network and a discriminator network. The original image is gradually denoised through a forward diffusion process, and the image is gradually restored through a reverse denoising process. The texture denoising network adopts a denoising network based on UNet to perform restoration by denoising and denoising the input image. The structure denoising network is used to guide the texture denoising process, and the discriminator network is used to measure the semantic correlation between the structure denoising result and the texture denoising result.
5. The method for image artifact restoration based on counterfactual diffusion model according to claim 1, characterized in that: The training of the image artifact restoration model using a standardized data set includes: learning the anatomical structure and context information of the artifact-free image by training only the artifact-free image.
6. The method for image artifact restoration based on counterfactual diffusion model according to claim 5, characterized in that: The method of learning the anatomical structure and contextual information of artifact-free images by training only artifact-free images includes: inputting the artifact-free images obtained by screening in the training stage to train the texture denoising network, the structure denoising network and the discriminator network, respectively training the texture denoising network and the structure denoising network through a joint optimization strategy, gradually injecting Gaussian noise into the artifact-free images in the forward process of the texture denoising network, and in the reverse process, using the denoising result of the structure denoising network to guide the texture denoising network to predict noise, gradually reconstructing the image, and using the discriminator network as a supervision mechanism to optimize the semantic consistency between the two.
7. The method for image artifact restoration based on counterfactual diffusion model according to claim 6, characterized in that: The training process of the discriminator network is: By training the discriminator network D, the result y of the structure denoising network is calculated. t-1 And the result of texture denoising network x t-1 The semantic relevance score between them is calculated using the discriminator loss L dis and triplet loss L tri To optimize the discriminator network; An adaptive resampling strategy is used during training. t-1 and x t-1 When the semantic relevance score between t-1 Adding noise generation Then through the After denoising, the updated Afterwards use To guide the texture denoising network to generate updated Then, evaluate and Repeat the above steps to adaptively adjust the semantic correlation between the denoised original image and the structure.
8. The method for image artifact restoration based on counterfactual diffusion model according to claim 1, characterized in that: The artifact restoration of the image contaminated by artifacts by using the trained image artifact restoration model includes: performing denoising and resampling in the artifact area, and retaining the original texture and anatomical structure of the artifact-free area.
9. The method for image artifact restoration based on counterfactual diffusion model according to claim 8, characterized in that: The artifact repairing comprises: Determine the counterfactual probability according to the goal of artifact restoration, and redetermine the counterfactual probability according to the fact that the influence of the artifact is limited to a specific area, while other areas remain unchanged; In the inference stage, the reverse diffusion resampling operation is performed only on the artifact area, while the original texture of the artifact-free area remains unchanged. The artifact image to be repaired and the corresponding artifact labeled image are input, and the sum of the diffused artifact-free area at time t-1 and the artifact area of the denoising resampling result at time t is taken as the repair result at time t-1; the repair result at time t-1 is taken as the input for the next denoising step, that is, from time t-1 to t-2, and the repair result obtained at each time step is taken as the input for the next time step.
10. An image artifact restoration system based on a counterfactual diffusion model, characterized in that: include: Data acquisition module: used to acquire multimodal data and perform data screening on the multimodal data; Preprocessing module: used to preprocess the filtered data and build a standardized data set; Model building module: used to build an image artifact restoration model based on the counterfactual diffusion model; Model training module: used to train the image artifact restoration model using a standardized data set; Artifact repair module: used to repair artifacts of images contaminated by artifacts through the trained image artifact repair model.
Citation Information
Patent Citations
Depth counterfeit image detection method and system combined with multi-scale features
CN114724008A
CT (Computed Tomography) metal artifact removal method based on convergent diffusion model
CN116012478A
Method and system for removing metal artifacts of CT image based on diffusion model
CN116402911A
Metal artifact removal method based on probability diffusion model (DDMP)
CN117252943A
Knowledge graph completion method and system based on generation diffusion
CN117540797A
Cited By
Contrast image intelligent analysis method and system based on generative anti-fact interpretation
CN120598939A
Shielding artifact restoration method for medical image
CN120746898A
A method for repairing occlusion artifacts of medical images
CN120746898B
Medical image report generation method, system and device and storage medium
CN121439069A