Image noise removal method and device, computer equipment and storage medium

By using an improved diffusion model and artifact masking technology, artifacts in medical images are accurately located and removed, solving the problem of artifact recognition and removal in existing technologies, improving image quality and robustness, and making it suitable for applications in the medical and financial fields.

CN120852212APending Publication Date: 2025-10-28PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510950552.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately identify and remove artifacts in medical images, especially in complex scenarios, leading to decreased image quality and impacting the accuracy of AI diagnosis. Meanwhile, image denoising methods in the financial sector struggle to effectively handle interference from optical effects such as specular reflection.

Method used

An improved diffusion model is adopted, which extracts image features through an encoder. Combined with the OEA module and cross-attention mechanism, artifact masking is used to accurately locate and remove artifacts. The training sample set includes images with artifacts, artifact masks and clean images. The model is optimized using reconstruction loss, attention loss and background fidelity loss.

Benefits of technology

It achieves precise positioning and seamless removal of artifacts, improves image reconstruction quality and accuracy, possesses strong robustness and generalization ability, is applicable to artifact removal of various types of medical images, and ensures high-fidelity image output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120852212A_ABST
    Figure CN120852212A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, medical health and finance, and discloses an image noise removal method and device, computer equipment and a storage medium. The method comprises the following steps: acquiring an image to be processed and a corresponding artifact mask to obtain initial data; inputting the initial data into a noise removal model for noise removal to obtain a target image; wherein the noise removal model is obtained by training an improved diffusion model by taking an image with an artifact, a corresponding artifact mask and a corresponding clean image as a sample set; and outputting the target image. By implementing the method provided by the invention, accurate positioning and traceless removal of the artifacts can be realized, information loss caused by excessive repair is effectively avoided, the accuracy of artifact removal and the quality of image reconstruction are improved, and the method also has strong robustness and good generalization ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence, medical and health care and financial technology, and more specifically to image noise removal methods, apparatus, computer equipment and storage media. Background Technology

[0002] In the field of medical imaging technology, such as X-rays, CT (Computed Tomography), and MRI (Magnetic Resonance Imaging), image quality is crucial for accurate diagnosis. However, limitations of imaging equipment, varying levels of patient cooperation, and the presence of implanted metal devices often lead to artifacts or interfering objects in the images, such as catheters, identification tags, and non-target areas like bed outlines. These factors can obscure critical tissue structures, interfere with physician judgment, and even affect the accuracy of AI-based automated diagnostic systems.

[0003] Furthermore, in the fintech sector, as the financial industry's requirements for data security and accuracy continue to increase, the trend of cross-industry technological integration is becoming increasingly apparent. Fintech companies are exploring how to apply advanced medical image processing technologies to areas such as risk assessment, customer verification, and fraud detection. For example, in the identity verification process, by analyzing high-resolution images of identity documents, similar techniques can be used to remove background noise or unnecessary elements, improving recognition efficiency and accuracy. Simultaneously, to ensure the security of the transaction environment, denoising algorithms similar to those used in medical image processing are employed to clean up interfering information in surveillance videos, thereby enabling more accurate behavioral analysis and anomaly detection. Traditional image denoising methods or image restoration techniques often struggle to accurately identify interfering objects and their boundary features in such complex scenarios, especially when specular reflections or other optical effects are involved, such as reflections in a mirror.

[0004] Therefore, it is necessary to design a new method to achieve precise localization and seamless removal of artifacts, effectively avoiding information loss caused by over-repair. This not only improves the accuracy of artifact removal and the quality of image reconstruction, but also has strong robustness and good generalization ability. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of the prior art and provide an image noise removal method, apparatus, computer equipment, and storage medium.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: an image noise removal method, comprising:

[0007] Obtain the image to be processed and the corresponding artifact mask to obtain the initial data;

[0008] The initial data is input into a noise removal model for noise removal to obtain the target image; wherein the noise removal model is obtained by training an improved diffusion model using an image with artifacts, the corresponding artifact mask, and the corresponding clean image as a sample set;

[0009] Output the target image.

[0010] The further technical solution is as follows: the improved diffusion model includes an encoder, an OEA module, and a cross-attention mechanism.

[0011] The further technical solution is as follows: inputting the initial data into a noise removal model for noise removal to obtain the target image includes:

[0012] The initial data is input into the noise removal model, and the encoder is used to extract the latent features corresponding to the initial data.

[0013] The OEA module is used to extract the foreground features of artifacts and the influence features of artifacts on the surrounding space from the latent features to obtain intermediate features;

[0014] The intermediate features are processed to obtain contextual information;

[0015] The target image is obtained by using a cross-attention mechanism with contextual information.

[0016] The further technical solution is as follows: the potential features include the semantic features of the image to be processed and the structural features of the artifact mask.

[0017] The further technical solution is as follows: the extraction of latent features corresponding to the initial data using the encoder includes:

[0018] The semantic features of the image to be processed are extracted using a ResNet network, and the structural features of the artifact mask are obtained using a CNN network.

[0019] The further technical solution is as follows: The OEA module is used to extract foreground features of artifacts and the influence features of artifacts on the surrounding space from the latent features to obtain intermediate features, including:

[0020] The OEA module is used to extract the foreground features of artifacts from the latent features;

[0021] The artifact mask is dilated to obtain an expanded artifact mask;

[0022] The foreground features of the artifact and the latent features corresponding to the dilated artifact mask are convolved and then concatenated to obtain combined features;

[0023] The combined features are processed using an attention mechanism to extract the influence features of artifacts on the surrounding space, in order to generate intermediate features.

[0024] The further technical solution is as follows: the loss function used in the training process of the noise removal model includes a function formed by averaging the reconstruction loss, attention loss and background fidelity loss.

[0025] The present invention also provides an image noise removal apparatus, comprising:

[0026] The acquisition unit is used to acquire the image to be processed and the corresponding artifact mask to obtain initial data;

[0027] The noise removal unit is used to input the initial data into the noise removal model for noise removal to obtain the target image; wherein the noise removal model is obtained by training an improved diffusion model using an image with artifacts, the corresponding artifact mask, and the corresponding clean image as a sample set;

[0028] The output unit is used to output the target image.

[0029] The present invention also provides a computer device, the computer device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the above-described method.

[0030] The present invention also provides a storage medium storing a computer program that, when executed by a processor, implements the above-described method.

[0031] The advantages of this invention compared to existing technologies are as follows: This invention obtains the image to be processed and its corresponding artifact mask as initial data, and inputs this data into a specially trained noise removal model, achieving accurate artifact localization and seamless removal. This model is built based on an improved diffusion model, using the image with artifacts, the corresponding artifact mask, and the clean target image as the training sample set. It can effectively learn and separate artifact features and their spatial effects, while preserving important structural information and details of the original image, avoiding information loss due to over-repair. With this sophisticated design, this method not only significantly improves the accuracy of artifact removal and the quality of image reconstruction, but also exhibits strong robustness and good generalization ability, applicable to artifact removal of various types of medical images, ensuring high-fidelity image output. Finally, the processed target image is output, providing high-quality data support for subsequent applications such as clinical diagnosis.

[0032] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. Attached Figure Description

[0033] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0034] Figure 1 This is a schematic diagram illustrating an application scenario of the image noise removal method provided in this embodiment of the invention;

[0035] Figure 2 This is a schematic flowchart of the image noise removal method provided in an embodiment of the present invention;

[0036] Figure 3 This is a schematic diagram of a sub-process of the image noise removal method provided in an embodiment of the present invention;

[0037] Figure 4 This is a schematic diagram of a sub-process of the image noise removal method provided in an embodiment of the present invention;

[0038] Figure 5 This is a schematic block diagram of an image noise removal apparatus provided in an embodiment of the present invention;

[0039] Figure 6 A schematic block diagram of the noise removal unit of the image noise removal apparatus provided in an embodiment of the present invention;

[0040] Figure 7 A schematic block diagram of the second extraction subunit of the image noise removal apparatus provided in an embodiment of the present invention;

[0041] Figure 8 A schematic block diagram of a computer device provided for an embodiment of the present invention. Detailed Implementation

[0042] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0043] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0044] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0045] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0046] Please see Figure 1 and Figure 2 , Figure 1 This is a schematic diagram illustrating an application scenario of the image noise removal method provided in an embodiment of the present invention. Figure 2 This is a schematic flowchart illustrating the image noise removal method provided in this embodiment of the invention. The method is applied in a server. The server interacts with the terminal, acquiring the image to be processed and its artifact mask as initial data, which is then input into a noise removal model comprising an encoder, an OEA (Object Extraction and Analysis) module, and a cross-attention mechanism. The encoder extracts semantic and structural features, the OEA module further analyzes the foreground features of artifacts and their impact on the surrounding space, combines dilation operations and an attention mechanism to generate intermediate features, and finally utilizes contextual information through the cross-attention mechanism to obtain an artifact-free target image. The entire process is optimized by a comprehensive loss function including reconstruction loss, attention loss, and background fidelity loss, achieving accurate artifact localization and seamless removal, effectively avoiding information loss due to over-reconstruction, improving the accuracy of artifact removal and image reconstruction quality, and enhancing the model's robustness and generalization ability.

[0047] Figure 2 This is a schematic flowchart of the image noise removal method provided in an embodiment of the present invention. Figure 2 As shown, the method includes the following steps S110 to S130.

[0048] S110. Obtain the image to be processed and the corresponding artifact mask to obtain the initial data.

[0049] In this embodiment, the initial data refers to data pairs containing the original medical image and its corresponding artifact mask. Specifically:

[0050] Ifull (raw medical image with artifacts): This refers to the original medical image containing artifacts without any processing. Examples include metallic artifacts that may appear in CT scans, or motion blur in MRI imaging. These images are derived directly from actual medical examination equipment or generated through simulators to reflect various artifacts present in the real world.

[0051] Martifact: This is a marker map corresponding to Ifull, used to clearly identify the location and extent of artifacts in an image. The artifact mask can be a binary map, clearly distinguishing artifact regions from other parts; or it can be a soft map, providing information on the gradual changes in artifact intensity. For clinical images, this mask usually requires manual annotation by professionals; while for simulated images, it can be automatically generated by algorithms.

[0052] These two data sets form the foundation for model training, representing two key components in each training sample. They not only provide the model with a learning objective but also guide the OEA module in identifying and separating artifact features from their spatial influence regions, thereby assisting in the subsequent denoising process. Based on this, the system can more accurately locate artifact positions, analyze their diffusion effects, and ultimately achieve effective artifact removal without compromising the important structural information of the original image.

[0053] This step is crucial in ensuring the accuracy of the data input into subsequent processing flows, which is essential for improving the overall system performance. It also lays the foundation for subsequent data preprocessing and feature extraction.

[0054] S120. Input the initial data into the noise removal model to remove noise and obtain the target image.

[0055] In this embodiment, the target image refers to the medical image Igen generated after artifact removal through the improved diffusion model. This image is as close as possible to the original clean image Iclean, while retaining important structural information and details from the original image, avoiding information loss caused by "over-repair".

[0056] The noise removal model is obtained by training an improved diffusion model using an image with artifacts (Ifull), the corresponding artifact mask (Martifact), and the corresponding clean image (Iclean) as a sample set. The improved diffusion model includes an encoder, an OEA module, and a cross-attention mechanism.

[0057] The training process of the noise removal model uses a loss function that is the average of the reconstruction loss, attention loss, and background fidelity loss.

[0058] In one embodiment, please refer to Figure 3 The above-mentioned step S120 may include steps S121 to S124.

[0059] S121. Input the initial data into the noise removal model and use the encoder to extract the latent features corresponding to the initial data.

[0060] In this embodiment, latent features refer to key information extracted from the input raw medical image Ifull and its corresponding artifact mask Martifact. Specifically:

[0061] Semantic features of the image to be processed: High-level semantic information extracted from Ifull using the ResNet network.

[0062] Structural features of artifact masks: Low-level structural information extracted from artifacts using a CNN network.

[0063] These latent features collectively constitute a deep understanding of the original image and its artifact regions, providing a foundation for subsequent steps.

[0064] Specifically, the semantic features of the image to be processed are extracted using a ResNet network, and the structural features of the artifact mask are obtained using a CNN network.

[0065] The latent features include the semantic features of the image to be processed and the structural features of the artifact mask.

[0066] S122. Using the OEA module, extract the foreground features of the artifacts and the influence features of the artifacts on the surrounding space from the latent features to obtain intermediate features.

[0067] In this embodiment, intermediate features refer to a more refined description of the artifacts obtained after processing by the OEA module, including the foreground features of the artifacts themselves and their impact on the surrounding space.

[0068] In one embodiment, please refer to Figure 4 The above step S122 may include steps S1221 to S1224.

[0069] S1221. Use the OEA module to extract the foreground features of the artifacts from the latent features.

[0070] In this embodiment, the foreground features of artifacts refer to the salient regions or structures directly related to the artifacts extracted from the input image Ifull and the artifact mask Martifact by the OEA module. These features typically include the specific shape, size, and location information of the artifact in the image. The purpose of extracting these features is to accurately locate and process the artifacts subsequently, ensuring that the removal process is as accurate as possible without damaging the surrounding normal medical structures.

[0071] S1222. Perform a dilation operation on the artifact mask to obtain a dilated artifact mask.

[0072] In this embodiment, the dilated artifact mask is obtained by performing a morphological dilation operation on the original artifact mask Martifact. The dilation operation expands the boundaries of the artifact region, thereby capturing the extent of the artifact's influence on its surrounding space. This expansion helps the model to more comprehensively understand the spatial diffusion effect of the artifact, thus improving the accuracy of the denoising process. The dilated artifact mask not only contains the location information of the original artifact but also additionally covers the neighboring regions that the artifact may affect.

[0073] S1223. The foreground features of the artifact and the latent features corresponding to the dilated artifact mask are convolved and then concatenated to obtain combined features.

[0074] In this embodiment, the combined feature refers to the comprehensive representation formed through the following steps:

[0075] First, convolution operations are performed on the foreground features of the artifacts and the latent features corresponding to the dilated artifact mask. This step aims to further refine the features and enhance the model's focus on specific regions.

[0076] Next, the two convolutional results are concatenated along the channel dimension to form a combined feature containing more detailed information. This combined feature encompasses information about the core part of the artifact and its extended affected area, providing rich context for the subsequent attention mechanism.

[0077] S1224. The combined features are processed using an attention mechanism to extract the influence features of artifacts on the surrounding space in order to generate intermediate features.

[0078] In this embodiment, an attention mechanism is used to process combined features to extract the impact features of artifacts on the surrounding space and generate intermediate features. Specifically, the attention mechanism can dynamically adjust the model's attention to different regions, enabling the model to focus on the parts that most need repair while reducing interference with non-critical regions.

[0079] Based on the combined features obtained in the previous steps, the attention mechanism can identify and highlight the interaction patterns between artifacts and their surrounding environment, especially how artifacts alter or disrupt the structural properties of their neighboring regions.

[0080] The final output is an intermediate feature representation that integrates the foreground features of the artifact itself and its impact on the surrounding space. This provides guidance for the subsequent cross-attention module, helping the model to perform artifact removal tasks more accurately while preserving important background structural information.

[0081] Through this series of meticulously designed steps, the system proposed in this invention can effectively remove various types of medical image artifacts, such as CT metal artifacts and MRI motion blur, while ensuring high fidelity, greatly improving image quality and making it suitable for multiple application scenarios such as clinical diagnosis.

[0082] S123. Process the intermediate features to obtain context information.

[0083] In this embodiment, contextual information refers to a comprehensive representation that combines artifact features with information about its surrounding environment. It not only covers the specific location and shape of the artifact, but also its potential impact on neighboring areas, providing necessary clues for accurate artifact removal.

[0084] S124. Use the context information and employ a cross-attention mechanism to process the target image.

[0085] This step aims to use contextual information to guide the model to focus on the correct cleanup areas and reconstruct the structural background. Through a cross-attention mechanism, the model can effectively integrate global and local information, ensuring that artifacts are accurately removed while maintaining the integrity of the rest of the image. The final output target image, Igen, should be a high-quality, artifact-free, and high-fidelity medical image suitable for clinical diagnosis and other medical applications.

[0086] S130, Output the target image.

[0087] In this embodiment, the target image is output to the terminal for display.

[0088] By explicitly perceiving and modeling the spatial influence of artifact regions, the accuracy of artifact removal and the structural consistency of the generated image are effectively improved.

[0089] In this system, the input image `Ifull` and its corresponding artifact mask `Martifact` are first processed by an encoder to extract latent features. Specifically, the image is processed using ResNet for feature extraction, while the artifact mask is processed using a CNN. These extracted features are then fed into an OEA module inserted before each Cross-Attention step in the UNet diffusion process. The OEA module splits these features into two paths: one for extracting the foreground features of the artifact itself, and the other for identifying its spatial influence region. The contextual information generated by these two paths is then integrated and fed into a cross-attention mechanism to guide the model to focus on the correct cleanup region and reconstruct the background structure.

[0090] For the noise removal model, its training samples consist of three parts: (Ifull, Marifact, Iclean). Ifull represents the original medical image containing interference, such as a CT scan with metallic artifacts; Marifact is a manually or automatically generated artifact mask, which can be binary or a soft map; Iclean is the image after removing the artifacts, which can be a real control or simulated data. Data sources include hospital clinical images (requiring manual mask annotation) or simulated artifact generators (capable of adding motion blur, metallic stripes, mosaic effects, etc.).

[0091] The Loss function consists of three main parts, and the final value is the average of them:

[0092] Reconstruction Loss is used to measure the difference between Ifull and Iclean.

[0093] Attention Loss is used to force the attention distribution of OEA to closely resemble the human mask, ensuring that the model focuses on the correct regions.

[0094] Background fidelity loss is used to emphasize the preservation of structure in areas unaffected by artifacts, ensuring that true details in non-artifact areas are not lost.

[0095] The specific working principle of the OEA module is as follows: After extracting foreground features from the image to be processed and the artifact mask, a dilation operation is performed on the artifact mask to expand the artifact region. Then, the two features are convolved separately and connected together to form a combined feature. Finally, the attention mechanism is used to process the combined feature to extract the influence features of the artifact on the surrounding space.

[0096] The method in this embodiment embeds the OEA module before the cross-attention mechanism of each UNet layer, realizing dual modeling of artifact regions and their diffusion effects; it uses ResNet to extract image semantic features and CNN to extract mask structure features, thus decoupling artifact information from background features; it proposes a method that combines attention supervision with background fidelity loss, which enhances the ability to remove artifacts while ensuring the authenticity of medical structures.

[0097] The method described in this embodiment, by accurately locating the interference region and avoiding information loss due to over-repair, is particularly suitable for interference types with spatial diffusion characteristics, such as CT metal artifacts and MRI motion blur. This system not only maintains high-fidelity reconstruction quality but also demonstrates strong robustness and good generalization ability, and is expected to play an important role in fields such as clinical auxiliary diagnosis and medical image cleaning.

[0098] The method described in this embodiment can be applied to the claims processing workflow of medical insurance companies. When a patient submits a medical image containing artifacts as evidence for claims, the method can automatically remove artifacts from the image and generate a high-quality target image (Igen), helping the insurance company to more accurately assess the condition and determine the claim amount. The method obtains the patient's original medical image (Ifull) and its corresponding artifact mask (Martifact) from the hospital. The artifact mask can be manually annotated by professionals or automatically generated by an algorithm. A noise removal model is used to process the image, generating an artifact-free target image (Igen). The Igen is compared with a standard clean image (Iclean) to confirm that artifacts have been effectively removed and that no important structural information has been lost.

[0099] With high-quality medical images, insurance companies can make more accurate claims decisions, while improving customer satisfaction and service efficiency.

[0100] The method described in this embodiment can be used to optimize the identity verification process of financial institutions. By removing artifacts from document photos, it not only improves the success rate of verification but also enhances system security and reduces the probability of fraud. Furthermore, it is also applicable to improving the quality of scanned financial documents, ensuring that key information is clearly readable and facilitating subsequent data extraction and analysis.

[0101] Collect user-uploaded identity verification document images (Ifull) and their corresponding artifact masks (Martifacts). Artifact masks can be automatically detected and generated using algorithms; apply a noise removal model to process the images, outputting clear, artifact-free photo (Igen); check if the processed images meet recognition requirements, ensuring all necessary details are preserved; use the high-quality processed images for further identity verification or document digitization, improving overall service quality and user experience.

[0102] Suppose a user uploads a photo of their ID card with light reflections when opening an online bank account. The artifacts caused by these reflections can affect the accuracy of identity verification. Financial institutions use the artifact removal system of this invention to process the photo, successfully removing the light reflection artifacts and generating a clear target image (Igen). Based on this high-quality image, the system can accurately identify the user, smoothly complete the account opening process, and also improve the user experience and trust.

[0103] The image noise removal method described above acquires the image to be processed and its corresponding artifact mask as initial data, and inputs this data into a specially trained noise removal model to achieve accurate artifact localization and seamless removal. This model is built based on an improved diffusion model and uses the image with artifacts, the corresponding artifact mask, and the clean target image as training sample sets. It can effectively learn and separate artifact features and their spatial effects, while preserving important structural information and details of the original image and avoiding information loss due to over-repair. With this sophisticated design, the method not only significantly improves the accuracy of artifact removal and the quality of image reconstruction, but also exhibits strong robustness and good generalization ability, making it suitable for artifact removal in various types of medical images and ensuring high-fidelity image output. Finally, the processed target image is output, providing high-quality data support for subsequent applications such as clinical diagnosis.

[0104] Figure 5 This is a schematic block diagram of an image noise removal device 300 provided in an embodiment of the present invention. Figure 5 As shown, corresponding to the above image noise removal method, the present invention also provides an image noise removal apparatus 300. This image noise removal apparatus 300 includes a unit for performing the above image noise removal method, and the apparatus can be configured in a server. Specifically, please refer to... Figure 5 The image noise removal device 300 includes an acquisition unit 301, a removal unit 302, and an output unit 303.

[0105] The acquisition unit 301 is used to acquire the image to be processed and the corresponding artifact mask to obtain initial data; the removal unit 302 is used to input the initial data into the noise removal model for noise removal to obtain the target image; wherein, the noise removal model is obtained by training an improved diffusion model using the image with artifacts, the corresponding artifact mask and the corresponding clean image as a sample set; the output unit 303 is used to output the target image.

[0106] In one embodiment, such as Figure 6As shown, the removal unit 302 includes a first extraction subunit 3021, a second extraction subunit 3022, a processing subunit 3023, and a cross-attention processing subunit 3024.

[0107] The first extraction subunit 3021 is used to input the initial data into the noise removal model and extract the latent features corresponding to the initial data using the encoder; the second extraction subunit 3022 is used to extract the foreground features of artifacts and the influence features of artifacts on the surrounding space from the latent features using the OEA module to obtain intermediate features; the processing subunit 3023 is used to process the intermediate features to obtain context information; the cross-attention processing subunit 3024 is used to process the context information using a cross-attention mechanism to obtain the target image.

[0108] In one embodiment, the first extraction subunit 3021 is used to extract semantic features of the image to be processed using a ResNet network and to obtain structural features of the artifact mask using a CNN network.

[0109] In one embodiment, please refer to Figure 7 The second extraction subunit 3022 includes a foreground feature extraction module 30221, a dilation module 30222, a convolutional connection module 30223, and a feature extraction module 30224.

[0110] The foreground feature extraction module 30224 (30221) is used to extract the foreground features of the artifacts from the latent features using the OEA module; the dilation module 30222 is used to dilate the artifact mask to obtain a dilated artifact mask; the convolution connection module 30223 is used to convolve the foreground features of the artifacts and the latent features corresponding to the dilated artifact mask respectively, and then connect them to obtain combined features; the feature extraction module 30224 is used to process the combined features using an attention mechanism to extract the influence features of the artifacts on the surrounding space to generate intermediate features.

[0111] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the above-mentioned image noise removal device 300 and each unit can be referred to the corresponding description in the foregoing method embodiments. For the sake of convenience and brevity, it will not be repeated here.

[0112] The aforementioned image noise removal device 300 can be implemented as a computer program, which can, for example... Figure 8 It runs on the computer device shown.

[0113] Please see Figure 8 , Figure 8This is a schematic block diagram of a computer device provided in an embodiment of this application. The computer device 500 can be a server, wherein the server can be a standalone server or a server cluster composed of multiple servers.

[0114] See Figure 8 The computer device 500 includes a processor 502, a memory, and a network interface 505 connected via a system bus 501. The memory may include a non-volatile storage medium 503 and internal memory 504.

[0115] The non-volatile storage medium 503 may store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions that, when executed, cause the processor 502 to perform an image noise removal method.

[0116] The processor 502 provides computing and control capabilities to support the operation of the entire computer device 500.

[0117] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can perform an image noise removal method.

[0118] This network interface 505 is used for network communication with other devices. Those skilled in the art will understand that... Figure 8 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device 500 to which the present application is applied. The specific computer device 500 may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0119] The processor 502 is used to run a computer program 5032 stored in the memory to perform the following steps:

[0120] Obtain the image to be processed and the corresponding artifact mask to obtain initial data; input the initial data into the noise removal model for noise removal to obtain the target image; wherein, the noise removal model is obtained by training an improved diffusion model using the image with artifacts, the corresponding artifact mask, and the corresponding clean image as a sample set; output the target image.

[0121] The improved diffusion model includes an encoder, an OEA module, and a cross-attention mechanism.

[0122] The training process of the noise removal model uses a loss function that is the average of the reconstruction loss, attention loss, and background fidelity loss.

[0123] In one embodiment, when the processor 502 implements the step of inputting the initial data into the noise removal model for noise removal to obtain the target image, the following steps are specifically implemented:

[0124] The initial data is input into the noise removal model, and the encoder is used to extract the latent features corresponding to the initial data. The OEA module is used to extract the foreground features of artifacts and the influence features of artifacts on the surrounding space from the latent features to obtain intermediate features. The intermediate features are processed to obtain context information. The context information is processed using a cross-attention mechanism to obtain the target image.

[0125] The latent features include the semantic features of the image to be processed and the structural features of the artifact mask.

[0126] In one embodiment, when implementing the step of extracting the latent features corresponding to the initial data using the encoder, the processor 502 specifically implements the following steps:

[0127] The semantic features of the image to be processed are extracted using a ResNet network, and the structural features of the artifact mask are obtained using a CNN network.

[0128] In one embodiment, when the processor 502 implements the step of extracting foreground features of artifacts and the influence features of artifacts on the surrounding space using the OEA module to obtain intermediate features, the specific steps are as follows:

[0129] The OEA module is used to extract the foreground features of the artifacts from the latent features; the artifact mask is dilated to obtain a dilated artifact mask; the foreground features of the artifacts and the latent features corresponding to the dilated artifact mask are convolved and then concatenated to obtain combined features; the combined features are processed using an attention mechanism to extract the influence features of the artifacts on the surrounding space to generate intermediate features.

[0130] It should be understood that in the embodiments of this application, the processor 502 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0131] It will be understood by those skilled in the art that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program includes program instructions and can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.

[0132] Therefore, the present invention also provides a storage medium. This storage medium can be a computer-readable storage medium. The storage medium stores a computer program, wherein when executed by a processor, the computer program causes the processor to perform the following steps:

[0133] Obtain the image to be processed and the corresponding artifact mask to obtain initial data; input the initial data into the noise removal model for noise removal to obtain the target image; wherein, the noise removal model is obtained by training an improved diffusion model using the image with artifacts, the corresponding artifact mask, and the corresponding clean image as a sample set; output the target image.

[0134] The improved diffusion model includes an encoder, an OEA module, and a cross-attention mechanism.

[0135] The training process of the noise removal model uses a loss function that is the average of the reconstruction loss, attention loss, and background fidelity loss.

[0136] In one embodiment, when the processor executes the computer program to implement the step of inputting the initial data into a noise removal model for noise removal to obtain the target image, it specifically implements the following steps:

[0137] The initial data is input into the noise removal model, and the encoder is used to extract the latent features corresponding to the initial data. The OEA module is used to extract the foreground features of artifacts and the influence features of artifacts on the surrounding space from the latent features to obtain intermediate features. The intermediate features are processed to obtain context information. The context information is processed using a cross-attention mechanism to obtain the target image.

[0138] The latent features include the semantic features of the image to be processed and the structural features of the artifact mask.

[0139] In one embodiment, when the processor executes the computer program to implement the step of extracting the latent features corresponding to the initial data using the encoder, it specifically implements the following steps:

[0140] The semantic features of the image to be processed are extracted using a ResNet network, and the structural features of the artifact mask are obtained using a CNN network.

[0141] In one embodiment, when the processor executes the computer program to implement the step of extracting foreground features of artifacts and the influence features of artifacts on the surrounding space using the OEA module to obtain intermediate features, the processor specifically implements the following steps:

[0142] The OEA module is used to extract the foreground features of the artifacts from the latent features; the artifact mask is dilated to obtain a dilated artifact mask; the foreground features of the artifacts and the latent features corresponding to the dilated artifact mask are convolved and then concatenated to obtain combined features; the combined features are processed using an attention mechanism to extract the influence features of the artifacts on the surrounding space to generate intermediate features.

[0143] The storage medium can be any computer-readable storage medium capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), magnetic disk, or optical disk.

[0144] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0145] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of each unit is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.

[0146] The steps in the method of this invention can be adjusted, merged, or reduced in order according to actual needs. The units in the device of this invention can be merged, divided, or reduced according to actual needs. Furthermore, the functional units in the various embodiments of this invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0147] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.

[0148] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. An image noise removal method, characterized in that, include: Obtain the image to be processed and the corresponding artifact mask to obtain the initial data; The initial data is input into a noise removal model for noise removal to obtain the target image; wherein the noise removal model is obtained by training an improved diffusion model using an image with artifacts, the corresponding artifact mask, and the corresponding clean image as a sample set; Output the target image.

2. The image noise removal method according to claim 1, characterized in that, The improved diffusion model includes an encoder, an OEA module, and a cross-attention mechanism.

3. The image noise removal method according to claim 2, characterized in that, The step of inputting the initial data into a noise removal model for noise removal to obtain the target image includes: The initial data is input into the noise removal model, and the encoder is used to extract the latent features corresponding to the initial data. The OEA module is used to extract the foreground features of artifacts and the influence features of artifacts on the surrounding space from the latent features to obtain intermediate features; The intermediate features are processed to obtain contextual information; The target image is obtained by using a cross-attention mechanism with contextual information.

4. The image noise removal method according to claim 3, characterized in that, The latent features include the semantic features of the image to be processed and the structural features of the artifact mask.

5. The image noise removal method according to claim 3, characterized in that, The step of extracting latent features corresponding to the initial data using the encoder includes: The semantic features of the image to be processed are extracted using a ResNet network, and the structural features of the artifact mask are obtained using a CNN network.

6. The image noise removal method according to claim 3, characterized in that, The method of using the OEA module to extract foreground features of artifacts and the influence features of artifacts on the surrounding space from the latent features to obtain intermediate features includes: The OEA module is used to extract the foreground features of artifacts from the latent features; The artifact mask is dilated to obtain an expanded artifact mask; The foreground features of the artifact and the latent features corresponding to the dilated artifact mask are convolved and then concatenated to obtain combined features; The combined features are processed using an attention mechanism to extract the influence features of artifacts on the surrounding space, in order to generate intermediate features.

7. The image noise removal method according to claim 2, characterized in that, The training process of the noise removal model uses a loss function that is the average of the reconstruction loss, attention loss, and background fidelity loss.

8. An image noise removal device, characterized in that, include: The acquisition unit is used to acquire the image to be processed and the corresponding artifact mask to obtain initial data; The noise removal unit is used to input the initial data into the noise removal model for noise removal to obtain the target image; wherein the noise removal model is obtained by training an improved diffusion model using an image with artifacts, the corresponding artifact mask, and the corresponding clean image as a sample set; The output unit is used to output the target image.

9. A computer device, characterized in that, The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method as described in any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium stores a computer program that, when executed by a processor, implements the method as described in any one of claims 1 to 7.