Image restoration method and device, electronic equipment and storage medium

By combining the diffusion model with the SMAS and SMAD algorithms to obtain global perceptual information and salient features of images, the shortcomings of diffusion-based models in image generation in complex scenes are solved, the detail fidelity and stability of image restoration are improved, and the adaptability of the model is enhanced.

CN121961935APending Publication Date: 2026-05-01UNIV OF MACAU
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
UNIV OF MACAU
Filing Date
2026-01-29
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Self-supervised learning methods based on diffusion models suffer from insufficient detail fidelity and stability in generating images when dealing with complex scenes, and have low adaptability to different scenarios.

Method used

By using a diffusion model combined with image region similarity search mechanism (SMAS algorithm) and image region difference search mechanism (SMAD algorithm), global perception information, regional similarity relationships and salient features are obtained during the image restoration process, which assists in image restoration and improves the flexibility and adaptability of the model.

Benefits of technology

The method improves the detail fidelity and stability of image generation in complex scenes by improving the self-supervised learning method based on diffusion model, and enhances its adaptability to different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121961935A_ABST
    Figure CN121961935A_ABST
Patent Text Reader

Abstract

The invention provides an image restoration method and device, electronic equipment and a storage medium, and relates to the technical field of computer vision. The method comprises the following steps: adding a mask to a to-be-restored image through a preset mask algorithm to obtain a mask image for inputting a trained neural network model, and performing noise processing on the mask image according to a preset time step; and according to the trained neural network model, obtaining global perception information, a region similarity relationship and significant features of the mask image in noise processing so as to obtain the mask image after noise processing, namely a repaired image. According to the method, through a neural network model comprising a diffusion model, an SMAS algorithm and an SMAD algorithm, global perception information, a region similarity relationship and significant features are acquired and utilized in an image restoration process, so that the detail fidelity and stability of a generated image are improved when a self-supervised learning method based on the diffusion model processes a complex scene, and the image restoration efficiency is improved. And the adaptability to different scenes is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Image restoration methods, devices, electronic equipment and storage media Technical Field

[0001] This application relates to the field of computer vision technology, and in particular to an image restoration method, apparatus, electronic device and storage medium. Background Technology

[0002] Image inpainting is a key technology in computer vision, aiming to intelligently infer and fill in missing or damaged areas based on pixel information of known regions in an image, in order to generate a visually coherent and semantically reasonable complete image. This technology has been widely applied in many important fields such as digital cultural heritage restoration, old photograph restoration, image editing, medical image analysis, and autonomous driving scene understanding.

[0003] Current image inpainting solutions are increasingly replacing deep learning-based inpainting methods and traditional inpainting methods that rely on local pixel diffusion or sample matching logic with self-supervised learning methods based on diffusion models. Self-supervised learning methods based on diffusion models learn the probability distribution of data as it moves from ordered to disordered (noise addition) and then from disordered to ordered (denoising), thus restoring image content without the need for registration references.

[0004] However, although self-supervised learning methods based on diffusion models are more suitable for open real-world environments, the diffusion inpainting models currently used in this method do not consider the inherent internal correlations of image content, such as texture similarity and structural differences between different regions. This results in insufficient detail fidelity and stability of the generated images when dealing with complex scenes, and low adaptability to different scenarios. Summary of the Invention

[0005] The main objective of this application is to propose an image inpainting method, apparatus, electronic device, and storage medium, which aims to improve the detail fidelity and stability of the generated images when the self-supervised learning method based on the diffusion model is used to process complex scenes, and to enhance its adaptability to different scenes.

[0006] In a first aspect, the present invention provides an image restoration method, comprising: acquiring an image to be restored; adding a mask to the image to be restored using a preset masking algorithm to obtain a masked image; inputting the masked image into a trained neural network model and performing noise processing on the masked image according to a preset time step; acquiring global perception information, regional similarity relationships, and salient features of the masked image during noise processing based on the trained neural network model; wherein the trained neural network model includes: a diffusion model, an image region similarity search mechanism SMAS algorithm, and an image region difference search mechanism SMAD algorithm; and acquiring the noise-processed masked image based on the global perception information, regional similarity relationships, and salient features, wherein the noise-processed masked image is the restored image.

[0007] In an optional implementation, before inputting the masked image into the trained neural network model and performing noise processing on the masked image according to a preset time step, the method further includes: acquiring multiple complete images; adding a mask to each of the complete images using the preset masking algorithm to obtain multiple training masked images; performing unsupervised training on the initial neural network model using the training masked images to obtain a pre-trained neural network model; and optimizing the pre-trained neural network model according to a preset optimization algorithm to obtain the trained neural network model.

[0008] In an optional implementation, the diffusion model includes a feedforward network and a feedback network; the preset time step includes a first preset time step and a second preset time step; the noise processing includes noise addition and noise removal; the step of inputting the masking image into the trained neural network model and performing noise processing on the masking image according to the preset time step includes: adding noise to the masking image through the feedforward network at the first preset time step to obtain a noisy masking image; and removing noise from the noisy masking image through the feedback network at the second preset time step to obtain a denoised masking image, wherein the denoised masking image is the noise-processed masking image.

[0009] In an optional implementation, the step of obtaining the global perception information and regional similarity relationship of the masked image in noise processing according to the trained neural network model includes: calculating and obtaining the global perception information and regional similarity relationship of the masked image in denoising processing at a second preset time step according to the SMAS algorithm.

[0010] In an optional implementation, obtaining the salient features of the masked image in noise processing based on the trained neural network model includes: calculating and obtaining the salient features of the masked image in denoising processing at a second preset time step according to the SMAD algorithm; obtaining the masked image after noise processing based on the global perception information, region similarity relationship, and salient features includes: when denoising the masked image after noise addition at the second preset time step through the feedback network, obtaining the masked image after denoising processing corresponding to each second preset time step according to the global perception information, region similarity relationship, and salient features corresponding to each second preset time step, wherein the masked image after denoising processing corresponding to the last second preset time step is the masked image after denoising processing.

[0011] In an optional implementation, the step of calculating and obtaining the global perception information and the region similarity relationship of the masked image in the denoising process at a second preset time step according to the SMAS algorithm includes: calculating and obtaining the region similarity relationship of the masked image in the denoising process at a second preset time step according to the following SMAS algorithm formula:

[0012] in, The region similarity relationship, and The indices within the mask image during the denoising process are respectively: The position and index are The pixel value at the location, Indicates the and stated Feature relation values, For the preset normalization function, for The corresponding vector mapping result.

[0013] In an optional implementation, the step of calculating and obtaining the salient features of the masked image in the denoising process at a second preset time step according to the SMAD algorithm includes: calculating and obtaining the salient features of the masked image in the denoising process at a second preset time step according to the following SMAD algorithm formula:

[0014] in, For the aforementioned salient features, To pre-define the static basic convolution kernel, Indicates a dynamically adjusted item. For weight generation function, The result is the affinity feature mapping of the location. This indicates that a convolution operation is to be performed.

[0015] Secondly, the present invention provides an image restoration apparatus, comprising: an acquisition module for acquiring an image to be restored and adding a mask to the image to be restored using a preset masking algorithm to obtain a masked image; a processing module for inputting the masked image into a trained neural network model and performing noise processing on the masked image according to a preset time step; an extraction module for acquiring global perception information, regional similarity relationships, and salient features of the masked image in the noise processing according to the trained neural network model; wherein the trained neural network model includes: a diffusion model, an image region similarity search mechanism SMAS algorithm, and an image region difference search mechanism SMAD algorithm; and a restoration module for acquiring the noise-processed masked image according to the global perception information, regional similarity relationships, and salient features, wherein the noise-processed masked image is the restored image.

[0016] Thirdly, the present invention provides an electronic device, comprising: a processor, a storage medium, and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the steps of the method described in the first aspect above.

[0017] Fourthly, the present invention provides a computer-readable storage medium on which a computer program is stored, the computer program being executed by a processor to perform the steps of the method described in the first aspect above.

[0018] The beneficial effects of this application are as follows: The image restoration method provided in this application, after obtaining the image to be restored, adds a mask to the damaged part of the image to be restored through a preset masking algorithm to obtain the mask image corresponding to the image to be restored, and inputs the mask image into a trained neural network model. The diffusion model in the neural network model performs noise processing on the mask image according to a preset time step. During the noise processing, the global perception information, regional similarity relationship and salient features of the mask image are also obtained according to the SMAS algorithm and SMAD algorithm in the neural network model according to a preset time step to assist in the restoration of the mask image. Finally, the noise-processed mask image is obtained, that is, the restored image. In this embodiment, a neural network model including a diffusion model, SMAS algorithm, and SMAD algorithm is used to simultaneously process noise in the mask image corresponding to the image to be repaired during the image restoration process, while acquiring and utilizing global perception information, regional similarity relationships, and salient features to assist in image restoration. The SMAS algorithm mainly utilizes the global perception information and regional similarity relationships of the image to guide the model to fill in more refined image details; the SMAD algorithm dynamically extracts salient features by learning the differences between different pixel regions to adapt to different scenarios and improve the flexibility of the model. This improves the detail fidelity and stability of the generated image by the self-supervised learning method based on the diffusion model when dealing with complex scenes, and enhances its adaptability to different scenarios. Attached Figure Description

[0019] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 is a schematic flowchart of an image restoration method according to an embodiment of this application; Figure 2 is a schematic flowchart of an image restoration method according to another embodiment of this application; Figure 3 is a schematic flowchart of an image restoration method according to yet another embodiment of this application; Figure 4 is a schematic structural diagram of an image restoration device according to an embodiment of this application; Figure 5 is a schematic structural diagram of an electronic device according to an embodiment of this application. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0022] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0023] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0024] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0025] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0026] Traditional restoration methods, which rely on local pixel diffusion or sample matching logic, struggle to effectively reconstruct high-level semantic information when faced with large-area or complex structural defects. The restoration results often exhibit blurring, artifacts, or structural discontinuities. Deep learning-based restoration methods, trained within a supervised learning framework, rely heavily on paired datasets containing both damaged and intact reference images for performance. In many real-world applications, such as historical image restoration and forensic evidence analysis, obtaining such paired reference images is extremely difficult or even impossible.

[0027] Therefore, current image inpainting solutions are increasingly replacing deep learning-based inpainting methods and traditional methods that rely on local pixel diffusion or sample matching logic with self-supervised learning methods based on diffusion models. These diffusion-based methods learn the probability distribution of data as it moves from ordered to disordered (noise addition) and then back from disordered to ordered (denoising) to restore image content without the need for registration references. However, while diffusion-based methods are more suitable for open real-world environments, the diffusion inpainting models currently used in these methods rarely consider the inherent internal relationships within image content, such as texture similarity and structural differences between different regions. This results in insufficient detail fidelity and stability of the generated images when handling complex scenes, and low adaptability to different scenarios.

[0028] To address the aforementioned issues, the main objective of this application is to propose an image restoration method that aims to improve the detail fidelity and stability of the generated image when the self-supervised learning method based on the diffusion model is used in complex scenarios, such as the aforementioned historical image restoration and judicial evidence analysis, or scenarios with many objects or color elements in the image and a large damaged area, and to enhance its adaptability to different scenarios.

[0029] Figure 1 is a schematic flowchart of an image restoration method provided in an embodiment of this application. The execution subject of this method can be a computer, server or other device with computing power, but is not limited thereto. Referring to Figure 1, the method includes: S101, obtaining the image to be restored, adding a mask to the image to be restored through a preset masking algorithm to obtain a masked image.

[0030] For example, the aforementioned image to be repaired may refer to an image that is partially dirty, blurry, or damaged. The sources of these images to be repaired may be old photos, unclear photos, or partially damaged photos in scenarios such as historical image restoration, forensic evidence analysis, medical image analysis, and intelligent driving environment analysis. They may also be electronic images with partially damaged or lost data due to encryption, hard drive damage, etc. The specific examples above are limited. The aforementioned preset masking algorithm may be an algorithm that can analyze and identify the dirty, blurry, or damaged areas in the aforementioned image to be repaired, and add a mask to the dirty, blurry, or damaged areas in the aforementioned image to be repaired. That is, the aforementioned masked image is the image after adding a mask to the dirty, blurry, or damaged areas in the aforementioned image to be repaired. The aforementioned mask may refer to a pure black area with the same shape and size as the dirty, blurry, or damaged areas in the aforementioned image to be repaired, but it is not limited to this.

[0031] S102. Input the masked image into the trained neural network model and perform noise processing on the masked image according to a preset time step.

[0032] For example, the trained neural network model mentioned above can be a neural network model trained using an unsupervised learning method. In addition to the traditional diffusion model, the neural network model can also include one or more other algorithms, such as feature extraction algorithms, similarity calculation algorithms, difference calculation algorithms, etc. However, the specific composition of the neural network model and the specific types and number of algorithms are not limited here.

[0033] The aforementioned preset time step can refer to multiple time segments of a fixed preset duration, such as 1000 segments, which can be called 1000 preset time steps. However, it is understood that the aforementioned noise processing can include noise addition processing and noise reduction processing, etc. For different noise processing stages, there can be different numbers and fixed preset time steps. For example, noise addition processing corresponds to 1000 preset time steps, and noise reduction processing corresponds to 250 preset time steps. The total duration of noise addition processing can be the same as or different from that of noise reduction processing. Of course, the above is only a possible example. The actual time segment duration corresponding to the preset time step, the actual number of time steps in each noise processing stage, etc. can be adjusted and determined according to the actual situation, and are not limited here.

[0034] S103. Based on the trained neural network model, obtain the global perception information, regional similarity relationship, and salient features of the masked image in noise processing.

[0035] The trained neural network models include: a diffusion model, the SMAS algorithm (image region similarity search mechanism), and the SMAD algorithm (image region difference search mechanism).

[0036] For example, the aforementioned diffusion model could be a probabilistic generative model whose core idea is to learn the data distribution through a progressive "forward noise addition" and "backward denoising" process. In the forward process, the model progressively adds Gaussian noise to the original data until the data becomes entirely random noise; in the backward process, the model learns to progressively denoise from the noise to reconstruct the original data. This method does not require paired training data, is suitable for unsupervised learning scenarios, and demonstrates strong creativity and fidelity in image generation and restoration tasks. Specifically, the diffusion model could be, for example, a denoising diffusion probabilistic model, or a continuous-time diffusion model defined by differential equations, but is not limited to these.

[0037] The aforementioned SMAS (Searching Mechanism of Image Area Similarities) algorithm can be an example of an algorithm designed to capture long-range dependencies and self-similarity within an image. Its core principle is "non-local operation," which calculates the feature similarity between any location in the feature map and all other locations, and injects global contextual information into each local location through weighted aggregation and other methods. This helps the trained neural network model reference semantically or texturally similar distant regions in the image to generate or repair the content of the current region, thereby ensuring the structural coherence and visual plausibility of the repaired result.

[0038] The aforementioned SMAD (Searching Mechanism of Image Area Differences) algorithm can be, for example, a dynamic feature modulation algorithm designed to enhance the adaptability of a trained neural network model to different image content. Its core idea is to dynamically generate or adjust the weights of the convolution kernels based on the features of the input image content itself, rather than using fixed parameters. Specifically, this mechanism generates an "affinity map" related to the input content by analyzing the differences (such as semantics and texture) between different regions within the feature map, and dynamically adjusts the parameters of the convolution operation accordingly. This allows the model to employ differentiated processing strategies for different objects or regions in the image, thereby extracting and reconstructing salient features more accurately.

[0039] The acquisition of global perceptual information and regional similarity relationships of the masked image in noise processing can be achieved, for example, through the SMAS algorithm described above, while the acquisition of salient features of the masked image in noise processing can be achieved, for example, through the SMAD algorithm described above.

[0040] Specifically, the aforementioned global perception information, regional similarity relationships, and salient features can be extracted from the same or different features of the masked image in noise processing, and calculated based on these features (or the vector values ​​converted from these features) and the aforementioned SMAS and SMAD algorithms. However, the specific types, quantities, and calculation formulas of features can be adjusted and determined according to the actual situation, and are not limited here.

[0041] The aforementioned globally perceived information can refer, for example, to contextual knowledge about the overall structure and semantic relationships of an image, calculated using the SMAS algorithm. Specifically, it can take the form of a dynamic attention weight distribution that encodes the dependencies between any location in the image and all other locations, ensuring that the newly generated content maintains global consistency with the known parts of the image in terms of structure, texture, and semantics.

[0042] The aforementioned region similarity relationship can refer, for example, to the degree of similarity between two different regions within an image in the feature space. This relationship is quantified by calculating nonlocal similarity. This region similarity relationship enables a trained neural network model to discover and utilize repetitive patterns, symmetrical structures, or similar textures present in the image, thereby guiding the filling of details in missing regions.

[0043] The aforementioned salient features can refer to key visual attributes in a masked image that define and distinguish different objects or regions, such as object edges, contours, unique textures, or high-contrast areas. The core task of the SMAD algorithm is to dynamically extract these features. These salient features enable the trained neural network model to adaptively focus on and enhance the core discriminative information of different regions, thereby avoiding the generation of blurry or homogeneous content when processing complex scenes and improving the clarity and discriminative power of the restoration results.

[0044] S104. Based on the above global perception information, regional similarity relationship, and salient features, obtain the noise-processed mask image. The noise-processed mask image is the repaired image.

[0045] The above-mentioned method of obtaining a noise-processed masked image based on global perception information, regional similarity relationships, and salient features can refer to combining the above-mentioned global perception information, regional similarity relationships, and salient features to perform noise processing on the masked image in order to improve the accuracy, reliability, and applicability of the noise processing. However, the specific method of combining the above-mentioned global perception information, regional similarity relationships, and salient features to perform noise processing on the masked image can be selected and determined according to the actual situation, and is not limited here.

[0046] Based on the above, the trained neural network model can be represented as follows:

[0047] Among them, the above The masked image after noise processing, i.e., the repaired image, is directly output by the trained neural network model. The mask image is used as input to the trained neural network model. This indicates that noise is added to the masked image. The function used to represent the above diffusion model, the above The function used to represent the SMAS algorithm described above, the above The function used to represent the SMAD algorithm described above, the above This represents a preset fusion operation, which is the masked image obtained after noise processing through the trained neural network model. For example, it can be an image that has been repaired by simultaneously obtaining global perception information, regional similarity relationships, and salient features based on the noise processing of the diffusion model and the global perception information, regional similarity relationships, and salient features obtained by the SMAS and SMAD algorithms.

[0048] However, it is understandable that the above This is merely an illustration; the actual forms of the functions corresponding to the diffusion model, the SMAS algorithm, and the SMAD algorithm are not limited to those described above. Instead, it can be determined based on the actual situation.

[0049] The image restoration method provided in this application, after acquiring the image to be restored, adds a mask to the damaged parts of the image to be restored using a preset masking algorithm to obtain a mask image corresponding to the image to be restored. This mask image is then input into a trained neural network model. The diffusion model in the neural network model performs noise processing on the mask image according to a preset time step. During the noise processing, the SMAS and SMAD algorithms in the neural network model are used to acquire global perception information, regional similarity relationships, and salient features of the mask image according to a preset time step to assist in restoring the mask image. Finally, the noise-processed mask image is obtained, which is the restored image. In this embodiment, through a neural network model including a diffusion model, the SMAS algorithm, and the SMAD algorithm, noise processing is performed on the mask image corresponding to the image to be restored while simultaneously acquiring and utilizing global perception information, regional similarity relationships, and salient features to assist in image restoration. The SMAS algorithm mainly utilizes the global perception information and regional similarity relationships of the image to guide the model to fill in finer image details. The SMAD algorithm dynamically extracts salient features by learning the differences between different pixel regions to adapt to different scenarios and improve the flexibility of the model. This improves the detail fidelity and stability of the generated images when dealing with complex scenes, and enhances the adaptability to different scenarios.

[0050] Figure 2 is a schematic flowchart of an image restoration method provided in another embodiment of this application. As shown in Figure 2, optionally, before inputting the masked image into the trained neural network model and performing noise processing on the masked image according to a preset time step, the method further includes: S201, acquiring multiple complete images, and adding a mask to each of the complete images respectively through the preset masking algorithm to obtain multiple training masked images.

[0051] For example, the above-mentioned masking algorithm is used to add a mask to each of the above complete images. This may differ from the preset masking algorithm used for the image to be repaired. This is because the image to be repaired is an image that already has some dirt, blur, or damage, while the complete image is an image that does not have some dirt, blur, or damage. Therefore, when the preset masking algorithm is used to add a mask to each of the above complete images, it is not necessary to analyze and identify the areas of dirt, blur, or damage. Instead, regular or irregular masks of different sizes and shapes can be randomly generated according to preset parameters to simulate the areas of dirt, blur, or damage in the image to be repaired in actual situations, and to be used for the initial image repair training of the neural network model.

[0052] S202. The initial neural network model is trained unsupervised using the above training mask images to obtain a preliminarily trained neural network model.

[0053] S203. Optimize the above-preliminary trained neural network model according to the preset optimization algorithm to obtain the above-preliminary trained neural network model.

[0054] For example, before optimizing the pre-trained neural network model according to the preset optimization algorithm, the output of the pre-trained neural network model can be evaluated by a preset loss function. This preset loss function can be a hybrid loss function, which can be expressed in the following form:

[0055] Among them, the above The mixture loss is used to represent the degree of error in the output of the initially trained neural network model. To rebuild the losses, the above For divergence loss, the above These are the preset loss parameters.

[0056] Among them, the aforementioned reconstruction losses For example, it can be calculated using the following formula:

[0057] Among them, the above This is real noise, which can be directly measured and obtained. This represents the predicted noise, which is a preset value.

[0058] The aforementioned divergence loss For example, it can be calculated using the following formula:

[0059] Among them, the above This represents the predicted noise distribution, which is a preset value. This represents the true noise distribution and can be directly measured. express The noise-damaged image corresponding to the time point, as described above This represents the original image.

[0060] The aforementioned optimization of the pre-trained neural network model according to a preset optimization algorithm can, for example, refer to optimizing the parameters in the pre-trained neural network model using a preset optimizer. This optimizer could be, for example, the Adam (Adaptive Moment Estimation) optimizer, and the optimization objective could be, for example, finding the mixture loss of the pre-trained neural network model. Model parameters that are less than the preset loss threshold can also be optimized model parameters obtained directly after a fixed number of optimization iterations, such as after 50,000 optimization iterations. The initial learning rate of the Adam optimizer can be set to 2.0e-4, for example, but it can be adjusted and determined according to the actual situation and is not limited to 2.0e-4.

[0061] Figure 3 is a schematic flowchart of an image restoration method according to another embodiment of this application. Referring to Figure 3, optionally, based on the embodiment in Figure 1, the diffusion model includes a feedforward network and a feedback network. The preset time step includes a first preset time step and a second preset time step. The noise processing includes noise addition processing and noise reduction processing.

[0062] The above-mentioned masking image is input into the trained neural network model, and noise processing is performed on the masking image according to a preset time step, including: S301, the above-mentioned feedforward network is used to add noise to the masking image at a first preset time step to obtain a noisy masking image.

[0063] S302. The above-mentioned noise-added masking image is denoised through the above-mentioned feedback network at a second preset time step to obtain a denoised masking image. The denoised masking image is the masking image after noise processing.

[0064] For example, the first preset time step and the second preset time step mentioned above may be preset time steps with the same or different numbers and fixed preset durations as indicated in the embodiment of Figure 1. For example, the fixed preset duration of the first preset time step and the fixed preset duration of the second preset time step may be 1:4, etc., but it is not limited to this.

[0065] Optionally, based on the embodiment in Figure 3 above, the above-mentioned acquisition of global perception information and regional similarity relationship of the masked image in noise processing according to the trained neural network model includes: calculating and acquiring the above-mentioned global perception information and regional similarity relationship of the masked image in denoising processing at a second preset time step according to the above-mentioned SMAS algorithm.

[0066] For example, the SMAS algorithm described above can be deployed in each downsampling module of the feedback network (the feedback network may contain three downsampling blocks, and the SMAS algorithm can be specifically set between the second and third downsampling blocks). This SMAS algorithm can, for example, calculate non-local similarity by comparing the features at the current location with the features of all pixels in the image, thereby capturing the internal structure and texture relationships of the image to guide the restoration process. Of course, the above is only one possible example; the specific deployment location and function of the SMAS algorithm can be adjusted and determined according to the actual situation and are not limited to the examples described above.

[0067] The above-mentioned calculation of the global perception information and regional similarity relationship of the masked image in the denoising process according to the SMAS algorithm at the second preset time step can, for example, mean that in the process of denoising the masked image after adding noise at the second preset time step, a denoising process is performed once at each time step. At the same time as the denoising process at each time step, the global perception information and regional similarity relationship of the masked image in the denoising process at the current time step are calculated and obtained according to the SMAS algorithm, so as to combine the global perception information and regional similarity relationship to complete the denoising process at the current time step.

[0068] Furthermore, based on the above embodiments, the acquisition of salient features of the masked image in noise processing according to the trained neural network model includes: calculating and acquiring the salient features of the masked image in denoising processing at a second preset time step according to the above SMAD algorithm.

[0069] For example, the SMAD algorithm described above can be deployed symmetrically between the downsampling and upsampling paths of the feedback network (specifically, between the output of the second downsampling layer and the input of the third upsampling layer, and between the output of the fifth downsampling layer and the input of the first upsampling layer). This SMAD algorithm can, for example, utilize affinity maps to calculate the differences between pixels in different regions and dynamically update the weights of the convolutional kernels accordingly. This allows the neural network model to dynamically learn significant common information from different input images, improving its adaptability. Of course, the above is merely a possible example; the specific deployment location and function of the SMAD algorithm can be adjusted and determined according to actual circumstances and are not limited to the examples described above.

[0070] Similar to the global perception information and region similarity relationship of the masked image in the denoising process calculated and obtained at the second preset time step according to the SMAS algorithm, the salient features of the masked image in the denoising process calculated and obtained at the second preset time step according to the SMAD algorithm can, for example, mean that in the process of denoising the masked image after adding noise at the second preset time step, a denoising process is performed at each time step. At the same time as the denoising process at each time step, the salient features of the masked image in the denoising process at the current time step are calculated and obtained according to the SMAS algorithm, so as to complete the denoising process at the current time step by combining the salient features.

[0071] The above-mentioned acquisition of the noise-processed mask image based on global perception information, regional similarity relationship, and salient features includes: when the above-mentioned noise-added mask image is denoised through the above-mentioned feedback network at the above-mentioned second preset time step, the denoised mask image corresponding to each above-mentioned second preset time step is acquired based on the above-mentioned global perception information, the above-mentioned regional similarity relationship, and the above-mentioned salient features corresponding to each above-mentioned second preset time step, wherein the denoised mask image corresponding to the last above-mentioned second preset time step is the denoised mask image.

[0072] For example, since the above denoising process is completed step by step according to the second preset time step, after the denoising process of each second preset time step is completed, a denoised mask image corresponding to the second preset time step will be obtained. For the denoised mask image corresponding to the second preset time step, the denoising process corresponding to the next second preset time step will continue until the last second preset time step is completed and the denoising process ends. The mask image will theoretically be completely denoised. At this time, the denoised mask image is obtained, which is also the image that has been repaired.

[0073] Optionally, based on the foregoing embodiments, the above-mentioned calculation of the global perception information and the region similarity relationship of the masked image in the denoising process using the SMAS algorithm at a second preset time step includes: calculating the region similarity relationship of the masked image in the denoising process using the following SMAS algorithm formula at a second preset time step:

[0074] in, The above-mentioned regions have similarity relationships. and The indices of the masked images in the above denoising process are respectively: The position and index are The pixel value at the position, the index is The position and index are The pixel value at a given location can be represented as a multi-dimensional vector, such as (R, G, B), but is not limited to this. The above... The above indicates and the above The characteristic relationship value, for example, can be based on the feature relationship value, for example, according to and The similarity is calculated using a combination of dot product similarity algorithms, Gaussian similarity algorithms, etc. Specific methods are not limited here. This indicates the preset normalization function. express The corresponding vector mapping result.

[0075] Furthermore, based on the aforementioned embodiments, the above-mentioned significant features of the masked image in the denoising process are calculated and obtained at a second preset time step according to the SMAD algorithm, including: calculating and obtaining the significant features of the masked image in the denoising process at a second preset time step according to the following SMAD algorithm formula:

[0076] in, For the above-mentioned significant features, This is a preset static base convolution kernel, a preset value that can be adjusted and determined according to actual conditions. Indicates a dynamically adjusted item. For weight generation function, The result is the affinity feature mapping of the location. This indicates that a convolution operation is to be performed.

[0077] Based on the above embodiments, this solution leverages the unsupervised learning capabilities of the aggregation diffusion model, combined with the SMAS and SMAD algorithms, to effectively handle masks of different scales, from narrow to wide. This enables the repair of large-area defects and complex structural regions. Specifically, by integrating global similarity guidance with local dynamic feature modulation, this solution improves the structural coherence, texture fidelity, and visual naturalness of the generated repair results. It also avoids the dependence of traditional methods on paired training data, enhancing the model's generalization ability in real open scenes.

[0078] Figure 4 is a schematic diagram of an image restoration device provided in an embodiment of this application. The image restoration device can perform the above-described image restoration method. The device can be integrated into the above-described computer, server or other devices with computing power. As shown in Figure 4, the device may include: an acquisition module 410, used to acquire the image to be restored, and add a mask to the image to be restored through a preset masking algorithm to obtain a masked image.

[0079] The processing module 420 is used to input the masked image into the trained neural network model and perform noise processing on the masked image according to a preset time step.

[0080] The extraction module 430 is used to obtain global perception information, region similarity relationships, and salient features of the masked image in noise processing based on the trained neural network model mentioned above. The trained neural network model includes: a diffusion model, the SMAS algorithm (image region similarity search mechanism), and the SMAD algorithm (image region difference search mechanism).

[0081] The repair module 440 is used to obtain a noise-processed mask image based on the aforementioned global perception information, regional similarity relationship, and salient features. The noise-processed mask image is the repaired image.

[0082] The image restoration method provided in this application, after acquiring the image to be restored, adds a mask to the damaged parts of the image to be restored using a preset masking algorithm to obtain a mask image corresponding to the image to be restored. This mask image is then input into a trained neural network model. The diffusion model in the neural network model performs noise processing on the mask image according to a preset time step. During the noise processing, the SMAS and SMAD algorithms in the neural network model are used to acquire global perception information, regional similarity relationships, and salient features of the mask image according to a preset time step to assist in restoring the mask image. Finally, the noise-processed mask image is obtained, which is the restored image. In this embodiment, through a neural network model including a diffusion model, the SMAS algorithm, and the SMAD algorithm, noise processing is performed on the mask image corresponding to the image to be restored while simultaneously acquiring and utilizing global perception information, regional similarity relationships, and salient features to assist in image restoration. The SMAS algorithm mainly utilizes the global perception information and regional similarity relationships of the image to guide the model to fill in finer image details. The SMAD algorithm dynamically extracts salient features by learning the differences between different pixel regions to adapt to different scenarios and improve the flexibility of the model. This improves the detail fidelity and stability of the generated images when dealing with complex scenes, and enhances the adaptability to different scenarios.

[0083] Optionally, the image restoration device may further include: a training module, configured to acquire multiple complete images, and add a mask to each of the complete images using the aforementioned preset masking algorithm to obtain multiple training masked images. An initial neural network model is then subjected to unsupervised training using the training masked images to obtain a pre-trained neural network model. Finally, the pre-trained neural network model is optimized using a preset optimization algorithm to obtain the trained neural network model.

[0084] Optionally, the above diffusion model includes a feedforward network and a feedback network. The above preset time step includes a first preset time step and a second preset time step. The above noise processing includes noise addition processing and noise reduction processing.

[0085] The aforementioned processing module 420 is specifically used to add noise to the masking image through the aforementioned feedforward network at a first preset time step to obtain a noisy masking image. Then, it uses the aforementioned feedback network to denoise the noisy masking image at a second preset time step to obtain a denoised masking image. The denoised masking image is the masking image after noise processing.

[0086] Optionally, the extraction module 430 is specifically used to calculate and obtain the global perception information and the region similarity relationship of the masked image in the denoising process at a second preset time step according to the SMAS algorithm.

[0087] Optionally, the extraction module 430 is specifically used to calculate and obtain the salient features of the masked image in the denoising process at a second preset time step according to the SMAD algorithm.

[0088] The aforementioned repair module 440 is specifically used to, when performing denoising processing on the aforementioned noisy mask image at the aforementioned second preset time step through the aforementioned feedback network, obtain the denoised mask image corresponding to each of the aforementioned second preset time steps based on the aforementioned global perception information, the aforementioned regional similarity relationship, and the aforementioned salient features corresponding to each of the aforementioned second preset time steps, wherein the denoised mask image corresponding to the last of the aforementioned second preset time steps is the aforementioned denoised mask image.

[0089] Optionally, the extraction module 430 is specifically used to calculate and obtain the region similarity relationship of the masked image in the denoising process at a second preset time step according to the following SMAS algorithm formula:

[0090] in, The above-mentioned regions have similarity relationships. and The indices of the masked images in the above denoising process are respectively: The position and index are The pixel value at the location, The above indicates and the above Feature relation values, For the preset normalization function, for The corresponding vector mapping result.

[0091] Optionally, the extraction module 430 is specifically used to calculate and obtain the aforementioned salient features of the masked image in the denoising process at a second preset time step according to the following SMAD algorithm formula:

[0092] in, For the above-mentioned significant features, To pre-define the static basic convolution kernel, Indicates a dynamically adjusted item. For weight generation function, The result is the affinity feature mapping of the location. This indicates that a convolution operation is to be performed.

[0093] The above-described device is used to execute the method provided in the foregoing embodiments, and its implementation principle and technical effect are similar, so they will not be described again here.

[0094] Figure 5 is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device can be a computer, server or other device with computing processing capabilities. As shown in Figure 5, the device 500 includes a processor 510, a storage medium 520 and a bus 530. The processor 510 and the storage medium 520 are connected to each other through the bus 530.

[0095] The storage medium 520 stores machine-readable instructions that can be executed by the processor 510. When the electronic device is running, the processor 510 executes the machine-readable instructions to perform the image restoration method.

[0096] It should be understood that the structure shown in Figure 5 is only a schematic diagram of the electronic device. The electronic device may include more or fewer components than shown in Figure 5, or have a different configuration than shown in Figure 5. The components shown in Figure 5 may be implemented using hardware, software, or a combination thereof.

[0097] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the image restoration method described in the above method embodiments.

[0098] Computer-readable storage media can be electronic storage devices such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Optionally, computer-readable storage media includes non-transitory computer-readable storage medium. The computer-readable storage medium has storage space for program code that performs any of the method steps described above. This program code can be read from or written to one or more computer program exhibits. The program code can be compressed, for example, in a suitable form.

[0099] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program exhibits according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0100] In addition, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0101] If the functionality is implemented as a software module and sold or used as an independent exhibit, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software exhibit. This computer software exhibit is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0102] The above description is merely a preferred embodiment of this application and does not limit the patent scope of this application. Any equivalent structural transformations made based on the inventive concept of this application and the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included within the patent protection scope of this application.

Claims

1. An image restoration method, characterized in that, include: The image to be repaired is obtained, and a mask is added to the image to be repaired using a preset masking algorithm to obtain a masked image; The masked image is input into a trained neural network model, and noise processing is performed on the masked image according to a preset time step. Based on the trained neural network model, global perception information, regional similarity relationships, and salient features of the masked image during noise processing are obtained. The trained neural network model includes a diffusion model, an image region similarity search mechanism SMAS algorithm, and an image region difference search mechanism SMAD algorithm. Based on the global perception information, regional similarity relationships, and salient features, a noise-processed masked image is obtained, which is the repaired image.

2. The image restoration method according to claim 1, characterized in that, Before inputting the masked image into the trained neural network model and performing noise processing on the masked image according to a preset time step, the method further includes: acquiring multiple complete images; adding a mask to each complete image using the preset masking algorithm to obtain multiple training masked images; performing unsupervised training on the initial neural network model using the training masked images to obtain a pre-trained neural network model; and optimizing the pre-trained neural network model according to a preset optimization algorithm to obtain the trained neural network model.

3. The image restoration method according to claim 1, characterized in that, The diffusion model includes a feedforward network and a feedback network; the preset time step includes a first preset time step and a second preset time step; the noise processing includes noise addition and noise removal; the step of inputting the masking image into the trained neural network model and performing noise processing on the masking image according to the preset time step includes: adding noise to the masking image through the feedforward network at the first preset time step to obtain a noisy masking image; and removing noise from the noisy masking image through the feedback network at the second preset time step to obtain a denoised masking image, wherein the denoised masking image is the noise-processed masking image.

4. The image restoration method according to claim 3, characterized in that, The step of obtaining the global perception information and region similarity relationship of the masked image in noise processing according to the trained neural network model includes: calculating and obtaining the global perception information and region similarity relationship of the masked image in denoising processing at a second preset time step according to the SMAS algorithm.

5. The image restoration method according to claim 4, characterized in that, The step of obtaining salient features of the masked image in noise processing according to the trained neural network model includes: calculating and obtaining the salient features of the masked image in denoising processing at a second preset time step according to the SMAD algorithm; the step of obtaining the masked image after noise processing according to the global perception information, region similarity relationship, and salient features includes: when denoising the masked image after noise addition at the second preset time step through the feedback network, obtaining the masked image after denoising processing corresponding to each second preset time step according to the global perception information, region similarity relationship, and salient features corresponding to each second preset time step, wherein the masked image after denoising processing corresponding to the last second preset time step is the masked image after denoising processing.

6. The image restoration method according to claim 4, characterized in that, The step of calculating and obtaining the global perception information and the region similarity relationship of the masked image in the denoising process at a second preset time step according to the SMAS algorithm includes: calculating and obtaining the region similarity relationship of the masked image in the denoising process at a second preset time step according to the following SMAS algorithm formula: in, The region similarity relationship, and The indices within the mask image during the denoising process are respectively: The position and index are The pixel value at the location, Indicates the Japanese Feature relation values, For the preset normalization function, for The corresponding vector mapping result.

7. The image restoration method according to claim 5, characterized in that, The step of calculating and obtaining the salient features of the masked image in the denoising process at a second preset time step according to the SMAD algorithm includes: calculating and obtaining the salient features of the masked image in the denoising process at a second preset time step according to the following SMAD algorithm formula: in, For the aforementioned salient features, To pre-define the static basic convolution kernel, Indicates a dynamically adjusted item. For weight generation function, The result of the affinity feature mapping for location. This indicates that a convolution operation is to be performed.

8. An image restoration device, characterized in that, include: The acquisition module is used to acquire the image to be repaired and add a mask to the image to be repaired using a preset masking algorithm to obtain a masked image; The processing module is used to input the masked image into the trained neural network model and perform noise processing on the masked image according to a preset time step; An extraction module is used to obtain global perception information, regional similarity relationships, and salient features of the masked image in noise processing based on the trained neural network model; wherein, the trained neural network model includes: a diffusion model, an image region similarity search mechanism SMAS algorithm, and an image region difference search mechanism SMAD algorithm; a repair module is used to obtain the noise-processed masked image based on the global perception information, regional similarity relationships, and salient features, wherein the noise-processed masked image is the repaired image.

9. An electronic device, characterized in that, include: The device includes a processor, a storage medium, and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is in operation, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the steps of the method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the method as described in any one of claims 1-7.