Image restoration method and device, electronic equipment and storage medium
By introducing an adaptive switching mechanism and using the loss function defined by edge and grayscale image features to train the model, the problem of restoration accuracy caused by texture area differences in video restoration is solved, and the restoration effect of clear edges in high-texture areas and natural colors in low-texture areas is achieved.
Patent Information
- Application Number
- CN202510866159.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-09-26
AI Technical Summary
Existing technologies have difficulty in restoring images in both high-texture and low-texture areas during video restoration, resulting in reduced restoration accuracy.
An adaptive switching mechanism based on average gradient amplitude is introduced. Different loss functions are defined by edge image features and grayscale image features. The second model and the first model are trained respectively for image restoration in high-texture and low-texture areas.
In high-texture areas, the edge structure is ensured to be clear, accurate and complete, and in low-texture areas, the grayscale/color transition is ensured to be natural and uniform, thereby improving the accuracy of image restoration.
Smart Images

Figure CN120707443A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to an image restoration method, device, electronic device, and storage medium. Background Art
[0002] Video restoration requires the restoration of multiple frames based on a restoration model. The restoration regions corresponding to each frame may include high-texture regions (those rich in edges and details, such as faces and fabrics) or low-texture regions (those relatively smooth and lacking significant structure, such as backgrounds and skies). Therefore, how to restore each frame has become a hot research topic in image processing. Summary of the Invention
[0003] The present application provides an image restoration method, device, electronic device and storage medium, which introduces an adaptive switching mechanism based on average gradient amplitude to achieve switching between a first model and a second model, which can take into account the texture conditions of different images and improve the accuracy of image restoration.
[0004] In a first aspect, the present application provides an image restoration method, the method comprising: Obtaining an average gradient magnitude between each pixel position in the repaired area of the first image; If the average gradient amplitude is not greater than a preset threshold, repairing the first image according to the first model to obtain a second image, where the first model is obtained by model training based on a first loss function, the first loss function includes a first sub-loss function determined based on a difference between a grayscale image of the third image and a grayscale image of the fourth image, the third image is used to represent a training target corresponding to an input image of the model during the model training process, and the fourth image is used to represent an output image of the model during the model training process; If the average gradient amplitude is greater than a preset threshold, the first image is repaired according to the second model to obtain a second image, where the second model is obtained by model training based on a second loss function, and the second loss function includes a second sub-loss function determined based on the difference between the edge image of the third image and the edge image of the fourth image.
[0005] As can be seen, this application introduces an adaptive switching mechanism based on average gradient amplitude. In high-texture areas (corresponding to higher average gradient amplitudes), a second loss function is defined based on the characteristics of the edge image. A second model trained based on this second loss function is used for image restoration, ensuring that the edge structure of the restoration result is clear, accurate, and complete. In low-texture areas (corresponding to lower average gradient amplitudes), a first loss function is defined based on the characteristics of the grayscale image. A first model trained based on this first loss function is used for image restoration, ensuring that the grayscale / color transition of the restoration result is natural and uniform. This allows for a balance between different texture conditions and improves the accuracy of image restoration.
[0006] In a second aspect, the present application provides an image restoration device, comprising: an acquisition unit, configured to acquire an average gradient magnitude between each pixel position in the repaired area of the first image; If the average gradient amplitude is not greater than a preset threshold, the processing unit is configured to repair the first image according to the first model to obtain a second image, where the first model is obtained by model training based on a first loss function, the first loss function including a first sub-loss function determined based on a difference between a grayscale image of the third image and a grayscale image of the fourth image, the third image being used to represent a training target corresponding to an input image of the model during the model training process, and the fourth image being used to represent an output image of the model during the model training process; If the average gradient amplitude is greater than a preset threshold, the processing unit is further used to repair the first image according to the second model to obtain a second image, where the second model is obtained by model training based on a second loss function, and the second loss function includes a second sub-loss function determined based on the difference between the edge image of the third image and the edge image of the fourth image.
[0007] In a third aspect, the present application provides an electronic device comprising a processor, a memory, and a communication interface. The processor, memory, and communication interface are interconnected and perform communication with each other. The memory stores executable program code, the communication interface is used for wireless communication, and the processor is used to retrieve the executable program code stored in the memory and execute some or all of the steps described in any method of the first aspect.
[0008] In a fourth aspect, the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements some or all of the steps described in the first aspect of the present application.
[0009] In a fifth aspect, the present application provides a computer program product, including a computer program, which, when processed and executed, implements some or all of the steps described in the first aspect of the present application. The computer program product may be a software installation package. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0011] Figure 1 A schematic structural diagram of an image restoration system provided in an embodiment of the present application; Figure 2 A schematic diagram of a process for an image restoration method provided in an embodiment of the present application; Figure 3 A schematic diagram of a process flow of another image restoration method provided in an embodiment of the present application; Figure 4 A schematic diagram of the structure of an encoder provided in an embodiment of the present application; Figure 5 A flowchart of another image restoration method provided in an embodiment of the present application; Figure 6 A flowchart of another image restoration method provided in an embodiment of the present application; Figure 7 A block diagram of the functional units of an image restoration device provided in an embodiment of the present application; Figure 8 A block diagram of the functional units of another image restoration device provided in an embodiment of the present application; Figure 9 This is a structural block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0012] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0013] The terms "first," "second," and so on, in the specification and claims of this application and the accompanying drawings are used to distinguish between different objects, not to describe a specific order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps is not limited to the listed steps but may optionally include steps not listed, or may optionally include other steps inherent to the process, method, product, or apparatus.
[0014] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0015] Currently, when performing video restoration, a single restoration model is usually used to repair multiple frames of images in the entire video. However, since the restoration areas corresponding to multiple frames of images may include high-texture areas or low-texture areas, it is difficult to take into account different texture conditions using a single model, which reduces the accuracy of image restoration.
[0016] Based on this, the present application provides an image restoration method that introduces an adaptive switching mechanism based on average gradient amplitude. In high-texture areas (corresponding to higher average gradient amplitudes), a second loss function is defined by the characteristics of the edge image. A second model is trained based on this second loss function to perform image restoration, ensuring that the edge structure of the restoration result is clear, accurate, and complete. In low-texture areas (corresponding to lower average gradient amplitudes), a first loss function is defined by the characteristics of the grayscale image. A first model is trained based on this first loss function to perform image restoration, ensuring that the grayscale / color transition of the restoration result is natural and uniform.
[0017] The following describes the prior art involved in this application: Optical Flow: Pixel-level motion estimation between multiple frames of a video, which can be used to guide video restoration.
[0018] Deformable Convolution: A convolutional layer with learnable sampling offsets that can flexibly adapt to complex motion or non-rigid deformations between frames.
[0019] Canny edge detector: A classic image edge detection algorithm that can extract significant structural contour information.
[0020] Encoder: A neural network module that extracts features from input data.
[0021] Pixel images are composed of pixels, the smallest, discrete unit in an image. Each pixel occupies a specific location (coordinate) in the image. Pixel images are the most basic form of image representation, typically stored in a matrix format. Each pixel contains information describing the appearance of that point. Pixel images can include color pixel images, grayscale images, and edge images.
[0022] Color pixel images are a primary and common type of pixel image. Each pixel uses multiple values (usually three) to represent its color information. Currently, the most common color model is RGB (red, green, blue). Each pixel consists of three independent values: (R, G, B). Each value represents the intensity of the corresponding color channel (usually in the range of 0-255). For example, consider a color landscape photo taken with a mobile phone.
[0023] An edge image is an image extracted from a color pixel or grayscale image using an edge detection algorithm, retaining only areas of the image where brightness or color changes dramatically (i.e., edges). Edges are the demarcation lines between different image regions, typically corresponding to object outlines or texture boundaries. Edge images are typically binary images (0 indicates no edge, 1 indicates an edge). Edge images retain key structural information but lose color and detail. Examples include the outline of a mountain or the contours of a tree's branches.
[0024] Grayscale images are images in which each pixel uses only a single value to represent its brightness (grayscale), without any color information. 0 represents pure black, 255 represents pure white, and intermediate values represent varying shades of gray. Grayscale images can be converted from color pixel images, such as a scanned black-and-white photo.
[0025] ProPainter: A video restoration framework that combines optical flow completion, feature propagation, and Transformer refinement modules. It mainly includes the following steps: Loop Completion: This stage aims to repair optical flow interruptions caused by occlusions and missing regions. By using a pre-trained RAFT model to estimate the original optical flow and repeatedly optimizing it in the loop completion module, the motion information of the occluded regions is gradually restored, thereby estimating the global dynamic field of the video.
[0026] Image propagation: Using the completed optical flow, pixels in the visible frame are propagated along the time axis to the missing region, completing the initial completion of the image layer. This stage provides a fundamental guarantee for visual plausibility and temporal coherence.
[0027] Feature propagation: The propagated frames are fed into the encoder to extract deep semantic features. The completed optical flow and mask information are then combined to perform deformable alignment and fusion in the feature space to enhance the contextual semantic connection of the occluded area.
[0028] MSVT (Mask-guided Sparse Video Transformer): By introducing the MSVT module, the propagated features are finely modeled based on spatiotemporal attention, focusing on the repair area and significantly improving processing efficiency through sparse token screening.
[0029] Decoding and reconstruction: Finally, the refined features are restored to the pixel space to generate multiple frames of the restored video.
[0030] The following describes the system architecture involved in this application: See also Figure 1 , Figure 1 A schematic diagram of the structure of an image restoration system provided in an embodiment of the present application is shown in FIG. Figure 1 As shown, the image restoration system 100 includes a terminal device 101 and a server 102 .
[0031] The terminal device 101 is used to collect videos that the user needs to modify, and then transmit the videos to the server 102. Optionally, the terminal device 101 can be a desktop computer, a laptop computer, a tablet computer, a smart phone, etc.
[0032] The server 102 is used to repair each frame of the video transmitted by the terminal device 101, thereby repairing the video. Optionally, the server 102 can be a single server, a server cluster, a cloud server, a cloud computing service center, or other forms of devices with computing capabilities.
[0033] Specifically, after receiving the video transmitted by the terminal device 101, the server 102 performs preliminary processing on the video, and for the first image in the video, obtains the average gradient amplitude between each pixel position in the repair area of the first image; if the average gradient amplitude is not greater than the preset threshold, the first image is repaired according to the first model to obtain the second image, the first model is obtained by model training based on the first loss function, the first loss function includes a first sub-loss function determined based on the difference between the grayscale image of the third image and the grayscale image of the fourth image, the third image is used to represent the training target corresponding to the input image in the model training process, and the fourth image is used to represent the output image in the model training process; if the average gradient amplitude is greater than the preset threshold, the first image is repaired according to the second model to obtain the second image, the second model is obtained by model training based on the second loss function, and the second loss function includes a second sub-loss function determined based on the difference between the edge image of the third image and the edge image of the fourth image.
[0034] It can be seen that by introducing an adaptive switching mechanism based on average gradient amplitude, in high-texture areas (corresponding to higher average gradient amplitudes), a second loss function is defined by the features of the edge image, and a second model is trained based on this second loss function for image restoration, ensuring that the edge structure of the restoration result is clear, accurate, and complete. In low-texture areas (corresponding to lower average gradient amplitudes), a first loss function is defined by the features of the grayscale image, and a first model is trained based on this first loss function for image restoration, ensuring that the grayscale / color transition of the restoration result is natural and uniform.
[0035] Based on this, an embodiment of the present application provides an image restoration method, which is described in detail below with reference to the accompanying drawings.
[0036] Embodiment 1: The main process of the image restoration method is described below.
[0037] See also Figure 2 , Figure 2 This is a flow chart of an image restoration method provided in an embodiment of the present application, which is applied to the above-mentioned server, such as Figure 2 As shown, the method includes the following steps.
[0038] Step S201: Obtain the average gradient amplitude between each pixel position in the repair area of the first image.
[0039] The average gradient magnitude refers to the average intensity of the pixel gradient vector within a local area of the image, reflecting the saliency of the image edge. High-gradient regions typically correspond to high-texture areas (such as faces and fabric), while low-gradient regions correspond to low-texture areas (such as the sky and background). The inpainted area can refer to areas in the video where information is missing due to occlusion or damage. For example, when a camera transmission is interrupted, green or black squares appear on the screen. When a video conference experiences network lag, the face area is covered by purple blocks. Another example is when a video shows a flickering color mosaic area, or when part of the video is fragmented into snowflakes during playback. The inpainted area can be determined within a single image using pixel outlier detection algorithms (such as Z-Score and DBSCAN), or the inpainted area can be determined for each frame of the video using optical flow combined with temporal consistency analysis.
[0040] Optionally, obtaining the average gradient amplitude between each pixel position in the repaired area of the first image includes: converting the first image into a grayscale image to obtain a first grayscale image; determining the gradient amplitude of each pixel position in the first grayscale image; and determining the average gradient amplitude between the gradient amplitudes of each pixel position in the repaired area of the first grayscale image.
[0041] Grayscale conversion is a process that converts a color image into a single-channel luminance image. This eliminates the interference of the color channel on subsequent calculations, allowing gradient calculations to focus on luminance changes rather than color differences. This conversion can be implemented using standard algorithms, such as weighted averaging, to preserve the original image's luminance information and simplify the structural representation.
[0042] Gradient magnitude is a quantitative metric that describes the intensity of local brightness changes in an image. It can be calculated pixel by pixel using a gradient operator (such as the Sobel operator). The Sobel operator uses horizontal and vertical convolution kernels to calculate the horizontal and vertical gradient responses of a pixel, respectively, and then synthesizes the gradient magnitude using the Pythagorean theorem. For example, combining this calculation with Gaussian blur preprocessing can further suppress spurious gradient responses caused by noise, ensuring that the gradient magnitude accurately reflects actual structural features.
[0043] The mean gradient magnitude is the statistical mean of the gradient magnitudes of all pixels within the inpainted area. A binary mask is used to restrict the pixels in the inpainted area to the calculation. This process uses a binary mask to filter out the inpainted area from the first image, avoiding gradient interference from non-inpainted areas and accurately quantifying the structural complexity of the inpainted area.
[0044] For example, when the first image is the t-th frame image in the video , first convert it into a grayscale image and perform Gaussian blur processing to suppress local noise interference, and obtain the first grayscale image. The calculation formula is as follows:
[0045] in, is used to represent the first grayscale image, This parameter represents the RGB-to-Gray conversion algorithm. σ represents the standard deviation of the Gaussian blur kernel, controlling the blur intensity. A larger value for σ results in a stronger blur and more thorough noise suppression. Used to represent the function used to implement Gaussian blur in image processing.
[0046] Then the Sobel gradient magnitude is calculated for each pixel position using the following formula:
[0047] in, Used to represent the gradient magnitude map corresponding to the first grayscale image, Used to represent the gradient component of the first grayscale image calculated by the Sobel operator in the horizontal direction (x-axis). Used to represent the gradient component of the first grayscale image calculated by the Sobel operator in the vertical direction (y-axis).
[0048] In addition, let the binary mask corresponding to the first image be , then calculate the average gradient amplitude in the repair area, and the calculation formula is as follows:
[0049] in, It is used to represent the average gradient amplitude of the gradient amplitude map corresponding to the first grayscale image in the repair area. With preset threshold The comparison results determine the choice of model. The comparison formula is as follows:
[0050] It can be seen that the grayscale conversion eliminates color interference to focus on brightness changes, so that the texture height of the image can be quantified according to the average gradient amplitude of the repaired area.
[0051] Step S202: If the average gradient amplitude is not greater than the preset threshold, the first image is repaired according to the first model to obtain a second image.
[0052] Among them, the first model is obtained by model training based on the first loss function, the first loss function includes a first sub-loss function determined based on the difference between the grayscale image of the third image and the grayscale image of the fourth image, the third image is used to represent the training target corresponding to the input image of the model during the model training process, and the fourth image is used to represent the output image of the model during the model training process.
[0053] Optionally, the first loss function is determined by weighted summation of the first sub-loss function, the third sub-loss function and the fourth sub-loss function, the third sub-loss function is determined based on the difference between the color pixel image of the third image and the color pixel image of the fourth image, and the fourth sub-loss function is determined based on the difference between the third image and the fourth image determined by the generative adversarial network.
[0054] It can be seen that by constraining structural smoothness with grayscale consistency loss, ensuring global color and texture matching with pixel-level reconstruction loss, and improving the authenticity of high-frequency details with adversarial loss, it is possible to achieve the technical effect of ensuring both global consistency and local authenticity in low-texture area restoration tasks.
[0055] For example, in the training phase of the first model, the fourth image corresponding to the t-th frame image of a single video in the training data is and the third image , using the RGB-to-Gray conversion algorithm Perform grayscale image conversion to obtain the grayscale image in the repair area. The calculation formula is as follows: ,
[0056] in, A grayscale image representing the fourth image in the repaired area, It is used to represent the grayscale image of the third image in the repair area. It is used to represent the binary mask corresponding to the repaired area of the t-th frame image in the aforementioned single video. In order to ensure that the overall structure and brightness distribution within the repaired area are reasonable, the grayscale consistency loss formula is defined as follows:
[0057] Among them, T is used to represent the total number of video frames, H and W are used to represent the height and width of the frame resolution. Then, the third sub-loss function is And the fourth sub-loss function And the first sub-loss function Perform weighted summation and obtain the calculation formula of the first loss function as follows:
[0058] in, Used to represent the first loss function, 、 、 They are used to represent the weights corresponding to the first sub-loss function, the third sub-loss function, and the fourth sub-loss function respectively.
[0059] Step S203: If the average gradient amplitude is greater than a preset threshold, the first image is repaired according to the second model to obtain a second image.
[0060] The second model is obtained by model training based on a second loss function, and the second loss function includes a second sub-loss function determined based on the difference between the edge image of the third image and the edge image of the fourth image.
[0061] Optionally, the second loss function can also be combined with the third sub-loss function and the fourth sub-loss function. That is, the second loss function is determined by weighted summation of the second sub-loss function, the third sub-loss function, and the fourth sub-loss function.
[0062] For example, in the training phase of the second model, the fourth image corresponding to the t-th frame image of a single video in the training data is and the third image , apply Canny edge detection with a binary mask to obtain the edge image in the repair area. The calculation formula is as follows: ,
[0063] in, Used to represent the Canny edge detection algorithm, Used to represent the edge image of the fourth image in the repair area, Used to represent the edge image of the third image in the repair area, The binary mask corresponding to the repaired area of the t-th frame image in the aforementioned single video is used. Subsequently, the difference between the extracted edge images is compared and the edge consistency loss formula is defined as follows:
[0064] Among them, T is used to represent the total number of video frames, H and W are used to represent the height and width of the frame resolution. Finally, the third sub-loss function And the fourth sub-loss function And the second sub-loss function Perform weighted summation and obtain the calculation formula of the second loss function as follows:
[0065] in, Used to represent the second loss function, 、 、 They are used to represent the weights corresponding to the second sub-loss function, the third sub-loss function, and the fourth sub-loss function respectively.
[0066] As can be seen, the embodiment of the present application introduces an adaptive switching mechanism based on average gradient amplitude. In high-texture areas (corresponding to higher average gradient amplitudes), a second loss function is defined by the features of the edge image, and a second model is trained based on the second loss function to perform image restoration. This ensures that the edge structure of the restoration result is clear, accurate, and complete. In low-texture areas (corresponding to lower average gradient amplitudes), a first loss function is defined by the features of the grayscale image, and a first model is trained based on the first loss function to perform image restoration. This ensures that the grayscale / color transition of the restoration result is natural and uniform.
[0067] In the second embodiment, the image restoration method is described in detail below in the case where the first image is any frame image in the first video.
[0068] See also Figure 3 , Figure 3 This is a flow chart of another image restoration method provided in an embodiment of the present application, which is applied to the above-mentioned server, such as Figure 3 As shown, the method includes the following steps.
[0069] Step S301: Obtain image features of each frame of image in the second video.
[0070] The image features include any one or more of the following: edge features, first external features, inpainted areas, optical flow features, optical flow valid areas, and second external features. The first external features include color features, texture features, and shape features. The second external features are features obtained by inversely deforming the first external features of each frame of image into the next frame of image based on the optical flow features.
[0071] Edge features can be image boundary structure information extracted by the first encoder, used to constrain the geometric shape and coherence of the repair area. For example, edge features can be obtained by a Canny edge detector or the edge detection branch of a deep neural network. The first external features can be visual representations extracted by the appearance encoder, used to preserve the visual consistency of the repair area. For example, color features can be represented by RGB channel statistics, texture features can be obtained by wavelet transform or convolution feature map, and shape features can be calculated based on contour closure and region area. The repair area can be the image area that needs to be repaired, determined by the defect detection algorithm in the preprocessing step.
[0072] The optical flow feature can be a vector field that describes the pixel motion trajectory between adjacent frames in the video, which is generated by an optical flow estimation algorithm (such as RAFT or PWCNet) and is used to capture the inter-frame motion pattern. The optical flow feature in this application can be an optical flow feature that is completed according to the aforementioned cyclic optical flow completion step. The optical flow valid area can be a binary image, which is generated by the confidence output of the optical flow algorithm or post-processing screening. The second external feature can be a feature map after the first external feature of the next frame image is reversely deformed, which is achieved by reverse interpolation of the optical flow feature and is used to introduce semantic information of future frames to enhance the spatiotemporal correlation of the current frame repair.
[0073] Optionally, the image features include edge features, and the edge features of each frame image are obtained by the following steps: obtaining a first edge image corresponding to each frame image in the second video according to an edge detection algorithm; determining a second edge image located in the repair area in the first edge image corresponding to each frame image in the second video; inputting the second edge image corresponding to each frame image in the second video into the first encoder to obtain edge features corresponding to each frame image in the second video.
[0074] Among them, the edge detection algorithm can be a mathematical method for extracting the boundaries of image structures, such as the Canny algorithm. The first edge image is an image obtained after processing by the edge detection algorithm, which highlights the significant boundary information in the original image, such as object contours or texture boundaries. The second edge image can be a sub-region image of the first edge image that only retains the edge information in the repaired area after spatial screening. The first encoder can be a neural network module specially designed to extract edge features, such as a lightweight encoder containing four convolutional layers and an output channel number of 128, which encodes structural features such as edge direction and contour closure through local pattern capture.
[0075] It can be seen that by using the edge detection algorithm to extract the global edge information of each frame image in the second video, and inputting the edge information corresponding to the repaired area into the encoder, the structured expression of the edge features can be achieved.
[0076] Optionally, the image features include a first external feature, and the first external feature of each frame image is obtained by the following steps: determining the fifth image located in the repair area in each frame image of the second video; inputting the fifth image in each frame image of the second video into the second encoder to obtain the first external feature corresponding to each frame image in the second video, and the number of convolution layers included in the second encoder is greater than the number of convolution layers included in the first encoder.
[0077] Among them, the repair area is also a specific spatial area that needs to be repaired. The fifth image can be a sub-image extracted in the repair area. The second encoder can be a deep neural network structure specially designed for extracting the first external feature, for example, it can include a stacked structure including nine convolutional layers. Since the complexity of the first external feature is higher than that of the edge feature, the number of parameters included in the second encoder should be greater than the number of parameters included in the first encoder. For example, the first encoder can include four convolutional layers, while the second encoder can include nine convolutional layers.
[0078] As can be seen, the second encoder with a higher number of parameters and the first encoder with a lower number of parameters are determined based on the extraction complexity of the first external feature and edge feature, respectively. This reduces the computational complexity of each forward inference during edge feature extraction and improves overall inference efficiency.
[0079] For example, first, the Canny edge detection algorithm is used Get the tth frame image in the second video The corresponding first edge image , the calculation formula is as follows:
[0080] Then, the edge information of the repaired area is extracted by combining the binary mask to construct the second edge image. , the calculation formula is as follows:
[0081] in, Used to represent the t-th frame image in the second video The binary mask corresponding to the original repair area can be understood to construct the t-th frame image The corresponding fifth image The calculation formula is similar, as follows:
[0082] Assuming that the input of the first encoder is three-channel data and the input of the second encoder is five-channel data, the formula for constructing the input form of the first encoder is as follows:
[0083] The formula for constructing the input form of the second encoder is as follows:
[0084] in, Used to represent the input of the first encoder, Used to represent the input of the second encoder, Used to represent the t-th frame image in the second video The binary mask corresponding to the repaired area is updated after the aforementioned image propagation step. Used to represent feature fusion algorithm.
[0085] Then, the two inputs and They are sent to the first encoder and the second encoder for independent processing. The formula for obtaining edge features is as follows:
[0086] The formula for obtaining the first external characteristic is as follows:
[0087] Among them, C is used to represent the number of channels, Used to represent edge features, Used to represent the first external feature, Used to indicate the first encoder, Used to indicate the second encoder.
[0088] For example, see Figure 4 , Figure 4 A schematic diagram of the structure of an encoder provided in an embodiment of the present application is shown in FIG. Figure 4 As shown in the figure, it includes a first encoder and a second encoder. The first encoder consists of four layers of 3×3 convolutions (3×3Conv), two with 64 channels (3×3Conv, 64) and two with 128 channels (3×3Conv, 128), each followed by an activation function (LeakyReLU). The first encoder has a 3-channel input (Input, 3) and a 128-channel output (Output, 128). The second encoder consists of nine layers of 3×3 convolutions, two with 64 channels (3×3Conv, 64), two with 128 channels (3×3Conv, 128), two with 256 channels (3×3Conv, 256), two with 384 channels (3×3Conv, 384), and one with 512 channels (3×3Conv, 512).
[0089] Step S302: Input the image features of each frame in the second video into a variable convolutional network to obtain the first video.
[0090] For example, the input condition of the deformable convolutional network is constructed as follows:
[0091] in, , used to represent the first external feature of the t-th frame image in the second video, Used to represent the second external feature of the t-th frame image in the second video, Used to represent the optical flow features of the t-th frame image in the second video, It is used to represent the binary image corresponding to the effective optical flow area of the t-th frame image in the second video. It is used to represent the binary mask corresponding to the original repaired area of the t-th frame image in the second video, Used to represent the edge features of the t-th frame image in the second video. Used to represent feature fusion algorithm. Used to represent the image features of the t-th frame image in the second video.
[0092] It can be seen that by using edge features to constrain the structure, external features to ensure visual consistency, and optical flow features to guide spatiotemporal alignment, and at the same time achieving adaptive fusion of cross-frame features through dynamic offset adjustment, the technical effect of enhancing the spatiotemporal continuity and detail fidelity of the repaired area can be achieved.
[0093] Step S303: for the first image in the first video, obtain the average gradient amplitude between each pixel position in the repair area of the first image.
[0094] Step S304: If the average gradient amplitude is not greater than the preset threshold, the first image is repaired according to the first model to obtain a second image.
[0095] Among them, the first model is obtained by model training based on the first loss function, the first loss function includes a first sub-loss function determined based on the difference between the grayscale image of the third image and the grayscale image of the fourth image, the third image is used to represent the training target corresponding to the input image of the model during the model training process, and the fourth image is used to represent the output image of the model during the model training process.
[0096] Step S305: If the average gradient amplitude is greater than a preset threshold, the first image is repaired according to the second model to obtain a second image.
[0097] The second model is obtained by model training based on a second loss function, and the second loss function includes a second sub-loss function determined based on the difference between the edge image of the third image and the edge image of the fourth image.
[0098] In the third embodiment, the image restoration method is described in detail below, also in the case where the first image is any frame image in the first video.
[0099] See also Figure 5 , Figure 5 A flowchart of another image restoration method provided in an embodiment of the present application is shown, which is applied to the above-mentioned server. Figure 5 As shown, the method includes the following steps.
[0100] Step S501: Obtain the optical flow features of each frame image in the third video, and determine the corresponding second pixel position of the first pixel position of the t-th frame image in the t+1-th frame image in the third video according to the optical flow features.
[0101] Among them, the first pixel position refers to the pixel position in the repair area of the t-th frame image, 1≤t≤T, T is the total number of frames of the third video. The optical flow feature can be a vector field that describes the pixel motion trajectory between video frames. This feature can characterize the displacement direction and distance of each pixel in the t-th frame in the t+1-th frame, and can be obtained by performing optical flow field analysis on the t-th frame and the t+1-th frame of the third video. Exemplarily, the optical flow feature may include an optical flow vector field, in which each vector corresponds to the displacement parameter of the pixel at the t-th frame position in the next frame. The second pixel position may be the corresponding coordinate of the first pixel position in the t-th frame in the t+1-th frame, which is calculated by reverse tracing of the optical flow field.
[0102] Step S502 : Overwrite the pixel value of the corresponding second pixel position in the t+1th frame image in the third video to the first pixel position of the tth frame image to obtain the second video.
[0103] Among them, in the pixel value overwriting process, the pixel value of the second pixel position is first extracted from the t+1 frame. If the coordinate is non-integer, the sub-pixel pixel value is obtained by bilinear interpolation to ensure a smooth transition of the repaired area. In the case where the second pixel position in the t+1 frame is still invisible due to continuous occlusion, the optical flow information of the area must be restored first, and then the pixel value migration is performed. Finally, the extracted pixel value is directly written into the first pixel position of the t frame to form a preliminary repaired image of the t frame, and this process is repeated to process the images of all frames in the third video to generate the second video.
[0104] It can be seen that the first pixel position is determined by optical flow feature extraction and reverse tracking, and the first pixel position is preliminarily repaired based on the optical flow feature. This, combined with the subsequent image repair of each frame through the first model or the second model, can improve the accuracy of video repair.
[0105] Step S503: Obtain image features of each frame of image in the second video.
[0106] It can be understood that the optical flow features in the image features of each frame image in the third video are the same as the optical flow features of each frame image in the second video.
[0107] Step S504: input the image features of each frame of the second video into a variable convolutional network to obtain the first video.
[0108] Step S505 : for the first image in the first video, obtain the average gradient amplitude between each pixel position in the repaired area of the first image.
[0109] Step S506: If the average gradient amplitude is not greater than the preset threshold, the first image is repaired according to the first model to obtain a second image.
[0110] Among them, the first model is obtained by model training based on the first loss function, the first loss function includes a first sub-loss function determined based on the difference between the grayscale image of the third image and the grayscale image of the fourth image, the third image is used to represent the training target corresponding to the input image during the model training process, and the fourth image is used to represent the output image during the model training process.
[0111] Step S507: If the average gradient amplitude is greater than a preset threshold, the first image is repaired according to the second model to obtain a second image.
[0112] The second model is obtained by model training based on a second loss function, and the second loss function includes a second sub-loss function determined based on the difference between the edge image of the third image and the edge image of the fourth image.
[0113] For example, see Figure 6 , Figure 6 A flow chart of another image restoration method provided in an embodiment of the present application is shown as follows: Figure 6As shown, the occluded optical flow information (Masked Flows) is first input into the recurrent flow completion module (Recurrent Flow Compeletion) to obtain the completed optical flow information (Completed Flows) (i.e., the aforementioned optical flow features). This information is then combined with the repaired region of each frame in the third video (Masked Frames) and input into the image propagation module (ImageProp) to obtain the updated second video (Updated Frames). Subsequently, the binary mask corresponding to the original repaired region of each frame in the second video, the binary mask corresponding to the repaired region updated after the aforementioned image propagation model for each frame, and the fifth image corresponding to each frame are combined and input into the second encoder (Encoder). The obtained first external feature of each frame is input into the feature propagation module (Feature Prop). In addition, a third video (Input Frames) is selected and input into a Canny edge detector (Canny Detector) to obtain edge images (EdgeMaps(Input)) for each frame in the third video. These images are combined with the original binary masks corresponding to the inpainted regions (OriginalMasks) in each frame of the third video to obtain edge images (Masked Edge Maps) corresponding to the inpainted regions. The image propagation module also outputs updated binary masks corresponding to the inpainted regions (Updated Masks). These updated binary masks, the original binary masks, and the edge images corresponding to the inpainted regions are input into the first encoder (Encoder(Edge)). The output edge features are then fed into the feature propagation module. The features output by the feature propagation module are then refined using spatiotemporal attention by an MSVT module (MSVT Blocks×N), focusing on the inpainted regions (Masks). Sparse token filtering significantly improves processing efficiency. Finally, the decoder (Decoder) restores the refined features to pixel space to generate the inpainted video (OutputFrames). It is understood that the edge features determined above can also be edge features of each frame in the second video.
[0114] It is further explained that for the model training of the decoder, edge consistency loss can be introduced. Specifically, the edge image (Edge Maps (Output)) corresponding to each frame of the restored video is obtained by inputting the restored video into the Canny edge detector (Canny Detector), and the edge image corresponding to the restored area of each frame of the restored video is obtained by combining the original binary mask (Original Masks). According to the training target corresponding to the edge image (Edge Maps (Input)) of each frame of the third video, the edge image corresponding to the restored area of the training target of each frame of the third video is obtained by combining the original occlusion mask (Original Masks). Finally, a second loss function (Edge Loss) is constructed based on the edge image corresponding to the restored area of the training target of each frame of the third video and the edge image corresponding to the restored area of each frame of the restored video, and the decoder (corresponding to the second model in this application) is trained according to the second loss function.
[0115] In accordance with the above-mentioned embodiment, please refer to Figure 7 , Figure 7 This is a block diagram of the functional units of an image restoration device provided in an embodiment of the present application. The image restoration device is the above-mentioned server or a part of the server, such as Figure 7 As shown, the image restoration device 70 includes: An acquiring unit 701 is configured to acquire an average gradient magnitude between each pixel position in the repaired area of the first image; If the average gradient amplitude is not greater than the preset threshold, the processing unit 702 is configured to repair the first image according to the first model to obtain a second image, where the first model is obtained by model training based on a first loss function, the first loss function including a first sub-loss function determined based on a difference between a grayscale image of the third image and a grayscale image of the fourth image, the third image being used to represent a training target corresponding to an input image of the model during the model training process, and the fourth image being used to represent an output image of the model during the model training process; If the average gradient amplitude is greater than a preset threshold, the processing unit 702 is further used to repair the first image according to the second model to obtain a second image, where the second model is obtained by model training based on a second loss function, and the second loss function includes a second sub-loss function determined based on the difference between the edge image of the third image and the edge image of the fourth image.
[0116] In a feasible embodiment, in terms of obtaining the average gradient magnitude between each pixel position in the repaired area of the first image, the obtaining unit 701 is specifically configured to: Converting the first image into a grayscale image to obtain a first grayscale image; determining a gradient magnitude at each pixel position in the first grayscale image; An average gradient magnitude between the gradient magnitudes at each pixel position in the inpainted region of the first grayscale image is determined.
[0117] In a feasible embodiment, the first image is any frame image in the first video, and before the first image is restored according to the second model, the acquiring unit 701 is further configured to: Obtaining image features of each frame of the second video, where the image features include any one or more of the following: an edge feature, a first external feature, a repaired area, an optical flow feature, an optical flow valid area, and a second external feature, where the first external feature includes a color feature, a texture feature, and a shape feature, and the second external feature is a feature obtained by inversely deforming the first external feature of the next frame of each frame according to the optical flow feature; The processing unit 702 is further configured to input the image features of each frame of the second video into the variable convolutional network to obtain the first video.
[0118] In a feasible embodiment, the acquiring unit 701 is specifically configured to: Acquire a first edge image corresponding to each frame image in the second video according to an edge detection algorithm; Determine a second edge image located in the repair area in the first edge image corresponding to each frame image in the second video; The second edge image corresponding to each frame image in the second video is input into the first encoder to obtain the edge features corresponding to each frame image in the second video.
[0119] In a feasible embodiment, the acquiring unit 701 is specifically configured to: determining a fifth image located in the repair area in each frame of the second video; The fifth image in each frame of the second video is input into the second encoder to obtain the first external feature corresponding to each frame of the second video, and the parameter amount included in the second encoder is greater than the parameter amount included in the first encoder.
[0120] In a feasible embodiment, the processing unit 702 is further configured to: Determine, based on the optical flow feature, a corresponding second pixel position of a first pixel position of the t-th frame image in the t+1-th frame image in the third video, where the first pixel position refers to a pixel position within the repaired area of the t-th frame image, 1≤t≤T, where T is the total number of frames of the third video; The pixel value of the corresponding second pixel position in the t+1th frame image in the third video is overwritten to the first pixel position of the tth frame image to obtain the second video.
[0121] In a feasible embodiment, the first loss function is determined by weighted summation of the first sub-loss function, the third sub-loss function and the fourth sub-loss function, the third sub-loss function is determined based on the difference between the color pixel image of the third image and the color pixel image of the fourth image, and the fourth sub-loss function is determined based on the difference between the third image and the fourth image determined by the generative adversarial network.
[0122] It can be understood that since the method embodiment and the device embodiment are different presentation forms of the same technical concept, the content of the method embodiment part in this application should be synchronously adapted to the device embodiment part and will not be repeated here.
[0123] In the case of integrated units, such as Figure 8 As shown, Figure 8 This is a block diagram of the functional units of another image restoration device provided in an embodiment of the present application. Figure 8 In the embodiment, the image restoration apparatus 70 includes: a processing module 812 and a communication module 811. The processing module 812 is used to control and manage the actions of the image restoration apparatus 70, for example, the steps of the acquisition unit 701 and the processing unit 702, and / or other processes for executing the technology described herein. The communication module 811 is used to support the interaction between the image restoration apparatus 70 and other devices. Figure 8 As shown, the image restoration device 70 may further include a storage module 813 , which is used to store program codes and data of the image restoration device 70 .
[0124] The processing module 812 may be a processor or controller, such as a central processing unit (CPU), a general-purpose processor, a digital signal processor (DSP), an ASIC, an FPGA, or other programmable logic device, a transistor logic device, a hardware component, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like. The communication module 811 may be a transceiver, an RF circuit, or a communication interface, and the like. The storage module 813 may be a memory.
[0125] Among them, all relevant contents of each scene involved in the above method embodiment can be referred to the functional description of the corresponding functional module, and will not be repeated here. Figure 2 The image restoration method shown.
[0126] The above embodiments can be implemented in whole or in part via software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. A computer program product comprises one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. For example, computer instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. A computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server or data center that contains a collection of one or more available media. Available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media. Semiconductor media can be solid-state drives.
[0127] Figure 9 This is a structural block diagram of an electronic device provided in an embodiment of the present application. Figure 9 As shown, the electronic device 900 may include one or more of the following components: a processor 901, a memory 902 and a communication interface 903. The processor 901, the memory 902 and the communication interface 903 are interconnected and perform communication with each other. The memory 902 may store one or more computer programs, and the one or more computer programs may be configured to implement the methods described in the above embodiments when executed by one or more processors 901.
[0128] The processor 901 may include one or more processing cores. The processor 901 uses various interfaces and lines to connect the various parts of the entire electronic device 900, and performs various functions and processes data of the electronic device 900 by running or executing instructions, programs, code sets or instruction sets stored in the memory 902, and calling data stored in the memory 902. Optionally, the processor 901 can be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array (PLA). The processor 901 can integrate one or more combinations of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. It is understandable that the above-mentioned modem may not be integrated into the processor 901, but may be implemented separately through a communication chip.
[0129] The memory 902 may include a random access memory (RAM) or a read-only memory (ROM). The memory 902 may be used to store instructions, programs, codes, code sets, or instruction sets. The memory 902 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc. The data storage area may also store data created by the electronic device 900 during use.
[0130] It is understandable that the electronic device 900 may include more or fewer structural elements than those in the above structural block diagram, for example, a power module, physical buttons, a WiFi (Wireless Fidelity) module, a speaker, a Bluetooth module, a sensor, etc., which are not limited here.
[0131] The electronic device 900 may be a server or a part of a server.
[0132] An embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, which, when executed by a processor, implements part or all of the steps of any one of the image restoration methods described in the above method embodiments.
[0133] The present application also provides a computer program product, including a computer program, which, when executed by a processor, implements some or all of the steps of any of the image restoration methods described in the above method embodiments. The computer program product may be a software installation package.
[0134] It should be noted that for the method embodiments of any of the aforementioned image restoration methods, for the sake of simplicity, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions involved are not necessarily required by this application.
[0135] Although the present application is described herein in conjunction with various embodiments, in the process of implementing the claimed application, those skilled in the art may understand and implement other variations of the disclosed embodiments by reviewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality of components or steps. The fact that certain measures are recited in different dependent claims does not mean that these measures cannot be combined to produce good results.
[0136] Those skilled in the art will appreciate that all or part of the steps in the various methods of any of the above-mentioned image restoration method embodiments may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable memory, which may include a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.
[0137] The above is a detailed introduction to the embodiments of the present application. Specific examples are used herein to illustrate the principles and implementation methods of an image restoration method, device, electronic device, and storage medium of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for those skilled in the art, based on the idea of an image restoration method, device, electronic device, and storage medium of the present application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
[0138] The present application is described with reference to the flowcharts and / or block diagrams of the methods, hardware products, and computer program products of the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0139] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0140] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0141] It can be understood that any product that is controlled or configured to execute the processing method of the flowchart described in the method embodiment of an image restoration method of the present application, such as the terminal and computer program product in the above flowchart, falls within the scope of the related products described in this application.
[0142] Obviously, those skilled in the art may make various modifications and variations to the image restoration method, apparatus, electronic device, and storage medium provided herein without departing from the spirit and scope of the present application. Thus, if such modifications and variations fall within the scope of the present claims and their equivalents, the present application is intended to encompass such modifications and variations.
Claims
1. An image restoration method, characterized in that: The method comprises: Obtaining an average gradient magnitude between each pixel position in the repaired area of the first image; If the average gradient amplitude is not greater than a preset threshold, repairing the first image according to a first model to obtain a second image, wherein the first model is obtained by model training based on a first loss function, the first loss function including a first sub-loss function determined based on a difference between a grayscale image of the third image and a grayscale image of the fourth image, the third image being used to represent a training target corresponding to an input image of the model during the model training process, and the fourth image being used to represent an output image of the model during the model training process; If the average gradient amplitude is greater than the preset threshold, the first image is repaired according to the second model to obtain a second image, where the second model is obtained by model training based on a second loss function, and the second loss function includes a second sub-loss function determined based on the difference between the edge image of the third image and the edge image of the fourth image.
2. The method according to claim 1, characterized in that The obtaining of the average gradient magnitude between each pixel position in the repaired area of the first image includes: Converting the first image into a grayscale image to obtain a first grayscale image; determining a gradient magnitude at each pixel position in the first grayscale image; An average gradient magnitude between the gradient magnitudes at each pixel position in the repaired area of the first grayscale image is determined.
3. The method according to claim 1, characterized in that The first image is any frame image in the first video. Before repairing the first image according to the second model, the method further includes: Obtaining image features of each frame of the second video, the image features including any one or more of the following: an edge feature, a first external feature, a repaired area, an optical flow feature, an optical flow valid area, and a second external feature, wherein the first external feature includes a color feature, a texture feature, and a shape feature, and the second external feature is a feature obtained by inversely deforming the first external feature of the next frame of each frame according to the optical flow feature; The image features of each frame of the second video are input into a deformable convolutional network to obtain the first video.
4. The method according to claim 3, characterized in that The image features include edge features, and the edge features of each frame of image are obtained by the following steps: Acquire a first edge image corresponding to each frame of the second video according to an edge detection algorithm; Determine a second edge image located in the repair area in the first edge image corresponding to each frame of image; The second edge image corresponding to each frame image in the second video is input into the first encoder to obtain the edge features corresponding to each frame image in the second video.
5. The method according to claim 4, characterized in that The image features include a first external feature, and the first external feature of each frame of image is obtained by the following steps: determining a fifth image located in a repair area in each frame of the second video; The fifth image in each frame is input into a second encoder to obtain a first external feature corresponding to each frame in the second video, and the parameter amount included in the second encoder is greater than the parameter amount included in the first encoder.
6. The method according to claim 3, characterized in that The method further comprises: Determine, based on the optical flow feature, a corresponding second pixel position of a first pixel position of the t-th frame image in the t+1-th frame image of the third video, where the first pixel position refers to a pixel position within the repaired area of the t-th frame image, 1≤t≤T, where T is the total number of frames of the third video; Overwrite the pixel value of the second pixel position in the t+1th frame image to the first pixel position of the tth frame image to obtain a second video.
7. The method according to claim 1, characterized in that The first loss function is determined by weighted summation of the first sub-loss function, the third sub-loss function and the fourth sub-loss function, the third sub-loss function is determined based on the difference between the color pixel image of the third image and the color pixel image of the fourth image, and the fourth sub-loss function is determined based on the difference between the third image and the fourth image determined by the generative adversarial network.
8. An image restoration device, characterized in that: The device comprises: an acquisition unit, configured to acquire an average gradient magnitude between each pixel position in the repaired area of the first image; If the average gradient amplitude is not greater than a preset threshold, a processing unit is configured to repair the first image according to a first model to obtain a second image, where the first model is obtained by model training based on a first loss function, the first loss function including a first sub-loss function determined based on a difference between a grayscale image of a third image and a grayscale image of a fourth image, the third image being used to represent a training target corresponding to an input image of the model during model training, and the fourth image being used to represent an output image of the model during model training; If the average gradient amplitude is greater than the preset threshold, the processing unit is further used to repair the first image according to a second model to obtain a second image, where the second model is obtained by model training based on a second loss function, and the second loss function includes a second sub-loss function determined based on the difference between the edge image of the third image and the edge image of the fourth image.
9. An electronic device comprising a processor, a memory, and an executable program code stored in the memory, wherein: The processor is configured to retrieve the executable program code stored in the memory to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.