Image inpainting method

By using deep network sampling and deconvolution processing, the problems of blurry and low-resolution images in cloud storage were solved, improving image resolution and clarity, and enhancing the user experience.

CN119809989BActive Publication Date: 2026-08-04CHINA MOBILE INTERNET CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA MOBILE INTERNET CO LTD
Filing Date
2025-01-03
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

The image quality in cloud storage varies greatly; early photos are blurry and have low resolution, resulting in a poor user experience.

Method used

By using deep networks for sampling, feature information of adjacent regions is obtained, and pixel amplification and deconvolution are performed to delete insufficiently repaired pixel blocks and synthesize pixel blocks with qualified repair, thereby improving image resolution and clarity.

Benefits of technology

It increases the receptive field of the image, improves the image resolution and clarity, and enhances the user's viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119809989B_ABST
    Figure CN119809989B_ABST
Patent Text Reader

Abstract

The application discloses an image repairing method. The image repairing method comprises the following steps: acquiring an original image to be processed; performing sampling processing on the original image; obtaining feature information between adjacent regions in the original image by using a deep network during the sampling process; performing pixel point amplification processing on the feature information, wherein the relative spatial positions of each pixel point before and after amplification are unchanged; performing inverse convolution processing on the pixel point amplified feature information to obtain a target image, wherein the target image comprises a plurality of pixel blocks; deleting a pixel block with a repairing degree less than a repairing degree threshold in the obtained target image; and synthesizing a pixel block with a repairing degree greater than or equal to the repairing degree threshold to obtain a repaired image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and in particular to an image restoration method. Background Technology

[0002] Cloud storage is a professional internet storage tool, a product of internet cloud technology. It provides services such as information storage, retrieval, and downloading for businesses and individuals via the internet, featuring security, stability, and massive storage capacity. It offers users storage and backup services for digital assets such as files, mobile phone pictures, and contacts, as well as multi-device synchronization, online management, and group sharing. For user convenience, cloud storage also provides image enhancement features.

[0003] The image enhancement function in cloud storage technology mainly involves beautifying photos containing portraits to improve their aesthetic appeal. However, as users store more and more images in their cloud storage, the quality of these uploaded images has become inconsistent. For example, images taken with low-resolution cameras may be blurry and have low resolution, resulting in a poor viewing experience for users browsing these images on the cloud storage. Summary of the Invention

[0004] This application provides an image restoration method to address the problem that image enhancement techniques in related technologies cannot solve the issues of blurry or low-resolution photos.

[0005] In a first aspect, embodiments of this application provide an image restoration method, including:

[0006] Obtain the original image to be processed;

[0007] The original image is sampled. During the sampling process, a deep network is used to obtain the feature information between adjacent regions in the original image. The feature information is then amplified into pixels, wherein the relative spatial position of each pixel remains unchanged before and after amplification.

[0008] The feature information after pixel amplification is deconvolved to obtain the target image, wherein the target image includes multiple pixel blocks;

[0009] Pixel blocks with a repair degree less than the repair degree threshold in the obtained target image are deleted, and pixel blocks with a repair degree greater than or equal to the repair degree threshold are synthesized to obtain the repaired image.

[0010] Secondly, embodiments of this application provide an image restoration apparatus, comprising:

[0011] The acquisition module is used to acquire the original image to be processed;

[0012] The sampling module is used to sample the original image. During the sampling process, a deep network is used to obtain the feature information between adjacent regions in the original image, and the feature information is amplified into pixels. The relative spatial position of each pixel remains unchanged before and after amplification.

[0013] The processing module is used to perform deconvolution processing on the feature information after pixel amplification to obtain a target image, wherein the target image includes multiple pixel blocks;

[0014] The determination module is used to delete pixel blocks in the obtained target image whose repair degree is less than the repair degree threshold, and to synthesize pixel blocks whose repair degree is greater than or equal to the repair degree threshold to obtain the repaired image.

[0015] Thirdly, embodiments of this application provide an electronic device, including a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the method described in the first aspect.

[0016] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the method described in the first aspect.

[0017] Fifthly, embodiments of this application provide a computer program product, the computer program product including a computer program stored on a non-transitory computer-readable storage medium, the computer program including program instructions, which, when executed by a computer, implement the steps of the method described in the first aspect.

[0018] In this embodiment, the original image to be processed is first acquired, and then sampled. During the sampling process, a deep network is used to obtain feature information between adjacent regions in the original image, and the feature information is then augmented with pixels. The relative spatial positions of the pixels before and after augmentation remain unchanged. The augmented feature information is then deconvolved to obtain the target image, which includes multiple pixel blocks. Finally, pixel blocks with a repair score less than a repair score threshold are deleted, and pixel blocks with a repair score greater than or equal to the repair score threshold are synthesized to obtain the repaired image. This embodiment, by sampling the original image to obtain feature information between adjacent regions and augmenting the feature information with pixels, can increase the receptive field of the original image, thereby improving image resolution while enhancing the image. Attached Figure Description

[0019] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0020] Figure 1 This is a flowchart of the image restoration method provided in the embodiments of this application;

[0021] Figure 2 This is a schematic diagram of the structure of the image restoration model provided in the embodiments of this application;

[0022] Figure 3 This is a schematic diagram of the structure of the fusion coding module provided in the embodiments of this application;

[0023] Figure 4 This is a schematic diagram of the structure of the high-density coding module provided in the embodiments of this application;

[0024] Figure 5 This is a schematic diagram of the pixel amplification module provided in an embodiment of this application;

[0025] Figure 6 This is a schematic diagram of the structure of the noise removal module provided in the embodiments of this application;

[0026] Figure 7 This is a schematic diagram of the structure of the learning network provided in the embodiments of this application;

[0027] Figure 8 This is a schematic diagram of the discriminant network provided in the embodiments of this application;

[0028] Figure 9-10 This is a schematic diagram of the user interface provided in an embodiment of this application;

[0029] Figure 11 This is a schematic diagram comparing images before and after using the image restoration method provided in the embodiments of this application;

[0030] Figure 12 This is a schematic diagram of the image restoration apparatus provided in the embodiments of this application;

[0031] Figure 13 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0032] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0033] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0034] The following is in conjunction with the appendix Figures 1 to 13 This application provides a detailed description of an image restoration method through specific embodiments and application scenarios. In related technologies, the beautification function in cloud storage primarily enhances the appearance of images containing portraits, improving their aesthetics. However, with the increasing number of cloud storage users, further development is needed to enrich the functionality of cloud storage and address the beautification function. As the number of images stored in cloud storage increases, the quality of uploaded images becomes inconsistent. For example, images taken with low-resolution cameras may be blurry and have low resolution, resulting in a poor viewing experience for users browsing these images on cloud storage. To address these issues, this application provides an image restoration method that increases the receptive field of the original image to increase its resolution and improve its clarity.

[0035] like Figure 1 The diagram shown is a flowchart of an image restoration method provided in an embodiment of this application. Figure 1 As shown, the image restoration method is applied to an electronic device, and the image restoration method may include the contents shown in S101 to S104.

[0036] In S101, the original image to be processed is obtained.

[0037] The original image can be a picture taken by the user using a low-resolution camera, which can be repaired using the image restoration model provided in this embodiment.

[0038] In S102, the original image is sampled. During the sampling process, a deep network is used to obtain the feature information between adjacent regions in the original image, and the feature information is amplified into pixels. The relative spatial position of each pixel remains unchanged before and after amplification.

[0039] The sampling process is an upsampling process, which uses deep networks to obtain more feature information between adjacent regions, and then performs pixel amplification processing to expand the receptive field without losing resolution and to keep the relative spatial position of the pixels unchanged.

[0040] In S103, the feature information after pixel amplification is deconvolved to obtain the target image, which includes multiple pixel blocks.

[0041] In this example, the image can be decoded and the target image can be obtained through the deconvolution operation.

[0042] In S104, pixel blocks with a repair degree less than the repair degree threshold in the obtained target image are deleted, and pixel blocks with a repair degree greater than or equal to the repair degree threshold are synthesized to obtain the repaired image.

[0043] In other words, parts that do not achieve the desired restoration effect can be removed, thereby improving the overall clarity of the image.

[0044] In this embodiment, the original image to be processed is first acquired, and then sampled. During the sampling process, a deep network is used to obtain feature information between adjacent regions in the original image, and the feature information is then augmented with pixels. The relative spatial positions of the pixels before and after augmentation remain unchanged. The augmented feature information is then deconvolved to obtain the target image, which includes multiple pixel blocks. Finally, pixel blocks with a repair score less than a repair score threshold are deleted, and pixel blocks with a repair score greater than or equal to the repair score threshold are synthesized to obtain the repaired image. This embodiment, by sampling the original image to obtain feature information between adjacent regions and augmenting the feature information with pixels, can increase the receptive field of the original image, thereby improving image resolution while enhancing the image.

[0045] In one possible implementation of this application, the original image is sampled. During the sampling process, a deep network is used to obtain feature information between adjacent regions in the original image, and the feature information is then amplified by pixel density. This process may include: sampling the original image, passing the original image through a convolution operation of a first fusion coding module to obtain long-range feature information of the original image, where the computational complexity of the long-range feature information is lower than that of the feature information in the original image; processing the long-range feature information through shallow network convolutional layers and deep network convolutional layers in a first high-density coding module to obtain feature information between multiple adjacent regions; and amplifying the feature information between multiple adjacent regions through a first pixel density amplification module to obtain feature information with an increased receptive field.

[0046] Furthermore, after the first pixel density amplification module, a first noise removal module may also be included.

[0047] In other words, the modules in the upsampling process may include a first fusion encoding module, a first high-density encoding module, a first pixel density amplification module, and a first noise removal module connected in sequence.

[0048] The first fusion coding module is used to obtain long-range feature information of the original image after convolution operation; the first high-density coding module is used to process the long-range feature information output by the first fusion coding module through shallow network convolutional layers and deep network convolutional layers to obtain feature information between multiple adjacent regions; the first pixel density augmentation module is used to perform pixel augmentation processing on the feature information between multiple adjacent regions output by the first high-density coding module to obtain feature information that increases the receptive field; the first noise removal module is used to suppress and fill noise features in the feature information that increases the receptive field output by the first pixel density augmentation module, and to enhance and fill important features in the feature information that increases the receptive field to obtain enhanced feature information.

[0049] like Figure 2 The left half of the image represents the upsampling model. In its overall structure, this model follows a process where, during upsampling, a fusion encoding module first reduces the complexity of feature computation, facilitating the restoration of texture details in the image to be repaired. Simultaneously, during upsampling, a high-density encoding module utilizes a deep network to obtain more feature information between adjacent regions from the processed low-feature channels. Subsequently, a pixel density augmentation module is used to augment pixels, expanding the receptive field without losing resolution and maintaining the relative spatial positions of pixels. This data is then input into a noise removal module. Under multi-feature channel input conditions, noise features are suppressed and filled in, while important features are enhanced and filled, resulting in a clearer and more natural image.

[0050] In one possible implementation of this application, the first fusion coding module includes multiple convolutional layers and average pooling layers.

[0051] The first fusion encoding module is used to: input the original image into two convolutional layers to obtain the first image feature; input the first image feature into the average pooling layer and the max pooling module respectively for convolution processing to obtain the second image feature and the third image feature; perform a Hadamard product operation on the second image feature and the third image feature, and perform a matrix product with the first image feature to obtain the fourth image feature; and use residual connections to fuse the fourth image feature with the original image to obtain long-range feature information.

[0052] In one instance, such as Figure 3 As shown.

[0053] (1) The input original image is first passed through two 3×3 convolutions to increase the receptive field of the encoder, which is beneficial for the encoder to extract rich multi-scale detail information in the damaged image. The receptive field of the second 3×3 convolution in the two 3×3 convolutions is equivalent to that of a 5×5 convolution, which is beneficial for saving model memory.

[0054] F1=Conv 3×3 (Conv 3×3 (F in ))

[0055] (2) The convolutional image features F1 are input into the average pooling layer for average pooling processing, which can produce a smoother feature map. The pooled content is then input into the channel attention module to obtain the output F2.

[0056] F2=σ(δ(F fc (F ap (F1))))

[0057] Among them, F fc It is a fully connected convolution operation, F ap It is an average pooling operation, δ is the PReLU function, and σ is the Sigmoid activation function.

[0058] (3) At the same time, the image features F1 after convolution are input into the max pooling operation module to perform max pooling operation. The max pooled content is input into a 7×7 convolution to perform convolution processing, and finally the output F3 is obtained.

[0059] F3=σ((F fc (Max(F1)))

[0060] Here, Max is the operation for selecting the maximum value.

[0061] (4) Perform a Hadamard product operation on F2 and F3, and then perform a matrix multiplication operation with feature F1 to obtain the output weighted attention mechanism feature map of F4;

[0062] F4 = F1 × (F2 ⊙ F3)

[0063] Where ⊙ represents the Hadama product.

[0064] After obtaining the weighted attention mechanism feature map F4, residual connections are used to fuse feature F4 with the input feature Fin to obtain long-range feature information. Finally, a 4×4 convolution with a stride of 2 is applied to obtain the outputs of two fused encoding modules, Fout11 and Fout12, where Fout11 is... Figure 2 The output of the first fusion coding module, Fout12, is Figure 2The output of the second fusion coding module. In this embodiment, using two fusion coding modules can better restore the texture details of the original image.

[0065] In one possible implementation of this application, the first high-density coding module includes a high-density coding layer and a decoupling layer.

[0066] The high-density coding layer includes multiple interconnected convolutional kernels with strides of 1 and 2. The convolutional kernels with strides of 1 and 2 are connected by a gradient dissipation processing module. The decoupling layer includes multiple upsampling modules. The convolutional kernels with strides of 1 in the high-density coding layer are input to the upsampling modules for upsampling processing, and the output is a restored feature map. The restored map includes feature information between multiple adjacent regions.

[0067] In this embodiment, when the image is feature extracted by the convolutional layer of the high-density coding layer, the outputs of the shallow network convolutional layer and the deep network convolutional layer contain image information with different characteristics. When the image is extracted by the shallow network, since the receptive field of the convolutional layer overlaps very little, more detailed image texture information and more image details can be obtained. As the convolutional layer continuously samples the image features, the receptive field of the convolutional layer will gradually increase, and the overlap of the receptive field will also increase. Therefore, the deep network can obtain more feature information between adjacent regions.

[0068] In one instance, such as Figure 4 As shown.

[0069] (1) The input of the first high-density coding module is the output Fout12 of the second fusion coding module;

[0070] (2) The features are upsampled by the convolutional layer of the high-density coding layer. This layer consists of 12 convolutional kernels with strides of 1 and 2. The gradient vanishing processing module is used to connect the convolutional kernels with strides of 1 and 2 to overcome the gradient vanishing problem.

[0071] (3) After the 3rd and 7th convolutional kernels with a stride of 1, the output is fed into the upsampling module through the "gradient dissipation processing module" and then into the convolutional kernel with a stride of 1 for processing. The 9th convolutional kernel with a stride of 1 is directly fed into the upsampling module. In addition, the output of the entire high-density coding layer is also fed into the upsampling module and then into the convolutional kernel with a stride of 1 for processing.

[0072] Among them, after the 3rd and 7th convolutional kernels with a stride of 1, the gradient dissipation processing module is used to input the data into the upsampling module, which can collect better features.

[0073] The 9th convolutional kernel with a stride of 1 is directly input to the upsampling module, which can backtrack the input to make the result more accurate and converge quickly.

[0074] The gradient dissipation processing module takes the data processed by the convolution kernel with a stride of 1 as input. First, the zero-padding module filters the image edges. Then, the data is input into the convolution kernel with a stride of 1. Finally, the data is input into the BN layer for normalization. The processed data is then combined with the data that has only been filtered by the zero-padding module and then convolved with a stride of 1. The outer product operation is then performed, and the result processed by the gradient dissipation processing module is output.

[0075] (4) Finally, the decoupling layer combines three convolutional kernels with a stride of 1 to upsample the image information output by the high-density coding layer based on the obtained input, and restores the image.

[0076] Fout21 is Figure 2 The output of the first high-density coding module, i.e., the high-density coding module of the upsampling model part, Fout22 is... Figure 2 The output of the second high-density coding module, which is the high-density coding module of the downsampling model part.

[0077] In one possible implementation of this application, the first pixel density augmentation module includes multiple feature extraction layers and multiple convolutional layers.

[0078] The first pixel density augmentation module is used to: input the feature map output by the high-density encoding module into multiple feature extraction layers for feature extraction to obtain a first feature map; input the first feature map into multiple convolutional layers for convolution processing to obtain a second feature map; perform an outer product operation between the second feature map and the first feature map to obtain a spatial weight feature map; and perform a summation operation between the spatial weight feature map and the feature map output by the high-density encoding module, followed by activation function processing to obtain a feature map with an increased receptive field, wherein the feature map with an increased receptive field has more pixels than the feature map output by the high-density encoding module.

[0079] In this embodiment, the pixel amplification module takes the output Fout21 of the high-density encoding module as its input to perform pixel amplification processing, so as to expand the receptive field without losing resolution and keep the relative spatial position of the pixels unchanged.

[0080] In this embodiment, Figure 2 Taking the first pixel density augmentation module (i.e., the first pixel density augmentation module in the upsampling model) as an example, this will be explained as follows: Figure 5 As shown.

[0081] (1) The output Fout21 of the high-density coding module flows into the feature extraction layer consisting of three different convolutions and ReLU activation functions, and combined with the activation operation, the feature map L1 is obtained;

[0082] The three convolutions in this layer, from left to right, are:

[0083] The first convolutional operation consists of a convolutional layer with a kernel of 1, a stride of 1, padding of 0, and a dilation rate of 1, along with a ReLU activation operation.

[0084] The second convolutional operation consists of a convolutional layer with a kernel size of 3, a stride of 1, padding of 1, and a dilation rate of 1, along with a ReLU activation operation.

[0085] The second convolutional operation consists of a convolutional layer with a kernel of 1, a stride of 1, padding of 0, and a dilation rate of 1, along with a ReLU activation operation.

[0086] (2) Feature map L1 is subjected to 4 more convolutional operations to obtain feature map L2;

[0087] The four convolutions in these four convolutional layers, from left to right, are:

[0088] The first layer of convolution consists of a convolution with a kernel of 1, a stride of 1, padding of 0, and a dilation of 1, and a convolution with a kernel of 3, a stride of 1, padding of 1, and a dilation of 1.

[0089] The second convolutional layer consists of a convolution with a kernel of 1, a stride of 1, padding of 0, and a dilation of 1, and a convolution with a kernel of 3, a stride of 1, padding of 1, and a dilation of 2.

[0090] The third convolutional layer consists of a convolution with a kernel of 1, a stride of 1, padding of 0, and a dilation of 1, and a convolution with a kernel of 3, a stride of 1, padding of 1, and a dilation of 4.

[0091] The fourth layer of convolution consists of a convolution with a kernel of 1, a stride of 1, padding of 0, and a dilation of 1, and a convolution with a kernel of 3, a stride of 1, padding of 1, and a dilation of 8.

[0092] This convolution operation reduces the number of parameters in the image while expanding the receptive field of view. The visual field features extracted by the four convolutional layers are stacked to obtain a four-channel image.

[0093] In other words, the number of channels is doubled by using a convolution with a kernel of 1.

[0094] (3) Perform an outer product operation on the spatial weight feature maps L1 and L2 to obtain the spatial weight feature map L3;

[0095] (4) After the spatial weight feature map L3 is summed and superimposed with the input image, it is then processed by the LeakyReLU activation function to obtain the final outputs of the four pixel augmentation modules, namely Fout31, Fout32, Fout33, and Fout34. Fout31 is... Figure 2 The output of the first pixel density augmentation module, Fout32, is Figure 2 The output of the second pixel density augmentation module, Fout33, is Figure 2 The output of the third pixel density augmentation module, Fout34, is Figure 2 The output of the fourth pixel density amplification module. This embodiment employs multiple pixel density amplification modules to further expand the receptive field.

[0096] In one possible implementation of this application, the first noise removal module includes a convolutional layer, a gated layer, a feature layer, an average pooling layer, and a fully connected layer.

[0097] The first noise removal module is used to: input the feature map output by the first pixel density augmentation module into a convolutional layer for convolution processing to obtain a third feature map; input the third feature map into a gating layer and a feature layer respectively for feature selection to obtain image channel features and spatial location features; multiply the image channel features and spatial location features to obtain a first output feature map; pass the output feature map through an average pooling layer and a fully connected layer, then through an activation function, then through a fully connected layer, and then through an activation function to obtain a second output feature map; and perform an outer product operation on the first output feature map and the second output feature map to obtain an enhanced feature image.

[0098] In this embodiment, the noise removal module mainly suppresses and fills noise features in the image under the input state of multiple feature channels, and strengthens and fills important features, so that the image is clearer and more natural.

[0099] This embodiment uses Figure 2 Taking the first noise removal module (i.e., the first noise removal module in the upsampling model) as an example, the explanation is as follows: Figure 6 As shown.

[0100] (1) The input is the input Fout31 of the pixel amplification module. It first passes through a convolutional layer with a kernel of 2 to perform a convolution operation and obtain the output Fz.

[0101] Fz = Conv(Fout31)

[0102] (2) The inputs are respectively fed into the gated kernel and the feature layer for dynamic feature selection, and the outputs of image channel features and spatial location features are obtained;

[0103] F 图像通道特征 =

[0104] F 空间位置特征 =

[0105] Where Wg and Wf are filters, Fz is the output after convolution, N is the number of feature layers, i is the feature of the i-th layer, and j is the feature of the j-th layer.

[0106] (3) Multiply the image channel features and spatial location features to obtain the corresponding output V;

[0107] V= (F) 图像通道特征 )⊙σ(F 空间位置特征 )

[0108] in, σ is the exponential linear unit activation function, and σ is the sigmoid activation function.

[0109] (4) The output V of the image channel and spatial location features is input to the average pooling layer for average pooling, and then input to the fully connected layer for fully connected operation. Then, ReLU is used for activation, and the activated features are input to the second fully connected layer, followed by Sigmoid activation to obtain the output of this branch.

[0110] (5) The output of the image channel and spatial location features is multiplied by the output in (3) to obtain the output of the noise removal module. The outputs of the two noise removal modules are Fout41 and Fout42, respectively.

[0111] In one possible implementation of this application, the downsampling model includes a second noise removal module, a second pixel density augmentation module, a second high-density encoding module, and a first fusion decoding module connected in sequence.

[0112] The first fusion decoding module is used to perform deconvolution operation on the feature information output by the second high-density encoding module to obtain the decoded image of the enhanced feature image.

[0113] This embodiment uses a fusion decoding module to perform deconvolution operations to decode the image.

[0114] It should be noted that the structures of the second noise removal module, the second pixel density augmentation module, and the second high-density encoding module are the same as those of the corresponding modules in the upsampling model. This embodiment will not repeat them here, and the above embodiment shall prevail.

[0115] In one possible implementation of this application, the feature information after pixel amplification is deconvolved to obtain a target image, including: processing the feature information after pixel amplification through shallow network convolutional layers and deep network convolutional layers in a second high-density coding module to obtain feature information between multiple adjacent regions; performing deconvolution processing on the feature information between multiple adjacent regions through a main decoding module and multiple sub-decoding modules in a first fusion decoding module, and then performing channel superposition processing to obtain the target image, wherein the number of network layers in the main decoding module is the same as the number of network layers in the first fusion coding module during the sampling process, the multiple sub-decoding modules are used to decode features of different scales to the same size as the target features, and channel superposition is the superposition of the channels of the features decoded by the main decoding module and the features decoded by the multiple sub-decoding modules.

[0116] That is, the first fusion decoding module includes a main decoding module and multiple sub-decoding modules, and the number of network layers in the main decoding module is the same as the number of network layers in the first fusion coding module in the upsampling model.

[0117] The first fusion decoding module is used to: perform deconvolution processing on features of different scales in the feature information output by the second high-density encoding module, and then perform channel superposition processing to obtain a decoded image of the enhanced feature image.

[0118] In one example, the fusion decoding module mainly includes a main decoding module and multiple sub-decoding modules, all of which are superpositions of deconvolutional networks.

[0119] The main decoding module has the same number of network layers as the fusion coding module. Its main function is to recover the feature maps corresponding to each scale layer by layer. It has L decoding layers. For this main decoding module, the feature maps of each layer correspond to the encoding parts, and are represented as follows: , where d is the number of gated kernel feature layers.

[0120] The first layer feature mapping of the main decoding module The formation depends only on the last layer in the encoder. :

[0121]

[0122] For the decoding layer number features of the Lth layer in the main decoding module, the aggregation process is as follows: first, the sub-decoding module aggregates the features of different scales in the encoded part. Decoding to the target at the current scale With the same size, the decoded feature fd( ),...,fd( ) and the feature fd decoded by the main decoding module ( The final aggregated target is obtained by combining the main and auxiliary decoding modules. , is represented as:

[0123] = ⊕fd( ... fd ( )⊕ fd( )= ⊕fd( ,..., , )

[0124] Where ⊕ represents the channel-dimensional stacking operation, fd(·) represents the deconvolution operation performed on features of different scales, and e is the total number of feature layers, i.e., the decoding process.

[0125] The outputs of the two fusion decoding modules are Fout51 and Fout52, respectively.

[0126] To further extract more effective low-level feature information from the backbone network and recover the true details and edge information of the image, this paper proposes a learning network. In this network, the information from the left branch of the high-density coding module, pixel augmentation module, and noise removal module is extracted and learned, and then output to the corresponding right-side module. By utilizing the gating module and feature distillation fusion module in the learning network, the number of parameters is reduced while maintaining the network reconstruction accuracy. In addition, to address the problem that the backbone network is not efficient enough for channel segmentation in this scheme, the feature distillation fusion module couples the information from the original input module together.

[0127] In one possible embodiment of this application, the image restoration method may further include: inputting the feature information from the first high-density encoding module and the first pixel density amplification module in the sampling process into a learning network for processing, and then inputting it into the second pixel density amplification module and the second high-density encoding module in the deconvolution process, wherein the learning network is used to condense and extract the input feature information and input the learned feature information into subsequent modules.

[0128] In one possible implementation of this application, the learning network includes at least one gating module and at least one feature distillation fusion module.

[0129] The gating module includes two convolutional layers. Feature information from the first high-density encoding module, the first pixel density augmentation module, and the first noise removal module is input into the two convolutional layers for convolution processing. The processed feature information is then activated, and the activated feature information is subjected to outer product processing to obtain gated features. The feature distillation and fusion module includes multiple convolutional layers with different kernels. Feature information from the first high-density encoding module, the first pixel density augmentation module, and the first noise removal module is input into the multiple convolutional layers with different kernels of the feature distillation and fusion module for convolution processing to obtain feature distillation and fusion features. The gating features and the feature distillation and fusion features are then processed by channel reduction through convolutional layers to obtain the learned feature information.

[0130] In one instance, such as Figure 7 As shown.

[0131] When the first high-density encoding module, the first pixel augmentation module, and the first noise removal module are used as inputs to feed image features into the learning network, the image features flow into the gate control module and the feature distillation and fusion module, respectively. The feature data flowing into the gate control module undergoes two 3x3 convolutions, activated by LReLU and Sigmoid respectively, followed by outer product processing. The feature data flowing into the feature distillation and fusion module first undergoes one 3x3 convolution, then the left branch uses three 1x1 convolutions for channel reduction, and the right branch uses four 3x3 convolutions for further refinement of coarse features. The right branch extracts the main part of the deep features, the left branch uses 1×1 convolution to separate 25% of the channels and continue to pass them on, and the right branch uses 3×3 convolution to retain 75% of the channels and continue to extract detailed features. The features processed by the left and right branches are then processed by the channel spatial attention mechanism and summed and superimposed with the input processed by 3*3 convolution. The feature data processed by the gating module and the feature distillation fusion module are then processed by 1*1 convolution to reduce the channels and are then input into the gating module and the feature distillation fusion module respectively for layered extraction. The process of (2) to (4) is repeated until the preset output result is met, and the output of the learning network is obtained.

[0132] After integrating modules such as fusion encoding, high-density encoding, pixel augmentation, noise removal, and fusion decoding, the output image improves the quality of the original image to a certain extent by removing noise and masking. However, during the restoration and enhancement process, there may be a situation where unsatisfactory restored images are mixed with well-restored images.

[0133] In one possible implementation of this application, after the gated features and feature distillation fusion features are processed by channel reduction through a convolutional layer, the image restoration method may further include: inputting the channel-reduced gated features and feature distillation fusion features into the gated module and the feature distillation fusion module respectively for overlay processing.

[0134] In this embodiment, images that fail to meet repair standards are removed, and repair information that can be used for iteration is output to the learning network for subsequent model iteration and optimization.

[0135] If there are relatively few features to be repaired in the damaged area of ​​the image, such as the color ring around the pigeon's eyes or the outline of the number edge, the aforementioned feature information is easily lost during downsampling. In this embodiment, image features, semantic information, and texture information at different scales are extracted and judged at the output layer to obtain an image that meets the repair and beautification standards and can be output.

[0136] In one possible implementation of this application, before deleting pixel blocks in the obtained target image whose repair degree is less than the repair degree threshold, the image repair method may further include: inputting the target image into a discriminant network for evaluation, wherein the discriminant network is used to extract features from the input image to obtain image features, semantic information and texture information at different scales.

[0137] In one possible implementation of this application, the discriminant network includes a first branch, a second branch, and a third branch, wherein the kernel size and number of layers of the first branch, the second branch, and the third branch are all different.

[0138] The discriminant network is used to: input the target image into the first branch, the second branch, and the third branch for feature extraction, and obtain features of the same size but different scales; fuse the features extracted by the first branch, the second branch, and the third branch to obtain fused features; and obtain the repair degree of multiple local regions based on the fused features.

[0139] In one instance, such as Figure 8 As shown.

[0140] The discriminator consists of three branches with different kernel sizes and number of layers: shallow branch 1, medium branch 2, and deep branch 3.

[0141] (1) Shallow branch 1: It consists of three convolutional layers with a total of 7 kernels stacked together, which is used to enhance the ability to extract geometric and texture information from the input image;

[0142] (2) Middle Branch 2: It consists of 4 convolutional layers with 5 kernels each, which enhances the ability to extract semantic information from the input image and improves the network's ability to extract information between adjacent regions of the receptive field.

[0143] (3) Deep branch 3: It consists of 5 convolutional layers with 5 kernels. This layer has more convolutional layers and fewer convolutional kernels, which can improve the stability of network training and ensure more comprehensive feature extraction of input images.

[0144] During the discrimination process, the discriminator distinguishes features extracted from the three branches that are of the same size but different scales.

[0145] After fusion, convolutional layers are used to evaluate multiple local regions of the features to improve the ability to distinguish image details, thereby outputting a high-quality image.

[0146] In this embodiment, the overall structure follows a process where, during upsampling, the complexity of computational features is first reduced by a fusion mode encoding module, which is beneficial for restoring texture details in the image to be repaired. Simultaneously, during upsampling, a high-density encoding module is used to obtain more feature information between adjacent regions from the processed low-feature channels using a deep network. Subsequently, a pixel density amplification module is used to amplify pixels, thereby expanding the receptive field without losing resolution and maintaining the relative spatial position of pixels. Then, the image is input to a noise removal module. Under the input state of multiple feature channels, noise features are suppressed and filled, and important features are enhanced and filled, making the image clearer and more natural.

[0147] The outputs of the high-density coding module, pixel augmentation module, and noise removal module in the backbone network are learned by using a learning network. The content of the left branch is input into the right branch for learning, so as to improve the output performance and accuracy of the right branch.

[0148] The final discriminator network adjusts the output by distinguishing between high-quality and low-quality restored images, in order to output the image of optimal quality.

[0149] In one instance, such as Figure 9-10 As shown, the user selects the toolbox in the cloud drive and clicks the "Image Enhancement" function; then the user imports the image to be enhanced and selects the enhancement style; if the user selects enhancement repair, the image repair method provided in this application will be used to repair the image for the user.

[0150] In one example, the enhancement and repair function is a new image processing feature to be launched in the cloud storage product. This function can improve image resolution and repair issues such as occlusion, white spots, dark spots, and blurring caused by lens smudges in early photos, ultimately improving image detail and quality. Figure 11As shown, the left side is the original image, and the right side is the restored image. The image is clearer and the colors are brighter.

[0151] like Figure 12 The image shown is a schematic diagram of an image restoration device provided in this application. Figure 12 As shown, the image restoration device can be applied to electronic devices. The image restoration device may include: an acquisition module 1201, a sampling module 1202, a processing module 1203, and a determination module 1204.

[0152] The system comprises the following modules: an acquisition module 1201, used to acquire the original image to be processed; a sampling module 1202, used to sample the original image, using a deep network to obtain feature information between adjacent regions in the original image during the sampling process, and performing pixel amplification on the feature information, wherein the relative spatial position of each pixel remains unchanged before and after amplification; a processing module 1203, used to perform deconvolution processing on the feature information after pixel amplification to obtain a target image, wherein the target image includes multiple pixel blocks; and a determination module 1204, used to delete pixel blocks in the obtained target image whose repair degree is less than the repair degree threshold, and to synthesize pixel blocks whose repair degree is greater than or equal to the repair degree threshold to obtain a repaired image.

[0153] In this embodiment, the acquisition module 1201 first acquires the original image to be processed. Then, the sampling module 1202 samples the original image, using a deep network to obtain feature information between adjacent regions in the original image. This feature information is then amplified using pixels, with the relative spatial positions of the pixels remaining unchanged before and after amplification. The processing module 1203 then performs deconvolution on the amplified feature information to obtain the target image, which includes multiple pixel blocks. Finally, the determination module 1204 deletes pixel blocks in the target image with a repair degree less than a repair degree threshold and synthesizes pixel blocks with a repair degree greater than or equal to the repair degree threshold to obtain the repaired image. This embodiment, by sampling the original image to obtain feature information between adjacent regions and amplifying this feature information using pixels, can increase the receptive field of the original image, thereby enhancing the image's beauty while improving its resolution.

[0154] In one possible embodiment of this application, the sampling module 1202 is used to sample the original image. During the sampling process, the original image is passed through the convolution operation of the first fusion coding module to obtain long-range feature information of the original image. The computational complexity of the long-range feature information is lower than that of the feature information in the original image. The long-range feature information is processed through the shallow network convolutional layer and the deep network convolutional layer in the first high-density coding module to obtain feature information between multiple adjacent regions. The feature information between multiple adjacent regions is processed through the first pixel density amplification module to obtain feature information that increases the receptive field.

[0155] In one possible implementation of this application, the first fusion encoding module includes multiple convolutional layers and average pooling layers. The first fusion encoding module is used to: input the original image into two convolutional layers to obtain a first image feature; input the first image feature into an average pooling layer and a max pooling module respectively for convolution processing to obtain a second image feature and a third image feature; perform a Hadamard product operation on the second and third image features, and then perform a matrix product with the first image feature to obtain a fourth image feature; and use residual connections to fuse the fourth image feature with the original image to obtain long-range feature information.

[0156] In one possible implementation of this application, the first high-density coding module includes a high-density coding layer and a decoupling layer. The high-density coding layer includes multiple interconnected convolutional kernels with strides of 1 and 2, wherein the convolutional kernels with strides of 1 and 2 are connected through a gradient dissipation processing module. The decoupling layer includes multiple upsampling modules. The convolutional kernels with strides of 1 in the high-density coding layer are input to the upsampling modules for upsampling processing, and the output is a restored feature map. The restored map includes feature information between multiple adjacent regions.

[0157] In one possible implementation of this application, the first pixel density augmentation module includes multiple feature extraction layers and multiple convolutional layers. The first pixel density augmentation module is used to: input the feature map output by the high-density encoding module into multiple feature extraction layers for feature extraction to obtain a first feature map; input the first feature map into multiple convolutional layers for convolution processing to obtain a second feature map; perform an outer product operation between the second feature map and the first feature map to obtain a spatial weight feature map; and perform a summation operation between the spatial weight feature map and the feature map output by the high-density encoding module, followed by activation function processing, to obtain a feature map with an increased receptive field. The feature map with the increased receptive field has more pixels than the feature map output by the high-density encoding module.

[0158] In one possible embodiment of this application, the processing module 1203 is used to process the feature information of the amplified pixels through the shallow network convolutional layer and the deep network convolutional layer in the second high-density coding module to obtain feature information between multiple adjacent regions; the feature information between multiple adjacent regions is deconvolutionally processed through the main decoding module and multiple sub-decoding modules in the first fusion decoding module, and then processed by channel superposition to obtain the target image. The number of network layers in the main decoding module is the same as the number of network layers in the first fusion coding module in the sampling process. The multiple sub-decoding modules are used to decode features of different scales to the same size as the target features. Channel superposition is the channel superposition of the features decoded by the main decoding module and the features decoded by the multiple sub-decoding modules.

[0159] In one possible embodiment of this application, the sampling module 1202 is used to: input the feature information from the first high-density encoding module and the first pixel density amplification module in the sampling process into the learning network for processing, and then input it into the second pixel density amplification module and the second high-density encoding module in the deconvolution process, wherein the learning network is used to condense and extract the input feature information and input the learned feature information into the subsequent modules.

[0160] In one possible implementation of this application, the learning network includes at least one gating module and at least one feature distillation and fusion module. The gating module includes two convolutional layers. Feature information from the first high-density encoding module, the first pixel density augmentation module, and the first noise removal module are respectively input into the two convolutional layers for convolution processing. The processed feature information is then activated, and the activated feature information is subjected to outer product processing to obtain gated features. The feature distillation and fusion module includes multiple convolutional layers with different kernels. Feature information from the first high-density encoding module, the first pixel density augmentation module, and the first noise removal module is input into the multiple convolutional layers with different kernels of the feature distillation and fusion module for convolution processing to obtain feature distillation and fusion features. The gating features and the feature distillation and fusion features are then processed by channel reduction through the convolutional layers to obtain the learned feature information.

[0161] In one possible embodiment of this application, the image restoration apparatus may further include an evaluation module.

[0162] The evaluation module is used to input the target image into the discriminant network for evaluation. The discriminant network is used to extract features from the input image to obtain image features, semantic information and texture information at different scales.

[0163] In one possible implementation of this application, the discriminant network includes a first branch, a second branch, and a third branch, wherein the kernel size and number of layers of the first branch, the second branch, and the third branch are all different.

[0164] The evaluation module is used to: input the target image into the first branch, the second branch, and the third branch for feature extraction, and obtain features of the same size but different scales; fuse the features extracted from the first branch, the second branch, and the third branch to obtain fused features; and obtain the repair degree of multiple local regions based on the fused features.

[0165] The image restoration device provided in this application embodiment can achieve... Figures 1 to 11 The various processes implemented in the method embodiments achieve the same technical effect, and will not be described again here to avoid repetition.

[0166] like Figure 13 As shown, this application embodiment also provides an electronic device 1300, including a processor 1301, a memory 1302, and a program or instructions stored in the memory 1302 and executable on the processor 1301. When the program or instructions are executed by the processor 1301, they implement the various processes of the above-described image restoration method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0167] Optionally, embodiments of this application also provide a computer-readable storage medium storing a computer program. When executed by a processor, this computer program implements the various processes of the above-described image restoration method embodiments and achieves the same technical effect. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.

[0168] Optionally, this application also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions, which, when executed by a computer, implement the various processes of the above-described image restoration method embodiments and achieve the same technical effects. To avoid repetition, they will not be described again here.

[0169] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0170] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0171] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. An image restoration method, characterized in that, include: Obtain the original image to be processed; The original image is sampled. During the sampling process, a deep network is used to obtain the feature information between adjacent regions in the original image. The feature information is then amplified into pixels, wherein the relative spatial position of each pixel remains unchanged before and after amplification. The feature information after pixel amplification is deconvolved to obtain the target image, wherein the target image includes multiple pixel blocks; Pixel blocks with a repair degree less than the repair degree threshold in the obtained target image are deleted, and pixel blocks with a repair degree greater than or equal to the repair degree threshold are synthesized to obtain the repaired image; The sampling process of the original image, which involves using a deep network to obtain feature information between adjacent regions in the original image and then performing pixel augmentation on the feature information, includes: The original image is sampled. During the sampling process, the original image is convolved by the first fusion coding module to obtain the long-range feature information of the original image. The computational complexity of the long-range feature information is lower than that of the feature information in the original image. The long-distance feature information is processed through shallow network convolutional layers and deep network convolutional layers in the first high-density coding module to obtain feature information between multiple adjacent regions; The feature information between the multiple adjacent regions is processed by the first pixel density amplification module to obtain feature information that increases the receptive field.

2. The method according to claim 1, characterized in that, The first fusion coding module includes multiple convolutional layers and average pooling layers, wherein the first fusion coding module is used for: The original image is input into two convolutional layers to obtain the first image features; The first image features are respectively input into the average pooling layer and the max pooling module for convolution processing to obtain the second image features and the third image features; Perform a Hadamard product operation on the second image feature and the third image feature, and then perform a matrix product with the first image feature to obtain the fourth image feature; The fourth image features are fused with the original image using residual connections to obtain the long-range feature information.

3. The method according to claim 1, characterized in that, The first high-density coding module includes a high-density coding layer and a decoupling layer. The high-density coding layer includes multiple interconnected convolutional kernels with strides of 1 and 2, wherein the convolutional kernels with stride of 1 and stride of 2 are connected by a gradient dissipation processing module. The decoupling layer includes multiple upsampling modules. The convolutional kernel with a stride of 1 in the high-density coding layer is input to the upsampling module for upsampling processing, and the output is a restored feature map. The restored feature map includes feature information between multiple adjacent regions.

4. The method according to claim 1, characterized in that, The first pixel density augmentation module includes multiple feature extraction layers and multiple convolutional layers, wherein the first pixel density augmentation module is used for: The feature map output by the high-density coding module is input into the multiple feature extraction layers for feature extraction to obtain the first feature map; The first feature map is input into the plurality of convolutional layers for convolution processing to obtain the second feature map; The second feature map is combined with the first feature map by an outer product operation to obtain the spatial weight feature map. After summing and superimposing the spatial weighted feature map with the feature map output by the high-density coding module, and then processing it through an activation function, a feature map with an increased receptive field is obtained. The feature map with an increased receptive field has more pixels than the feature map output by the high-density coding module.

5. The method according to claim 1, characterized in that, The step of performing deconvolution processing on the feature information after pixel amplification to obtain the target image includes: The feature information after pixel amplification is processed by the shallow network convolutional layer and the deep network convolutional layer in the second high-density coding module to obtain the feature information between multiple adjacent regions. The feature information between the multiple adjacent regions is deconvolutionally processed by the main decoding module and multiple sub-decoding modules in the first fusion decoding module, and then processed by channel superposition to obtain the target image. The number of network layers in the main decoding module is the same as the number of network layers in the first fusion encoding module in the sampling process. The multiple sub-decoding modules are used to decode features of different scales to the same size as the target features. The channel superposition is the superposition of the channels of the features decoded by the main decoding module and the features decoded by the multiple sub-decoding modules.

6. The method according to claim 5, characterized in that, The method further includes: The feature information from the first high-density encoding module, the first pixel density augmentation module, and the first noise removal module during the sampling process is input into the learning network for processing, and then input into the second pixel density augmentation module and the second high-density encoding module during the deconvolution process. The learning network is used to condense and extract the input feature information and input the learned feature information into subsequent modules. The first noise removal module is used to suppress and fill the noise features in the feature information of the increased receptive field output by the first pixel density augmentation module.

7. The method according to claim 6, characterized in that, The learning network includes at least one gating module and at least one feature distillation fusion module. The gating module includes two convolutional layers. The feature information from the first high-density encoding module, the first pixel density augmentation module, and the first noise removal module is respectively input into the two convolutional layers for convolution processing. The processed feature information is then activated, and the activated feature information is then subjected to outer product processing to obtain the gating feature. The feature distillation fusion module includes multiple convolutional layers with different convolutional kernels. The feature information from the first high-density encoding module, the first pixel density augmentation module, and the first noise removal module is input into the multiple convolutional layers with different convolutional kernels of the feature distillation fusion module for convolution processing to obtain the feature distillation fusion feature. The gated features and the feature distillation fusion features are processed by channel reduction in a convolutional layer to obtain the learned feature information.

8. The method according to claim 1, characterized in that, Before deleting pixel blocks with a repair score less than the repair score threshold in the obtained target image, the method further includes: The target image is input into a discriminant network for evaluation. The discriminant network is used to extract features from the input image to obtain image features, semantic information, and texture information at different scales.

9. The method according to claim 8, characterized in that, The discriminant network includes a first branch, a second branch, and a third branch. The convolutional kernel size and number of layers of the first branch, the second branch, and the third branch are all different. The discriminant network is used for: The target image is input into the first branch, the second branch, and the third branch for feature extraction to obtain features of the same size but different scales. The features extracted from the first branch, the second branch, and the third branch are fused to obtain fused features; Based on the fusion features, the repair degree of multiple local regions is obtained.

10. An electronic device, characterized in that, It includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the method as claimed in any one of claims 1 to 9.

11. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, which, when executed by a processor, implements the steps of the method as described in any one of claims 1 to 9.

12. A computer program product, characterized in that, The computer program product includes a computer program stored on a non-transitory computer-readable storage medium, the computer program including program instructions that, when executed by a computer, implement the steps of the method as described in any one of claims 1 to 9.