Training Method for Facial Image Restoration Model

By combining the first feature extraction network, the second feature extraction network and the generation network in the face image restoration model, using training samples and face key point sample images, the problem of difficult to achieve high-quality restoration based on only a single prior information in the prior art is solved, and high-quality face image restoration is achieved.

CN119784634BActive Publication Date: 2025-06-13HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510273701.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-13
Estimated Expiration
2045-03-10

AI Technical Summary

Technical Problem

Existing facial image restoration models are usually based on only one prior information, making it difficult to achieve high-quality facial image restoration.

Method used

A face image restoration model including a first feature extraction network, a second feature extraction network and a generation network is adopted. By obtaining the training samples and the corresponding face key point sample image, inputting them into the initial model to trigger the generation network's constraint-based face identity features and key point features, predicting the restored sample image, and training the model based on the restored sample image and reference image.

Benefits of technology

It achieves a high-quality restoration effect while keeping the facial identity characteristics unchanged. By capturing the location of the key points on the face, focusing on the restoration effect of the key points on the face, improving the quality of facial image restoration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119784634B_ABST
    Figure CN119784634B_ABST
Patent Text Reader

Abstract

The present application is applicable to the field of image processing technology, and provides a training method for a facial image restoration model. The training method includes: obtaining training samples and corresponding facial key point sample images; inputting the facial sample images and the facial key point sample images into an initial facial image restoration model, so as to trigger a generation network to predict a restored sample image corresponding to the facial sample image based on the facial identity features of the facial sample image constrained by a first feature extraction network and the facial key point features of the facial key point sample image constrained by a second feature extraction network; training the initial facial image restoration model based on the restored sample image and a restored reference image to obtain a target facial image restoration model, and the target facial image restoration model is used to perform facial image restoration on a to-be-restored facial image. High-quality facial image restoration can be achieved through the present application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of image processing, and particularly relates to a method for training a face image restoration model. Background Art

[0002] Face image restoration aims to restore a high-quality face image from a low-quality face image, and usually uses a face image restoration model for face image restoration. However, there are still some limitations in the existing face image restoration models. For example, the existing face image restoration models usually perform face image restoration only based on a priori information, and it is difficult to achieve high-quality face image restoration. Summary of the Invention

[0003] An embodiment of this application provides a method for training a face image restoration model, which can achieve high-quality face image restoration.

[0004] In a first aspect, an embodiment of this application provides a method for training a face image restoration model. The initial face image restoration model includes: a first feature extraction network, a second feature extraction network respectively connected to the first feature extraction network, and a generation network. The training method includes:

[0005] Obtain a training sample and a face key point sample image corresponding to the training sample. The training sample includes a face sample image and a restoration reference image corresponding to the face sample image, and the image quality of the restoration reference image is higher than that of the face sample image;

[0006] Input the face sample image and the face key point sample image into the initial face image restoration model to trigger the generation network to predict a restoration sample image corresponding to the face sample image based on the face identity feature of the face sample image constrained by the first feature extraction network and the face key point feature of the face key point sample image constrained by the second feature extraction network;

[0007] Train the initial face image restoration model based on the restoration sample image and the restoration reference image to obtain a target face image restoration model, and the target face image restoration model is used to perform face image restoration on a face image to be restored.

[0008] In the embodiments of the present application, the initial facial image restoration model includes a first feature extraction network, a second feature extraction network respectively connected to the first feature extraction network, and a generation network. By obtaining training samples and facial key point sample images corresponding to the training samples, and inputting the facial sample images and facial key point sample images in the training samples into the initial facial image restoration model, the generation network can be triggered to predict a restored sample image corresponding to the facial sample image based on the facial identity features of the facial sample image constrained by the first feature extraction network and the facial key point features of the facial key point sample image constrained by the second feature extraction network. Based on the restored sample image and the restored reference image in the training sample, the initial facial image restoration model can be trained to obtain a target facial image restoration model. In this solution, through the facial identity features of the facial sample image constrained by the first feature extraction network, the generation network can provide a high-quality restoration effect while keeping the facial identity features unchanged. Through the facial key point features constrained by the second feature extraction network, the generation network can be assisted to capture the positions of facial key points and focus on the restoration effect of facial key points. Therefore, when the generation network trains the initial facial image restoration model based on the facial identity features constrained by the extraction network and the facial key point features constrained by the second feature extraction network, the target facial image restoration model can provide a high-quality restoration effect when performing facial image restoration on the facial image to be restored. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0010] Figure 1 is a schematic flowchart of the method for training a facial image restoration model provided by an embodiment of the present application;

[0011] Figure 2 is a structural example diagram of the initial facial image restoration model provided by an embodiment of the present application;

[0012] Figure 3 is a structural example diagram of the downsampling residual block provided by an embodiment of the present application;

[0013] Figure 4 is a structural example diagram of the upsampling residual block provided by an embodiment of the present application;

[0014] Figure 5 is a structural example diagram of the style convolutional network provided by an embodiment of the present application;

[0015] Figure 6 It is a schematic structural diagram of a condition control network provided by an embodiment of the present application;

[0016] Figure 7 It is a schematic structural diagram of a fusion module provided by an embodiment of the present application;

[0017] Figure 8 It is another schematic structural diagram of a style convolutional network provided by an embodiment of the present application;

[0018] Figure 9 It is a schematic flowchart of a method for generating training samples provided by an embodiment of the present application;

[0019] Figure 10-1 It is a schematic diagram of a face image;

[0020] Figure 10-2 It is a schematic diagram of a standard template;

[0021] Figure 10-3 It is a schematic diagram of an aligned face image;

[0022] Figure 10-4 It is a schematic diagram of a face sample image;

[0023] Figure 10-5 It is a schematic diagram of a restored reference image;

[0024] Figure 11 It is a schematic flowchart of a face image restoration method provided by an embodiment of the present application;

[0025] Figure 12 It is a schematic structural diagram of a training device for a face image restoration model provided by an embodiment of the present application;

[0026] Figure 13 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0027] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are set forth in order to provide a thorough understanding of the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0028] It should be understood that when used in the specification of the present application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0029] It should be noted that the information collection process (such as the face image collection process) / feature extraction process involved in this application is executed with the user's knowledge and permission. That is, the information collection process / feature extraction process complies with the requirements of laws and regulations and does not belong to acts that harm the public interest.

[0030] Existing face image restoration models usually include prior networks for each part of the human face. Given a low-quality face image, three part sub-images are obtained, which are respectively mapped to the high-quality feature z spaces corresponding to each part through their respective encoders. Subsequently, the restored part results are obtained through the part prior networks, and the restored parts are restored to the overall face image to obtain the final restored result. The problem with this method is that it is not an end-to-end method. At the same time, because it needs to combine the prior results of different parts and depends on the accuracy of part detection, when there are accuracy problems, abnormalities are likely to occur.

[0031] The face image restoration model provided by the embodiments of this application includes a first feature extraction network, a second feature extraction network and a generation network respectively connected to the first feature extraction network. The face identity features of the face sample image constrained by the first feature extraction network can enable the generation network to provide a high-quality restoration effect while maintaining the face identity features unchanged. The face key point features constrained by the second feature extraction network can assist the generation network in capturing the positions of face key points and focus on the restoration effect of face key points. Therefore, the generation network trains the initial face image restoration model based on the face identity features constrained by the extraction network and the face key point features constrained by the second feature extraction network, which can enable the target face image restoration model to provide a high-quality restoration effect when performing face image restoration on the face image to be restored, without relying on the detection of different human face parts. Moreover, the face image restoration model provided by the embodiments of this application is an end-to-end model, which can automate feature learning, reduce manual intervention, simplify model design, improve the overall performance of the system, and has better generalization ability.

[0032] Existing solutions usually need to fuse the three-dimensional shape information or two-dimensional segmentation map information of the image (that is, classifying each pixel point in the image into a certain part). The three-dimensional shape information and two-dimensional segmentation map information have a relatively large impact on the final restoration result, that is, the final result is relatively dependent on the prediction accuracy of the three-dimensional shape information and two-dimensional segmentation map information, and it is relatively easy to have problems with inaccurate results due to prediction errors. The feature point map (such as face identity features and face key point features) in the embodiments of this application is a relatively rough geometric constraint and there will be a certain degree of deviation of the feature points in actual training, making the model more robust to prediction errors. Moreover, the face image restoration model provided by the embodiments of this application adopts a learnable latent coding form to provide high-quality image features and can implement the model on embedded devices.

[0033] The training method of the facial image restoration model provided by the embodiments of the present application can be applied to electronic devices such as mobile phones, tablet computers, laptop computers, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), servers, imaging devices (such as cameras, video cameras, etc.). The embodiments of the present application do not impose any restrictions on the specific types of electronic devices.

[0034] Please refer to Figure 1 , Figure 1 which shows a schematic flowchart of the training method of the facial image restoration model provided by the embodiments of the present application. The initial facial image restoration model to be trained includes a first feature extraction network, a second feature extraction network respectively connected to the first feature extraction network, and a generation network. As an example rather than a limitation, this method is applied to an electronic device, and the method includes the following steps:

[0035] Step 101, obtain a training sample and a facial key point sample image corresponding to the training sample.

[0036] Among them, the training sample includes a facial sample image and a restoration reference image corresponding to the facial sample image, and the image quality of the restoration reference image is higher than that of the facial sample image. The facial sample image may refer to low-quality data for training and evaluating the facial image restoration model. The restoration reference image may refer to the accurate label or high-quality data for training and evaluating the facial image restoration model.

[0037] To facilitate the distinction between the facial image restoration model to be trained and the trained facial image restoration model, the facial image restoration model to be trained can be called the initial facial image restoration model, and the trained facial image restoration model (i.e., the trained facial image restoration model) can be called the target facial image restoration model. It should be understood that other names can also be given to the facial image restoration model to be trained and the trained facial image restoration model, and the present application does not limit this.

[0038] To better optimize the model parameters of the initial facial image restoration model, reduce the risk of overfitting, and improve the generalization ability of the model, multiple training samples can be used to train the initial facial image restoration model. Of course, it can be understood that one training sample can also be used to train the initial facial image restoration model, and the present application does not limit this. One training sample includes a facial sample image and a restoration reference image corresponding to the facial sample image.

[0039] It should be noted that the face corresponding to the training sample can be any face with identity features, such as a human face or the face of an animal, etc.

[0040] As an example rather than a limitation, if a face restoration model is used to restore a human face image, then the face corresponding to the training sample is a human face. The face sample images of different training samples can be the face sample images of the same human face or the face sample images of different human faces. Training the initial face image restoration model with the training samples corresponding to the human face can enable the trained target face image restoration model to restore the human face image. If a face restoration model is used to restore the face image of a certain type of animal, then the face corresponding to the training sample is the face of that type of animal. The face sample images of different training samples can be the face sample images of the same animal in that type of animal or the face sample images of different animals in that type of animal. Training the initial face image restoration model with the training samples corresponding to the face of the animal can enable the trained target face image restoration model to restore the animal face image.

[0041] Among them, the above-mentioned training sample and the face key point sample image corresponding to it correspond to the same face and have the same size. The face key points in the face key point sample image can be key points such as the eyes, nose, and mouth of the face. In order to reduce the influence of other regions of the face (i.e., the regions other than the regions where the face key points are located) on model training, the pixel values of the pixels in the regions where the face key points are located in the face key point sample image are 255, and the pixel values of the pixels in the regions other than the regions where the face key points are located in the face key point sample image are 0.

[0042] Step 102: Input the face sample image and the face key point sample image into the initial face image restoration model to trigger the generation network to predict the restored sample image corresponding to the face sample image based on the face identity features of the face sample image constrained by the first feature extraction network and the face key point features of the face key point sample image constrained by the second feature extraction network.

[0043] Among them, the above-mentioned inputting the face sample image and the face key point sample image into the initial face image restoration model can refer to inputting the face sample image into the first feature extraction network and inputting the face key point sample image into the second feature extraction network. Inputting the face sample image into the first feature extraction network can extract and constrain the face identity features of the face sample image, so that the face identity features of the restored sample image are consistent with the face identity features of the face sample image, that is, the face identity features remain unchanged. Inputting the face key point sample image into the second feature extraction network can extract and constrain the face key point features of the face key point sample image, which can assist the generation network to capture the positions of the face key points and focus on the restoration effect of the face key points.

[0044] The generation network restores the initial face image model based on the face identity features constrained by the first feature extraction network and the face key point features constrained by the second feature extraction network. It can restore the face image while keeping the face identity features unchanged, thus providing a high-quality restoration effect. The face key point features constrained by the second feature extraction network can assist the generation network in capturing the positions of face key points and focus on the restoration effect of face key points.

[0045] In a possible implementation manner, the first feature extraction network includes a downsampling residual block, a first fusion module, and an upsampling residual block. The generation network includes a convolutional layer, and the convolutional layer corresponds to a latent code, which is used to adjust the weights of the convolutional layer of the generation network. On this basis, step 102 above includes:

[0046] Input the face sample image into the downsampling residual block to obtain a downsampled feature map; input the face key point sample image into the second feature extraction network to obtain a face key point feature map, and perform attention enhancement on the face key points of the downsampled feature map based on the face key point feature map to obtain an image enhancement feature map. The face key point feature map includes the face key point features of the face key point sample image; input the image enhancement feature map into the first fusion module to trigger the first fusion module to fuse the image enhancement feature map and the downsampled feature map to obtain a fused feature map; input the downsampled feature map and the fused feature map into the upsampling residual block to obtain an upsampled feature map; input the downsampled feature map, the upsampled feature map, and the latent code into the generation network to obtain a restored sample image.

[0047] Among them, the second feature extraction network can multiply the face key point feature map and the downsampled feature map in the face key point sample image to achieve attention enhancement of the face key points of the downsampled feature map. The first fusion module fusing the image enhancement feature map and the downsampled feature map may refer to splicing the image enhancement feature map and the downsampled feature map.

[0048] In another possible implementation manner, the first feature extraction network includes a downsampling residual block, a second fusion module, and an upsampling residual block. The generation network includes a convolutional layer, and the convolutional layer corresponds to a latent code, which is used to adjust the weights of the convolutional layer of the generation network. On this basis, step 102 above includes:

[0049] Input the facial sample image into the downsampling residual block to obtain a downsampled feature map; input the facial key point sample image into the second feature extraction network to obtain a facial key point feature map, where the facial key point feature map includes the facial key point features of the facial key point sample image; input the facial key point feature map into the second fusion module to trigger the second fusion module to enhance the attention of the facial key points on the downsampled feature map based on the facial key point feature map, obtain an image-enhanced feature map, and fuse the image-enhanced feature map and the downsampled feature map to obtain a fused feature map; input the downsampled feature map and the fused feature map into the upsampling residual block to obtain an upsampled feature map; input the downsampled feature map, the upsampled feature map, and the latent code into the generation network to obtain a restored sample image.

[0050] Among them, the second fusion module can multiply the facial key point feature map and the downsampled feature map in the facial key point sample image to enhance the attention of the facial key points on the downsampled feature map. The fusion of the image-enhanced feature map and the downsampled feature map by the second fusion module above can refer to splicing the image-enhanced feature map and the downsampled feature map.

[0051] It should be understood that the above downsampled feature map can refer to the feature map extracted by the downsampling residual block, and this feature map is an image that can include the global features of the facial sample image. The global features of the facial sample image include but are not limited to color features, shape features, facial identity features, etc. The facial identity features can refer to features that can be used for individual identification, including but not limited to texture features, structural features of facial key points, etc. The above upsampled feature map can refer to the feature map extracted by the upsampling residual block.

[0052] As an example rather than a limitation, the structure of the first feature extraction network can be a Unet structure or other network structures, and this application does not limit this. The generation network can be a pre-trained StyleGAN network, which can provide high-quality image features including low-frequency information, medium-frequency information, and high-frequency information, so as to achieve a better image restoration effect. The above generation network can include a style convolutional network and a conditional control network. The convolutional layers included in the generation network can refer to the convolutional layers included in the style convolutional network.

[0053] The latent code provided in this embodiment is a learnable parameter. The optimal latent code (i.e., the trained latent code) can be obtained through model training, which can enable the model to retain a high-quality generation prior and be applied on embedded devices.

[0054] By way of example and not limitation, the first feature extraction network may include M downsampling residual blocks, M-1 first fusion modules or M-1 second fusion modules respectively having the same resolution as the first M-1 downsampling residual blocks, and M-1 upsampling residual blocks corresponding to the first M-1 downsampling residual blocks, where M is an integer greater than 1. On this basis, the second feature extraction network may include M-1 downsampling residual blocks, and the generation network includes M-1 network blocks, each network block including a style convolutional network and a conditional control network. Among them, the resolutions of the M downsampling residual blocks in the first feature extraction network decrease in sequence, the resolutions of the M-1 first fusion modules or M-1 second fusion modules decrease in sequence, the resolutions of the M-1 upsampling residual blocks increase in sequence, and the first M-1 downsampling residual blocks may refer to the downsampling residual blocks with resolutions ranked at the M-1 position.

[0055] By way of example and not limitation, M is 4, the first feature extraction network includes four downsampling residual blocks, three second fusion modules and three upsampling residual blocks. Based on this first feature extraction network, the facial sample image can be downsampled four times, that is, four downsampling residual blocks and three upsampling residual blocks are used. This means that in this embodiment, some weights in the underlying generation priors will be discarded, and only high-scale generation priors are needed, which can control the computational amount of the first feature extraction network.

[0056] As Figure 2 shown is a structural example diagram of the initial facial image restoration model provided by the embodiment of the present application. For the convenience of distinction, A1, A2, A3, A4 are used to represent the four downsampling residual blocks in the first feature extraction network, and the resolutions of these four downsampling residual blocks decrease in sequence. B1, B2, B3 are used to represent the three second fusion modules in the first feature extraction network, C1, C2, C3 are used to represent the three upsampling residual modules in the first feature extraction network, D1, D2, D3 are used to represent the three downsampling residual blocks in the second feature extraction network, E1, E2, E3 are used to represent the three network blocks in the generation network, E11 is used to represent the style convolutional network in E1, E12 is used to represent the conditional control network in E1, E21 is used to represent the style convolutional network in E2, E22 is used to represent the conditional control network in E2, E31 is used to represent the style convolutional network in E3, and E32 is used to represent the conditional control network in E3. The resolutions of A1, B1, C3, D1, E31 and E32 are the same, the resolutions of A2, B2, C2, D2, E21 and E22 are the same, and the resolutions of A3, B3, C1, D3, E31 and E32 are the same. When training the initial facial image restoration model as Figure 2 shown, the facial sample image is input into A1, and the downsampled feature output by A1 Figure 1Input B1 and A2, input the facial key point sample image into D1, and use the facial key point features output by D1 Figure 1 Input B1 and D2, and B1 is based on the facial key point features Figure 1 Perform downsampling on the features Figure 1 Enhance the attention of the facial key points to obtain image enhancement features Figure 1 , and use the image enhancement features Figure 1 and the downsampled features Figure 1 Perform fusion to obtain fused features Figure 1 , and use the fused features Figure 1 Input into C3; A2 processes the input downsampled features Figure 1 and then outputs the downsampled features Figure 1 , and use the downsampled features Figure 1 Input into B2 and A3, and D2 processes the input facial key point features Figure 1 and then outputs the facial key point features Figure 2 , and use the facial key point features Figure 2 Input into B2, and B2 is based on the facial key point features Figure 2 Perform downsampling on the features Figure 2 Enhance the attention of the facial key points to obtain image enhancement features Figure 2 , and use the image enhancement features Figure 2 and the downsampled features Figure 2 Perform fusion to obtain fused features Figure 2 , and use the fused features Figure 2 Input into C2; A3 processes the input downsampled features Figure 2 and then outputs the downsampled features Figure 3 , and use the downsampled features Figure 3 Input into B3 and A4, and D3 processes the input facial key point features Figure 2 and then outputs the facial key point features Figure 3 , and use the facial key point features Figure 3 Input into B3, and B3 is based on the facial key point features Figure 3 Perform downsampling on the features Figure 3 Enhance the attention of the facial key points to obtain image enhancement features Figure 3 , and use the image enhancement features Figure 3 and the downsampled features Figure 3 Perform fusion to obtain fused features Figure 3 , and use the fused features Figure 3 Input into C1; A4 processes the input downsampled features Figure 3 and then outputs the downsampled features Figure 4 , and use the downsampled features Figure 4 Input into C1 and E11; C1 processes the downsampled features Figure 4 and the fused features Figure 3After processing, upsampled features are obtained. Figure 1 , and the upsampled features Figure 1 are input into C2 and E12; C2 processes the upsampled features Figure 1 and the fused features Figure 2 to obtain upsampled features Figure 2 , and the upsampled features Figure 2 are input into C3 and E22; C3 processes the upsampled features Figure 2 and the fused features Figure 1 to obtain upsampled features Figure 3 , and the upsampled features Figure 3 are input into E32; E11 processes the input downsampled features Figure 4 and the latent encoding corresponding to E11 to obtain the first high-quality feature Figure 1 , and the first high-quality feature Figure 1 is input into E12; E12 processes the input first high-quality feature Figure 1 and the upsampled features Figure 1 to obtain the second high-quality feature Figure 1 , and the second high-quality feature Figure 1 is input into E21; E21 processes the input second high-quality feature Figure 1 and the latent encoding corresponding to E21 to obtain the first high-quality feature Figure 2 , and the first high-quality feature Figure 2 is input into E22; E22 processes the input first high-quality feature Figure 2 and the upsampled features Figure 2 to obtain the second high-quality feature Figure 2 , and the second high-quality feature Figure 2 is input into E31; E31 processes the input second high-quality feature Figure 2 and the latent encoding corresponding to E31 to obtain the first high-quality feature Figure 3 , and the first high-quality feature Figure 3 is input into E32; E32 processes the input first high-quality feature Figure 3 and the upsampled features Figure 3 to obtain the restored sample image.

[0057] As Figure 3 shown is a structural example diagram of the downsampling residual block (i.e., the downsampling residual block in the first feature extraction network and the second feature extraction network) provided by an embodiment of the present application. The downsampling residual block is completed using a convolutional layer with a stride equal to 2 during downsampling. As Figure 4The figure shows a structural example diagram of the upsampling residual block provided by an embodiment of the present application. The upsampling residual block uses an upsample layer with the nearest neighbor interpolation method during upsampling. Selecting a downsampling residual block that uses a convolutional layer with a stride of 2 during downsampling and an upsample layer with the nearest neighbor interpolation method during upsampling can facilitate subsequent model deployment without sacrificing high-quality effects.

[0058] As Figure 5 The figure shows a structural example diagram of the style convolutional network provided by an embodiment of the present application. The role of the style convolutional network is to adjust the style of the features extracted by the generation network, so that a variety of different styles of high-quality prior information can be obtained. At the same time, the richness of the texture is enhanced by adding random noise, and the latent code is transformed through a fully connected layer (FC layer) to control the weights of the convolutional layer. As Figure 6 The figure shows a structural example diagram of the conditional control network provided by an embodiment of the present application. The conditional control network performs a convolutional transformation on the feature maps extracted by the first extraction network at different resolutions, which can restrict the generation network to provide high-quality features, so as to achieve the purpose of keeping the facial identity features unchanged. As Figure 7 The figure shows a structural example diagram of the fusion module provided by an embodiment of the present application. The fusion module can focus on restoring the key point details of the face by fusing the geometric prior (i.e., the facial key point features extracted by the second feature extraction network) and the global features extracted by the first feature extraction network, and uses a splicing method for fusion, which can make the fusion module more flexible in fusing the above two parts of the extracted features. Optionally, the Concat function can be used to implement feature splicing.

[0059] Step 103: Based on the restored sample image and the restored reference image, train the initial facial image restoration model to obtain a target facial image restoration model.

[0060] Among them, the target facial image restoration model is used to perform facial image restoration on the facial image to be restored.

[0061] During the training process of the initial facial image restoration model, the model parameters of the initial facial image restoration model can be continuously updated based on the difference between the restored sample image and the restored reference image. The model parameters corresponding to the minimum difference between the restored sample image and the restored reference image are the model parameters of the target facial image restoration model.

[0062] Among them, training the initial facial image restoration model can be to train the first feature extraction network, the second feature extraction network, and the generation network. In this case, the model parameters of the initial facial image restoration model include the model parameters of the first feature extraction network, the model parameters of the second feature extraction network, and the model parameters of the generation network; it can also be to train the first feature extraction network and the second feature extraction network. In this case, the generation network in the initial facial image restoration model is pre-trained, and the model parameters of the initial facial image restoration model include the model parameters of the first feature extraction network and the model parameters of the second feature extraction network; when there is a latent code corresponding to the convolutional layer of the generation network, and the latent code is used to adjust the weights of the convolutional layer of the generation network, the first feature extraction network, the second feature extraction network, and the latent code can be trained to obtain the trained first feature extraction network, the trained target feature extraction network, and the trained latent code. The model parameters of the initial facial image restoration model include the model parameters of the first feature extraction network and the model parameters of the second feature extraction network. After obtaining the trained latent code, the trained latent code can be merged into the corresponding convolutional layer, that is, first use the trained latent code to adjust the weights of the corresponding convolutional layer to obtain the adjusted weights, and update the weights of the convolutional layer to the adjusted weights, so that the structure of the generation network can be made simpler.

[0063] To facilitate the use of the model in embedded devices, in this embodiment, the latent code is designed as a learnable parameter. When the trained latent code is obtained and deployed, the trained latent code can be merged into the corresponding convolutional layer. When adding random noise during deployment, it can be set as a fixed value as the bias of the style convolutional network, so that the structure of the target facial image restoration model can be made simpler while achieving high-quality restoration. As Figure 8 shown is another structural example diagram of the style convolutional network provided by the embodiment of the present application, which is a structural example diagram of deploying random noise and the trained latent code after the style convolutional network.

[0064] The deployment process of random noise and the trained latent code can be represented by the following formulas (1) and (2). Formula (1) represents Figure 5 the feature map processing process of the style convolutional network shown. Formula (2) represents Figure 8 the feature map processing process of the style convolutional network shown.

[0065] (1)

[0066] (2)

[0067] Among them, represents convolution, denotes multiplication. In formula (1), and respectively represent the weights and biases of the FC layer, which are used to calculate the transformed latent code and adjust the original (i.e., the weights of the convolutional layer). After adjustment, it is convolved with (the feature map received by the convolutional layer), and then upsampled. Finally, random noise is added for activation (using the Leaky ReLu activation function shown in Figure 5 ) to obtain the result Y. Since the embodiments of the present application use learnable latent codes for training, during actual deployment, the transformation of in formula (1) can be calculated in advance, and at the same time, noise can also be incorporated into the biases of the convolutional layer. In this way, the computational complexity can be greatly reduced, and the deployment difficulty can be reduced. By removing the convolutional layer with variable weights, formula (1) is transformed into formula (2), and the calculation results of these two formulas are the same. In formula (2), is the adjusted weight obtained after adjusting using the trained latent code, , represents the trained latent code. N in formula (1) and formula (2) represents the upsampling multiple, and N is an integer greater than 1. By way of example and not limitation, N is 2.

[0068] In a possible implementation manner, when training the initial face image restoration model based on the restored sample image and the restored reference image, the first feature extraction network, the second feature extraction network, and the generation network can be trained. In this case, the model parameters of the above initial face image restoration model include the model parameters of the first feature extraction network, the model parameters of the second feature extraction network, and the model parameters of the generation network; it is also possible to only train the first feature extraction network and the second feature extraction network. In this case, the generation network is pre-trained, and the model parameters of the above initial face image restoration model include the model parameters of the first feature extraction network, the model parameters of the second feature extraction network, and the latent code. The above model parameters may include the weights and biases of the corresponding network.

[0069] Based on the pre-trained generation network, the generation network can provide a generation prior, thereby providing high-quality low, medium, and high-frequency image features and achieving a better image restoration effect. The generation prior generally refers to the feature information extracted from high-quality images. For example, a generation network is used to complete the mapping of the Gaussian distribution and the high-quality data distribution, so as to obtain a high-quality generation prior on the known distribution.

[0070] The face image restoration model provided by the embodiments of this application is a deep learning-based face image restoration model that combines geometric priors (i.e., face key point features) and generative priors, providing stable high-quality image restoration results. At the same time, a learnable latent code is used to optimize an optimal latent code during training, retaining part of the high-quality generative priors. After adjusting the weights of the convolutional layers of the generative network, the target face image restoration model can be used on some embedded devices that are not easy to deploy and have poor performance.

[0071] In a possible implementation, the difference between the restored sample image and the restored reference image can be characterized by a loss value. On this basis, step 103 above includes: calculating a first loss value, where the first loss value is the loss value between the restored sample image and the restored reference image; obtaining a first perceptual feature map and a second perceptual feature map, where the first perceptual feature map includes the perceptual features of the restored sample image, and the second perceptual feature map includes the perceptual features of the restored reference image; calculating a second loss value, where the second loss value is the loss value between the first perceptual feature map and the second perceptual feature map; obtaining a first face identity feature map and a second face identity feature map, where the first face identity feature map includes the face identity features of the restored sample image, and the second identity feature map includes the face identity features of the restored reference image; calculating a third loss value, where the third loss value is the loss value between the first face identity feature map and the second face identity feature map; and training the initial face image restoration model based on the first loss value, the second loss value, and the third loss value.

[0072] Among them, the above first loss value characterizes the pixel-level difference between the restored sample image and the restored reference image. The second loss value is a perceptual loss. The perceptual loss focuses on the high-level features of the image (such as texture features, shape features, etc.). Through the perceptual loss, the subtle differences perceived by the human visual system can be better captured, thereby generating high-quality images. The third loss value is an identity feature loss, which is used to maintain the original face identity features of the face sample image, so that the restored sample image is as consistent as possible with the face sample image in terms of face identity features.

[0073] Optionally, the Visual Geometry Group (VGG) can be used to extract the first perceptual feature map and the second perceptual feature map respectively. The Arcface network can be used to extract the first face identity feature map and the second face identity feature map respectively.

[0074] Optionally, a preset loss function can be used to calculate the first loss value, the second loss value, and the third loss value. This application does not limit the type of the loss function.

[0075] As an example and not by way of limitation, the first loss value can be calculated by formula (3), the second loss value can be calculated by formula (4), and the third loss value can be calculated by formula (5).

[0076] (3)

[0077] Wherein, represents the first loss value, and respectively represent the height and width of the restored sample image or the restored reference image (the restored sample image and the restored reference image have the same height and the same width), represents the restored sample image, represents the restored reference image, represents the face sample image, represents the coordinates of the pixels in the restored sample image and the restored reference image, represents the restored sample image at the coordinates of the pixel value of the pixel, represents the restored reference image at the coordinates of the pixel value of the pixel.

[0078] (4)

[0079] Wherein, represents the second loss value, and respectively represent the height and width of the first perceptual feature map or the second perceptual feature map (the first perceptual feature map and the second perceptual feature map have the same height and the same width), represents the first perceptual feature map, represents the second perceptual feature map, represents the coordinates of the pixels in the first perceptual feature map and the second perceptual feature map, represents the first perceptual feature map at the coordinates of the pixel value of the pixel, represents the second perceptual feature map at the coordinates of the pixel value of the pixel.

[0080] (5)

[0081] Wherein, represents the third loss value, and respectively represent the height and width of the first face identity feature map or the second face identity feature map (the first face identity feature map and the second face identity feature map have the same height and the same width), represents the first face identity feature map, represents the second face identity feature map, represent the coordinates of pixels in the first facial identity feature map and the second facial identity feature map, represent the pixel value of the pixel with coordinates in the first facial identity feature map, represent the pixel value of the pixel with coordinates in the second facial identity feature map.

[0082] In one embodiment, when training the initial facial image restoration model based on the first loss value, the second loss value, and the third loss value, the first loss value, the second loss value, and the third loss value may be weighted and summed first based on the weights of the first loss value, the second loss value, and the third loss value to obtain a total loss value, and the initial facial image restoration model may be trained based on the total loss value. Optionally, the weights of the first loss value, the second loss value, and the third loss value may be set according to the scenario requirements. By way of example and not limitation, the weight of the first loss value is 1, the weight of the second loss value is 1, and the weight of the third loss value is 10.

[0083] By way of example and not limitation, during the training process of the initial facial image restoration model, the optimizer used may be the Adam optimizer, the learning rate is 3e-4, and the batch size is 16 to train the initial model. Optionally, other optimizers, other learning rates, and other batch sizes may also be selected for training according to actual requirements.

[0084] During the training process, oscillation needs to be controlled. When the total loss value is too large, the gradient used for parameter update during backpropagation will be relatively large. Adjusting a smaller learning rate can solve the problem of training collapse caused by such large gradient updates, but it will greatly slow down the training speed. In one embodiment, the gradient clipping algorithm may be used to balance the problem of training collapse and the problem of slow training speed, stabilize the training process, and obtain a facial image restoration model applicable to full gain. Specifically, the gradient of each parameter may be scaled by the maximum norm to prevent it from affecting normal training, as shown in formula (6):

[0085] (6)

[0086] where, represents the preset maximum gradient. By way of example and not limitation, it may be set to 10. represents the gradient with the largest absolute value in a batch of gradients, represents the gradient of the total loss value with respect to the parameter w in one gradient descent, represents the clipped gradient (i.e., for The gradient obtained after cropping). It should be understood that the parameter w can be the model parameter of the first feature extraction network and the second feature extraction network, or the latent code corresponding to the convolutional layer in the generation network.

[0087] In the embodiment of the present application, the facial identity features of the facial sample image constrained by the first feature extraction network can enable the generation network to provide a high-quality restoration effect while keeping the facial identity features unchanged. The facial key point features constrained by the second feature extraction network can assist the generation network to capture the positions of facial key points and focus on the restoration effect of facial key points. Therefore, when the generation network trains the initial facial image restoration model based on the facial identity features constrained by the extraction network and the facial key point features constrained by the second feature extraction network, the target facial image restoration model can provide a high-quality restoration effect when restoring the facial image to be restored.

[0088] It should be noted that for a facial image restoration model based on deep learning, the matching degree of the image has a great impact on the restoration effect. For example: whether a low-quality image conforms to the actual acquisition image domain distribution, and whether a high-quality image conforms to the subjective visual perception of the evaluator. According to the actual image acquisition scenario, since the captured human faces in reality will be affected by various conditions such as light, noise, and strong light, the actual situations that cause low-quality images to be processed are not unified. Different situations also have different impacts on the restoration process. For example: the impact of daytime light and noise on the image effect is smaller than that at night; the face with a longer exposure time is more likely to be affected by motion blur than the image with a shorter exposure time; when the sharpening intensity of the image is increased subsequently, the edges on the image are enhanced and the noise is amplified. Conversely, when the noise reduction parameter is larger, the image noise is smaller, but more details are lost. In the embodiment of the present application, through the Figure 9 training sample generation method shown as follows, when making training samples, various situations that may be faced in the actual use scenario are considered, and the low-quality images used in the training samples can be made as close as possible to the actually acquired images.

[0089] Please refer to Figure 9 , Figure 9 which shows a schematic flow diagram of the training sample generation method provided by the embodiment of the present application. As an example but not a limitation, this method is applied to an electronic device, and the method includes the following steps:

[0090] Step 901, obtain a facial image.

[0091] The present application does not limit the source of the facial image. As an example but not a limitation, the above facial image can be a facial image captured by an electronic device, or a facial image obtained from other devices, or a facial image stored in the electronic device itself.

[0092] Step 902: Determine a corresponding facial sample image based on the facial image.

[0093] The image quality of the facial sample image is lower than the image quality of the facial image. The facial sample image and the facial image correspond to the same face.

[0094] In a possible implementation, the determining the corresponding facial sample image based on the facial image includes: performing K degradation processes on the facial image to obtain K degraded images, wherein the K degradation processes have different degrees of degradation, and K is an integer greater than 1; and fusing different regions of the K degraded images to obtain the facial sample image.

[0095] It should be noted that the present application does not limit the manner of degrading the facial image. As an example and not a limitation, the facial image may be subjected to blurring, downsampling, noise addition, compression, noise reduction, upsampling, etc. in sequence to obtain a degraded image.

[0096] In a possible implementation, a method for determining any degraded image may include: adding noise to a facial image to obtain a noisy image; performing non-local mean denoising on the noisy image to obtain a first denoised image; performing total variation denoising on the first denoised image to obtain a second denoised image; and determining a corresponding degraded image based on the second denoised image.

[0097] After obtaining the second denoised image, the electronic device may determine the second denoised image as the corresponding degraded image, or may perform random blurring and random sharpening on the second denoised image to obtain the corresponding degraded image. In the above determination method, different degrees of degradation may be achieved by adding noise of different intensities, performing different degrees of noise reduction, etc., to obtain degraded images with different degrees of degradation.

[0098] The non-local means (NLM) denoising algorithm utilizes the non-local similarity in the image to remove noise by weighted averaging similar pixel blocks while retaining the image details as much as possible. The total variation (TV) denoising algorithm smoothes the image to remove noise while maintaining the edge features of the image. This embodiment can simulate the traditional image signal processing algorithms that may be applied in actual devices by using different types of denoising algorithms. The above-mentioned random blurring and random sharpening are also a kind of simulation. By simulating the actual image processing process, the degraded image can be made more consistent with the actual captured image.

[0099] In a possible implementation, the above-mentioned fusion of different regions of K degraded images to obtain a face sample image includes: obtaining K first mask images, where the regions with pixel value 1 in the K first mask images are different, and the sizes of the K first mask images are the same as the sizes of the K degraded images; fusing the K first mask images with the K degraded images to obtain K fused images, where the degraded images corresponding to different first mask images are different; and fusing the K fused images to obtain a face sample image.

[0100] When fusing the K first mask images with the K degraded images, one first mask image is fused with one degraded image, and the degraded images fused with different first mask images are different. Therefore, K fused images can be obtained after fusion. It should be understood that when fusing the K first mask images with the K degraded images, the first mask image and the degraded image to be fused together can be randomly selected, and the present application does not limit this. By way of example and not limitation, K is 3, the three first mask images are the first mask image 1, the first mask image 2, and the first mask image 3, and the three degraded images are the degraded image 1, the degraded image 2, and the degraded image 3. The first mask image 1 can be fused with the degraded image 1 to obtain the fused image 1, the first mask image 2 can be fused with the degraded image 2 to obtain the fused image 2, and the first mask image 3 can be fused with the degraded image 3 to obtain the fused image 3.

[0101] Optionally, the regions with pixel value 1 in the K first mask images correspond to the entire regions of the degraded images (that is, if the regions obtained by combining the regions with pixel value 1 in the K first mask images have the same size as the size of the degraded images). On this basis, it can be ensured that the finally obtained fused image and the degraded image correspond to the same face and have the same size. By way of example and not limitation, K is 2, the region with pixel value 1 in the first mask image 1 corresponds to the region 1 in the degraded image, and the region with pixel value 1 in the second mask image 2 corresponds to the region 2 in the degraded image, where the region 2 is the region in the degraded image other than the region 1.

[0102] In a possible implementation, the above-mentioned obtaining of K first mask images includes: generating K second mask images, where the pixel values in the second mask images are all 0, and the sizes of the K second mask images are the same as the sizes of the K degraded images; generating L groups of target coordinate information, each group of target coordinate information corresponding to a region, and the regions corresponding to the L groups of target coordinate information are different, where L is an integer greater than or equal to K; distributing the L groups of target coordinate information to the K second mask images, with each second mask image being assigned at least one group of target coordinate information; for any second mask image, updating the pixel values of the at least one group of target coordinate information in the corresponding region of the second mask image to 1 to obtain the corresponding first mask image.

[0103] Among them, the pixel values, sizes, and coordinate information of the K second mask images are the same. L groups of target coordinate information can be generated based on the coordinate information of the second mask image (for example, the coordinate information of the second mask image can be divided into L groups of target coordinate information, which can be randomly divided or the division rules can be specified in advance). The regions corresponding to the L groups of target coordinate information correspond to the entire region of the second mask image (that is, if the regions corresponding to the L groups of target coordinate information are combined, the size of the resulting region is the same as the size of the second mask image).

[0104] When the electronic device distributes the L groups of target coordinate information to the K second mask images, it can be randomly distributed or distributed according to the preset distribution rules. This application does not limit this. Optionally, the distribution rules can be set according to the scenario requirements.

[0105] The added noise includes but is not limited to the noise in the actually collected flat area and Gaussian noise.

[0106] In an actual scenario, the degradation degrees corresponding to different regions in an image may be different. For example, the quality of a face image taken at night is relatively low, and this situation is more serious. To simulate this situation, multiple different degraded images can be generated using different degradation parameters, and then different regions of the multiple degraded images can be fused through a randomly generated mask.

[0107] Optionally, the face image can be degraded to different degrees by setting degradation parameters such as different blur degrees, different downsampling multiples, adding different intensities of noise, different compression degrees during compression processing, different noise reduction degrees, and different upsampling multiples when performing blurring processing.

[0108] To reduce the learning difficulty of the initial face image restoration model, before performing degradation processing and noise reduction processing on the face image, the face image can be aligned to a standard template first. The standard template can refer to a preset face key point template. As Figure 10-1 shown is an example diagram of a face image, and as Figure 10-2 shown is an example diagram of the standard template, and as Figure 10-3 shown is an example diagram of the face image after alignment. The positions of the face key points in the face image after alignment are the same as the positions of the face key points in the standard template. As Figure 10-4 shown is an example diagram of a face sample image, which is obtained by fusing different regions of two degraded images. Figure 10-4 The area within the red frame in

[0109] Optionally, the face image can be aligned to a standard template through an affine transformation. If there are missing parts that cannot be interpolated in the affine transformation, a reflection filling algorithm can be used to fill them.

[0110] Step 903: Denoise the face image to obtain a first denoised image and a second denoised image.

[0111] Among them, the denoising intensity of the first denoised image is greater than that of the second denoised image.

[0112] To improve the mid - low frequency performance of the face data and not further increase the high - frequency texture of the model (because too much high - frequency texture details on the face image are likely to bring anomalies to the final result), this embodiment can use TV denoising with two different denoising intensities to achieve the above effect. TV denoising mainly smooths the image to remove noise while maintaining the edge features of the image. TV denoising can be achieved by minimizing the energy functional in formula (7).

[0113] (7)

[0114] Among them, represents the energy functional, represents the denoised image, represents the L1 norm of the gradient of the denoised image, represents the face image, represents the regularization coefficient, used to balance the trade - off between denoising and fidelity, represents the pixel, and 2 represents the L2 norm.

[0115] Step 904: Based on the first denoised image and the second denoised image, determine the mid - low frequency information of the first denoised image.

[0116] In one embodiment, subtracting the pixel value of the pixel in the first denoised image from the pixel value of the corresponding pixel in the second denoised image can obtain the mid - low frequency information in the first denoised image. Among them, the sizes of the first denoised image and the second denoised image are the same.

[0117] Step 905: Based on the mid - low frequency information, perform frequency enhancement on the face image to obtain a corresponding restored reference image.

[0118] Among them, the image quality of the restored reference image is higher than that of the face image, and the face sample image and the restored reference image are training samples corresponding to the face image.

[0119] After performing frequency enhancement on the face image based on the mid - low frequency information, the restored reference image can be made more three - dimensional than the face image. As Figure 10-5The following is an example diagram of a restored reference image, which is more three-dimensional than the Figure 10-3 shown face image.

[0120] In a possible implementation, after generating the training samples corresponding to the face image, the face image restoration model can be trained so that the trained face image restoration model can better restore the face image to be restored. As an example rather than a limitation, the training samples generated in this embodiment can be used to train the Figure 1 initial face image restoration model of the embodiment shown, so that the target face image restoration model can better restore the face image to be restored.

[0121] The face sample images generated in the embodiments of the present application conform to the actual capture scenario, and the generated restored reference images can improve the mid-low frequency performance of the face data without further enhancing the high-frequency texture of the model.

[0122] Please refer to Figure 11 , Figure 11 which shows a schematic flowchart of the face image restoration method provided by the embodiments of the present application. As an example rather than a limitation, this method is applied to an electronic device in which a target face image restoration model is deployed. The method includes the following steps:

[0123] Step 1101, obtain the face image to be restored and the face key point image corresponding to the face image to be restored.

[0124] The present application does not limit the source of the face image to be restored. As an example rather than a limitation, the above face image to be restored may be a face image captured by an electronic device, or may refer to a face image obtained from other devices, or may also be a face image stored in the electronic device itself.

[0125] The above face image to be restored and the face key point image corresponding to it correspond to the same face and have the same size. The face key points in the face key point image may be key points such as the eyes, nose, and mouth of the face. In order to reduce the influence of other regions of the face (i.e., the regions other than the face key point regions) on the face image restoration, the pixel values of the pixels in the face key point regions in the face key point image are 255, and the pixel values of the pixels in the regions other than the face key point regions in the face key point sample image are 0. The present application does not limit the acquisition method of the face key point image.

[0126] Step 1102, based on the face image to be restored, the face key point image corresponding to the face image to be restored, and the target face image restoration model, determine the restored face image corresponding to the face image to be restored.

[0127] Among them, the above-mentioned target face image restoration model includes a first feature extraction network, a second feature extraction network respectively connected to the first feature extraction network, and a generation network.

[0128] As an example, the target face image restoration model can be trained by intelligent devices such as mobile phones and tablet computers or servers through Figure 1 the method shown, or obtained by the intelligent device or server from device A, where device A is a device that executes Figure 1 the method shown to train the target face image restoration model.

[0129] Optionally, the above step 1102 includes:

[0130] In one case, if the target face image restoration model is deployed in the intelligent device or server, the face image to be restored and the face key point image are input into the target face image restoration model to obtain the restored face image. In another case, if the target face image restoration model is not deployed in the intelligent device or server, the intelligent device sends the face image to be restored and the face key point image to the device where the target face image restoration model is deployed, and then the device where the target face image restoration model is deployed obtains the restored face image based on the face image to be restored and the face key point image, and sends it to the intelligent device or server.

[0131] As an example rather than a limitation, in practical applications, the target face image restoration model can be used to perform face image restoration on real-time captured face images.

[0132] It should be noted that the solutions not detailed in this embodiment can be referred to the descriptions of the relevant solutions in the foregoing embodiments, and will not be elaborated here. It should be understood that the magnitudes of the sequence numbers of the above steps in the embodiments do not mean the order of execution is prior or posterior, and the execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0133] Corresponding to the training method of the face image restoration model described in the above embodiments, Figure 12 the structural schematic diagram of the training device of the face image restoration model provided by the embodiments of the present application is shown. The initial face image restoration model includes: a first feature extraction network, a second feature extraction network respectively connected to the first feature extraction network, and a generation network. For the sake of convenience of description, only the parts related to the embodiments of the present application are shown.

[0134] Referring to Figure 12 , the device includes:

[0135] A sample acquisition module 1201 for acquiring training samples and facial key point sample images corresponding to the training samples. The training samples include facial sample images and restoration reference images corresponding to the facial sample images, and the image quality of the restoration reference images is higher than that of the facial sample images.

[0136] An image processing module 1202 for inputting the facial sample images and the facial key point sample images into the initial facial image restoration model, so as to trigger the generation network to predict a restored sample image corresponding to the facial sample image based on the facial identity features of the facial sample image constrained by the first feature extraction network and the facial key point features of the facial key point sample image constrained by the second feature extraction network.

[0137] A model training module 1203 for training the initial facial image restoration model based on the restored sample image and the restoration reference image to obtain a target facial image restoration model, and the target facial image restoration model is used for performing facial image restoration on the facial image to be restored.

[0138] Optionally, the convolutional layer of the generation network corresponds to a latent code, and the latent code is used to adjust the weights of the convolutional layer of the generation network. Specifically, the model training module 1203 is configured to:

[0139] Train the first feature extraction network, the second feature extraction network, and the latent code based on the restored sample image and the restoration reference image to obtain a trained first feature extraction network, a trained target feature extraction network, and a trained latent code.

[0140] Optionally, the first feature extraction network includes a downsampling residual block, a first fusion module, and an upsampling residual block. Specifically, the image processing module 1202 is configured to:

[0141] Input the facial sample image into the downsampling residual block to obtain a downsampled feature map, and the downsampled feature map includes the facial identity features of the facial sample image.

[0142] Input the facial key point sample image into the second feature extraction network to obtain a facial key point feature map, and perform attention enhancement on the downsampled feature map based on the facial key point feature map to obtain an image enhancement feature map. The facial key point feature map includes the facial key point features of the facial key point sample image.

[0143] Input the image enhancement feature map into the first fusion module to trigger the first fusion module to fuse the image enhancement feature map and the downsampled feature map to obtain a fused feature map.

[0144] Input the downsampled feature map and the fused feature map into the upsampling residual block to obtain an upsampled feature map;

[0145] Input the downsampled feature map, the upsampled feature map, and the latent code into the generation network to obtain the restored sample image.

[0146] Optionally, the first feature extraction network includes a downsampling residual block, a second fusion module, and an upsampling residual block. The above image processing module 1202 is specifically configured to:

[0147] Input the face sample image into the downsampling residual block to obtain a downsampled feature map, where the downsampled feature map includes the face identity features of the face sample image;

[0148] Input the face key point sample image into the second feature extraction network to obtain a face key point feature map, where the face key point feature map includes the face key point features of the face key point sample image;

[0149] Input the face key point feature map into the second fusion module to trigger the second fusion module to enhance the attention of the face key points to the downsampled feature map based on the face key point feature map, obtain an image enhancement feature map, and fuse the image enhancement feature map and the downsampled feature map to obtain a fused feature map;

[0150] Input the downsampled feature map and the fused feature map into the upsampling residual block to obtain an upsampled feature map;

[0151] Input the downsampled feature map, the upsampled feature map, and the latent code into the generation network to obtain the restored sample image.

[0152] Optionally, the above model training module 1203 is specifically configured to:

[0153] Calculate a first loss value, where the first loss value is the loss value between the restored sample image and the restored reference image;

[0154] Obtain a first perceptual feature map and a second perceptual feature map, where the first perceptual feature map includes the perceptual features of the restored sample image, and the second perceptual feature map includes the perceptual features of the restored reference image;

[0155] Calculate a second loss value, where the second loss value is the loss value between the first perceptual feature map and the second perceptual feature map;

[0156] Obtain a first facial identity feature map and a second facial identity feature map, where the first facial identity feature map includes the facial identity features of the restored sample image, and the second facial identity feature map includes the facial identity features of the restored reference image;

[0157] Calculate a third loss value, where the third loss value is the loss value between the first facial identity feature map and the second facial identity feature map;

[0158] Based on the first loss value, the second loss value, and the third loss value, train the initial facial image restoration model.

[0159] Optionally, the above sample acquisition module 1201 includes:

[0160] An image acquisition unit for acquiring facial images;

[0161] A sample determination unit for determining a corresponding facial sample image based on the facial image, where the image quality of the facial sample image is lower than that of the facial image;

[0162] A noise reduction processing unit for performing noise reduction processing on the facial image to obtain a first noise reduction image and a second noise reduction image, where the noise reduction intensity of the first noise reduction image is greater than that of the second noise reduction image;

[0163] A mid-low frequency determination unit for determining the mid-low frequency information of the first noise reduction image based on the first noise reduction image and the second noise reduction image;

[0164] A frequency enhancement unit for performing frequency enhancement on the facial image based on the mid-low frequency information to obtain a corresponding restored reference image, where the image quality of the restored reference image is higher than that of the facial image.

[0165] Optionally, the above sample determination unit includes:

[0166] A degradation subunit for performing K times of degradation processing on the facial image to obtain K degraded images, where the degradation degrees of the K times of degradation processing are different, and K is an integer greater than 1;

[0167] A fusion subunit for fusing different regions of the K degraded images to obtain the facial sample image.

[0168] Optionally, the above degradation subunit is specifically used for:

[0169] Adding noise to the facial image to obtain a noise image;

[0170] Performing non-local mean noise reduction on the noise image to obtain a first noise reduction image;

[0171] Perform total variation denoising on the first denoised image to obtain a second denoised image;

[0172] Based on the second denoised image, determine the corresponding degraded image.

[0173] Optionally, the above-mentioned fusion subunit is specifically configured to:

[0174] Obtain K first mask images, where the regions with pixel value 1 in the K first mask images are different, and the sizes of the K first mask images are the same as the sizes of the K degraded images;

[0175] Fuse the K first mask images with the K degraded images to obtain K fused images, and the degraded images corresponding to different first mask images are different;

[0176] Fuse the K fused images to obtain the face sample image.

[0177] Optionally, the above-mentioned fusion subunit is specifically configured to:

[0178] Generate K second mask images, where the pixel values in the second mask images are all 0, and the sizes of the K second mask images are the same as the sizes of the K degraded images;

[0179] Generate L groups of target coordinate information, each group of target coordinate information corresponds to a region, the regions corresponding to the L groups of target coordinate information are different, and L is an integer greater than or equal to K;

[0180] Allocate the L groups of target coordinate information to the K second mask images, and each second mask image is allocated at least one group of target coordinate information;

[0181] For any one of the second mask images, update the pixel values of at least one group of target coordinate information in the region corresponding to the second mask image to 1 to obtain the corresponding first mask image.

[0182] It should be noted that for the information interaction, execution process, etc. between the above-mentioned devices / units, since they are based on the same concept as the method embodiments of the present application, their specific functions and the technical effects brought about can be specifically referred to the method embodiment part, and will not be elaborated here.

[0183] Figure 13 This is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 13 shown, the electronic device 13 in this embodiment includes: at least one processor 1300 ( Figure 13only one is shown), a memory 1301, and a computer program 1302 stored in the memory 1301 and executable on the at least one processor 1300. When the processor 1300 executes the computer program 1302, the steps in any of the above method embodiments are implemented.

[0184] The electronic device may include, but is not limited to, a processor 1300 and a memory 1301. Those skilled in the art can understand that Figure 13 merely an example of the electronic device 13, which does not constitute a limitation on the electronic device 13. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0185] The so-called processor 1300 may be a central processing unit (CPU), and the processor 1300 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0186] In some embodiments, the memory 1301 may be an internal storage unit of the electronic device 13, such as the hard disk or memory of the electronic device 13. In other embodiments, the memory 1301 may also be an external storage device of the electronic device 13, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 13. Further, the memory 1301 may also include both the internal storage unit and the external storage device of the electronic device 13. The memory 1301 is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of the computer program. The memory 1301 may also be used to temporarily store data that has been output or will be output.

[0187] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of this application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the device / electronic device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc.

[0188] In the above embodiments, the descriptions of the various embodiments have their own focuses. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0189] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0190] In the embodiments provided in this application, it should be understood that the disclosed device / electronic device and method can be implemented in other ways. For example, the device / electronic device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in an electrical, mechanical, or other form.

[0191] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0192] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included within the protection scope of the present application.

Claims

1. A training method for a facial image restoration model, characterized in that: The initial facial image restoration model includes: a first feature extraction network, a second feature extraction network and a generation network respectively connected to the first feature extraction network, and the training method includes: Acquire a training sample and a facial key point sample image corresponding to the training sample, wherein the training sample includes a facial sample image and a restored reference image corresponding to the facial sample image, and the image quality of the restored reference image is higher than that of the facial sample image; Inputting the facial sample image and the facial key point sample image into the initial facial image restoration model to trigger the generation network to predict a restored sample image corresponding to the facial sample image based on the facial identity features of the facial sample image constrained by the first feature extraction network and the facial key point features of the facial key point sample image constrained by the second feature extraction network; Based on the restored sample image and the restored reference image, the initial facial image restoration model is trained to obtain a target facial image restoration model, wherein the target facial image restoration model is used to perform facial image restoration on the facial image to be restored; The convolution layer of the style convolution network in the generation network corresponds to a potential code, and the potential code is used to adjust the weight of the convolution layer of the style convolution network. The initial facial image restoration model is trained based on the restored sample image and the restored reference image to obtain a target facial image restoration model, including: Based on the restored sample image and the restored reference image, the first feature extraction network, the second feature extraction network and the latent code are trained to obtain a trained first feature extraction network, a trained second feature extraction network and a trained latent code; After getting the trained latent code, it also includes: deploying random noise and the trained latent code in the style convolutional network; The deployment process of the random noise and the trained latent code is represented by the input and output relationship of the style convolution network before deployment and the input and output relationship of the style convolution network after deployment; The input and output relationship of the style convolutional network described before deployment is expressed as follows: After deployment, the input and output relationship of the style convolutional network is expressed as follows: in, represents convolution, represents multiplication, represents the output of the style convolutional network, represents the input of the style convolutional network, represents the latent code after training, represents the weights of the convolutional layer of the style convolutional network, Represents the potential encoding pair after training using The adjusted weights obtained after adjustment are and They represent the weights and biases of the FC layer of the style convolutional network described before deployment, represents the Leaky ReLu activation function, represents the upsample layer of the style convolutional network, represents random noise, Indicates the upsampling multiple, is an integer greater than 1.

2. The training method according to claim 1, characterized in that: The first feature extraction network includes a downsampling residual block, a first fusion module and an upsampling residual block, and the step of inputting the facial sample image into the initial facial image restoration model to trigger the generation network to predict a restored sample image corresponding to the facial sample image based on the facial identity features of the facial sample image constrained by the first feature extraction network and the facial key point features of the facial key point sample image constrained by the second feature extraction network, comprises: Inputting the facial sample image into the down-sampling residual block to obtain a down-sampling feature map, wherein the down-sampling feature map includes facial identity features of the facial sample image; Inputting the facial key point sample image into the second feature extraction network to obtain a facial key point feature map, and performing attention enhancement of facial key points on the downsampled feature map based on the facial key point feature map to obtain an image enhancement feature map, wherein the facial key point feature map includes facial key point features of the facial key point sample image; Inputting the image enhancement feature map into the first fusion module to trigger the first fusion module to fuse the image enhancement feature map and the down-sampling feature map to obtain a fused feature map; Inputting the downsampled feature map and the fused feature map into the upsampled residual block to obtain an upsampled feature map; The down-sampled feature map, the up-sampled feature map and the potential code are input into the generation network to obtain the restored sample image.

3. The training method according to claim 1, characterized in that: The first feature extraction network includes a downsampling residual block, a second fusion module and an upsampling residual block, and the step of inputting the facial sample image into the initial facial image restoration model to trigger the generation network to predict a restored sample image corresponding to the facial sample image based on the facial identity features of the facial sample image constrained by the first feature extraction network and the facial key point features of the facial key point sample image constrained by the second feature extraction network, comprises: Inputting the facial sample image into the down-sampling residual block to obtain a down-sampling feature map, wherein the down-sampling feature map includes facial identity features of the facial sample image; Inputting the facial key point sample image into the second feature extraction network to obtain a facial key point feature map, wherein the facial key point feature map includes facial key point features of the facial key point sample image; Inputting the facial key point feature map into the second fusion module to trigger the second fusion module to perform attention enhancement of facial key points on the downsampled feature map based on the facial key point feature map to obtain an image enhancement feature map, and fusing the image enhancement feature map with the downsampled feature map to obtain a fused feature map; Inputting the downsampled feature map and the fused feature map into the upsampled residual block to obtain an upsampled feature map; The down-sampled feature map, the up-sampled feature map and the potential code are input into the generation network to obtain the restored sample image.

4. The training method according to any one of claims 1 to 3, characterized in that: The training of the initial facial image restoration model based on the restored sample image and the restored reference image comprises: Calculating a first loss value, where the first loss value is a loss value between the restored sample image and the restored reference image; Acquire a first perceptual feature map and a second perceptual feature map, wherein the first perceptual feature map includes perceptual features of the restored sample image, and the second perceptual feature map includes perceptual features of the restored reference image; Calculate a second loss value, where the second loss value is a loss value between the first perceptual feature map and the second perceptual feature map; Acquire a first facial identity feature map and a second facial identity feature map, wherein the first facial identity feature map includes facial identity features of the restored sample image, and the second facial identity feature map includes facial identity features of the restored reference image; Calculating a third loss value, where the third loss value is a loss value between the first facial identity feature map and the second facial identity feature map; The initial facial image restoration model is trained based on the first loss value, the second loss value and the third loss value.

5. The training method according to any one of claims 1 to 3, characterized in that: The obtaining of training samples comprises: Get a facial image; Based on the facial image, determining a corresponding facial sample image, wherein the image quality of the facial sample image is lower than the image quality of the facial image; Performing denoising processing on the facial image to obtain a first denoised image and a second denoised image, wherein the denoising intensity of the first denoised image is greater than the denoising intensity of the second denoised image; Determining mid- and low-frequency information of the first denoised image based on the first denoised image and the second denoised image; The frequency of the facial image is enhanced based on the medium and low frequency information to obtain a corresponding restored reference image, wherein the image quality of the restored reference image is higher than that of the facial image.

6. The training method according to claim 5, characterized in that: The determining, based on the facial image, a corresponding facial sample image comprises: Performing K times of degradation processing on the facial image to obtain K degraded images, wherein the K times of degradation processing have different degrees of degradation, and K is an integer greater than 1; Different regions of the K degraded images are fused to obtain the face sample image.

7. The training method according to claim 6, characterized in that: Any method for determining the degraded image includes: adding noise to the facial image to obtain a noise image; Performing non-local mean denoising on the noisy image to obtain a first denoised image; Performing total variation denoising on the first denoised image to obtain a second denoised image; Based on the second denoised image, the corresponding degraded image is determined.

8. The training method according to claim 6, characterized in that: The different regions of the K degraded images are fused to obtain the face sample image, including: Acquire K first mask images, wherein regions with pixel values ​​of 1 in the K first mask images are different, and sizes of the K first mask images are the same as sizes of the K degraded images; fusing the K first mask images with the K degraded images to obtain K fused images, where different first mask images correspond to different degraded images; The K fused images are fused to obtain the face sample image.

9. The training method according to claim 8, characterized in that: The step of acquiring K first mask images comprises: Generate K second mask images, wherein pixel values ​​in the second mask images are all 0, and sizes of the K second mask images are the same as sizes of the K degraded images; Generate L groups of target coordinate information, each group of target coordinate information corresponds to an area, the L groups of target coordinate information correspond to different areas, and L is an integer greater than or equal to K; Allocating the L groups of target coordinate information to the K second mask images, each of the second mask images being allocated at least one group of target coordinate information; For any of the second mask images, the pixel value of the allocated at least one set of target coordinate information in the area corresponding to the second mask image is updated to 1 to obtain the corresponding first mask image.

Citation Information

Patent Citations

  • Human face restoration model training method, restoration method, device and equipment and medium

    CN111507914A

  • Face super-resolution system based on multi-scale convolution and receptive field feature fusion

    CN112507997A

  • Progressive face image restoration method, system and device and storage medium

    CN117391995A

  • Human image restoration method and apparatus, electronic device, storage medium and program product

    WO2022110638A1

Cited By

  • Pathological prognosis modeling method based on intra-tumor nerve neighborhood cell level graph convolution

    CN120997576A