A method and system for face image restoration

By using high-quality and low-quality feature extraction networks, cross-quality transfer estimation networks and pre-trained image recovery networks in facial image repair, the difficult problem of unknown degradation in the prior art is solved, and more accurate identity information and flexible repair is achieved.

CN114219728BActive Publication Date: 2025-05-30SHANGHAI JIAOTONG UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111496917.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-09
Publication Date
2025-05-30
Estimated Expiration
2041-12-09

AI Technical Summary

Technical Problem

In the repair of face image, the prior art requires pairs of training images, which are difficult to deal with unknown degradation situations in the real world, and the identity information generated in the image is not accurate enough.

Method used

Feature expressions of images are acquired through high-quality feature extraction networks and low-quality feature extraction networks, transfer vectors between high-quality and low-quality expressions are estimated using the cross-quality transfer estimation network, and the edited expressions are mapped into output images through the pre-trained image recovery network.

Benefits of technology

Without pairs of training image pairs, it can be used for unknown degradation in real images, the identity information generated in the image is more accurate, and allows the user to adjust the degree of repair of the output image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114219728B_ABST
    Figure CN114219728B_ABST
Patent Text Reader

Abstract

The present invention provides a face image restoration method, including: obtaining a high-quality expression of an input high-quality face image in a feature space by using a high-quality feature extraction network; obtaining a low-quality expression of an input low-quality face image in the feature space by using a low-quality feature extraction network; using a cross-quality transfer estimation network to estimate a transfer vector between the high-quality expression and the low-quality expression in the feature space, and editing the expression by using the transfer vector; using a pre-trained image restoration network to map the edited expression to an output image; training under the joint loss constraint of the entire network; and using the trained network for human image restoration. The present invention does not require paired training image pairs, can be applicable to unknown degradations in real images, and improves the problem that the prior art requires paired training image pairs and does not conform to the actual usage scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method in the fields of computer vision and image processing, and particularly to a method and system for face image restoration. Background Art

[0002] Face restoration aims to restore a low-quality face image (LQ) to a corresponding high-quality face image (HQ). In the past few years, many image restoration methods based on deep neural networks have achieved great success. However, image restoration is an ill-posed problem, and multiple high-quality images can degrade to the same low-quality image, that is, one low-quality image corresponds to multiple high-quality images. During training, the network is also affected by this one-to-many relationship, and what is fitted is the average of one low-quality image corresponding to multiple high-quality images, which results in a blurred output image. Considering this, some methods use pre-trained generative models. Since these pre-trained models are trained on high-quality image datasets, their network parameters have the characteristics of generating high-quality images. However, most of these methods require paired training images to train the network parameters supervised. In the real world, there are low-quality images that have undergone various unknown degradations. When their degradations exceed the training degradation range, these supervised methods cannot handle these low-quality images. Summary of the Invention

[0003] Aiming at the defects in the prior art, the purpose of the present invention is to provide a method and system for face image restoration.

[0004] According to one aspect of the present invention, a method for face image restoration is provided, including:

[0005] Obtaining a high-quality expression of an input high-quality face image in the feature space by using a high-quality feature extraction network;

[0006] Obtaining a low-quality expression of an input low-quality face image in the feature space by using a low-quality feature extraction network;

[0007] Estimating a transfer vector between the high-quality expression and the low-quality expression in the feature space by using a cross-quality transfer estimation network, and editing the expression by using the transfer vector;

[0008] Mapping the edited expression to an output image by using a pre-trained image restoration network;

[0009] Preferably, it further includes training under the joint loss constraint of the entire network; using the trained network for human image restoration.

[0010] Preferably, the obtaining a high-quality expression of an input high-quality face image in the feature space by using a high-quality feature extraction network includes:

[0011] Input a high-quality image where C is the number of image channels, W is the width of the image, and H is the height of the image. Use the high-quality feature extraction network E hq to extract the visual features of the image where k is the feature dimension and N corresponds to the number of network layers of the pre-trained image restoration network.

[0012] Preferably, obtaining the low-quality representation of the input low-quality face image in the feature space by using the low-quality feature extraction network includes:[[]]

[0013] Input a low-quality image where C is the number of image channels, W is the width of the image, and H is the height of the image. Use the low-quality feature extraction network E lq to extract the visual features of the image where k is the feature dimension and N corresponds to the number of network layers of the pre-trained image restoration network.

[0014] Preferably, using the cross-quality transfer estimation network to estimate the transfer vector between the high-quality representation and the low-quality representation in the feature space, and using the transfer vector to edit the low-quality representation and the high-quality representation includes:[[]]

[0015] According to the low-quality representation w l and the high-quality representation w h in the feature space, estimate the transfer vector Δ by the cross-quality transfer estimation network;

[0016] Make the image corresponding to w l +Δ in the high-quality image domain discriminated by the discriminator, and the image corresponding to w h -Δ in the low-quality image domain discriminated by the discriminator;

[0017] Edit the low-quality representation w l and the high-quality representation w h through the transfer vector Δ to obtain the edited representation.

[0018] Preferably, using the pre-trained image restoration network to map the edited representation to the output image includes:[[]]

[0019] The image restoration network G is a pre-trained StyleGAN model, which maps the edited representation to the output image, l h = G(w h -Δ), h l = G(w l +Δ). They are respectively the high-quality form corresponding to the low-quality input l and the low-quality form corresponding to the high-quality input h.

[0020] Preferably, using the pre-trained image restoration network to map the edited expression to an output image further includes:

[0021] Using the high-quality feature extraction network E hq Obtain the expression of the high-quality face image in the feature space, edit it using the estimated transfer vector Δ, and finally use the image restoration network G to map the edited expression to an output image, including:

[0022]

[0023] Using the low-quality feature extraction network E lq Obtain the expression of the low-quality face image in the feature space, edit it using the estimated transfer vector Δ, and finally use the image restoration network G to map the edited expression to an output image, including:

[0024]

[0025] Preferably, using the pre-trained image restoration network to map the edited expression to an output image further includes:

[0026] Through the feature extraction network E lq And E hq The extracted low-quality expression w l And the high-quality expression w h Can be restored to the low-quality image l rec And the high-quality image h rec :

[0027] l rec =G(w l ),h rec =G(w h )。

[0028] Preferably, the loss function of the feature extraction network includes:

[0029] L down =1-<R(l↓),R(h l ↓)>

[0030] L rep =||h-h rec || 1 +||l-l rec || 1

[0031]

[0032] L adv =-log(D lq (G(w h-Δ)))-log(D hq (G(w l +Δ))),

[0033]

[0034] wherein, L down is the downsampling loss, R is a face recognition method, ↓ is downsampling, and L cc is the cycle consistency loss, and is the output image, and L adv is the adversarial loss, and L perc is the high-quality loss, is a classification method.

[0035] Preferably, the loss function L of the entire network is:

[0036] L = λ cc L cc + λ rep L rep + λ adv L adv + λ perc L perc + λ down L down ,

[0037] wherein, L rec , L vgg , L vgg , L vgg and L vgg are the loss functions of the image restoration network, and λ cc , λ rep , λ adv , λ perc and λ down are the weights for balancing several losses.

[0038] According to the second aspect of the present invention, a face image restoration system is provided, including:

[0039] A high-quality feature extraction module that extracts the expression of the input high-quality image in the feature space;

[0040] A low-quality feature extraction module that extracts the expression of the input low-quality image in the feature space;

[0041] A cross-quality transfer estimation module that estimates the transfer vector between the high-quality expression and the low-quality expression in the feature space and uses it to edit the low-quality expression and the high-quality expression.

[0042] An image restoration module, which uses an image restoration network to map the edited low-quality expression to an output image.

[0043] Compared with the prior art, the present invention has the following beneficial effects:

[0044] (1) A face restoration method and system provided by the present invention do not require paired training image pairs and can be applied to unknown degradations in real images, improving the problem of the prior art that requires paired training image pairs and does not conform to the actual usage scenario.

[0045] (2) A face restoration method and system provided by the present invention estimate the gap between the low-quality image and the high-quality image in the feature space, rather than estimating its degradation based on the low-quality image, which makes the identity information of the generated image more accurate.

[0046] (3) A face restoration method and system provided by the present invention allow users to change the estimated cross-quality transfer vector to adjust the restoration degree of the output image. Description of the Drawings

[0047] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objectives, and advantages of the present invention will become more apparent:

[0048] Figure 1 It is a flowchart of a face image restoration method according to an embodiment provided by the present invention;

[0049] Figure 2 It is a block diagram of a face image restoration system according to an embodiment provided by the present invention. Detailed Embodiments

[0050] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that those of ordinary skill in the art can make several modifications and improvements without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0051] As Figure 1 shown, it is a flowchart of a face restoration method according to an embodiment of the present invention. It includes:

[0052] S11. Using a high-quality feature extraction network for the input high-quality face image to obtain its high-quality expression in the feature space;

[0053] S12. Using a low-quality feature extraction network for the input low-quality face image to obtain its low-quality expression in the feature space;

[0054] S13. Estimate the transfer vector between high-quality and low-quality expressions in the feature space using a cross-quality transfer estimation network, and edit the expression using the transfer vector;

[0055] S14. Map the edited expression to an output image using a pre-trained image restoration network.

[0056] It also includes:

[0057] S15. Train under the joint loss constraint of the entire network;

[0058] S16. Use the trained network for human image restoration.

[0059] The above-mentioned feature space refers to the feature space of the face image restoration model, and the above-mentioned entire network is the entire face image restoration model. This embodiment uses unpaired high-quality images and low-quality images to train the model parameters, which means that this embodiment can train on images with different degradations to cope with various degradations existing in the real world.

[0060] To obtain a better high-quality expression of a high-quality face image, a preferred embodiment is provided to execute S11. Using a high-quality feature extraction module, obtain the expression of the high-quality image in the feature space. Input the high-quality image where C is the number of image channels, W is the width of the image, and H is the height of the image. Using the convolutional neural network E hq Extract the visual features of the image where k is the feature dimension, and N corresponds to the number of network layers of the pre-trained image restoration network.

[0061] To obtain a better low-quality expression of a low-quality face image, a preferred embodiment is provided to execute S12. Using a low-quality feature extraction module, obtain the expression of the low-quality image in the feature space. Input the low-quality image where C is the number of image channels, W is the width of the image, and H is the height of the image. Using the convolutional neural network E lq Extract the visual features of the image where k is the feature dimension, and N corresponds to the number of network layers of the pre-trained image restoration network.

[0062] Based on S11 and S12 of the above embodiments, execute S13. Use the cross-quality transfer estimation network to estimate the transfer vector between high-quality expressions and low-quality expressions in the feature space, and use it to edit expressions of different qualities. Estimate the transfer vector according to the low-quality expression and the high-quality expression in the feature space. Make the corresponding image in the high-quality image domain discriminated by the discriminator of the cross-quality transfer estimation network, and the corresponding image in the low-quality image domain discriminated by the discriminator of the cross-quality transfer estimation network. Then, edit the low-quality expression and the high-quality expression through the transfer vector to obtain the edited expression.

[0063] To obtain a better image output, a preferred embodiment is provided to execute S14. In this embodiment, the pre-trained image restoration network selects the pre-trained Stylegan model to map the fused features into an output image. Since the Stylegan model is pre-trained, the images it generates have high-quality details.

[0064] Specifically, the model maps the edited expression into an output image, l h = G(w h -Δ), h l = G(w l +Δ). They are respectively the low-quality form corresponding to the high-quality input h and the high-quality form corresponding to the low-quality input l.

[0065] Use the high-quality feature extraction network E hq to obtain the expression of the high-quality face image in the feature space, and edit it using the estimated transfer vector Δ. Finally, use the image restoration network G to map the edited expression into an output image, including:

[0066]

[0067] Use the low-quality feature extraction network E lq to obtain the expression of the low-quality face image in the feature space, and edit it using the estimated transfer vector Δ. Finally, use the image restoration network G to map the edited expression into an output image, including:

[0068]

[0069] Thus, it can be seen that in this embodiment, the edited expressions are specifically: w h -Δ, w l +Δ, E hq (h l )-Δ and E lq (l h )+Δ.

[0070] Through the feature extraction networks E lq and E hq extract the low-quality expression wl with high-quality expression w h can be restored to a low-quality image l through the image restoration network G rec and high-quality image h rec :

[0071] l rec = G(w l ), h rec = G(w h ).

[0072] Based on the above embodiments, S15 is executed. The face image restoration system does not require paired training data for the training module parameters, and the feature extraction module has the following loss function L down :

[0073] L down = 1 - <R(l↓), R(h l ↓)>.

[0074] L rep = ||h - h rec || 1 + ||l - l rec || 1 ,

[0075]

[0076] L adv = -log(D lq (G(w h - Δ))) - log(D hq (G(w l + Δ))),

[0077]

[0078] The face image restoration module can be trained using unpaired images to adapt to various unknown degradations in the real world, and it can generate high-quality images. The loss function L of the entire network is:

[0079] L = λ cc L cc + λ adv L adv + λ perc L perc + λ down L down ,

[0080] where L cc , L adv , L perc and L down are the loss functions of the image restoration module, and λ cc, λ adv , λ perc and λ down are used to balance the weights of several losses.

[0081] Based on the same concept of the above embodiments, another embodiment is also provided. As Figure 2 shown, it is a framework diagram of a face image restoration system according to this embodiment, including: a high-quality feature extraction module, a low-quality feature extraction module, a cross-quality transfer estimation module, and an image restoration module. The high-quality feature extraction module extracts the expression of the input high-quality image in the feature space; the low-quality feature extraction module extracts the expression of the input low-quality image in the feature space; the cross-quality transfer estimation module estimates the transfer vector between the high-quality expression and the low-quality expression in the feature space and uses it to edit the low-quality expression and the high-quality expression. The image restoration module uses an image restoration network to map the edited low-quality expression into an output image.

[0082] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various deformations or modifications within the scope of the claims, which do not affect the essence of the present invention. The above preferred features can be used in any combination without conflict.

Claims

1. A method for face image restoration, characterized in that, comprising: Obtaining the high-quality expression of the input high-quality face image in the feature space by using a high-quality feature extraction network; Obtaining the low-quality expression of the input low-quality face image in the feature space by using a low-quality feature extraction network; Using a cross-quality transfer estimation network to estimate the transfer vector between the high-quality expression and the low-quality expression in the feature space, and using the transfer vector for editing the expression; Using a pre-trained image restoration network to map the edited expression to an output image; Wherein, obtaining the high-quality expression of the input high-quality face image in the feature space by using a high-quality feature extraction network includes: Input high-quality images where C is the number of image channels, W is the width of the image, and H is the height of the image. Using the high-quality feature extraction network E hq Extract image visual features where k is the feature dimension, and N corresponds to the number of network layers of the pre-trained image restoration network; where w h is called high-quality expression; Obtaining the low-quality expression of the input low-quality face image in the feature space by using a low-quality feature extraction network includes: Input low-quality image where C is the number of image channels, W is the width of the image, and H is the height of the image. Using the low-quality feature extraction network E lq Extract the visual features of the image where k is the feature dimension, and N corresponds to the number of network layers of the pre-trained image restoration network; where w l is called the low-quality representation; Wherein, using a cross-quality transfer estimation network to estimate the transfer vector between the high-quality expression and the low-quality expression in the feature space, and using the transfer vector for editing the expression includes: According to the low-quality expression w l and the high-quality expression w h estimate the transfer vector Δ by a cross-quality transfer estimation network in the feature space; Make w l +Δ corresponds to an image in the high-quality image domain discriminated by the discriminator of the cross-quality estimation network, and w h -Δ corresponds to an image in the low-quality image domain discriminated by the discriminator of the cross-quality estimation network; Edit the low-quality expression w by means of the transfer vector Δ l and the high-quality expression w h , and obtain the edited expression w l +Δ, w h -Δ; Wherein, using a pre-trained image restoration network to map the edited expression to an output image includes: The image restoration network G is a pre-trained StyleGAN model that maps the edited expression to the output image, l h = G(w h -Δ), h l = G(w l +Δ).

2. The method for face image restoration according to claim 1, characterized in that, further comprising: Training under the joint loss constraint of the entire model; Using the trained model for face image restoration.

3. The method for face image restoration according to claim 1, characterized in that, Using a pre-trained image restoration network to map the edited expression to an output image further includes: Using the high-quality feature extraction network E hq Obtain the expression of the high-quality face image in the feature space, edit it using the estimated transfer vector Δ, and finally use the image restoration network G to map the edited expression into the output image: Using the low-quality feature extraction network E lq Obtain the expression of the low-quality face image in the feature space, edit it using the estimated transfer vector Δ, and finally use the image restoration network G to map the edited expression to the output image:

4. The method for face image restoration according to claim 3, characterized in that, Using a pre-trained image restoration network G to map the edited expression to an output image further includes: Through the feature extraction network E lq and E hq The low-quality expression w extracted l and the high-quality expression w h can be restored to the low-quality image l through the image restoration network G rec and the high-quality image h rec : l rec = G(w l ), h rec = G(w h ).

5. The method for face image restoration according to claim 4, characterized in that, The loss function of the feature extraction network includes: L down = 1 - <R(l↓), R(h l ↓)> L adv = -log(D lq (G(w h - Δ)))-log(D hq (G(w l + Δ))), Among them, L down is the downsampling loss, R is a face recognition method, ↓ is downsampling, and L cc is the cycle consistency loss, and is the output image, and L adv is the adversarial loss, and L perc is the high-quality loss, is a classification method.

6. The method for face image restoration according to claim 5, characterized in that, The loss function L of the entire network is: L = λ cc L cc + λ adv L adv + λ perc L perc + λ down L down , Among them, L cc , L adv , L perc and L down are the loss functions of the image restoration network, and λ cc , λ adv , λ perc and λ down are the weights for balancing several losses.

7. A face image restoration system, characterized in that, comprising: A high-quality feature extraction module, which extracts the expression of the input high-quality image in the feature space; A low-quality feature extraction module, which extracts the expression of the input low-quality image in the feature space; A cross-quality transfer estimation module, which estimates the transfer vector between the high-quality expression and the low-quality expression in the feature space, and uses it to edit the low-quality expression and the high-quality expression; An image restoration module, which uses a pre-trained image restoration network to map the edited low-quality expression to an output image; Wherein, the high-quality feature extraction module extracting the expression of the input high-quality image in the feature space includes: Input high-quality images where C is the number of image channels, W is the width of the image, and H is the height of the image. Using the high-quality feature extraction network E hq Extract the visual features of the image where k is the feature dimension, and N corresponds to the number of network layers of the pre-trained image restoration network; where w h is called high-quality expression; Wherein, the low-quality feature extraction module extracting the expression of the input low-quality image in the feature space includes: Input low-quality image where C is the number of image channels, W is the width of the image, and H is the height of the image. Using the low-quality feature extraction network E lq Extract the visual features of the image where k is the feature dimension, and N corresponds to the number of network layers of the pre-trained image restoration network; where w l is called the low-quality representation; Wherein, the cross-quality transfer estimation module estimating the transfer vector between the high-quality expression and the low-quality expression in the feature space, and using it to edit the low-quality expression and the high-quality expression includes: According to the low-quality expression w l and the high-quality expression w h estimate the transfer vector Δ by a cross-quality transfer estimation network in the feature space; Make w l The image corresponding to w + Δ is in the high-quality image domain discriminated by the discriminator of the cross-quality estimation network, and w h The image corresponding to w - Δ is in the low-quality image domain discriminated by the discriminator of the cross-quality estimation network; Edit the low-quality expression w by means of the transfer vector Δ l and the high-quality expression w h to obtain the edited expressions w l +Δ, w h -Δ; Among them, the image restoration module uses a pre-trained image restoration network to map the edited low-quality expression into an output image, including: The image restoration network G is a pre-trained StyleGAN model that maps the edited expression to an output image, l h = G(w h -Δ), h l = G(w l +Δ).