Methods and related equipment for processing facial images

By subdividing the face synthesis feature space into subspaces and controlling style vectors, the problem of poor face image quality restoration in existing technologies is solved, achieving high-quality face image restoration and fidelity of face attributes.

CN114648787BActive Publication Date: 2025-10-31HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210130599.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-11
Publication Date
2025-10-31
Estimated Expiration
2042-02-11

AI Technical Summary

Technical Problem

In existing technologies, the quality restoration of face images is poor, which cannot meet the needs of improving face detection and recognition tasks, and does not make full use of prior face knowledge and synthetic features.

Method used

By subdividing the face synthesis feature space, fusing multiple face prior features, and combining a feature encoding module controlled by style vectors, the quality of face images is improved while ensuring the fidelity of face attributes.

Benefits of technology

It achieves high-quality face image restoration with natural details and unchanged face identity and pose information, thus improving the restoration and generalization capabilities of the face generator.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114648787B_ABST
    Figure CN114648787B_ABST
Patent Text Reader

Abstract

This application discloses a method and related apparatus for processing facial images in the field of artificial intelligence. The method includes: acquiring a low-quality facial image and a first clustering label; extracting a first target facial feature and a second target facial feature from the low-quality facial image; classifying each of P third target facial features into R categories of first facial sub-features according to the first clustering label, wherein the P third target facial features are the output of the target convolutional neural network module of a face generator, and the input of the target convolutional neural network module corresponding to the P third target facial features is obtained based on the first target facial features; combining the divided first facial sub-features into a first combined facial feature according to the second target facial feature and the first clustering label; and obtaining a first synthetic facial image based on the first combined facial feature. Using this application can improve the quality of facial images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method and related equipment for processing human face images. Background Technology

[0002] Limited by the imaging hardware performance and image signal processing (ISP) algorithm performance of electronic devices, the quality of currently acquired images remains insufficient. This is especially true for critical image content such as faces, which suffers from low resolution, loss of detail, and blurriness. Furthermore, image compression, downsampling, and interpolation are commonly performed during image storage and transmission, further degrading image quality, particularly facial images. In the consumer product sector, facial image quality restoration is a pressing need, significantly improving the visual appearance of faces and enhancing the accuracy of subsequent face detection and recognition tasks. However, current face restoration (or face enhancement) technologies are ineffective in improving facial image quality and fail to meet the required standards. Summary of the Invention

[0003] This application discloses a method and related equipment for processing facial images, which can improve the quality of facial images.

[0004] In a first aspect, embodiments of this application provide a method for processing face images, comprising: acquiring a low-quality face image and a first clustering label; performing feature extraction on the low-quality face image to obtain a first target face feature and a second target face feature; dividing each of P third target face features into R categories of first face sub-features according to the first clustering label to obtain a set of P first face sub-features, wherein any one of the P first face sub-feature sets includes R categories of first face sub-features, where P is a positive integer and R is an integer greater than 1; the P third target face features are the output of a target convolutional neural network module of a face generator, and the input of the target convolutional neural network module corresponding to the P third target face features is obtained based on the first target face feature; combining the first face sub-features in the set of P first face sub-features into a first combined face feature according to the second target face feature and the first clustering label; and obtaining a first synthetic face image based on the first combined face feature. The first target face feature and the second target face feature have different sizes; optionally, the first target face feature is smaller than the second target face feature. The second target face feature and the first face sub-feature have the same size. The P third target face features correspond to the P sets of first face sub-features, and the set of first face sub-features corresponding to any one of the P third target face features includes R categories of first face sub-features divided from that any one third target face feature. It should be noted that the target convolutional neural network module can have multiple inputs; the input to the target convolutional neural network module obtained from the first target face feature can be a portion of all the inputs to the target convolutional neural network module.

[0005] In this embodiment, for a low-quality face image, feature extraction is performed to obtain a first target face feature and a second target face feature of the low-quality face image. The first target face feature is used as the input to the target convolutional neural network module of the face generator. Based on this input, the target convolutional neural network module can output P third target face features. Then, each of the P third target face features is divided into R categories of first face sub-features according to a first clustering label, resulting in a set of P first face sub-features, where any set of first face sub-features includes R categories of first face sub-features. The first face sub-features in the P sets of first face sub-features are then combined into a first combined face feature based on the second target face feature and the first clustering label. Finally, an enhanced first synthetic face image can be obtained based on the first combined face features. For example, the first combined face features are input into a subsequent module of the target convolutional neural network module in the face generator for processing, ultimately outputting an enhanced, high-quality first synthetic face image. It should be understood that the third target face features constitute a face synthesis feature space. After being divided into R categories of first face sub-features, the first face sub-features of each of the R categories constitute a face synthesis feature subspace. Therefore, the R categories of first face sub-features each constitute R face synthesis feature subspaces. Furthermore, since there are P sets of first face sub-features, and each set of first face sub-features includes the first face sub-features of the R categories, each face synthesis feature subspace contains P first face features, that is, each face synthesis feature subspace contains multiple face prior sub-features. Moreover, the first face sub-features in the P sets of first face sub-features are combined to obtain the first combined face features, that is, the multiple face prior sub-features in each face synthesis feature subspace are fused to obtain more effective face prior features. Therefore, the first synthesized face image recovered based on the first combined face features is the enhanced face image. Thus, in this embodiment, the face synthesis feature space is divided into subspaces to obtain multiple face prior sub-features for each face synthesis feature subspace. Then, the multiple face prior sub-features for each face synthesis feature subspace are combined to obtain more effective face prior features. Face restoration (or face enhancement) is then performed based on the combined face prior features, thereby realizing the use of face prior features during face restoration. This not only improves the quality of face images (e.g., restoring natural details), but also ensures the fidelity and invariance of face attributes (e.g., face identity, pose, etc.).

[0006] In one possible implementation, the P third target face features are obtained by performing convolutional modulation on the target convolutional neural network module based on the first target face features and P first random vectors. The P third target face features correspond to the P first random vectors.

[0007] In this implementation, the target convolutional neural network module is convolved and modulated based on the first target face features and P first random vectors to obtain P third target face features. That is, the P third target face features are the output of the convolutional modulation of the target convolutional neural network module. Since the weights of the convolution kernels in the target convolutional neural network module can be corrected by convolving and modulating the target convolutional neural network module, face restoration based on the P third target face features output by convolutional modulation of the target convolutional neural network module can improve the quality of the face image while ensuring that the face attributes are preserved and unchanged during the face restoration process.

[0008] In one possible implementation, the P third target face features are obtained by convolutional modulation of the target convolutional neural network module based on P target style vectors, and the P target style vectors are obtained based on the first target face features and P first random vectors. Specifically, the P third target face features correspond to the P target style vectors, and the P target style vectors correspond to the P first random vectors.

[0009] In this implementation, the face generator (or target convolutional neural network module) is based on style vector control. For example, P target style vectors are obtained based on the first target face features and P first random vectors. Then, the target convolutional neural network module is convolved and modulated based on the P target style vectors to obtain P third target face features. Finally, face restoration is performed based on the P third target face features. In this way, the controllability, diversity, and robustness of the face prior features can be improved, and the face prior features can be fully utilized during face restoration, thereby improving the face restoration capability of the face generator (e.g., making the detail restoration of the face image richer) and the generalization capability of the face generator.

[0010] In one possible implementation, the P target style vectors are obtained from P first concatenated vectors, which are obtained by concatenating a first feature vector with each of the P first random vectors. The first feature vectors are obtained based on the first target face features. Here, the P target style vectors correspond to the P first concatenated vectors, and the P first concatenated vectors correspond to the P first random vectors.

[0011] In this implementation, the first target face feature is first transformed into a first feature vector; then, the first feature vector is concatenated with P first random vectors to obtain P first concatenated vectors; then, P target style vectors are obtained based on the P first concatenated vectors, for example, by inputting the P first concatenated vectors into the same first fully connected layer to obtain P target style vectors; thus, P target style vectors can be obtained based on the first target face feature and P first random vectors, which is beneficial for the face generator (or target convolutional neural network module) to control based on style vectors.

[0012] In one possible implementation, combining the first face sub-features from the P sets of first face sub-features into a first combined face feature based on the second target face feature and the first clustering label includes: obtaining P sets of first combined weights based on the second target face feature and the P sets of first face sub-features, wherein the P sets of first combined weights correspond to the P sets of first face features, and any one of the P sets of first combined weights includes R first combined weights, wherein the R first combined weights correspond to R categories in the first target face feature set. The first facial feature corresponds to the first target facial feature set, which is the first facial feature set among the P sets of first facial features that corresponds to any one of the first combination weight sets. Any one of the R first combination weights is obtained based on the second target facial feature and the first facial feature of the category corresponding to any one of the first combination weights in the first target facial feature set. The first facial features in the P sets of first facial features are combined into the first combined facial feature based on the first clustering label and the P sets of first combination weights. Specifically, any one of the first combination weights is obtained by performing convolution and pooling operations on the first concatenated feature, with the output of the convolution operation serving as the input of the pooling operation. The first concatenated feature is obtained by concatenating the second target facial feature and the first facial feature corresponding to any one of the first combination weights.

[0013] In this implementation, a first combination weight corresponding to each first face sub-feature is obtained based on the second target face feature and each first face sub-feature. For example, the second target face feature is concatenated with each first face sub-feature, and then the concatenation result of the second target face feature and each first face sub-feature is subjected to convolution and pooling operations to obtain the first combination weight corresponding to each first face feature. Then, each first face sub-feature is combined into a first combination face feature based on the first cluster label and the first combination weight corresponding to each first face feature. In this way, since the first combination weight corresponding to each first face feature is obtained based on the second target face feature and the first face feature, it can be guaranteed that the combined first combination face feature is a more effective face prior feature.

[0014] In one possible implementation, combining the first face features from the P sets of first face sub-features into the first combined face features based on the first clustering label and the P sets of first combination weights includes: obtaining P sets of second face sub-features based on the P sets of first face sub-features and the P sets of first combination weights, wherein the P sets of first face sub-features correspond to the P sets of second face features, any one of the P sets of second face sub-features includes R categories of second face features, the R categories of second face features correspond to R categories of first face features in a second target face sub-feature set, and the second target face sub-feature set is the set of first face features from the P sets of first face features that corresponds to the first face features in the second target face sub-feature set. The first face feature set corresponding to any second face feature set is defined as follows: any second face feature of any of the R categories is obtained by multiplying a first target face feature and a first target combination weight; the first target face feature is the first face feature of the category corresponding to the second face feature of any given category; the first target combination weight is the first combination weight corresponding to the first target face feature; the second face features of the same category in the P sets of second face features are added together to obtain R third face features; the first clustering label is multiplied by the R third face features to obtain R fourth face features; the R fourth face features are combined to form the first combined face feature.

[0015] In this implementation, each first face feature in the P sets of first face features is multiplied by its corresponding first combination weight to obtain a second face feature corresponding to each first face feature. Since there are R categories of first face features, there are R categories of second face features, and each of the R categories has P second face features. The second face features of the same category in the R categories are added together to obtain R third face features. The first cluster label is multiplied by the R third face features to obtain R fourth face features. The R fourth face features are combined to form a first combination face feature. In this way, the first face features in the P sets of first face features can be combined to form a first combination face feature.

[0016] In one possible implementation, the first clustering label is obtained by one-hot encoding of the second clustering label. The second clustering label is obtained by processing the similarity matrix using a preset clustering method. The similarity matrix is ​​obtained based on a first self-expression matrix, which is trained on the second self-expression matrix using multiple first facial features. The multiple first facial features are obtained by inputting multiple second random vectors into the face generator, and these multiple first facial features are the output of the target convolutional neural network module. The multiple first facial features correspond to the multiple second random vectors.

[0017] In this implementation, the first cluster label is obtained by one-hot encoding the second cluster label. The second cluster label is obtained by processing the similarity matrix using a preset clustering method. The similarity matrix is ​​obtained based on the first self-expression matrix, which is trained. Thus, the first cluster label is trained, which is beneficial for segmenting the facial features of the third target.

[0018] In one possible implementation, the first self-expression matrix is ​​obtained through the following operations: For the plurality of first face features, the following operations are performed to obtain the first self-expression matrix: S11: Multiply the fourth target face feature by the first target self-expression matrix to obtain a fourth face feature, wherein the fourth target face feature is one of the plurality of first face features; S12: Obtain a second synthetic face image based on the fourth face feature; S13: Obtain a first loss based on the fourth target face feature and the second synthetic face image; S14: If the first loss is less than a first preset threshold, then... The first target self-expression matrix is ​​the first self-expression matrix; otherwise, adjust the elements in the first target self-expression matrix according to the first loss to obtain the second target self-expression matrix, and execute step S15; S15: take the fifth target face feature as the fourth target face feature, and take the second target self-expression matrix as the first target self-expression matrix, and continue to execute steps S11 to S14, wherein the fifth target face feature is the first face feature among the plurality of first face features that has not yet been used for training; wherein, when executing step S11 for the first time, the first target self-expression matrix is ​​the second self-expression matrix.

[0019] In this implementation, multiple first face features output by the face generator are used to iteratively train the second self-expression matrix, that is, to optimize the second self-expression matrix to obtain the first self-expression matrix; thus, it is beneficial to obtain a suitable second cluster label, and then obtain a suitable first cluster label.

[0020] Secondly, embodiments of this application provide a face image processing apparatus, including a processing unit, configured to: acquire a low-quality face image and a first clustering label; extract features from the low-quality face image to obtain a first target face feature and a second target face feature; divide each of P third target face features into R categories of first face sub-features according to the first clustering label to obtain a set of P first face sub-features, wherein any one of the P first face sub-feature sets includes R categories of first face sub-features, where P is a positive integer and R is an integer greater than 1; the P third target face features are the output of a target convolutional neural network module of a face generator, and the input of the target convolutional neural network module corresponding to the P third target face features is obtained based on the first target face feature; combine the first face sub-features in the set of P first face sub-features into a first combined face feature according to the second target face feature and the first clustering label; and obtain a first synthetic face image based on the first combined face feature.

[0021] In one possible implementation, the P third target face features are obtained by performing convolutional modulation on the target convolutional neural network module based on the first target face features and P first random vectors.

[0022] In one possible implementation, the P third target face features are obtained by performing convolutional modulation on the target convolutional neural network module based on the P target style vectors, and the P target style vectors are obtained based on the first target face features and the P first random vectors.

[0023] In one possible implementation, the P target style vectors are obtained from P first concatenation vectors, which are obtained by concatenating a first feature vector with each of the P first random vectors. The first feature vectors are obtained from the first target face features.

[0024] In one possible implementation, the processing unit is specifically configured to: obtain P sets of first combined weights based on the second target face feature and the P sets of first face sub-features, wherein the P sets of first combined weights correspond to the P sets of first face sub-features, any one of the P sets of first combined weights includes R first combined weights, the R first combined weights correspond to R categories of first face sub-features in the first target face sub-feature set, the first target face sub-feature set is the set of first face features in the P sets of first face features that corresponds to any one of the first combined weights, and any one of the R first combined weights is obtained based on the second target face feature and the first face sub-features in the first target face sub-feature set that correspond to the category of any one of the first combined weights; and combine the first face features in the P sets of first face features into the first combined face feature based on the first clustering label and the P sets of first combined weights.

[0025] In one possible implementation, the processing unit is specifically configured to: obtain P sets of second face features based on the P sets of first face sub-features and the P sets of first combined weights, wherein the P sets of first face features correspond to the P sets of second face features, any one of the P sets of second face features includes R categories of second face features, the R categories of second face features correspond to R categories of first face features in a second target face feature set, and the second target face feature set is the first face feature set in the P sets of first face features that corresponds to any one of the second face feature sets. The second face feature of any one of the R categories of second face features is obtained by multiplying the first target face feature and the first target combination weight. The first target face feature is the first face feature of the category corresponding to the second face feature of any one of the R categories, and the first target combination weight is the first combination weight corresponding to the first target face feature. The second face features of the same category in the P sets of second face features are added together to obtain R third face features. The first clustering label is multiplied by the R third face features to obtain R fourth face features. The R fourth face features are combined to form the first combined face feature.

[0026] In one possible implementation, the first clustering label is obtained by one-hot encoding of the second clustering label, the second clustering label is obtained by processing the similarity matrix using a preset clustering method, the similarity matrix is ​​obtained based on the first self-expression matrix, the first self-expression matrix is ​​obtained by training the second self-expression matrix based on multiple first face features, the multiple first face features are obtained by inputting multiple second random vectors into the face generator respectively, and the multiple first face features are the output of the target convolutional neural network module.

[0027] In one possible implementation, the first self-expression matrix is ​​obtained through the following operations: For the plurality of first face features, the following operations are performed to obtain the first self-expression matrix: S11: Multiply the fourth target face feature by the first target self-expression matrix to obtain a fourth face feature, wherein the fourth target face feature is one of the plurality of first face features; S12: Obtain a second synthetic face image based on the fourth face feature; S13: Obtain a first loss based on the fourth target face feature and the second synthetic face image; S14: If the first loss is less than a first preset threshold, then... The first target self-expression matrix is ​​the first self-expression matrix; otherwise, adjust the elements in the first target self-expression matrix according to the first loss to obtain the second target self-expression matrix, and execute step S15; S15: take the fifth target face feature as the fourth target face feature, and take the second target self-expression matrix as the first target self-expression matrix, and continue to execute steps S11 to S14, wherein the fifth target face feature is the first face feature among the plurality of first face features that has not yet been used for training; wherein, when executing step S11 for the first time, the first target self-expression matrix is ​​the second self-expression matrix.

[0028] It should be noted that the beneficial effects of the second aspect can be found in the description of the first aspect, and will not be repeated here.

[0029] Thirdly, embodiments of this application provide an electronic device including a processor, a memory, a transceiver, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the processor, and the programs include instructions for performing steps in the method as described in any one of the first aspects above.

[0030] Fourthly, embodiments of this application provide a chip, including: a processor, configured to call and run a computer program from a memory, causing a device equipped with the chip to perform the method described in any of the first aspects above.

[0031] Fifthly, embodiments of this application provide a computer-readable storage medium storing a computer program for electronic data interchange, wherein the computer program causes a computer to perform the method as described in any one of the first aspects above.

[0032] In a sixth aspect, embodiments of this application provide a computer program product that causes a computer to perform the method as described in any one of the first aspects above. Attached Figure Description

[0033] The accompanying drawings used in the embodiments of this application are described below.

[0034] Figure 1 This is a schematic diagram of the structure of a face generator based on a generative adversarial network (GAN) provided in an embodiment of this application;

[0035] Figure 2 yes Figure 1 The diagram shows the structure of a generative adversarial network module.

[0036] Figure 3 Based on Figure 1 The diagram shows the structure of the face reconstruction network in the face generator.

[0037] Figure 4 This is a comparative diagram of face restoration solutions;

[0038] Figure 5 This is a schematic diagram of the system architecture provided in the embodiments of this application;

[0039] Figure 6 This is a schematic flowchart of a face image processing method provided in an embodiment of this application;

[0040] Figure 7 This is a schematic diagram of the data flow for processing a face image provided in an embodiment of this application;

[0041] Figure 8 This is a schematic diagram of the training phase of a face recovery network provided in an embodiment of this application;

[0042] Figure 9 yes Figure 8 A schematic diagram of the inference phase of a face reconstruction network is shown.

[0043] Figure 10 yes Figure 8 A schematic diagram of the training phase of an exemplary structure of a face reconstruction network shown;

[0044] Figure 11 yes Figure 10 A schematic diagram of the inference phase of a face reconstruction network is shown.

[0045] Figure 12 This is a schematic diagram of the structure of a face image processing device provided in an embodiment of this application;

[0046] Figure 13 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;

[0047] Figure 14 This is a schematic diagram of the structure of a computer program product provided in an embodiment of this application. Detailed Implementation

[0048] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.

[0049] The terms "comprising" and "having," and any variations thereof, in the specification, claims, and accompanying drawings of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the steps or units listed, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such process, method, product, or apparatus.

[0050] In this specification, the term "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0051] First, some of the technical terms used in this application will be explained to help those skilled in the art understand this application.

[0052] (1) Peak Signal-to-Noise Ratio (PSNR): An engineering term representing the ratio of the maximum possible power of a signal to the power of destructive noise that affects its representation accuracy. PSNR is frequently used as a measurement method for signal reconstruction quality in fields such as image processing, and is usually simply defined by mean square error.

[0053] (2) Structural Similarity Index (SSIM): This index measures the similarity between two images and is used to evaluate the quality of the output image processed by the algorithm. From the perspective of image composition, the SSIM defines structural information as an attribute reflecting the structure of objects in a scene, independent of brightness and contrast, and models distortion as a combination of three different factors: brightness, contrast, and structure. It uses the mean as an estimate of brightness, the standard deviation as an estimate of contrast, and the covariance as a measure of structural similarity.

[0054] (3) Learned Perceptual Image Patch Similarity (LPIPS): This metric measures the difference between two images. It learns the inverse mapping from the generated image to the ground truth image, forcing the generator to learn the inverse mapping to reconstruct the real image from the fake image, and prioritizing the perceptual similarity between them. A lower LPIPS value indicates a more similar pair of images; conversely, a higher LPIPS value indicates a greater difference between the two images.

[0055] (4) Natural Image Quality Evaluator (NIQE): This is a non-referenced evaluation metric for measuring image quality. It extracts features from natural landscapes to test the image. These features are fitted into a multivariate Gaussian model. This Gaussian model actually measures the difference of a test image on a multivariate distribution, which is constructed from a series of features extracted from normal natural images.

[0056] (5) Fréchet Inception Distance (FID): This is an objective metric used to evaluate the quality of images created by generative models. It measures the similarity between two sets of images by the statistical similarity of the computer vision features of the original images, which are calculated using an image classification model based on Convolutional Neural Networks (CNN). The lower the Fréchet Inception Distance, the more similar the two sets of images are.

[0057] (6) Face restoration (also known as face enhancement): refers to the technology of processing the color and texture of images containing human faces to meet specific indicators.

[0058] (7) Artifacts: In image quality enhancement tasks, obvious errors or anomalies appear in the image after neural network enhancement; among them, the errors or anomalies include: obvious color abnormalities in areas that should have normal color and natural details, and obvious errors in image details, etc.

[0059] (8) Face Generator: A generative model based on neural networks. A random or fixed vector is input into the face generator, which then outputs realistic, natural, high-quality face images. The face synthesis feature space refers to the space composed of face synthesis features (also called face generator features), that is, the feature space containing all face generator features. The face generator is a multi-layer convolutional neural network, and the face synthesis features are the feature tensors generated by convolutional operations in each layer of the convolutional neural network.

[0060] (9) Style vector: A vector of intermediate generated variables commonly found in generative networks, used to scale the weights of convolutional kernels.

[0061] Secondly, some problems with deep learning methods in face restoration tasks are analyzed to facilitate understanding of this application.

[0062] Deep learning methods, especially those based on convolutional neural networks, have achieved industry-leading performance in image restoration and enhancement, gradually surpassing traditional algorithms. However, for face restoration tasks, deep learning methods still face some unresolved issues. These are analyzed in detail below:

[0063] (1) Prior knowledge of faces is not fully utilized. There is a lot of prior knowledge about faces (such as the relatively fixed structure of faces and the unchanging relative positions of facial features). However, many existing general image enhancement and super-resolution methods based on convolutional neural networks do not utilize this prior knowledge of faces, resulting in problems such as poor face restoration details, serious artificial traces, and poor enhancement effects.

[0064] (2) Poor generalization of convolutional neural network models. In real-world applications, the facial images obtained after acquisition, image processing and transmission have complex and varied degradation patterns. Convolutional neural network models trained on limited data cannot effectively recover facial images after different degradations, resulting in convolutional neural network models being unable to adapt to open and diverse scenarios.

[0065] (3) Face restoration methods based on convolutional neural networks suffer from performance issues such as long processing time, large storage requirements, and high power consumption. Some methods combine traditional algorithms with deep learning methods, introducing prior knowledge of faces based on dictionary matching. However, this method only targets specific facial features, and is severely affected by facial pose and lighting, resulting in long online matching times and high memory consumption. Meanwhile, some methods utilize face generators to map low-quality input faces to a face synthesis feature space to obtain effective face synthesis features and restore the face image. However, the distribution of synthesized faces is inconsistent with that of real faces, making it difficult to obtain features that match the input in the face synthesis feature space, leading to problems such as changes in face identity and obvious artificial traces.

[0066] In summary, there is an urgent need to propose more effective and comprehensive ways to utilize prior facial knowledge in real-world scenarios and to design high-quality and efficient face recovery solutions for all scenarios, addressing both real-world scenarios and unknown complex degradation.

[0067] Furthermore, to facilitate understanding of the embodiments of this application, several related technical solutions for face recovery are exemplarily introduced.

[0068] Related technical solution 1: Based on a general image enhancement convolutional neural network.

[0069] The core idea of ​​the face enhancement method based on a general image enhancement convolutional neural network is relatively simple. The network structure consists of N cascaded convolutional layers, and the network output size is K times the input size (K≥1). It utilizes visual perception loss and adversarial loss to enhance the details and texture of the final output. However, this method does not utilize prior knowledge of faces, resulting in insufficient restoration of facial details, significant artifacts, and poor restoration effects.

[0070] Related technical solution 2: Convolutional neural network based on offline dictionary matching.

[0071] The core idea of ​​the face enhancement method based on offline dictionary matching using convolutional neural networks is as follows: In the face dictionary generation stage, VGG (Visual Geometry Group Network) features of high-quality face images are extracted, and a feature dictionary for the facial feature regions is generated offline. In the face restoration stage, a ground truth (Unet) structure is used to extract VGG features from degraded face images, and these features are matched with the generated feature dictionary to correct the features of the facial features, ultimately obtaining the restored face. This method has several shortcomings: First, it can only generate dictionaries for some facial organs, resulting in poor restoration of areas such as hair and skin; second, dictionary loading and online matching consume significant time and memory resources; and third, it is based on traditional matching methods and lacks robustness, with changes in facial pose and lighting severely affecting the method's performance.

[0072] Related technical solution 3: Convolutional neural network based on face generator.

[0073] Face enhancement methods based on face generators represent the latest technological trend. Figure 1This is a schematic diagram of a face generator based on a generative adversarial network (GAN). The network structure of this face generator includes a mapping network (M network) and a GAN block (G network). The M network generates intermediate hidden variables ω from the hidden variable z, where ω controls the style of the synthesized image. The hidden variable z is a random vector z following a Gaussian distribution. The G network generates the synthesized image. Figure 2 As shown, the Generative Adversarial Network (GAN) module provides each sub-network layer with inputs A and B; A is an affine transformation obtained from the ω transformation, used to control the style of the generated image; B is the transformed random noise broadcast, which is used to enrich the details of the generated image, meaning each convolutional layer can adjust its style based on the input A. Figure 3 As shown, the core idea of ​​this method is to pre-train a face generator, then input the degraded face into a feature extraction module, and use the extracted features to control the face generator to obtain a generated face as the final restored result. The network parameters of the face generator may change during this process. This method has several shortcomings: First, the distribution of the generated face is inconsistent with that of the real face, making it difficult to effectively utilize the synthesized face features; second, faces that have undergone complex degradation processes are difficult to map to the synthesized face feature space, leading to changes in the identity and other information of the finally restored face; furthermore, the utilization and fusion of synthesized face features are relatively simple and require further optimization.

[0074] In view of the problems that the face restoration schemes provided by the above-mentioned related technologies have difficulties in utilizing prior knowledge of faces and insufficient utilization of face synthesis features, the embodiments of this application provide a face restoration scheme based on multi-subspace prior synthesis.

[0075] Specifically, addressing the difficulty of utilizing prior face knowledge in current face restoration methods, this application designs a face restoration network framework based on multiple mappings of the face synthesis feature subspace. The face synthesis feature space is subdivided to obtain multiple prior face features in each subspace, which are then fused to obtain more effective prior face features. This improves the quality of face restoration while ensuring the fidelity and invariance of information such as face identity and pose. For example, it ensures the fidelity and invariance of information such as face identity and pose while making the restored face image details natural and rich. Furthermore, addressing the problem of insufficient utilization of prior face features, this application designs a feature encoding module based on style vector control, which can improve the controllability, diversity, and robustness of prior face features, thereby enhancing the face restoration capability and generalization ability of the face generator.

[0076] like Figure 4As shown, the core technologies of the technical solution provided in this application include at least the following: First, unlike the aforementioned related technical solutions which are based on a single feature mapping of the face synthesis feature space, this application provides a face restoration network framework based on multiple feature mappings of the face synthesis feature subspace, ensuring the face restoration quality of the face restoration scheme based on the face generator in real open scenarios. Second, unlike the feature encoding module in the aforementioned related technical solutions which is used to generate random variables or latent variables, the feature encoding module in this application is used to generate style vectors, ensuring that the obtained synthesis space features are more diverse, more effective, and more controllable.

[0077] The technical solution provided in this application will be described in detail below with reference to specific implementation methods.

[0078] Please see Figure 5 , Figure 5 This application provides a system architecture 50. As shown in the system architecture 50, the data acquisition device 56 is used to acquire training data. In this application embodiment, the training data includes at least one of the following: a second random vector, a first face feature, a first face image, and a third random vector. The training data is stored in a database 53, and the training device 52 trains a target model / rule 513 based on the training data maintained in the database 53. The following describes in more detail how the training device 52 obtains the target model / rule 513 based on the training data. This target model / rule 513 can be used to implement the face image processing method provided in this application embodiment, that is, by inputting a low-quality face image and the first random vector into the target model / rule 513, a first synthetic face image can be obtained. In this application embodiment, the target model / rule 513 can specifically be a face recovery network. It should be noted that in actual applications, the training data maintained in the database 53 may not all come from the acquisition by the data acquisition device 56, but may also be received from other devices. It should also be noted that the training device 52 may not necessarily train the target model / rule 513 entirely based on the training data maintained by the database 53. It may also obtain training data from the cloud or other places for model training. The above description should not be construed as a limitation on the embodiments of this application.

[0079] The target model / rule 513 trained using training device 52 can be applied to different systems or devices, such as... Figure 3 The execution device 51 shown can be a terminal, such as a mobile phone terminal, tablet computer, laptop computer, AR / VR, vehicle terminal, etc., or it can be a server or cloud platform. Figure 3In the process, the execution device 51 is equipped with an I / O interface 512 for data interaction with external devices. The user can input data to the I / O interface 512 through the client device 54. The input data in this embodiment may include: a low-quality face image, a first random vector, and other random vectors.

[0080] During the calculation and related processing performed by the calculation module 511 of the execution device 51, the execution device 51 can call data, code, etc. in the data storage system 55 for corresponding processing, and can also store the data, instructions, etc. obtained from the corresponding processing into the data storage system 55.

[0081] Finally, the I / O interface 512 returns the processing result, such as the first synthetic face image obtained above, to the client device 54, thereby providing it to the user.

[0082] It is worth noting that the training device 52 can generate corresponding target models / rules 513 based on different training data for different objectives or tasks. The corresponding target models / rules 513 can be used to achieve the above objectives or complete the above tasks, thereby providing the user with the required results.

[0083] In the appendix Figure 3 In the scenario shown, the user can manually provide input data, which can be done through the interface provided by I / O interface 512. Alternatively, the client device 54 can automatically send input data to I / O interface 512. If user authorization is required for the client device 54 to automatically send input data, the user can set the corresponding permissions in the client device 54. The user can view the output results of execution device 51 on the client device 54, which can be presented in various forms such as display, sound, or animation. The client device 54 can also act as a data acquisition terminal, collecting the input data and output results of the input I / O interface 512 as shown in the figure, and storing them as new sample data in database 53. Alternatively, data can be collected directly from the I / O interface 512 without going through the client device 54, using the input data and output results of the input I / O interface 512 as shown in the figure, and storing them as new sample data in database 53.

[0084] It is worth noting that, attached Figure 3 This is merely a schematic diagram of a system architecture provided in an embodiment of this application. The positional relationships between the devices, components, modules, etc., shown in the diagram do not constitute any limitation. For example, in the attached diagram... Figure 3 In this context, the data storage system 55 is an external memory relative to the execution device 51. In other cases, the data storage system 55 may also be placed within the execution device 51.

[0085] like Figure 3 As shown, a target model / rule 513 is obtained by training the training device 52. In this embodiment of the application, the target model / rule 513 may be a face recovery network, etc.

[0086] Optionally, in this application, the execution device 51 and the training device 52 can be the same electronic device.

[0087] Please see Figure 6 , Figure 6 This application provides a method for processing facial images, which can be executed by an electronic device. The method is described as a series of steps or operations. It should be understood that the method can be executed in various orders and / or occur simultaneously, and is not limited to... Figure 6 The execution order is shown. Furthermore, Figure 6 The method shown can be combined Figure 7 To understand, Figure 7 This is a schematic diagram of the data flow for processing a face image provided in an embodiment of this application; Figure 6 The methods shown include, but are not limited to, the following steps or operations:

[0088] 601: Obtain low-quality face images and first cluster labels.

[0089] 602: Extract features from the low-quality face image to obtain the first target face features and the second target face features.

[0090] The first target face feature and the second target face feature have different sizes; optionally, the first target face feature is smaller than the second target face feature; further optionally, feature extraction from a low-quality face image can yield two or more face features, with the first target face feature being the smallest among them.

[0091] It should be understood that the size of the first target facial feature or the second target facial feature refers to the width and height of the first target facial feature or the second target facial feature, expressed as width × height; in addition, the size of features or images described elsewhere in this application refers to width × height.

[0092] 603: Based on the first clustering label, each of the P third target face features is divided into R categories of first face sub-features to obtain a set of P first face sub-features. Any one of the P first face sub-feature sets includes R categories of first face sub-features, where P is a positive integer and R is an integer greater than 1. The P third target face features are the output of the target convolutional neural network module of the face generator, and the input of the target convolutional neural network module corresponding to the P third target face features is obtained based on the first target face features.

[0093] Wherein, the P third target face features correspond to the P sets of first face sub-features, and the set of first face sub-features corresponding to any one of the P third target face features includes R categories of first face sub-features obtained by dividing the any one third target face feature.

[0094] It should be noted that the target convolutional neural network module can also have multiple inputs. The input of the target convolutional neural network module obtained from the first target face features can be a part of all the inputs of the target convolutional neural network module. The target convolutional neural network module processes all its inputs (including the input obtained from the first target face features) to obtain P third target face features.

[0095] The process of dividing each of the P third-target face features into R categories of first-face sub-features based on the first cluster label can be found in [reference needed]. Figure 7 .

[0096] 604: Based on the second target face features and the first clustering label, combine the first face features in the P sets of first face features into a first combined face feature.

[0097] The second target facial feature can be the same size as the first facial sub-feature.

[0098] 605: Obtain the first synthetic face image based on the first combination of facial features.

[0099] The first synthesized face image is a high-quality face image recovered from the aforementioned low-quality face image, or in other words, the first synthesized face image is a face image enhanced from the aforementioned low-quality face image.

[0100] It should be noted that the face generator has a multi-module or multi-layer structure. The target convolutional neural network module can be one of the modules or one layer of the face generator. The face generator also includes subsequent structures connected to the target convolutional neural network module. This application can input the first combination of face features into the subsequent structures connected to the target convolutional neural network module in the face generator. The final output of the face generator is the first synthetic face image.

[0101] In this embodiment, for a low-quality face image, feature extraction is performed to obtain a first target face feature and a second target face feature of the low-quality face image. The first target face feature is used as the input to the target convolutional neural network module of the face generator. Based on this input, the target convolutional neural network module can output P third target face features. Then, each of the P third target face features is divided into R categories of first face sub-features according to a first clustering label, resulting in a set of P first face sub-features, where any set of first face sub-features includes R categories of first face sub-features. The first face sub-features in the P sets of first face sub-features are then combined into a first combined face feature based on the second target face feature and the first clustering label. Finally, an enhanced first synthetic face image can be obtained based on the first combined face features. For example, the first combined face features are input into a subsequent module of the target convolutional neural network module in the face generator for processing, ultimately outputting an enhanced, high-quality first synthetic face image. It should be understood that the third target face features constitute a face synthesis feature space. After being divided into R categories of first face sub-features, the first face sub-features of each of the R categories constitute a face synthesis feature subspace. Therefore, the R categories of first face sub-features each constitute R face synthesis feature subspaces. Furthermore, since there are P sets of first face sub-features, and each set of first face sub-features includes the first face sub-features of the R categories, each face synthesis feature subspace contains P first face features, that is, each face synthesis feature subspace contains multiple face prior sub-features. Moreover, the first face sub-features in the P sets of first face sub-features are combined to obtain the first combined face features, that is, the multiple face prior sub-features in each face synthesis feature subspace are fused to obtain more effective face prior features. Therefore, the first synthesized face image recovered based on the first combined face features is the enhanced face image. Thus, in this embodiment, the face synthesis feature space is divided into subspaces to obtain multiple face prior sub-features for each face synthesis feature subspace. Then, the multiple face prior sub-features for each face synthesis feature subspace are combined to obtain more effective face prior features. Face restoration (or face enhancement) is then performed based on the combined face prior features, thereby realizing the use of face prior features during face restoration. This not only improves the quality of face images (e.g., restoring natural details), but also ensures the fidelity and invariance of face attributes (e.g., face identity, pose, etc.).

[0102] In one possible implementation, the P third target face features are obtained by performing convolutional modulation on the target convolutional neural network module based on the first target face features and P first random vectors.

[0103] like Figure 7 As shown, the input of the target convolutional neural network module to the target convolutional neural network module includes using the input to perform convolution modulation on the target convolutional neural network module; for example, the input of the target convolutional neural network module is obtained according to the first target face features and P first random vectors, and the input is used to perform convolution modulation on the target convolutional neural network module, so that the target convolutional neural network module outputs P third target face features.

[0104] Among them, P third target facial features correspond to P first random vectors.

[0105] The first random vector is the intermediate hidden variable ω output by the M network of the face generator; for example, by inputting a random vector z that follows a Gaussian distribution into the M network of the face generator, the output of the M network of the face generator is the first random vector.

[0106] In this implementation, the target convolutional neural network module is convolved and modulated based on the first target face features and P first random vectors to obtain P third target face features. That is, the P third target face features are the output of the convolutional modulation of the target convolutional neural network module. Since the weights of the convolution kernels in the target convolutional neural network module can be corrected by convolving and modulating the target convolutional neural network module, face restoration based on the P third target face features output by convolutional modulation of the target convolutional neural network module can improve the quality of the face image while ensuring that the face attributes are preserved and unchanged during the face restoration process.

[0107] In one possible implementation, the P third target face features are obtained by performing convolutional modulation on the target convolutional neural network module based on the P target style vectors, and the P target style vectors are obtained based on the first target face features and the P first random vectors.

[0108] Among them, P third target face features correspond to P target style vectors, and P target style vectors correspond to P first random vectors.

[0109] like Figure 7 As shown, the input to the target convolutional neural network module includes a target style vector. P target style vectors are obtained based on the first target face features and P first random vectors. The P target style vectors are input into the target convolutional neural network module to perform convolution modulation on the target convolutional neural network module, thereby obtaining P third target face features.

[0110] In this implementation, the face generator (or target convolutional neural network module) is based on style vector control. For example, P target style vectors are obtained based on the first target face features and P first random vectors. Then, the target convolutional neural network module is convolved and modulated based on the P target style vectors to obtain P third target face features. Finally, face restoration is performed based on the P third target face features. In this way, the controllability, diversity, and robustness of the face prior features can be improved, and the face prior features can be fully utilized during face restoration, thereby improving the face restoration capability of the face generator (e.g., making the detail restoration of the face image richer) and the generalization capability of the face generator.

[0111] In one possible implementation, the P target style vectors are obtained from P first concatenation vectors, which are obtained by concatenating a first feature vector with each of the P first random vectors. The first feature vectors are obtained from the first target face features.

[0112] Among them, the P target style vectors correspond to the P first concatenation vectors, and the P first concatenation vectors correspond to the P first random vectors.

[0113] In this implementation, the first target face feature is first transformed into a first feature vector; then, the first feature vector is concatenated with P first random vectors to obtain P first concatenated vectors; then, P target style vectors are obtained based on the P first concatenated vectors, for example, by inputting the P first concatenated vectors into the same first fully connected layer to obtain P target style vectors; thus, P target style vectors can be obtained based on the first target face feature and P first random vectors, which is beneficial for the face generator (or target convolutional neural network module) to control based on style vectors.

[0114] In one possible implementation, combining the first face sub-features from the P sets of first face sub-features into a first combined face feature based on the second target face feature and the first clustering label includes: obtaining P sets of first combined weights based on the second target face feature and the P sets of first face sub-features, wherein the P sets of first combined weights correspond to the P sets of first face features, and any one of the P sets of first combined weights includes R first combined weights, wherein the R first combined weights correspond to R categories in the first target face feature set. The first face feature corresponds to the first target face feature set, which is the first face feature set among the P sets of first face features that corresponds to any one of the first combination weight sets. Any one of the R first combination weights is obtained based on the second target face feature and the first face feature of the category corresponding to any one of the first combination weights in the first target face feature set. The first face features in the P sets of first face features are combined into the first combination face feature based on the first clustering label and the P sets of first combination weights.

[0115] Wherein, any one of the first combined weights is obtained by performing convolution and pooling operations on the first concatenated feature, the output of the convolution operation is the input of the pooling operation, and the first concatenated feature is obtained by concatenating the second target face feature and the first face sub-feature corresponding to any one of the first combined weights.

[0116] In this implementation, a first combination weight corresponding to each first face sub-feature is obtained based on the second target face feature and each first face sub-feature. For example, the second target face feature is concatenated with each first face sub-feature, and then the concatenation result of the second target face feature and each first face sub-feature is subjected to convolution and pooling operations to obtain the first combination weight corresponding to each first face feature. Then, each first face sub-feature is combined into a first combination face feature based on the first cluster label and the first combination weight corresponding to each first face feature. In this way, since the first combination weight corresponding to each first face feature is obtained based on the second target face feature and the first face feature, it can be guaranteed that the combined first combination face feature is a more effective face prior feature.

[0117] In one possible implementation, combining the first face features from the P sets of first face sub-features into the first combined face features based on the first clustering label and the P sets of first combination weights includes: obtaining P sets of second face sub-features based on the P sets of first face sub-features and the P sets of first combination weights, wherein the P sets of first face sub-features correspond to the P sets of second face features, any one of the P sets of second face sub-features includes R categories of second face features, the R categories of second face features correspond to R categories of first face features in a second target face sub-feature set, and the second target face sub-feature set is the set of first face features from the P sets of first face features that corresponds to the first face features in the second target face sub-feature set. The first face feature set corresponding to any second face feature set is defined as follows: any second face feature of any of the R categories is obtained by multiplying a first target face feature and a first target combination weight; the first target face feature is the first face feature of the category corresponding to the second face feature of any given category; the first target combination weight is the first combination weight corresponding to the first target face feature; the second face features of the same category in the P sets of second face features are added together to obtain R third face features; the first clustering label is multiplied by the R third face features to obtain R fourth face features; the R fourth face features are combined to form the first combined face feature.

[0118] In this implementation, each first face feature in the P sets of first face features is multiplied by its corresponding first combination weight to obtain a second face feature corresponding to each first face feature. Since there are R categories of first face features, there are R categories of second face features, and each of the R categories has P second face features. The second face features of the same category in the R categories are added together to obtain R third face features. The first cluster label is multiplied by the R third face features to obtain R fourth face features. The R fourth face features are combined to form a first combination face feature. In this way, the first face features in the P sets of first face features can be combined to form a first combination face feature.

[0119] In one possible implementation, the first clustering label is obtained by one-hot encoding of the second clustering label, the second clustering label is obtained by processing the similarity matrix using a preset clustering method, the similarity matrix is ​​obtained based on the first self-expression matrix, the first self-expression matrix is ​​obtained by training the second self-expression matrix based on multiple first face features, the multiple first face features are obtained by inputting multiple second random vectors into the face generator respectively, and the multiple first face features are the output of the target convolutional neural network module.

[0120] The plurality of first facial features correspond to the plurality of second random vectors.

[0121] It should be noted that the first random vector and the second random vector are different random vectors; the first random vector is the intermediate hidden variable ω output by the M network of the face generator, that is, the first random vector is a random vector processed by the M network of the face generator; while the second random vector is a random vector that has not been processed by the M network of the face generator. For example, the second random vector is a random vector z that follows a Gaussian distribution.

[0122] In this implementation, the first cluster label is obtained by one-hot encoding the second cluster label. The second cluster label is obtained by processing the similarity matrix using a preset clustering method. The similarity matrix is ​​obtained based on the first self-expression matrix, which is trained. Thus, the first cluster label is trained, which is beneficial for segmenting the facial features of the third target.

[0123] In one possible implementation, the first self-expression matrix is ​​obtained through the following operations: For the plurality of first face features, the following operations are performed to obtain the first self-expression matrix: S11: Multiply the fourth target face feature by the first target self-expression matrix to obtain a fourth face feature, wherein the fourth target face feature is one of the plurality of first face features; S12: Obtain a second synthetic face image based on the fourth face feature; S13: Obtain a first loss based on the fourth target face feature and the second synthetic face image; S14: If the first loss is less than a first preset threshold, then... The first target self-expression matrix is ​​the first self-expression matrix; otherwise, adjust the elements in the first target self-expression matrix according to the first loss to obtain the second target self-expression matrix, and execute step S15; S15: take the fifth target face feature as the fourth target face feature, and take the second target self-expression matrix as the first target self-expression matrix, and continue to execute steps S11 to S14, wherein the fifth target face feature is the first face feature among the plurality of first face features that has not yet been used for training; wherein, when executing step S11 for the first time, the first target self-expression matrix is ​​the second self-expression matrix.

[0124] In this implementation, multiple first face features output by the face generator are used to iteratively train the second self-expression matrix, that is, to optimize the second self-expression matrix to obtain the first self-expression matrix; thus, it is beneficial to obtain a suitable second cluster label, and then obtain a suitable first cluster label.

[0125] It should be noted that, Figure 6 The face image processing method shown can be implemented based on a face reconstruction network. The following is an example of a method for implementing this method. Figure 6 The face image processing method shown is a face restoration network.

[0126] Please see Figure 8 , Figure 8 This is a schematic diagram of the structure of a face restoration network provided in an embodiment of this application. The face restoration network includes a feature encoder 100, a style vector control module 200, a face generator 300, a face subspace clustering and partitioning module 400, a multi-face feature mapping module 500, and a multi-face feature combination module 600. Among them, the face subspace clustering and partitioning module 400 includes a face subspace partitioning unit 410, a similarity matrix learning unit 420, and a face subspace clustering unit 430.

[0127] The training of this face reconstruction network is conducted in two phases, as follows: The face subspace clustering and partitioning module 400 participates in the first training phase, while the feature encoder 100, style vector control module 200, face generator 300, multi-face feature mapping module 500, and multi-face feature combination module 600 participate in the second training phase; alternatively, the similarity matrix learning unit 420 and the face subspace clustering unit 430 participate in the first training phase, while the face subspace partitioning unit 410, feature encoder 100, style vector control module 200, face generator 300, multi-face feature mapping module 500, and multi-face feature combination module 600 participate in the second training phase. These are described in detail below.

[0128] I. First Training Phase:

[0129] The training samples in the first training phase include multiple first face features output by the face generator 300. These multiple first face features are several intermediate results output by the face generator 300. The multiple first face features are obtained by inputting multiple second random vectors into the face generator 300, with each first face feature corresponding one-to-one with a second random vector. For example, the multiple second random vectors can be multiple random vectors z following a Gaussian distribution. The first training phase also involves using these multiple first face features to iteratively train the face subspace clustering and partitioning module 400 multiple times to obtain first cluster labels; or, the first training phase also involves using these multiple first face features to iteratively train the similarity matrix learning unit 420 and the face subspace clustering unit 430 multiple times to obtain second cluster labels.

[0130] Among them, the face generator 300 can be a pre-trained face generator, including but not limited to style-based generator network (stylegan), second-generation style-based generator network (stylegan2), etc.

[0131] It should be noted that, since the face generator 300 is a multi-layer network structure, the first face feature includes the output of one or more intermediate layers in the face generator 300; and whether the first face feature specifically includes the output of one intermediate layer or multiple intermediate layers in the face generator 300 is determined according to actual needs; furthermore, which intermediate layer or layers of the face generator 300 the first face feature specifically includes is also determined according to actual needs.

[0132] The first training phase is described below using the example of multiple first face features being the output of one of the intermediate layers in the face generator 300:

[0133] Step 1: For multiple first facial features, perform the following operations to obtain the first self-expression matrix:

[0134] S11: Input the fourth target face feature into the similarity matrix learning unit 420 to obtain the fourth face feature, which is one of multiple first face features.

[0135] The similarity matrix learning unit 420 is used to: receive features and features Perform a matrix multiplication operation with the self-expression matrix C to obtain the features. Among them, features The output features of the face generator are represented by k∈{1,2,…,P}, i∈{1,2,…,Q}; the self-expression matrix C has a matrix dimension of N. g ×N g N g Features The clustering dimension, such as channel dimension or spatial dimension, is used to describe the features. The self-expression matrix C is used to describe the features. The degree of similarity in the channel dimension or spatial dimension.

[0136] It should be noted that if multiple features If all the outputs are from one of the intermediate layers in the face generator 300, then there is only one self-expression matrix C. The output of this intermediate layer corresponds to this single self-expression matrix C. During the first training phase, multiple features... Each of the features is multiplied by the self-expression matrix C; if multiple features If the output of a face generator 300 consists of multiple intermediate layers, then there are multiple self-expression matrices C, and the output of each intermediate layer corresponds one-to-one with the multiple self-expression matrices C, representing multiple features. Each of them is multiplied by its corresponding self-expression matrix C.

[0137] For example, since multiple first face features are the output of one of the intermediate layers (e.g., the target convolutional neural network module) in the face generator 300, the similarity matrix learning unit 420 is specifically used to: multiply the fourth target face feature with the first target self-expression matrix to obtain the fourth face feature. At this time, the features... For the fourth target's facial features, the self-expression matrix C is the self-expression matrix of the first target, and the features... This is the fourth facial feature.

[0138] It should be understood that the training objective of the similarity matrix learning unit 420 in the first training phase is to obtain the first self-expression matrix; that is, to continuously adjust the elements in the first target self-expression matrix to obtain the first self-expression matrix. Specifically, during the first training of the similarity matrix learning unit 420, the first target self-expression matrix is ​​the initial self-expression matrix, for example, the second self-expression matrix.

[0139] S12: Input the fourth facial feature into the face generator 300 to obtain the second synthetic face image.

[0140] The similarity matrix learning unit 420 is also used to: transfer features Input face generator 300, and face generator 300 outputs a synthesized face image.

[0141] For example, features For the fourth face feature, the similarity matrix learning unit 420 is also used to: input the fourth face feature into the face generator 300 to obtain the second synthetic face image. At this time, the features... The fourth facial feature is used as the basis for the synthetic facial image output by the face generator 300, which is the second synthetic facial image.

[0142] S13: Obtain the first loss based on the fourth target face features and the second synthesized face image.

[0143] Among them, the first loss is through features It uses synthetic face images for calculation.

[0144] For example, features Given the fourth target face feature and the second synthesized face image, the first loss is calculated based on the fourth target face feature and the second synthesized face image.

[0145] S14: If the first loss is less than the first preset threshold, then the first target self-expression matrix is ​​the first self-expression matrix; otherwise, adjust the elements in the first target self-expression matrix according to the first loss to obtain the second target self-expression matrix, and execute step S15.

[0146] The first loss is used to adjust the elements in the self-expression matrix C to obtain the updated self-expression matrix C.

[0147] For example, if the self-expression matrix C is the first target self-expression matrix, then the first loss is used to adjust the elements in the first target self-expression matrix to obtain the second target self-expression matrix, which is also the updated self-expression matrix C.

[0148] S15: Take the fifth target face feature among the multiple first face features as the fourth target face feature, and take the second target self-expression matrix as the first target self-expression matrix, and continue to execute steps S11 to S14. The fifth target face feature is the first face feature among the multiple first face features that has not yet been input into the face generator 300.

[0149] Step 2: Input the first self-expression matrix into the face subspace clustering unit 430 to obtain the second clustering label.

[0150] The face subspace clustering unit 430 is used to: receive the self-expression matrix C output by the similarity matrix learning unit 420, that is, to take the self-expression matrix C as input; process the self-expression matrix C to obtain the similarity matrix A; and process the similarity matrix through a preset clustering method to obtain the second clustering label.

[0151] The process of obtaining the similarity matrix A from the self-expression matrix C is as follows:

[0152] A = 1 / 2(|C| + |C|) T )

[0153] The preset clustering method includes, but is not limited to, spectral clustering algorithm, k-means clustering algorithm, etc.

[0154] For example, the self-expression matrix C is the first self-expression matrix, and the face subspace clustering unit 430 is specifically used to: obtain the similarity matrix A based on the first self-expression matrix, and then process the similarity matrix A using a preset clustering method to obtain the second clustering label.

[0155] Step 3: Input the second cluster label into the face subspace partitioning unit 410 to obtain the first cluster label.

[0156] The face subspace partitioning unit 410 is used to: receive the second clustering label output by the face subspace clustering unit 430, and perform one-hot encoding on the second clustering label, the one-hot encoding of which yields the first clustering label.

[0157] For example, if the second cluster label of a certain feature channel is 5, the first cluster label obtained after one-hot encoding is [0,0,0,0,1].

[0158] It should be noted that step three is optional to be executed in the first training phase; if the entire face subspace clustering and partitioning module 400 participates in the first training phase, then step three is executed in the first training phase; if the similarity matrix learning unit 420 and the face subspace clustering unit 430 participate in the first training phase, then step three is not executed in the first training phase.

[0159] II. Second Training Phase:

[0160] If step three is not executed in the first training phase, the samples in the second training phase include multiple first face images, multiple sets of random vectors, and a second cluster label. Any one of the multiple random vector sets includes P third random vectors, where P is a positive integer. Since the second training phase includes step three, it needs to be executed first to convert the second cluster label into a first cluster label. If step three is executed in the first training phase, the samples in the second training phase include the aforementioned multiple first face images, the aforementioned multiple sets of random vectors, and the first cluster label. Optionally, the first face image can be a low-quality face image.

[0161] For example, the structure of the face generator 300 can be as follows: Figures 1 to 3 As shown, the face generator 300 consists of two parts: an M network and a G network. The input to the M network is a one-dimensional vector (e.g., a one-dimensional vector with dimension 512), and the output of the M network is also a one-dimensional vector (e.g., a one-dimensional vector with dimension 512). The third random vector is obtained by inputting a random vector that follows a Gaussian distribution into the M network of the face generator 300.

[0162] In addition to step three, the second training phase also includes step four, as shown below:

[0163] Step 4: For multiple first face images, multiple sets of random vectors, and the first cluster label, perform the following operations to obtain the face reconstruction network:

[0164] S21: Input the second face image into the feature encoder 100 to obtain M fifth face features and a second feature vector. The second face image is one of multiple first face images, and M is an integer greater than 1.

[0165] The feature encoder 100 is used to: receive the input face image I input For the input face image I input Feature extraction is performed to obtain several features. Among these features The dimensions are all different, and the larger j is, the greater the value of j. The smaller the size; and based on several characteristics The smallest feature in the middle size Obtain the control vector ω e .

[0166] The feature encoder 100 may include M first feature extraction modules. The input of the (j+1)th first feature extraction module among the M first feature extraction modules is the output of the j-th first feature extraction module among the M first feature extraction modules. The output of any one of the M first feature extraction modules is a feature. Optionally, the input of the first feature extraction module among the M first feature extraction modules is the input face image I. input Or for the input face image I iinput Features obtained through feature extraction.

[0167] Among them, the feature encoder 100 is based on features Obtain the control vector ω e The process can be: to extract features Input several second fully connected layers to obtain the control vector ω. e It should be understood that the feature encoder 100 may or may not include the aforementioned second fully connected layers; for example, when the feature encoder 100 does not include the aforementioned second fully connected layers, the aforementioned second fully connected layers may be several M-networks of the face generator 300, that is, the feature encoder 100 will generate features... Input several M-networks to obtain the control vector ω e .

[0168] It should be noted that the feature encoder 100 is a multi-layer network structure with several features. These are the output results of several layers in the multilayer structure of the feature encoder 100.

[0169] For example, input a face image I input For the second face image, the feature encoder 100 is specifically used to: extract features from the second face image to obtain M fifth face features; and input the sixth target face feature into several second fully connected layers to obtain a second feature vector; the M fifth face features have different sizes; the sixth target face feature is one of the M fifth face features, optionally, the sixth target face feature is the fifth face feature with the smallest size among the M fifth face features. At this time, several features... For M fifth facial features, the features For the sixth target facial features, control vector ω e This is the second eigenvector.

[0170] S22: Input the P third random vectors and second feature vectors from the first target random vector set into the style vector control module 200 to obtain P second style vector sets. The first target random vector set is one of multiple random vector sets. The P third random vectors in the first target random vector set correspond one-to-one with the P second style vector sets. Any second style vector set in the P second style vector sets includes Q second style vectors, where Q is a positive integer.

[0171] The style vector control module 200 is used to: receive the control vector ω output by the feature encoder 100. e and receiving random vectors In the channel dimension, the control vector ω e and random vectors Perform concatenation and combine the control vector ω e and random vectors The concatenated results are input into Q first fully connected layers to obtain Q style vectors. k∈{1,2,…,P}, i∈{1,2,…,Q}, Q style vectors Each of the Q first fully connected layers corresponds to one of the Q style vectors. any style vector in This is the output of the corresponding first fully connected layer. It should be understood that the style vector control module 200 includes Q first fully connected layers.

[0172] Among them, random vector The intermediate hidden variable ω is the output of the M-network of the face generator; that is, by inputting a random vector z following a Gaussian distribution into the M-network of the face generator, the output of the M-network of the face generator is a random vector.

[0173] For example, random vectors Let ω be any one of the P third random vectors in the first target random vector set, and let ω be the control vector. e The second feature vector is obtained by concatenating any third random vector with the second feature vector to form a second concatenated vector, which is the result of concatenating the any third random vector with the second feature vector. This second concatenated vector is then input into Q first fully connected layers to obtain Q second style vectors corresponding to the any third random vector. These Q second style vectors constitute the set of second style vectors corresponding to the any third random vector. At this point, the Q style vectors... Let Q be the second style vectors corresponding to any given third random vector. Similarly, since there are P third random vectors in the first target random vector set, and each of these P third random vectors is concatenated with a second feature vector, and the second concatenated vector obtained by concatenating each third random vector with the second feature vector is input into Q first fully connected layers, each third random vector corresponds to a set of second style vectors; thus, these P third random vectors correspond to P sets of second style vectors, and each of these P sets of second style vectors includes Q second style vectors.

[0174] S23: Input P sets of second style vectors into the face generator 300 to perform convolution modulation (Mod) operation on the face generator 300 to obtain P sets of second face features. The P sets of second face features correspond one-to-one with the P sets of second style vectors. Any one of the P sets of second face features includes Q sixth face features. The Q sixth face features correspond one-to-one with the Q second style vectors in the second target style vector set. The second target style vector set is the set of second style vectors in the P sets of second style vectors that corresponds to any one of the sets of second face features. Any one of the Q sixth face features is obtained by performing convolution modulation on the face generator 300 based on the second style vector corresponding to the any one of the sixth face features.

[0175] Among them, the face generator 300 is used to: receive style vectors Style vectors Convolutional modulation is performed with a constant (Const) as input, and the output features are...

[0176] It should be noted that during convolutional modulation, the constant (Const) is a fixed input to the face generator 300; the face generator 300 includes Q convolutional neural network modules (e.g., generative adversarial network modules), the input of the first convolutional neural network module among the Q convolutional neural network modules includes the constant (Const), and the input of the i-th convolutional neural network module among the Q convolutional neural network modules includes the output of the (i-1)-th convolutional neural network module among the Q convolutional neural network modules; furthermore, for any k, its corresponding style vector There are Q style vectors. These Q convolutional neural network modules correspond one-to-one, that is, these Q style vectors Each style vector in This serves as the input to the corresponding convolutional neural network module; thus, the style vector... The input of the i-th convolutional neural network module out of Q convolutional neural network modules is the feature.

[0177] The face generator 300 performs convolution modulation operations, which can modify the weights of the convolution kernels in each convolutional layer of the face generator 300. The convolution modulation process can be represented by the following formula:

[0178]

[0179] In the above formula, w represents the style vector. abc w′ represents the weights of the convolution kernel before convolution modulation. abc This represents the weights of the convolution kernel after convolution modulation. 'a' indicates the layer number of the convolution kernel, and 'b' and 'c' indicate the spatial location of the weights in the convolution kernel. For example, 'b' indicates the row of the convolution kernel, and 'c' indicates the column of the convolution kernel.

[0180] For example, style vectors For any second style vector in a set of P second style vectors, the feature Let the sixth face feature be the one corresponding to any second style vector; then, the arbitrary second style vector is input into the face generator 300, and the sixth face feature corresponding to the arbitrary second style vector is output; and, the face generator 300 is subjected to a convolution modulation to obtain the corrected face generator 300.

[0181] It should be understood that the face generator 300 includes Q convolutional neural network modules. The input of the first convolutional neural network module among the Q convolutional neural network modules includes a constant (Const), and the input of the i-th convolutional neural network module among the Q convolutional neural network modules includes the output of the (i-1)-th convolutional neural network module among the Q convolutional neural network modules. Furthermore, any one of the P sets of second style vectors includes Q second style vectors, and these Q second style vectors correspond one-to-one with the Q convolutional neural network modules; that is, each of the Q second style vectors is the input of its corresponding convolutional neural network module. In addition, the set of second face features corresponding to any one of the P sets of second face features includes Q sixth face features. Thus, the i-th second style vector among the Q second style vectors is the input of the i-th convolutional neural network module among the Q convolutional neural network modules, and the output of the i-th convolutional neural network module among the Q convolutional neural network modules is the i-th sixth face feature among the Q sixth face features. Figure 5The target convolutional neural network module described in the illustrated embodiment is any one of Q convolutional neural network modules.

[0182] S24: Input each of the P seventh target face features into the multi-face feature mapping module 500 to obtain P sets of third face sub-features. The P seventh target face features are the P sixth face features in the P sets of second face features. The P seventh target face features are the sixth face features in different sets of second face features. The P seventh target face features are obtained by convolutional modulation of the face generator 300 (specifically the target convolutional neural network module) based on the second style vector output by the same first fully connected layer. The P seventh target face features correspond one-to-one with the P sets of third face sub-features. Any third face sub-feature set in the P sets of third face features includes R categories of fifth face sub-features. The R categories of fifth face sub-features included in any third face feature set are obtained by dividing the seventh target face features corresponding to the any third face feature set. R is an integer greater than 1.

[0183] The multi-face feature mapping module 500 is used to: receive features output by the face generator 300 And based on the first clustering label output by the face subspace partitioning unit 410, the features are... Sub-features divided into R categories in either the channel dimension or the spatial dimension.

[0184] It should be noted that the first cluster label includes R categories, therefore the features Sub-features divided into R categories In addition, features Corresponding to a feature space, sub-features This corresponds to a subspace of the feature space.

[0185] For example, features For any seventh target face feature, sub-feature The first cluster label represents the fifth face feature. Therefore, any seventh target face feature can be divided into R categories of fifth face features based on the first cluster label. These R categories of fifth face features constitute the set of third face features corresponding to the seventh target face feature. Similarly, since there are P seventh target face features, each of these P features is divided in the aforementioned way to obtain a set of third face features corresponding to each seventh target face feature. Therefore, after feature division of the P seventh target face features, P sets of third face features are obtained. It should be understood that the P seventh target face features are obtained by convolutionally modulating the face generator 300 using the second style vector output from the same first fully connected layer; that is, the P seventh target face features are represented as features. When i takes the same value, k∈{1,2,…,P}.

[0186] It should be further explained that, regarding features When performing feature partitioning, since i∈{1,2,…,Q}, each time from all features... In this process, features with the same value for i are selected for feature partitioning. For each value of i, since k∈{1,2,…,P}, there are a total of P features. Each of these P features is divided into R categories of sub-features, resulting in P groups of R categories of sub-features. These P groups of R categories of sub-features are also P sets of sub-features, each consisting of R categories of sub-features. Since there are Q values ​​for i, when selecting a value for i from the Q values, the specific value of i can be determined according to actual needs. Furthermore, when selecting a value for i from the Q values, the number of values ​​of i selected can also be determined according to actual needs, that is, selecting features corresponding to one or more values ​​of i for feature partitioning. This application is only an example of describing the selection of one value of i. It should be understood that the more values ​​of i are selected, the higher the accuracy of the face restoration network, that is, the more features are used for feature segmentation, the better the face restoration ability of the trained face restoration network, but the computational cost will increase. Therefore, an appropriate number of values ​​of i can be selected so that the face image quality can be improved without increasing the computational cost too much.

[0187] For example, the P seventh target face features are one of the P sets of second face features and one of the P sets of second face features. That is, step S24 is only an example of feature segmentation for some of the sixth face features, and does not segment all of the sixth face features. This application does not limit the number of sixth face features to be segmented, and can dynamically determine the number of sixth face features to be segmented according to actual needs.

[0188] S25: Input the eighth target face feature and P sets of third face sub-features into the multi-face feature combination module 600 to obtain the second combined face feature. The eighth target face feature is one of the M fifth face features, and the eighth target face feature is not the sixth target face feature. The sixth target face feature is the smallest in size among the M fifth face features.

[0189] The multi-face feature combination module 600 is used to: receive features output by the feature encoder 100 Sub-features output by the multi-face feature mapping module 500 And based on characteristics Features of the sub-score Obtain the combined weights And based on sub-features and combined weights Perform weighted summation to obtain the combined features Specifically:

[0190] (1) Obtain the combined weights Features Features of the sub-score Connect the features in the channel dimension or spatial dimension. Features of the sub-score The concatenated results are subjected to several convolution and pooling operations to obtain the corresponding combined weights.

[0191] (2) Obtain the combined features

[0192] a. Sub-features along the superscript k dimension and combined weights Perform multiplication and sum the results;

[0193] b. Multiply the summation result from step a by the first cluster label, and combine the multiplication result along the index r dimension to obtain the combined feature.

[0194] For example, features For the eighth target facial features, sub-features For any fifth face feature in a set of P third face features, the combined weights are... The second combination weight, combination features This is the second set of facial features; thus, the process of obtaining the second set of facial features by combining the eighth target facial features and P sets of third facial sub-features is as follows:

[0195] (1) Obtain the second combined weight corresponding to each fifth face feature: First, concatenate the eighth target face feature with each fifth face feature in the P sets of third face features in the channel dimension or spatial dimension to obtain the concatenation result of the eighth target face feature and each fifth face feature, which is called the second concatenation feature; then perform several convolution and pooling operations on the second concatenation feature to obtain the second combined weight corresponding to each fifth face feature.

[0196] That is, any second combined weight is obtained by performing convolution and pooling operations on the second concatenated feature, with the output of the convolution operation serving as the input of the pooling operation. The second concatenated feature is obtained by concatenating the eighth target face feature and the fifth face sub-feature corresponding to any first combined weight.

[0197] Therefore, based on the eighth target face feature and P sets of third face sub-features, P sets of second combination weights can be obtained. The P sets of second combination weights correspond to the P sets of third face sub-features. Any one of the P sets of second combination weights includes R sets of second combination weights. The R sets of second combination weights correspond to the fifth face sub-features of the R categories in the third target face sub-feature set. The third target face sub-feature set is the third face feature set in the P sets of third face features that corresponds to any one of the sets of second combination weights. Any one of the R sets of second combination weights is obtained based on the fifth face sub-features of the categories corresponding to any one of the second combination weights in the eighth target face feature and the third target face sub-feature set.

[0198] (2) Obtain the second set of facial features:

[0199] a. First, multiply each fifth face feature in the P sets of third face features by its corresponding second combination weight to obtain the sixth face feature corresponding to each fifth face feature; since there are R categories of fifth face features, there are R categories of sixth face features; for these R categories of sixth face features, add all the sixth face features of the same category to obtain a seventh face feature of that category; since there are R categories, there are R seventh face features.

[0200] b. Multiply the first cluster label by R seventh face features to obtain R eighth face features; then, the R eighth face features form a second combined face feature in the channel dimension or spatial dimension.

[0201] Therefore, based on the P sets of third-person face features and the P sets of second-person combination weights, we can obtain P sets of fourth-person face features. The P sets of third-person face features correspond to the P sets of fourth-person face features. Any one of the P sets of fourth-person face features includes R categories of sixth-person face features. These R categories of sixth-person face features correspond to the R categories of fifth-person face features in the fourth target face feature set. The fourth target face feature set is the set of third-person face features in the P sets of third-person face features that corresponds to any one of the P sets of fourth-person face features. Any one of the R categories of sixth-person face features... The sixth face feature of a category is obtained by multiplying the second target face feature and the weight of the second target combination. The second target face feature is the fifth face feature of the category corresponding to the sixth face feature of any category. The weight of the second target combination is the weight of the second combination corresponding to the second target face feature. Then, the sixth face features of the same category in the P sets of fourth face features are added together to obtain R seventh face features. The first cluster label is then multiplied by the R seventh face features to obtain R eighth face features. Finally, the R eighth face features are combined to obtain the second combined face features.

[0202] S26: Input the second combination of facial features into the face generator 300 to obtain the third synthetic face image.

[0203] In step S25, the combined features are... Input face generator 300 to obtain the restored face image I. rec .

[0204] For example, combined features The second set of facial features, the restored facial image I rec The third synthesized face image is obtained by inputting the second combined face features into the face generator 300.

[0205] S27: Calculate the second loss based on the ground truth image corresponding to the second face image and the third synthesized face image.

[0206] During training, each input face image I input There is a corresponding ground truth image, which is the input face image I. input It is an input face image, the input face image I input The corresponding ground truth image is the input face image I. input A high-quality version, that is, the input face image I input Corresponding ground truth image and the input face image I input The content of the image is the same as the input face image I.input Corresponding ground truth image and the input face image I input The image quality differs; while the restored face image I rec For the input face image I input The image recovered by the face reconstruction network can therefore be used as the input face image I. input The restored face image is evaluated using the ground truth image. rec The quality of the image. Specifically, based on the input face image I input The corresponding ground truth image and the restored face image I rec To calculate the loss in the second training phase (i.e., the second loss).

[0207] For example, any one of the multiple first face images corresponds to a ground truth image, so the second face image also corresponds to a ground truth image. Thus, the second loss for this training can be calculated based on the ground truth image corresponding to the second face image and the third synthesized face image.

[0208] S28: If the second loss is less than the second preset threshold, then the training ends and the face restoration network at this time is the final face restoration network, which can be used for inference; otherwise, adjust the parameters in the face restoration network according to the second loss to obtain the updated face restoration network, and execute step S29.

[0209] In the second training phase, the modules whose parameters need to be updated according to the second loss include the feature encoder 100, the style vector control module 200, the face generator 300, and the multi-face feature combination module 600. It should be noted that as the parameters of the aforementioned modules are updated, the parameters of the first and second fully connected layers are also updated; however, the parameters of the face generator 300 are optional to be updated, while the face subspace partitioning unit 410 and the multi-face feature mapping module 500 do not have their parameters updated.

[0210] S29: Using the third face image as the second face image and the second target random vector set as the first target random vector set, continue to execute steps S21 to S28 to train the aforementioned updated face recovery network; the third face image is the first face image among multiple first face images that has not yet been used for training, and the second target random vector set is the random vector set among multiple random vector sets that has not yet been used for training.

[0211] Please see Figure 9 , Figure 9 yes Figure 8 The diagram illustrates the inference phase of a face reconstruction network, as shown below:

[0212] The face subspace partitioning unit 410 is used to output the first clustering label to the multi-face feature mapping module 500.

[0213] The feature encoder 100 is used to: receive an input face image I input For the input face image I input Feature extraction is performed to obtain several features. Among these features The dimensions are all different, and the larger j is, the greater the value of j. The smaller the size; and based on several characteristics The smallest feature in the middle size Obtain the control vector ω e .

[0214] For example, the feature encoder 100 receives the input face image I input For low-quality face images, the feature encoder 100 extracts features from the low-quality face image, obtaining M features. There are M second face features, each with a different size. These M second face features include both a first target face feature and a second target face feature. Optionally, the feature with the smallest size among the M second face features... The first target is the facial features. The feature encoder 100 obtains the control vector ω based on the first target facial features. e This is the first eigenvector.

[0215] Style vector control module 200 is used to: receive the control vector ω output by feature encoder 100 e and receiving random vectors In the channel dimension, the control vector ω e and random vectors Perform concatenation and combine the control vector ω e and random vectors The concatenated results are input into Q first fully connected layers to obtain Q style vectors. Q style vectors Each of the Q first fully connected layers corresponds to one of the Q style vectors. any style vector in This is the output of the corresponding first fully connected layer.

[0216] Among them, random vector The intermediate hidden variable ω is the output of the M-network of the face generator; that is, by inputting a random vector z following a Gaussian distribution into the M-network of the face generator, the output of the M-network of the face generator is a random vector.

[0217] For example, the control vector ω received by the style vector control module 200 e The first feature vector and the random vector received by the style vector control module 200. Given P first random vectors; the style vector control module 200 concatenates the first feature vector with any one of the P first random vectors to obtain the first concatenated vector; then, the first concatenated vector is input into Q first fully connected layers to obtain Q style vectors. There are Q first style vectors, which correspond to Q first fully connected layers, and any one of these Q first style vectors is the output of its corresponding first fully connected layer.

[0218] It should be understood that, since there are P first random vectors, the style vector control module 200 concatenates the first feature vector with each of the P first random vectors to obtain P first concatenated vectors; then, each of the P first concatenated vectors is input into Q first fully connected layers to obtain Q first style vectors corresponding to each first concatenated vector; therefore, after processing the first feature vector and the P first random vectors, the style vector control module 200 can obtain a set of P first style vectors, and any set of P first style vectors includes Q first style vectors.

[0219] Among them, the P sets of first style vectors include P target style vectors, that is, the P target style vectors are the P first style vectors in the P sets of first style vectors, and the P target style vectors are the first style vectors in different sets of first style vectors in the P sets of first style vectors. The P target style vectors are obtained by inputting the P first concatenated vectors into the same first fully connected layer in the Q first fully connected layers.

[0220] Face generator 300 is used to: receive style vectors Style vectors Convolutional modulation is performed with a constant (Const) as input, and the output features are...

[0221] For example, the style vector received by the face generator 300 Given any first-style vector from a set of P first-style vectors, the face generator 300 takes any first-style vector from the set of P first-style vectors and a constant (Const) as input, and outputs features. Let be the third face feature corresponding to any one of the first style vectors.

[0222] It should be understood that each set of first style vectors contains Q first style vectors. Therefore, inputting the Q first style vectors into the face generator 300 for convolutional modulation yields Q third face features. These Q third face features constitute the first face feature set corresponding to that set of first style vectors. Furthermore, since there are P sets of first style vectors, using these P sets of first style vectors to perform convolutional modulation on the face generator 300 yields P sets of first face features. These P sets of first face features correspond one-to-one with the P sets of first style vectors, and any one of the P sets of first face features includes Q third face features.

[0223] The multi-face feature mapping module 500 is used to: receive features output by the face generator 300 And based on the first clustering label output by the face subspace partitioning unit 410, the features are... Sub-features divided into R categories in either the channel dimension or the spatial dimension.

[0224] For example, the features received by the multi-face feature mapping module 500 This includes P third target face features, where each P third target face feature is a third face feature from a set of P first face features, and the P third target face features are third face features from different sets of first face features. These P third target face features are obtained by convolutional modulation of the face generator 300 (specifically, the target convolutional neural network module) based on the first style vector output from the same first fully connected layer. The multi-face feature mapping module 500 divides any one of the P third target face features into sub-features based on the first clustering label, either in the channel dimension or the spatial dimension. The first face feature is defined as follows: any third target face feature is divided into R categories of first face features. Each of the R categories includes one first face feature. Therefore, the R categories of first face features are also the same as the R first face features. The R categories of first face features obtained from the division of any third target face feature constitute the set of first face features corresponding to any third target face feature.

[0225] It should be understood that since there are P third target face features, the multi-face feature mapping module 500 processes the P third target face features to obtain P sets of first face sub-features, and any one of the P sets of first face sub-features includes R categories of first face sub-features.

[0226] The multi-face feature combination module 600 is used to: receive features output by the feature encoder 100 Sub-features output by the multi-face feature mapping module 500 And based on characteristics Features of the sub-score Obtain the combined weights And based on sub-features and combined weights Perform weighted summation to obtain the combined features Specifically:

[0227] (1) Obtain the combined weights Features Features of the sub-score Connect the features in the channel dimension or spatial dimension. Features of the sub-score The concatenated results are subjected to several convolution and pooling operations to obtain the corresponding combined weights.

[0228] (2) Obtain the combined features

[0229] a. Sub-features along the superscript k dimension and combined weights Perform multiplication and sum the results;

[0230] b. Multiply the summation result from step a by the first cluster label, and combine the multiplication result along the index r dimension to obtain the combined feature.

[0231] For example, the features received by the multi-face feature combination module 600 The second target's facial features, and the received sub-features. Let P be any set of P first-face feature sets, and let P be any category of first-face feature.

[0232] (1) Obtaining the first combined weights: First, concatenate the second target face feature with the first face sub-feature of any category in the channel dimension or spatial dimension. Then, perform several convolution and pooling operations on the concatenation results of the second target face feature and the first face sub-feature of any category to obtain the combined weights. The first combined weight is the first facial feature corresponding to any category.

[0233] It should be understood that since a first face feature set includes R categories of first face features, the R categories of first face features can yield R first combination weights, and these R first combination weights constitute the first combination weight set corresponding to the first face feature set; furthermore, if there are P first face feature sets, then there are P first combination weight sets, and any one of the P first combination weight sets includes R first combination weights.

[0234] (2) Obtain the first set of facial features:

[0235] a. First, multiply each first face feature in the P sets of first face features by its corresponding first combination weight to obtain the second face feature corresponding to each first face feature; since there are R categories of first face features, there are R categories of second face features; for these R categories of second face features, add all the second face features of the same category to obtain a third face feature of that category; since there are R categories, there are R third face features.

[0236] b. Multiply the first cluster label by R third face features to obtain R fourth face features; then, the R fourth face features are combined into a first combined face feature in the channel dimension or spatial dimension.

[0237] The face generator 300 is also used to: receive combined features Combined features The input is the recovered face image I. rec .

[0238] For example, the face generator 300 receives combined features The face generator 300 generates the restored face image I obtained from the first combination of facial features. rec This is the first synthesized human face image.

[0239] Please see Figure 10 , Figure 10 yes Figure 8 The diagram shown is an exemplary structural diagram of a face recovery network. The following section will discuss... Figure 10 The first and second training phases of the face reconstruction network shown are introduced.

[0240] I. First Training Phase:

[0241] It should be noted beforehand that the face generator 300 can use the StyleGAN2 network, which includes 23 convolutional neural network modules. Figure 10Only a portion of the 23 convolutional neural network (CNN) modules are shown, namely CNN module G_4 (output feature size 4×4), CNN module G_8 (output feature size 8×8), CNN module G_16 (output feature size 16×16), CNN module G_32 (output feature size 32×32), CNN module G_64 (output feature size 64×64), CNN module G_128 (output feature size 128×128), CNN module G_256 (output feature size 256×256), and CNN module G_512 (output feature size 512×512). Therefore, Figure 10 The connections between convolutional neural network modules G_4, G_8, G_16, G_32, G_64, G_128, G_256, and G_512 shown do not necessarily represent direct connections; they could also represent interface connections, for example... Figure 10 There may be one or more convolutional neural network modules between the two convolutional neural network modules shown in the diagram.

[0242] As an example, Figure 6 The target convolutional neural network module described in the illustrated embodiment can be Figure 10 The convolutional neural network module G_16 or convolutional neural network module G_128 shown.

[0243] As an example, in Figure 10 In the diagram, convolutional neural network module G_16 is the 5th convolutional neural network module out of 23 convolutional neural network modules, and convolutional neural network module G_128 is the 11th convolutional neural network module out of 23 convolutional neural network modules.

[0244] Step 1: Optimize the self-expression matrix C1 and the self-expression matrix C2 to obtain the final self-expression matrix C1 and the final self-expression matrix C2.

[0245] In the first training phase, a random vector z (e.g., a second random vector) following a Gaussian distribution is input into the face generator 300. Features output from the convolutional neural network module G_16 of the face generator 300 are extracted, and these features are used to train the self-expression matrix C1 (512×512 dimensions) to obtain the final self-expression matrix C1. Similarly, features output from the convolutional neural network module G_128 of the face generator 300 are extracted, and these features are used to train the self-expression matrix C2 (16384×16384 dimensions) to obtain the final self-expression matrix C2. The specific process for training the final self-expression matrices C1 and C2 is as follows:

[0246] (1) A random vector z following a Gaussian distribution is input into the convolutional neural network module G_4 of the face generator 300, and then processed sequentially by convolutional neural network module G_4, convolutional neural network module G_8, and convolutional neural network module G_16. At this time, the similarity matrix learning unit 420 is a channel self-expression layer, specifically used to: extract the features output by convolutional neural network module G_16 (denoted as features). The dimensions are 16×16×512), and the features are... Performing a matrix multiplication operation with the self-expression matrix C1 yields the features. The dimension of a feature refers to width × height × number of channels.

[0247] (2) Features The face generator 300's convolutional neural network module G_32 is re-inputted and processed sequentially by convolutional neural network module G_32, convolutional neural network module G_64, and convolutional neural network module G_128. At this point, the similarity matrix learning unit 420 is a spatial self-expression layer, specifically used to: receive the features output by convolutional neural network module G_128 (denoted as features). (Dimensions are 128×128×64), and features Perform a matrix multiplication operation with the self-expression matrix C2 to obtain the features.

[0248] (3) Features The subsequent structures of the input face generator 300, such as the input convolutional neural network module G_256, are processed by the convolutional neural network module G_256 and the convolutional neural network module G_512 to obtain the synthesized face image.

[0249] Based on the loss function of the first training stage, the self-expression matrix C1 is first optimized, then C1 is fixed and the self-expression matrix C2 is optimized, thus obtaining the final self-expression matrices C1 and C2. The loss function of the first training stage is shown below:

[0250] loss1 = ||G1(z) - G1(z)C i ‖2+λ1||G2(G1(z))-G2(G1(z)C i )||1+λ2‖C i ||1, i = 1, 2

[0251] In the above formula, loss1 is the first loss; the face generator is divided into two parts, G1 and G2; z is a random vector following a Gaussian distribution, such as the second random vector; G1(z) is the intermediate feature obtained by taking the random vector z following a Gaussian distribution as input; G2(G1(z)) is the face image obtained by taking the random vector z following a Gaussian distribution as input; G2(G1(z)C i ) is related to the self-expression matrix C i The face image obtained after matrix multiplication, where λ1 and λ2 are the weights of the loss components.

[0252] Step 2: Input the final self-expression matrix C1 into the face subspace clustering unit 430 to obtain the second cluster label_1; and input the final self-expression matrix C2 into the face subspace clustering unit 430 to obtain the second cluster label_2.

[0253] The face subspace clustering unit 430 is used to: obtain a similarity matrix A1 by taking the final self-expression matrix C1 as input; and process the similarity matrix A1 using a preset clustering method (e.g., spectral clustering) to obtain a second cluster label_1. Furthermore, the face subspace clustering unit 430 is also used to: obtain a similarity matrix A2 by taking the final self-expression matrix C2 as input; and process the similarity matrix A2 using a preset clustering method (e.g., spectral clustering) to obtain a second cluster label_2.

[0254] Step 3: Obtain the first cluster label_1 based on the second cluster label_1, and obtain the first cluster label_1 based on the second cluster label_1.

[0255] The face subspace partitioning unit 410 is used to: perform one-hot encoding on the second cluster label_1 to obtain the first cluster label_1, such as... Figure 10 As shown, the first cluster label_1 includes m 1 m 2 m 3 ... m R ; and perform one-hot encoding on the second cluster label_2 to obtain the first cluster label_2, the first cluster label_2 not being in Figure 10 As shown in the image.

[0256] It should be noted that the first cluster label_1 is used to divide the feature into R sub-features in the channel dimension, and the first cluster label_2 is used to divide the feature into R sub-features in the spatial dimension.

[0257] II. Second Training Phase:

[0258] Step 4: Transfer the quality of the face image I input Input feature encoder 100 to obtain features (Dimensions are 128×128×64), Features (Dimensions are 16×16×512), Features (dimensions are 4×4×512) and control vector ω e (Dimensions are 512×1).

[0259] like Figure 10 As shown, the feature encoder 100 includes seven first feature extraction modules and one second feature extraction module. Figure 10 (Not all are shown in the image). Each of the seven first feature extraction modules includes a cascaded convolutional layer (Conv), an activation layer (ReLU), and a downsampling layer. The downsampling factor of each of the seven first feature extraction modules is different. The second feature extraction module includes a cascaded convolutional layer (Conv) and an activation layer (ReLU). The input to this second feature extraction module is the input face image I. input The output of the second feature extraction module is the input of the first feature extraction module out of the seven first feature extraction modules. The input of the j-th first feature extraction module out of the seven first feature extraction modules is the output of the (j-1)-th first feature extraction module out of the seven first feature extraction modules. The output of the seventh first feature extraction module out of the seven first feature extraction modules is the input of two second fully connected layers. The output of the two second fully connected layers is the control vector ω. e .

[0260] The feature encoder 100 is used to: receive the input face image I input (Dimensions are 512×512×3), feature extraction is performed using one second feature extraction module and seven first feature extraction modules, resulting in the feature output by the third of the seven first feature extraction modules. (Dimensions are 128×128×64), features output by the 6th first feature extraction module (dimensions are 16×16×512), and the features output by the 7th first feature extraction module. (dimensions are 4×4×512), and features are... The control vector ω is obtained by passing through two second fully connected layers.e (Dimensions are 512×1).

[0261] Step 5: Set the control vector ω e (dimension 512×1) and random vectors (512×1 dimension) Input style vector control module 200 to obtain style vector. (Dimension is 512×1), k∈{1,2,…,10}, i∈{1,2,…,23}.

[0262] The style vector control module 200 is used to: receive the control vector ω e and random vectors (dimension 512×1), k∈{1,2,…,P}; the control vector ω e Each random vector Concatenate along the channel dimension and combine the control vector ω e and each random vector The splicing results are respectively input into Q first fully connected layers ( Figure 10 (not shown in the image), to obtain each random vector The corresponding Q style vectors (Dimension is 512×1), k∈{1,2,…,P}, i∈{1,2,…,Q}.

[0263] For example, P = 10, Q = 23, therefore there are 10 random vectors. 23 first fully connected layers; control vector ω e and each random vector The splicing results are respectively input into 23 first fully connected layers ( Figure 10 (not shown in the image), to obtain each random vector The corresponding 23 style vectors (Dimension is 512×1), k∈{1,2,…,10}, i∈{1,2,…,23}.

[0264] Step Six: Transfer Style Vectors Input the face generator 300 to perform convolution modulation operation on the convolutional neural network module G_16 of the face generator 300, and obtain the features output by the convolutional neural network module G_16.

[0265] Among them, the face generator 300 is used to: convert style vectors As input to the 300 convolutional modulation operation of the face generator, the output features

[0266] For example, Q = 23, therefore the face generator includes 23 convolutional neural network modules. Figure 10 Only a portion of the 23 convolutional neural network (CNN) modules are shown; CNN module G_16 is the 5th of the 23 CNN modules. Therefore, the input to CNN module G_16 for convolutional modulation includes the output of the 4th of the 23 CNN modules and the style vector. The output of the convolutional neural network module G_16 is the feature. It should be noted that the process of performing convolution modulation operation on the other convolutional neural network modules among the 23 convolutional neural network modules is the same as the process of performing convolution modulation operation on convolutional neural network module G_16, and will not be described again here.

[0267] Step 7: Feature Input the multi-face feature mapping module 500 to obtain sub-features.

[0268] The multi-face feature mapping module 500 is used to: receive features output by the face generator 300 {1,2,…,P}, i∈{1,2,…,Q}; and based on the first clustering label output by the face subspace partitioning unit 410, the features are... Sub-features divided into R categories in either the channel dimension or the spatial dimension. Specifically, for any feature output by a convolutional neural network module, the value of i is fixed, and a feature... (The value of i is fixed) corresponds to a face mapping. Since k∈{1,2,…,P}, there are P features. These P features Corresponding to P individual face mappings; P features are assigned based on the first cluster label. Each of the features is divided into R categories.

[0269] like Figure 10 As shown, the features output by the convolutional neural network module G_16 are based on the first cluster label_1. Sub-features divided into R categories along the channel dimension For example, P = 10, R = 5, therefore there are 10 features. The first cluster label_1 includes m 1 m 2 m 3 ... m 5 Based on the first cluster label_1, 10 features are selected. Each sub-feature is divided into 5 categories.

[0270] Step 8: Feature Features of the sub-score Input the multi-face feature combination module 600 to obtain the combined features It should be understood that the characteristics AND characteristics They are the same size.

[0271] like Figure 10 As shown, for any sub-feature Since k∈{1,2,…,P}, there are P subspaces; and since r∈{1,2,…,R}, for any one of the P values ​​of k, based on the sub-features... We can obtain R combined weights. For example, for sub-features... Therefore, there are 10 subspaces, and each subspace has 5 combined weights.

[0272] For example, the multi-face feature combination module 600 is used to: combine features Features of the sub-score Concatenate along the channel dimension and combine the features. Features of the sub-score The concatenated result is input into two cascaded first preset network modules to obtain the combined weights. The first preset network module includes cascaded convolutional layers (Conv), activation layers (ReLU), and downsampling layers (downsampling factor of 4); features are first processed along the superscript k dimension. and combined weights Multiply and combine features and combined weights The summation of the multiplication results is then performed; the summation result along the k-th dimension is multiplied by the first cluster label _1, and the summation result along the k-th dimension is combined with the first cluster label _1 along the r-th dimension to obtain the combined feature.

[0273] Step 9: Combine features As input to the convolution modulation operation of the convolutional neural network module G_64, according to Figure 10 The face recovery network shown sequentially performs convolutional modulation operations on convolutional neural network modules G_32, G_64, and G_128 to obtain the features output by convolutional neural network module G_128.

[0274] Step 10: Feature Input the multi-face feature mapping module 500 to obtain sub-features.

[0275] For example, features are grouped based on the first cluster label_2. Sub-features are divided into 5 categories in the spatial dimension.

[0276] Step 11: Add features Features of the sub-score Input the multi-face feature combination module 600 to obtain the combined features It should be understood that the characteristics AND characteristics They are the same size.

[0277] For example, the multi-face feature combination module 600 is used to: combine features Features of the sub-score The features are stitched together in the spatial dimension. Features of the sub-score The concatenated result is input into four cascaded second preset network modules, and then the output of the last of the four second preset network modules is input into one third preset network module to obtain the combined weights. k∈{1,2,…,10}, r∈{1,2,…,5}, where the second preset network module includes cascaded convolutional layers (Conv), activation layers (ReLU), and downsampling layers (downsampling factor of 4), and the third preset network module includes cascaded convolutional layers (Conv), activation layers (ReLU), and downsampling layers (downsampling factor of 2); first, the features are mapped along the superscript k dimension. and combined weights Multiply and combine features and combined weights The summation results of the multiplication along the k-th dimension are then multiplied by the first cluster label _2, and the summation results along the k-th dimension are combined with the first cluster label _2 along the r-th dimension to obtain the combined features.

[0278] Step 12: Combine features As input to the convolution modulation operation of the convolutional neural network module G_256, according to Figure 10 The face restoration network shown sequentially performs convolutional modulation operations on convolutional neural network module G_256 and convolutional neural network module G_512, outputting the restored face image I. rec .

[0279] Among them, the restored face image I based on the output of step twelfth. recCalculate the second loss. If the second loss is not less than the second preset threshold, adjust the parameters of the face recovery network based on the second loss and change the training samples. Repeat the above steps four to twelve until the second loss is less than the second preset threshold, and the second training phase ends.

[0280] The formula for calculating the second loss is as follows:

[0281] loss2 = ||I rec -GT‖1+λ3‖VGG(I rec )-VGG(GT)‖1+λ4‖log(1-D(I rec ))‖1

[0282] In the above formula, loos2 is the second loss, I rec The image represents the restored face, GT represents the ground truth image, VGG represents the VGG model, D represents the discriminant network or discriminator, and λ3 and λ4 represent the weights of the loss components.

[0283] Figure 11 yes Figure 10 The diagram shows the inference stage of a face restoration network. This face restoration network can receive low-quality or complexly degraded face images and generate high-quality face images with rich details, correct colors, and no artificial artifacts; for example, it can receive a low-quality face image and output a high-quality first synthetic face image.

[0284] To facilitate understanding of the beneficial effects of the embodiments of this application, the performance of the embodiments of this application is compared with the following seven benchmark algorithms:

[0285] Benchmark Algorithm 1: ESRGAN method, detailed reference "ESRGAN: Enhanced Super-Resolution Generative Adversarial Networks, ECCVW 2018".

[0286] Benchmark Algorithm 2: DFDNET method, detailed reference "Blind Face Restoration via DeepMulti-scale Component Dictionaries, ECCV 2020".

[0287] Benchmark Algorithm 3: GLEAN method, detailed reference "GLEAN: Generative Latent Bank for Large-Factor Image Super-Resolution, CVPR 2021".

[0288] Benchmark Algorithm 4: GFPGAN method, detailed reference "Towards Real-World Blind Face Restoration with Generative Facial Prior, CVPR 2021".

[0289] Benchmark Algorithm 5: GPEN method, see the detailed reference "GAN Prior Embedded Network for BlindFace Restoration in the Wild".

[0290] Benchmark Algorithm 6: PULSE method, detailed reference "PULSE: Self-Supervised PhotoUpsampling via Latent Space Exploration of Generative Models, CVPR 2020".

[0291] Benchmark Algorithm 7: mGANprior method, detailed reference "Image Processing Using Multi-Code GAN Prior, CVPR 2020".

[0292] The performance comparison results on the given training and test sets are shown in Table 1.

[0293] Table 1. Algorithm performance comparison results

[0294] algorithm PSNR SSIM LPIPS NIQE FID ESRGAN 28.1088 0.7808 0.3256 15.2320 68.4088 DFDNET 26.8188 0.7769 0.2561 9.7146 44.6026 GLEAN 24.5390 0.6389 0.3378 12.9772 67.3824 GFPGAN 26.9351 0.7807 0.2431 11.0229 37.7252 GPEN 26.5649 0.7698 0.2706 11.6622 50.1208 PULSE 21.4504 0.5413 0.5324 13.0708 147.6991 mGANprior 21.3004 0.5435 0.5381 13.4579 153.3856 This application 27.5722 0.7872 0.2317 9.4669 36.2616

[0295] In Table 1, PSNR represents Peak Signal-to-Noise Ratio, SSIM represents Structural Similarity, LPIPS represents Learned Perceptual Block Similarity, NIQE represents Natural Image Index, and FID represents Fréchet Initial Distance. Experiments show that the method provided in this application significantly outperforms the seven benchmark methods compared on this test dataset in terms of SSIM, LPIPS, NIQE, and FID. It is worth noting that the ESRGAN method achieves a better PSNR than this application because the former results in excessively blurred face reconstruction; although the PSNR is high, the visual quality deteriorates.

[0296] It should be noted that the embodiments of this application have a wide range of applications and can also be applied to other image restoration or enhancement tasks, such as building, home decoration, and portrait images. The modules in the embodiments of this application can also be adapted to other tasks. For example, the face synthesis subspace clustering and partitioning module 400 can be applied to face style transfer, face editing, and other tasks; similarly, the style vector control module 200 can be applied to face image restoration tasks. Furthermore, the embodiments of this application have high robustness to real-world open scenes and can adapt to degraded images obtained from different mobile phone models, different shooting scenarios, different ISP paths, and different transmission methods.

[0297] Please see Figure 12 , Figure 12 This is a schematic diagram of the structure of a face image processing device 1200 provided in an embodiment of this application. The face image processing device 1200 is applied to an electronic device and may include a processing unit 1201 and a communication unit 1202. The processing unit 1201 is used to perform actions such as... Figure 6 In any step of the method embodiment shown, and when performing data transmission such as acquisition, the communication unit 1202 may be selectively invoked to complete the corresponding operation. A detailed description follows.

[0298] The processing unit 1201 is configured to: acquire a low-quality face image and a first clustering label; extract features from the low-quality face image to obtain a first target face feature and a second target face feature; divide each of the P third target face features into R categories of first face sub-features according to the first clustering label to obtain a set of P first face sub-features, wherein any one of the P first face sub-feature sets includes R categories of first face sub-features, where P is a positive integer and R is an integer greater than 1; the P third target face features are the output of the target convolutional neural network module of the face generator, and the input of the target convolutional neural network module corresponding to the P third target face features is obtained based on the first target face feature; combine the first face sub-features in the set of P first face sub-features into a first combined face feature according to the second target face feature and the first clustering label; and obtain a first synthetic face image based on the first combined face feature.

[0299] In one possible implementation, the P third target face features are obtained by performing convolutional modulation on the target convolutional neural network module based on the first target face features and P first random vectors.

[0300] In one possible implementation, the P third target face features are obtained by performing convolutional modulation on the target convolutional neural network module based on the P target style vectors, and the P target style vectors are obtained based on the first target face features and the P first random vectors.

[0301] In one possible implementation, the P target style vectors are obtained from P first concatenation vectors, which are obtained by concatenating a first feature vector with each of the P first random vectors. The first feature vectors are obtained from the first target face features.

[0302] In one possible implementation, the processing unit 1201 is specifically configured to: obtain P sets of first combined weights based on the second target face feature and the P sets of first face sub-features, wherein the P sets of first combined weights correspond to the P sets of first face sub-features, any one of the P sets of first combined weights includes R first combined weights, the R sets of first combined weights correspond to R categories of first face sub-features in the first target face sub-feature set, the first target face sub-feature set is the set of first face features in the P sets of first face features that corresponds to any one of the first combined weights, and any one of the R sets of first combined weights is obtained based on the second target face feature and the first face sub-features in the first target face sub-feature set that correspond to the category of any one of the first combined weights; and combine the first face features in the P sets of first face features into the first combined face feature based on the first clustering label and the P sets of first combined weights.

[0303] In one possible implementation, the processing unit 1201 is specifically configured to: obtain P sets of second face features based on the P sets of first face sub-features and the P sets of first combined weights, wherein the P sets of first face features correspond to the P sets of second face features, any one of the P sets of second face features includes R categories of second face features, the R categories of second face features correspond to R categories of first face features in a second target face feature set, and the second target face feature set is the first face feature set in the P sets of first face features that corresponds to any one of the second face feature sets. In summary, any one of the R categories of second face features is obtained by multiplying a first target face feature and a first target combination weight. The first target face feature is the first face feature of the category corresponding to the second face feature of any one of the R categories, and the first target combination weight is the first combination weight corresponding to the first target face feature. The second face features of the same category in the P sets of second face features are added together to obtain R third face features. The first clustering label is multiplied by each of the R third face features to obtain R fourth face features. The R fourth face features are then combined to form the first combined face feature.

[0304] In one possible implementation, the first clustering label is obtained by one-hot encoding of the second clustering label, the second clustering label is obtained by processing the similarity matrix using a preset clustering method, the similarity matrix is ​​obtained based on the first self-expression matrix, the first self-expression matrix is ​​obtained by training the second self-expression matrix based on multiple first face features, the multiple first face features are obtained by inputting multiple second random vectors into the face generator respectively, and the multiple first face features are the output of the target convolutional neural network module.

[0305] In one possible implementation, the first self-expression matrix is ​​obtained through the following operations: For the plurality of first face features, the following operations are performed to obtain the first self-expression matrix: S11: Multiply the fourth target face feature by the first target self-expression matrix to obtain a fourth face feature, wherein the fourth target face feature is one of the plurality of first face features; S12: Obtain a second synthetic face image based on the fourth face feature; S13: Obtain a first loss based on the fourth target face feature and the second synthetic face image; S14: If the first loss is less than a first preset threshold, then... The first target self-expression matrix is ​​the first self-expression matrix; otherwise, adjust the elements in the first target self-expression matrix according to the first loss to obtain the second target self-expression matrix, and execute step S15; S15: take the fifth target face feature as the fourth target face feature, and take the second target self-expression matrix as the first target self-expression matrix, and continue to execute steps S11 to S14, wherein the fifth target face feature is the first face feature among the plurality of first face features that has not yet been used for training; wherein, when executing step S11 for the first time, the first target self-expression matrix is ​​the second self-expression matrix.

[0306] The face image processing device 1200 may further include a storage unit 1203 for storing program code and data of the electronic device. The processing unit 1201 may be a processor, the communication unit 1202 may be a transceiver, and the storage unit 1203 may be a memory.

[0307] It should be noted that the implementation of each unit can also be referenced accordingly. Figure 6 The corresponding description of the method embodiments shown; Figure 12 The beneficial effects of the described face image processing device 1200 can also be referred to accordingly. Figure 6 The corresponding description of the method embodiments shown.

[0308] Please see Figure 13 , Figure 13 This is a schematic diagram of the structure of an electronic device 1310 provided in an embodiment of this application. The electronic device 1310 includes a transceiver 1311, a processor 1312, and a memory 1313. The transceiver 1311, the processor 1312, and the memory 1313 are interconnected through a bus 1314.

[0309] The memory 1313 includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or compact disc read-only memory (CD-ROM), which is used for related instructions and data.

[0310] Transceiver 1311 is used to receive and send data.

[0311] The processor 1312 can be one or more central processing units (CPUs). When the processor 1312 is a CPU, the CPU can be a single-core CPU or a multi-core CPU.

[0312] The processor 1312 in the electronic device 1310 is used to read the program code stored in the memory 1313 and execute it. Figure 6 The method shown.

[0313] It should be noted that the implementation of each operation can also be referenced accordingly. Figure 6 The corresponding description of the illustrated embodiments; Figure 13 The beneficial effects of the described electronic device 1310 can also be referred to accordingly. Figure 6 The corresponding description of the method embodiments shown.

[0314] In some embodiments, the disclosed method may be implemented as computer program instructions encoded in a machine-readable format on a computer-readable storage medium or on other non-transitory media or articles of art. Figure 14 A conceptual partial view schematically illustrates an example computer program product arranged according to at least some embodiments shown herein, the example computer program product comprising a computer program for executing computer processes on a computing device. In one embodiment, the example computer program product 1400 is provided using a signal carrying medium 1401. The signal carrying medium 1401 may include one or more program instructions 1402 that, when executed by one or more processors, can provide the above-described instructions for… Figure 6 The described function or part of the function. Therefore, for example, refer to... Figure 6 In the embodiment shown, one or more features of blocks 601-605 can be represented by one or more instructions associated with the signal carrying medium 1401. Furthermore, Figure 14 The program instruction 1402 in the document also describes example instructions.

[0315] In some examples, signal-bearing medium 1401 may include computer-readable medium 1403, such as, but not limited to, hard disk drive, compact disc (CD), digital video optical disc (DVD), digital magnetic tape, memory, read-only memory (ROM), or random access memory (RAM), etc. In some embodiments, signal-bearing medium 1401 may include computer-recordable medium 1404, such as, but not limited to, memory, read / write (R / W) CD, R / W DVD, etc. In some embodiments, signal-bearing medium 1401 may include communication medium 1405, such as, but not limited to, digital and / or analog communication media (e.g., fiber optic cable, waveguide, wired communication link, wireless communication link, etc.). Therefore, for example, signal-bearing medium 1401 may be conveyed by wireless communication medium 1405 (e.g., wireless communication medium conforming to the IEEE 802.11 standard or other transmission protocols). One or more program instructions 1402 may be, for example, computer-executable instructions or logical implementation instructions. In some examples, such as for... Figure 13 The described electronic device can be configured to provide various operations, functions, or actions in response to program instructions 1402 transmitted to a computing device via one or more of computer-readable media 1403, computer-recordable media 1404, and / or communication media 1405. It should be understood that the arrangements described herein are merely illustrative. Therefore, those skilled in the art will understand that other arrangements and other elements (e.g., machines, interfaces, functions, sequences, and functional groups, etc.) can be used instead, and some elements can be omitted depending on the desired result. Furthermore, many of the described elements are functional entities that can be implemented as discrete or distributed components, or in any suitable combination and location in conjunction with other components.

[0316] This application also provides a chip, which includes at least one processor, a memory, and an interface circuit. The memory, the transceiver, and the at least one processor are interconnected via circuits. The at least one memory stores a computer program. When the computer program is executed by the processor... Figure 6 The method and flow shown are thus implemented.

[0317] This application also provides a computer-readable storage medium storing a computer program that, when run on an electronic device. Figure 6 The method and flow shown are thus implemented.

[0318] It should be understood that the processor mentioned in the embodiments of this application can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0319] It should also be understood that the memory mentioned in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).

[0320] It should be noted that when the processor is a general-purpose processor, DSP, ASIC, FPGA, or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component, the memory (storage module) is integrated into the processor.

[0321] It should be noted that the memories described in this specification are intended to include, but are not limited to, these and any other suitable types of memories.

[0322] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0323] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this specification can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0324] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0325] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0326] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0327] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0328] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0329] The steps in the methods of this application can be adjusted, combined, or deleted according to actual needs. Furthermore, the terminology and explanations in the embodiments of this application can be referred to the corresponding descriptions in other embodiments.

[0330] The modules in the device of this application embodiment can be merged, divided, and deleted according to actual needs.

[0331] The above description and embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. A method for processing facial images, characterized in that, include: Obtain low-quality face images and first cluster labels; Feature extraction is performed on the low-quality face image to obtain the first target face feature and the second target face feature; Based on the first clustering label, each of the P third target face features is divided into R categories of first face sub-features to obtain a set of P first face sub-features. Any one of the P first face sub-feature sets includes R categories of first face sub-features, where P is a positive integer and R is an integer greater than 1. The P third target face features are the output of the target convolutional neural network module of the face generator, and the input of the target convolutional neural network module corresponding to the P third target face features is obtained based on the first target face features. Based on the second target face features and the first clustering label, the first face features in the P sets of first face sub-features are combined into a first combined face feature; A first synthetic face image is obtained based on the first combination of facial features.

2. The method according to claim 1, characterized in that, The P third target face features are obtained by performing convolution modulation on the target convolutional neural network module based on the first target face features and P first random vectors.

3. The method according to claim 1 or 2, characterized in that, The P third target face features are obtained by performing convolutional modulation on the target convolutional neural network module based on the P target style vectors. The P target style vectors are obtained based on the first target face features and the P first random vectors.

4. The method according to claim 3, characterized in that, The P target style vectors are obtained based on the P first concatenation vectors. The P first concatenation vectors are obtained by concatenating the first feature vector with the P first random vectors respectively. The first feature vector is obtained based on the first target face features.

5. The method according to claim 1, characterized in that, The step of combining the first face features from the P sets of first face features into a first combined face feature based on the second target face features and the first clustering label includes: P sets of first combined weights are obtained based on the second target face feature and the P sets of first face sub-features. The P sets of first combined weights correspond to the P sets of first face sub-features. Any one of the P sets of first combined weights includes R first combined weights. The R first combined weights correspond to R categories of first face sub-features in the first target face sub-feature set. The first target face sub-feature set is the first face feature set in the P sets of first face features that corresponds to any one of the first combined weights. Any one of the R sets of first combined weights is obtained based on the second target face feature and the first face sub-features in the first target face sub-feature set that correspond to the category of any one of the first combined weights. The first facial features in the P sets of first facial sub-features are combined into the first combined facial features based on the first clustering label and the P sets of first combined weights.

6. The method according to claim 5, characterized in that, The step of combining the first face features from the P sets of first face sub-features into the first combined face features based on the first clustering label and the P sets of first combined weights includes: P sets of second face features are obtained based on the P sets of first face features and the P sets of first combination weights. The P sets of first face features correspond to the P sets of second face features. Any one of the P sets of second face features includes R categories of second face features. The R categories of second face features correspond to the R categories of first face features in the second target face feature set. The second target face feature set is the first face feature set in the P sets of first face features that corresponds to any one of the second face feature sets. Any one of the R categories of second face features is obtained by multiplying a first target face feature and a first target combination weight. The first target face feature is the first face feature of the category corresponding to the second face feature of any one of the categories. The first target combination weight is the first combination weight corresponding to the first target face feature. Add the second face features of the same category in the P sets of second face features to obtain R third face features; The first cluster label is multiplied by the R third face features to obtain R fourth face features; The R fourth facial features are combined to form the first combined facial features.

7. The method according to claim 1, characterized in that, The first clustering label is obtained by one-hot encoding the second clustering label. The second clustering label is obtained by processing the similarity matrix using a preset clustering method. The similarity matrix is ​​obtained based on the first self-expression matrix. The first self-expression matrix is ​​obtained by training the second self-expression matrix based on multiple first face features. The multiple first face features are obtained by inputting multiple second random vectors into the face generator respectively, and the multiple first face features are the output of the target convolutional neural network module.

8. The method according to claim 7, characterized in that, The first self-expression matrix is ​​obtained through the following operation: For the plurality of first facial features, perform the following operations to obtain the first self-expression matrix: S11: Multiply the fourth target face feature with the first target self-expression matrix to obtain the fourth face feature, wherein the fourth target face feature is one of the plurality of first face features; S12: Obtain the second synthetic face image based on the fourth facial feature; S13: Obtain the first loss based on the fourth target face features and the second synthesized face image; S14: If the first loss is less than the first preset threshold, then the first target self-expression matrix is ​​the first self-expression matrix; otherwise, adjust the elements in the first target self-expression matrix according to the first loss to obtain the second target self-expression matrix, and execute step S15. S15: Using the fifth target face feature as the fourth target face feature and the second target self-expression matrix as the first target self-expression matrix, continue to execute steps S11 to S14. The fifth target face feature is the first face feature among the plurality of first face features that has not yet been used for training. Specifically, during the first execution of step S11, the first target self-expression matrix is ​​the second self-expression matrix.

9. A facial image processing apparatus, characterized in that, Includes a processing unit for: Obtain low-quality face images and first cluster labels; Feature extraction is performed on the low-quality face image to obtain the first target face feature and the second target face feature; Based on the first clustering label, each of the P third target face features is divided into R categories of first face sub-features to obtain a set of P first face sub-features. Any one of the P first face sub-feature sets includes R categories of first face sub-features, where P is a positive integer and R is an integer greater than 1. The P third target face features are the output of the target convolutional neural network module of the face generator, and the input of the target convolutional neural network module corresponding to the P third target face features is obtained based on the first target face features. Based on the second target face features and the first clustering label, the first face features in the P sets of first face sub-features are combined into a first combined face feature; A first synthetic face image is obtained based on the first combination of facial features.

10. The apparatus according to claim 9, characterized in that, The P third target face features are obtained by performing convolution modulation on the target convolutional neural network module based on the first target face features and P first random vectors.

11. The apparatus according to claim 9 or 10, characterized in that, The P third target face features are obtained by performing convolutional modulation on the target convolutional neural network module based on the P target style vectors. The P target style vectors are obtained based on the first target face features and the P first random vectors.

12. The apparatus according to claim 11, characterized in that, The P target style vectors are obtained based on the P first concatenation vectors. The P first concatenation vectors are obtained by concatenating the first feature vector with the P first random vectors respectively. The first feature vector is obtained based on the first target face features.

13. The apparatus according to claim 9, characterized in that, The processing unit is specifically used for: P sets of first combined weights are obtained based on the second target face feature and the P sets of first face sub-features. The P sets of first combined weights correspond to the P sets of first face sub-features. Any one of the P sets of first combined weights includes R first combined weights. The R first combined weights correspond to R categories of first face sub-features in the first target face sub-feature set. The first target face sub-feature set is the first face feature set in the P sets of first face features that corresponds to any one of the first combined weights. Any one of the R sets of first combined weights is obtained based on the second target face feature and the first face sub-features in the first target face sub-feature set that correspond to the category of any one of the first combined weights. The first facial features in the P sets of first facial sub-features are combined into the first combined facial features based on the first clustering label and the P sets of first combined weights.

14. The apparatus according to claim 13, characterized in that, The processing unit is specifically used for: P sets of second face features are obtained based on the P sets of first face features and the P sets of first combination weights. The P sets of first face features correspond to the P sets of second face features. Any one of the P sets of second face features includes R categories of second face features. The R categories of second face features correspond to the R categories of first face features in the second target face feature set. The second target face feature set is the first face feature set in the P sets of first face features that corresponds to any one of the second face feature sets. Any one of the R categories of second face features is obtained by multiplying a first target face feature and a first target combination weight. The first target face feature is the first face feature of the category corresponding to the second face feature of any one of the categories. The first target combination weight is the first combination weight corresponding to the first target face feature. Add the second face features of the same category in the P sets of second face features to obtain R third face features; The first cluster label is multiplied by the R third face features to obtain R fourth face features; The R fourth facial features are combined to form the first combined facial features.

15. The apparatus according to claim 9, characterized in that, The first clustering label is obtained by one-hot encoding the second clustering label. The second clustering label is obtained by processing the similarity matrix using a preset clustering method. The similarity matrix is ​​obtained based on the first self-expression matrix. The first self-expression matrix is ​​obtained by training the second self-expression matrix based on multiple first face features. The multiple first face features are obtained by inputting multiple second random vectors into the face generator respectively, and the multiple first face features are the output of the target convolutional neural network module.

16. The apparatus according to claim 15, characterized in that, The first self-expression matrix is ​​obtained through the following operation: For the plurality of first facial features, perform the following operations to obtain the first self-expression matrix: S11: Multiply the fourth target face feature with the first target self-expression matrix to obtain the fourth face feature, wherein the fourth target face feature is one of the plurality of first face features; S12: Obtain the second synthetic face image based on the fourth facial feature; S13: Obtain the first loss based on the fourth target face features and the second synthesized face image; S14: If the first loss is less than the first preset threshold, then the first target self-expression matrix is ​​the first self-expression matrix; otherwise, adjust the elements in the first target self-expression matrix according to the first loss to obtain the second target self-expression matrix, and execute step S15. S15: Using the fifth target face feature as the fourth target face feature and the second target self-expression matrix as the first target self-expression matrix, continue to execute steps S11 to S14. The fifth target face feature is the first face feature among the plurality of first face features that has not yet been used for training. Specifically, during the first execution of step S11, the first target self-expression matrix is ​​the second self-expression matrix.

17. An electronic device, characterized in that, The method includes a processor, a memory, a transceiver, and one or more programs, said one or more programs being stored in the memory and configured to be executed by the processor, said programs including instructions for performing the steps of the method as described in any one of claims 1-8.

18. A chip, characterized in that, include: A processor for retrieving and running a computer program from memory, causing a device on which the chip is mounted to perform the method as described in any one of claims 1-8.

19. A computer-readable storage medium, characterized in that, It stores a computer program for electronic data interchange, wherein the computer program causes the computer to perform the method as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Data processing method and device and storage medium

    CN110046586A

  • Face image synthesis method and system, electronic equipment and storage medium

    CN112651915A