Face beautification method, device, electronic device and computer-readable storage medium

By dynamically adjusting the ratio of template faces and original face features in the face-changing image, and using encoding and decoding networks to beautify faces, the problem of lack of realism and personalization of face-changing faces in the existing technology is solved, and the face beautification effect is achieved that is more in line with mainstream aesthetics and retains personal characteristics, which enhances the attractiveness of live broadcast content.

CN118350983BActive Publication Date: 2025-06-06GUANGZHOU HUYA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410407788.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-07
Publication Date
2025-06-06
Estimated Expiration
2044-04-07

AI Technical Summary

Technical Problem

When existing face-changing technology replaces faces in real time in images or videos, it is difficult to maintain realism and difference, and lacks personalization, resulting in homogeneity of live broadcast content.

Method used

By dynamically adjusting the proportion of the template face features and the original face features to beautified in the face-changing image, the original face of the anchor is analyzed and intelligently coded using the encoding network and the decoding network to generate a face beautified image that conforms to mainstream aesthetics and retains personal characteristics.

Benefits of technology

It has achieved seamless adjustment of the anchor’s facial features during the live broadcast process to make it more in line with the mainstream aesthetics, while fully retaining personal characteristics, enhancing the integrity, coordination and authenticity of the face-changing image, thereby greatly enhancing the attractiveness and user stickiness of the live broadcast content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118350983B_ABST
    Figure CN118350983B_ABST
Patent Text Reader

Abstract

The embodiments of the present invention propose a face beautification method, device, electronic device and computer-readable storage medium, which relate to the field of image processing technology. The method dynamically adjusts the ratio of template facial features and the original facial features of the anchor to be beautified in the face-changing image by the face-changing ratio, and realizes face beautification by using the encoding network, the decoding network of the target template face, the decoding network of the original face to be beautified and the target template image, which can ensure the seamless adjustment of the anchor's facial features during the live broadcast, so that the anchor is more in line with the mainstream aesthetics and fully retains personal characteristics. At the same time, the target template facial features are corrected by using the features of the original face to be beautified, which can effectively improve the integrity, coordination and authenticity of the face-changing image, thereby greatly improving the attractiveness of the live broadcast content, providing a stable and smooth live broadcast experience, and promoting the healthy development of the live broadcast business.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a face beautification method, device, electronic device and computer-readable storage medium. Background Art

[0002] With the rapid development of the Internet and artificial intelligence technology, online live broadcast platforms have become a popular form of entertainment and communication in modern society. In various popular live broadcast platforms, the appearance and expressiveness of the anchor have become important factors in attracting viewers. Therefore, improving the image of the anchor and the visual appeal of the live broadcast room has a positive role in increasing user stickiness, increasing viewing time, and promoting consumption conversion.

[0003] At present, common face-changing technology usually replaces one face with another face in an image or video in real time, which will cause the images or videos after the face-changing to be monotonous and lack of realism and difference. Summary of the invention

[0004] In view of this, the purpose of the present invention is to provide a face beautification method, device, electronic device and computer-readable storage medium, which can dynamically adjust the proportion of the host's features in the face-swapped image, so as to generate a face-swapped image that conforms to mainstream aesthetics and retains personal characteristics.

[0005] In order to achieve the above purpose, the technical solution adopted by the embodiment of the present invention is as follows:

[0006] In a first aspect, the present invention provides a method for beautifying a face, the method comprising:

[0007] Inputting the template face image into the encoding network to obtain template features; the template features include the facial features of the target template face; the face includes the template face and the original face;

[0008] Obtaining a modified feature according to the template feature, a feature mean of the original face to be beautified, and a feature variance of the original face to be beautified;

[0009] Determining a target decoding network according to the decoding network of the target template face, the decoding network of the original face to be beautified, and the face swap ratio; wherein the decoding network corresponds to the face one by one;

[0010] The modified features are input into the target decoding network to obtain a face-changing image of the original face to be beautified.

[0011] Optionally, obtaining the modified features according to the template features, the feature mean of the original face to be beautified and the feature variance of the original face to be beautified includes:

[0012] Calculating the mean of the template feature and the variance of the template feature according to the template feature;

[0013] The modified feature is obtained according to the template feature, the mean of the template feature, the variance of the template feature, the feature mean of the original human face to be beautified and the feature variance of the original human face to be beautified.

[0014] Optionally, determining the target decoding network according to the decoding network of the target template face, the decoding network of the original face to be beautified, and the face swap ratio includes:

[0015] Determining a retention ratio according to the face-changing ratio;

[0016] The product of the face-changing ratio and the decoding network of the target template face, and the product of the retention ratio and the decoding network of the original face to be beautified are added to obtain the target decoding network.

[0017] Optionally, the decoding network is obtained by:

[0018] Obtaining a training image set and a plurality of decoding networks to be trained; the training image set includes a plurality of training images of each template face and a plurality of training images of each original face;

[0019] Inputting each training image in the training image set into the encoding network in sequence to obtain training features;

[0020] Inputting the training features into a decoding network to be trained corresponding to the training image to obtain a beautified image;

[0021] Determine total loss information of the training image according to the training image, the training features and the beautified image; the total loss information represents the difference between the training image and the beautified image;

[0022] Iteratively update the parameters of the to-be-trained decoding network corresponding to the training image according to the total loss information of the training image to obtain a plurality of trained decoding networks.

[0023] Optionally, determining the total loss information of the training image according to the training image, the training feature and the beautified image includes:

[0024] Determine reconstruction loss information according to the training image and the beautified image; the reconstruction loss information represents a pixel difference between the training image and the beautified image;

[0025] Determine face recognition loss information according to the beautified image and the training features; the face recognition loss information represents the feature vector difference between the training image and the beautified image;

[0026] The beautified image, the beautified image label, the training image and the training image label are input into a discriminator to determine discrimination loss information; the discriminator is used to determine whether the input image has not been beautified; the discrimination loss information represents the accuracy of the discrimination result output by the discriminator; the parameters of the discriminator are iteratively updated according to the total loss information of the training image;

[0027] The total loss information of the training image is determined according to the reconstruction loss information, the face recognition loss information and the discrimination loss information.

[0028] Optionally, determining face recognition loss information according to the beautified image and the training features includes:

[0029] Inputting the beautified image into the encoding network to obtain beautified image features;

[0030] According to the similarity between the training features and the beautified image features, face recognition loss information is determined.

[0031] Optionally, inputting the beautified image, the beautified image label, the training image and the training image label into a discriminator to determine the discrimination loss information includes:

[0032] Inputting the beautified image, the beautified image label, the training image and the training image label into a discriminator to obtain a recognition result of the beautified image and a recognition result of the training image;

[0033] The discrimination loss information is determined according to the recognition result of the beautified image, the beautified image label, the recognition result of the training image, and the training image label.

[0034] In a second aspect, the present invention provides a face beautification device, the device comprising:

[0035] The encoding module is used to input the template face image into the encoding network to obtain the template features; the template features include the facial features of the target template face; the face includes the template face and the original face;

[0036] A processing module, used for obtaining a modified feature according to the template feature, a feature mean of the original face to be beautified and a feature variance of the original face to be beautified;

[0037] The decoding module is used to determine the target decoding network according to the decoding network of the target template face, the decoding network of the original face to be beautified and the face-changing ratio; the decoding network corresponds to the face one-to-one; the correction feature is input into the target decoding network to obtain the face-changing image of the original face to be beautified.

[0038] In a third aspect, the present invention provides an electronic device, comprising a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to execute the face beautification method as described in any one of the aforementioned implementations when calling the computer program.

[0039] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the face beautification method as described in any of the aforementioned embodiments.

[0040] Compared with the prior art, the face beautification method, device, electronic device and computer-readable storage medium provided by the embodiments of the present invention input a template face image into an encoding network to obtain template features; the template features include facial features of the target template face. According to the template features, the feature mean of the original face to be beautified and the feature variance of the original face to be beautified, a correction feature is obtained. According to the decoding network of the target template face, the decoding network of the original face to be beautified and the face-changing ratio, a target decoding network is determined. The correction feature is input into the target decoding network to obtain a face-changing image of the original face to be beautified.

[0041] The present invention dynamically adjusts the ratio of the template face features and the original face features to be beautified in the face-changing image through the face-changing ratio, which can ensure that the host's facial features are seamlessly adjusted during the host's live broadcast, so that the host is more in line with mainstream aesthetics and fully retains personal characteristics. At the same time, the target template face features are corrected using the features of the original face to be beautified, which can effectively improve the integrity, coordination and authenticity of the face-changing image, thereby greatly improving the attractiveness of the live broadcast content, providing a stable and smooth live broadcast experience, and further promoting the healthy development of the live broadcast business.

[0042] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.

[0044] Figure 1 A schematic diagram of a process of a face beautification method provided by an embodiment of the present invention is shown.

[0045] Figure 2 Another schematic flow chart of a face beautification method provided in an embodiment of the present invention is shown.

[0046] Figure 3 A block diagram of a face beautification device provided by an embodiment of the present invention is shown.

[0047] Figure 4 A block diagram of an electronic device provided by an embodiment of the present invention is shown.

[0048] Icon: 100 - electronic device; 110 - memory; 120 - processor; 130 - communication module; 200 - face beautification device; 201 - encoding module; 202 - processing module; 203 - decoding module. DETAILED DESCRIPTION

[0049] The following will be combined with the accompanying drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.

[0050] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present invention.

[0051] It should be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0052] With the development of modern information technology, live streaming has become a new form of entertainment, playing an important role in promoting cultural exchange and entertainment diversity. In such a live streaming environment, the image of the anchor has become one of the key factors affecting the live streaming effect and user participation. Although the anchor's image is not an indicator of the quality of the anchor's content, research shows that the image has a significant impact on the user's initial impression and continued attention.

[0053] In recent years, deep learning has provided revolutionary solutions for a variety of visual and audio tasks, especially in the field of computer vision, where deep learning makes image and visual facial recognition, analysis, and editing possible. Face swapping technology, as a branch of this category, can replace a face with another face in an image or video in real time. In order to reduce excessive reliance on high-looking anchors with good images and enhance the anchors' confidence and competitiveness in live broadcasts, face swapping technology is currently commonly used to replace the anchor's face in a live video with another face.

[0054] The inventors have found that the following problems exist when the anchor directly uses face swapping technology to change faces: First, it is difficult to achieve a high degree of realism without sacrificing the dynamics of facial features and natural expressions while transferring facial features. This is especially important in live broadcasting scenarios, because the audience is extremely sensitive to the changes in the anchor's expression and emotional expression. Second, using face swapping technology to directly replace the anchor's face with a human face often produces stereotyped results, lacking personalization and differentiation. Although the anchor's face becomes better after the face swap, the lack of personal charm may lead to the homogenization of the live broadcast content.

[0055] Based on this, the face beautification method, device, electronic device and computer-readable storage medium provided by the embodiments of the present invention have the core idea of ​​dynamically adjusting the ratio of the template face features and the original face features to be beautified in the face-changing image through the face-changing ratio, which can ensure that the host's facial features are seamlessly adjusted during the host's live broadcast, so that the host is more in line with mainstream aesthetics and fully retains personal characteristics. At the same time, the target template face features are corrected using the features of the original face to be beautified, which can effectively improve the integrity, coordination and authenticity of the face-changing image, thereby greatly improving the attractiveness of the live broadcast content, providing a stable and smooth live broadcast experience, and promoting the healthy development of the live broadcast business.

[0056] The embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0057] Please refer to Figure 1 , Figure 1 A flow chart of a face beautification method provided by an embodiment of the present invention is shown. The method is applied to an electronic device, on which an encoding network and a plurality of pre-trained decoding networks are recorded, and the method comprises the following steps:

[0058] Step 10, input the template face image into the encoding network to obtain template features; the template features include the facial features of the target template face; the face includes the template face and the original face.

[0059] In an embodiment of the present invention, after receiving a beautification instruction, the target template face identifier, the original face identifier to be beautified, the template face image and the face-changing ratio carried in the beautification instruction are obtained. Among them, the template face identifier is used to uniquely identify the template face, the original face identifier is used to uniquely identify the original face, and the face-changing ratio represents the proportion of the facial features of the target template face in the face-changing image.

[0060] The encoding network encodes the input image into low-dimensional key features. The key features contain important information in the image. If the input image is a face image, the key information includes the facial features. Facial features are an important part of characterizing human appearance, including eye features, nose features, mouth features, ear features, and eyebrow features.

[0061] Step 20, obtaining the modified features according to the template features, the feature mean of the original face to be beautified and the feature variance of the original face to be beautified.

[0062] As an implementation method, a training image of the original face to be beautified is obtained according to the identification of the original face to be beautified, the training image of the original face to be beautified is input into the encoding network to obtain the features of the original face to be beautified, and the feature mean and feature variance of the original face to be beautified are determined according to the features of the original face to be beautified.

[0063] As another implementation, in the process of pre-training the decoding network, the feature mean and feature variance of each original face can be saved in the electronic device, so that when performing face beautification, the feature mean of the original face to be beautified and the feature variance of the original face to be beautified can be quickly obtained, which can effectively improve the efficiency of face beautification. The specific method for obtaining the feature mean of the original face to be beautified and the feature variance of the original face to be beautified is not limited by the present invention.

[0064] In the embodiment of the present invention, since the template features and the features of the original face to be beautified may be quite different, for example, if the relatively compact facial features of the target template face are replaced with the large face of the original face to be beautified, the face-changing effect may not be ideal. In order to ensure the integrity, coordination and authenticity of the face-changing image, the template features and the features of the face to be beautified are aligned to obtain the aligned corrected features.

[0065] Step 30, determining a target decoding network according to the decoding network of the target template face, the decoding network of the original face to be beautified, and the face swap ratio; the decoding network corresponds to the face one-to-one.

[0066] Step 40, input the corrected features into the target decoding network to obtain the face-changing image of the original face to be beautified.

[0067] In an embodiment of the present invention, a corresponding decoding network is pre-trained for each face, and the decoding network of the target template face and the decoding network of the original face to be beautified are integrated through the face-changing ratio, so as to control the ratio of the target template face features and the original face features to be beautified in the face-changing image. Different degrees of face beautification effects can be generated by setting different face-changing ratios, thereby enriching the visual performance of the anchor during live broadcast.

[0068] In summary, the face beautification method provided by the embodiment of the present invention inputs the template face image into the encoding network to obtain the template features; the template features include the facial features of the target template face. According to the template features, the feature mean of the original face to be beautified and the feature variance of the original face to be beautified, the correction features are obtained. According to the decoding network of the target template face, the decoding network of the original face to be beautified and the face-changing ratio, the target decoding network is determined. The correction features are input into the target decoding network to obtain the face-changing image of the original face to be beautified.

[0069] The present invention dynamically adjusts the ratio of the template face features and the original face features to be beautified in the face-changing image through the face-changing ratio, which can ensure that the host's facial features are seamlessly adjusted during the host's live broadcast, so that the host is more in line with mainstream aesthetics and fully retains personal characteristics. At the same time, the target template face features are corrected using the features of the original face to be beautified, which can effectively improve the integrity, coordination and authenticity of the face-changing image, thereby greatly improving the attractiveness of the live broadcast content, providing a stable and smooth live broadcast experience, and further promoting the healthy development of the live broadcast business.

[0070] Optionally, in practical applications, the template features are adjusted by feature alignment to ensure that the face-swapped image is more coordinated and natural. Figure 1 The sub-steps of step 20 may include:

[0071] The mean of the template feature and the variance of the template feature are calculated according to the template feature; the modified feature is obtained according to the template feature, the mean of the template feature, the variance of the template feature, the feature mean of the original face to be beautified and the feature variance of the original face to be beautified.

[0072] In the embodiment of the present invention, the correction feature can be obtained according to the feature alignment formula, which can be expressed as:

[0073]

[0074] Among them, Fea′ is the correction feature; Fea 1 is the template feature; μ(Fea 1 ) is the mean value of the template feature; σ(Fea 1 ) is the variance of the template feature; μ(Fea 2 ) is the feature mean of the original face to be beautified; σ(Fea 2) is the feature variance of the original face to be beautified.

[0075] It should be noted that the feature mean and feature variance of the original face to be beautified are calculated according to the features of the original face to be beautified during the decoding network training process. The features of the original face to be beautified are obtained by inputting multiple training images of the original face to be beautified into the encoding network for encoding.

[0076] Assuming that the feature dimensions of the template feature and the original face feature to be beautified are both 128*8*8, the feature dimension is 128*1*1 according to the mean and variance processing. In this way, the feature alignment can be performed according to the feature alignment formula to obtain the corrected feature. It is worth mentioning that the feature dimension of the feature vector obtained is different for different image sizes, and needs to be set according to the actual application.

[0077] Optionally, in practical applications, a specified face-changing ratio can be used to achieve a feature balance between the target template face and the original face to be beautified to varying degrees, thereby improving the appearance attractiveness of the original face to be beautified while ensuring the uniqueness and difference of the face-changing image. Figure 1 The sub-steps of step 30 may include:

[0078] According to the face-changing ratio, the retention ratio is determined; the product of the face-changing ratio and the decoding network of the target template face, and the product of the retention ratio and the decoding network of the original face to be beautified are added to obtain the target decoding network.

[0079] In the embodiment of the present invention, the face swap ratio is used to adjust the ratio of the target template face, and the value range is 0 to 1. The target decoding network can be determined by the following formula:

[0080] decoder 目标 =α×decoder 模板 +(1-α)×decoder 原始

[0081] Among them, decoder 目标 Decode the network for the target; decoder 模板 Decoder is the decoding network of the target template face; 原始 is the decoding network of the original face to be beautified; α is the face-changing ratio.

[0082] It should be noted that the decoding network of the target template face and the decoding network of the original face to be beautified are composed of multiple feature vectors of the same number, and the feature vectors of the decoding network of the target template face correspond one-to-one to the feature vectors of the decoding network of the original face to be beautified, and the one-to-one corresponding feature vectors in the decoding network of the target template face and the decoding network of the original face to be beautified have the same dimension. The face-changing ratio is multiplied in sequence with the feature vector in the decoding network of the target template face, and then the retention ratio is multiplied in sequence with the feature vector in the decoding network of the original face to be beautified, and finally the corresponding products are summed to obtain the target decoding network.

[0083] As an implementation method, during the live broadcast of the anchor, by setting the face-changing ratio to fuse the features of the anchor's original face and the features of the template face to varying degrees, the anchor's facial features can be seamlessly adjusted in real-time live broadcast, which can not only increase the attractiveness of the anchor's appearance, but also ensure the uniqueness of each anchor's appearance, and avoid the anchor losing his or her individuality after the face-changing. By beautifying the face, the threshold caused by appearance not meeting the mainstream aesthetic standards is eliminated, so that more potential anchors can get a fair chance to show themselves, making the live broadcast content more diversified and personalized, thereby enhancing the user's viewing experience, while encouraging the anchor to focus more on the instructions of the live broadcast content, and promoting the healthy development of the live broadcast business.

[0084] Optionally, before using the decoding network for face beautification, it is necessary to train multiple decoding networks to be trained to obtain trained decoding networks. The following is an explanation of the training process of the decoding network. Figure 2 , Figure 2 Another schematic diagram of the process of the face beautification method provided by the embodiment of the present invention is shown. The decoding network is obtained in the following manner:

[0085] Step 50, obtaining a training image set and a plurality of decoding networks to be trained; the training image set includes a plurality of training images of each template face and a plurality of training images of each original face.

[0086] In an embodiment of the present invention, in order to ensure the accuracy and efficiency of the decoding network training, a database containing multiple high-definition training images of template faces and original faces can be created. Each high-definition training image in the database needs to have sufficient resolution to ensure that fine facial features can be extracted. The same face is provided with as diverse training patterns as possible, such as training images covering different shooting angles, expressions and lighting conditions, so as to ensure that the generalization ability of the decoding network recognition and learning is enhanced through the diverse training images.

[0087] As an implementation method, frames are extracted from the video data of each face to obtain the corresponding training images, and multiple training images of each face are stored in a separate folder. This storage method facilitates the subsequent selection of image recognition and decoding networks, and is conducive to the decoding network learning individual differences. In order to make face beautification more targeted and personalized, the template face can be a celebrity face, and the original face can be the face of the anchor in the live broadcast.

[0088] Step 60, input each training image in the training image set into the encoding network in sequence to obtain training features.

[0089] Step 70: input the training features into the to-be-trained decoding network corresponding to the training image to obtain a beautified image.

[0090] In an embodiment of the present invention, the encoding network may be a convolutional neural network, and the decoding network to be trained may be a deconvolutional neural network or a generative adversarial network. The encoding network is used to preprocess each received training image, and a high-dimensional feature vector, i.e., a training feature, can be extracted from each training image. The training feature includes facial features, such as the position and shape features of the eyes, nose, and mouth. Training the decoding network to be trained with the training features enables the decoding network to be trained to distinguish subtle individual differences and understand the uniqueness of different faces.

[0091] Step 80, determining total loss information of the training image according to the training image, the training features and the beautified image; the total loss information represents the difference between the training image and the beautified image.

[0092] Step 90, iteratively updating the parameters of the decoding network to be trained corresponding to the training image according to the total loss information of the training image, to obtain a plurality of decoding networks after training.

[0093] In the embodiment of the present invention, in order to improve the accuracy and training efficiency of the decoding network, the discriminator can be used to identify the training image and the beautified image, and the total loss information of the training image is determined by the training image, the training features, the beautified image and the recognition result, and the parameters of the discriminator and the decoding network to be trained corresponding to the training image are iteratively updated according to the total loss information of each training image. In other words, the decoding network to be trained for each face is trained according to multiple training images of the same face, and the discriminator can be trained according to all training images.

[0094] Specifically, if the training image of the original face is encoded through a shared encoding network to obtain the training features of the original face, and then the training features of the original face are output to the decoding network to be trained corresponding to the original face, an image that carries the facial features of the original face to the greatest extent is obtained. Similarly, if the training image of the template face is encoded through a shared encoding network to obtain the training features of the template face, and then the training features of the template face are input to the decoding network to be trained corresponding to the template face, an image that carries the facial features of the template face to the greatest extent is obtained.

[0095] As an implementation method, when the total loss information corresponding to the beautified images output by each decoding network to be trained oscillates smoothly within a preset time range, the training ends to obtain multiple decoding networks after training. The training end condition can be preset according to the actual application scenario, and the present invention is not limited to this.

[0096] It should be noted that the trained decoding network can retain the facial lighting, expression, posture and face shape of the corresponding face, and then replace the facial features of the original face with the facial features of the template face, thereby obtaining a high-quality face-changing image.

[0097] Optionally, in practical applications, the total loss information of each training image can be determined by the pixel difference, feature difference and beautification discrimination result between the training image and the beautified image, thereby ensuring a more comprehensive and accurate evaluation of the loss value of the beautified image. Figure 2 The sub-steps of step 80 may include:

[0098] Reconstruction loss information is determined based on the training image and the beautified image; the reconstruction loss information represents the pixel difference between the training image and the beautified image; face recognition loss information is determined based on the beautified image and the training features; the face recognition loss information represents the feature vector difference between the training image and the beautified image.

[0099] The beautified image, the beautified image label, the training image and the training image label are input into the discriminator to determine the discriminant loss information; the discriminator is used to determine whether the input image has not been beautified; the discriminant loss information represents the accuracy of the discriminant result output by the discriminator; the parameters of the discriminator are iteratively updated according to the total loss information of the training image; the total loss information of the training image is determined according to the reconstruction loss information, the face recognition loss information and the discriminant loss information.

[0100] In the embodiment of the present invention, the total loss information of each training image is determined based on the pixel difference, feature vector difference and whether the training image has been beautified or not between the training image and the beautified image. When calculating the reconstruction loss information, the pixel mean square error of the training image and the beautified image is often used as the loss function.

[0101] It should be noted that the present invention uses a coding network and a decoding network to perform in-depth analysis and intelligent encoding and decoding of the host's original face, thereby ensuring highly realistic face beautification in the live stream, and providing a stable and smooth live experience without increasing too much computing burden. It not only greatly enhances the attractiveness and personalization of the live content, but also brings unprecedented audience experience and commercial potential to the live platform, bringing technological innovation and value enhancement to the entire live industry.

[0102] Optionally, in practical applications, face recognition loss information may be determined based on feature differences between the beautified image and the training image, and the feature differences of the beautified image may be minimized through continuous training. The step of determining face recognition loss information based on the beautified image and the training features may include:

[0103] The beautified image is input into the encoding network to obtain the beautified image features; and the face recognition loss information is determined according to the similarity between the training features and the beautified image features.

[0104] In an embodiment of the present invention, an efficient face recognition network is used as an encoding network. The facial features of the training image and the corresponding beautified image are accurately captured through the encoding network to obtain training features and beautified image features, and the similarity between the training features and the beautified image features is obtained based on the cosine values ​​of the training features and the beautified image features. The larger the cosine value, the higher the similarity.

[0105] Optionally, in practical applications, a discriminator may be used to perform high-dimensional feature recognition on the beautified image and the training image, and determine the discrimination loss information based on the recognition result. The steps of inputting the beautified image, the beautified image label, the training image and the training image label into the discriminator and determining the discrimination loss information may include:

[0106] The beautified image, the beautified image label, the training image and the training image label are input into the discriminator to obtain the recognition result of the beautified image and the recognition result of the training image; the discriminant loss information is determined according to the recognition result of the beautified image, the beautified image label, the recognition result of the training image and the training image label.

[0107] In an embodiment of the present invention, a beautified image, a beautified image label, a training image and a training image label are input into a discriminator, the discriminator is used to refine the features of the beautified image and the features of the training image, and the features of the beautified image and the features of the training image are recognized based on the beautified image label and the training image label, and the recognition results of the beautified image and the recognition results of the training image are output.

[0108] Assume that the beautified image label is set to 0 to indicate that the beautified image is an image that has been beautified. Set the training image label to 1 to indicate that the training image is an original image that has not been beautified. Assume that the discriminator identifies the beautified image, the beautified image label, the training image, and the training image label, and the recognition results of the beautified image and the training image are both 1. This means that the training image is correctly identified, while the beautified image is incorrectly identified, and the discrimination loss information is 0.5. The discrimination loss information is improved by iteratively updating the parameters of the discriminator, thereby ensuring that the face-swapped image has a more natural expression and posture.

[0109] Based on the same inventive concept, the embodiment of the present invention also provides a face beautification device. Its basic principle and technical effects are the same as those of the above embodiment. For the sake of brief description, for parts not mentioned in this embodiment, reference can be made to the corresponding contents in the above embodiment.

[0110] Please refer to Figure 3 , Figure 3 A block diagram of a face beautification device 200 provided by an embodiment of the present invention is shown. The face beautification device 200 is applied to an electronic device, and an encoding network and a plurality of pre-trained decoding networks are recorded on the electronic device. The face beautification device 200 includes an encoding module 201, a processing module 202 and a decoding module 203.

[0111] The encoding module 201 is used to input the template face image into the encoding network to obtain template features; the template features include the facial features of the target template face; the face includes the template face and the original face.

[0112] The processing module 202 is used to obtain the modified features according to the template features, the feature mean of the original face to be beautified and the feature variance of the original face to be beautified.

[0113] The decoding module 203 is used to determine the target decoding network according to the decoding network of the target template face, the decoding network of the original face to be beautified and the face-changing ratio; the decoding network corresponds to the face one-to-one; the corrected features are input into the target decoding network to obtain the face-changing image of the original face to be beautified.

[0114] In summary, the face beautification device provided by the embodiment of the present invention inputs the template face image into the encoding network to obtain the template features; the template features include the facial features of the target template face. According to the template features, the feature mean of the original face to be beautified and the feature variance of the original face to be beautified, the correction features are obtained. According to the decoding network of the target template face, the decoding network of the original face to be beautified and the face-changing ratio, the target decoding network is determined. The correction features are input into the target decoding network to obtain the face-changing image of the original face to be beautified.

[0115] The present invention dynamically adjusts the ratio of the template face features and the original face features to be beautified in the face-changing image through the face-changing ratio, which can ensure that the host's facial features are seamlessly adjusted during the host's live broadcast, so that the host is more in line with mainstream aesthetics and fully retains personal characteristics. At the same time, the target template face features are corrected using the features of the original face to be beautified, which can effectively improve the integrity, coordination and authenticity of the face-changing image, thereby greatly improving the attractiveness of the live broadcast content, providing a stable and smooth live broadcast experience, and further promoting the healthy development of the live broadcast business.

[0116] Optionally, the processing module 202 is specifically used to calculate the mean of the template features and the variance of the template features based on the template features; and obtain the modified features based on the template features, the mean of the template features, the variance of the template features, the feature mean of the original face to be beautified, and the feature variance of the original face to be beautified.

[0117] Optionally, the decoding module 203 is specifically used to determine the retention ratio according to the face-changing ratio; add the product of the face-changing ratio and the decoding network of the target template face, and the product of the retention ratio and the decoding network of the original face to be beautified to obtain the target decoding network.

[0118] Optionally, the encoding module 201 is also used to obtain a training image set and multiple decoding networks to be trained; the training image set includes multiple training images of each template face and multiple training images of each original face; each training image in the training image set is input into the encoding network in turn to obtain training features.

[0119] The decoding module 203 is further used to input the training features into the decoding network to be trained corresponding to the training image to obtain a beautified image.

[0120] The processing module 202 is also used to determine the total loss information of the training image based on the training image, the training features and the beautified image; the total loss information represents the difference between the training image and the beautified image; and the parameters of the decoding network to be trained corresponding to the training image are iteratively updated according to the total loss information of the training image to obtain multiple decoding networks after training.

[0121] Optionally, the processing module 202 is specifically used to determine reconstruction loss information based on the training image and the beautified image; the reconstruction loss information represents the pixel difference between the training image and the beautified image; determine face recognition loss information based on the beautified image and the training features; the face recognition loss information represents the feature vector difference between the training image and the beautified image; input the beautified image, the beautified image label, the training image and the training image label into the discriminator to determine the discrimination loss information; the discriminator is used to determine whether the input image has not been beautified; the discrimination loss information represents the accuracy of the discrimination result output by the discriminator; the parameters of the discriminator are iteratively updated according to the total loss information of the training image; the total loss information of the training image is determined based on the reconstruction loss information, the face recognition loss information and the discrimination loss information.

[0122] Optionally, the processing module 202 is specifically configured to input the beautified image into the encoding network to obtain beautified image features; and determine face recognition loss information according to the similarity between the training features and the beautified image features.

[0123] Optionally, the processing module 202 is specifically used to input the beautified image, the beautified image label, the training image and the training image label into the discriminator to obtain the recognition result of the beautified image and the recognition result of the training image; and determine the discrimination loss information based on the recognition result of the beautified image, the beautified image label, the recognition result of the training image and the training image label.

[0124] Please refer to Figure 4 , Figure 4 A block diagram of an electronic device 100 provided in an embodiment of the present invention. The electronic device 100 may be any device having an image processing function, such as a personal computer, a laptop computer, a server, etc. The electronic device 100 includes a memory 110, a processor 120, and a communication module 130. The memory 110, the processor 120, and the communication module 130 are electrically connected to each other directly or indirectly to achieve data transmission or interaction. For example, these components may be electrically connected to each other via one or more communication buses or signal lines.

[0125] The memory 110 is used to store programs or data. The memory 110 may be, but is not limited to, a random access memory (RAM), a read only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc.

[0126] The processor 120 is used to read / write data or programs stored in the memory 110 and execute corresponding functions. For example, when the computer program stored in the memory 110 is executed by the processor 120, the face beautification method disclosed in the above embodiments can be implemented.

[0127] The communication module 130 is used to establish a communication connection between the electronic device 100 and other communication terminals through a network, and to send and receive data through the network.

[0128] It should be understood that Figure 4 The structure shown is only a schematic diagram of the structure of the electronic device 100. The electronic device 100 may also include Figure 4 More or fewer components as shown, or with Figure 4 Different configurations are shown. Figure 4 Each component shown in the figure can be implemented by hardware, software or a combination thereof.

[0129] The embodiment of the present invention further provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by the processor 120, the face beautification method disclosed in the above embodiments is implemented.

[0130] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, a program segment or a part of a code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.

[0131] In addition, the functional modules in the various embodiments of the present invention may be integrated together to form an independent part, or each module may exist independently, or two or more modules may be integrated to form an independent part.

[0132] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0133] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A face beautification method, characterized in that: The method comprises: Inputting the template face image into the encoding network to obtain template features; the template features include the facial features of the target template face; the face includes the template face and the original face; Obtaining a modified feature according to the template feature, a feature mean of the original face to be beautified, and a feature variance of the original face to be beautified; Determining a target decoding network according to the decoding network of the target template face, the decoding network of the original face to be beautified, and the face-changing ratio, including: determining a retention ratio according to the face-changing ratio; adding the product of the face-changing ratio and the decoding network of the target template face, and the product of the retention ratio and the decoding network of the original face to be beautified, to obtain the target decoding network; the decoding network corresponds to the face one-to-one; The modified features are input into the target decoding network to obtain a face-changing image of the original face to be beautified.

2. The face beautification method according to claim 1, characterized in that: The step of obtaining the modified features according to the template features, the feature mean of the original face to be beautified and the feature variance of the original face to be beautified comprises: Calculating the mean of the template feature and the variance of the template feature according to the template feature; The modified feature is obtained according to the template feature, the mean of the template feature, the variance of the template feature, the feature mean of the original human face to be beautified and the feature variance of the original human face to be beautified.

3. The face beautification method according to claim 1, characterized in that: The decoding network is obtained in the following way: Obtaining a training image set and a plurality of decoding networks to be trained; the training image set includes a plurality of training images of each template face and a plurality of training images of each original face; Inputting each training image in the training image set into the encoding network in sequence to obtain training features; Inputting the training features into a decoding network to be trained corresponding to the training image to obtain a beautified image; Determine total loss information of the training image according to the training image, the training features and the beautified image; the total loss information represents the difference between the training image and the beautified image; Iteratively update the parameters of the to-be-trained decoding network corresponding to the training image according to the total loss information of the training image to obtain a plurality of trained decoding networks.

4. The face beautification method according to claim 3, characterized in that: The determining the total loss information of the training image according to the training image, the training feature and the beautified image comprises: Determine reconstruction loss information according to the training image and the beautified image; the reconstruction loss information represents a pixel difference between the training image and the beautified image; Determine face recognition loss information according to the beautified image and the training features; the face recognition loss information represents the feature vector difference between the training image and the beautified image; The beautified image, the beautified image label, the training image and the training image label are input into a discriminator to determine discrimination loss information; the discriminator is used to determine whether the input image has not been beautified; the discrimination loss information represents the accuracy of the discrimination result output by the discriminator; the parameters of the discriminator are iteratively updated according to the total loss information of the training image; The total loss information of the training image is determined according to the reconstruction loss information, the face recognition loss information and the discrimination loss information.

5. The face beautification method according to claim 4, characterized in that: The determining of face recognition loss information according to the beautified image and the training features includes: Inputting the beautified image into the encoding network to obtain beautified image features; According to the similarity between the training features and the beautified image features, face recognition loss information is determined.

6. The face beautification method according to claim 4, characterized in that: The step of inputting the beautified image, the beautified image label, the training image and the training image label into a discriminator to determine the discriminant loss information includes: Inputting the beautified image, the beautified image label, the training image and the training image label into a discriminator to obtain a recognition result of the beautified image and a recognition result of the training image; The discrimination loss information is determined according to the recognition result of the beautified image, the beautified image label, the recognition result of the training image, and the training image label.

7. A face beautification device, characterized in that: The device comprises: The encoding module is used to input the template face image into the encoding network to obtain the template features; the template features include the facial features of the target template face; the face includes the template face and the original face; A processing module, used for obtaining a modified feature according to the template feature, a feature mean of the original face to be beautified and a feature variance of the original face to be beautified; The decoding module is used to determine the retention ratio according to the face-changing ratio; add the product of the face-changing ratio and the decoding network of the target template face, and the product of the retention ratio and the decoding network of the original face to be beautified to obtain a target decoding network; the decoding network corresponds to the face one-to-one; the corrected features are input into the target decoding network to obtain a face-changing image of the original face to be beautified.

8. An electronic device, characterized in that: The electronic device comprises a memory and a processor, the memory is used to store a computer program, and the processor is used to execute the face beautification method according to any one of claims 1 to 6 when calling the computer program.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the face beautification method as described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Face correction model training method and device, electronic equipment and storage medium

    CN112164002A

  • Face changing model training method and device, equipment, storage medium and program product

    CN115565238A