Face generation method and apparatus
Through style transfer technology, personalized lip characteristics are generated by combining target audio and style feature sequences, the problem of consistent lip styles in the existing technology is solved, and a higher authenticity of facial features and personalized style expression is achieved.
Patent Information
- Application Number
- PCT/CN2024/126654
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-21
- Filing Date
- 2024-10-23
- Publication Date
- 2025-06-26
AI Technical Summary
Existing digital life generation technology is difficult to achieve personalized lip styles in different images, resulting in the consistent lip styles of the model when reasoning and lack of personalization.
By obtaining the target audio and target style feature sequences, the target lip typographic features are predicted, and the overall facial feature sequence is generated based on the target style feature sequence, so that style transfer can be achieved to generate personalized lip typographic features.
It significantly improves the authenticity of the generated facial features, realizes the expression of personalized style, and avoids the distortion problem when changing faces at large angles.
Smart Images

Figure CN2024126654_26062025_PF_FP_ABST
Abstract
Description
Face generation method and device Technical Field
[0001] The present invention relates to the field of digital humans, and more particularly to a method and apparatus for generating a face. Background Art
[0002] Digital human technology continues to advance, gradually replacing real people in scenarios such as introductions and interactions. In the field of digital human generation, related technologies often use large-scale spoken data sets to train audio-driven lip-sync models, aiming to achieve higher generalization. However, this approach results in consistent lip-syncing styles when the model infers different avatars, hindering the realization of personalized styles.
[0003] Summary of the Invention
[0004] The present application aims to solve at least one of the technical problems existing in the prior art. To this end, the present application proposes a facial generation method and device that can generate different lip features with different styles and significant differences in style, which is conducive to the realization of personalized style.
[0005] In a first aspect, the present application provides a face generation method, the method comprising:
[0006] Obtaining target audio and target style feature sequence; the target style feature sequence is a facial feature sequence corresponding to the target style object when any audio is spoken;
[0007] Based on the target audio and the target style feature sequence, a target lip shape feature is predicted; the target lip shape feature is a lip shape feature that matches the target audio under the lip shape style corresponding to the target style object;
[0008] An overall facial feature sequence is generated based on the target lip shape feature and the target style feature sequence.
[0009] According to the facial generation method of the present application, by performing style transfer on the target style features, the target lip shape features of the target style object when broadcasting the target audio are obtained. On the basis of generating the lip shape features corresponding to each audio frame included in the target audio based on the target audio, the lip shape style of the target style object to be generated can be further combined to generate the target lip shape features that match the lip shape style. For the same target audio, different lip shape features with different styles and large style differences can be generated, which significantly improves the authenticity of the generated facial features and is conducive to the realization of personalized style.
[0010] According to one embodiment of the present application, generating an overall facial feature sequence based on the target lip shape feature and the target style feature sequence includes:
[0011] The other facial features extracted from the initial facial features corresponding to the target style feature sequence are fused with the target lip shape features to obtain the overall facial feature sequence; the other facial features are the features in the initial facial features except the lip features.
[0012] According to one embodiment of the present application, the step of fusing other facial features extracted from the initial facial features corresponding to the target style feature sequence with the target lip shape features to obtain the overall facial feature sequence includes:
[0013] Masking the lip region in the initial facial feature to obtain a first facial feature;
[0014] The target lip shape feature and the first facial feature are fused to obtain the overall facial feature sequence.
[0015] According to one embodiment of the present application, fusing the target lip shape feature and the first facial feature to obtain the overall facial feature sequence includes:
[0016] Input the target mouth shape feature and the first facial feature into a rendering module, and obtain the overall facial feature sequence output by the rendering module; wherein,
[0017] The rendering module is trained based on a first objective loss function, using a sample lip shape feature and a sample facial mask feature sequence as samples and a sample overall facial feature sequence corresponding to the sample lip shape feature and the sample facial mask feature sequence as sample labels.
[0018] According to one embodiment of the present application, predicting target lip shape features based on the target audio and the target style feature sequence includes:
[0019] The target audio and the target style feature sequence are input into the lip-sync style transfer module to obtain the target lip-sync features output by the lip-sync style transfer module; wherein,
[0020] The lip-shape style transfer module is trained based on a second objective loss function, where the second objective loss function includes at least one of a lip-shape feature loss function and a lip-shape key point loss function.
[0021] According to one embodiment of the present application, the lip style transfer module is trained by the following steps:
[0022] Obtaining sample audio and a spoken word dataset corresponding to the sample audio under multiple different sample style features;
[0023] Constructing a training sample based on the sample audio and a spoken data set corresponding to a target sample style feature among a plurality of different sample style features to obtain a plurality of training samples;
[0024] The lip style transfer module is trained based on the multiple training samples.
[0025] In a second aspect, the present application provides a face generation device, the device comprising:
[0026] The first processing module is used to obtain a target audio and a target style feature sequence; the target style feature sequence is a facial feature sequence corresponding to the target style object when any audio is played;
[0027] A second processing module is configured to obtain a target lip shape feature based on the target audio and the target style feature sequence; the target lip shape feature is a lip shape feature that matches the target audio under the lip shape style corresponding to the target style object;
[0028] The third processing module is configured to generate an overall facial feature sequence based on the target lip shape feature and the target style feature sequence.
[0029] According to the facial generation device of the present application, by performing style transfer on the target style features, the target lip shape features of the target style object when broadcasting the target audio are obtained. On the basis of generating the lip shape features corresponding to each audio frame included in the target audio based on the target audio, the lip shape style of the target style object to be generated can be further combined to generate the target lip shape features that match the lip shape style. For the same target audio, different lip shape features with different styles and large style differences can be generated, which significantly improves the authenticity of the generated facial features and is conducive to the realization of personalized style.
[0030] In a third aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the face generation method as described in the first aspect above is implemented.
[0031] In a fourth aspect, the present application provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the face generation method as described in the first aspect above.
[0032] In a fifth aspect, the present application provides a computer program product, comprising a computer program, which, when executed by a processor, implements the face generation method as described in the first aspect above.
[0033] The above one or more technical solutions in the embodiments of the present application have at least one of the following technical effects:
[0034] By performing style transfer on the target style features, the target lip shape features of the target style object when broadcasting the target audio are obtained. On the basis of generating lip shape features corresponding to each audio frame included in the target audio based on the target audio, the lip shape style of the target style object to be generated can be further combined to generate target lip shape features that match the lip shape style. For the same target audio, different lip shape features with different styles and large style differences can be generated, which significantly improves the authenticity of the generated facial features and is conducive to the realization of personalized style.
[0035] Furthermore, by masking the lip area in the initial facial features to obtain other facial features, and then fusing the masked image with the target lip shape features, the operation is simple and convenient, and can avoid problems such as severe distortion caused by face-changing at large angles.
[0036] Furthermore, by training the lip-shape style transfer module through the lip-shape feature loss function and the lip-shape key point loss function, the processing accuracy of the lip-shape style transfer module can be improved, the distortion in the lip-shape synthesis process can be effectively reduced, and the opening and closing of the lip shape can be ensured to be synchronized with the target audio and the lip shape continuity is high.
[0037] Furthermore, by training the rendering module with the reconstruction loss function and the perceptual loss function, the rendering smoothness of the rendering module can be improved, and the distortion in the lip synthesis process can be effectively reduced, so that in the final generated overall facial features, the opening and closing of the mouth is synchronized with the target audio, the lip shape is highly continuous, and has a high degree of realism.
[0038] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become obvious from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:
[0040] FIG1 is a flow chart of a face generation method according to an embodiment of the present invention;
[0041] FIG2 is a second flow chart of the face generation method provided in an embodiment of the present application;
[0042] FIG3 is a third flow chart of the face generation method provided in an embodiment of the present application;
[0043] FIG4 is a schematic diagram of the structure of a face generation device provided in an embodiment of the present application;
[0044] FIG5 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0045] The following will be combined with the accompanying drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.
[0046] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims refers to at least one of the connected objects, and the character " / " generally indicates that the objects connected are in an "or" relationship.
[0047] The face generation method, face generation device, electronic device, and readable storage medium provided in the embodiments of the present application are described in detail below with reference to the accompanying drawings through specific embodiments and their application scenarios.
[0048] The face generation method may be applied to a terminal, and may be specifically executed by hardware or software in the terminal.
[0049] The terminal includes but is not limited to portable communication devices such as mobile phones or tablet computers. It should also be understood that in some embodiments, the terminal may not be a portable communication device, but a desktop computer.
[0050] The facial generation method provided in the embodiments of the present application may be executed by an electronic device or a functional module or functional entity in the electronic device that can implement the facial generation method. The electronic devices mentioned in the embodiments of the present application include but are not limited to mobile phones, tablet computers, computers, cameras, and wearable devices. The facial generation method provided in the embodiments of the present application is described below using an electronic device as an example of the execution entity.
[0051] As shown in FIG1 , the face generation method includes steps 110 , 120 and 130 .
[0052] Step 110: Obtain target audio and target style feature sequence; the target style feature sequence is a facial feature sequence corresponding to the target style object when any audio is spoken;
[0053] In this step, the target audio is an audio file that needs to be synthesized by a digital human, such as a sentence or multiple sentences.
[0054] The target audio includes multiple audio frames.
[0055] The target style object is the style feature that the final synthesized digital human needs to present.
[0056] The target style objects can be different people, animals, or different cartoon characters, etc., and this application does not limit this.
[0057] The target style feature sequence includes a lip shape feature sequence corresponding to the target style object when speaking a certain audio; each lip shape feature in the lip shape feature sequence corresponds one-to-one to each audio frame included in the audio.
[0058] It is understandable that the audio corresponding to the target style feature sequence and the target audio are different audio files.
[0059] The target style features corresponding to the target style feature sequence may include style features of the digital human's facial features and style features when speaking, including: the size and shape of the facial features, the texture of the skin, the degree of opening of the lips when speaking, and the unique lip shape style corresponding to the specific content when speaking.
[0060] Step 120: predicting target lip shape features based on the target audio and the target style feature sequence; the target lip shape features are lip shape features that match the target audio under the lip shape style corresponding to the target style object;
[0061] In this step, the target lip shape feature is the lip shape feature that matches the target audio under the lip shape style corresponding to the target style feature sequence.
[0062] Each audio frame included in the target audio corresponds to at least one lip shape feature, thereby obtaining a target lip shape feature sequence.
[0063] Lip style may include but is not limited to: lip shape, lip size, lip skin texture, teeth, and the opening and closing range of the mouth when speaking.
[0064] It is understandable that different stylistic features may correspond to different lip shape styles. For example, different people may have different mouth opening degrees when pronouncing the phoneme "O".
[0065] By migrating the style features of the target style object when speaking other audio, the target lip shape features of the target style object when speaking the target audio can be predicted.
[0066] For example, in actual implementation, a style transfer model can be used to learn and process the target audio and target style feature sequence to predict the target lip shape features.
[0067] Among them, the style transfer model can include any realizable artificial intelligence model, which is not limited in this application.
[0068] The style transfer model is pre-trained.
[0069] In some embodiments, step 120 may include:
[0070] Input the target audio and target style feature sequence into the lip-sync style transfer module to obtain the target lip-sync features output by the lip-sync style transfer module; wherein,
[0071] The lip-shape style transfer module is trained based on a second objective loss function, where the second objective loss function includes at least one of a lip-shape feature loss function and a lip-shape key point loss function.
[0072] In this embodiment, the lip-sync style transfer module may be a deep learning model or any other achievable model, and this application does not limit this.
[0073] The second objective loss function is a loss function used to train the lip-sync style transfer module, so that the output of the lip-sync style transfer module is closer to the true value.
[0074] During the actual implementation process, the target audio and target style feature sequence are input into the lip-sync style transfer module, and the lip-sync style transfer module can directly output the target lip-sync features.
[0075] In some embodiments, the lip-sync style transfer module may include: an audio content encoder Ea, a style encoder Es, and a lip-sync style transfer decoder Ds.
[0076] The audio content encoder Ea is used to encode the input target audio, and the audio content encoder Ea can use a pre-trained Hubert encoder.
[0077] The style encoder Es and the lip-sync style transfer decoder Ds can be trained accordingly based on sample data in actual scenarios.
[0078] The style encoder Es is used to encode style features.
[0079] In some embodiments, the lip-shape style transfer module may be trained based on a second objective loss function, and the second objective loss function may include at least one of a lip-shape feature loss function and a lip-shape key point loss function.
[0080] In some embodiments, the lip shape feature loss function can be expressed as:
[0081] L feat =||D s (E s (x i ),E a (y))-F e ||2
[0082] Among them, L feat is the lip shape feature loss function; D s is the lip-sync style transfer decoder; D s () represents the lip-shape features reconstructed using the lip-shape style transfer decoder; E s is the style encoder; E s () represents the features extracted by the style encoder; E a F is the audio content encoder; e is the sample mouth shape feature; x i are aligned video frames of different lip shapes (i.e., facial features); y is the sample audio input to the audio content encoder.
[0083] In some embodiments, the lip keypoint loss function can be expressed as:
[0084] L lmk =||K(D s (E s (x i ),E a (y)))-K(F e )||2
[0085] Among them, L lmk is the lip shape key point loss function; D s is the lip-sync style transfer decoder; E s is the style encoder; E a F is the audio content encoder; e is the sample mouth shape feature; K(*) is used to render the sample mouth shape feature into the mouth shape key point function; x i is the i-th sample style feature of the input style encoder; y is the sample audio of the input audio content encoder.
[0086] In some embodiments, Deep3DMM technology may be used to extract lip shape features from a sample video, and the extracted lip shape features may be rendered to obtain a lip shape key point function.
[0087] In some embodiments, the lip-sync style transfer module may be trained based on a second objective loss function, which may be:
[0088] L total =L feat +λLlmk
[0089] Among them, L total is the second objective loss function; L feat is the lip shape feature loss function; L lmk is the lip-shape key point loss function; λ is the second target weight.
[0090] In this embodiment, the second objective loss function may include a lip shape feature loss function and a lip shape key point loss function.
[0091] The second target weight may be obtained based on experimental verification, or may be determined based on historical experience. For example, the value range of the second target weight may be set to be between 0.01 and 0.04.
[0092] During the training process, the lip shape feature loss function and the lip shape key point loss function can be integrated to perform overall training on the lip shape style transfer module.
[0093] According to the facial generation method provided in the embodiment of the present application, the lip style transfer module is trained by using the lip feature loss function and the lip key point loss function, which can improve the processing accuracy of the lip style transfer module, effectively reduce the distortion in the lip synthesis process, and ensure that the opening and closing of the lip shape is synchronized with the target audio and the lip shape continuity is high.
[0094] The following describes the training method of the lip-sync style transfer module.
[0095] As shown in FIG2 , in some embodiments, the lip-sync style transfer module can be trained by the following steps:
[0096] Obtain sample audio and a spoken word dataset with various sample style features corresponding to the sample audio;
[0097] Building a training sample based on a spoken word data set corresponding to a target sample style feature among the sample audio and a plurality of different sample style features, thereby obtaining a plurality of training samples;
[0098] Based on multiple training samples, train the lip-sync style transfer module.
[0099] In this embodiment, the sample audio may be any audio file, and the sample audio may be multiple audio files.
[0100] The oral data set includes: the facial features corresponding to a certain sample style feature when the sample audio is spoken, and the sample lip shape features corresponding to each audio frame included in the sample audio.
[0101] During the actual implementation process, the oral data set can be obtained by extracting the sample overall facial feature sequence corresponding to a certain sample style feature when the sample audio is spoken, wherein the number of image frames included in the sample overall facial feature sequence is the same as the number of audio frames included in the sample audio.
[0102] The target sample style feature may be any style feature among a variety of different sample style features.
[0103] Constructing training samples based on sample audio and spoken data corresponding to target sample style features among multiple different sample style features can include: establishing positive samples based on target audio frames in the sample audio and video frames corresponding to the target audio frames, and establishing negative samples based on target audio frames in the sample audio and the other video frames.
[0104] It is understandable that in the actual execution process, different image oral video data can be obtained. By separating the audio and video of each image oral video data, the sample audio and oral data set corresponding to the image oral video data can be obtained.
[0105] The following description will be made using a certain image spoken video data as an example.
[0106] During the actual execution process, audio data aligned with the video data included in the image's spoken video data and aligned face images of other different lip shapes under the image can be obtained as input to the lip style transfer module to train the lip style transfer module.
[0107] Step 130: Generate an overall facial feature sequence based on the target lip shape feature and the target style feature sequence.
[0108] In this step, the target lip shape feature may be expressed as a feature based on a time series change, and the target lip shape feature changes accordingly with the change of the audio frames included in the target audio.
[0109] The overall facial feature sequence is a sequence that includes facial features of the style corresponding to the target style feature sequence, and the lip shape can change accordingly based on the change of the audio, as shown in the final facial image in Figure 3.
[0110] Continuing to refer to Figure 3, in the actual implementation process, for the audio file to be synthesized, such as a speech, the speech can be input into the audio content encoder Ea for encoding to obtain the encoded target audio frame sequence; the required target style feature sequence is selected, such as obtaining any target image lip style sequence as shown in Figure 3, and the target image lip style sequence is input into the style encoder Es for encoding to obtain the lip style sequence of the target image when speaking any audio segment.
[0111] The target audio frame sequence and lip style sequence are then input into the lip style transfer decoder Ds for processing, so as to transfer the lip style of the target image to the lip features corresponding to the target audio frame sequence, thereby obtaining the target lip features.
[0112] By fusing the target lip shape features and the target style feature sequence corresponding to the target image, the overall facial features can be generated.
[0113] It is understandable that the same sentence may have different effects when said by different people.
[0114] During the research and development process, the inventors discovered that in related technologies, audio-driven lip models are mainly trained on large-scale oral broadcast data sets. This method easily leads to the average face problem in the model. That is, when the model is inferring, the lip opening and closing amplitude and lip and tooth similarity between different images are extremely high, which is not conducive to the realization of personalized style, thus affecting the user experience.
[0115] In this application, style transfer is introduced on the basis of lip shape prediction. On the basis of generating lip shape features corresponding to each audio frame included in the target audio based on the target audio, the lip shape style of the image to be generated is further combined to generate target lip shape features that match the image style. For the same target audio, different lip shape features with different styles and large differences can be generated, which is conducive to the realization of personalized style.
[0116] In addition, it can also make the lip and teeth synthesis more similar to the original image, eliminating the need for image fine-tuning and improving the realism of the generated digital human.
[0117] According to the facial generation method provided in the embodiment of the present application, by performing style transfer on the target style features, the target lip shape features of the target style object when broadcasting the target audio are obtained. On the basis of generating the lip shape features corresponding to each audio frame included in the target audio based on the target audio, the lip shape style of the target style object to be generated can be further combined to generate the target lip shape features that match the lip shape style. For the same target audio, different lip shape features with different styles and large style differences can be generated, which significantly improves the authenticity of the generated facial features and is conducive to the realization of personalized style.
[0118] In some embodiments, step 130 may include:
[0119] The other facial features extracted from the initial facial features corresponding to the target style feature sequence are fused with the target lip shape features to obtain the overall facial features; the other facial features are the features in the initial facial features except the lip features.
[0120] In this embodiment, other facial features are features of the initial facial features except for the lip features, such as eyes, nose, and ears.
[0121] After obtaining other facial features, the extracted facial features are fused with the lip shape features corresponding to the predicted target audio to obtain the overall facial features.
[0122] In some embodiments, other facial features extracted from the initial facial features corresponding to the target style feature sequence are fused with the target lip shape features to obtain the overall facial features, which may include:
[0123] Masking the lip area in the initial facial feature to obtain the first facial feature;
[0124] The target mouth shape feature and the first facial feature are fused to obtain the overall facial feature.
[0125] In this embodiment, the first facial feature is the feature of the initial facial features except the lip feature.
[0126] By masking the lip area in the initial facial features to obtain other facial features, and then fusing the masked image with the target lip shape features, the operation is simple and convenient, and can avoid the problems such as serious distortion caused by large-angle face swapping caused by the existing technology of performing an overall face swap operation on the face in the video after generating the video.
[0127] In actual application, there is no need to retrain the model for new audio data and image data. The model can adaptively transfer the lip style and lip shape features according to the new image features to reconstruct the facial image and obtain the overall facial features.
[0128] In some embodiments, the reconstructed face image can be expressed as:
[0129] Among them, D s is the lip-sync style transfer decoder; E s is the style encoder; E a E is the audio content encoder; v () indicates that identity features are extracted using the identity encoder; I s is a sequence of masked facial features (ie, the first facial feature).
[0130] Among them, E v The input of () is a sequence of masked facial features, E v () is used to blend and connect the input value with other areas such as the background (such as the hair and background of the target style object) and perform posture simulation.
[0131] In the actual implementation process, other facial features can be integrated with the target mouth shape features through a neural network model or other feasible models to obtain the overall facial features. Please do not train it here.
[0132] In some embodiments, fusing the target lip shape feature and the first facial feature to obtain an overall facial feature includes:
[0133] Input the target mouth shape feature and the first facial feature into the rendering module, and obtain the overall facial feature output by the rendering module; wherein,
[0134] The rendering module is trained based on the first objective loss function, using the sample lip shape feature and the sample facial mask feature sequence as samples and the sample overall facial feature sequence corresponding to the sample lip shape feature and the sample facial mask feature sequence as sample labels.
[0135] In this embodiment, the sample lip shape features can be obtained by extracting the lip shape features in the sample video using the Deep3DMM technology.
[0136] The sample facial mask feature sequence can be obtained by masking the lip region in the sample overall facial feature sequence.
[0137] The rendering module may include: an identity encoder Ev and a rendering decoder G.
[0138] Among them, both the identity encoder Ev and the rendering decoder G need to be trained.
[0139] In some embodiments, the first objective loss function may be expressed as:
[0140] L render =L recon +βL vgg
[0141] Among them, L render is the first objective loss function; L recon is the reconstruction loss function; L vgg is the perceptual loss function; β is the first target weight.
[0142] In this embodiment, the first target weight may be obtained based on experimental verification, or may be determined based on historical experience. For example, the value range of the first target weight may be set to between 0 and 1.
[0143] In some embodiments, the reconstruction loss function may be expressed as:
[0144] Among them, L recon is the reconstruction loss function; is the facial features reconstructed by the rendering decoder G; Ii is the true facial feature of the i-th frame of the original video (i.e., the i-th frame in the overall facial feature sequence of the sample); M is the total number of frames of the original video, which is the same as the number of frames of the input sample audio; i is the index of each frame in the original video.
[0145] In some embodiments, the perceptual loss function may be expressed as:
[0146] Among them, L vgg is the perceptual loss function; is the facial features reconstructed by the rendering decoder G; I i is the true facial feature of the i-th frame of the original video (i.e., the i-th frame in the overall facial feature sequence of the sample); Vn() represents the n-th layer feature extractor of the VGG19 pre-trained model; M is the total number of frames of the original video, which is the same as the number of audio frames of the input sample; N is the total number of feature extractors in the VGG19 pre-trained model; and i is the index of each frame in the original video.
[0147] During the training process, the reconstruction of the face image based on the rendering module can be expressed as:
[0148] Among them, F e is the sample mouth shape feature; E v () indicates that identity features are extracted using the identity encoder; I s is a sequence of masked facial features; G(*) represents the reconstruction of the face using the rendering decoder.
[0149] According to the facial generation method provided in the embodiment of the present application, by training the rendering module through the reconstruction loss function and the perceptual loss function, the rendering smoothness of the rendering module can be improved, and the distortion in the lip synthesis process can be effectively reduced, so that in the overall facial features finally generated, the opening and closing of the mouth is synchronized with the target audio, the lip continuity is high, and it has a high degree of realism.
[0150] The method provided in the embodiment of the present application can realize the adaptive transfer of lip style and lip shape features when only audio and partial lip shape pictures of the image are provided; it enables the universal audio-driven lip model to have different lip styles for different images while having high generalization, thereby improving the realism and style diversity of the lip shapes generated by virtual digital humans; in addition, only one training is required, and in subsequent applications, for new style features, there is no need to re-train, and the style sequence only needs to be input into the pre-trained model to obtain the style sequence and the overall facial features under the audio, which is simple and convenient to operate.
[0151] The face generation method provided in the embodiment of the present application can be executed by a face generation device. In the embodiment of the present application, the face generation device provided in the embodiment of the present application is described by taking the face generation method executed by the face generation device as an example.
[0152] An embodiment of the present application also provides a face generation device.
[0153] As shown in FIG. 4 , the face generation device includes a first processing module 410 , a second processing module 420 and a third processing module 430 .
[0154] The first processing module 410 is used to obtain a target audio and a target style feature sequence; the target style feature sequence is a facial feature sequence corresponding to the target style object when any audio is spoken;
[0155] The second processing module 420 is configured to predict target lip shape features based on the target audio and the target style feature sequence; the target lip shape features are lip shape features that match the target audio under the lip shape style corresponding to the target style object;
[0156] The third processing module 430 is configured to generate an overall facial feature sequence based on the target lip shape feature and the target style feature sequence.
[0157] According to the facial generation device provided in the embodiment of the present application, by performing style transfer on the target style features, the target lip shape features of the target style object when broadcasting the target audio are obtained. On the basis of generating the lip shape features corresponding to each audio frame included in the target audio based on the target audio, the lip shape style of the target style object to be generated can be further combined to generate the target lip shape features that match the lip shape style. For the same target audio, different lip shape features with different styles and large style differences can be generated, which significantly improves the authenticity of the generated facial features and is conducive to the realization of personalized style.
[0158] In some embodiments, the third processing module 430 may also be used to:
[0159] The other facial features extracted from the initial facial features corresponding to the target style feature sequence are fused with the target lip shape features to obtain the overall facial feature sequence; the other facial features are the features in the initial facial features except the lip features.
[0160] In some embodiments, the third processing module 430 may also be used to:
[0161] Masking the lip area in the initial facial feature to obtain the first facial feature;
[0162] The target mouth shape features and the first facial features are fused to obtain the overall facial feature sequence.
[0163] In some embodiments, the third processing module 430 may also be used to:
[0164] Input the target mouth shape feature and the first facial feature into the rendering module, and obtain the overall facial feature sequence output by the rendering module; wherein,
[0165] The rendering module is trained based on the first objective loss function, using the sample lip shape feature and the sample facial mask feature sequence as samples and the sample overall facial feature sequence corresponding to the sample lip shape feature and the sample facial mask feature sequence as sample labels.
[0166] In some embodiments, the second processing module 420 may also be used to:
[0167] Input the target audio and target style feature sequence into the lip-sync style transfer module to obtain the target lip-sync features output by the lip-sync style transfer module; wherein,
[0168] The lip-shape style transfer module is trained based on a second objective loss function, where the second objective loss function includes at least one of a lip-shape feature loss function and a lip-shape key point loss function.
[0169] In some embodiments, the apparatus may further include a fourth processing module configured to:
[0170] Obtain sample audio and a spoken word dataset with various sample style features corresponding to the sample audio;
[0171] Building a training sample based on a spoken word data set corresponding to a target sample style feature among the sample audio and a plurality of different sample style features, thereby obtaining a plurality of training samples;
[0172] Based on multiple training samples, train the lip-sync style transfer module.
[0173] The facial generation device in the embodiments of the present application can be an electronic device or a component of an electronic device, such as an integrated circuit or chip. The electronic device can be a terminal or other device other than a terminal. For example, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, a mobile internet device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. It can also be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc., and the embodiments of the present application do not specifically limit this.
[0174] The face generation device in the embodiment of the present application may be a device having an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.
[0175] The face generation device provided in the embodiment of the present application can implement each process implemented in the method embodiments of Figures 1 to 3. To avoid repetition, they are not described here.
[0176] In some embodiments, as shown in Figure 5, the embodiment of the present application further provides an electronic device 500, including a processor 501, a memory 502, and a computer program stored in the memory 502 and executable on the processor 501. When the program is executed by the processor 501, the various processes of the above-mentioned face generation method embodiment are implemented, and the same technical effects can be achieved. To avoid repetition, they will not be described here.
[0177] It should be noted that the electronic devices in the embodiments of the present application include the mobile electronic devices and non-mobile electronic devices mentioned above.
[0178] The embodiment of the present application further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the various processes of the above-mentioned face generation method embodiment are implemented, and the same technical effects can be achieved. To avoid repetition, they are not described here.
[0179] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0180] An embodiment of the present application further provides a computer program product, including a computer program, which implements the above-mentioned face generation method when executed by a processor.
[0181] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0182] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned face generation method embodiment, and can achieve the same technical effects. To avoid repetition, they are not described here.
[0183] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.
[0184] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be noted that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0185] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0186] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.
[0187] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "illustrative embodiments," "examples," "specific examples," or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0188] Although the embodiments of the present application have been shown and described, those skilled in the art will appreciate that various changes, modifications, substitutions, and variations may be made to the embodiments without departing from the principles and intent of the present application, and that the scope of the present application is defined by the claims and their equivalents.
Claims
1. A face generation method, characterized in that: include: Obtain target audio and target style feature sequence; The target style feature sequence is a facial feature sequence corresponding to the target style object when any audio is played orally; Based on the target audio and the target style feature sequence, predict a target lip shape feature; The target lip shape feature is a lip shape feature that matches the target audio under the lip shape style corresponding to the target style object; Based on the target lip shape feature and the target style feature sequence, an overall facial feature sequence is generated.
2. The face generation method according to claim 1, characterized in that: The step of generating an overall facial feature sequence based on the target lip shape feature and the target style feature sequence comprises: The other facial features extracted from the initial facial features corresponding to the target style feature sequence are fused with the target lip shape features to obtain the overall facial feature sequence; the other facial features are the features in the initial facial features except the lip features.
3. The face generation method according to claim 2, characterized in that: The other facial features extracted from the initial facial features corresponding to the target style feature sequence are fused with the target lip shape features to obtain the overall facial feature sequence, including: Masking the lip region in the initial facial feature to obtain a first facial feature; The target mouth shape feature and the first facial feature are fused to obtain the overall facial feature sequence.
4. The face generation method according to claim 3, characterized in that: The step of fusing the target mouth shape feature and the first facial feature to obtain the overall facial feature sequence includes: The target mouth shape feature and the first facial feature are input into a rendering module to obtain an overall facial feature sequence output by the rendering module; wherein, The rendering module is trained based on a first target loss function, using sample lip shape features and sample facial mask feature sequences as samples and a sample overall facial feature sequence corresponding to the sample lip shape features and the sample facial mask feature sequence as sample labels.
5. The face generation method according to any one of claims 1 to 4, characterized in that: The step of predicting a target lip shape feature based on the target audio and the target style feature sequence comprises: The target audio and the target style feature sequence are input into the lip style transfer module to obtain the target lip features output by the lip style transfer module; wherein, The lip style transfer module is trained based on the second target loss function. The function includes at least one of a lip feature loss function and a lip key point loss function.
6. The face generation method according to claim 5, characterized in that: The lip style transfer module is trained by the following steps: Obtain sample audio and a spoken data set under a plurality of different sample style features corresponding to the sample audio; Building a training sample based on the sample audio and a spoken data set corresponding to a target sample style feature among a plurality of different sample style features to obtain a plurality of training samples; Based on the multiple training samples, the lip style transfer module is trained.
7. A facial generation device, characterized in that: include: A first processing module, used to obtain a target audio and a target style feature sequence; The target style feature sequence is a facial feature sequence corresponding to the target style object when any audio is played orally; A second processing module, configured to predict a target lip shape feature based on the target audio and the target style feature sequence; The target lip shape feature is a lip shape feature that matches the target audio under the lip shape style corresponding to the target style object; The third processing module is used to generate an overall facial feature sequence based on the target lip shape feature and the target style feature sequence.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the face generation method according to any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the face generation method according to any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the face generation method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Semantic-based audio-driven digital human generation method and system
CN112562722A
Expression generation model training method, expression generation method and device
CN116468826A
Video generation method and device, equipment, storage medium and product
CN116994307A
Face generation method and device
CN117765950A
Image processing method, apparatus, equipment, and storage medium
US20210365710A1