Lip shape generation method and device and training method and device of lip shape generation model
By separately generating lip shape and texture, the lip shape generation model and training method are used to solve the problem of lip generation error in the prior art, and achieve higher generation accuracy and adaptability.
Patent Information
- Application Number
- CN202510439759.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-25
AI Technical Summary
The existing lip shape generation algorithms are prone to problems such as chromatic aberration, mismatch in speaking habits and changes in appearance when generating lip shapes, resulting in errors in the generated lip shapes.
The method of generating lip shape and lip texture separately is adopted to generate lip shape and regenerate lip texture. Through the lip generation model and training method, the lip pattern parameter sequence, cross attention mechanism and mask feature sequence are used to predict lip texture, and the generation accuracy is improved by combining a three-dimensional face model and a single-person adapter.
It improves the accuracy and accuracy of lip shape generation, reduces the impact of lip texture on shape, adapts to the generation needs of different faces, and reduces the training cost and reasoning speed.
Smart Images

Figure CN120374845A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of image generation, and particularly relates to a lip shape generation method and device, and a training method and device for a lip shape generation model. Background Art
[0002] The lip shape generation algorithm is an algorithm that uses computer technology and artificial intelligence methods to accurately replicate and simulate the lip shape of a specific person as a prototype. Such algorithms have extensive applications in scenarios such as film and television production, game development, virtual social interaction, and digital anchors. This algorithm focuses on replacing the mouth area of a human face according to the driving audio when generating the lip shape.
[0003] In related technologies, the lip shape generation algorithm includes: inputting a source human face and speech into the lip shape generation algorithm, completely covering the lip part of the source human face, and then generating a human face image containing lip pixels based on the speech and the source human face with the lip part covered.
[0004] However, although this method can achieve the matching of lip shape and speech, it inevitably has problems such as color difference between the generated lip shape and the source human face, mismatch in speaking habits, and change in appearance, that is, there are errors in the lip shape generated in related technologies. Summary of the Invention
[0005] The present disclosure provides a lip shape generation method and device, and a training method and device for a lip shape generation model, which can improve the accuracy of the generated lip shape. The technical solutions at least include the following: In a first aspect, a lip shape generation method is provided, including: obtaining a first speech, where the first speech is the speech input into the lip shape generation model; inputting the first speech and a first human face image into the lip shape generation model to obtain a second sequence of human face images, where the second sequence of human face images is used to simulate the lip shape and lip texture when the human body in the first human face image speaks the first speech; wherein, during the process of the lip shape generation model generating the second sequence of human face images, first generate the lip shape in the second sequence of human face images, and then generate the lip texture in the second sequence of human face images based on the lip shape in the second sequence of human face images.
[0006] Optionally, the lip shape generation model generates the second sequence of human face images in the following manner: based on the first speech, obtain a sequence of lip shape parameters corresponding to the first speech; use the sequence of lip shape parameters to generate a third sequence of human face images, where the human face in the third sequence of human face images is a three-dimensional human face model, and the third sequence of human face images is used to indicate the change in lip shape when the human body speaks the first speech; based on the lip part of the third sequence of human face images, predict the lip texture of the first human face image in the third sequence of human face images to obtain the second sequence of human face images.
[0007] Optionally, for the lip part based on the third face image sequence, predicting the lip texture of the first face image in the third face image sequence to obtain the second face image sequence includes: using the lip part of the third face image sequence as a mask for the first face image to obtain a mask image sequence; obtaining a mask feature sequence, where the mask feature sequence includes multiple mask feature sets, and any one mask feature set corresponds to one mask image in the mask image sequence; predicting the lip texture of the mask part in the mask image sequence based on the mask feature sequence to obtain the second face image sequence.
[0008] Optionally, obtaining the lip shape parameter sequence corresponding to the first voice includes: obtaining the voice feature sequence of the first voice; obtaining a lip shape semantic vector set, where the lip shape semantic vector set includes multiple lip shape semantic vectors, and the lip shape semantic vectors are determined based on the benchmark expressions related to lip shapes in Blendshape expressions; using a cross-attention mechanism to process the voice feature sequence and the lip shape semantic set to obtain the lip shape parameter sequence.
[0009] In a second aspect, a method for training a lip shape generation model is further provided. The method includes: obtaining the lip shape parameter sequence corresponding to a sample voice; using the lip shape parameter sequence to generate a third face image sequence; using the third face image sequence and a sample face image sequence to train the lip shape generation model, where the sample face image sequence is multiple frames of face images when a human body speaks the sample voice; where, during the process of training the lip shape generation model, using the lip part of the third face image sequence as a mask for the lip part of the first face image sequence to train the ability of the lip shape generation model to generate lip texture.
[0010] Optionally, after using the lip part of the third face image sequence as a mask for the lip part of the first face image sequence, obtaining the mask image sequence corresponding to the sample voice, the lip shape generation model further includes a single-person adapter, the single-person adapter is connected to an encoder and a decoder, the encoder is used to obtain the mask feature sequence corresponding to the mask image sequence, the decoder is used to predict the lip texture of the mask part in the mask image sequence based on the mask feature sequence, and the method further includes: fine-tuning the single-person adapter based on a first dataset, where the first dataset includes multiple voice-face image pair data of a single person, and each face image in the first dataset includes a lip shape.
[0011] In a third aspect, a lip shape generation device is further provided, including: an acquisition module, configured to acquire a first voice, where the first voice is the voice input to the lip shape generation model; a generation module, configured to input the first voice and a first face image into the lip shape generation model to obtain a second face image sequence, where the second face image sequence is used to simulate the lip shape and lip texture when the human body in the first face image speaks the first voice; wherein, in the process of the lip shape generation model generating the second face image sequence, the lip shape in the second face image sequence is first generated, and then the lip texture in the second face image sequence is generated based on the lip shape in the second face image sequence.
[0012] Optionally, the generation module is further configured to obtain a lip shape parameter sequence corresponding to the first voice based on the first voice; generate a third face image sequence using the lip shape parameter sequence, where the faces in the third face image sequence are three-dimensional face models, and the third face image sequence is used to indicate the change in the lip shape when the human body speaks the first voice; predict the lip texture of the first face image in the third face image sequence based on the lip part of the third face image sequence to obtain the second face image sequence.
[0013] Optionally, the generation module is further configured to use the lip part of the third face image sequence as a mask for the first face image to obtain a mask image sequence; obtain a mask feature sequence, where the mask feature sequence includes multiple mask feature sets, and any one mask feature set corresponds to a mask image in the mask image sequence; predict the lip texture of the mask part in the mask image sequence based on the mask feature sequence to obtain the second face image sequence.
[0014] Optionally, the acquisition module is further configured to obtain a voice feature sequence of the first voice; obtain a lip shape semantic vector set, where the lip shape semantic vector set includes multiple lip shape semantic vectors, and the lip shape semantic vectors are determined based on the benchmark expressions related to lip shapes in Blendshape expressions; process the voice feature sequence and the lip shape semantic set using a cross-attention mechanism to obtain the lip shape parameter sequence.
[0015] Fourthly, a training device for a lip shape generation model is also provided. The device includes: an acquisition module for acquiring a sequence of lip shape parameters corresponding to a sample voice; a generation module for generating a sequence of third-person face images by using the sequence of lip shape parameters; and a training module for training the lip shape generation model by using the sequence of third-person face images and a sequence of sample face images, where the sequence of sample face images is a multi-frame face image when a human body utters the sample voice. During the process of training the lip shape generation model, the lip part of the sequence of third-person face images is used as a mask for the lip part of the sequence of first-person face images to train the ability of the lip shape generation model to generate lip textures.
[0016] Optionally, the lip shape generation model further includes a single-person adapter. After using the lip part of the sequence of third-person face images as a mask for the lip part of the sequence of first-person face images, a sequence of masked images corresponding to the sample voice is obtained. The lip shape generation model further includes a single-person adapter, and the single-person adapter is connected to an encoder and a decoder. The encoder is used to obtain a sequence of masked features corresponding to the sequence of masked images, and the decoder is used to predict the lip textures of the masked part in the sequence of masked images based on the sequence of masked features. The device further includes: a fine-tuning module for fine-tuning the single-person adapter based on a first data set, where the first data set includes multiple voice-face image pair data of a single person, and each face image in the first data set includes a lip shape.
[0017] Fifthly, a computer device is also provided, including: a memory and a processor. At least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to execute the lip shape generation method and the training method of the lip shape generation model in the above embodiments.
[0018] Sixthly, a computer-readable storage medium is also provided. At least one computer program is stored in the computer-readable storage medium, and the at least one computer program is loaded and executed by a processor to execute the lip shape generation method and the training method of the lip shape generation model in the above embodiments.
[0019] Seventhly, a computer program product is provided, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the method described in the first aspect is implemented.
[0020] The beneficial effects brought by the technical solutions provided in the embodiments of the present disclosure at least include: In the related art, when generating the lip shape, the lips are completely blocked and then the lip image is directly generated. That is to say, the lip shape and the lip texture are generated simultaneously, and it is very easy to have the situation that the generated lip texture affects the accuracy of the lip shape, resulting in errors in the finally generated lip shape. In the embodiments of the present disclosure, the lip shape and the lip texture are generated separately (generating the lip shape first and then generating the lip texture). Therefore, when generating the lip texture, it will not affect the previous lip shape, thereby improving the accuracy of the finally obtained second face image sequence. Compared with the lip shape generation method in the related art, the lip shape generation method in the embodiments of the present disclosure has higher accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0022] Figure 1 The flowchart of the lip shape generation method provided by an exemplary embodiment of the present disclosure is shown; Figure 2 It is a schematic structural diagram of a lip shape generation model; Figure 3 It is a schematic structural diagram of a lip shape parameter generation module; Figure 4 It is a schematic structural diagram of a lip texture generation module; Figure 5 It is a schematic structural diagram of a lip texture generation module including a single-person adapter; Figure 6 The flowchart of the training method of the lip shape generation model provided by another exemplary embodiment of the present disclosure is shown; Figure 7 The schematic structural diagram of a lip shape generation device provided by an exemplary embodiment of the present disclosure is shown; Figure 8 The schematic structural diagram of a training device of a lip shape generation model provided by an exemplary embodiment of the present disclosure is shown; Figure 9 It is a schematic structural diagram of a computer device provided by the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] Unless otherwise defined, the technical terms or scientific terms used herein shall have the ordinary meanings as understood by those of ordinary skill in the art to which this disclosure pertains. The terms "first", "second", "third" and similar terms used in the specification and claims of this patent application of the present disclosure do not denote any order, quantity or importance, but are only used to distinguish different components. Similarly, terms such as "a" or "an" do not denote a quantity limitation, but mean that there is at least one. Terms such as "comprising" or "including" mean that the elements or objects appearing before "comprising" or "including" cover the elements or objects listed after "comprising" or "including" and their equivalents, and do not exclude other elements or objects. Terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect.
[0024] To make the objectives, technical solutions and advantages of the present disclosure more clear, the following will further describe the embodiments of the present disclosure in detail with reference to the accompanying drawings.
[0025] Figure 1 The flowchart of a lip shape generation method provided by an exemplary embodiment of the present disclosure is shown. This method can be executed by a computer device. Refer to Figure 1 , this method includes: In step 101, a first voice is obtained.
[0026] The first voice is the voice input to the lip shape generation model.
[0027] In step 102, the first voice and the first face image are input into the lip shape generation model to obtain a second face image sequence.
[0028] The second face image sequence is used to simulate the lip shape and lip texture when the human body in the first face image speaks the first voice.
[0029] Wherein, in the process of the lip shape generation model generating the second face image sequence, the lip shape in the second face image sequence is first generated, and then the lip texture in the second face image sequence is generated based on the lip shape in the second face image sequence.
[0030] Here, the first face image is a real face, equivalent to the source face, the second face image sequence is a face image sequence generated based on the first face image, and the face images in the second face image sequence are also complete faces.
[0031] In the related art, when generating a lip shape, the lips are completely blocked and then a lip image is directly generated. This is equivalent to generating the lip shape and lip texture simultaneously, and it is very easy for the generated lip texture to affect the accuracy of the lip shape, resulting in errors in the finally generated lip shape. In the embodiments of the present disclosure, the lip shape and lip texture are generated separately (generating the lip shape first and then the lip texture). Therefore, when generating the lip texture, it will not affect the previous lip shape, thereby improving the accuracy of the finally obtained second face image sequence. Compared with the lip shape generation method in the related art, the lip shape generation method in the embodiments of the present disclosure has higher accuracy.
[0032] Optionally, step 102 includes the following steps A-C. Figure 2 It is a schematic structural diagram of a lip shape generation model. The following will describe steps A-C in combination with the structure of the lip shape generation model. As Figure 2 shown, the lip shape generation model includes a lip shape parameter generation module 201, a rendering module 202, and a lip texture generation module 203 that are connected in sequence.
[0033] Step A: Based on the first speech, obtain a lip shape parameter sequence corresponding to the first speech.
[0034] The lip shape parameter generation module 201 is used to obtain a lip shape parameter sequence corresponding to the first speech based on the first speech, that is, the lip shape parameter generation module 201 is used to execute step A. Figure 3 It is a schematic structural diagram of the lip shape parameter generation module. As Figure 3 shown, the lip shape parameter generation module 201 includes a speech encoding unit 2011 and a cross-attention unit 2012 that are connected in sequence.
[0035] Optionally, step A includes the following three steps: The first step: Obtain a speech feature sequence of the first speech.
[0036] Figure 3 In, the speech encoding unit 2011 is used to obtain a speech feature sequence of the first speech, that is, the speech encoding unit is used to execute the first step of step A.
[0037] Exemplarily, a pre-trained Wav2vec2 model is used to obtain the speech feature sequence of the first speech. Through the pre-trained Wav2vec2 model, the first speech (usually a single-channel Wav format with a sampling rate of 16000 Hz) can be mapped into a speech feature sequence. The speech feature sequence can be expressed as , where respectively represent the features of the first speech corresponding to the first frame image, the features of the first speech corresponding to the second frame image... the features of the first speech corresponding to the Tth frame image.
[0038] The process of the first step is essentially a process of encoding the first voice. When implemented, the Wav2vec2 model can also be replaced with speech encoding models such as Hubert and Whisper.
[0039] In the second step, a set of lip semantic vectors is obtained.
[0040] The set of lip semantic vectors includes multiple lip semantic vectors, and the lip semantic vectors are determined based on the benchmark expressions related to lip shapes in the Blendshape expression.
[0041] In the Blendshape expression, the facial expression is divided into movements of 52 regions, such as JawOpen, MouthClose, MouthFunnel, etc., that is, the Blendshape expression has a total of 52 benchmark expressions. In the embodiments of the present disclosure, the benchmark expressions related to lip shapes in the Blendshape expression are encoded to obtain a set of lip semantic vectors. Here, a benchmark expression related to a lip shape can be encoded into a lip semantic vector.
[0042] Exemplarily, there are 27 benchmark expressions related to lip shapes in the Blendshape expression, that is, the set of lip semantic vectors includes 27 lip semantic vectors.
[0043] When encoding 27 benchmark expressions related to lip shapes into 27 lip semantic vectors, first encode these 27 benchmark expressions in onehot form, then use a semantic encoding network to extract the semantic information of each benchmark expression, and finally map the semantic information of each benchmark expression to the corresponding onehot encoding to obtain 27 lip semantic vectors.
[0044] In the third step, a cross-attention mechanism is used to process the speech feature sequence and the lip semantic set to obtain a lip parameter sequence.
[0045] Figure 3 In this, the cross-attention unit 2012 is used to process the speech feature sequence and the lip semantic set by using a cross-attention mechanism to obtain a lip parameter sequence, that is, the cross-attention unit 2012 is used to execute the third step of step A.
[0046] When using a cross-attention mechanism (Cross Attention, CA) to process the speech feature sequence and the lip semantic set, the query is the speech feature sequence, and the key and value are the lip semantic set.
[0047] In addition to the speech corresponding to the lip shape, the first speech often includes irrelevant features (such as timbre, background noise, etc.). Through the cross-attention mechanism in the third step, the lip shape parameters and speech features can be interacted, so as to dynamically extract the information related to the current frame lip shape from the speech and ignore the information irrelevant to the lip shape. For example, when the first speech contains the word "apple", by processing the first speech through the cross-attention mechanism, it is possible to focus on the lip shapes corresponding to the pronunciations of "p-ing-guo" (for example, "p" requires closing the lips), and ignore whether the timbre is sharp or thick when the first speech pronounces the word "apple".
[0048] Optionally, the cross-attention mechanism can also be combined with the self-attention mechanism (Self Attention, SA), for example, the output of the cross-attention mechanism is further processed using the self-attention mechanism.
[0049] The self-attention mechanism helps to understand the temporal structure of the speech signal (such as phonemes, intonation, rhythm). The self-attention mechanism eliminates the influence of noise or redundant information by capturing long-range dependencies. For example: the pronunciation of the previous phoneme may affect the lip shape change of the subsequent phoneme, and the self-attention mechanism can associate these long-distance speech segments. For example, in the first speech, "ping" and "guo" are consecutive, and the self-attention mechanism will mark them as a whole; there is a transition from "opening the mouth" to "closing the mouth" in the lip shape, and the self-attention mechanism can ensure the continuity of the process from "opening the mouth" to "closing the mouth" in the lip shape.
[0050] When implementing the third step, a multi-level cascaded attention mechanism can be adopted. Any layer of the attention mechanism includes cross-attention and self-attention, and the structure of each layer of the attention mechanism is the same.
[0051] As Figure 3 shown, in the i-th layer attention mechanism 301, the output of the (i - 1)-th layer attention mechanism and the lip shape semantic set are input into the cross-attention mechanism. At this time, the output of the (i - 1)-th layer attention mechanism is the query (Query), and the lip shape semantic set is the key (Key) and value (Value). Then the output of the cross-attention mechanism is input into the self-attention mechanism for processing, so as to obtain the output of the i-th layer attention mechanism.
[0052] In the case of a multi-level cascaded attention mechanism, the output of the last layer attention mechanism is the lip shape parameter sequence corresponding to the first speech.
[0053] Here, since the speech feature sequence extracted from the first speech corresponds to multiple frames of images, the lip parameter sequence extracted based on this speech feature sequence also corresponds to multiple frames of images. It is equivalent to that each speech feature in the speech feature sequence extracted from the first speech extracts a lip parameter, and each lip parameter corresponds to one frame of image, and these lip parameters constitute the lip parameter sequence.
[0054] Step B: Generate a third face image sequence using the lip parameter sequence.
[0055] The faces in the third face image sequence are three-dimensional face models, and the third face image sequence is used to indicate the lip shape changes when the human body speaks the first speech.
[0056] The rendering module 202 is used to generate a third face image sequence using the lip parameter sequence.
[0057] Exemplarily, the three-dimensional face model in the third face image sequence is FLAME (Faces Learned with an Articulated Model and Expressions, a three-dimensional face model based on an articulated model and expression learning). In this case, the rendering module 202 is the FLAME renderer.
[0058] When generating the third face image sequence from the lip parameter sequence, it is equivalent to that each lip parameter in the lip parameter sequence generates a third face image.
[0059] The three-dimensional face model does not have skin texture, but the three-dimensional face model can reflect the facial muscle movements of the face, such as lip movements, etc. Therefore, the third face image sequence can reflect the lip shape changes when the human body speaks the first speech.
[0060] Step C: Predict the lip texture of the first face image in the third face image sequence based on the lip part of the third face image sequence to obtain a second face image sequence.
[0061] The second face image sequence is the multiple frames of images corresponding to the first speech output by the lip generation model, and these multiple frames of images can be played in the form of a video.
[0062] Optionally, step C includes the following steps C1 - C3. Figure 4 It is a schematic structural diagram of the lip texture generation module, as Figure 4 shown, the lip texture generation module 203 includes a mask unit 2031, an encoder 2032, a texture generation unit 2033, and a decoder 2034 connected in sequence.
[0063] Step C1: Use the lip part of the third face image sequence as a mask for the first face image to obtain a mask image sequence.
[0064] The masking unit 2031 is used to use the lip part of the third face image sequence as a mask for the first face image to obtain a masked image sequence, that is, the masking module 2031 is used to perform step C1.
[0065] As Figure 4 shown, when implementing step C1, the masking module 2031 can splice the lip part of the third face image sequence 401 with the part above the lips of the first face image 402 based on a face key point matching algorithm to obtain a masked image sequence 403.
[0066] Since lip shape parameters are implicitly expressed, directly embedding lip shape parameters into the first face image for training will increase the learning burden of the network. Therefore, in the embodiments of the present disclosure, the lip shape parameter sequence is first rendered into a third image sequence, and then the lip part of the third face image sequence is used as a mask for the first face image, thereby reducing the learning burden of the network and improving learning efficiency.
[0067] When a conventional lip shape generation algorithm generates a lip shape, it also masks the lip part of the mask of the source face (i.e., the first face image), but this mask carries no information and is just a mask. Then, when the lip shape generation model generates a lip shape based on such a mask, it needs to start from scratch, and directly generating from scratch will increase the learning burden of the lip shape generation model.
[0068] In the embodiments of the present disclosure, the lip part of the third face image sequence is used as a mask for the first face image. Since the lip part of the third face image sequence already has a complete lip shape, the lip shape generation model only needs to generate lip textures on the basis of the lip part of the third face image sequence. It is equivalent to that the lip part of the third face image sequence serves as both a mask and has a geometric prior guiding effect on the lip shape generation model, thereby effectively reducing the learning burden of the lip shape generation model, improving the efficiency of lip shape generation, enabling replication even with a small amount of data, and not causing overfitting.
[0069] In the related art, when generating a lip shape, the lips are completely occluded and then a lip image is directly generated. It is equivalent to generating the lip shape and lip textures simultaneously, and it is very easy to have a situation where the generated lip textures affect the accuracy of the lip shape, resulting in errors in the finally generated lip shape. In the embodiments of the present disclosure, the lip shape and lip textures are generated separately (generating the lip shape first and then the lip textures), so generating lip textures will not affect the previous lip shape, thereby improving the accuracy of the finally obtained second face image sequence. Compared with the lip shape generation method in the related art, the lip shape generation method in the embodiments of the present disclosure has higher accuracy.
[0070] When generating the second sequence of facial images, the part above the lips of the source face remains unchanged. Therefore, the lip generation model only needs to generate the lip part of the third sequence of facial images to obtain the second sequence of facial images.
[0071] Step C2: Obtain the mask feature sequence.
[0072] The mask feature sequence includes multiple mask feature sets, and any one mask feature set corresponds to one mask image in the mask image sequence.
[0073] The encoder 2032 is used to convert the mask image sequence into a mask feature sequence, that is, the encoder 2032 is used to execute step C2. As Figure 4 shown, after the mask image sequence 403 is input into the encoder 2032, the encoder 2032 outputs the mask feature sequence 404.
[0074] The essence of step C2 is to use an encoder to extract features from each mask image in the mask image sequence. Since a mask feature set can be extracted from any mask image, multiple mask feature sets can be extracted from multiple mask images in the mask image sequence, that is, the mask feature sequence.
[0075] Step C3: Predict the lip texture of the masked part in the mask image sequence based on the mask feature sequence to obtain the second sequence of facial images.
[0076] Optionally, step C3 includes the following three steps.
[0077] The first step: Obtain the lip texture feature set.
[0078] The lip texture feature set can be jointly trained with the lip shape generation model, or the lip texture feature set can be trained separately. When obtaining the lip texture feature set, the lip texture data set can be obtained first, and then multiple feature vectors are extracted from the lip texture data set. These multiple feature vectors are the multiple lip texture features in the lip texture feature set.
[0079] The lip texture feature set includes various possible lip texture details (such as lip lines, reflections, wrinkles, etc.). The lip texture feature set can provide the "skin texture" of the lips, rather than controlling how the lips move.
[0080] The lip texture feature set is reusable and has nothing to do with the texture of the first voice and the content of the first facial image.
[0081] The second step: For the i-th feature vector in the first mask feature set, obtain the i-th lip texture feature in the lip texture feature set that is closest to the i-th feature vector.
[0082] The first mask feature set is any one of the mask feature sets in the mask feature sequence; The texture generation unit 2033 is configured to obtain the i-th lip texture feature closest to the i-th feature vector in the first mask feature set from the lip texture feature set, that is, the texture generation unit 2033 is configured to execute the second step of step C3.
[0083] When implementing the texture generation unit 2033, the nearest neighbor algorithm can be used to obtain the i-th lip texture feature closest to the i-th feature vector in the lip texture feature set. For each feature vector in the first mask feature set, a closest lip texture feature can be found in the lip texture feature set, and then the third step can be executed.
[0084] Thirdly, based on the lip texture features corresponding to each feature vector in the first mask feature set, generate the lip texture of the masked part in the mask image corresponding to the first mask feature set.
[0085] The decoder 2034 is configured to generate the lip texture of the masked part in the mask image corresponding to the first mask feature set based on the lip texture features corresponding to each feature vector in the first mask feature set. After the decoder generates the lip texture of the masked part in the mask image corresponding to the first mask feature set, this image is one of the images in the second face image sequence. The essence of the decoder 2034 is to restore the lip texture features corresponding to each feature vector in the first mask feature set to an image form.
[0086] For any mask feature set in the mask feature sequence, the decoder 2034 can process it using the first to third steps of step C3 above. Finally, the decoder 2034 can output the second face image sequence 405.
[0087] When a traditional lip shape generation model generates lip texture, it often provides lip texture features through some preset reference face images. However, the reference face images often also contain lip shape information, and the lip shape information will interfere with the generation of lip texture, resulting in errors in the generated lip texture. In the embodiments of the present disclosure, lip texture features are provided through a lip texture feature set, and this lip texture feature set does not contain lip shape information, avoiding the interference of lip shape information on lip texture generation, and thus improving the accuracy of the finally generated lip texture.
[0088] The lip shape generation model in the above steps A to C can relatively accurately generate a lip shape and has good versatility.
[0089] When generating lip shapes using a lip shape generation model, although the lip feature set contains various mouth features, it may still not be fine enough for faces that the lip feature set has not learned. Therefore, a single-person adapter can be additionally set in the lip texture generation module 203. When generating lip shapes for a single person, only the single-person adapter needs to be fine-tuned, thereby improving the accuracy of the lip shape generation model when generating lip shapes for a single person.
[0090] Figure 5 It is a schematic structural diagram of a lip texture generation module including a single-person adapter. In Figure 5 it, the single-person adapter 501 is respectively connected to the encoder and the decoder. The single-person adapter is, for example, a convolutional neural network. The single-person adapter 501 receives the input of multi-scale features from the encoder 2032, and through convolution, obtains adapted features and injects them into the decoder, thereby enhancing the details for a single person and achieving high-precision reproduction.
[0091] Here, before using the above lip shape generation model including a single-person adapter, the single-person adapter needs to be fine-tuned first. The method includes: fine-tuning the single-person adapter based on a first data set, where the first data set includes multiple speech-face image pairs of a single person, and each face image in the first data set includes a lip shape.
[0092] In the case of generating lip shapes for a single person, the methods in the related art often have problems such as slow inference speed and high training costs. For example, the method based on Neural Radiance Field (Nerf) can relatively well achieve generating lip shapes for a single person, but it requires per-pixel rendering and has a slow inference speed. The method based on 3DGS (3D Gaussian Splatting) improves the rendering speed, but its effect highly depends on the distribution and quality of the input data, and its ability to capture complex details (such as fine textures or fast-changing lighting effects) is relatively weak, which may lead to problems such as blurred edges or lost details.
[0093] In the embodiments of the present disclosure, when generating lip shapes for a single person through a single-person adapter, only the single-person adapter needs to be fine-tuned when applied to different people, without training the entire model, reducing the training cost when generating lip shapes for a single person. And the single-person adapter uses a conventional convolutional neural network, with a faster inference speed and being more conducive to deployment.
[0094] In related technologies, there is also a general lip shape generation method that realizes lip shape generation by using an encoder-decoder architecture. However, in order to achieve generalization performance, the general lip shape generation method often adopts a complex model architecture and a large number of parameters, resulting in a slow inference speed; the replication effect of the general lip shape generation method is insufficient, and directly using it in a single-person scenario will cause problems such as color difference, loss of speaking habits, inability to generate details of lips and teeth, and damage to identity information. If directly fine-tuning the general model for a single person, the fine-tuning strategy is complex, requiring a large number of experiments, and in related technologies, the general lip shape generation method does not explicitly decouple lip movement and texture, which easily causes overfitting and loss of the speech and lip synchronization ability of the general model itself.
[0095] The lip shape generation method in the embodiments of the present disclosure decouples the lip shape and lip texture (that is, separately generates the lip shape and lip texture), and only needs to adjust the ability of the model to generate lip texture during fine-tuning. Therefore, only a single-person adapter is required, which is not easy to cause overfitting and will not lose the speech and lip synchronization ability of the general model itself.
[0096] Figure 6 The flowchart of the training method of the lip shape generation model provided by an exemplary embodiment of the present disclosure is shown. This method can be executed by a computer device. See Figure 6 , this method includes: In step 601, obtain the lip shape parameter sequence corresponding to the sample speech.
[0097] For the relevant content of step 601, refer to the foregoing step A, which is omitted here for detailed description.
[0098] In the embodiments of the present disclosure, the structure of the lip shape generation model is the same as that of the lip shape generation model in the foregoing Figure 2 . In the lip shape generation model, step 601 is implemented through the lip shape parameter generation module 201. Therefore, it is necessary to train the lip shape parameter generation module 201 first.
[0099] In a possible implementation manner, the following steps D-F are adopted to train the lip shape parameter generation module 201.
[0100] Step D, obtain the second data set.
[0101] The second data set includes multiple speeches and multiple lip shape parameter sequences, and each speech corresponds to a lip shape parameter sequence.
[0102] Step E, input the second speech into the lip shape parameter generation module to obtain the first lip shape parameter sequence output by the lip shape parameter generation module.
[0103] The second speech is any speech data in the second data set.
[0104] Step F: Train the lip parameter generation module based on the mean squared error between the lip parameter sequence corresponding to the second voice in the second dataset and the first lip parameter sequence.
[0105] The essence of step F is to use the mean squared error (MSE) between the lip parameter sequence corresponding to the second voice in the second dataset and the first lip parameter sequence as the loss, and train the lip generation module 201 with the goal of minimizing the loss.
[0106] In another possible implementation, the following steps G - I are used to train the lip parameter generation module 201.
[0107] Step G: Generate the fourth sequence of face images using the first lip parameter sequence.
[0108] The faces in the fourth sequence of face images are three - dimensional face models.
[0109] Step H: Generate the fifth sequence of face images using the lip parameter sequence corresponding to the second voice in the second dataset.
[0110] The faces in the fifth sequence of face images are three - dimensional face models.
[0111] Here, in both step G and step H, the FLAME renderer is used to generate the sequence of face images.
[0112] Step I: Train the lip parameter generation module based on the mean squared error between the fourth image sequence and the fifth image sequence.
[0113] The essence of steps G - I is to use the mean squared error between the geometric projection images as the loss, and train the lip generation module 201 with the goal of minimizing the loss.
[0114] Optionally, when implementing the training of the lip parameter generation module 201, steps D - F and steps G - I can be used simultaneously, that is, use the mean squared error between the lip parameter sequences and the mean squared error between the geometric projection images together as the loss to implement the training of the lip parameter generation module 201.
[0115] After the lip parameter generation module 201 is trained, the mapping from voice to parameters can be realized (that is, after inputting the voice into the lip parameter generation module 201, the lip parameter generation module 201 will output the corresponding lip parameter sequence). And the process of training the lip parameter generation module 201 is independent of a specific person. Therefore, it only needs to be trained once and can be widely used for different voice inputs without further fine - tuning. When performing step 603 later, only the lip texture generation module 203 needs to be trained, and there is no need to train the lip parameter generation module 201.
[0116] In step 602, a third sequence of face images is generated using the lip parameter sequence.
[0117] For the relevant content of step 602, refer to the foregoing step B, which is omitted here for details.
[0118] In step 603, a lip generation model is trained using the third sequence of face images and the sequence of sample face images.
[0119] The sequence of sample face images is a multi-frame face image when a human body speaks the sample speech. The sample speech and the sequence of sample face images can be data in a pre-acquired training set. In this training set, any sample speech corresponds to a sequence of sample face images.
[0120] Among them, in the process of training the lip generation model, the lip part of the third sequence of face images is used as a mask for the lip part of the sequence of sample face images to train the ability of the lip generation model to generate lip textures.
[0121] Optionally, step 603 includes the following three steps.
[0122] The first step is to generate a sequence of mask images based on the third sequence of face images and the sequence of sample face images.
[0123] The second step is to generate a second sequence of face images based on the sequence of mask images.
[0124] Both the first step and the second step are implemented by the lip texture generation module 203 of the lip generation model.
[0125] The third step is to train the lip generation model using the mean square error, perceptual loss, and generative adversarial network training loss between the sequence of sample face images and the second sequence of face images.
[0126] The perceptual loss is, for example, LPIPS (Learned Perceptual Image Patch Similarity), and the generative adversarial network training loss is, for example, LSGAN (Least Squares Generative Adversarial Network).
[0127] Regarding the acquisition methods of the mean square error, perceptual loss, and generative adversarial network training loss, there are many in the related technologies, which are omitted here for details.
[0128] By training the lip generation model using the above steps 601 to 603, the accuracy of the second image sequence generated by the lip generation model can be effectively improved.
[0129] The following is a device embodiment of the present application. For details not described in detail in the device embodiment, reference may be made to the above method embodiment.
[0130] Figure 7 The structural schematic diagram of a lip shape generation device provided by an exemplary embodiment of the present disclosure is shown. Refer to Figure 7 , the lip shape generation device 700 includes: an acquisition module 701 and a generation module 702.
[0131] The acquisition module 701 is used to acquire a first voice, and the first voice is the voice input into the lip shape generation model.
[0132] The generation module 702 is used to input the first voice and the first face image into the lip shape generation model to obtain a second face image sequence, and the second face image sequence is used to simulate the lip shape and lip texture of the human body in the first face image when speaking the first voice; wherein, during the process of the lip shape generation model generating the second face image sequence, the lip shape in the second face image sequence is generated first, and then the lip texture in the second face image sequence is generated based on the lip shape in the second face image sequence.
[0133] Optionally, the generation module 702 is further used to obtain a lip shape parameter sequence corresponding to the first voice based on the first voice; generate a third face image sequence using the lip shape parameter sequence, the face in the third face image sequence is a three-dimensional face model, and the third face image sequence is used to indicate the lip shape change of the human body when speaking the first voice; predict the lip texture of the first face image in the third face image sequence based on the lip part of the third face image sequence to obtain the second face image sequence.
[0134] Optionally, the generation module 702 is further used to use the lip part of the third face image sequence as a mask for the first face image to obtain a mask image sequence; obtain a mask feature sequence, the mask feature sequence includes a plurality of mask feature sets, and any one mask feature set corresponds to a mask image in the mask image sequence; predict the lip texture of the mask part in the mask image sequence based on the mask feature sequence to obtain the second face image sequence.
[0135] Optionally, the acquisition module 701 is further used to obtain a voice feature sequence of the first voice; obtain a lip shape semantic vector set, the lip shape semantic vector set includes a plurality of lip shape semantic vectors, and the lip shape semantic vectors are determined based on the reference expressions related to lip shape in the Blendshape expression; process the voice feature sequence and the lip shape semantic set using a cross-attention mechanism to obtain a lip shape parameter sequence.
[0136] Figure 8 The structural schematic diagram of a training device for a lip shape generation model provided by an exemplary embodiment of the present disclosure is shown. Refer to Figure 8, the training device 800 of the lip shape generation model includes: an acquisition module 801, a generation module 802, and a training module 803.
[0137] The acquisition module 801 is configured to acquire a sequence of lip shape parameters corresponding to a sample voice.
[0138] The generation module 802 is configured to generate a sequence of third-person face images using the sequence of lip shape parameters.
[0139] The training module 803 is configured to train the lip shape generation model using the sequence of third-person face images and the sequence of sample face images. The sequence of sample face images is a multi-frame face image when a human body speaks the sample voice. Among them, during the process of training the lip shape generation model, the lip part of the sequence of third-person face images is used as a mask for the lip part of the sequence of first-person face images to train the ability of the lip shape generation model to generate lip textures. The first voice, the sequence of lip shape parameters, and the sequence of third-person face images are obtained using the lip shape generation method in the first aspect.
[0140] Optionally, after using the lip part of the sequence of third-person face images as a mask for the lip part of the sequence of first-person face images, a sequence of mask images corresponding to the sample voice is obtained. The lip shape generation model further includes a single-person adapter. The single-person adapter is connected to the encoder and the decoder. The encoder is configured to acquire a sequence of mask features corresponding to the sequence of mask images, and the decoder is configured to predict the lip texture of the mask part in the sequence of mask images based on the sequence of mask features. The device further includes: a fine-tuning module 804. The fine-tuning module 804 is configured to fine-tune the single-person adapter based on a first data set. The first data set includes multiple voice-face image pair data of a single person, and each face image in the first data set includes a lip shape.
[0141] It should be noted that: when the lip shape generation device provided in the above embodiment performs lip shape generation, or when the training device of the lip shape generation model trains the lip shape generation model, only the above division of each functional module is used for illustration. In practical applications, the above functions can be assigned to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the lip shape generation device provided in the above embodiment and the lip shape generation method embodiment belong to the same concept. The training device of the lip shape generation model provided in the above embodiment and the lip shape generation model training method embodiment belong to the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0142] The division of modules in the embodiments of the present disclosure is illustrative, merely a logical function division. In actual implementation, there may be other division methods. Additionally, in each embodiment of the present disclosure, the various functional modules can be integrated in a processor, can exist separately physically, or two or more modules can be integrated into one module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of a software functional module.
[0143] If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a terminal device (which can be a personal computer, a mobile phone, or a communication device, etc.) or a processor to execute all or part of the steps of the method in each embodiment of the present disclosure. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0144] Figure 9 It is a schematic structural diagram of the computer device provided by the embodiments of the present disclosure. As Figure 9 shown, the computer device 900 includes: a processor 901 and a memory 902.
[0145] The processor 901 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 901 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 901 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 901 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 901 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0146] The memory 902 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 902 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 902 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 901 to implement the lip shape generation method and the training method of the lip shape generation model provided in the embodiments of the present disclosure.
[0147] Those skilled in the art can understand that Figure 9 the structure shown in does not constitute a limitation on the computer device 900, and it may include more or fewer components than shown in the figure, or combine certain components, or adopt different component arrangements.
[0148] The embodiments of the present disclosure also provide a non-temporary computer-readable storage medium. When the instructions in the storage medium are executed by the processor of the computer device, the computer device can execute the lip shape generation method and the training method of the lip shape generation model provided in the embodiments of the present disclosure.
[0149] The embodiments of the present disclosure also provide a computer program product, including a computer program / instructions. When the computer program / instructions are executed by the processor, the lip shape generation method and the training method of the lip shape generation model provided in the embodiments of the present disclosure are implemented.
[0150] The above are only optional embodiments of the present disclosure, and are not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A lip shape generation method, characterized in that, The method includes: Obtain a first speech, where the first speech is the speech input into the lip shape generation model; Input the first speech and a first face image into the lip shape generation model to obtain a second face image sequence, where the second face image sequence is used to simulate the lip shape and lip texture when the human body in the first face image speaks the first speech; Among them, in the process of the lip shape generation model generating the second face image sequence, first generate the lip shape in the second face image sequence, and then generate the lip texture in the second face image sequence based on the lip shape in the second face image sequence.
2. The lip shape generation method according to claim 1, wherein The lip shape generation model realizes the generation of the second face image sequence in the following manner: Based on the first speech, obtain the lip shape parameter sequence corresponding to the first speech; Generate a third face image sequence using the lip shape parameter sequence, where the face in the third face image sequence is a three-dimensional face model, and the third face image sequence is used to indicate the change in the lip shape when the human body speaks the first speech; Based on the lip part of the third face image sequence, predict the lip texture of the first face image in the third face image sequence to obtain the second face image sequence.
3. The lip shape generation method according to claim 2, wherein The step of based on the lip part of the third face image sequence, predicting the lip texture of the first face image in the third face image sequence to obtain the second face image sequence includes: Use the lip part of the third face image sequence as the mask of the first face image to obtain a mask image sequence; Obtain a mask feature sequence, where the mask feature sequence includes multiple mask feature sets, and any one mask feature set corresponds to a mask image in the mask image sequence; Predict the lip texture of the mask part in the mask image sequence based on the mask feature sequence to obtain the second face image sequence.
4. The lip shape generation method according to claim 2 or 3, characterized in that, The step of based on the first speech, obtaining the lip shape parameter sequence corresponding to the first speech includes: Obtain the speech feature sequence of the first speech; Obtain a lip shape semantic vector set, where the lip shape semantic vector set includes multiple lip shape semantic vectors, and the lip shape semantic vectors are determined based on the benchmark expressions related to lip shape in the Blendshape expression; Use a cross-attention mechanism to process the speech feature sequence and the lip shape semantic set to obtain the lip shape parameter sequence.
5. A training method for a lip shape generation model, characterized in that, The method includes: Obtain the lip shape parameter sequence corresponding to the sample speech; Generate a third face image sequence using the lip shape parameter sequence; Train the lip shape generation model using the third face image sequence and a sample face image sequence, where the sample face image sequence is multiple frames of face images when the human body speaks the sample speech; Among them, in the process of training the lip shape generation model, use the lip part of the third face image sequence as the mask of the lip part of the first face image sequence to train the ability of the lip shape generation model to generate lip texture.
6. The method according to claim 5, characterized in that After using the lip part of the third face image sequence as the mask for the lip part of the first face image sequence, a mask image sequence corresponding to the sample speech is obtained. The lip shape generation model further includes a single-person adapter, which is connected to the encoder and the decoder. The encoder is used to obtain a mask feature sequence corresponding to the mask image sequence, and the decoder is used to predict the lip texture of the mask part in the mask image sequence based on the mask feature sequence. The method further includes: Fine-tuning the single-person adapter based on a first dataset, where the first dataset includes multiple speech-face image pair data of a single person, and each face image in the first dataset includes a lip shape.
7. A lip shape generating device, characterized in that, The device includes: An acquisition module, configured to acquire a first speech, where the first speech is the speech input to the lip shape generation model; A generation module, configured to input the first speech and a first face image into the lip shape generation model to obtain a second face image sequence, where the second face image sequence is used to simulate the lip shape and lip texture when the person in the first face image speaks the first speech; Wherein, during the process of the lip shape generation model generating the second face image sequence, the lip shape in the second face image sequence is first generated, and then the lip texture in the second face image sequence is generated based on the lip shape in the second face image sequence.
8. A training device for a lip shape generation model, characterized in that, The device includes: An acquisition module, configured to acquire a lip parameter sequence corresponding to a sample speech; A generation module, configured to generate a third face image sequence using the lip parameter sequence; A training module, configured to train the lip shape generation model using the third face image sequence and a sample face image sequence, where the sample face image sequence is multiple frames of face images when a person speaks the sample speech; Wherein, during the process of training the lip shape generation model, the lip part of the third face image sequence is used as the mask for the lip part of the first face image sequence to train the ability of the lip shape generation model to generate lip texture.
9. A computer device, characterized in that, The computer device includes: a memory and a processor. At least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to implement the method according to any one of claims 1 to 4 or any one of claims 5 to 6.
10. A computer-readable storage medium, characterized in that, At least one computer program is stored in the computer-readable storage medium, and the at least one computer program is loaded and executed by the processor to implement the method according to any one of claims 1 to 4 or any one of claims 5 to 6.
Citation Information
Patent Citations
Lip shape generation method and device, equipment and medium
CN114820891A