Method and device for generating speaking portrait video and training face rendering model
By generating a lip feature sequence corresponding to speech and combining preset expression features, the problem that speaking portrait expressions cannot be artificially controlled in the prior art is solved, and high-quality speaking portrait video generation is achieved.
Patent Information
- Application Number
- CN202210201928.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-03
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-03-03
AI Technical Summary
The prior art cannot realize artificial controllability of speaking portrait expressions, mainly focusing on the correspondence between mouth shape and audio and voice judgment expression categories, and cannot realize detailed expression control.
By first using the pre-trained lip-generating model to generate a sequence of lip-type features corresponding to the voice, then input it into the face rendering model, and by setting the preset expression features, a speaking portrait video is generated that uses voice to manipulate the target portrait.
The speech portrait video generation with artificially controlled expressions is realized, improving the authenticity of expressions and the realism of human-computer interactions.
Smart Images

Figure CN114581980B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computer technology, and in particular to methods, devices, electronic devices, and media for generating speaking portrait videos and training face rendering models. Background Art
[0002] With the development of artificial intelligence technology, the speaking portrait generation technology, which generates a speaking video of a person’s face based on a given audio by controlling a face image, has also shown a wide range of application prospects. For example, film and television industry workers can directly generate actor performance shots based on actor portraits and voices; game and other entertainment industry personnel can manipulate the facial movements of virtual characters through voice, and can achieve more realistic human-computer interaction by combining human-computer dialogue; in online conference software, this technology can restore speaking portrait video frames that are missing due to network failures based on audio.
[0003] The existing technology mainly focuses on whether the generated mouth shape corresponds to the audio, or judges the expression category through voice and adjusts the portrait expression accordingly. Therefore, it cannot achieve human controllable expression of the speaking portrait. Summary of the invention
[0004] Embodiments of the present disclosure provide methods and devices for generating a speaking person portrait video and for training a face rendering model.
[0005] In a first aspect, an embodiment of the present disclosure provides a method for generating a speaking portrait video, the method comprising: inputting a pre-acquired audio feature sequence of speech into a pre-trained lip shape generation model to generate a lip shape feature sequence corresponding to the speech; inputting the lip shape feature sequence into a face rendering model trained based on a pre-acquired target portrait to generate a speaking portrait video of the target portrait manipulated by speech, wherein the face rendering model includes a first decoder, the first decoder is used to characterize the correspondence between portrait features and the speaking portrait, the portrait features including preset expression features and lip shape features in the lip shape feature sequence.
[0006] In some embodiments, the above-mentioned lip shape generation model includes a second encoder and a second decoder; and the above-mentioned inputting the audio feature sequence of the pre-acquired speech into the pre-trained lip shape generation model to generate a lip shape feature sequence corresponding to the speech includes: inputting the audio feature sequence into the pre-trained second encoder to generate an audio feature based on the attention mechanism; inputting the lip shape features of the previous moment before the current moment, the audio features based on the attention mechanism and the posture features of the current moment into the pre-trained second decoder to generate the lip shape features of the current moment, wherein the posture features include features of other key points of the face except the mouth; based on the obtained multiple lip shape features at the current moment, generating a lip shape feature sequence corresponding to the duration of the speech.
[0007] In some embodiments, the second encoder and the second decoder are a Transformer encoder and a Transformer decoder respectively, and the Transformer decoder uses a method of mapping a posture feature sequence composed of posture features to a dimension matching the lip feature sequence as position encoding.
[0008] In some embodiments, the above-mentioned portrait features also include other features extracted from the target portrait using the first encoder corresponding to the first decoder.
[0009] In a second aspect, an embodiment of the present disclosure provides a method for training a face rendering model, the method comprising: obtaining a training sample set, wherein the training samples in the training sample set include a sample portrait; obtaining an initial face rendering model, wherein the initial face rendering model includes an initial encoder, an initial decoder, and an initial discriminator; inputting the sample portrait into the initial encoder to obtain corresponding face image features; inputting the face image features into the initial decoder to generate a generated portrait corresponding to the sample portrait; generating a loss value using a preset loss function, wherein the preset loss function includes a reconstruction loss function and an adversarial loss function, and the reconstruction loss function is used to characterize the difference between the sample portrait and the generated portrait; based on the generated loss value, adjusting the parameters of the initial face rendering model.
[0010] In some embodiments, the above-mentioned facial image features include mouth shape sub-features and expression sub-features, and the above-mentioned preset loss function also includes at least one of the following: expression loss function, mouth shape loss function; and the above-mentioned generation of loss value using the preset loss function includes: inputting the generated portrait into a pre-trained classifier to generate mouth shape generation features and expression generation features; performing at least one of the following: based on the difference between the mouth shape generation features and the mouth shape sub-features, generating a mouth shape loss value using the mouth shape loss function; based on the difference between the expression generation features and the expression sub-features, generating an expression loss value using the expression loss function; based on at least one of the generated mouth shape loss value and expression loss value, generating a total loss value.
[0011] In a third aspect, an embodiment of the present disclosure provides a device for generating a speaking portrait video, the device comprising: a lip shape generation unit, configured to input a pre-acquired audio feature sequence of speech into a pre-trained lip shape generation model, and generate a lip shape feature sequence corresponding to the speech; a video generation unit, configured to input the lip shape feature sequence into a face rendering model trained based on a pre-acquired target portrait, and generate a speaking portrait video in which the target portrait is controlled by voice, wherein the face rendering model comprises a first decoder, and the first decoder is used to characterize the correspondence between portrait features and the speaking portrait, and the portrait features comprise preset expression features and lip shape features in the lip shape feature sequence.
[0012] In some embodiments, the above-mentioned lip shape generation model includes a second encoder and a second decoder; and the above-mentioned lip shape generation unit is further configured to: input the audio feature sequence into a pre-trained second encoder to generate audio features based on the attention mechanism; input the lip shape features of the previous moment before the current moment, the audio features based on the attention mechanism and the posture features of the current moment into the pre-trained second decoder to generate the lip shape features of the current moment, wherein the posture features include features of other key points of the face except the mouth; based on the obtained multiple lip shape features of the current moment, generate a lip shape feature sequence corresponding to the duration of the speech.
[0013] In some embodiments, the second encoder and the second decoder are a Transformer encoder and a Transformer decoder respectively, and the Transformer decoder uses a method of mapping a posture feature sequence composed of posture features to a dimension matching the lip feature sequence as position encoding.
[0014] In some embodiments, the above-mentioned portrait features also include other features extracted from the target portrait using the first encoder corresponding to the first decoder.
[0015] In a fourth aspect, an embodiment of the present disclosure provides a device for training a face rendering model, the device comprising: a first acquisition unit, configured to acquire a training sample set, wherein the training samples in the training sample set include a sample portrait; a second acquisition unit, configured to acquire an initial face rendering model, wherein the initial face rendering model includes an initial encoder, an initial decoder and an initial discriminator; a feature generation unit, configured to input the sample portrait into the initial encoder to obtain corresponding face image features; a portrait generation unit, configured to input the face image features into the initial decoder to generate a generated portrait corresponding to the sample portrait; a loss determination unit, configured to generate a loss value using a preset loss function, wherein the preset loss function includes a reconstruction loss function and an adversarial loss function, and the reconstruction loss function is used to characterize the difference between the sample portrait and the generated portrait; an adjustment unit, configured to adjust the parameters of the initial face rendering model based on the generated loss value.
[0016] In some embodiments, the above-mentioned facial image features include mouth shape sub-features and expression sub-features, and the preset loss function also includes at least one of the following: expression loss function, mouth shape loss function; and the above-mentioned loss determination unit is further configured to: input the generated portrait into a pre-trained classifier to generate mouth shape generation features and expression generation features; perform at least one of the following: based on the difference between the mouth shape generation features and the mouth shape sub-features, generate a mouth shape loss value using the mouth shape loss function; based on the difference between the expression generation features and the expression sub-features, generate an expression loss value using the expression loss function; based on at least one of the generated mouth shape loss value and expression loss value, generate a total loss value.
[0017] In a fifth aspect, an embodiment of the present disclosure provides an electronic device, comprising: one or more processors; a storage device on which one or more programs are stored; when the one or more programs are executed by one or more processors, the one or more processors implement the method described in any implementation manner in the first aspect.
[0018] In a sixth aspect, an embodiment of the present disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the method described in any implementation manner in the first aspect.
[0019] The embodiments of the present disclosure provide a method, an apparatus, an electronic device, and a medium for generating a speaking portrait video and for training a face rendering model. The method first uses a pre-trained lip shape generation model to generate a lip shape feature sequence corresponding to the voice, then inputs the generated lip shape feature sequence into the face rendering model, and generates a speaking portrait video that uses voice to control the target portrait through the setting of preset expression features, thereby achieving the generation of a speaking portrait video with artificially controlled expressions. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Other features, objects and advantages of the present disclosure will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:
[0021] Figure 1 is an exemplary system architecture diagram in which an embodiment of the present disclosure may be applied;
[0022] Figure 2 is a flow chart of an embodiment of a method for generating a speaking person portrait video according to the present disclosure;
[0023] Figure 3 is a schematic diagram of an application scenario of a method for generating a speaking person portrait video according to an embodiment of the present disclosure;
[0024] Figure 4is a flowchart of another embodiment of a method for training a face rendering model according to the present disclosure;
[0025] Figure 5 is a structural schematic diagram of an embodiment of a device for generating a speaking person portrait video according to the present disclosure;
[0026] Figure 6 is a structural schematic diagram of an embodiment of an apparatus for training a face rendering model according to the present disclosure;
[0027] Figure 7 It is a schematic diagram of the structure of an electronic device suitable for implementing the embodiments of the present disclosure. DETAILED DESCRIPTION
[0028] The present disclosure is further described in detail below in conjunction with the accompanying drawings and embodiments. It is understood that the specific embodiments described herein are only used to explain the relevant invention, rather than to limit the invention. It is also necessary to explain that, for ease of description, only the parts related to the relevant invention are shown in the accompanying drawings.
[0029] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other. The present disclosure will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0030] Figure 1 An exemplary architecture 100 is shown to which the method for generating a speaking person portrait video and training a face rendering model or the apparatus for generating a speaking person portrait video and training a face rendering model of the present disclosure can be applied.
[0031] like Figure 1 As shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104 and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links or optical fiber cables, etc.
[0032] The terminal devices 101, 102, 103 interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, video editing applications, machine learning model training applications, etc.
[0033] Terminal devices 101, 102, 103 can be hardware or software. When terminal devices 101, 102, 103 are hardware, they can be various electronic devices with display screens and supporting human-computer interaction, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, etc. When terminal devices 101, 102, 103 are software, they can be installed in the electronic devices listed above. It can be implemented as multiple software or software modules (for example, software or software modules used to provide distributed services), or it can be implemented as a single software or software module. No specific limitation is made here.
[0034] The server 105 may be a server that provides various services, such as a background server that provides support for video editing applications on the terminal devices 101, 102, and 103. The background server may train the lip generation model and the face rendering model, and provide the trained lip generation model and the face rendering model to the terminals 101, 102, and 103, so that the terminals 101, 102, and 103 generate a speaking person video using the trained lip generation model and the face rendering model.
[0035] It should be noted that, optionally, the terminal devices 101, 102, and 103 can also send the acquired audio feature sequence and preset expression features of the speech to the background server, so that the background server can generate a speaking portrait video using the trained lip generation model and face rendering model.
[0036] It should be noted that the server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or it can be implemented as a single server. When the server is software, it can be implemented as multiple software or software modules (for example, software or software modules used to provide distributed services), or it can be implemented as a single software or software module. No specific limitation is made here.
[0037] It should be noted that the method for generating a speaking person portrait video provided in the embodiment of the present disclosure is generally performed by the terminal devices 101, 102, and 103, and accordingly, the apparatus for generating a speaking person portrait video can also be set in the terminal devices 101, 102, and 103. The method for training a face rendering model provided in the embodiment of the present disclosure is generally performed by the server 105, and accordingly, the apparatus for training a face rendering model is generally set in the server 105.
[0038] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.
[0039] Continue to refer Figure 2, shows a process 200 of an embodiment of a method for generating a speaking person portrait video according to the present disclosure. The method for generating a speaking person portrait video comprises the following steps:
[0040] Step 201: input the pre-acquired audio feature sequence of speech into the pre-trained lip shape generation model to generate a lip shape feature sequence corresponding to the speech.
[0041] In this embodiment, the execution subject of the method for generating a speaking person portrait video (such as Figure 1 The terminal devices 101, 102, 103 or the server 105 shown in the figure can first obtain the audio feature sequence of the speech through a wired connection or a wireless connection. The above audio feature sequence may include a sequence composed of various features that can reflect the audio characteristics. For example, the above audio feature may be MFCC (Mel-Frequency Cepstral Coefficients).
[0042] Therefore, the above-mentioned execution subject can directly obtain the audio feature sequence of the speech, or can first obtain the speech and then generate the audio feature sequence of the obtained speech using various audio feature extraction methods.
[0043] Afterwards, the above-mentioned execution subject can input the acquired audio feature sequence of the speech into the pre-trained lip shape generation model to generate a lip shape feature sequence corresponding to the above-mentioned speech. Among them, the lip shape features in the above-mentioned lip shape feature sequence can be used to characterize the position information of the key points of the mouth in the face image. Thus, the lip shape features in the above-mentioned lip shape feature sequence can be used to characterize the changes in the position information of the key points of the mouth that are continuous in time. As an example, the above-mentioned lip shape generation model can include various pre-trained RNNs (Recurrent Neural Networks), such as GRU, LSTM (Long Short-Term Memory, Long Short-Term Memory Network), etc. Optionally, when the frame rate of the output audio (e.g., 100FPS) and the frame rate of the lip shape feature (e.g., 30FPS) do not belong to a corresponding integer ratio relationship, the lip shape feature can be expanded using an interpolation method so that the frame rate of the lip shape feature is in an integer ratio with the frame rate of the audio.
[0044] In some optional implementations of this embodiment, the above-mentioned lip shape generation model may include a second encoder and a second decoder. The above-mentioned execution subject may generate a lip shape feature sequence corresponding to the above-mentioned speech according to the following steps:
[0045] In the first step, the audio feature sequence is input into the pre-trained second encoder to generate audio features based on the attention mechanism.
[0046] In these implementations, the execution subject may input the audio feature sequence pre-acquired in step 201 into a pre-trained second encoder to generate audio features based on the attention mechanism. The second encoder may be any model trained based on the attention mechanism that can achieve feature dimensionality reduction.
[0047] In the second step, the lip shape features of the previous moment before the current moment, the audio features based on the attention mechanism and the posture features of the current moment are input into the pre-trained second decoder to generate the lip shape features of the current moment.
[0048] In these implementations, the execution entity may input the lip shape features of the previous moment (e.g., moment T-1) of the current moment, the audio features based on the attention mechanism generated in the first step, and the posture features of the current moment (e.g., moment T) into a pre-trained second decoder to generate the lip shape features of the current moment. The posture features may include features of other key points of the face except the mouth. The posture features may usually be acquired in advance. Since there is a relative positional relationship between the key points of a person's face, adding posture features can help enhance the authenticity of the expression of the generated portrait.
[0049] The third step is to generate a lip shape feature sequence corresponding to the duration of the speech based on the obtained multiple lip shape features at the current moment.
[0050] In these implementations, the execution entity may execute the second step in a loop to generate, for example, lip shape features corresponding to time 0 to the end time of the audio, thereby forming a lip shape feature sequence corresponding to the duration of the speech.
[0051] Based on the above optional implementation methods, this solution can utilize a lip shape generation model composed of a second encoder and a second decoder, combined with an attention mechanism and posture features to generate a lip shape feature sequence corresponding to the duration of speech, enriching the generation method of the lip shape feature sequence, and helping to generate a speaking person with natural expressions and more consistent lip shape and speech.
[0052] Optionally, based on the above optional implementation, the second encoder and the second decoder may be a Transformer encoder and a Transformer decoder, respectively. The Transformer decoder may use a method of mapping the posture feature sequence composed of the posture features to a dimension matching the lip shape feature sequence as position encoding.
[0053] In these implementations, the Transformer encoder is responsible for processing the global features of the audio (such as MFCC features) and passing the audio features to the Transformer decoder in the form of cross attention. The Transformer decoder is responsible for combining the audio features with the mouth shape features of the previous timestamp and outputting the mouth shape features of the current frame.
[0054] Since the Transformer mechanism has the flexibility of self-attention and cross-attention, it can perform time-unsynchronized audio-to-lip shape prediction without any interpolation method. That is, even when there is no integer ratio correspondence between the audio frame rate and the lip shape frame rate (that is, the video frame rate), the corresponding generation of lip shape feature sequences can be achieved.
[0055] In these implementations, as an example, the pre-acquired audio feature sequence of speech can be used, for example, Indicates. Among them, the above L s and D s They can be used to characterize the length of the audio feature sequence and the dimension of the audio feature, respectively. In the Transformer encoder, the above-mentioned execution subject can map it to Query, Key, and Value. In the multi-head attention module, the above-mentioned execution subject can group the Query, Key, and Value vectors, and then perform self-attention operations. Afterwards, for the obtained lip features (for example, the lip features corresponding to the timestamp before the current moment), the sequence is mapped to Query through the Transformer decoder. Then, the Value and Key are output through the last layer of the Transformer encoder. Similarly, in the multi-head attention module, the above-mentioned execution subject can group the Query, Key, and Value vectors, and then perform cross-attention operations. Then, continue to enter the next layer of Transformer decoder operations.
[0056] It should be noted that in natural speech processing and speech recognition tasks, the Transformer-based model needs to specifically define the start word (Start Token), the mask word (Mask Token) and the end word (End Token). Based on this, the start word of the above Transformer encoder and Transformer decoder is defined as an all-zero vector, and the mask word is defined as an all-one vector. At the same time, since the timestamp correspondence between the audio and the mouth shape is known, the prediction can be stopped at the mouth shape timestamp corresponding to the audio stop, so there is no need to define the end word. In addition, the classic Transformer uses sin / cos encoding as the position encoding of the input of the Transformer encoder and the Transformer decoder to represent the distance relationship in the time series. In these implementations, the above Transformer encoder can still use position encoding. In the Transformer decoder, the above execution subject can use the method of mapping the posture feature sequence composed of the above posture features to the dimension matching the above mouth shape feature sequence as the position encoding. As an example, the above execution subject can multiply the posture sequence with a preset matrix and then add it to the above mouth shape feature vector as the input of the Transformer encoder. The preset matrix may be used to characterize the transformation relationship between the dimensions of the posture feature sequence composed of the posture features and the lip shape feature sequence.
[0057] Based on the above optional implementation methods, this solution can innovatively apply the Transformer traditionally used in speech recognition tasks to the generation of lip feature sequences, and further improve the consistency of the generated lip features and audio through the improvement of position encoding.
[0058] Step 202: input the lip shape feature sequence into a face rendering model trained based on a pre-acquired target portrait, and generate a speaking portrait video of the target portrait controlled by voice.
[0059] In this embodiment, the execution subject may input the lip shape feature sequence generated in step 201 into a face rendering model trained based on a pre-acquired target portrait, and generate a speaking portrait video that uses voice to control the target portrait. The face rendering model may include a first decoder. The first decoder may be used to characterize the correspondence between portrait features and speaking portraits. The portrait features may include preset expression features and lip shape features in the lip shape feature sequence.
[0060] In this embodiment, the face rendering model may include a decoder part in a generator in a generative adversarial network trained using a target portrait as the first decoder. The generator in the generative adversarial network may include a first encoder and a first decoder. The first encoder is used to extract features of the target portrait. The first decoder is used to restore the portrait based on the features extracted by the first encoder, so that the restored portrait is as consistent as possible with the target portrait.
[0061] In this embodiment, the preset expression feature may be, for example, a one-hot encoding, such as "001" for "happy", "010" for "calm", and "100" for "angry". The execution subject may input the lip feature sequence and the preset expression feature generated in step 201 into the first decoder in the face rendering model, thereby generating a speaking portrait video that uses the voice to control the target portrait.
[0062] In some optional implementations of this embodiment, the above-mentioned portrait features also include other features extracted from the above-mentioned target portrait using the first encoder corresponding to the above-mentioned first decoder.
[0063] In these implementations, the execution subject may also decouple the features of the target portrait extracted by the first encoder in various ways, thereby forming other features in addition to the preset expression features and lip shape features.
[0064] Based on the above optional implementation methods, this solution can decouple the features, thereby facilitating human control in the expression dimension.
[0065] In some optional implementations of this embodiment, the execution subject may further use the speaking person portrait video generated in step 202 as an expansion sample to train the lip reading recognition model.
[0066] Based on the above optional implementation methods, this solution can use the method of generating speaking portrait videos to expand the data samples required for training the lip reading recognition model, so as to improve the effect of the lip reading recognition model.
[0067] In some optional implementations of this embodiment, the above-mentioned face rendering model can be based on the following Figure 4 The described method for training a face rendering model is trained.
[0068] Continue to see Figure 3 , Figure 3 FIG. 1 is a schematic diagram of an application scenario of a method for generating a speaking person portrait video according to an embodiment of the present disclosure. Figure 3In the application scenario, user 301 can record a voice through terminal 302. Terminal 302 can extract the audio feature sequence 303 of the above voice and input it into the pre-trained lip generation model 304 to generate a corresponding lip feature sequence 305. Afterwards, terminal 302 can input the lip feature sequence 305 into a face rendering model 306 trained based on a pre-acquired target portrait (such as a cartoon character) to generate a speaking portrait video 307. Among them, the above-mentioned preset expression features 308 can be specified by user 301 or by default, which is not limited here. Optionally, the above-mentioned lip generation model 304 and face rendering model 306 can be trained by a server 309 that is communicatively connected to terminal 302.
[0069] At present, one of the existing technologies is usually to focus on whether the generated mouth shape corresponds to the audio, or to judge the expression category through voice and adjust the portrait expression accordingly, which results in the inability to realize the artificial control of the expression of the speaking portrait. The method provided by the above-mentioned embodiment of the present disclosure first generates a mouth shape feature sequence corresponding to the voice using a pre-trained mouth shape generation model, then inputs the generated mouth shape feature sequence into the face rendering model, and generates a speaking portrait video that uses voice to control the target portrait through the setting of preset expression features, thereby realizing the generation of a speaking portrait video with artificially controlled expression.
[0070] Further references Figure 4 , which shows a process 400 of another embodiment of a method for training a face rendering model. The process 400 of the method for training a face rendering model comprises the following steps:
[0071] Step 401: Obtain a training sample set.
[0072] In this embodiment, the execution body (eg, Figure 1 The server 105 shown in the figure can obtain the training sample set from the local or communication-connected electronic device through a wired or wireless connection. The training samples in the training sample set may include sample portraits.
[0073] Step 402: Obtain an initial face rendering model.
[0074] In this embodiment, the execution subject may obtain the initial face rendering model from a local or communication-connected electronic device through a wired or wireless connection. The initial face rendering model may include an initial encoder, an initial decoder, and an initial discriminator. The initial encoder is used to achieve feature dimensionality reduction, and the initial decoder is used to achieve feature dimensionality increase. The initial discriminator is used to determine whether the output of the initial decoder is "true" or "false" relative to the sample portrait.
[0075] Step 403: input the sample portrait into the initial encoder to obtain the corresponding facial image features.
[0076] In this embodiment, the execution subject may input the sample portrait in the training sample set obtained in step 401 into the initial encoder obtained in step 402, so as to obtain the corresponding facial image features.
[0077] In some optional implementations of this embodiment, the execution subject may replace the mouth shape vector with the corresponding mouth shape of another data point with a certain probability during training; similarly, the expression vector may be replaced with the corresponding expression of another data point with a certain probability. Thus, the mouth shape feature, expression feature and other features may be decoupled from the portrait renderer.
[0078] Step 404: input the facial image features into the initial decoder to generate a generated portrait corresponding to the sample portrait.
[0079] In this embodiment, the execution subject may input the facial image features generated in step 403 into the initial decoder obtained in step 402 to generate a generated portrait corresponding to the sample portrait.
[0080] Step 405: Generate a loss value using a preset loss function.
[0081] In this embodiment, the execution subject may generate a loss value in various ways using a preset loss function. The preset loss function may include a reconstruction loss function and an adversarial loss function. The reconstruction loss function may be used to characterize the difference between the sample portrait and the generated portrait.
[0082] As an example, the above adversarial loss function can be:
[0083]
[0084] The above De and D can be used to represent the above decoder and discriminator respectively. The above i and i' can be used to represent the sample portrait and the generated portrait respectively.
[0085] In some optional implementations of this embodiment, the facial image features may include mouth shape sub-features and expression sub-features. The preset loss function may also include at least one of the following: expression loss function, mouth shape loss function. The execution subject may generate a loss value using the preset loss function according to the following steps:
[0086] In the first step, the generated portrait is input into a pre-trained classifier to generate lip shape generation features and expression generation features.
[0087] In these implementations, the classifier can be used to characterize the correspondence between the lip shape generation feature and the generated portrait, and the correspondence between the expression generation feature and the generated portrait. Optionally, the classifier can also generate the probabilities of the obtained lip shape generation feature and expression generation feature corresponding to each other.
[0088] In the second step, at least one of the following is performed: based on the difference between the lip shape generation feature and the lip shape sub-feature, a lip shape loss value is generated using a lip shape loss function; based on the difference between the expression generation feature and the expression sub-feature, an expression loss value is generated using an expression loss function;
[0089] In these implementations, as an example, the lip shape loss function may be a least squares loss; and the expression loss function may be a cross entropy loss.
[0090] The third step is to generate a total loss value based on at least one of the generated lip loss value and expression loss value.
[0091] In these implementations, based on at least one of the generated lip loss value and expression loss value, the execution subject may generate a total loss value in various ways. As an example, the execution subject may perform a weighted summation of at least one of the lip loss value and expression loss value generated in the second step and the loss values corresponding to the reconstruction loss function and the adversarial loss function in step 405, respectively, to generate a total loss value.
[0092] Based on the above optional implementation methods, this solution can supervise the generated images from two dimensions, expression and lip shape, on the basis of making the overall details of the portrait as realistic as possible, thereby further improving the effect of the face rendering model.
[0093] Step 406: Adjust the parameters of the initial face rendering model based on the generated loss value.
[0094] In this embodiment, based on the generated loss value, the execution subject may adjust the parameters of the initial face rendering model obtained in step 402 using a machine learning method.
[0095] In some optional implementations of this embodiment, the execution subject may determine whether the training end condition is met, and when the training end condition is met, the training is terminated and the initial face rendering model after parameter adjustment is determined as the trained face rendering model. When the training end condition is not met, the execution subject may re-determine the initial face rendering model after parameter adjustment as the initial face rendering model in step 403, and continue to perform steps 403 to 406 to iteratively perform training.
[0096] from Figure 4As can be seen, the process 400 of the method for training a face rendering model in this embodiment embodies the steps of using a sample portrait to train an initial face rendering model including an initial encoder, an initial decoder, and an initial discriminator. Therefore, the solution described in this embodiment enriches the generation method of the face rendering model.
[0097] Further references Figure 5 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a device for generating a speaking person portrait video. Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.
[0098] like Figure 5 As shown, the apparatus 500 for generating a speaking person portrait video provided by this embodiment includes a lip shape generation unit 501 and a video generation unit 502. The lip shape generation unit 501 is configured to input a pre-acquired audio feature sequence of a speech into a pre-trained lip shape generation model to generate a lip shape feature sequence corresponding to the speech; the video generation unit 502 is configured to input the lip shape feature sequence into a face rendering model trained based on a pre-acquired target portrait to generate a speaking person portrait video that uses speech to control the target portrait, wherein the face rendering model includes a first decoder, and the first decoder is used to characterize the correspondence between portrait features and the speaking person portrait, and the portrait features include preset expression features and lip shape features in the lip shape feature sequence.
[0099] In the apparatus 500 for generating a speaking person portrait video, the specific processing of the lip shape generation unit 501 and the video generation unit 502 and the technical effects thereof can be referred to in Figure 2 The relevant descriptions of step 201 and step 202 in the corresponding embodiment are not repeated here.
[0100] In some optional implementations of this embodiment, the above-mentioned lip shape generation model may include a second encoder and a second decoder. The above-mentioned lip shape generation unit 501 may be further configured to: input the audio feature sequence into a pre-trained second encoder to generate audio features based on the attention mechanism; input the lip shape features of the previous moment of the current moment, the audio features based on the attention mechanism and the posture features of the current moment into the pre-trained second decoder to generate the lip shape features of the current moment, wherein the posture features may include features of other key points of the face except the mouth; based on the obtained multiple lip shape features of the current moment, generate a lip shape feature sequence corresponding to the duration of the speech.
[0101] In some optional implementations of this embodiment, the second encoder and the second decoder may be a Transformer encoder and a Transformer decoder, respectively. The Transformer decoder may use a method of mapping a posture feature sequence composed of posture features to a dimension matching the lip feature sequence as position encoding.
[0102] In some optional implementations of this embodiment, the above-mentioned portrait features may also include other features extracted from the target portrait by using the first encoder corresponding to the first decoder.
[0103] The device provided by the above-mentioned embodiment of the present disclosure first generates a lip shape feature sequence corresponding to the voice using a pre-trained lip shape generation model through the lip shape generation unit 501, and then inputs the generated lip shape feature sequence into the face rendering model through the video generation unit 502, and generates a speaking portrait video of the target portrait controlled by voice through the setting of preset expression features, thereby realizing the generation of a speaking portrait video with artificially controlled expressions.
[0104] Further references Figure 6 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a device for training a face rendering model. Figure 4 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.
[0105] like Figure 6 As shown, the apparatus 600 for training a face rendering model provided by the present embodiment includes a first acquisition unit 601, configured to acquire a training sample set, wherein the training samples in the training sample set include a sample portrait; a second acquisition unit 602, configured to acquire an initial face rendering model, wherein the initial face rendering model includes an initial encoder, an initial decoder, and an initial discriminator; a feature generation unit 603, configured to input the sample portrait into the initial encoder to obtain a corresponding face image feature; a portrait generation unit 604, configured to input the face image feature into the initial decoder to generate a generated portrait corresponding to the sample portrait; a loss determination unit 605, configured to generate a loss value using a preset loss function, wherein the preset loss function includes a reconstruction loss function and an adversarial loss function, and the reconstruction loss function is used to characterize the difference between the sample portrait and the generated portrait; an adjustment unit 606, configured to adjust the parameters of the initial face rendering model based on the generated loss value.
[0106] In the apparatus 600 for training a face rendering model in this embodiment, the specific processing of the first acquisition unit 601, the second acquisition unit 602, the feature generation unit 603, the portrait generation unit 604, the loss determination unit 605 and the adjustment unit 606 and the technical effects thereof can be referred to respectively. Figure 4 The relevant descriptions of steps 401 to 406 in the corresponding embodiment are not repeated here.
[0107] In some optional implementations of this embodiment, the above-mentioned facial image features may include lip shape sub-features and expression sub-features. The preset loss function may also include at least one of the following: expression loss function, lip shape loss function. The above-mentioned loss determination unit 605 may be further configured to: input the generated portrait into a pre-trained classifier to generate lip shape generation features and expression generation features; perform at least one of the following: based on the difference between the lip shape generation features and the lip shape sub-features, generate a lip shape loss value using a lip shape loss function; based on the difference between the expression generation features and the expression sub-features, generate an expression loss value using an expression loss function; based on at least one of the generated lip shape loss value and expression loss value, generate a total loss value.
[0108] The apparatus provided by the above embodiment of the present disclosure realizes the use of sample portraits to train an initial face rendering model including an initial encoder, an initial decoder and an initial discriminator to generate a face rendering model through the first acquisition unit 601, the second acquisition unit 602, the feature generation unit 603, the portrait generation unit 604, the loss determination unit 605 and the adjustment unit 606. This enriches the generation method of the face rendering model.
[0109] Reference below Figure 7 , which shows an electronic device (eg, Figure 1 Schematic diagram of the structure of server 105)700. Figure 7 The server shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0110] like Figure 7 As shown, the electronic device 700 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the electronic device 700 are also stored. The processing device 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0111] Typically, the following devices may be connected to the I / O interface 705: an input device 706 including, for example, a touch screen, a touch pad, a keyboard, a mouse, etc.; an output device 707 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 708 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 709. The communication device 709 may allow the electronic device 700 to communicate with other devices wirelessly or by wire to exchange data. Figure 7 The electronic device 700 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead. Figure 7 Each block shown in the figure may represent one device, or may represent multiple devices as required.
[0112] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 709, or installed from the storage device 708, or installed from the ROM 702. When the computer program is executed by the processing device 701, the above-mentioned functions defined in the method of the embodiment of the present application are executed.
[0113] It should be noted that the computer-readable medium described in the embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, device or device. In the embodiments of the present disclosure, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.
[0114] The above-mentioned computer-readable medium may be included in the above-mentioned electronic device; or it may exist independently without being assembled into the electronic device. The above-mentioned computer-readable medium carries one or more programs. When the above-mentioned one or more programs are executed by the electronic device, the electronic device: inputs the pre-acquired audio feature sequence of the speech into a pre-trained lip shape generation model to generate a lip shape feature sequence corresponding to the speech; inputs the lip shape feature sequence into a face rendering model trained based on a pre-acquired target portrait to generate a speaking portrait video of the target portrait controlled by voice, wherein the face rendering model includes a first decoder, and the first decoder is used to characterize the correspondence between portrait features and speaking portraits, and the portrait features include preset expression features and lip shape features in the lip shape feature sequence; or, Electronic device: obtaining a training sample set, wherein the training samples in the training sample set include a sample portrait; obtaining an initial face rendering model, wherein the initial face rendering model includes an initial encoder, an initial decoder and an initial discriminator; inputting the sample portrait into the initial encoder to obtain corresponding face image features; inputting the face image features into the initial decoder to generate a generated portrait corresponding to the sample portrait; generating a loss value using a preset loss function, wherein the preset loss function includes a reconstruction loss function and an adversarial loss function, and the reconstruction loss function is used to characterize the difference between the sample portrait and the generated portrait; adjusting the parameters of the initial face rendering model based on the generated loss value.
[0115] Computer program code for performing the operations of the embodiments of the present disclosure may be written in one or more programming languages or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C", Python, or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0116] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0117] The units involved in the embodiments described in the present disclosure may be implemented by software or by hardware. The described units may also be arranged in a processor, for example, may be described as: a processor including a lip shape generation unit and a video generation unit; or may be described as: a processor including a first acquisition unit, a second acquisition unit, a feature generation unit, a portrait generation unit, a loss determination unit, and an adjustment unit. The names of these units do not constitute a limitation on the units themselves in certain cases. For example, the lip shape generation unit may also be described as "a unit that inputs a pre-acquired audio feature sequence of speech into a pre-trained lip shape generation model to generate a lip shape feature sequence corresponding to the speech".
[0118] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with the technical features with similar functions disclosed in the embodiments of the present disclosure (but not limited to) to form a technical solution.
Claims
1. A method for generating a speaking person portrait video, include: Inputting the pre-acquired audio feature sequence of speech into a pre-trained lip shape generation model to generate a lip shape feature sequence corresponding to the speech; Inputting the lip shape feature sequence into a face rendering model trained based on a pre-acquired target portrait, generating a speaking portrait video of the target portrait controlled by the voice, wherein the face rendering model includes a first decoder, the first decoder is used to characterize the correspondence between portrait features and the speaking portrait, the portrait features including preset expression features and the lip shape features in the lip shape feature sequence; Among them, the lip shape generation model includes a second decoder, which is a Transformer decoder. The Transformer decoder uses a method of mapping a posture feature sequence composed of posture features to a dimension matching the lip shape feature sequence as position encoding, and the posture features include features of other key points of the face except the mouth.
2. The method according to claim 1, in, The lip shape generation model includes a second encoder; and The step of inputting the pre-acquired audio feature sequence of speech into a pre-trained lip shape generation model to generate a lip shape feature sequence corresponding to the speech comprises: Inputting the audio feature sequence into the pre-trained second encoder to generate audio features based on the attention mechanism; Inputting the lip shape features at the previous moment of the current moment, the audio features based on the attention mechanism, and the posture features at the current moment into the pre-trained second decoder to generate the lip shape features at the current moment; Based on the obtained multiple lip shape features at the current moment, a lip shape feature sequence corresponding to the duration of the speech is generated.
3. The method according to claim 2, in, The second encoder is a Transformer encoder.
4. The method according to any one of claims 1 to 3, in, The portrait features also include other features extracted from the target portrait using a first encoder corresponding to the first decoder.
5. The method according to claim 1, in, The face rendering model is trained by the following steps: Acquire a training sample set, wherein the training samples in the training sample set include sample portraits; Acquire an initial face rendering model, wherein the initial face rendering model includes an initial encoder, an initial decoder and an initial discriminator; Inputting the sample portrait into the initial encoder to obtain corresponding facial image features; Inputting the facial image features into the initial decoder to generate a generated portrait corresponding to the sample portrait; Generate a loss value using a preset loss function, wherein the preset loss function includes a reconstruction loss function and an adversarial loss function, and the reconstruction loss function is used to characterize the difference between the sample portrait and the generated portrait; Based on the generated loss value, parameters of the initial face rendering model are adjusted.
6. The method according to claim 5, in, The facial image features include mouth shape sub-features and expression sub-features, and the preset loss function also includes at least one of the following: expression loss function, mouth shape loss function; as well as The generating the loss value by using the preset loss function includes: Inputting the generated portrait into a pre-trained classifier to generate lip shape generation features and expression generation features; Perform at least one of the following: based on the difference between the lip shape generation feature and the lip shape sub-feature, generate a lip shape loss value using the lip shape loss function; based on the difference between the expression generation feature and the expression sub-feature, generate an expression loss value using the expression loss function; A total loss value is generated based on at least one of the generated lip loss value and expression loss value.
7. A device for generating a video of a speaking person, include: A lip shape generation unit is configured to input a pre-acquired audio feature sequence of speech into a pre-trained lip shape generation model to generate a lip shape feature sequence corresponding to the speech; A video generating unit is configured to input the lip shape feature sequence into a face rendering model trained based on a pre-acquired target portrait, and generate a speaking portrait video of the target portrait being controlled by the voice, wherein the face rendering model includes a first decoder, and the first decoder is used to characterize the correspondence between portrait features and the speaking portrait, wherein the portrait features include preset expression features and the lip shape features in the lip shape feature sequence; Among them, the lip shape generation model includes a second decoder, which is a Transformer decoder. The Transformer decoder uses a method of mapping a posture feature sequence composed of posture features to a dimension matching the lip shape feature sequence as position encoding, and the posture features include features of other key points of the face except the mouth.
8. An electronic device, include: one or more processors; a storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.
9. A computer readable medium having a computer program stored thereon, in, When the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Portrait-based video generation method and device, and storage medium
CN111383307A
Digital human generation method and device, equipment and medium
CN113886642A