Facial video generation method, apparatus and electronic device

The method enhances face video generation by using a trained face-mouth-type driving model to extract features and generate styled face videos that accurately represent personalized mouth shape styles, addressing the limitations of general-purpose models.

JP7763298B2Active Publication Date: 2025-10-31BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024099048
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2024-02-01
Filing Date
2024-06-19
Publication Date
2025-10-31
Estimated Expiration
2044-06-19

AI Technical Summary

Technical Problem

Existing face and mouth shape driving models generate face videos in a general style, failing to reflect personalized mouth shape styles of different target objects and resulting in low accuracy.

Method used

A method involving obtaining a mouth-shape multimedia resource and a reference face image, extracting features, generating a styled face image based on driving features and a reference style vector, and determining a stylized face video to reflect personalized mouth shape styles, using a trained face-mouth-type driving model with a feature extraction and face driving network.

Benefits of technology

Ensures that generated face videos accurately reflect personalized mouth shape styles, improving the accuracy of the generated face videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007763298000001
    Figure 0007763298000001
  • Figure 0007763298000002
    Figure 0007763298000002
  • Figure 0007763298000003
    Figure 0007763298000003
Patent Text Reader

Abstract

To provide a face video generation method, apparatus, and electronic device, which can reflect a personalized mouth style of a target object, ensure that a generated style face video can reflect the personalized mouth style of the target object, and further improve the accuracy of the generated style face video.SOLUTION: The method includes: acquiring a mouth shape multimedia resource and a reference face image of a target object; acquiring a reference style vector of the target object; performing feature extraction processing on each resource frame in the mouth shape multimedia resource, to acquire mouth shape driving features; on the basis of the mouth shape driving features, the reference face image, and the reference style vector, generating a style face image corresponding to the resource frame; and determining a style face video of the target object.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to the field of artificial intelligence technology, particularly to technical fields such as deep learning, big data, computer vision, and voice technology, and in particular to a facial video generation method, apparatus, and electronic device. [Background technology]

[0002] The current face and mouth type driving solution mainly involves obtaining a face and mouth type driving model, obtaining a facial image and audio or video of a target object, inputting the audio or video and the facial image of the target object into the face and mouth type driving model, and obtaining a facial video of the target object output from the face and mouth type driving model.

[0003] In the above technical proposal, the face and mouth shape driving model is a general-purpose face and mouth shape driving model, and the output face video is a face video of the target object in a general style, which makes it difficult to reflect the personalized mouth shape styles of different target objects, and the accuracy of the generated face video is low. Summary of the Invention [Problem to be solved by the invention]

[0004] The present disclosure provides a facial video generation method, apparatus, and electronic device.

[0005] According to one aspect of the present disclosure, there is provided a face video generation method, the method including: obtaining a mouth-shape multimedia resource and a reference face image of a target object; obtaining a reference style vector of the target object; for each resource frame in the mouth-shape multimedia resource, performing a feature extraction process on the resource frame to obtain mouth-shape driving features; generating a stylized face image corresponding to the resource frame based on the mouth-shape driving features, the reference face image, and the reference style vector; and determining a stylized face video of the target object based on the stylized face image corresponding to each resource frame in the mouth-shape multimedia resource.

[0006] According to another aspect of the present disclosure, there is provided a method for training a face-mouth-type driving model, the method including the steps of: obtaining a pre-trained face-mouth-type driving model and an encoding network, wherein the face-mouth-type driving model includes a feature extraction network and a face driving network connected in sequence; obtaining sample mouth-type driving features of each sample resource frame in a sample mouth-type multimedia resource, a sample reference face image, and a sample style face video, wherein the sample resource frames in the sample mouth-type multimedia resource correspond one-to-one to the sample video frames in the sample style face video; and for each sample resource frame in the sample mouth-type multimedia resource, inputting a sample video frame and a sample mouth shape driving feature corresponding to the frame into an initial encoding network to obtain a predicted style vector output from the encoding network; inputting the predicted style vector, the sample mouth shape driving feature, and the sample reference face image into the face driving network to obtain a predicted style face image output from the face driving network; and performing a parameter adjustment process on the encoding network and the face driving network in the face mouth shape driving model based on a distribution to which the predicted style vector belongs, a Gaussian distribution, the predicted style face image, and a sample video frame corresponding to the sample resource frame, to obtain a trained face mouth shape driving model.

[0007] According to another aspect of the present disclosure, there is provided a face video generation apparatus, the apparatus including: a first acquisition module for acquiring a mouth-shape multimedia resource and a reference face image of a target object; a second acquisition module for acquiring a reference style vector of the target object; a feature extraction module for, for each resource frame in the mouth-shape multimedia resource, performing a feature extraction process on the resource frame to obtain mouth-shape driving features; a generation module for generating a stylized face image corresponding to the resource frame based on the mouth-shape driving features, the reference face image, and the reference style vector; and a determination module for determining a stylized face video of the target object based on the stylized face image corresponding to each resource frame in the mouth-shape multimedia resource.

[0008] According to another aspect of the present disclosure, there is provided an apparatus for training a face-mouth-type driving model, the apparatus including: a first acquisition module for acquiring a pre-trained face-mouth-type driving model and an encoding network, the face-mouth-type driving model including a feature extraction network and a face driving network connected in sequence; a second acquisition module for acquiring sample mouth-type driving features, a sample reference face image, and a sample styled face video for each sample resource frame in a sample mouth-type multimedia resource, the sample resource frames in the sample mouth-type multimedia resource corresponding one-to-one to sample video frames in the sample styled face video; a third acquisition module for inputting a sample mouth shape driving feature and a sample video frame corresponding to a sample resource frame into an initial encoding network to obtain a predicted style vector output from the encoding network; a fourth acquisition module for inputting the predicted style vector, the sample mouth shape driving feature, and the sample reference face image into the face driving network to obtain a predicted style face image output from the face driving network; and a training module for performing a parameter adjustment process on the encoding network and the face driving network in the face mouth shape driving model based on a distribution to which the predicted style vector belongs, a Gaussian distribution, the predicted style face image, and a sample video frame corresponding to the sample resource frame, to obtain a trained face mouth shape driving model.

[0009] According to another aspect of the present disclosure, there is provided an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor such that the at least one processor can perform the facial video generation method proposed above in the present disclosure or the facial mouth shape drive model training method proposed above in the present disclosure.

[0010] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium having stored thereon computer instructions, the computer instructions causing a computer to perform the facial video generation method proposed by the above of the present disclosure or the facial mouth shape driving model training method proposed by the above of the present disclosure.

[0011] According to another aspect of the present disclosure, a computer program M When the computer program is executed by a processor, the steps of the facial video generation method proposed above in the present disclosure are realized, or the method for training a facial mouth shape drive model proposed above in the present disclosure is realized.

[0012] It should be understood that the contents described in this section are not intended to identify key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will be readily apparent from the following description. [Brief explanation of the drawings]

[0013] The drawings are used for better understanding of the present technical solution and are not intended to limit the present disclosure. [Figure 1] FIG. 1 is a schematic diagram according to a first embodiment of the present disclosure. [Figure 2] FIG. 10 is a schematic diagram according to a second embodiment of the present disclosure. [Figure 3]FIG. 10 is a schematic diagram according to a third embodiment of the present disclosure. [Figure 4] FIG. 1 is a schematic diagram of the training of the face-mouth drive model. [Figure 5] FIG. 10 is a schematic diagram according to a fourth embodiment of the present disclosure. [Figure 6] FIG. 10 is a schematic diagram according to a fifth embodiment of the present disclosure. [Figure 7] FIG. 1 is a block diagram of an electronic device for implementing a facial video generation method or a facial mouth shape driving model training method according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0014] The following describes exemplary embodiments of the present disclosure in conjunction with the drawings. For ease of understanding, various details of the embodiments of the present disclosure are included therein and should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, the following description will omit descriptions of well-known functions and structures.

[0015] The current face and mouth type driving solution mainly obtains a face and mouth type driving model, obtains a facial image and audio or video of a target object, inputs the audio or video and the facial image of the target object into the face and mouth type driving model, and obtains a facial video of the target object output from the face and mouth type driving model.

[0016] In the above technical proposal, the face and mouth shape driving model is a general-purpose face and mouth shape driving model, and the output face video is a face video of the target object in a general style, which makes it difficult to reflect the personalized mouth shape styles of different target objects, and the accuracy of the generated face video is low.

[0017] In response to the above problems, the present disclosure proposes a facial video generation method, apparatus and electronic device.

[0018] FIG. 1 is a schematic diagram of a first embodiment of the present disclosure, in which the facial video generation method of the embodiment of the present disclosure is applicable to a facial video generation apparatus, which can be provided in an electronic device so that the electronic device can perform a facial video generation function.

[0019] Here, the electronic device may be any device having any computing capability, such as a personal computer (abbreviated as PC), a mobile terminal, a server, etc. The mobile terminal may be a hardware device having various operating systems, a touch panel, and / or a display, such as an in-vehicle device, a mobile phone, a tablet computer, a personal digital assistant, a wearable device, a smart speaker, etc. In the following embodiments, an example will be described in which the execution entity is an electronic device.

[0020] As shown in FIG. 1, the face video generating method can include the following steps 101 to 105.

[0021] In step 101, a mouth shape multimedia resource and a reference face image of a target object are obtained.

[0022] In the embodiment of the present disclosure, the mouth-shaped multimedia resource is a mouth-shaped multimedia resource of a non-target object or a mouth-shaped multimedia resource obtained by synthesis.

[0023] In one example, the mouth-shaped multimedia resource is a mouth-shaped multimedia resource of a non-target object. The non-target object is an object different from the target object. Here, if the target object belongs to a specific object set, the non-target object may be an object other than the target object in this object set. The mouth-shaped multimedia resource may be a mouth-shaped multimedia resource of one non-target object, or may be a resource obtained by splicing the mouth-shaped multimedia resources of multiple non-target objects.

[0024] In another example, the mouth-shaped multimedia resource is a synthesized mouth-shaped multimedia resource, which may be a synthesized mouth-shaped multimedia resource for a specific character, or may be a resource obtained by splicing the synthesized mouth-shaped multimedia resources of multiple characters, such as an animated character.

[0025] The diversification of the acquisition methods of the portable multimedia resources makes it easier for electronic devices to quickly and conveniently acquire portable multimedia resources at lower cost, thereby reducing the acquisition costs of portable multimedia resources.

[0026] In the embodiments of the present disclosure, the mouth-shaped multimedia resource may be mouth-shaped audio or mouth-shaped video. The mouth-shaped video may be a video without audio and including the speaking action of an object, or may be a video with audio and including the speaking action of an object. The mouth-shaped audio is audio corresponding to a series of speaking actions of an object. That is, some audio frames in the audio coincide with the speaking action of the object and occur after the speaking action of the object.

[0027] The setting of mouth-shaped audio or mouth-shaped video facilitates the electronic device to select appropriate mouth-shaped multimedia resources as needed, and further reduces the acquisition cost of mouth-shaped multimedia resources.

[0028] In step 102, the reference style vector of the target object is obtained.

[0029] In an embodiment of the present disclosure, the reference style vector of the target object may be determined by combining it with any one facial image of the target object. Here, the facial image may include a mouth shape image region of the target object, from which a reference style vector capable of reflecting the mouth shape style of the target object may be extracted. Correspondingly, the process in which the electronic device performs step 102 may, for example, input the facial image of the target object into a style vector extraction model and obtain a reference style vector output from the style vector extraction model. Here, the facial image of the target object may be any one facial image of the target object or a reference facial image of the target object.

[0030] Here, the style vector extraction model can be obtained by training a combination of positive sample pairs and negative sample pairs, where the positive sample pairs can include two face images of the same object, and the negative sample pairs can include two face images of different objects.

[0031] In step 103, for each resource frame in the mouth-shape multimedia resource, a feature extraction process is performed on the resource frame to obtain a mouth-shape driving feature.

[0032] In the embodiments of the present disclosure, the process of the electronic device performing step 103 may be, for example, for each resource frame in the mouth-shape multimedia resource, input the resource frame into a feature extraction network in the face-mouth-shape driving model, and obtain mouth-shape driving features output from the feature extraction network.

[0033] Wherein, if the mouth-type multimedia resource is mouth-type audio, the resource frame may be an audio frame; if the mouth-type multimedia resource is mouth-type video, the resource frame may be a video frame.

[0034] In step 104, a styled face image corresponding to the resource frame is generated based on the mouth shape driving features, the reference face image, and the reference style vector.

[0035] In an embodiment of the present disclosure, in one example, the process in which the electronic device performs step 104 may be, for example, inputting the mouth shape driving features, the reference face image, and the reference style vector into a face driving network in a face mouth shape driving model, and obtaining a style face image output from the face driving network.

[0036] In another example, the process by which the electronic device performs step 104 may be, for example, determining a style mouth shape driving feature based on the mouth shape driving feature and the reference style vector, and generating a style facial image corresponding to the resource frame based on the style mouth shape driving feature and the reference facial image.

[0037] Here, the style mouth shape driving feature can reflect the personalized mouth shape style of the target object, ensure that the generated stylized face video can reflect the personalized mouth shape style of the target object, and further improve the accuracy of the generated stylized face video.

[0038] In step 105, a styled face video of the target object is determined according to the styled face image corresponding to each resource frame in the mouth shape multimedia resource.

[0039] In the embodiments of the present disclosure, the process of the electronic device performing step 105 may be, for example, performing a sorting and combining process on each style face image according to the sorted order of each resource frame in the mouth-shaped multimedia resource to obtain a style face image of the target object.

[0040] In the face video generation method of the embodiment of the present disclosure, a mouth-shaped multimedia resource and a reference face image of a target object are obtained, a reference style vector of the target object is obtained, and for each resource frame in the mouth-shaped multimedia resource, a feature extraction process is performed on the resource frame to obtain mouth-shaped driving features, and a styled face image corresponding to the resource frame is generated based on the mouth-shaped driving features, the reference face image, and the reference style vector, and a styled face video of the target object is determined based on the styled face image corresponding to each resource frame in the mouth-shaped multimedia resource, and the reference style vector of the target object can reflect the personalized mouth-shaped style of the target object, ensuring that the generated styled face video can reflect the personalized mouth-shaped style of the target object, and improving the accuracy of the generated styled face video.

[0041] Here, the electronic device can select a target Gaussian distribution from each candidate Gaussian distribution based on the sample resource frame, the sample reference face image, and the sample video frame, and further set the style vector that satisfies the target Gaussian distribution as the reference style vector of the target object, thereby accurately and quickly obtaining the reference style vector of the target object and reducing the amount of data processing when determining the reference style vector. As shown in Figure 2, Figure 2 is a schematic diagram of a second embodiment of the present disclosure, and the embodiment shown in Figure 2 can include the following steps 201 to 208.

[0042] In step 201, a mouth shape multimedia resource and a reference face image of a target object are obtained.

[0043] In step 202, each candidate Gaussian distribution is obtained.

[0044] Here, each candidate Gaussian distribution has a different mean and / or variance.

[0045] In step 203, obtain a sample resource frame in the sample mouth-shaped multimedia resource, a sample reference face image of the target object, and a sample video frame corresponding to the sample resource frame in the sample style face video of the target object.

[0046] In an embodiment of the present disclosure, the mouth-shape features of the sample resource frames in the sample mouth-shape multimedia resource match the mouth-shape features of the sample video frames in the sample styled facial video. Correspondingly, the process performed by the electronic device in step 203 may be, for example, to obtain a plurality of mouth-shape multimedia resources and a plurality of styled facial videos of a target object, determine a first mouth-shape feature of each resource frame for each mouth-shape multimedia resource, and determine a second mouth-shape feature of each video frame for each styled facial video of the target object, obtain a plurality of video combinations, each video combination including one mouth-shape multimedia resource and one styled facial video of the target object, and for each video combination, determine a mouth-shape feature matching degree between the mouth-shape multimedia resource and the styled facial video in the video combination based on the plurality of first mouth-shape features of the mouth-shape multimedia resources and the plurality of second mouth-shape features of the styled facial video in the video combination, and if the mouth-shape feature matching degree meets a predetermined matching degree condition, the mouth-shape multimedia resource in this video combination is the sample mouth-shape multimedia resource, and the styled facial video in this video combination is the sample styled facial video.

[0047] In step 204, a target Gaussian distribution is selected from each candidate Gaussian distribution based on the sample resource frame, the sample reference face image, and the sample video frame.

[0048] In an embodiment of the present disclosure, the process by which the electronic device performs step 204 may be, for example, determining sample mouth-shape driving features of a sample resource frame, determining a candidate style vector that fits the candidate Gaussian distribution for each candidate Gaussian distribution in turn, generating a predicted style face image based on the candidate style vector, a sample reference face image, and the sample mouth-shape driving features, and determining the candidate Gaussian distribution as a target Gaussian distribution if the similarity between the predicted style face image and the sample video frame satisfies a similarity condition.

[0049] Here, the candidate style vector may include values ​​of multiple dimensions. A candidate style vector that fits a candidate Gaussian distribution means that values ​​of multiple dimensions in the candidate style vector fit a Gaussian distribution.

[0050] Here, the number of candidate Gaussian distributions is small, and for each candidate Gaussian distribution, a candidate style vector that matches the candidate Gaussian distribution is determined in turn, and a predicted style face image is generated based on the candidate style vector, a sample reference face image, and a sample mouth shape driving feature, and a target Gaussian distribution is then selected, thereby shortening the time required to obtain the target Gaussian distribution, reducing the amount of data processing required when determining the target Gaussian distribution, and improving the accuracy of the selected target Gaussian distribution.

[0051] In step 205, the style vector that satisfies the target Gaussian distribution is set as the reference style vector of the target object.

[0052] In step 206, for each resource frame in the mouth-shape multimedia resource, a feature extraction process is performed on the resource frame to obtain a mouth-shape driving feature.

[0053] In step 207, a styled face image corresponding to the resource frame is generated based on the mouth shape driving features, the reference face image, and the reference style vector.

[0054] In step 208, a styled face video of the target object is determined based on the styled face image corresponding to each resource frame in the mouth shape multimedia resource.

[0055] For detailed explanations of step 201 and steps 206 to 208, please refer to the detailed explanations of step 101 and steps 103 to 105 in the embodiment of FIG. 1, and therefore further explanations will be omitted here.

[0056] The face video generating method of the embodiment of the present disclosure includes: obtaining a mouth-shaped multimedia resource and a reference face image of a target object; obtaining each candidate Gaussian distribution; obtaining a sample resource frame in the sample mouth-shaped multimedia resource, a sample reference face image of the target object, and a sample video frame corresponding to the sample resource frame in the sample style face video of the target object; selecting a target Gaussian distribution from each candidate Gaussian distribution based on the sample resource frame, the sample reference face image, and the sample video frame; determining a style vector that satisfies the target Gaussian distribution as the reference style vector of the target object; and performing a feature extraction process on each resource frame in the mouth-shaped multimedia resource. to obtain mouth-shape driving features, generate a styled face image corresponding to a resource frame based on the mouth-shape driving features, the reference face image, and the reference style vector, determine a styled face video of a target object based on the styled face image corresponding to each resource frame in the mouth-shape multimedia resource, set the reference style vector of the target object, and based on the sample resource frame, the sample reference face image, and the sample video frame, select a target Gaussian distribution from each candidate Gaussian distribution, and further obtain a reference style vector that fits the target Gaussian distribution, thereby ensuring that the generated styled face video can reflect the personalized mouth-shape style of the target object and improving the accuracy of the generated styled face video.

[0057] FIG. 3 is a schematic diagram of a third embodiment of the present disclosure, and the face-mouth drive model training method of the embodiment of the present disclosure is applicable to a face-mouth drive model training device, which can be provided in an electronic device so that the electronic device can perform the face-mouth drive model training function.

[0058] Here, the electronic device may be any device having any computing capability, such as a personal computer (abbreviated as PC), a mobile terminal, a server, etc. The mobile terminal may be a hardware device having various operating systems, a touch panel, and / or a display, such as an in-vehicle device, a mobile phone, a tablet computer, a personal digital assistant, a wearable device, a smart speaker, etc. In the following embodiments, an example will be described in which the execution entity is an electronic device.

[0059] As shown in FIG. 3, the face and mouth shape driving model training method can include the following steps 301 to 305.

[0060] In step 301, a pre-trained face and mouth shape driving model and an encoding network are obtained, and the face and mouth shape driving model includes a feature extraction network and a face driving network connected in sequence.

[0061] In an embodiment of the present disclosure, the input of the encoding network can be connected to the output of the feature extraction network, and the output of the encoding network can be connected to the input of the face driving network.

[0062] Here, the encoding network may be, for example, an encoder in a Conditional Variational Autoencoder (CVAE). Here, the CVAE can include one encoder and one decoder. The encoder extracts features from the input and transforms it into a specific Gaussian distribution, then selects a random vector as the output. The decoder deconvolves the input and reconstructs a brain volume structure image as the output from the features of the latent distribution.

[0063] In step 302, sample mouth-shaped driving features, a sample reference face image, and a sample style face video of each sample resource frame in the sample mouth-shaped multimedia resource are obtained, and the sample resource frames in the sample mouth-shaped multimedia resource correspond one-to-one to the sample video frames in the sample style face video.

[0064] In an embodiment of the present disclosure, the process by which the electronic device performs step 302 may be, for example, to obtain a sample mouth-shaped multimedia resource, a sample reference face image, and a sample style face video; and for each sample resource frame in the sample mouth-shaped multimedia resource, input the sample resource frame into a feature extraction network in a face mouth-shaped driving model, and obtain the sample mouth-shaped driving features of the sample resource frame output from the feature extraction network.

[0065] Here, the sample mouth-shaped multimedia resource may be a mouth-shaped multimedia resource of any one or more objects. The sample styled face video may be a styled face video of any one or more objects. Here, the objects corresponding to the sample mouth-shaped multimedia resource and the objects corresponding to the sample styled face video may be the same or different. The number of objects corresponding to the sample mouth-shaped multimedia resource and the number of objects corresponding to the sample styled face video may or may not match.

[0066] Here, by combining with the feature extraction network in the face-mouth-type driving model and performing feature extraction processing on the sample resource frame in the sample mouth-type multimedia resource, the accuracy of the determined and obtained sample mouth-type driving features can be improved.

[0067] In an embodiment of the present disclosure, the process by which an electronic device acquires a sample mouth-type multimedia resource, a sample reference face image, and a sample style face video may be, for example, acquiring a sample face video, setting the sample face video as a sample style face video, setting the sample face video or audio in the sample face video as a sample mouth-type multimedia resource, and setting any video frame in the sample face video as a sample reference face image.

[0068] Here, by respectively determining a sample mouth-shaped multimedia resource, a sample style face video, and a sample reference face image based on the sample face video, the cost of acquiring the sample mouth-shaped multimedia resource, the sample style face video, and the sample reference face image can be reduced, the acquisition efficiency can be improved, and the training speed of the face-mouth-shaped driving model can be further improved.

[0069] In step 303, for each sample resource frame in the sample resource frame type multimedia resource, the sample resource frame type driving feature and the sample video frame corresponding to the sample resource frame are input into an initial encoding network to obtain a prediction style vector output from the encoding network.

[0070] In step 304, the predicted style vector, the sample mouth shape driving features, and the sample reference face image are input into the face-driven network to obtain a predicted style face image output from the face-driven network.

[0071] In step 305, based on the distribution to which the predicted style vector belongs, the Gaussian distribution, the predicted style face image, and the sample video frame corresponding to the sample resource frame, a parameter adjustment process is performed on the encoding network and the face driving network in the face-mouth type driving model to obtain a trained face-mouth type driving model.

[0072] In an embodiment of the present disclosure, the process in which the electronic device performs step 305 may be, for example, determining a value of a first sub-loss function based on a distribution to which the predicted style vector belongs, a Gaussian distribution, and a first sub-loss function; determining a value of a second sub-loss function based on a predicted style face image, a sample video frame corresponding to the sample resource frame, and a second sub-loss function; determining a value of a loss function based on the value of the first sub-loss function and the value of the second sub-loss function; and performing a parameter adjustment process on the encoding network and the face driving network in the face-mouth-type driving model based on the value of the loss function to obtain a trained face-mouth-type driving model.

[0073] The distribution to which the predictive style vector belongs may be a probability distribution. The probability distribution may include at least one of a binomial distribution, a multinomial distribution, a hypergeometric distribution, a Poisson distribution, a normal distribution, an exponential distribution, and a uniform distribution. Here, the process by which the electronic device determines the distribution to which the predictive style vector belongs may include, for example, determining at least one distribution hypothesis based on the predictive style vector, determining, for each distribution hypothesis, distribution parameters for the distribution hypothesis and a degree of fit between the distribution hypothesis and the predictive style vector for the distribution parameters, selecting a target distribution hypothesis for the predictive style vector from the multiple distribution hypotheses based on the degrees of fit among the multiple distribution hypotheses, and determining the distribution to which the predictive style vector belongs based on the target distribution hypothesis and the distribution parameters for the target distribution hypothesis.

[0074] Here, the first sub-loss function may be, for example, KL dispersion, which is a measurement method for determining the difference between two probability distributions. When combined with the KL dispersion formula, it is possible to determine the difference between a Gaussian distribution and the distribution to which the predicted style vector belongs.

[0075] Here, the distribution to which the predicted style vector belongs, the Gaussian distribution, the predicted style face image, and the sample video frame corresponding to the sample resource frame are combined to determine the value of the first sub-loss function and the second sub-loss function, and parameter adjustment processing is further performed on the encoding network and the face-driven network, so that the trained face-driven network can combine with the style vector conforming to the Gaussian distribution to generate a style face video with a personalized style, and improve the accuracy of the generated style face video.

[0076] The face mouth shape driving model training method of the embodiment of the present disclosure includes obtaining a pre-trained face mouth shape driving model and an encoding network, the face mouth shape driving model including a feature extraction network and a face driving network connected in sequence, obtaining sample mouth shape driving features, a sample reference face image, and a sample style face video of each sample resource frame in the sample mouth shape multimedia resource, the sample resource frames in the sample mouth shape multimedia resource corresponding one-to-one to the sample video frames in the sample style face video, and inputting the sample mouth shape driving features and the sample video frame corresponding to each sample resource frame into the initial encoding network, and outputting the prediction from the encoding network. A measured style vector is obtained, and the predicted style vector, the sample mouth shape driving feature, and the sample reference face image are input into a face driving network to obtain a predicted style face image output from the face driving network. Based on the distribution to which the predicted style vector belongs, the Gaussian distribution, the predicted style face image, and the sample video frame corresponding to the sample resource frame, a parameter adjustment process is performed on the face driving network in the encoding network and the face mouth shape driving model to obtain a trained face mouth shape driving model, and a style face video generation process is further performed, and the trained face driving network can generate a style face video with a personalized style in combination with the style vector conforming to the Gaussian distribution, and the accuracy of the generated style face video can be improved.

[0077] An example will be given below. As shown in FIG. 4, FIG. 4 is a schematic diagram of the training of a face-mouth-type driving model. In FIG. 4, (1) an image / audio frame (a sample resource frame in a sample mouth-type multimedia resource) is input into a feature extraction network in the face-mouth-type driving model to obtain driving features (sample mouth-type driving features) output from the feature extraction network. (2) The driving features and a face image (true value) are input into an encoder (encoding network) to obtain a style vector (predicted style vector) output from the encoder. Here, the face image (true value) is a sample video frame corresponding to the sample resource frame. (3) The driving features and the style vector are input into a face driving network in the face-mouth-type driving model to obtain a face image (predicted style face image) output from the face driving network. (4) Combine with the Gaussian distribution and style vector to determine the KL loss (value of the first sub-loss function), combine with the face image output from the face-driven network and the face image (true value) to determine the pixel loss (value of the second sub-loss function), and then perform a training process on the encoder and face-driven network to obtain a trained face-mouth-shaped driven model.

[0078] To realize the above embodiments, the present disclosure further provides a facial video generating device. As shown in Fig. 5, Fig. 5 is a schematic diagram of a fourth embodiment of the present disclosure. The facial video generating device 50 may include a first acquisition module 501, a second acquisition module 502, a feature extraction module 503, a generation module 504, and a determination module 505.

[0079] Here, the first acquisition module 501 acquires a mouth-shaped multimedia resource and a reference face image of a target object, the second acquisition module 502 acquires a reference style vector of the target object, the feature extraction module 503 performs feature extraction processing on each resource frame in the mouth-shaped multimedia resource to obtain mouth-shaped driving features, the generation module 504 generates a styled face image corresponding to the resource frame based on the mouth-shaped driving features, the reference face image, and the reference style vector, and the determination module 505 determines a styled face video of the target object based on the styled face image corresponding to each resource frame in the mouth-shaped multimedia resource.

[0080] In one possible implementation form of the embodiments of the present disclosure, the reference style vector conforms to a Gaussian distribution, and the second acquisition module 502 includes a first acquisition unit, a second acquisition unit, a selection unit, and a determination unit, wherein the first acquisition unit acquires each candidate Gaussian distribution, the second acquisition unit acquires a sample resource frame in a sample mouth-shaped multimedia resource, a sample reference face image of the target object, and a sample video frame corresponding to the sample resource frame in a sample style face video of the target object, the selection unit selects a target Gaussian distribution from each candidate Gaussian distribution based on the sample resource frame, the sample reference face image, and the sample video frame, and the determination unit determines the style vector that satisfies the target Gaussian distribution as the reference style vector of the target object.

[0081] In one possible implementation form of the embodiments of the present disclosure, the selection unit specifically determines sample mouth-shape driving features of the sample resource frame, and for each candidate Gaussian distribution, determines a candidate style vector that matches the candidate Gaussian distribution in turn, generates a predicted style face image based on the candidate style vector, the sample reference face image, and the sample mouth-shape driving features, and determines the candidate Gaussian distribution as the target Gaussian distribution if the similarity between the predicted style face image and the sample video frame satisfies a similarity condition.

[0082] In one possible implementation form of the embodiments of the present disclosure, the generation module 504 specifically determines a style mouth shape driving feature based on the mouth shape driving feature and the reference style vector, and generates a style face image corresponding to the resource frame based on the style mouth shape driving feature and the reference face image.

[0083] In one possible implementation of the embodiment of the present disclosure, the mouth-shaped multimedia resource is mouth-shaped audio or mouth-shaped video.

[0084] In one possible implementation form of the embodiment of the present disclosure, the mouth-shaped multimedia resource is a mouth-shaped multimedia resource of a non-target object, or is a mouth-shaped multimedia resource obtained by synthesis.

[0085] A facial video generating device in an embodiment of the present disclosure obtains a mouth-shaped multimedia resource and a reference facial image of a target object, obtains a reference style vector of the target object, and for each resource frame in the mouth-shaped multimedia resource, performs a feature extraction process on the resource frame to obtain mouth-shaped driving features, generates a stylized facial image corresponding to the resource frame based on the mouth-shaped driving features, the reference facial image, and the reference style vector, and determines a stylized facial video of the target object based on the stylized facial image corresponding to each resource frame in the mouth-shaped multimedia resource, so that the reference style vector of the target object can reflect the personalized mouth-shaped style of the target object, ensuring that the generated stylized facial video can reflect the personalized mouth-shaped style of the target object, and improving the accuracy of the generated stylized facial video.

[0086] To realize the above embodiments, the present disclosure further provides a face-mouth drive model training device. As shown in FIG. 6, FIG. 6 is a schematic diagram of a fifth embodiment of the present disclosure. The face-mouth drive model training device 60 may include a first acquisition module 601, a second acquisition module 602, a third acquisition module 603, a fourth acquisition module 604, and a training module 605.

[0087] Here, the first acquisition module 601 acquires a pre-trained face and mouth shape driving model and an encoding network, and the face and mouth shape driving model includes a feature extraction network and a face driving network connected in sequence; the second acquisition module 602 acquires sample mouth shape driving features, a sample reference face image, and a sample style face video of each sample resource frame in the sample mouth shape multimedia resource, and the sample resource frames in the sample mouth shape multimedia resource correspond one-to-one to the sample video frames in the sample style face video; and the third acquisition module 603, for each sample resource frame in the sample mouth shape multimedia resource, acquires the sample mouth shape corresponding to the sample resource frame. The driving features and the sample video frame are input into an initial encoding network to obtain a predicted style vector output from the encoding network; a fourth acquisition module 604 inputs the predicted style vector, the sample mouth shape driving features, and the sample reference face image into the face driving network to obtain a predicted style face image output from the face driving network; a training module 605 performs a parameter adjustment process on the encoding network and the face driving network in the face mouth shape driving model based on the distribution to which the predicted style vector belongs, the Gaussian distribution, the predicted style face image, and the sample video frame corresponding to the sample resource frame, to obtain a trained face mouth shape driving model.

[0088] In one possible implementation form of the embodiments of the present disclosure, the second acquisition module 602 includes a first acquisition unit and a second acquisition unit, where the first acquisition unit acquires the sample mouth-shaped multimedia resource, the sample reference face image, and the sample style face video, and the second acquisition unit, for each sample resource frame in the sample mouth-shaped multimedia resource, inputs the sample resource frame into a feature extraction network in the face mouth-shaped driving model, and obtains the sample mouth-shaped driving features of the sample resource frame output from the feature extraction network.

[0089] In one possible implementation form of an embodiment of the present disclosure, the first acquisition unit specifically acquires a sample face video, the sample face video is the sample style face video, the sample face video or the audio in the sample face video is the sample mouth-type multimedia resource, and any video frame in the sample face video is the sample reference face image.

[0090] In one possible implementation form of the embodiments of the present disclosure, the training module 605 specifically determines a value of a first sub-loss function based on the distribution to which the predicted style vector belongs, the Gaussian distribution, and a first sub-loss function; determines a value of the second sub-loss function based on the predicted style face image, a sample video frame corresponding to the sample resource frame, and a second sub-loss function; determines a value of a loss function based on the value of the first sub-loss function and the value of the second sub-loss function; and performs a parameter adjustment process on the encoding network and the face driving network in the face and mouth shape driving model based on the value of the loss function to obtain a trained face and mouth shape driving model.

[0091] The face and mouth shape driving model training device of the embodiment of the present disclosure obtains a pre-trained face and mouth shape driving model and an encoding network, where the face and mouth shape driving model includes a feature extraction network and a face driving network connected in sequence; obtains sample mouth shape driving features, a sample reference face image, and a sample style face video of each sample resource frame in the sample mouth shape multimedia resource, where the sample resource frames in the sample mouth shape multimedia resource correspond one-to-one to the sample video frames in the sample style face video, and for each sample resource frame in the sample mouth shape multimedia resource, inputs the sample mouth shape driving features and the sample video frame corresponding to the sample resource frame into an initial encoding network, and outputs a prediction signal from the encoding network. A style vector is obtained, and the predicted style vector, the sample mouth shape driving feature, and the sample reference face image are input into a face driving network to obtain a predicted style face image output from the face driving network. Based on the distribution to which the predicted style vector belongs, the Gaussian distribution, the predicted style face image, and the sample video frame corresponding to the sample resource frame, a parameter adjustment process is performed on the face driving network in the encoding network and the face mouth shape driving model to obtain a trained face mouth shape driving model, and a style face video generation process is further performed, and the trained face driving network can generate a style face video with a personalized style in combination with the style vector that conforms to the fitted Gaussian distribution, and the accuracy of the generated style face video can be improved.

[0092] In addition, in the technical solution disclosed herein, the collection, storage, use, processing, transmission, provision, disclosure, etc. of relevant user personal information will all be carried out with the user's consent, will comply with the provisions of relevant laws and regulations, and will not violate public order and morals.

[0093] According to an embodiment of the present disclosure, the present disclosure provides an electronic device, and Readable storage medium Body More to offer. According to an embodiment of the present disclosure, the present disclosure provides a computer program, which, when executed by a processor, realizes the facial video generation method proposed by the present disclosure, or realizes the facial mouth shape driving model training method proposed by the present disclosure.

[0094] 7 is a schematic block diagram of an exemplary electronic device 700 for implementing embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, mobile phones, smartphones, wearable devices, and other similar computing devices. The components, their connections and relationships, and their functions shown herein are merely examples and are not intended to limit the description herein and / or the practice of the present disclosure as desired.

[0095] 7, electronic device 700 includes a computing unit 701 that can perform various appropriate operations and processes in accordance with a computer program stored in read-only memory (ROM) 702 or loaded from storage unit 708 into random access memory (RAM) 703. RAM 703 may also store various programs and data necessary for the operation of electronic device 700. Computing unit 701, ROM 702, and RAM 703 are connected to one another via a bus 704. An input / output (I / O) interface 705 is also connected to bus 704.

[0096] The components of the electronic device 700 are connected to an I / O interface 705, which includes an input unit 706 such as a keyboard, a mouse, etc., an output unit 707 such as various types of displays, speakers, etc., a storage unit 708 such as a magnetic disk, an optical disk, etc., and a communication unit 709 such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 enables the electronic device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0097] The computing unit 701 may be various general-purpose and / or special-purpose processing components having processing and computation capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various machine driving learning model algorithm computing units, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs each of the methods and processes described above, such as the facial video generation method or the facial mouth shape driving model training method. For example, in some embodiments, the facial video generation method or the facial mouth shape driving model training method may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 708. In some embodiments, some or all of the computer program may be loaded and / or installed into the electronic device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, the facial video generation method or the facial mouth shape driving model training method described above may be performed. Alternatively, in other embodiments, the computing unit 701 may be configured in any other suitable manner (e.g., via firmware) to perform the facial video generation method or the facial mouth shape drive model training method.

[0098] Various embodiments of the systems and techniques described herein above may be implemented in digital electronic circuitry systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include being implemented in one or more computer programs that may be executed and / or interpreted by a programmable system including at least one programmable processor, which may be an application specific or general purpose programmable processor, and that may receive data and instructions from, and transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0099] Program code for carrying out the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing apparatus such that, when executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are performed. The program code may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine, partially on a remote machine, or entirely on a remote machine or server.

[0100] In the context of this disclosure, a machine-readable medium may be a tangible medium that contains or can store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the above. More specific examples of machine-readable storage media include one or more line-based electrical connections, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0101] To provide interaction with a user, the systems and techniques described herein can be executed on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to a user, and a keyboard and pointing device (e.g., a mouse or trackball) through which a user can provide input to the computer. Other types of devices can also provide interaction with a user; for example, the feedback provided to the user can be any form of sensing feedback (e.g., visual feedback, auditory feedback, or haptic feedback) and can receive input from the user in any form (including acoustic, speech, or tactile input).

[0102] The systems and techniques described herein may be implemented in a computing system including back-end components (e.g., a data server), or a computing system including middleware components (e.g., an application server), or a computing system including front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with embodiments of the systems and techniques described herein), or a computing system including any combination of such back-end, middleware, and front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0103] The computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The relationship between the client and the server is created by computer programs running on corresponding computers and having a client-server relationship to each other. The server may be a cloud server, a server in a distributed system, or a server incorporating a blockchain.

[0104] It should be understood that steps can be rearranged, added, or deleted using the various types of flows shown above. For example, the steps described in the present disclosure may be performed in parallel, sequentially, or in a different order, but this specification is not limited thereto as long as the technical solution disclosed in the present disclosure can achieve the desired results.

[0105] The above specific embodiments do not limit the scope of protection of the present disclosure. It should be understood that those skilled in the art can make various modifications, combinations, subcombinations, and substitutions according to design requirements and other factors. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present disclosure should be included within the scope of protection of the present disclosure.

Claims

1. 1. A facial video generation method, comprising: obtaining a mouth shape multimedia resource and a reference face image of a target object; obtaining a reference style vector of the target object, the reference style vector reflecting a personalized mouth style of the target object; For each resource frame in the mouth shape multimedia resource, performing feature extraction processing on the resource frame to obtain mouth shape driving features; generating a styled face image corresponding to the resource frame based on the mouth shape driving features, the reference face image, and the reference style vector; and determining a styled face video of the target object according to a styled face image corresponding to each resource frame in the mouth shape multimedia resource. A method for facial video generation.

2. the reference style vector is fitted to a Gaussian distribution; The step of obtaining a reference style vector of the target object comprises: obtaining each candidate Gaussian distribution; Obtaining a sample resource frame in a sample mouth-style multimedia resource, a sample reference face image of the target object, and a sample video frame corresponding to the sample resource frame in a sample style face video of the target object; selecting a target Gaussian distribution from each candidate Gaussian distribution based on the sample resource frame, the sample reference face image, and the sample video frame; and determining a style vector that satisfies the target Gaussian distribution as a reference style vector of the target object. The facial video generation method of claim 1 .

3. selecting a target Gaussian distribution from each candidate Gaussian distribution based on the sample resource frame, the sample reference face image, and the sample video frame, determining a sample aperture driving characteristic of the sample resource frame; for each candidate Gaussian distribution in turn, determining a candidate style vector that fits said candidate Gaussian distribution; generating a predicted style face image based on the candidate style vector, the sample reference face image, and the sample mouth shape driving features; determining the candidate Gaussian distribution as the target Gaussian distribution if the similarity between the predicted style face image and the sample video frame satisfies a similarity condition; The facial video generation method of claim 2 .

4. generating a styled face image corresponding to the resource frame based on the mouth shape driving features, the reference face image, and the reference style vector, determining style mouth shape driving features based on the mouth shape driving features and the reference style vector; generating a styled face image corresponding to the resource frame based on the styled mouth shape driving features and the reference face image; The facial video generation method of claim 1 .

5. The mouth-shaped multimedia resource is a mouth-shaped audio or a mouth-shaped video; The facial video generation method of claim 1 .

6. The mouth-shaped multimedia resource is a mouth-shaped multimedia resource of a non-target object or a mouth-shaped multimedia resource obtained by synthesis; The facial video generation method of claim 1 .

7. A method for training a face-mouth shape drive model, comprising: Obtaining a pre-trained face and mouth shape driving model and an encoding network, wherein the face and mouth shape driving model includes a feature extraction network and a face driving network connected in sequence; Obtaining sample mouth-type driving features, a sample reference face image, and a sample style face video for each sample resource frame in a sample mouth-type multimedia resource, wherein the sample resource frames in the sample mouth-type multimedia resource correspond one-to-one to the sample video frames in the sample style face video; For each sample resource frame in the sample resource frame multimedia resource, input the sample resource frame driving feature and the sample video frame corresponding to the sample resource frame into an initial coding network to obtain a prediction style vector output from the coding network; inputting the predicted style vector, the sample mouth shape driving features, and the sample reference face image into the face-driven network to obtain a predicted style face image output from the face-driven network; According to a distribution to which the prediction style vector belongs, a Gaussian distribution, the prediction style face image, and a sample video frame corresponding to the sample resource frame, performing a parameter adjustment process on the encoding network and the face driving network in the face mouth shape driving model to obtain a trained face mouth shape driving model. Training method for face-mouth driven model.

8. The step of obtaining sample mouth-shaped driving features, a sample reference face image, and a sample style face video of each sample resource frame in the sample mouth-shaped multimedia resource includes: obtaining the sample mouth-shaped multimedia resource, the sample reference face image, and the sample style face video; For each sample resource frame in the sample mouth shape multimedia resource, input the sample resource frame into a feature extraction network in the face and mouth shape driving model to obtain the sample mouth shape driving features of the sample resource frame output from the feature extraction network; The method for training a face-mouth drive model according to claim 7.

9. The step of obtaining the sample mouth-shaped multimedia resource, the sample reference face image, and the sample style face video includes: obtaining a sample face video; the sample face video as the sample style face video; the sample face video or the audio in the sample face video as the sample mouth type multimedia resource; and selecting one of the video frames in the sample face video as the sample reference face image. The method for training a face-mouth drive model according to claim 8.

10. According to a distribution to which the prediction style vector belongs, a Gaussian distribution, the prediction style face image, and a sample video frame corresponding to the sample resource frame, performing a parameter adjustment process on the encoding network and the face driving network in the face and mouth shape driving model to obtain a trained face and mouth shape driving model, determining a value of the first sub-loss function based on a distribution to which the prediction style vector belongs, the Gaussian distribution, and a first sub-loss function; determining a value of the second sub-loss function based on the predicted style face image, a sample video frame corresponding to the sample resource frame, and a second sub-loss function; determining a value of a loss function based on the value of the first sub-loss function and the value of the second sub-loss function; and performing a parameter adjustment process on the encoding network and the face driving network in the face and mouth shape driving model based on the value of the loss function to obtain a trained face and mouth shape driving model. The method for training a face-mouth drive model according to claim 7.

11. A facial video generation device, comprising: a first acquisition module for acquiring a mouth shape multimedia resource and a reference face image of a target object; a second acquisition module for acquiring a reference style vector of the target object, the reference style vector reflecting a personalized mouth style of the target object; For each resource frame in the mouth shape multimedia resource, a feature extraction module is used to perform feature extraction processing on the resource frame to obtain mouth shape driving features; a generation module for generating a styled face image corresponding to the resource frame based on the mouth shape driving features, the reference face image, and the reference style vector; a determination module for determining a styled face video of the target object according to a styled face image corresponding to each resource frame in the mouth shape multimedia resource; Facial video generator.

12. the reference style vector is fitted to a Gaussian distribution; the second acquisition module includes a first acquisition unit, a second acquisition unit, a selection unit, and a determination unit; the first acquisition unit acquires each candidate Gaussian distribution; The second obtaining unit obtains a sample resource frame in a sample mouth-type multimedia resource, a sample reference face image of the target object, and a sample video frame corresponding to the sample resource frame in a sample style face video of the target object; the selection unit selects a target Gaussian distribution from each candidate Gaussian distribution based on the sample resource frame, the sample reference face image, and the sample video frame; the determining unit determines a style vector that satisfies the target Gaussian distribution as a reference style vector of the target object; The facial video generating device of claim 11.

13. The selection unit determining a sample aperture driving characteristic of the sample resource frame; for each candidate Gaussian distribution in turn, determining a candidate style vector that fits said candidate Gaussian distribution; generating a predicted style face image based on the candidate style vector, the sample reference face image, and the sample mouth shape driving features; determining the candidate Gaussian distribution as the target Gaussian distribution if the similarity between the predicted style face image and the sample video frame satisfies a similarity condition; The facial video generating device of claim 12.

14. The generation module: determining style mouth shape driven features based on the mouth shape driven features and the reference style vector; generating a styled face image corresponding to the resource frame based on the styled mouth shape driving features and the reference face image; The facial video generating device of claim 11.

15. The mouth-shaped multimedia resource is a mouth-shaped audio or a mouth-shaped video; The facial video generating device of claim 11.

16. The mouth-shaped multimedia resource is a mouth-shaped multimedia resource of a non-target object or a mouth-shaped multimedia resource obtained by synthesis; The facial video generating device of claim 11.

17. A face-mouth driven model training device, a first acquisition module for acquiring a pre-trained face-mouth shape driving model and an encoding network, the face-mouth shape driving model including a feature extraction network and a face driving network connected in sequence; a second acquisition module for acquiring sample mouth-type driving features, sample reference face images, and sample style face videos of each sample resource frame in the sample mouth-type multimedia resource, wherein the sample resource frames in the sample mouth-type multimedia resource correspond one-to-one to the sample video frames in the sample style face video; a third obtaining module for inputting, for each sample resource frame in the sample resource frame type multimedia resource, the sample resource frame type driving feature and the sample video frame corresponding to the sample resource frame into an initial encoding network to obtain a predictive style vector output from the encoding network; a fourth acquisition module for inputting the predicted style vector, the sample mouth shape driving features, and the sample reference face image into the face-driven network to obtain a predicted style face image output from the face-driven network; a training module for performing a parameter adjustment process on the encoding network and the face driving network in the face and mouth shape driving model according to a distribution to which the prediction style vector belongs, a Gaussian distribution, the prediction style face image, and a sample video frame corresponding to the sample resource frame, to obtain a trained face and mouth shape driving model; A face-to-mouth driven model training device.

18. the second acquisition module includes a first acquisition unit and a second acquisition unit; The first acquisition unit acquires the sample mouth-shaped multimedia resource, the sample reference face image, and the sample style face video; The second obtaining unit, for each sample resource frame in the sample mouth-shaped multimedia resource, inputs the sample resource frame into a feature extraction network in the face mouth-shaped driving model, and obtains the sample mouth-shaped driving features of the sample resource frame output from the feature extraction network; 18. A training device for a face-and-mouth driven model according to claim 17.

19. The first acquisition unit: Get a sample face video, the sample face video is the sample style face video; The sample face video or the audio in the sample face video is the sample mouth type multimedia resource; any video frame in the sample face video is taken as the sample reference face image; 19. The face-and-mouth drive model training device according to claim 18.

20. The training module comprises: determining a value of the first sub-loss function based on a distribution to which the prediction style vector belongs, the Gaussian distribution, and a first sub-loss function; determining a value of the second sub-loss function based on the predicted style face image, a sample video frame corresponding to the sample resource frame, and a second sub-loss function; determining a value of a loss function based on the value of the first sub-loss function and the value of the second sub-loss function; and performing a parameter adjustment process on the encoding network and the face driving network in the face and mouth shape driving model based on the value of the loss function to obtain a trained face and mouth shape driving model.

18. A training device for a face-and-mouth driven model according to claim 17.

21. An electronic device, at least one processor; a memory communicatively coupled to the at least one processor; The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the facial video generation method according to any one of claims 1 to 6 or the face-mouth shape drive model training method according to any one of claims 7 to 10. electronic equipment.

22. A non-transitory computer-readable storage medium having computer instructions stored thereon, comprising: The computer instructions cause a computer to perform a facial video generation method according to any one of claims 1 to 6 or a method for training a face and mouth shape drive model according to any one of claims 7 to 10. A non-transitory computer-readable storage medium.

23. A computer program comprising: When the computer program is executed by a processor, the facial video generation method according to any one of claims 1 to 6 or the face-mouth shape drive model training method according to any one of claims 7 to 10 is realized. Computer program.

Citation Information

Patent Citations

  • Image processing method and device, model training method and device, electronic apparatus, storage medium, and computer program

    JP2023011742A