Data processing method and apparatus, device, medium, and product

By obtaining reference videos and driving audio, using expression changes and speech characteristics to represent data to predict expression states, and combining the lip-driven model to adjust the expression states in the video, the difficult problem of determining the audio expression state in the video is solved, and high-precision expression state adjustment is achieved.

WO2025213838A1PCT designated stage Publication Date: 2025-10-16BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/139911
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-08
Filing Date
2024-12-17
Publication Date
2025-10-16

AI Technical Summary

Technical Problem

Existing technologies have difficulty effectively determining the facial expression state corresponding to the audio in a video, especially in scenarios where the video language is switched, sentences are modified or replaced, and it is difficult to accurately adjust the facial expression state in the video, such as the lip shape.

Method used

By obtaining reference video and driving audio, using expression change representation data, speech characteristics representation data and iconic facial representation data, the expression state representation data corresponding to the driving audio is predicted, and combined with the lip-driven model for adjustment and rendering to generate an adapted generated video.

Benefits of technology

It enables accurate adjustment of facial expressions in the video when switching between video languages, modifying or replacing sentences, ensuring that facial expressions such as lip shape are compatible with the audio, thereby improving the accuracy and effect of video processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024139911_16102025_PF_FP_ABST
    Figure CN2024139911_16102025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in embodiments of the present disclosure are a data processing method and apparatus, a device, a medium, and a product. The method comprises: upon acquiring a reference video and a driving audio, such as an i-th audio frame in a certain audio segment, first obtaining speech characteristic representation data of the reference video on the basis of expression change representation data of the reference video, so that the speech characteristic representation data can represent speech characteristics, such as exaggerated mouth opening during speaking, presented in the reference video; and then, on the basis of the speech characteristic representation data, audio features of the driving audio, and distinctive facial representation data of the reference video, predicting expression state representation data corresponding to the driving audio, so that the expression state representation data can represent the expression state of a subject in the reference video under the driving audio, such as a lip shape state.
Need to check novelty before this filing date? Find Prior Art

Description

A data processing method, device, equipment, medium and product

[0001] The present application claims priority from the Chinese patent application No. 202410417923.0 filed on April 8, 2024 and entitled "A data processing method, device, equipment, medium and product", the whole content of which is incorporated herein by reference. TECHNICAL FIELD

[0002] Embodiments of the present disclosure relate to the technical field of data processing, and in particular, to a data processing method, device, equipment, medium and product. BACKGROUND

[0003] For some application scenarios, such as language switching processing of a video, sentence modification processing in a video, or sentence replacement processing in a video, the following requirements may exist: determining expression representation data corresponding to an audio according to an existing video, so that the expression representation data can represent an expression state, such as a mouth shape state, to which the audio is adapted, so as to subsequently complete an expression state adjustment task, such as a mouth shape modification task, based on the expression representation data. SUMMARY

[0004] At least one embodiment of the present disclosure provides a data processing method, which comprises:

[0005] obtaining a reference video and a driving audio;

[0006] obtaining speaking feature representation data of the reference video according to expression change representation data of the reference video;

[0007] predicting expression state representation data corresponding to the driving audio according to the speaking feature representation data, audio features of the driving audio, and identifying facial representation data of the reference video; the expression state representation data is used to represent a mouth shape state of an object in the reference video under the driving audio.

[0008] In a possible implementation, the speaking feature representation data is determined from a continuous distribution according to an encoding feature of the expression change representation data; the continuous distribution is used to describe speaking feature representation data corresponding to different encoding features.

[0009] In a possible implementation, the continuous distribution is learned from a plurality of sample videos; different sample videos present different speaking features.

[0010] In a possible implementation, for any one frame of image in the reference video, three-dimensional facial parameters of the image include identifying parameters, expression parameters and posture parameters;

[0011] The expression change representation data comprises three-dimensional face key points corresponding to the expression parameters of each frame image in the reference video.

[0012] The identifying face representation data is determined according to three-dimensional face key points corresponding to the identification parameters of at least one frame image in the reference video.

[0013] In a possible implementation, the driving audio refers to any one frame of audio in the audio sequence.

[0014] The driving audio corresponds to a target image in the reference video.

[0015] After predicting the expression state representation data corresponding to the driving audio, the method further comprises:

[0016] According to the three-dimensional face key points corresponding to the posture parameters of the target image, the expression state representation data corresponding to the driving audio is updated; the updated expression state representation data represents a posture state consistent with the posture state presented in the target image.

[0017] In a possible implementation, the determination process of the expression state representation data corresponding to the driving audio comprises:

[0018] According to the speaking feature representation data, the audio feature, and the identifying face representation data, the face adjustment representation data corresponding to the driving audio is predicted.

[0019] According to the face adjustment representation data, the identifying face representation data is adjusted to obtain the expression state representation data corresponding to the driving audio.

[0020] In a possible implementation, the determination process of the expression state representation data corresponding to the driving audio comprises:

[0021] According to the speaking feature representation data, the audio feature, and the identifying face representation data, the expression coefficient corresponding to the driving audio is predicted.

[0022] According to the expression coefficient, the expression state representation data corresponding to the driving audio is generated.

[0023] In a possible implementation, the driving audio refers to any one frame of audio in the audio sequence.

[0024] After predicting the expression state representation data corresponding to the driving audio, the method further comprises:

[0025] rendering processing is performed on the expression state representation data corresponding to the driving audio to obtain generated images corresponding to the driving audio;

[0026] A generated video corresponding to the audio sequence is constructed according to the generated images corresponding to each frame of audio in the audio sequence, and the generated video is used to describe the lip movement change state of the object under the audio sequence.

[0027] In a possible implementation, the expression state representation data corresponding to the driving audio is determined by using a lip movement driving model.

[0028] The driving audio refers to a frame of audio extracted from a sample video.

[0029] The reference video is determined according to the sample video.

[0030] After predicting the expression state representation data corresponding to the driving audio, the method further includes:

[0031] The lip movement driving model is updated according to the expression state representation data corresponding to the driving audio and an expression state label corresponding to the driving audio, and the expression state label is determined according to an image corresponding to the driving audio in the sample video.

[0032] In a possible implementation, the expression state representation data corresponding to the driving audio is determined by using a lip movement driving model.

[0033] The lip movement driving model includes an audio encoding module, a face encoding module, a speaking feature analysis module, and an expression prediction module.

[0034] The audio encoding module is configured to perform encoding processing on the driving audio to obtain audio features of the driving audio.

[0035] The face encoding module is configured to perform encoding processing on the expression change representation data to obtain encoding features of the expression change representation data, and perform encoding processing on the identifying face representation data to obtain encoding features of the identifying face representation data.

[0036] The speaking feature analysis module is configured to analyze the encoding features of the expression change representation data to obtain the speaking feature representation data.

[0037] The expression prediction module is configured to predict the expression state representation data corresponding to the driving audio according to the speaking feature representation data, the audio features, and the encoding features of the identifying face representation data.

[0038] At least one embodiment of the present disclosure provides a data processing apparatus, which comprises:

[0039] an acquisition unit configured to acquire a reference video and a driving audio;

[0040] an analysis unit configured to analyze, according to expression change representation data of the reference video, to obtain speaking feature representation data of the reference video;

[0041] a prediction unit configured to predict, according to the speaking feature representation data, audio features of the driving audio, and identifying facial representation data of the reference video, expression state representation data corresponding to the driving audio; the expression state representation data is used to represent a lip shape state of an object in the reference video under the driving audio.

[0042] At least one embodiment of the present disclosure provides an electronic device, the device comprising: a processor and a memory;

[0043] the memory is configured to store instructions or computer programs;

[0044] the processor is configured to execute the instructions or computer programs in the memory, so that the electronic device executes the data processing method provided by the embodiments of the present disclosure.

[0045] At least one embodiment of the present disclosure provides a computer readable medium, the computer readable medium stores instructions or computer programs, when the instructions or computer programs run on a device, the device executes the data processing method provided by the embodiments of the present disclosure.

[0046] At least one embodiment of the present disclosure provides a computer program product, which includes a computer program carried on a non-transitory computer readable medium, and the computer program includes program codes for executing the data processing method provided by the embodiments of the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the related art, the drawings needed to be used in the embodiments or related art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments described in the present disclosure, and for those skilled in the art, other drawings can be obtained without creative labor on the basis of these drawings.

[0048] FIG. 1 is a flowchart of a data processing method provided by an embodiment of the present disclosure;

[0049] FIG. 2 is a schematic diagram of a lip shape modification process provided by an embodiment of the present disclosure;

[0050] FIG. 3 is a schematic diagram of an expression prediction process provided by an embodiment of the present disclosure;

[0051] FIG. 4 is a structural schematic diagram of a data processing apparatus according to an embodiment of the present disclosure;

[0052] FIG. 5 is a structural schematic diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0053] In order to enable persons skilled in the art to better understand the solutions of the embodiments of the present disclosure, the technical solutions in the embodiments of the present disclosure will be described clearly and completely below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only some of the embodiments of the present disclosure, but not all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by persons skilled in the art without creative work fall within the scope of protection of the present disclosure.

[0054] In order to better understand the technical solutions provided by the embodiments of the present disclosure, the data processing method provided by the embodiments of the present disclosure will be described below in conjunction with some drawings. As shown in FIG. 1, the data processing method provided by the embodiments of the present disclosure includes the following S1-S3. Wherein, FIG. 1 is a flowchart of a data processing method according to an embodiment of the present disclosure.

[0055] S1: Obtain a reference video and a driving audio.

[0056] Wherein, the reference video refers to a video required for expression determination processing, such as the reference video shown in FIG. 2, so that the reference video is used to provide information other than lip shapes, such as facial identification information similar to facial contour features and facial feature distribution features, speaking features, etc.

[0057] In addition, the embodiments of the present disclosure do not limit the implementation of the above-mentioned reference video. In order to facilitate understanding, two scenarios will be described below.

[0058] Scenario one, when the data processing method provided by the embodiments of the present disclosure is used to perform a task involving expression determination processing, such as a lip modification task, the above-mentioned reference video can be a video specified by a user through certain means, such as a single-person lip-synching video, so that the reference video meets some needs of the user, such as needs in terms of facial identification information, speaking features, etc. It should be noted that the present disclosure does not limit the means, for example, the reference video can be a video manually uploaded by the user, or a video selected by the user from some candidate videos, or a video downloaded by the user with the help of certain means.

[0059] Scene two, when the data processing method provided by the embodiment of the present disclosure is used to implement the model training process, the above-mentioned reference video can refer to a sample video. Wherein, the sample video refers to a video required to be used in the model training process; and the embodiment of the present disclosure does not limit the implementation of the sample video.

[0060] In addition, the embodiment of the present disclosure does not limit the acquisition method of the above-mentioned reference video.

[0061] The driving audio refers to the audio required to be used in the expression determination process, such as the driving audio shown in FIG. 3, so that the driving audio is used to provide the driving information required to be used in the expression determination process, such as the lip shape driving information.

[0062] In addition, the embodiment of the present disclosure does not limit the implementation of the above-mentioned driving audio, and the following will be described in combination with two scenes.

[0063] Scene one, when the data processing method provided by the embodiment of the present disclosure is used to execute a task related to the expression determination process, such as a lip shape modification task, the above-mentioned driving audio can refer to any one frame of audio in an audio sequence specified by a user, so that subsequent expression determination processing for the driving audio can be used to complete the expression determination processing for all audios in the audio sequence. Wherein, the audio sequence refers to an audio sequence specified by a user through a certain means, such as the audio sequence shown in FIG. 2; and the embodiment of the present disclosure does not limit the implementation of the audio sequence, for example, in some application scenarios, such as a video translation scenario, the audio sequence satisfies the following constraints: the semantic information of the sentence described by the audio sequence is consistent with the semantic information of the sentence described in the reference video, but the language of the sentence described by the audio sequence is different from the language of the sentence described in the reference video. For another example, in some application scenarios, such as a video sentence modification scenario, the audio sequence at least satisfies the following constraints: the sentence described by the audio sequence is partially the same as the sentence described in the reference video. For another example, in some application scenarios, such as a video sentence replacement scenario, the audio sequence at least satisfies the following constraints: the sentence described by the audio sequence is completely different from the sentence described in the reference video.

[0064] Scene two, when the data processing method provided by the embodiment of the present disclosure is used to implement the model training process, if the above-mentioned reference video refers to a sample video, the above-mentioned driving audio refers to any one frame of audio in an audio sequence extracted from the sample video. It should be noted that the embodiment of the present disclosure does not limit the relationship between the audio sequence and the sample video, for example, the audio sequence can include part or all of the audios in the sample video.

[0065] In addition, the embodiment of the present disclosure does not limit the manner of obtaining the driving audio as described above, for example, after obtaining the audio sequence, the i-th frame of audio in the audio sequence is determined as the driving audio, so that subsequent expression determination processing for the i-th frame of audio can be completed by means of expression determination processing for the driving audio, i is a positive integer, i≤I, and I represents the total number of frames of the audio sequence.

[0066] Furthermore, the embodiment of the present disclosure does not limit the association relationship between the audio sequence and the reference video as described above, for example, the two can satisfy the following constraint: for the i-th frame of audio in the audio sequence, there is an image corresponding to the i-th frame of audio in the reference video, such as the i-th frame of image in the reference video. Wherein, i is a positive integer, i≤I, and I represents the total number of frames of the audio sequence.

[0067] It should be noted that the embodiment of the present disclosure does not limit the implementation of the corresponding relationship as shown in the above paragraph, for example, when the total number of frames of the audio sequence is equal to the total number of frames of the reference video, the corresponding relationship can specifically include: the corresponding relationship between the i-th frame of audio in the audio sequence and the i-th frame of image in the reference video. For another example, when the total number of frames of the audio sequence is greater than the total number of frames of the reference video, the corresponding relationship can specifically include: the corresponding relationship between the i-th frame of audio in the audio sequence and the i-th frame of image in the reference video after frame increasing. Wherein, the reference video after frame increasing is obtained by performing a certain frame increasing processing on the original reference video, so that the total number of frames of the reference video after frame increasing is equal to the total number of frames of the audio sequence; and the embodiment of the present disclosure does not limit the implementation of the frame increasing processing, for example, it can be implemented by using any method that can increase the total number of frames of a video, such as frame insertion processing or copy splicing processing on the original video. For another example, when the total number of frames of the audio sequence is less than the total number of frames of the reference video, the corresponding relationship can specifically include: the corresponding relationship between the i-th frame of audio in the audio sequence and the i-th frame of image in the reference video after frame decreasing. Wherein, the reference video after frame decreasing is obtained by performing a certain frame decreasing processing on the original reference video, so that the total number of frames of the reference video after frame decreasing is equal to the total number of frames of the audio sequence; and the embodiment of the present disclosure does not limit the implementation of the frame decreasing processing, for example, it can be implemented by using any method that can decrease the total number of frames of a video, such as sampling processing or cropping processing on the original video. Wherein, i is a positive integer, i≤I, and I represents the total number of frames of the audio sequence.

[0068] Based on the related content of S1 above, in some application scenarios, such as performing a certain task or model training scenario, a reference video and an audio sequence can be obtained, and the i-th frame of audio in the audio sequence is regarded as a driving audio, so that subsequent expression determination processing for all audios in the audio sequence can be completed by means of expression determination processing for the driving audio. Wherein, i is a positive integer, i≤I, I represents the total number of frames of the audio sequence.

[0069] S2: obtaining the speaking characteristic representation data of the reference video according to the expression change representation data of the reference video.

[0070] Wherein, the expression change representation data of the reference video is used to represent the expression change situation presented by the reference video, such as the lip shape change situation, so that the expression change representation data can represent the expression change situation of the object in the reference video; and the expression change representation data is determined according to the reference video. Wherein, the object refers to the object described by the reference video, and the embodiments of the object are not limited by the embodiments of the object, for example, the object can be implemented by any kind of animal or virtual image capable of presenting expression.

[0071] In addition, the embodiments of the expression change representation data are not limited by the embodiments of the expression change representation data, for example, when the driving audio refers to any one frame of audio in the audio sequence, and the total number of frames of the audio sequence is equal to the total number of frames of the reference video, the expression change representation data can include the expression state representation data corresponding to each frame of image in the reference video. Wherein, the expression state representation data corresponding to the i-th frame of image in the reference video is used to represent the expression state presented in the i-th frame of image, such as the lip shape state; and the embodiments of the expression state representation data corresponding to the i-th frame of image are not limited by the embodiments of the expression state representation data corresponding to the i-th frame of image. i is a positive integer, i≤I.

[0072] In addition, in the lip modification scenario, in order to avoid as much as possible the interference caused by other facial information such as facial identification information and posture information in addition to the lip, the present embodiment also provides a possible implementation of the expression state representation data corresponding to the i-th image, in which when the three-dimensional facial parameters of the i-th image include identity (Identity document, ID) parameters, expression parameters and posture parameters, the expression state representation data corresponding to the i-th image can be the three-dimensional facial key points corresponding to the i-th image under the expression parameters, so that the expression state representation data corresponding to the i-th image can represent the expression state such as the lip state presented in the i-th image in the case of no ID, no posture and expression, thereby effectively avoiding the interference of the facial identification information and the posture information on the expression state, and further enabling the expression state representation data corresponding to the i-th image to as accurately as possible represent the expression state such as the lip state presented in the i-th image. The three-dimensional facial parameters of the i-th image are used to describe the facial state of the face in the i-th image in three-dimensional space; and the present embodiment does not limit the implementation of the three-dimensional facial parameters, for example, it can be implemented by 3DMM (3D Morphable Model) parameters. The ID parameters are used to describe the facial identification information of the face in the i-th image, such as facial contour characteristics, facial feature distribution characteristics and the like, so that the three-dimensional facial key points corresponding to the i-th image under the ID parameters can represent the facial identification information in three-dimensional space. The expression parameters are used to describe the expression state such as the lip state presented by the face in the i-th image, so that the three-dimensional facial key points corresponding to the i-th image under the expression parameters can represent the expression state presented by the face in the i-th image in three-dimensional space. The posture parameters are used to describe the posture state such as the side face and the front face presented by the face in the i-th image, so that the three-dimensional facial key points corresponding to the i-th image under the posture parameters can represent the facial posture presented by the face in the i-th image in three-dimensional space.

[0073] Based on the above two paragraphs, in a possible implementation, if the three-dimensional face parameters of any image in the reference video include the identity parameter, the expression parameter and the posture parameter, the expression change representation data of the reference video can include the three-dimensional face key points corresponding to the expression parameter of each image in the reference video, so that the expression change representation data can represent the expression change situation presented by the reference video without ID and posture, so that the expression change representation data can more accurately represent the expression change situation presented by the reference video, such as the lip shape change situation, and further make the expression change representation data more accurately represent the speaking characteristics presented by the reference video, such as the mouth opening amplitude when speaking, whether the mouth is crooked when speaking, whether the eyebrows are raised when speaking, and the like.

[0074] In addition, the embodiments of the present disclosure do not limit the implementation of the expression change representation data of the reference video described above, for example, it can be implemented in a sequence manner. It can be seen that, in a possible implementation, if the three-dimensional face parameters of any image in the reference video include the identity parameter, the expression parameter and the posture parameter, the expression change representation data of the reference video can be a sequence. The sequence is constructed according to the three-dimensional face key points corresponding to the expression parameter of each image in the reference video, and the position of the three-dimensional face key points corresponding to the expression parameter of the i-th image in the sequence is the same as the position of the i-th image in the reference video.

[0075] In addition, the embodiments of the present disclosure do not limit the implementation of the expression change representation data of the reference video described above, for example, it can be implemented in a sequence manner. It can be seen that, in a possible implementation, if the three-dimensional face parameters of any image in the reference video include the identity parameter, the expression parameter and the posture parameter, the expression change representation data of the reference video can be a sequence. The sequence is constructed according to the three-dimensional face key points corresponding to the expression parameter of each image in the reference video, and the position of the three-dimensional face key points corresponding to the expression parameter of the i-th image in the sequence is the same as the position of the i-th image in the reference video.

[0076] The speaking characteristics representation data of the reference video is used to represent the speaking characteristics presented in the reference video, such as the mouth opening amplitude when speaking, whether the mouth is crooked when speaking, whether the eyebrows are raised when speaking, and the like, so that the speaking characteristics representation data can represent the speaking characteristics of the object in the reference video; and the speaking characteristics representation data is analyzed from the expression change representation data of the reference video.

[0077] In addition, in order to better improve the determination effect of the speaking feature, the embodiment of the disclosure further provides an analysis process of the speaking feature representation data, which can be specifically: determining the speaking feature representation data of the reference video from a continuous distribution according to the encoding feature of the expression change representation data of the reference video, so that the speaking feature representation data satisfies the continuous distribution. It should be noted that the embodiment of the disclosure does not limit the implementation of the encoding process, for example, it can be implemented by using any encoding method, such as a multilayer perceptron (MLP).

[0078] In addition, for the continuous distribution shown in the above paragraph, the continuous distribution is a continuous distribution, and the continuous distribution is used to describe the speaking feature representation data corresponding to different encoding features, so that the embodiment of the disclosure can determine the speaking feature representation data corresponding to any encoding feature by means of the continuous distribution; and the embodiment of the disclosure does not limit the implementation of the continuous distribution. In order to facilitate understanding, two scenarios are described below.

[0079] Scenario one, when the data processing method provided by the embodiment of the disclosure is used to perform a task involving expression determination processing, such as a lip modification task, the continuous distribution described above can refer to a continuous distribution learned from a plurality of sample videos, used to describe the speaking feature representation data corresponding to different encoding features. Among them, because the speaking features presented in different sample videos are different, the continuous distribution learned based on these sample videos can more accurately represent the speaking feature representation data corresponding to different encoding features, so that the speaking feature representation data determined based on the continuous distribution is more accurate, which is conducive to improving the speaking feature determination effect, thereby improving the expression determination effect. Also because the continuous distribution is continuous, the continuous distribution can provide the speaking feature representation data corresponding to any encoding feature, so that the continuous distribution can not only provide the speaking feature representation data corresponding to the encoding features involved in the learning process, but also provide the speaking feature representation data corresponding to the encoding features not involved in the learning process, thereby effectively avoiding the defects caused by the limitation of the training data involved in the learning process, which is conducive to improving the speaking feature determination effect, thereby improving the expression determination effect.

[0080] It should be noted that the embodiment of the disclosure does not limit the implementation of the learning process of the continuous distribution shown in the above paragraph, for example, the learning process of the continuous distribution can be implemented by means of the training process of the lip driving model. The related content of the lip driving model is described below.

[0081] In a second scenario, when the data processing method provided by the embodiments of the present disclosure is used to implement the model training process, the continuous distribution above refers to a continuous distribution that needs to be referred to by a speaking feature analysis module in the lip-sync model when performing speaking feature analysis processing. The initial value of the continuous distribution in the speaking feature analysis module can be set in advance based on the application scenario, and the continuous distribution in the speaking feature analysis module can be continuously optimized with the update of the lip-sync model.

[0082] In addition, the embodiments of the present disclosure do not limit the implementation of S2 above, for example, S2 can specifically be: obtaining speaking feature representation data of the reference video by analyzing the encoding features of the expression change representation data of the reference video. For another example, S2 can be implemented by means of a machine learning model with a speaking feature analysis function constructed in advance, such as a model constructed by means of a convolutional neural network (CNN) and an MLP.

[0083] Based on the related content of S2 above, for some application scenarios, after obtaining the reference video, the three-dimensional face parameters of each frame image in the reference video, such as ID parameters, expression parameters, and posture parameters, can be obtained first; then the expression change representation data of the reference video can be constructed according to the three-dimensional face key points corresponding to the expression parameters of each frame image in the reference video, so that the expression change representation data can more accurately represent the expression change situation, such as the lip shape change situation, when speaking presented by the reference video, so that the expression change representation data can describe the speaking features of the object to some extent; then the expression change representation data is encoded to obtain the encoding features of the expression change representation data, so that the encoding features can better represent the expression change situation presented by the reference video; finally, the speaking feature representation data of the reference video is determined from the continuous distribution according to the encoding features of the expression change representation data, so that the speaking feature representation data can more accurately represent the speaking features presented by the reference video.

[0084] S3: predicting expression state representation data corresponding to the driving audio according to the speaking feature representation data of the reference video, the audio features of the driving audio, and the identity face representation data of the reference video; the expression state representation data is used to represent the lip shape state of the object in the reference video under the driving audio.

[0085] The audio feature of the driving audio is used to represent the audio characteristics carried by the driving audio. The disclosure embodiments do not limit the manner of obtaining the audio feature, which can be implemented by using any existing or future audio feature extraction method, such as a content encoder. It should be noted that the disclosure embodiments do not limit the implementation of the content encoder, which can be implemented by using the encoding module 1 or the transformer shown in FIG. 3.

[0086] The identifying facial representation data of the reference video is used to represent the facial identification information presented by the reference video, such as facial contour and facial feature distribution, so that the identifying facial representation data can accurately represent the facial characteristics of the object in the reference video. The identifying facial representation data is determined based on the reference video.

[0087] In addition, the disclosure embodiments do not limit the manner of determining the identifying facial representation data of the reference video, which can be implemented by using any existing or future method for determining the facial ID of an object from a video, such as by using a pre-constructed machine learning model with facial ID analysis function.

[0088] In addition, in order to better improve the expression determination effect, the disclosure embodiments further provide a determination process of the identifying facial representation data of the reference video. If the three-dimensional facial parameters of any image in the reference video include identification parameters, expression parameters and posture parameters, the identifying facial representation data of the reference video can be determined based on the three-dimensional facial key points corresponding to the identification parameters of at least one image in the reference video.

[0089] In addition, the disclosure embodiments do not limit the implementation of the determination process of the identifying facial representation data in the above paragraph, which can be specifically implemented by: first randomly selecting an image from the reference video, and then determining the three-dimensional facial key points corresponding to the identification parameters of the selected image as the identifying facial representation data of the reference video.

[0090] For example, the determination process of the identifying facial representation data of the reference video can be specifically implemented by: performing average value calculation processing on the three-dimensional facial key points corresponding to the identification parameters of all images in the reference video to obtain the identifying facial representation data of the reference video, so that the identifying facial representation data can more accurately represent the facial identification information presented by the object in the reference video, which is conducive to improving the expression determination effect.

[0091] Based on the above content, in one possible implementation, after the reference video is acquired, the three-dimensional face parameters of each frame image in the reference video can be acquired first, such as the ID parameter, the expression parameter, and the posture parameter; then the average value of the three-dimensional face key points corresponding to the ID parameter of all images in the reference video is calculated as the identifying face representation data of the reference video, so that the identifying face representation data can more accurately represent the face state of the reference video in the case of no expression, no posture, and ID, thereby enabling the identifying face representation data to more accurately represent the face identifying information of the object in the reference video.

[0092] The expression state representation data corresponding to the driving audio is used to represent the expression state of the object in the reference video under the driving audio, such as the lip shape state, so that the expression state representation data can represent the expression state, such as the lip shape state, adapted by the object under the driving audio; and the embodiment of the disclosure does not limit the implementation of the expression state representation data corresponding to the driving audio. For example, when the expression change representation data and the identifying face representation data of the reference video are implemented by using the three-dimensional face key points, the expression state representation data corresponding to the driving audio can be implemented by using the three-dimensional face key points.

[0093] It can be seen that, in one possible implementation, the expression state representation data corresponding to the driving audio can include the three-dimensional face key points corresponding to the driving audio under the ID parameter and the three-dimensional face key points corresponding to the driving audio under the expression parameter, so that the face identifying information represented by the expression state representation data is consistent with the face identifying information represented by the identifying face representation data of the reference video, and the expression state represented by the expression state representation data, such as the lip shape state, is as consistent as possible with the driving audio. The three-dimensional face key points corresponding to the driving audio under the ID parameter are used to represent the face identifying information of the object in the reference video, so that the three-dimensional face key points corresponding to the driving audio under the ID parameter are consistent with the identifying face representation data of the reference video. The three-dimensional face key points corresponding to the driving audio under the expression parameter are used to represent the expression state, such as the lip shape state, adapted by the driving audio.

[0094] In addition, the embodiment of the disclosure does not limit the implementation of S3, and some cases are described below for ease of understanding.

[0095] In some application scenarios, the absolute coordinates of the three-dimensional facial key points corresponding to the driving audio can be directly predicted. Based on this, the embodiment of the present disclosure also provides a possible implementation of S3 above, in which S3 can be specifically: predicting the expression state representation data corresponding to the driving audio according to the speaking feature representation data of the reference video, the audio features of the driving audio, and the encoding features of the identifying facial representation data of the reference video, so that the expression state representation data can represent the absolute position of the three-dimensional facial key points corresponding to the driving audio in the three-dimensional space. The encoding features of the identifying facial representation data refer to the results obtained by encoding the identifying facial representation data, and the embodiment of the present disclosure does not limit the implementation of the encoding processing. For example, the implementation of the encoding processing is similar to the implementation of the encoding processing involved in the encoding features of the expression change representation data.

[0096] It should be noted that the embodiment of the present disclosure does not limit the implementation of the prediction in the above paragraph, for example, it can be implemented by using a machine learning model with an absolute coordinate prediction function of three-dimensional facial key points, such as MLP.

[0097] In some application scenarios, the offset of the three-dimensional facial key points corresponding to the driving audio relative to the identifying facial representation data above can be directly predicted. Based on this, the embodiment of the present disclosure also provides a possible implementation of S3 above, in which S3 can specifically include the following steps 11-12.

[0098] Step 11: predicting the facial adjustment representation data corresponding to the driving audio according to the speaking feature representation data of the reference video, the audio features of the driving audio, and the identifying facial representation data of the reference video.

[0099] The facial adjustment representation data corresponding to the driving audio refers to the data required for adjusting the identifying facial representation data above, and the embodiment of the present disclosure does not limit the implementation of the facial adjustment representation data. For example, when the identifying facial representation data is implemented by using three-dimensional facial key points, the facial adjustment representation data can refer to the offset of the three-dimensional facial key points corresponding to the driving audio relative to the identifying facial representation data, so that the three-dimensional facial key points corresponding to the driving audio can be determined based on the offset and the identifying facial representation data subsequently.

[0100] In addition, the embodiment of the present disclosure does not limit the implementation of step 11 above, for example, in order to better improve the expression determination effect, step 11 can be specifically: predicting the facial adjustment representation data corresponding to the driving audio according to the speaking characteristic representation data of the reference video, the audio features of the driving audio, and the encoding features of the identifying facial representation data of the reference video. For another example, step 11 can be implemented by using a pre-constructed machine learning model with offset prediction function of three-dimensional facial key points, such as MLP.

[0101] Step 12: adjusting the identifying facial representation data of the reference video according to the facial adjustment representation data corresponding to the driving audio to obtain the expression state representation data corresponding to the driving audio.

[0102] Based on the related content of steps 11 to 12 above, in some application scenarios, the process of obtaining the expression state representation data corresponding to the driving audio can be: first, predicting the offset of the three-dimensional facial key points corresponding to the driving audio relative to the identifying facial representation data above; then, adding the offset and the identifying facial representation data to obtain the expression state representation data corresponding to the driving audio, so that the expression state representation data can represent the absolute coordinates of the three-dimensional facial key points corresponding to the driving audio.

[0103] Case 3, in some application scenarios, the expression coefficient, such as blendshape, corresponding to the driving audio can be directly predicted. Based on this, the embodiment of the present disclosure also provides a possible implementation of S3 above, in which the S3 can specifically include steps 21-22 below.

[0104] Step 21: predicting the expression coefficient corresponding to the driving audio according to the speaking characteristic representation data of the reference video, the audio features of the driving audio, and the identifying facial representation data of the reference video.

[0105] The expression coefficient corresponding to the driving audio is used to represent the expression state of the object in the reference video under the driving audio; and the embodiment of the present disclosure does not limit the implementation of the expression coefficient, for example, it can be implemented by using any existing or future expression coefficient, such as blendshape.

[0106] In addition, the embodiment of the present disclosure does not limit the implementation of step 21 above, for example, in order to better improve the prediction effect, step 21 can be specifically: predicting the expression coefficient corresponding to the driving audio according to the speaking characteristic representation data of the reference video, the audio features of the driving audio, and the encoding features of the identifying facial representation data of the reference video. For another example, step 21 can be implemented by using a pre-constructed machine learning model with expression coefficient prediction function, such as MLP.

[0107] Step 22: generating the expression state representation data corresponding to the driving audio according to the expression coefficient corresponding to the driving audio.

[0108] It should be noted that the embodiments of the present disclosure do not limit the implementation of step 22 above, for example, it can be implemented by using any one of the existing or future methods that can convert the expression coefficient into the three-dimensional face key point.

[0109] Based on the related content of steps 21 to 22 above, in some application scenarios, the process of obtaining the expression state representation data corresponding to the driving audio can be: first, predicting the expression coefficient corresponding to the driving audio; and then converting the expression coefficient into the three-dimensional face key point to obtain the expression state representation data corresponding to the driving audio, so that the expression state representation data can represent the absolute coordinates of the three-dimensional face key point corresponding to the driving audio.

[0110] Based on the related content of S1 to S3 above, for the data processing method provided by the embodiments of the present disclosure, after obtaining the reference video and the driving audio, such as the i-th frame of audio in a certain audio segment, first, the expression change representation data of the reference video is analyzed to obtain the speaking feature representation data of the reference video, so that the speaking feature representation data can represent the speaking features presented in the reference video, such as the large opening amplitude of the mouth when speaking, etc. Then, the expression state representation data corresponding to the driving audio is predicted according to the speaking feature representation data, the audio features of the driving audio, and the identifying face representation data of the reference video, so that the expression state representation data is used to represent the lip shape state of the object in the reference video under the driving audio.

[0111] Among them, because the expression change representation data of the reference video is used to represent the expression change situation presented in the reference video, so that the expression change representation data can represent the speaking features possessed by the object in the reference video to a certain extent, such as the large opening amplitude of the mouth when speaking, etc., so that the speaking feature representation data obtained by analyzing the expression change representation data can more accurately represent the speaking features presented in the reference video, and then the expression state representation data predicted based on the speaking feature representation data conforms to the speaking features presented in the reference video, so that the expression state representation data can more accurately represent the expression state of the object in the reference video under the driving audio, such as the lip shape state, etc., thereby facilitating to improve the expression determination effect.

[0112] In addition, the identification facial feature data of the reference video is used to represent some identification facial features of the object in the reference video, such as facial contour features, facial feature distribution features, and the like, so that the expression state feature data predicted based on the identification facial feature data conforms to the identification facial features, thereby enabling the expression state feature data to more accurately represent the expression state of the object under the driving audio, such as the mouth shape state, and thus facilitating improvement of the expression determination effect.

[0113] In addition, the data processing method provided by the embodiments of the present disclosure is not limited to the execution subject, for example, the data processing method provided by the embodiments of the present disclosure can be applied to a terminal device or a server. For another example, the data processing method provided by the embodiments of the present disclosure can also be implemented by means of a data interaction process between a terminal device and a server. The terminal device can be a smart phone, a computer, a personal digital assistant (PDA), a tablet computer, and the like. The server can be a stand-alone server, a cluster server, or a cloud server.

[0114] In fact, in some application scenarios, such as a mouth shape modification scenario, in order to better meet the requirement that the driving audio only affects the expression state, the embodiments of the present disclosure further provide a possible implementation of the above data processing method, in which, when the above driving audio refers to any one frame of audio in the audio sequence, and the driving audio corresponds to a target image in the reference video, the data processing method not only includes S1-S3, but also includes the following step 31. The execution time of step 31 is later than the execution time of S3.

[0115] Step 31: updating the expression state feature data corresponding to the driving audio according to the three-dimensional facial key points corresponding to the target image under the pose parameter; the updated expression state feature data represents a pose state consistent with the pose state presented in the target image.

[0116] The target image refers to an image in the reference video corresponding to the driving audio. For example, when the total number of frames of the audio sequence is equal to the total number of frames of the reference video, if the driving audio is the i-th frame of audio in the audio sequence, the target image can be the i-th frame of image in the reference video, so that the subsequent expression determination processing for the driving audio can be used to modify the mouth shape presented in the i-th frame of image to the mouth shape adapted to the driving audio, and keep other facial states, such as facial identification information, pose, and the like, unchanged except the mouth shape.

[0117] In addition, for the target image, the three-dimensional face key points corresponding to the target image under the posture parameter are used to represent the posture of the face presented in the target image, such as a front face, a side face, and the like.

[0118] In addition, the embodiments of the present disclosure do not limit the implementation of step 31, for example, when the expression state representation data corresponding to the driving audio includes the three-dimensional face key points corresponding to the driving audio under the identity parameter and the three-dimensional face key points corresponding to the driving audio under the expression parameter, step 31 can be specifically: updating the expression state representation data corresponding to the driving audio according to the three-dimensional face key points corresponding to the target image under the posture parameter, so that the updated expression state representation data includes the three-dimensional face key points corresponding to the target image under the posture parameter, so that the posture state represented by the updated expression state representation data is consistent with the posture state presented in the target image, and then the face state represented by the updated expression state representation data is consistent with the face state presented in the target image in other aspects except for the lip shape.

[0119] Based on the related content of step 31, for some application scenarios, after predicting the expression state representation data corresponding to the driving audio, if the expression state representation data includes the three-dimensional face key points corresponding to the driving audio under the identity parameter and the three-dimensional face key points corresponding to the driving audio under the expression parameter, the expression state representation data can more accurately represent the lip shape of the object in the reference image adapted to the driving audio, so as to better achieve the requirement of modifying only the lip shape state in these scenarios., you can first determine the image corresponding to the driving audio from the reference video as a target image; then update the expression state representation data corresponding to the driving audio according to the three-dimensional face key points corresponding to the target image under the posture parameter, so that the updated expression state representation data can include the three-dimensional face key points corresponding to the driving audio under the identity parameter, the three-dimensional face key points corresponding to the driving audio under the expression parameter, and the three-dimensional face key points corresponding to the target image under the posture parameter, so that the face state represented by the updated expression state representation data is consistent with the face state presented in the target image in other aspects except for the lip shape. This can better meet the lip shape modification requirements of these scenarios, thereby facilitating the improvement of the lip shape modification effect for the reference video.

[0120] In fact, in some application scenarios, such as lip shape modification for videos and the like, the embodiments of the present disclosure also provide a possible implementation of the above data processing method, in which the data processing method can include the following steps 41-45.

[0121] Step 41: Obtain a reference video and a driving audio. The driving audio refers to any one frame of audio in the audio sequence.

[0122] It should be noted that the related content of step 41 can be referred to the related content of S1.

[0123] Step 42: Obtain speaking characteristic representation data of the reference video according to expression change representation data of the reference video. The driving audio refers to any one frame of audio in the audio sequence.

[0124] It should be noted that the related content of step 42 can be referred to the related content of S2.

[0125] Based on the related content of steps 41-42, in some scenarios, when the total number of frames of the reference video is equal to the total number of frames of the audio sequence, the i-th frame of audio in the audio sequence can be regarded as the driving audio, and the speaking characteristic representation data of the reference video can be obtained according to the expression change representation data of the reference video, so that the speaking characteristic representation data can represent the speaking characteristics presented by the reference video, so as to subsequently predict the corresponding mouth shape state of the driving audio based on the speaking characteristics. Wherein, i is a positive integer, i≤I.

[0126] Step 43: Predict the expression state representation data corresponding to the driving audio according to the speaking characteristic representation data of the reference video, the audio feature of the driving audio, and the identifying facial representation data of the reference video. The expression state representation data is used to represent the mouth shape of the object in the reference video adapted to the driving audio. The driving audio refers to any one frame of audio in the audio sequence.

[0127] It should be noted that the related content of step 43 can be referred to the related content of S3.

[0128] Based on the related content of step 43, in some scenarios, when the total number of frames of the reference video is equal to the total number of frames of the audio sequence, after regarding the i-th frame of audio in the audio sequence as the driving audio, the expression state representation data corresponding to the driving audio can be predicted according to the speaking characteristic representation data of the reference video, the audio feature of the driving audio, and the identifying facial representation data of the reference video, so that the expression state representation data can represent the predicted expression state of the object in the reference video under the i-th frame of audio, such as mouth shape state, so as to subsequently complete the mouth shape modification process for the i-th frame of image in the reference video based on the expression state representation data. Wherein, i is a positive integer, i≤I.

[0129] Step 44: Perform rendering processing according to the expression state representation data corresponding to the driving audio to obtain the generated image corresponding to the driving audio. The driving audio refers to any one frame of audio in the audio sequence.

[0130] wherein the generated image corresponding to the driving audio refers to an image generated according to the driving audio; and the generated image at least satisfies the following constraint: the lip shape state presented in the generated image is consistent with the lip shape state represented by the expression state representation data corresponding to the driving audio.

[0131] In addition, in some application scenarios, when the total number of frames of the reference video is equal to the total number of frames of the audio sequence, if the driving audio is the i-th frame of audio in the audio sequence, the generated image corresponding to the driving audio can satisfy the following constraint: the lip shape state presented in the generated image is consistent with the lip shape state represented by the expression state representation data corresponding to the driving audio; and the other states presented in the generated image, such as facial identification information, posture, etc., are consistent with the corresponding states presented in the i-th frame of image in the reference video.

[0132] In addition, the embodiments of the present disclosure do not limit the implementation of the above step 44, for example, it can be implemented by means of any one of the existing or future methods capable of generating images based on expression state representation data, such as pre-constructed machine learning models for rendering three-dimensional facial key points into images.

[0133] Based on the related content of the above step 44, in some scenarios, when the total number of frames of the reference video is equal to the total number of frames of the audio sequence, after predicting the expression state representation data corresponding to the i-th frame of audio in the audio sequence, the i-th frame of audio corresponding to the expression state representation data can be rendered to obtain the generated image corresponding to the i-th frame of audio, so that the lip shape state presented in the generated image is consistent with the lip shape state described by the expression state representation data corresponding to the i-th frame of audio, and the other states presented in the generated image are consistent with the corresponding states presented in the i-th frame of image in the reference video, so as to realize the lip shape modification processing of the i-th frame of audio according to the i-th frame of image in the reference video. Wherein i is a positive integer, i≤I.

[0134] Step 45: constructing a generated video corresponding to the audio sequence according to the generated images corresponding to each frame of audio in the audio sequence; the generated video is used to describe the lip shape change state of the object in the reference video under the audio sequence.

[0135] In the embodiments of the present disclosure, when the total number of frames of the reference video is equal to the total number of frames of the audio sequence, after obtaining the generated images corresponding to the audio of each frame in the audio sequence, the generated video corresponding to the audio sequence can be constructed according to the generated images corresponding to the audio, so that the generated video includes the generated images, so that the lip movement change presented by the generated video is consistent with the lip movement change required by the audio sequence, and the change presented by the generated video in aspects other than the lip movement is consistent with the change presented by the corresponding aspect in the reference video. Thus, the lip modification of the reference video according to the audio sequence can be realized.

[0136] Based on the related content of steps 41 to 45 above, in some application scenarios, the data processing method provided by the embodiments of the present disclosure can better realize the lip modification of a video according to an audio, so that the modified video is consistent with the audio in terms of lip movement, and the modified video is consistent with the video in terms of aspects other than lip movement. Thus, the lip modification demand can be better met, thereby facilitating to improve the lip modification effect of the video.

[0137] In fact, in order to better improve the expression determination effect, the embodiments of the present disclosure also provide a possible implementation of the above data processing method, in which the data processing method can be implemented by means of a lip driving model. The lip driving model has a lip driving function, and the lip driving model includes an audio encoding module, a face encoding module, a speaking feature analysis module, and an expression prediction module. These modules will be introduced respectively.

[0138] For the audio encoding module, such as the encoding module 1 shown in FIG. 3, the audio encoding module is used for audio encoding processing on the input data of the audio encoding module. Moreover, the embodiments of the present disclosure do not limit the implementation of the audio encoding module, for example, the audio encoding module can be implemented by a transformer. In addition, the input data of the audio encoding module can be an audio signal, such as the driving audio shown in FIG. 3, and the output data of the audio encoding module is an audio feature corresponding to the audio signal, such as a feature vector. Furthermore, the embodiments of the present disclosure do not limit the working principle of the audio encoding module, for example, the audio encoding module can be used for encoding processing on the driving audio to obtain the audio feature of the driving audio.

[0139] For the face encoding module, such as the encoding module 2 shown in FIG. 3, the face encoding module is configured to perform face encoding processing on input data of the face encoding module; and the embodiments of the present disclosure do not limit the implementation of the face encoding module, for example, the face encoding module can be implemented by using MLP. In addition, the input data of the face encoding module can be three-dimensional face key points, such as the expression change representation data of the reference video or the identifying face representation data of the reference video; and the output data of the face encoding module can be the encoded features of the three-dimensional face key points, such as feature vectors. In addition, the embodiments of the present disclosure do not limit the working principle of the face encoding module, for example, the face encoding module is configured to encode the expression change representation data to obtain the encoded features of the expression change representation data, and encode the identifying face representation data to obtain the encoded features of the identifying face representation data.

[0140] For the speaking feature analysis module, such as the speaking feature analysis module shown in FIG. 3, the speaking feature analysis module is configured to perform speaking feature analysis processing on input data of the speaking feature analysis module; and the embodiments of the present disclosure do not limit the implementation of the speaking feature analysis module, for example, the speaking feature analysis module can be implemented by using CNN+MLP. In addition, the input data of the speaking feature analysis module is the encoded features of the three-dimensional face key points, and the output data of the speaking feature analysis module is a statistical value under a continuous distribution. In addition, the embodiments of the present disclosure do not limit the working principle of the speaking feature analysis module, for example, the speaking feature analysis module can be configured to analyze the encoded features of the expression change representation data to obtain the speaking feature representation data, such as determining the statistical value corresponding to the encoded features from the continuous distribution as the speaking feature representation data.

[0141] For the expression prediction module, such as the expression prediction module shown in FIG. 3, the expression prediction module is used for expression prediction processing for input data of the expression prediction module, such as absolute coordinate prediction of three-dimensional face key points, relative offset prediction of three-dimensional face key points, or expression coefficient prediction, etc. Moreover, embodiments of the present disclosure do not limit the implementation of the expression prediction module, for example, the expression prediction module can be implemented by using MLP. In addition, the input data of the expression prediction module can include the output data of multiple encoding modules, and the output data of the expression prediction module can be absolute coordinates of three-dimensional face key points, relative offsets of three-dimensional face key points, or expression coefficients, etc. Furthermore, embodiments of the present disclosure do not limit the working principle of the expression prediction module, for example, the expression prediction module can be used to predict the expression state representation data corresponding to the driving audio according to the above speech characteristic representation data, the audio features of the above driving audio, and the encoding features of the above identifying facial representation data. As another example, the expression prediction module can be used to predict the facial adjustment representation data corresponding to the driving audio according to the speech characteristic representation data, the audio features of the driving audio, and the encoding features of the identifying facial representation data, so that subsequent adjustment processing can be performed on the identifying facial representation data based on the facial adjustment representation data to obtain the expression state representation data corresponding to the driving audio. As another example, the expression prediction module can be used to predict the expression coefficient corresponding to the driving audio according to the speech characteristic representation data, the audio features of the driving audio, and the encoding features of the identifying facial representation data, so that subsequent expression state representation data corresponding to the driving audio can be generated according to the expression coefficient.

[0142] Based on the above related content of the lip driving model, in one possible implementation, the working principle of the lip driving model can be as follows: after the driving audio, the expression change representation data of the reference video, and the identifying facial representation data of the reference video are input into the lip driving model, first, the audio encoding module in the lip driving model extracts the audio features of the driving audio, and the facial encoding module in the lip driving model extracts the encoding features of the expression change representation data and the encoding features of the identifying facial representation data. Then, the speech characteristic analysis module in the lip driving model analyzes the speech characteristic representation data of the reference video from the encoding features of the expression change representation data. Finally, after the audio features, the speech characteristic representation data, and the encoding features of the identifying facial representation data are spliced, the expression prediction module in the lip driving model predicts the expression state representation data corresponding to the driving audio according to the splicing result, so that the expression determination processing for the driving audio can be completed by means of the lip driving model.

[0143] In addition, when the data processing method provided by the embodiment of the present disclosure is used to implement the model training process, the data processing method can include the following steps 51-53.

[0144] Step 51: Obtain a reference video, a driving audio, and an expression state label corresponding to the driving audio. The driving audio refers to an audio frame extracted from a sample video; the reference video is determined according to the sample video; and the expression state label is determined according to an image corresponding to the driving audio in the sample video.

[0145] The driving audio refers to an audio required to be used in the current round, and the embodiment of the present disclosure does not limit the acquisition method of the driving audio. For example, for the current round, an audio frame not participating in the training process is randomly selected from the sample video as the driving audio.

[0146] In addition, when the driving audio refers to an audio frame extracted from a sample video, the reference video can be determined according to the sample video, so that the reference video includes part or all of the images in the sample video.

[0147] In addition, for the driving audio, the expression state label corresponding to the driving audio refers to the true value of the expression state corresponding to the driving audio; and when the driving audio refers to an audio frame extracted from a sample video, the expression state label can be determined according to an image corresponding to the driving audio in the sample video. The image corresponding to the driving audio in the sample video refers to an image in the sample video that has a corresponding relationship with the driving audio, so that the image corresponding to the driving audio in the sample video can represent the expression state actually under the driving audio. For example, when the driving audio refers to the i th audio frame in the sample video, the image corresponding to the driving audio in the sample video can refer to the i th image in the sample video. Wherein, i is a positive integer, i≤I.

[0148] In addition, the embodiment of the present disclosure does not limit the determination process of the expression state label corresponding to the driving audio. For example, it can specifically be: first, obtain the three-dimensional face parameters of the image corresponding to the driving audio in the sample video, and the three-dimensional face parameters include identification parameters, expression parameters and posture parameters; and then, according to the three-dimensional face parameters of the image, construct the expression state label corresponding to the driving audio.

[0149] Further, the embodiment of the present disclosure does not limit the implementation of the step of "constructing the expression state label corresponding to the driving audio according to the three-dimensional facial parameters of the image" in the above paragraph. For example, when the determination process of the expression state representation data corresponding to the driving audio does not include the step 31, the step can be specifically: constructing the expression state label corresponding to the driving audio according to the three-dimensional facial key points of the image corresponding to the expression parameters and the three-dimensional facial key points of the image corresponding to the identity parameters, so that the expression state label includes the three-dimensional facial key points under the two parameters. For another example, when the determination process of the expression state representation data corresponding to the driving audio includes the step 31, the step can be specifically: constructing the expression state label corresponding to the driving audio according to the three-dimensional facial key points of the image corresponding to the expression parameters, the three-dimensional facial key points of the image corresponding to the identity parameters, and the three-dimensional facial key points of the image corresponding to the posture parameters, so that the expression state label includes the three-dimensional facial key points under the three parameters.

[0150] Based on the related content of the step 51, for the current round, part or all of the sample video is regarded as the reference video, one frame of audio in the sample video that has not participated in the training is regarded as the driving audio, and the expression state label corresponding to the driving audio is determined according to the image corresponding to the driving audio in the sample video, so that the subsequent training of the current round can be completed based on the three kinds of data.

[0151] Step 52: predicting, by the lip-driven model, the expression state representation data corresponding to the driving audio according to the driving audio, the expression change representation data of the reference video, and the identity facial representation data of the reference video.

[0152] It should be noted that the implementation of the step 52 is similar to the implementation of the determination process of the expression state representation data shown in the above.

[0153] Based on the related content of the step 52, for the current round, after the reference video and the driving audio are obtained, the expression state representation data corresponding to the driving audio can be predicted by the lip-driven model according to the driving audio, the expression change representation data of the reference video, and the identity facial representation data of the reference video, so that the performance of the lip-driven model in the current round can be determined based on the expression state representation data in the subsequent process.

[0154] Step 53: updating the lip-driven model according to the expression state representation data corresponding to the driving audio and the expression state label corresponding to the driving audio, and returning to continue to execute the step 51 and the subsequent steps until a preset stop condition is reached.

[0155] The preset stop condition can be set according to an actual application scenario. For example, the preset stop condition can include that the loss value of the lip-driven model is lower than a pre-set loss threshold. For another example, the preset stop condition can include that a change rate of the loss value of the lip-driven model is lower than a pre-set change rate threshold. For yet another example, the preset stop condition can include that the number of updates of the lip-driven model reaches a pre-set number threshold.

[0156] The loss value of the lip-driven model is used to represent the performance of the lip-driven model, and is determined according to the gap between the expression state representation data corresponding to the driving audio and the expression state label corresponding to the driving audio. It should be noted that the disclosure does not limit the determination process of the loss value of the lip-driven model.

[0157] In addition, the disclosure does not limit the implementation of the "updating the lip-driven model" in step 53 above. For example, in some application scenarios, it can be that all modules of the lip-driven model are updated. For another example, in some application scenarios, it can be that part of the modules of the lip-driven model are updated, such as the face encoding module, the speaking feature analysis module, and the expression prediction module. It can be seen that in a possible implementation, for the lip-driven model, if the audio encoding module in the lip-driven model is a pre-trained module with good audio feature extraction function, only the other modules in the lip-driven model except the audio encoding module need to be updated when updating the lip-driven model, and the audio encoding module does not need to be updated, which is conducive to improving the model training effect.

[0158] Based on the related content of steps 51 to 53 above, it can be seen that the data processing method provided by the disclosure can be used to implement the training processing of the lip-driven model, so that the trained lip-driven model has good performance, thereby presenting good effect when the lip-driven model is used to perform some expression determination tasks, and thus being conducive to improving the expression determination effect.

[0159] It should be noted that the disclosure does not limit the implementation of the sample video. For example, in order to better ensure that the speaking feature analysis module in the lip-driven model can learn as accurate continuous distribution as possible, the sample video configured for the iterative training process of the lip-driven model should cover as many videos as possible for presenting various speaking features, so that the continuous distribution learned by the speaking feature analysis module in the lip-driven model trained based on these videos can be more accurate, thereby being conducive to improving the subsequent expression determination effect.

[0160] It is also to be noted that, through research, it is found that the way of directly predicting the relative offset by using the mouth shape driving model converges faster and is more robust than the way of directly predicting the absolute coordinates of three-dimensional face key points or expression coefficients by using the mouth shape driving model.

[0161] Based on the data processing method provided by the embodiments of the present disclosure, the embodiments of the present disclosure also provide a data processing device, which is explained and described below in combination with FIG. 4. FIG. 4 is a structural schematic diagram of a data processing device provided by the embodiments of the present disclosure. It should be noted that the technical details of the data processing device provided by the embodiments of the present disclosure are referred to the related content of the data processing method above.

[0162] As shown in FIG. 4, the data processing device 400 provided by the embodiments of the present disclosure includes:

[0163] The acquisition unit 401 is configured to acquire a reference video and a driving audio;

[0164] The analysis unit 402 is configured to obtain speaking feature representation data of the reference video according to expression change representation data of the reference video.

[0165] The prediction unit 403 is configured to predict expression state representation data corresponding to the driving audio according to the speaking feature representation data, an audio feature of the driving audio, and identity face representation data of the reference video; the expression state representation data is used to represent a mouth shape state of an object in the reference video under the driving audio.

[0166] In a possible implementation, the speaking feature representation data is determined from a continuous distribution according to an encoded feature of the expression change representation data; the continuous distribution is used to describe the speaking feature representation data corresponding to different encoded features.

[0167] In a possible implementation, the continuous distribution is learned from a plurality of sample videos; different sample videos present different speaking features.

[0168] In a possible implementation, for any one frame of image in the reference video, three-dimensional face parameters of the image include identity parameters, expression parameters and posture parameters.

[0169] The expression change representation data includes three-dimensional face key points corresponding to each frame of image in the reference video under the expression parameters.

[0170] The identity face representation data is determined according to three-dimensional face key points corresponding to at least one frame of image in the reference video under the identity parameters.

[0171] In a possible implementation, the driving audio refers to any one frame of audio in the audio sequence; and the driving audio corresponds to a target image in the reference video.

[0172] The data processing apparatus 400 further includes:

[0173] The data updating unit is configured to update the expression state representation data corresponding to the driving audio according to the three-dimensional face key points corresponding to the target image under the posture parameter after predicting the expression state representation data corresponding to the driving audio; and the updated expression state representation data represents a posture state consistent with the posture state represented in the target image.

[0174] In a possible implementation, the prediction unit 403 is specifically configured to: predict the face adjustment representation data corresponding to the driving audio according to the speaking feature representation data, the audio feature, and the identifying face representation data; and adjust the identifying face representation data according to the face adjustment representation data to obtain the expression state representation data corresponding to the driving audio.

[0175] In a possible implementation, the prediction unit 403 is specifically configured to: predict an expression coefficient corresponding to the driving audio according to the speaking feature representation data, the audio feature, and the identifying face representation data; and generate the expression state representation data corresponding to the driving audio according to the expression coefficient.

[0176] In a possible implementation, the driving audio refers to any one frame of audio in the audio sequence;

[0177] The data processing apparatus 400 further includes:

[0178] The generation unit is configured to perform rendering processing on the expression state representation data corresponding to the driving audio to obtain a generated image corresponding to the driving audio.

[0179] The construction unit is configured to construct a generated video corresponding to the audio sequence according to the generated images corresponding to the frames of audio in the audio sequence; and the generated video is used to describe the lip movement change state of the object under the audio sequence.

[0180] In a possible implementation, the expression state representation data corresponding to the driving audio is determined by using a lip movement driving model; the driving audio refers to a frame of audio extracted from a sample video; and the reference video is determined according to the sample video.

[0181] The data processing apparatus 400 further includes:

[0182] a model updating unit configured to update the lip driving model according to the expression state representation data corresponding to the driving audio and an expression state label corresponding to the driving audio, the expression state label being determined according to an image corresponding to the driving audio in the sample video.

[0183] In a possible implementation, the expression state representation data corresponding to the driving audio is determined by using the lip driving model.

[0184] The lip driving model comprises an audio encoding module, a face encoding module, a speaking feature analysis module and an expression prediction module.

[0185] The audio encoding module is configured to encode the driving audio to obtain audio features of the driving audio.

[0186] The face encoding module is configured to encode the expression change representation data to obtain encoded features of the expression change representation data, and encode the identifying face representation data to obtain encoded features of the identifying face representation data.

[0187] The speaking feature analysis module is configured to analyze the encoded features of the expression change representation data to obtain the speaking feature representation data.

[0188] The expression prediction module is configured to predict the expression state representation data corresponding to the driving audio according to the speaking feature representation data, the audio features and the encoded features of the identifying face representation data.

[0189] Based on the related content of the above data processing apparatus 400, for the data processing apparatus 400 provided by the embodiments of the present disclosure, after obtaining the reference video and the driving audio, such as the i-th frame of audio in a certain audio segment, first, the speaking characteristic representation data of the reference video is obtained according to the expression change representation data of the reference video, so that the speaking characteristic representation data can represent the speaking characteristics presented in the reference video, such as the characteristic that the mouth opening amplitude is relatively large when speaking; then, the expression state representation data corresponding to the driving audio is predicted according to the speaking characteristic representation data, the audio features of the driving audio, and the identifying facial representation data of the reference video, so that the expression state representation data is used to represent the lip shape state of the object in the reference video under the driving audio. Among them, because the expression change representation data of the reference video is used to represent the expression change situation presented in the reference video, so that the expression change representation data can represent the speaking characteristics possessed by the object in the reference video to a certain extent, such as the characteristic that the mouth opening amplitude is relatively large when speaking, so that the speaking characteristic representation data analyzed based on the expression change representation data can more accurately represent the speaking characteristics presented in the reference video, and then the expression state representation data predicted based on the speaking characteristic representation data conforms to the speaking characteristics presented in the reference video, so that the expression state representation data can more accurately represent the expression state of the object in the reference video under the driving audio, such as the lip shape state, thereby facilitating to improve the expression determination effect. Also, because the identifying facial representation data of the reference video is used to represent some identifying facial characteristics possessed by the object in the reference video, such as facial contour characteristics, facial feature distribution characteristics and the like, so that the expression state representation data predicted based on the identifying facial representation data conforms to these identifying facial characteristics, so that the expression state representation data can more accurately represent the expression state of the object under the driving audio, such as the lip shape state, so as to facilitate to improve the expression determination effect.

[0190] In addition, the embodiments of the present disclosure also provide an electronic device, which comprises a processor and a memory: the memory is used to store instructions or computer programs; the processor is used to execute the instructions or computer programs in the memory, so that the electronic device executes any implementation manner of the data processing method provided by the embodiments of the present disclosure.

[0191] Referring to FIG. 5, a structural diagram of an electronic device 500 suitable for implementing embodiments of the disclosure is illustrated. The terminal device in embodiments of the disclosure can include, but is not limited to, a mobile terminal such as a mobile phone, a notebook computer, a digital broadcast receiver, a PDA (Personal Digital Assistant), a PAD (Tablet Personal Computer), a PMP (Portable Multimedia Player), a car terminal (e.g., a car navigation terminal), and the like, as well as a stationary terminal such as a digital TV, a desktop computer, and the like. The electronic device illustrated in FIG. 5 is merely an example, and should not impose any limitation on the functions and use range of embodiments of the disclosure.

[0192] As illustrated in FIG. 5, the electronic device 500 can include a processing device (e.g., a central processor, a graphic processor, etc.) 501 that can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded into a random access memory (RAM) 503 from a storage device 508. In the RAM 503, various programs and data required for the operation of the electronic device 500 are also stored. The processing device 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0193] Generally, the following devices can be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; a storage device 508 including, for example, a magnetic tape, a hard disk, and the like; and a communication device 509. The communication device 509 can allow the electronic device 500 to communicate with other devices wirelessly or wired to exchange data. Although FIG. 5 illustrates the electronic device 500 having various devices, it should be understood that all of the illustrated devices are not required to be implemented or possessed. More or less devices can be alternatively implemented or possessed.

[0194] In particular, according to embodiments of the disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the disclosure include a computer program product including a computer program carried on a non-transitory computer readable medium, the computer program containing program code for executing the methods illustrated in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above-described functions defined in the methods of embodiments of the disclosure are performed.

[0195] The electronic device provided by the embodiments of the present disclosure and the method provided by the above embodiments belong to the same inventive concept, and the technical details not described in detail in the present embodiment can be referred to the above embodiments, and the present embodiment has the same beneficial effects as the above embodiments.

[0196] The embodiments of the present disclosure also provide a computer readable medium, wherein instructions or computer programs are stored in the computer readable medium, and when the instructions or computer programs are run on a device, the device is caused to execute any implementation of the data processing method provided by the embodiments of the present disclosure.

[0197] It should be noted that the computer readable medium of the present disclosure described above can be a computer readable signal medium or a computer readable storage medium or any combination of the above two. The computer readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or component. In the present disclosure, the computer readable signal medium can include a data signal carried in a baseband or as a part of a carrier wave, which carries computer readable program code. Such a propagated data signal can take many forms, including but not limited to an electromagnetic signal, an optical signal or any suitable combination of the above. The computer readable signal medium can also be any computer readable medium other than the computer readable storage medium, which can send, propagate or transmit a program for use by or in conjunction with an instruction execution system, device or component. The program code contained in the computer readable medium can be transmitted by any suitable medium, including but not limited to an electrical wire, an optical cable, an RF (radio frequency) or the like, or any suitable combination of the above.

[0198] In some embodiments, the client, server, or other computing machines utilized by the system can communicate information using any known or future developed end-to-end communication protocol, such as the Hyper Text Transfer Protocol (HTTP), and can be interconnected via any form or medium of digital data communication (for example, a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), the Internet, and peer-to-peer networks (for example, ad hoc peer-to-peer networks), as well as any current or future developed network.

[0199] The computer-readable medium described above can be included in the electronic device described above; alternatively, the computer-readable medium can exist as a standalone entity.

[0200] The computer-readable medium described above can be included in the electronic device described above; alternatively, the computer-readable medium can exist as a standalone entity.

[0201] Computer program code for carrying out operations of the present disclosure can be written in any one or combination of one or more programming languages or combinations thereof, including an object oriented programming language such as Java, Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network ("LAN") or a wide area network ("WAN"), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0202] The computer program product of the first aspect can include one or more non-transitory computer-readable media storing instructions that, when executed, cause one or more processors to perform the operations of the first aspect. The computer program product of the first aspect can include a non-transitory computer-readable medium storing code that, when executed, causes a computer to perform operations for the first aspect.

[0203] The units described in the embodiments of the present disclosure can be implemented by software, or by hardware, or by a combination of software and hardware. In some cases, the names of the units / modules do not constitute a limitation on the units themselves.

[0204] The functions described in this description above can be implemented in hardware, software, or any combination thereof. If implemented in software, the functions can be stored on or transmitted over as one or more instructions or code on a computer-readable medium. Computer-readable media include both computer storage media and communication media including any medium that facilitates transfer of a computer program from one place to another. A storage media can be any available media that can be accessed by a general purpose or special purpose computer. By way of example, and not limitation, such computer-readable media can comprise RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code means in the form of instructions or data structures and that can be accessed by a general-purpose or special-purpose computer, or a general-purpose or special-purpose processor. Also, functional computer- readable media that can store program code, when executed by a processor, software, or hardware components such as software agents, can direct computing equipment such as computers, server devices, or other hardware components to function in a particular manner, such that the functional computer-readable media can be broadly described as a processor, software, or hardware component that includes the instructions for implementing a function described herein and / or one or more signals that can be used by a processor, software, or hardware component to carry out a function described herein. Computer program or application code in the present context can mean any instructions, code or data that, when executed, cause a processor, software, or hardware component to perform a function. The instructions, code or data can be stored on a non-transitory computer-readable medium or transmitted from a computer program product or a signal.

[0205] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store program code for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium can include a tangible, non-transitory computer-readable storage medium that can be based on one or more lines of code, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0206] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the systems or devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.

[0207] It should be understood that in the embodiments of the present disclosure, "at least one (item)" refers to one or more, and "plurality" refers to two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can represent: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0208] It should also be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.

[0209] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0210] The above description of disclosed embodiments enables a person skilled in the art to implement or use the disclosure embodiments. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other embodiments without departing from the spirit or scope of the disclosure embodiments. Thus, the disclosure embodiments are not to be limited to the embodiments shown herein but are to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A data processing method, wherein: The method comprises: Get reference video and driving audio; Obtaining speech characteristic representation data of the reference video based on the expression change representation data of the reference video; Based on the speech characteristic representation data, the audio features of the driving audio, and the identifying facial representation data of the reference video, the expression state representation data corresponding to the driving audio is predicted; the expression state representation data is used to represent the lip shape state of the object in the reference video under the driving audio.

2. The method according to claim 1, wherein The speech characteristic representation data is determined from a continuous distribution based on the coding features of the expression change representation data; the continuous distribution is used to describe the speech characteristic representation data corresponding to different coding features.

3. The method according to claim 2, wherein: The continuous distribution is learned from multiple sample videos; different sample videos present different speech characteristics.

4. The method according to claim 1, wherein For any frame image in the reference video, the three-dimensional facial parameters of the image include identification parameters, expression parameters and posture parameters; The expression change representation data includes three-dimensional facial key points corresponding to each frame image in the reference video under the expression parameters; The identifying facial representation data is determined based on three-dimensional facial key points corresponding to at least one frame of image in the reference video under the identifying parameters.

5. The method according to claim 4, wherein The driving audio refers to any frame of audio in the audio sequence; The driving audio corresponds to the target image in the reference video; After predicting the facial expression state representation data corresponding to the driving audio, the method further includes: updating the facial expression state representation data corresponding to the driving audio according to the three-dimensional facial key points corresponding to the target image under the posture parameters; The posture state represented by the updated expression state representation data is consistent with the posture state presented in the target image.

6. The method according to claim 1, wherein The process of determining the facial expression state representation data corresponding to the driving audio includes: Predicting facial adjustment representation data corresponding to the driving audio based on the speech characteristic representation data, the audio features, and the indicative facial representation data; The identification facial representation data is adjusted based on the facial adjustment representation data to obtain facial expression state representation data corresponding to the driving audio.

7. The method according to claim 1, wherein The process of determining the facial expression state representation data corresponding to the driving audio includes: Predicting an expression coefficient corresponding to the driving audio based on the speech characteristic representation data, the audio features, and the identifier facial representation data; Generate expression state representation data corresponding to the driving audio according to the expression coefficient.

8. The method according to claim 1, wherein The driving audio refers to any frame of audio in the audio sequence; After predicting the facial expression state representation data corresponding to the driving audio, the method further includes: Performing rendering processing based on the facial expression state representation data corresponding to the driving audio to obtain a generated image corresponding to the driving audio; A generated video corresponding to the audio sequence is constructed based on the generated image corresponding to each frame of audio in the audio sequence; the generated video is used to describe the lip shape change state of the object under the audio sequence.

9. The method according to claim 1, wherein The facial expression state representation data corresponding to the driving audio is determined using a lip-activated model; The driving audio refers to a frame of audio extracted from the sample video; The reference video is determined based on the sample video; After predicting the facial expression state representation data corresponding to the driving audio, the method further includes: updating the lip-sync driving model according to the facial expression state representation data corresponding to the driving audio and the facial expression state label corresponding to the driving audio; The expression state label is determined according to the image corresponding to the driving audio in the sample video.

10. The method according to claim 1, wherein The facial expression state representation data corresponding to the driving audio is determined using a lip-activated model; The lip-activated model includes an audio encoding module, a facial encoding module, a speech characteristics analysis module, and an expression prediction module; The audio encoding module is used to encode the driving audio to obtain audio features of the driving audio; The facial encoding module is used to encode the expression change representation data to obtain encoding features of the expression change representation data, and to encode the identification facial representation data to obtain encoding features of the identification facial representation data; The speech characteristic analysis module is used to analyze the encoding features of the expression change representation data to obtain the speech characteristic representation data; The expression prediction module is used to predict the expression state representation data corresponding to the driving audio based on the speech characteristic representation data, the audio features, and the encoding features of the identification face representation data.

11. A data processing device, wherein: include: An acquisition unit, used for acquiring reference video and driving audio; an analyzing unit, configured to obtain speech characteristic representation data of the reference video based on the expression change representation data of the reference video; A prediction unit is configured to predict facial expression state representation data corresponding to the driving audio based on the speech characteristic representation data, the audio features of the driving audio, and the identifying facial representation data of the reference video; the facial expression state representation data is used to represent the lip shape state of the object in the reference video under the driving audio.

12. An electronic device, wherein: The device includes: a processor and a memory; The memory is used to store instructions or computer programs; The processor is configured to execute the instructions or computer program in the memory, so that the electronic device executes the method according to any one of claims 1 to 10.

13. A computer-readable medium, wherein: The computer-readable medium stores instructions or a computer program, and when the instructions or the computer program are executed on a device, the device is caused to execute the method according to any one of claims 1 to 10.

14. A computer program product, wherein The method comprises a computer program carried on a non-transitory computer-readable medium, the computer program comprising a program code for executing the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Method and device for generating speaking portrait video and training face rendering model

    CN114581980A

  • Expression data generation method and device, readable medium and electronic equipment

    CN114882155A

  • Digital face generation method and device, storage medium and electronic equipment

    CN116129003A

  • Speaking face video generation method, computer equipment and storage medium

    CN117789751A

  • Speaking video generation method and apparatus, and electronic device and storage medium

    WO2023088080A1