Expression migration method and device, electronic equipment and computer readable storage medium

By generating and applying face control parameters, adjusting the vertex position of the target face model to present expression details matching emotional information, solving the problems of high cost, flexibility and scalability in the prior art, and achieving high-fidelity migration of expression details.

CN120088823APending Publication Date: 2025-06-03NETEASE (HANGZHOU) NETWORK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411952170.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The prior art has high cost, low flexibility and scalability in the process of expression migration, and is difficult to capture high fidelity of expression details.

Method used

By obtaining the source face image, source audio and target face model, determine its corresponding expression feature data, audio feature data and identity feature data, generate face control parameters, and adjust the vertex position of the target face model to present expression details matching emotional information.

Benefits of technology

It realizes high fidelity to maintain expression details during expression migration, reduces costs, improves flexibility and scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088823A_ABST
    Figure CN120088823A_ABST
Patent Text Reader

Abstract

The invention provides an expression migration method and apparatus, an electronic device and a computer readable storage medium. The method comprises the steps of obtaining a source face image containing a first expression, a source audio and a target face model to be subjected to expression migration; determining expression feature data corresponding to the source face image, audio feature data corresponding to the source audio and identity feature data corresponding to the target face model; generating a face control parameter corresponding to the target face model according to the expression feature data, the audio feature data and the identity feature data; and according to the face control parameter, adjusting the position of a vertex in the target face model so as to generate an animation presenting a first expression with expression details matched with the emotion information for the target face model. Therefore, high fidelity of expression details can be kept in the expression migration process, the cost of expression migration is effectively reduced, and the flexibility and expandability of expression migration are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to a method, apparatus, electronic device, and computer-readable storage medium for expression transfer. Background Art

[0002] With the rapid development of digital entertainment, virtual reality (VR), augmented reality (AR), and human-computer interaction technologies, 3D facial animation transfer has become increasingly important. This technology aims to capture and reproduce real facial expressions and movements to create realistic and vivid facial animations.

[0003] In related technologies, expression transfer is usually carried out by means of deep learning-based facial reconstruction + redirection. First, a corresponding 3D facial model needs to be reconstructed for the source face. Then, according to the correspondence between the 3D facial model and the facial binding of the target face, a mapping model is established. Then, using the mapping model, the expression of the 3D facial model is transferred to the target face.

[0004] However, this method requires a semantically consistent binding between the target face and the source face, which often relies on a large amount of manual adjustment, resulting in too high a cost for expression transfer and reducing the flexibility and scalability of expression transfer. Summary of the Invention

[0005] This application provides a method, apparatus, electronic device, and computer-readable storage medium for expression transfer, which can effectively reduce the cost of expression transfer while maintaining high-fidelity expression details during the expression transfer process, and improve the flexibility and scalability of expression transfer. The specific solutions are as follows:

[0006] In a first aspect, an embodiment of this application provides a method for expression transfer, the method including:

[0007] Obtain a source face image containing a first expression, source audio, and a target face model to be subjected to expression transfer;

[0008] Determine expression feature data corresponding to the source face image, audio feature data corresponding to the source audio, and identity feature data corresponding to the target face model, where the audio feature data includes emotional information corresponding to the source audio;

[0009] Generate facial control parameters corresponding to the target face model according to the expression feature data, the audio feature data, and the identity feature data;

[0010] Adjust the positions of vertices in the target face model according to the facial control parameters to generate the first expression animation presenting expression details matching the emotional information for the target face model.

[0011] In a second aspect, an embodiment of the present application provides an expression transfer device, which includes:

[0012] An acquisition unit, configured to acquire a source face image including a first expression, a source audio, and a target face model to be subjected to expression transfer;

[0013] A determination unit, configured to determine expression feature data corresponding to the source face image, audio feature data corresponding to the source audio, and identity feature data corresponding to the target face model, where the audio feature data includes emotional information corresponding to the source audio;

[0014] A generation unit, configured to generate face control parameters corresponding to the target face model according to the expression feature data, the audio feature data, and the identity feature data;

[0015] An adjustment unit, configured to adjust the positions of vertices in the target face model according to the face control parameters, so as to generate an animation presenting the first expression with expression details matching the emotional information for the target face model.

[0016] In a third aspect, the present application further provides an electronic device, including:

[0017] A processor; and

[0018] A memory, configured to store a data processing program, and after the electronic device is powered on and runs the program through the processor, the method described in the first aspect is executed.

[0019] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, storing a data processing program, and when the program is run by a processor, the method described in the first aspect is executed.

[0020] Compared with the prior art, the present application has the following advantages:

[0021] The method for expression transfer provided by the embodiments of the present application includes the following steps: obtaining a source face image containing a first expression, a source audio, and a target face model to be subjected to expression transfer; determining expression feature data corresponding to the source face image, audio feature data corresponding to the source audio, and identity feature data corresponding to the target face model, wherein the audio feature data includes emotional information corresponding to the source audio; generating face control parameters corresponding to the target face model according to the expression feature data, the audio feature data, and the identity feature data; and adjusting the positions of vertices in the target face model according to the face control parameters to generate an animation presenting the first expression with expression details matching the emotional information for the target face model. It can be seen that in the present application, first, the expression feature data corresponding to the source face image containing the first expression, the audio feature data corresponding to the source audio, and the identity feature data corresponding to the target face model are determined. The expression feature data is used to represent the expression information corresponding to the first expression, and the audio feature data includes the emotional information corresponding to the source audio. This emotional information can supplement the visual performance details of the face. In this way, through the multi-modal data of the expression feature data and the audio feature data, combined with the identity feature data of the target face model, the generated face control parameters can represent fine-grained expression detail information, thereby driving the target face model to present the first expression with expression details matching the emotional information.

[0022] Therefore, the method for expression transfer provided by the embodiments of the present application can maintain high fidelity of expression details during the expression transfer process. In addition, the method for expression transfer provided by the embodiments of the present application does not require a costly facial capture system to capture the motion data corresponding to the source face image, and only a monocular camera, a camera and other facial acquisition devices are needed to collect the source face image, effectively reducing the cost in expression transfer. Moreover, the method for expression transfer provided by the embodiments of the present application does not require establishing a semantic consistency binding between the source face and the target face, which improves the flexibility and scalability of expression transfer while reducing the cost of expression transfer. Description of the Drawings

[0023] Figure 1 is the system architecture diagram corresponding to the expression transfer system provided by the embodiments of the present application;

[0024] Figure 2 is the flowchart of the method for expression transfer provided by the embodiments of the present application;

[0025] Figure 3 is the schematic diagram of the construction of the facial feature space in the method for expression transfer provided by the embodiments of the present application;

[0026] Figure 4It is an example diagram of optimizing a feature encoder through an expression triple in the method for expression migration provided by an embodiment of the present application;

[0027] Figure 5 It is a structural block diagram of an example of the device for expression migration provided by an embodiment of the present application;

[0028] Figure 6 It is a structural block diagram of an example of an electronic device for data processing provided by an embodiment of the present application. Detailed implementation manners

[0029] Many specific details are set forth in the following description in order to provide a thorough understanding of the present application. However, the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the connotation of the present application. Therefore, the present application is not limited by the specific implementations disclosed below.

[0030] It should be noted that the terms "first", "second", "third", etc. in the claims, the description and the drawings of the present application are used to distinguish similar objects and are not used to describe a specific order or sequence. Data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than that shown or described herein. In addition, the terms "comprising", "having" and their variants are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0031] It should be understood that in the embodiments of the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. "Including A, B, and / or C" means including any one or any two or all three of A, B, and C.

[0032] It should be understood that in the embodiments of the present application, "B corresponding to A", "B corresponding to A relatively", "A corresponding to B relatively" or "B corresponding to A relatively" means that B is associated with A, and B can be determined according to A. Determining B according to A does not mean determining B only according to A, and B can also be determined according to A and / or other information.

[0033] Before elaborating on the implementation manners of the present application in detail, the prior art will be further described first.

[0034] 3D face animation transfer is an important technology that can capture and reproduce human facial expressions and movements to create realistic and vivid digital avatars. This technology has a wide range of applications in the fields of digital humans, virtual reality (VR), augmented reality (AR), human-computer interaction, etc.

[0035] Currently, the commonly used methods in the industry are as follows:

[0036] Method 1: Capture the motion data of the human face through a facial capture system (such as multi-camera or sensors) to reproduce real expressions in a virtual environment.

[0037] However, although Method 1 can achieve high-quality expression transfer, it requires expensive professional equipment and software support, making the cost of expression transfer too high.

[0038] Method 2: Transfer expressions based on deep learning-based face reconstruction + redirection. First, a corresponding 3D face model needs to be reconstructed for the source face. Then, according to the correspondence between the 3D face model and the face rigging of the target face, a mapping model is established. After that, using the mapping model, the expressions of the 3D face model are transferred to the target face.

[0039] However, Method 2 requires a semantic-consistent rigging between the target face and the source face, which often relies on a large amount of manual adjustment, reducing the flexibility and scalability of expression transfer.

[0040] Method 3: Generate 3D control parameters of the target face image from the source face image based on a neural network. This method usually relies on a single-modal source face image input and uses geometric priors (such as facial key points) and expression features to maintain the expression semantic consistency between the input source face image and the target face image.

[0041] However, Method 3 has the following problems: On the one hand, this method cannot capture supplementary or contextual information provided by other modalities (such as audio, text, etc.), which makes the neural network lack a comprehensive understanding when performing expression transfer; on the other hand, the geometric prior method based on facial key points relies on facial landmarks to define and track facial features, and it is often difficult to capture subtle expression changes, such as slight frowning or lip squeezing, which are crucial for expressing real emotions, resulting in a lack of fine-grained detail performance in expression transfer.

[0042] For the above reasons, in order to maintain high-fidelity of facial expression details during the facial expression transfer process, effectively reduce the cost of facial expression transfer, and improve the flexibility and scalability of facial expression transfer, the first embodiment of this application provides a method for facial expression transfer. This method is applied to an electronic device, which can be a desktop computer, a laptop computer, a mobile phone, a tablet computer, a smartwatch, etc., or other electronic devices capable of performing facial expression transfer. The embodiments of this application do not specifically limit it.

[0043] The method for facial expression transfer provided by this application can be applied to film and television dramas to generate the facial expressions of virtual characters through the facial expressions of real actors. It can also be applied to game development to enable game characters to reproduce the facial expressions of real players. It can also be applied to fields such as human-computer interaction and artistic creation. This application does not specifically limit the application scenarios.

[0044] Before introducing the method for facial expression transfer provided by the first embodiment of this application, first introduce the facial expression transfer system architecture applied by the method for facial expression transfer provided by the embodiments of this application. As Figure 1 shown, it is the system architecture diagram corresponding to the facial expression transfer system provided by the embodiments of this application. The facial expression transfer system 10 includes a facial expression extraction model 101, an audio feature encoder 102, an identity feature encoder 103, and a parameter generation model 104. Among them, the facial expression extraction model 101 is used to extract corresponding facial expression feature data from the source facial image. The audio feature encoder 102 is used to extract corresponding audio feature data from the phoneme sequence obtained by converting the source audio. The identity encoder 103 is used to extract the identity feature data corresponding to the target facial model. Then, the extracted facial expression feature data, audio feature data, and identity feature data are spliced and input into the parameter generation model 104. The parameter generation model 104 is used to generate the facial control parameters corresponding to the target facial model based on the input facial expression feature data, audio feature data, and identity feature data. The facial control parameters are used to control the deformation of the target facial model to present the facial expression of the source face and the mouth shape corresponding to the source audio. It should be noted that Figure 1 each model in will be described in detail in the introduction of the first embodiment of this application.

[0045] It should be noted that in the embodiments of this application, the execution subject of the method for facial expression transfer can be a terminal device or a server. Among them, the terminal device can be a local terminal device. The embodiments of this application do not limit the type of the execution subject.

[0046] The technical solutions of this application will be described in detail below through specific embodiments. It should be noted that the following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0047] Next, in combination with Figures 2 to 4 introduce the method for expression transfer provided in the embodiments of the present application.

[0048] As Figure 2 shown, it is a flowchart of the method for expression transfer provided in the first embodiment of the present application, including the following steps S101 to step S104.

[0049] Step S101: Obtain a source face image containing a first expression, a source audio, and a target face model to be subjected to expression transfer.

[0050] The above first expression can be, for example, expressions such as smiling, crying, surprise, fear, etc.

[0051] Optionally, in this step, a source video of a person speaking can be obtained, and the source video can be a video collected in any scenario. In this case, the source face image is the face image corresponding to the person in the source video, and the source audio is the audio corresponding to the person in the source video.

[0052] Optionally, this step can also separately obtain the source face image presenting the first expression and separately obtain the source audio, and the source face image and the source audio can also be videos collected in any scenario. In this case, the emotional characteristics corresponding to the source audio match the first expression. For example, if the first expression is smiling, the content corresponding to the source audio can be "Today is really a wonderful day, and I feel very happy", and for another example, if the first expression is surprise, the content corresponding to the source audio can be "Really? This is incredible".

[0053] The above target face model is a three-dimensional face model to be subjected to expression transfer, and specifically can include at least one of the three-dimensional face models corresponding to virtual characters in virtual games, the three-dimensional face models corresponding to virtual characters in human-computer interaction, the three-dimensional face models corresponding to virtual characters in the field of film and television animation, etc.

[0054] Step S102: Determine the expression feature data corresponding to the source face image, the audio feature data corresponding to the source audio, and the identity feature data corresponding to the target face model, where the audio feature data includes the emotional information corresponding to the source audio.

[0055] The above expression feature data is used to characterize the expression information of the first expression corresponding to the source face image, the above audio feature data is used to characterize the audio information corresponding to the source audio, and specifically can include the emotional information and text information corresponding to the source audio, and the above identity feature data is used to characterize the facial appearance information corresponding to the target face model. It should be noted that the expression feature data, the audio feature data, and the identity feature data can specifically be data in vector form.

[0056] Optionally, it can be through attachmentFigure 1 The pre-trained expression extraction model 101 shown in FIG. 1 extracts expression feature data from a source facial image.

[0057] Specifically, in the case where a source facial image is acquired separately, the "determining the expression feature data corresponding to the source facial image" in step S102 may be: inputting the source facial image into a pre-trained expression extraction model, so as to output the corresponding expression feature data through the expression extraction model. When the source facial image is a facial image corresponding to a person in the acquired source video, the "determining the expression feature data corresponding to the source facial image" in step S102 may be: inputting the source video into a pre-trained expression extraction model, so as to output the corresponding expression feature data through the expression extraction model.

[0058] Optionally, you can attach Figure 1 The pre-trained audio feature encoder 102 shown in FIG. 1 extracts audio feature data from source audio.

[0059] As an optional specific implementation manner, the emotional information includes emotion information and rhythm information, and the step of determining the audio feature data corresponding to the source audio may include the following steps:

[0060] Determine the phoneme sequence and text information corresponding to the source audio;

[0061] Determining the emotion information and rhythm information corresponding to the phoneme sequence;

[0062] The audio feature data corresponding to the source audio is determined according to the emotion information, the rhythm information and the text information.

[0063] As you can understand, phonemes are the smallest units of speech divided according to the natural properties of speech, and divided according to the pronunciation actions in the syllable, and one pronunciation action constitutes one phoneme. For example, the Chinese syllable "啊(a1)" has only one phoneme, and "代(dai)" has two phonemes d and ai4, and so on. "a1" represents the tone of "a" as 1, and "ai4" represents the tone of "ai" as 4. Therefore, each text can be divided into corresponding phoneme sequences.

[0064] In this implementation, first, the audio signal corresponding to the source audio can be converted into a phoneme sequence and text information. Specifically, the text corresponding to the source audio can be obtained first, and then the text can be converted into a phoneme sequence. For example, if the text information corresponding to the source audio is "Welcome here", the corresponding phonemes are: h, uan1, y, ing2, l, ai2, d, ao4, zh, e4, l, i3. In this way, the phonemes can be converted into corresponding phoneme sequences according to the sequence numbers of the phonemes.

[0065] After that, the emotional information and prosodic information of the phoneme sequence can be predicted to obtain the emotional information and prosodic information corresponding to the phoneme sequence. The emotional information is used to represent the emotional category and emotional intensity of the source audio, and the emotional category can include anger, surprise, disgust, happiness, fear, sadness, etc. The prosodic information is used to represent the cadence in the source audio.

[0066] By predicting the emotional information and prosodic information of the phoneme sequence, the emotional information and prosodic information corresponding to the phoneme sequence are obtained, and then combined with the text information corresponding to the source audio, audio feature data including emotional information, prosodic information, and text information can be obtained.

[0067] In a specific implementation, the phoneme sequence can be input into a pre-trained audio feature encoder to extract the context features of the phoneme sequence through the audio feature encoder, perform efficient global feature extraction on the phoneme sequence, and thus accurately and efficiently predict the emotional information and prosodic information of the phoneme sequence to output audio feature data including emotional information, prosodic information, and text information.

[0068] In this application, a Transformer model (a deep learning based on the self-attention mechanism) can be used as the above-mentioned audio feature encoder. The self-attention mechanism in the Transformer model allows the model to consider the relationships between all positions in the sequence when processing sequence data, rather than just adjacent positions. In this way, using the Transformer model as the audio feature encoder can model the phoneme sequence corresponding to the source audio, improve the robustness to diverse speech features such as accents and speech rate changes, and obtain audio feature data that can accurately reflect the audio information corresponding to the source audio.

[0069] Through this specific implementation, the source audio is converted into a phoneme sequence. On the one hand, it can significantly reduce the complexity of audio data, remove irrelevant information such as redundant information and background noise in the source audio, and at the same time retain the key information of the source audio; on the other hand, it greatly reduces the computational amount and storage requirements, making the audio feature encoder more efficient during inference.

[0070] As another alternative specific implementation, determining the audio feature data corresponding to the source audio can also be: directly inputting the source audio into a specified model to output the audio feature data corresponding to the source audio through the specified model. The specified model can be the above-mentioned audio feature encoder.

[0071] Optionally, the identity feature data can be extracted from the target face model through the pre-trained identity encoder 103 shown in the appendix Figure 1

[0072] ​In a specific implementation manner, identity feature data can be extracted based on the above-mentioned target face model, and identity feature data can also be extracted based on the 2D target face image corresponding to the target face model. That is to say, the target face model can be input into a pre-trained identity encoder, or alternatively, the 2D target face image corresponding to the target face model can be input into the pre-trained identity encoder to obtain a feature vector of a preset length through encoding by the identity encoder, and this feature vector is the identity feature data corresponding to the target face model.

[0073] In this application, a Multilayer Perceptron (MLP) can be used as the above-mentioned identity encoder. A multilayer perceptron is a feedforward neural network composed of multiple layers of nodes (or called "neurons"), and each node is connected to all nodes in the next layer through weights.

[0074] Through step S102, the expression feature data representing the expression information of the first expression included in the source face image, the audio feature data representing the audio information corresponding to the source audio, and the identity feature data representing the facial appearance information corresponding to the target face model are determined.

[0075] Step S103: Generate facial control parameters corresponding to the target face model according to the expression feature data, the audio feature data, and the identity feature data.

[0076] This step is used to convert multi-modal features into facial control parameters corresponding to the target face model for precise control of the expression presentation of the target face model.

[0077] The above-mentioned facial control parameters are used to control the target face model to present the first expression and the mouth shape corresponding to the source audio. The facial control parameters are parameters that change over time, that is, the facial control parameters include facial control sub-parameters corresponding to each frame, and the facial control sub-parameters corresponding to every two frames may be different.

[0078] In this step, facial control parameters corresponding to the target face model can be generated according to the expression feature data, the audio feature data, and the identity feature data. Since multi-modal data corresponding to the expression dimension and the audio dimension are considered, the audio feature data is used to supplement and enrich the visual expression details of the expression. Therefore, the generated facial control parameters can represent fine-grained expression detail information.

[0079] In an optional implementation manner, step S103 can be specifically implemented through the following steps:

[0080] Generate initial facial control parameters corresponding to the target face model according to the expression feature data and the identity feature data;

[0081] According to the emotion information, prosody information, and text information included in the audio feature data, make detailed adjustments to the initial facial control parameters to generate the facial control parameters corresponding to the target facial model.

[0082] In this embodiment, the initial facial control parameters for roughly controlling the target facial model can be generated first through the expression feature data and identity feature data. The initial facial control parameters are only used to feedback the influence of the single-modal data of the expression dimension on the facial movements of the target facial model.

[0083] After that, according to the emotion information, prosody information, and text information included in the audio feature data, make detailed adjustments to the initial facial control parameters of the rough control.

[0084] Specifically, the emotion information and prosody information are used to control the movement patterns of facial muscles. Exemplarily, when the emotion category is happy, the movement pattern of facial muscles can be that the corners of the mouth turn up to form a smile, the eyes narrow, and the cheeks lift; when the emotion category is sad, the movement pattern of facial muscles can be that the corners of the mouth droop and the inner corners of the eyebrows are raised; when the emotion category is sad, the movement pattern of facial muscles can be that the eyebrows are furrowed and the eyes are wide open; when the emotion category is fear, the movement pattern of facial muscles can be that the eyes widen, the eyebrows are raised, and the mouth is slightly open; when the emotion category is surprise, the movement pattern of facial muscles can be that the eyebrows are raised high, the eyes are wide open, and the mouth opens in a round shape. Exemplarily, when the intonation is rising, it is usually accompanied by facial details such as raised eyebrows and widened eyes; when the volume and pitch are increased significantly, it is usually accompanied by facial details such as furrowed eyebrows; when the speech rate is fast, it is usually accompanied by facial details such as frequent blinking and rapid twitching of the corners of the mouth; when the speech rate is slow, it is usually accompanied by facial details such as furrowed eyebrows and lowered eyes.

[0085] The text information is used to control the changes in mouth shapes. The changes in mouth shapes include basic opening and closing movements, as well as the specific position changes of the tongue, teeth, and lips, and can accurately feedback the corresponding pronunciation features. Specifically, it can be achieved through the following steps:

[0086] First, determine the text semantics and grammatical structure corresponding to the text information. This step is used to obtain information such as sentence boundaries, stress positions, and pause times in the text;

[0087] After that, for each character in the text information, determine its corresponding mouth shape, and the mouth shape includes the lip shape and tooth position.

[0088] In the specific implementation steps, first, the corresponding minimum pronunciation unit is determined for each character, for example, the character "你", the corresponding minimum pronunciation units include n and i; then, the corresponding sub-mouth shape is determined for each minimum pronunciation unit, for example, the lip shape corresponding to n is slightly open lips, the tooth position is the differential of upper and lower teeth, and the lip shape corresponding to i is the upper and lower lips separated and flat, and the tooth position is the differential of upper and lower teeth; finally, the sub-mouth shape corresponding to each minimum pronunciation unit is connected in sequence to obtain the mouth shape corresponding to the character, for example, by connecting the sub-mouth shapes corresponding to n and i, the mouth shape corresponding to the character "你" can be obtained (the lip shape changes from slightly open lips to upper and lower lips separated and flat, and the tooth position is the differential of upper and lower teeth)

[0089] Afterwards, for the i-th and i+1-th characters in the text information, a lip transition animation is generated from the lip shape corresponding to the i-th character to the lip shape corresponding to the i+1-th character, and a tooth transition animation is generated from the tooth position corresponding to the i-th character to the tooth position corresponding to the i+1-th character, and the lip transition animation and the tooth transition animation are determined as the lip transition animation. Among them, i traverses 1 to M-1, and M is the total number of characters in the text information;

[0090] Finally, the lip shape and the lip shape transition animation are connected in sequence to obtain the lip shape corresponding to the text information.

[0091] In this way, the rough facial control parameters can be adjusted in detail through the emotion information, rhythm information and text information, so as to generate facial control parameters for fine control of the target facial model. The facial control parameters are used to feedback the influence of the multimodal data in the expression dimension and the audio dimension on the facial movements of the target facial model.

[0092] In an optional implementation, step S103 may be implemented by the following steps:

[0093] The expression feature data, the audio feature data and the identity feature data are input into a pre-trained parameter generation model, so as to generate facial control parameters corresponding to the target facial model through the parameter generation model.

[0094] Specifically, the expression feature data, audio feature data and identity feature data can be spliced ​​and integrated to obtain target feature data, and then the target feature data is input into a pre-trained parameter generation model, so as to output facial control parameters corresponding to the target facial model through the parameter generation model.

[0095] It should be noted that the target feature data obtained by splicing and integrating the facial expression feature data, audio feature data, and identity feature data is the data after integrating data of different dimensions, with comprehensive and rich detailed information. In this way, it provides more comprehensive and rich information for the parameter generation model, enabling the parameter generation model to better predict the facial control parameters corresponding to the target facial model.

[0096] Step S104: According to the facial control parameters, adjust the positions of the vertices in the target facial model to generate an animation of the first expression presenting expression details matching the emotion information for the target facial model.

[0097] After obtaining the facial control parameters corresponding to the target facial model, through these facial control parameters, the vertices of the target facial model can be driven to move, and the positions of the vertices in the target facial model can be adjusted. It can be understood that the vertices of the model are the basic elements in a 3D geometric figure, representing a point in three-dimensional space. Each vertex has specific position coordinates (usually represented by x, y, z), and the vertices are used to define and describe the shape, appearance, and other characteristics of the 3D model.

[0098] As can be seen from the foregoing introduction, the facial control parameters are parameters that change over time, including the facial control sub-parameters corresponding to each frame. In this step, the vertices of the target facial model can be driven to move respectively according to the facial control sub-parameters corresponding to each frame to adjust the position of each frame of vertices, thereby obtaining the animation of the target facial model. Since the facial control parameters are used to reflect the influence of the multi-modal data in the expression dimension and audio dimension on the facial actions of the target facial model, in this way, according to the facial control parameters, an animation of the first expression presenting expression details matching the emotion information can be generated for the target facial model. That is to say, in this animation, the target facial model generally presents the first expression contained in the source facial image, and the facial expression details match the emotion information corresponding to the source audio, with rich expression details.

[0099] In an optional implementation manner, the above-mentioned generating an animation of the first expression presenting expression details matching the emotion information for the target facial model may include: generating an animation of the first expression presenting expression details matching the emotion information and the mouth shape corresponding to the source audio for the target facial model.

[0100] The method for facial expression transfer provided by the embodiments of the present application includes the following steps: obtaining a source facial image containing a first expression, a source audio, and a target facial model to be subjected to facial expression transfer; determining expression feature data corresponding to the source facial image, audio feature data corresponding to the source audio, and identity feature data corresponding to the target facial model, wherein the audio feature data includes emotional information corresponding to the source audio; generating facial control parameters corresponding to the target facial model according to the expression feature data, the audio feature data, and the identity feature data; and adjusting the positions of vertices in the target facial model according to the facial control parameters to generate an animation presenting the first expression with expression details matching the emotional information for the target facial model. It can be seen that in the present application, first, the expression feature data corresponding to the source facial image containing the first expression, the audio feature data corresponding to the source audio, and the identity feature data corresponding to the target facial model are determined. The expression feature data is used to characterize the expression information corresponding to the first expression, and the audio feature data includes the emotional information corresponding to the source audio. This emotional information can supplement the visual expression details of the face. In this way, through the multi-modal data of the expression feature data and the audio feature data, combined with the identity feature data of the target facial model, the generated facial control parameters can characterize fine-grained expression detail information, thereby driving the target facial model to present the first expression with expression details matching the emotional information and the mouth shape corresponding to the source audio.

[0101] Therefore, the method for facial expression transfer provided by the embodiments of the present application can maintain high fidelity of expression details during the facial expression transfer process. In addition, the method for facial expression transfer provided by the embodiments of the present application does not require a costly facial capture system to capture the motion data corresponding to the source facial image. Only a monocular camera, a camera and other facial acquisition devices are needed to acquire the source facial image, effectively reducing the cost in facial expression transfer. Moreover, the method for facial expression transfer provided by the embodiments of the present application does not require establishing a semantic consistency binding between the source face and the target face, which reduces the cost of facial expression transfer while improving the flexibility and scalability of facial expression transfer.

[0102] The following provides a detailed introduction to the training of the expression extraction model and the parameter generation model used in the embodiments of the present application:

[0103] Training of the expression extraction model

[0104] Optionally, a pre-trained expression extraction model can be obtained through the following method:

[0105] Obtain sample facial images;

[0106] Randomly mask the sample facial images to obtain processed sample facial images;

[0107] Input the processed sample face image into the feature encoder to be trained, so as to output the sample expression feature data of the processed sample image through the feature encoder;

[0108] Determine the predicted face image corresponding to the sample expression feature data;

[0109] Adjust the model parameters of the feature encoder to make the difference between the predicted face image and the sample face image less than the first threshold, and obtain the pre-trained feature encoder;

[0110] Determine the pre-trained expression extraction model according to the pre-trained feature encoder.

[0111] In this embodiment, the sample face image can be a face expression image without a label (that is, not annotated by an annotator), and this sample face image can include multiple face images collected in multiple environments. The multiple environments can include environments with different backgrounds and different illuminations. The face expression image can include real-person face expression images and stylized face expression images of cartoon animations. The real-person face expression images can include real-person face expression images of various ethnic groups, genders, ages, and expressions.

[0112] By obtaining real-person face expression images and stylized face expression images of cartoon animations, the differences between cartoon characters and real people are fully considered, and the obtained sample face images have no labels, providing a basis for training an expression extraction model with strong generalization ability. In addition, since the sample face images include face images collected in multiple environments, the trained expression extraction model can be adapted to source face images collected in various environments, rather than being limited to source face images collected in a strictly illuminated background set in the laboratory, which can improve the flexibility and scalability of this solution.

[0113] After that, the obtained sample face images can be randomly masked to obtain the processed sample face images. Specifically, each sample face image can be divided into a preset number (such as 16*16, or 32*32, etc.) of non-overlapping regions, and a preset proportion (such as 70%, or 75%, etc.) of target non-overlapping regions are randomly selected from these non-overlapping regions, and the target non-overlapping regions are masked to obtain the processed sample face images.

[0114] After that, the processed sample face images can be input into the feature encoder to be trained, so as to encode the processed sample face images through the feature encoder to obtain the corresponding sample expression feature data.

[0115] After that, the predicted face image corresponding to the sample expression feature data can be determined. Specifically, the sample expression feature data can be decoded to predict the corresponding predicted face image. As an alternative implementation, the sample expression feature data can be input into a pre-trained feature decoder to output the corresponding predicted face image through the feature decoder.

[0116] After that, based on the training strategy that the difference between the predicted face image and the sample face image is less than the first threshold, the model parameters of the feature encoder to be trained can be adjusted to obtain a pre-trained feature encoder. In this way, the difference between the predicted face image corresponding to the sample expression feature data encoded by the pre-trained feature encoder for the processed sample face image without a label and after masking and the original sample face image is small. Therefore, the pre-trained feature encoder can learn the internal features and structure of the face image and has strong robustness and generalization ability.

[0117] As Figure 3 shown, it is a schematic diagram of the construction of the facial feature space in the method for expression transfer provided by the embodiment of the present application. The processed sample face image 107 with a certain degree of masking is input into the feature encoder 105 to be trained. The feature encoder 105 encodes the processed sample face image 107 to obtain the corresponding sample expression feature data. The sample expression feature data is input into the pre-trained feature decoder 106, and the feature decoder 106 decodes the sample expression feature data to obtain the predicted face image.

[0118] In a specific implementation, the model parameters of the feature encoder to be trained can be optimized by the L2 function. In the embodiment of the present application, an example of the optimization function used by the feature encoder is shown in the following formula (1):

[0119]

[0120] In the above formula (1), N is the number of sample face images, y i is the i-th sample face image, is the predicted face image corresponding to the i-th sample face image, and L recons represents the difference between the predicted face image and the sample face image in the L2 function.

[0121] After obtaining the pre-trained feature encoder, a pre-trained expression extraction model can be determined according to the pre-trained feature encoder.

[0122] In an alternative implementation, the pre-trained feature encoder can be determined as the pre-trained expression extraction model.

[0123] In an alternative embodiment, the above step of "determining the pre-trained expression extraction model according to the pre-trained feature encoder" may specifically include the following steps:

[0124] Construct expression triplets, where each expression triplet includes an anchor expression image, a positive sample expression image, and a negative sample expression image;

[0125] Input the anchor expression image, the positive sample expression image, and the negative sample expression image into the pre-trained feature encoder respectively, so as to output the first expression feature data corresponding to the anchor expression image, the second expression feature data corresponding to the positive sample expression image, and the third expression feature data corresponding to the negative sample expression image through the pre-trained feature encoder respectively;

[0126] Adjust the model parameters of the pre-trained feature encoder so that the difference between the second expression feature data and the first expression feature data is less than a second threshold, and the difference between the third expression feature data and the first expression feature data is greater than a third threshold, to obtain an optimized feature encoder; wherein, the second threshold is less than the third threshold;

[0127] Determine the optimized feature encoder as the pre-trained expression extraction model.

[0128] In this embodiment, the feature encoder can be optimized by constructing expression triplets, so as to obtain an expression extraction model that can accurately perceive various expressions.

[0129] The facial expression images of the constructed expression triplets can include various expressions. Each expression triplet contains three facial expression images, and each expression triplet has an annotation of expression similarity. Based on the corresponding annotation of expression similarity, the facial expression images included in each expression triplet can be divided into an anchor expression image, a positive sample expression image, and a negative sample expression image.

[0130] For each expression triplet, the facial expression image annotated as the least similar to the other two facial expression images is the negative sample expression image, and any one of the remaining two facial expression images is the anchor expression image, and the other is the positive sample expression image. For example, expression triplet A includes facial expression image a1, facial expression image a2, and facial expression image a3. The annotation result of the annotator is that a1 and a2 are relatively similar, a3 is not similar to a1, and a3 is not similar to a2 either. Then a3 is the negative sample expression image, and any one of a1 and a2 is the anchor expression image, and the other is the positive sample expression image.

[0131] In an alternative specific embodiment, the expression triplets can be constructed in the following manner:

[0132] Obtain a first expression triple with labeled expression similarity and facial expression images without labels;

[0133] Construct a second expression triple based on the facial expression images, and label the expression similarity of the second expression triple to obtain a second expression triple with labeled expression similarity. Among them, the first expression triple includes a first number of asymmetric expression samples, and the second expression triple includes a second number of asymmetric expression triples, and the first number is less than the second number;

[0134] Construct an expression triple according to the first expression triple and the second expression triple with labeled expression similarity.

[0135] In this embodiment, a first expression triple with labeled expression similarity can be obtained from a database, and the consistency rate of the labels corresponding to each expression triple in the first expression triple is greater than a first preset value (such as 60% or 70%, etc.). That is to say, each expression triple in the first expression triple has labels from multiple annotators, and the ratio of the target annotators with the same annotation result for this expression triple to the total number of multiple annotators is greater than the first preset value.

[0136] After that, facial expression images without labels can be obtained. These facial expression images can be the same as the sample facial images introduced above, or can be other obtained real-person facial expression images and cartoon animation facial expression images. This application does not limit whether the facial expression images and the sample facial images are the same.

[0137] After that, multiple second expression triples can be randomly constructed from these facial expression images. For example, if the facial expression images include facial expression image 1, facial expression image 2, facial expression image 3,..., facial expression image x-2, facial expression image x-1, facial expression image x, (facial expression image 1, facial expression image 2, facial expression image 3) can be used as a second expression triple,..., and (facial expression image x-1, facial expression image x-2, facial expression image x-3) can be used as a second expression triple, and so on.

[0138] After that, each second expression triple can be sent to multiple annotation terminals so that multiple annotators can respectively label the expression similarity of the three facial expression images included in each second expression triple on the multiple annotation terminals, thereby obtaining a second expression triple with labeled expression similarity.

[0139] After that, an expression triple can be constructed according to the first expression triple and the second expression triple with labeled expression similarity.

[0140] Optionally, the first expression triplet and the second expression triplet may be used as an expression triplet.

[0141] Optionally, a target second expression triplet with a labeling consistency rate greater than a second preset value (such as 60%, or 70%) can be determined from the second expression triplet, and the first expression triplet and the target second expression triplet are used as expression triplet. In this setting, the second expression triplet can be effectively screened by labeling consistency rate, thereby obtaining high-quality training data to train the expression extraction model efficiently and accurately.

[0142] It should be noted that the first expression triplet includes a first number of asymmetric expression samples, the second expression triplet includes a second number of asymmetric expression triplets, and the first number is less than the second number. In other words, asymmetric expression samples in the first expression triplet are relatively scarce, so the expression extraction model trained only by the first expression triplet is difficult to effectively express facial asymmetry. After adding a larger number of asymmetric expression images to the second expression triplet, the trained expression extraction model also performs well in facial asymmetry.

[0143] After constructing the expression triples, for each expression triple, its corresponding anchor expression image, positive sample expression image and negative sample expression image can be mapped to the same expression latent space. Specifically, the anchor expression image is input into the pre-trained feature encoder to obtain the first expression feature data corresponding to the anchor expression image, the positive sample expression image is input into the pre-trained feature encoder to obtain the second expression feature data corresponding to the positive sample expression image, and the negative sample expression image is input into the pre-trained feature encoder to obtain the third expression feature data corresponding to the negative sample expression image.

[0144] Afterwards, the model parameters of the pre-trained feature encoder can be further adjusted based on the training strategy that the difference between the second expression feature data and the first expression feature data is less than the second threshold, and the difference between the third expression feature data and the first expression feature data is greater than the third threshold, thereby obtaining an optimized feature encoder. That is, in the expression latent space, by shortening the distance between the anchor expression image and the positive sample expression image, and increasing the distance between the anchor expression image and the negative sample expression image, and ensuring that the distance between the anchor expression image and the positive sample expression image is less than the distance between the anchor expression image and the negative sample expression image.

[0145] In a specific implementation, an example optimization function used by the feature encoder is shown in the following formula (2):

[0146] L tri =w·Max(0,|f a -fp | 2 -|f a -f n | 2 +m) Formula (2)

[0147] In the above formula (2), w is the confidence score, which is the ratio of the number of agreed annotations to the total number of annotations for the sample. The number of agreed annotations refers to the maximum number of annotators with the same or consistent annotation results when multiple annotators annotate the same data sample. For example, for a sample, 100 annotators have made annotations, and among them, 80 annotators have consistent annotations, then the number of agreed annotations is 80, and the confidence score is 80%. f a is the first expression feature data, f p is the second expression feature data, f n is the third expression feature data. m is the value obtained by subtracting the second difference from the first difference in the critical state. The first difference is the difference between the first expression feature data and the second expression feature data, and the second difference is the difference between the first expression feature data and the third expression feature data. The critical state is the state corresponding to the boundary that ensures that the anchor expression image and the positive sample expression image are closer in the expression latent space than the anchor expression image and the negative sample expression image, that is, the state where the first difference just changes from being greater than the second difference to being less than the second difference.

[0148] As Figure 4 shown, it is an example diagram of optimizing the feature encoder through an expression triple in the expression migration method provided by an embodiment of the present application. The expression triple 109 includes an anchor expression image 109-1, a positive sample expression image 109-2, and a negative sample expression image 109-3. Inputting the anchor expression image 109-1 into the feature encoder 105 can obtain the first expression feature data. Inputting the positive sample expression image 109-2 into the feature encoder 105 can obtain the second expression feature data. Inputting the negative sample expression image 109-3 into the feature encoder 105 can obtain the third expression feature data. Then, the distance between the first expression feature data and the second expression feature data can be shortened, and the distance between the first expression feature data and the third expression feature data can be lengthened to optimize the feature encoder.

[0149] The optimized feature encoder can be used as a pre-trained expression extraction model for extracting expression feature data from the source face image.

[0150] Above, the training of the expression extraction model is completed.

[0151] It can be seen that through the above two-stage training of the expression extraction model, based on the first training to learn the internal features and structure of the face image, and then retraining in combination with the expression triples, the expression extraction model can achieve high-precision perception of various expressions.

[0152] Training of the parameter generation model

[0153] Optionally, a pre-trained parameter generation model can be obtained through the following method:

[0154] Obtain a sample video of a sample speaker speaking with at least one second expression;

[0155] According to the sample video, determine the sample face control parameters corresponding to the target face model, and the sample face control parameters are used to control the target face model to present the second expression;

[0156] According to the sample face control parameters, determine the sample expression feature data corresponding to the target face model presenting the second expression;

[0157] Determine the sample audio feature data corresponding to the sample audio of the sample video;

[0158] Input the sample expression feature data, the sample audio feature data, and the identity feature data corresponding to the target face model into the parameter generation model to be trained, so as to output the predicted face control parameters through the parameter generation model;

[0159] Adjust the model parameters of the parameter generation model so that the difference between the predicted face control parameters and the sample face control parameters is less than the fourth threshold to obtain a pre-trained parameter generation model.

[0160] The above sample video can be obtained by a facial acquisition device. A facial acquisition device refers to a hardware device used to capture and record face images or videos. There can be multiple acquired sample videos, and each sample video can be a video of at least one sample speaker speaking with one second expression. The second expression can include expressions such as smiling, crying, surprise, fear, etc.

[0161] After obtaining the sample video, the sample face control parameters corresponding to the target face model when presenting the second expression can be determined according to the sample video. Specifically, the sample face control parameters can be obtained by resolving from the sample video through facial capture software, and facial capture software refers to software used to analyze and process facial data.

[0162] In an optional specific implementation manner, the step of "According to the sample video, determine the sample face control parameters corresponding to the target face model" may include the following steps:

[0163] Identify the first facial key points corresponding to the sample speaker;

[0164] Determine the position data corresponding to each frame of the first facial key points in the sample video;

[0165] According to the correspondence between the position data, the first facial key points, and the second facial key points corresponding to the target face model, determine the sample face control parameters corresponding to the target face model presenting the second expression.

[0166] First, the sample video can be imported into facial capture software, which usually has an automatic tracking function built-in to identify and track the first facial key points corresponding to the sample speaker in the sample video. The first facial key points usually can include parts such as eyes, eyebrows, nose, and mouth. Usually, based on the imported sample video, the facial capture software can output a series of time-series data, which is used to describe the position data corresponding to each frame of the first facial key points in the sample video.

[0167] After that, the correspondence between the first facial key points and the second facial key points corresponding to the target face model can be established. Specifically, the first facial key points (such as eyes, eyebrows, nose, mouth, etc.) extracted from the sample video can be matched with the same or similar positions on the target face model. This step is used to ensure that the captured second expression can be accurately and naturally applied to the target face model.

[0168] After that, based on the position data of the first facial key points in the sample video and the correspondence between the first facial key points and the second facial key points, the sample face control parameters corresponding to the target face model presenting the second expression can be determined. It should be noted that the sample face control parameters include the sample face control sub-parameters corresponding to each frame.

[0169] After that, according to the sample face control parameters, the sample expression feature data corresponding to the target face model presenting the second expression can be determined. The sample face control parameters are used to control the target face model to present the second expression, and the sample expression feature data is used to characterize the expression information when the target face model presents the second expression.

[0170] Specifically, the sample expression feature data can be determined through the following steps:

[0171] According to the sample face control parameters, adjust the positions of the vertices in the target face model to obtain the sample animation of the target face model;

[0172] Obtain the sample face images corresponding to each frame in the sample animation of the target face model;

[0173] Determine the sample expression feature data corresponding to each frame of the sample facial image respectively.

[0174] In the embodiment of the present application, the vertices of the target facial model can be driven to move through the sample facial control parameters, so as to adjust the positions of the vertices in the target facial model and obtain the sample animation of the target facial model. As can be known from the foregoing introduction, the sample facial control parameters include sample facial control sub-parameters corresponding to each frame. In this way, the vertices of the target facial model can be driven to move respectively through the sample facial control sub-parameters corresponding to each frame, so as to adjust the positions of the vertices in the target facial model for each frame, and obtain the sample facial image corresponding to each frame in the sample animation to be generated by the target facial model.

[0175] It should be noted that in specific implementation, parameter configuration can be performed in the rendering engine. Specifically, the shooting angle (such as the front shooting angle) and camera parameters can be configured. The camera parameters can include at least one of focal length, viewing angle width, and illumination. After the parameter configuration is completed, the rendering engine can render the target facial model into the sample facial image corresponding to the fixed shooting angle and fixed camera parameters through the above sample facial control parameters.

[0176] After that, the sample expression feature data corresponding to each frame of the sample facial image can be determined. Specifically, the sample facial images corresponding to each frame can be respectively input into the pre-trained expression extraction model mentioned in the foregoing introduction, or the above sample animation can be input into the pre-trained expression extraction model, so as to output the sample expression feature data corresponding to each frame through the expression extraction model.

[0177] After that, the sample audio feature data corresponding to the sample audio of the above sample video can be determined. Specifically, the sample phoneme sequence corresponding to the sample audio can be determined, and the sample phoneme sequence can be input into the pre-trained audio feature encoder, so as to output the sample audio feature data corresponding to the sample audio through the audio feature encoder.

[0178] After obtaining the sample facial control parameters, sample expression feature data, and sample audio feature data, the sample expression feature data, sample audio feature data, and identity feature data corresponding to the target facial model can be input into the parameter generation model to be trained, so as to predict the facial control parameters corresponding to presenting the second expression indicated by the sample expression feature data and the mouth shape indicated by the sample audio feature data through the parameter generation model, and obtain the predicted facial control parameters.

[0179] After that, based on a training strategy where the difference between the predicted facial control parameters and the sample facial control parameters is less than a fourth threshold, the model parameters of the parameter generation model can be adjusted to obtain a pre-trained parameter generation model.

[0180] As can be seen from the foregoing introduction, the sample expression feature data is the sample expression feature data corresponding to each frame of the image. In this embodiment, the sample audio can be framed to align the sample audio feature data and the sample expression feature data. Therefore, the above step of "determining the sample audio feature data corresponding to the sample audio corresponding to the sample video" may include the following steps:

[0181] Cut the sample audio corresponding to the sample video into single-frame sample audio corresponding to each frame according to the frame rate of the sample video;

[0182] Determine the sample audio feature data corresponding to each frame of the single-frame sample audio respectively.

[0183] In this embodiment, the frame rate of the sample video can be obtained first. The frame rate refers to the number of frames of images displayed per second. After that, according to the frame rate of the sample video, the above sample audio can be cut into single-frame sample audio corresponding to each frame. For example, if the frame rate of the sample video is 30 FPS, which means 30 frames of images are displayed per second in the sample video, then the sample audio can also be cut into 30 frames of speech per second to obtain single-frame sample audio.

[0184] After that, the sample audio feature data corresponding to each frame of the single-frame sample audio can be determined by the above pre-trained audio feature encoder.

[0185] In an optional embodiment, after obtaining the sample expression feature data corresponding to each frame of the sample facial image and the sample audio feature data corresponding to each frame of the single-frame sample audio, the parameter generation model can be trained frame by frame in combination with the sample facial control sub-parameters corresponding to each frame. Specifically, the training method is as follows:

[0186] For the i-th frame, input the sample expression feature data corresponding to the i-th frame, the sample audio feature data corresponding to the i-th frame, and the identity feature data corresponding to the target facial model into the parameter generation model to be trained, so as to output the predicted facial control sub-parameters corresponding to the i-th frame through the parameter generation model, where i traverses 1 to N, and N is the number of video frames of the sample video;

[0187] Adjust the model parameters of the parameter generation model so that the difference between the predicted facial control sub-parameters corresponding to the i-th frame and the sample facial control sub-parameters corresponding to the i-th frame in the sample facial control parameters is less than a fourth threshold to obtain a pre-trained parameter generation model.

[0188] In this embodiment, for the first frame, the sample expression feature data corresponding to the first frame, the sample audio feature data corresponding to the first frame, and the identity feature data corresponding to the target face model can be input into the parameter generation model to be trained, so as to output the predicted face control sub-parameters corresponding to the first frame through this parameter generation model. And based on the training strategy that the difference between the predicted face control sub-parameters corresponding to the first frame and the sample face control sub-parameters corresponding to the first frame is less than the fourth threshold, the model parameters of the parameter generation model are adjusted;

[0189] For the second frame, the sample expression feature data corresponding to the second frame, the sample audio feature data corresponding to the second frame, and the identity feature data corresponding to the target face model can be input into the parameter generation model, so as to output the predicted face control sub-parameters corresponding to the second frame through this parameter generation model. And based on the training strategy that the difference between the predicted face control sub-parameters corresponding to the second frame and the sample face control sub-parameters corresponding to the second frame is less than the fourth threshold, the model parameters of the parameter generation model are adjusted again.

[0190] And so on...

[0191] For the Nth frame, the sample expression feature data corresponding to the Nth frame, the sample audio feature data corresponding to the Nth frame, and the identity feature data corresponding to the target face model can be input into the parameter generation model, so as to output the predicted face control sub-parameters corresponding to the Nth frame through this parameter generation model. And based on the training strategy that the difference between the predicted face control sub-parameters corresponding to the Nth frame and the sample face control sub-parameters corresponding to the Nth frame is less than the fourth threshold, the model parameters of the parameter generation model are adjusted again.

[0192] In this way, the parameter generation model can be trained frame by frame, so as to obtain a pre-trained parameter generation model.

[0193] It should be noted that in the embodiments of the present application, for a target face model, its corresponding parameter generation model can be trained separately. For example, for the target face model a, a parameter generation model suitable for the target face model a is trained, and for the target face model b, a parameter generation model suitable for the target face model b is trained. Since different target face models usually have unique facial structures, proportions and styles, in this way, training the corresponding parameter generation model for each target face model separately can accurately capture the facial detail information corresponding to the target face model.

[0194] The above is the completion of the training of the parameter generation model.

[0195] It should be noted that since the pre-trained parameter generation model can generate predicted facial control parameters with a difference less than the fourth threshold from the sample facial control parameters for the input sample facial expression feature data, sample audio feature data, and identity feature data corresponding to the target facial model, and the sample facial control parameters are used to control the target facial model to present the expression when the sample speaker speaks with the second expression, the predicted facial control parameters can transfer the expression details of the sample speaker speaking with the second expression to the target facial model with high fidelity. In this way, through the pre-trained parameter generation model provided by the embodiments of the present application, facial control parameters can be accurately generated through the expression feature data corresponding to the input source facial image, the audio feature data corresponding to the source audio, and the identity feature data corresponding to the target facial model, so as to transfer the expression details of the first expression of the source facial image and the mouth shape corresponding to the source audio to the target facial model with high fidelity through the facial control parameters, thereby achieving high fidelity of expression details in the expression transfer process.

[0196] It should be noted that in the embodiments of the present application, only during the training process of the parameter generation model, the correspondence between the facial key points of the sample speaker and the facial key points of the target facial model is established to obtain the sample facial control parameters through the redirection technology, which is contrasted with the predicted facial control parameters generated by the parameter generation model based on the sample facial expression feature data, sample audio feature data, and identity feature data of the target facial model, so as to adjust the model parameters of the parameter generation model. In the inference stage of the parameter generation model, there is no need to establish the correspondence between the facial key points of the source facial image and the facial key points of the target facial model. Only the expression feature data of the source facial image needs to be extracted, and then, combined with the audio feature data and identity feature data, the parameter generation model can be used to generate the facial control parameters without semantic consistency binding between the source face and the target face, thereby reducing the cost of expression transfer and improving the flexibility and scalability of expression transfer.

[0197] It can be seen that for the expression transfer method provided by the embodiments of the present application, since an expression extraction model capable of learning the internal features and structures of facial images and having high-precision perception of various expressions is trained, through this expression extraction model, the expression feature data of the source facial image can be extracted with high precision. Since the trained parameter generation model can generate facial control parameters for transferring the expression details of the first expression of the source facial image and the mouth shape corresponding to the source audio to the target facial model with high fidelity based on the expression feature data, audio feature data, and identity feature data, through the expression extraction model and the parameter generation model, high fidelity of expression details in the expression transfer process can be achieved.

[0198] Corresponding to the method for expression transfer provided in the first embodiment of the present application, the second embodiment of the present application further provides an expression transfer device, as Figure 5 shown. The expression transfer device 500 includes:

[0199] An acquisition unit 501, configured to acquire a source face image including a first expression, a source audio, and a target face model to be subjected to expression transfer;

[0200] A determination unit 502, configured to determine expression feature data corresponding to the source face image, audio feature data corresponding to the source audio, and identity feature data corresponding to the target face model, where the audio feature data includes emotional information corresponding to the source audio;

[0201] A generation unit 503, configured to generate face control parameters corresponding to the target face model according to the expression feature data, the audio feature data, and the identity feature data;

[0202] An adjustment unit 504, configured to adjust the positions of vertices in the target face model according to the face control parameters, so as to generate an animation presenting the first expression with expression details matching the emotional information for the target face model.

[0203] Optionally, the determination unit 502 is specifically configured to:

[0204] Input the source face image into a pre-trained expression extraction model, so as to output corresponding expression feature data through the expression extraction model.

[0205] Optionally, the expression transfer device 500 further includes a training unit, and the training unit is configured to train a pre-trained expression extraction model in the following manner:

[0206] Acquire sample face images;

[0207] Randomly mask the sample face images to obtain processed sample face images;

[0208] Input the processed sample face images into the to-be-trained feature encoder, so as to output sample expression feature data of the processed sample images through the feature encoder;

[0209] Determine a predicted face image corresponding to the sample expression feature data;

[0210] Adjust model parameters of the feature encoder, so that the difference between the predicted face image and the sample face image is less than a first threshold, to obtain a pre-trained feature encoder;

[0211] Determine a pre-trained expression extraction model according to the pre-trained feature encoder.

[0212] Optionally, the training unit is specifically configured to:

[0213] Input the sample expression feature data into a pre-trained feature decoder to output a corresponding predicted face image through the feature decoder.

[0214] Optionally, the training unit is specifically configured to:

[0215] Construct expression triples, where each expression triple includes an anchor expression image, a positive sample expression image, and a negative sample expression image;

[0216] Input the anchor expression image, the positive sample expression image, and the negative sample expression image into the pre-trained feature encoder respectively, to output first expression feature data corresponding to the anchor expression image, second expression feature data corresponding to the positive sample expression image, and third expression feature data corresponding to the negative sample expression image respectively through the pre-trained feature encoder;

[0217] Adjust the model parameters of the pre-trained feature encoder to make the difference between the second expression feature data and the first expression feature data less than a second threshold, and the difference between the third expression feature data and the first expression feature data greater than a third threshold, to obtain an optimized feature encoder; where the second threshold is less than the third threshold;

[0218] Determine the optimized feature encoder as the pre-trained expression extraction model.

[0219] Optionally, the training unit is specifically configured to:

[0220] Obtain a first expression triple with labeled expression similarity and facial expression images without labels;

[0221] Construct a second expression triple according to the facial expression images and label the expression similarity of the second expression triple to obtain a second expression triple with labeled expression similarity, where the first expression triple includes a first number of asymmetric expression samples, and the second expression triple includes a second number of asymmetric expression triples, and the first number is less than the second number;

[0222] Construct an expression triple according to the first expression triple and the second expression triple with labeled expression similarity.

[0223] Optionally, the determining unit 502 is specifically configured to:

[0224] Determine the phoneme sequence and text information corresponding to the source audio;

[0225] Determine the emotion information and prosody information corresponding to the phoneme sequence;

[0226] Determine the audio feature data corresponding to the source audio according to the emotion information, the prosody information and the text information.

[0227] Optionally, the generating unit 503 is specifically configured to:

[0228] Generate the initial facial control parameters corresponding to the target facial model according to the facial expression feature data and the identity feature data;

[0229] Perform detailed adjustment on the initial facial control parameters according to the emotion information, the prosody information and the text information included in the audio feature data, and generate the facial control parameters corresponding to the target facial model.

[0230] Optionally, the generating unit 503 is specifically configured to:

[0231] Input the facial expression feature data, the audio feature data and the identity feature data into a pre-trained parameter generation model, so as to generate the facial control parameters corresponding to the target facial model through the parameter generation model.

[0232] Optionally, the driving unit 504 is specifically configured to:

[0233] Generate an animation for the target facial model to present the first facial expression matching the emotion information and the mouth shape corresponding to the source audio.

[0234] Optionally, the training unit is specifically configured to train the pre-trained parameter generation model in the following manner:

[0235] Obtain a sample video of a sample speaker speaking with at least one second facial expression;

[0236] Determine the sample facial control parameters corresponding to the target facial model according to the sample video, where the sample facial control parameters are used to control the target facial model to present the second facial expression;

[0237] Determine the sample facial expression feature data corresponding to the target facial model when presenting the second facial expression according to the sample facial control parameters;

[0238] Determine the sample audio feature data corresponding to the sample audio of the sample video;

[0239] Input the sample expression feature data, the sample audio feature data, and the identity feature data corresponding to the target face model into a parameter generation model to be trained, so as to output predicted face control parameters through the parameter generation model;

[0240] Adjust the model parameters of the parameter generation model to make the difference between the predicted face control parameters and the sample face control parameters less than a fourth threshold, and obtain a pre-trained parameter generation model.

[0241] Optionally, the training unit is specifically configured to:

[0242] Identify the first facial key points corresponding to the sample speaker;

[0243] Determine the position data corresponding to each frame of the first facial key points in the sample video;

[0244] Determine the sample face control parameters corresponding to the target face model presenting the second expression according to the corresponding relationship between the position data, the first facial key points, and the second facial key points corresponding to the target face model.

[0245] Optionally, the training unit is specifically configured to:

[0246] Adjust the positions of the vertices in the target face model according to the sample face control parameters to obtain a sample animation of the target face model;

[0247] Obtain the sample face images corresponding to each frame in the sample animation of the target face model;

[0248] Determine the sample expression feature data corresponding to each frame of the sample face images respectively.

[0249] Optionally, the training unit is specifically configured to:

[0250] Slice the sample audio corresponding to the sample video into single-frame sample audio corresponding to each frame according to the frame rate of the sample video;

[0251] Determine the sample audio feature data corresponding to each frame of the single-frame sample audio respectively.

[0252] Optionally, the training unit is specifically configured to:

[0253] For the i-th frame, input the sample expression feature data corresponding to the i-th frame, the sample audio feature data corresponding to the i-th frame, and the identity feature data corresponding to the target face model into the parameter generation model to be trained, so as to output the predicted face control sub-parameters corresponding to the i-th frame through the parameter generation model, where i traverses 1 to N, and N is the number of video frames of the sample video;

[0254] Adjust the model parameters of the parameter generation model so that the difference between the predicted facial control sub-parameters corresponding to the i-th frame and the sample facial control sub-parameters corresponding to the i-th frame in the sample facial control parameters is less than a fourth threshold, and obtain a pre-trained parameter generation model.

[0255] Corresponding to the method for expression transfer provided in the first embodiment of the present application, the third embodiment of the present application further provides an electronic device for expression transfer.

[0256] As Figure 6 shown, it is a structural block diagram of an example of an electronic device for data processing provided by an embodiment of the present application.

[0257] In this embodiment, an optional hardware structure of the electronic device 600 can be as Figure 6 shown, including: at least one processor 601, at least one memory 602, and at least one communication bus 605; the memory 602 contains a program 603 and data 604.

[0258] The bus 605 can be a communication device for transmitting data between components inside the electronic device 600, such as an internal bus (for example, a CPU-memory bus, where the processor is the central processing unit, abbreviated as CPU), an external bus (for example, a universal serial bus port, a peripheral component interconnect express port), etc.

[0259] In addition, the electronic device further includes: at least one network interface 606, at least one peripheral interface 607. The network interface 606 provides wired or wireless communication related to an external network 608 (for example, the Internet, an intranet, a local area network, a mobile communication network, etc.); in some embodiments, the network interface 606 can include any combination of any number of network interface controllers (English: network interface controller, abbreviated as NIC), radio frequency (English: Radio Frequency, abbreviated as RF) modules, repeaters, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication (English: Near Field Communication, abbreviated as NFC) adapters, cellular network chips, etc.

[0260] The peripheral interface 607 is used to connect to peripherals, and the peripherals can be peripheral 1 ( Figure 6 in 609), peripheral 2 ( Figure 6 in 610), and peripheral 3 ( Figure 6in 611). A peripheral device is a device that is not part of the core computer system but is connected to it. Peripheral devices may include, but are not limited to, cursor control devices (such as mice, touchpads, or touchscreens), keyboards, displays (such as cathode ray tube displays, liquid crystal displays, light-emitting diode displays), video input devices (such as cameras or input interfaces communicatively coupled to video archives), etc.

[0261] The processor 601 may be a CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application.

[0262] The memory 602 may include high-speed RAM (Random Access Memory) memory, and may also include non-volatile memory, such as at least one disk memory.

[0263] Wherein, the processor 601 calls the programs and data stored in the memory 602 and executes the following steps:

[0264] Obtain a source face image containing a first expression, source audio, and a target face model to be subject to expression migration;

[0265] Determine the expression feature data corresponding to the source face image, the audio feature data corresponding to the source audio, and the identity feature data corresponding to the target face model, wherein the audio feature data includes the emotional information corresponding to the source audio;

[0266] Generate face control parameters corresponding to the target face model according to the expression feature data, the audio feature data, and the identity feature data;

[0267] According to the face control parameters, adjust the positions of the vertices in the target face model to generate an animation of the first expression presenting expression details matching the emotional information for the target face model.

[0268] Corresponding to the expression migration method provided in the first embodiment of the present application, the fourth embodiment of the present application provides a computer-readable storage medium storing a program for the expression migration method, and the program is run by a processor to execute the following steps:

[0269] Obtain a source face image containing a first expression, source audio, and a target face model to be subject to expression migration;

[0270] Determine the expression feature data corresponding to the source face image, the audio feature data corresponding to the source audio, and the identity feature data corresponding to the target face model, where the audio feature data includes the emotional information corresponding to the source audio;

[0271] Generate the face control parameters corresponding to the target face model according to the expression feature data, the audio feature data, and the identity feature data;

[0272] Adjust the positions of the vertices in the target face model according to the face control parameters, so as to generate an animation of the first expression presenting expression details matching the emotional information for the target face model.

[0273] It should be noted that for the detailed descriptions of the devices, electronic devices, and computer-readable storage media provided in the second, third, and fourth embodiments of the present application, reference can be made to the relevant descriptions of the first embodiment of the present application, which will not be repeated here.

[0274] Although the present application is disclosed above with preferred embodiments, it is not used to limit the present application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the protection scope of the present application should be subject to the scope defined by the claims of the present application.

[0275] In a typical configuration, the node devices in the blockchain include one or more processors (CPUs), input / output interfaces, network interfaces, and memories.

[0276] The memory may include non-permanent memory in the computer-readable medium, random access memory (RAM), and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash memory (flash RAM). The memory is an example of the computer-readable medium.

[0277] 1. A computer-readable medium includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), random access memory (RAM) of other properties, read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage media, or any other non-transitory medium that can be used to store information accessible by a computing device. As defined herein, a computer-readable medium does not include transitory media such as modulated data signals and carrier waves.

[0278] 2. Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0279] Although the present application is disclosed above in preferred embodiments, it is not intended to limit the present application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the protection scope of the present application should be determined by the scope defined by the claims of the present application.

Claims

1. A method for expression transfer, characterized in that: The method comprises: Acquire a source facial image including a first expression, a source audio, and a target facial model to be subjected to expression transfer; Determine expression feature data corresponding to the source facial image, audio feature data corresponding to the source audio, and identity feature data corresponding to the target facial model, wherein the audio feature data includes emotion information corresponding to the source audio; Generate facial control parameters corresponding to the target facial model according to the facial expression feature data, the audio feature data and the identity feature data; The positions of the vertices in the target facial model are adjusted according to the facial control parameters to generate an animation for the target facial model that presents the first expression with expression details matching the emotion information.

2. The method according to claim 1, characterized in that The determining of the expression feature data corresponding to the source facial image comprises: The source facial image is input into a pre-trained expression extraction model to output corresponding expression feature data through the expression extraction model.

3. The method according to claim 2, characterized in that The pre-trained expression extraction model is trained in the following way: Get a sample face image; Randomly masking the sample facial image to obtain a processed sample facial image; Inputting the processed sample facial image into the feature encoder to be trained, so as to output sample expression feature data of the processed sample image through the feature encoder; Determining a predicted facial image corresponding to the sample expression feature data; Adjusting the model parameters of the feature encoder so that the difference between the predicted facial image and the sample facial image is less than a first threshold, thereby obtaining a pre-trained feature encoder; A pre-trained expression extraction model is determined according to the pre-trained feature encoder.

4. The method according to claim 3, characterized in that The step of determining the predicted facial image corresponding to the sample expression feature data comprises: The sample expression feature data is input into a pre-trained feature decoder so as to output a corresponding predicted facial image through the feature decoder.

5. The method according to claim 3, characterized in that: The step of determining a pre-trained expression extraction model according to the pre-trained feature encoder comprises: Constructing expression triples, wherein each of the expression triples includes an anchor expression image, a positive sample expression image, and a negative sample expression image; Inputting the anchor expression image, the positive sample expression image and the negative sample expression image into the pre-trained feature encoder respectively, so as to output the first expression feature data corresponding to the anchor expression image, the second expression feature data corresponding to the positive sample expression image and the third expression feature data corresponding to the negative sample expression image respectively through the pre-trained feature encoder; Adjusting the model parameters of the pre-trained feature encoder so that the difference between the second expression feature data and the first expression feature data is less than a second threshold, and the difference between the third expression feature data and the first expression feature data is greater than a third threshold, to obtain an optimized feature encoder; wherein the second threshold is less than the third threshold; The optimized feature encoder is determined as a pre-trained expression extraction model.

6. The method according to claim 5, characterized in that The constructing of the expression triplet includes: Acquire a first expression triplet with annotated expression similarity and a facial expression image without a label; Constructing a second expression triplet according to the facial expression image, and marking the second expression triplet with expression similarity to obtain a second expression triplet with marked expression similarity, wherein the first expression triplet includes a first number of asymmetric expression samples, and the second expression triplet includes a second number of asymmetric expression triplets, and the first number is less than the second number; An expression triplet is constructed according to the first expression triplet and the second expression triplet with the annotated expression similarity.

7. The method according to claim 1, characterized in that The emotional information includes emotional information and rhythmic information. The step of determining the audio feature data corresponding to the source audio includes: Determine the phoneme sequence and text information corresponding to the source audio; Determining emotional information and rhythmic information corresponding to the phoneme sequence; The audio feature data corresponding to the source audio is determined according to the emotion information, the rhythm information and the text information.

8. The method according to claim 7, characterized in that The step of generating the facial control parameters corresponding to the target facial model according to the facial expression feature data, the audio feature data and the identity feature data comprises: Generating initial facial control parameters corresponding to the target facial model according to the facial expression feature data and the identity feature data; According to the emotion information, the rhythm information and the text information contained in the audio feature data, the initial facial control parameters are adjusted in detail to generate facial control parameters corresponding to the target facial model.

9. The method according to claim 1, characterized in that: The step of generating the facial control parameters corresponding to the target facial model according to the facial expression feature data, the audio feature data and the identity feature data comprises: The expression feature data, the audio feature data and the identity feature data are input into a pre-trained parameter generation model, so as to generate facial control parameters corresponding to the target facial model through the parameter generation model.

10. The method according to claim 1, characterized in that The step of generating an animation for presenting the first expression matching the emotion information for the target facial model comprises: An animation is generated for the target facial model to present the first expression matching the emotional information and the lip shape corresponding to the source audio.

11. The method according to claim 9, characterized in that The pre-trained parameter generation model is trained in the following way: Obtaining a sample video of a sample speaker speaking with at least one second expression; Determining, according to the sample video, sample facial control parameters corresponding to the target facial model, wherein the sample facial control parameters are used to control the target facial model to present the second expression; Determining, according to the sample face control parameters, sample expression feature data corresponding to when the target face model presents the second expression; Determining sample audio feature data corresponding to the sample audio corresponding to the sample video; Inputting the sample expression feature data, the sample audio feature data and the identity feature data corresponding to the target face model into a parameter generation model to be trained, so as to output predicted facial control parameters through the parameter generation model; The model parameters of the parameter generation model are adjusted so that the difference between the predicted facial control parameter and the sample facial control parameter is smaller than a fourth threshold, thereby obtaining a pre-trained parameter generation model.

12. The method according to claim 11, characterized in that Determining the sample face control parameters corresponding to the target face model according to the sample video includes: Identifying a first facial key point corresponding to the sample speaker; Determine the position data corresponding to each frame of the sample video of the first facial key point; According to the correspondence between the position data, the first facial key points and the second facial key points corresponding to the target facial model, the sample facial control parameters corresponding to when the target facial model presents the second expression are determined.

13. The method according to claim 11, characterized in that The step of determining, according to the sample face control parameter, sample expression feature data corresponding to when the target face model presents the second expression comprises: According to the sample face control parameters, adjusting the positions of vertices in the target face model to obtain a sample animation of the target face model; Obtaining a sample facial image corresponding to each frame of the target facial model in the sample animation; The sample expression feature data corresponding to each frame of the sample facial image is determined respectively.

14. The method according to claim 11, characterized in that The determining of the sample audio feature data corresponding to the sample audio corresponding to the sample video includes: According to the frame rate of the sample video, the sample audio corresponding to the sample video is divided into single-frame sample audio corresponding to each frame; The sample audio feature data corresponding to each frame of the single-frame sample audio is determined respectively.

15. The method according to claim 11, characterized in that The step of inputting the sample expression feature data, the sample audio feature data and the identity feature data corresponding to the target face model into a parameter generation model to be trained, so as to output predicted face control parameters through the parameter generation model, comprises: For the i-th frame, the sample expression feature data corresponding to the i-th frame, the sample audio feature data corresponding to the i-th frame, and the identity feature data corresponding to the target face model are input into the parameter generation model to be trained, so as to output the predicted facial control sub-parameters corresponding to the i-th frame through the parameter generation model, wherein i traverses 1 to N, and N is the number of video frames of the sample video; The model parameters of the parameter generation model are adjusted so that the difference between the predicted facial control parameter and the sample facial control parameter is less than a fourth threshold, to obtain a pre-trained parameter generation model: The model parameters of the parameter generation model are adjusted so that the difference between the predicted facial control sub-parameter corresponding to the i-th frame and the sample facial control sub-parameter corresponding to the i-th frame in the sample facial control parameters is less than a fourth threshold, thereby obtaining a pre-trained parameter generation model.

16. A device for expression transfer, characterized in that: The device comprises: An acquisition unit, used to acquire a source facial image containing a first expression, a source audio, and a target facial model to be subjected to expression migration; A determination unit, configured to determine expression feature data corresponding to the source facial image, audio feature data corresponding to the source audio, and identity feature data corresponding to the target facial model, wherein the audio feature data includes emotion information corresponding to the source audio; A generating unit, configured to generate a facial control parameter corresponding to the target facial model according to the facial expression feature data, the audio feature data and the identity feature data; An adjustment unit is used to adjust the positions of vertices in the target facial model according to the facial control parameters, so as to generate an animation for the target facial model that presents the first expression with expression details matching the emotion information.

17. An electronic device, characterized in that: include: processor; as well as The memory is used to store a data processing program. After the electronic device is powered on and the program is run by the processor, the method according to any one of claims 1 to 15 is executed.

18. A computer-readable storage medium, characterized in that: A data processing program is stored, and the program is run by a processor to execute the method according to any one of claims 1 to 15.