Animation generation method and device, electronic equipment and storage medium

By acquiring the speech features and fusion deformation parameter sequence of irrelevant semantics from speech data, facial animation is generated using a generative model and decoder. This solves the problems of insufficient matching between lip-sync animation and speech data and insufficient richness of facial expressions in existing technologies, and achieves high-quality animation generation.

CN121053261APending Publication Date: 2025-12-02BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410693291.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-30
Publication Date
2025-12-02

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve a good match between lip-sync animation and speech data, as well as to provide richness and diversity in facial expression animation when generating lip-sync animations that correspond to speech data.

Method used

By acquiring the speech features of the speech data and the fusion deformation parameter sequence that is semantically independent of the speech data, a pre-defined generation model is used to generate an encoding sequence, which is then decoded into a second fusion deformation parameter sequence by a pre-defined decoder, driving the object model to generate a facial animation corresponding to the speech data.

Benefits of technology

It achieves matching between lip-sync animation and voice data, while ensuring the richness and diversity of facial expression animation, thus improving the accuracy and variety of animation generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121053261A_ABST
    Figure CN121053261A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an animation generation method and device, electronic equipment and a storage medium. The method comprises the steps that voice features of voice data and a first fusion deformation parameter sequence are acquired; wherein the first fusion deformation parameter sequence is irrelevant to semantics of the voice data; generating a coding sequence according to the voice features and the first fusion deformation parameter sequence through a preset generation model; decoding the coding sequence into a second fusion deformation parameter sequence through a preset decoder; and driving the object model according to the second fusion deformation parameter sequence to generate a facial animation corresponding to the voice data. Animation generation based on fusion deformation parameters can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to an animation generation method, apparatus, electronic device, and storage medium. Background Technology

[0002] Currently, technologies for generating lip-sync animations corresponding to speech data are widely used in various fields. Existing technologies often generate animations based on vertex-driven methods. Summary of the Invention

[0003] This disclosure provides an animation generation method, apparatus, electronic device, and storage medium that can realize animation generation based on fused deformation parameters.

[0004] In a first aspect, embodiments of this disclosure provide an animation generation method, including:

[0005] Acquire speech features and a first fusion deformation parameter sequence from the speech data; wherein the first fusion deformation parameter sequence is semantically independent of the speech data;

[0006] A coding sequence is generated based on the speech features and the first fusion deformation parameter sequence using a preset generation model.

[0007] The encoded sequence is decoded into a second fused deformation parameter sequence using a preset decoder;

[0008] The object model is driven according to the second fusion deformation parameter sequence to generate a facial animation corresponding to the speech data.

[0009] Secondly, embodiments of this disclosure also provide an animation generation apparatus, comprising:

[0010] The acquisition module is used to acquire the speech features and a first fusion deformation parameter sequence of the speech data; wherein, the first fusion deformation parameter sequence is semantically independent of the speech data;

[0011] The encoding generation module is used to generate an encoding sequence based on the speech features and the first fusion deformation parameter sequence using a preset generation model;

[0012] The decoding module is used to decode the encoded sequence into a second fused deformation parameter sequence using a preset decoder;

[0013] An animation generation module is used to drive the object model according to the second fusion deformation parameter sequence to generate a facial animation corresponding to the voice data.

[0014] Thirdly, embodiments of this disclosure also provide an electronic device, the electronic device comprising:

[0015] One or more processors;

[0016] Storage device for storing one or more programs.

[0017] When the one or more programs are executed by the one or more processors, the one or more processors implement the animation generation method as described in any of the embodiments of this disclosure.

[0018] Fourthly, embodiments of this disclosure also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the animation generation method as described in any of the embodiments of this disclosure.

[0019] The technical solution of this disclosure embodiment involves acquiring speech features and a first fusion deformation parameter sequence from speech data; wherein the first fusion deformation parameter sequence is semantically independent of the speech data; generating an encoding sequence based on the speech features and the first fusion deformation parameter sequence using a preset generation model; decoding the encoding sequence into a second fusion deformation parameter sequence using a preset decoder; and driving an object model based on the second fusion deformation parameter sequence to generate a facial animation corresponding to the speech data.

[0020] By utilizing a pre-built generative model, an encoded sequence is generated based on speech features and a first fusion deformation parameter sequence. This encoded sequence is then decoded into a second fusion deformation parameter sequence using a pre-built decoder, enabling the creation of an object model driven by the fusion deformation parameter sequence. Because the pre-built generative model possesses strong generative capabilities, it can generate encoded sequences that adapt to lip movements and conform to the style of the first fusion deformation parameter sequence. This ensures that while lip-syncing animation matches the speech data, it also guarantees the richness and diversity of facial expression animations in dimensions other than lip movements. Attached Figure Description

[0021] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.

[0022] Figure 1 This is a flowchart illustrating an animation generation method provided in an embodiment of the present disclosure;

[0023] Figure 2 This is a schematic block diagram of the data flow for an animation generation method provided in an embodiment of the present disclosure;

[0024] Figure 3This is a schematic diagram illustrating the process of constructing a preset vector quantization model in an animation generation method provided in this embodiment of the present disclosure;

[0025] Figure 4 This is a schematic block diagram of the data flow in an animation generation method provided in this embodiment of the present disclosure, where a preset vector quantization model is constructed.

[0026] Figure 5 This is a schematic diagram illustrating the process of constructing a preset generation model in an animation generation method provided in this embodiment of the present disclosure;

[0027] Figure 6 This is a schematic block diagram of the data flow for constructing a preset generation model in an animation generation method provided in this embodiment of the present disclosure;

[0028] Figure 7 This is a schematic diagram of the structure of an animation generation apparatus provided in an embodiment of the present disclosure;

[0029] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0030] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0031] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0032] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0033] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0034] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0035] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0036] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0037] Figure 1 This is a schematic flowchart illustrating an animation generation method provided in an embodiment of this disclosure. This embodiment is applicable to generating facial animations that correspond to lip movements in speech data. The method can be executed by an animation generation device, which can be implemented in software and / or hardware and can be configured in an electronic device, such as a mobile phone or computer.

[0038] like Figure 1 As shown, the animation generation method provided in this embodiment may include:

[0039] S110. Acquire the speech features of the speech data and the first fusion deformation parameter sequence; wherein, the first fusion deformation parameter sequence is independent of the semantics of the speech data.

[0040] In this embodiment of the disclosure, the voice data may include voice data of any duration. Voice data input by the user can be acquired in real time via an audio acquisition module, or historically recorded voice data can be read from a preset storage space. Other methods of acquiring voice data are also possible, and these will not be exhaustively listed here.

[0041] In this embodiment of the disclosure, for speech data A, the corresponding speech features F can be obtained through an existing feature extractor. A ∈R T,C Where T represents the number of speech features corresponding to the speech data per second, for example, T can be 25; where C represents the dimension of each speech feature.

[0042] For example, Figure 2 This is a schematic block diagram of the data flow for an animation generation method provided in an embodiment of this disclosure. See also... Figure 2 In some implementations, the process of acquiring speech features from speech data may include: extracting speech features from speech data using at least two feature extraction algorithms; and determining the final speech features of speech data based on the extracted at least two speech features.

[0043] Among them, at least two feature extraction algorithms are included, which may include existing audio feature extraction algorithms. For example... Figure 2 This study utilizes three feature extraction algorithms: Automatic Speech Recognition (ASR) streaming feature extraction, ASR non-streaming feature extraction, and lip-sync feature extraction. The ASR streaming / non-streaming feature extraction algorithms can be implemented based on the encoder in the neural network model corresponding to the ASR task; the lip-sync feature extraction algorithm can be implemented based on the encoder in the neural network model corresponding to the audio visualization task.

[0044] By extracting features from the speech data using at least two feature extraction algorithms, at least two speech features can be obtained. Based on these at least two speech features, the final speech features of the speech data can be determined. For example, the at least two speech features can be concatenated to obtain the final speech features; alternatively, the at least two speech features can be further encoded to obtain the final speech features, and so on. Obtaining the final speech features of the speech data based on multiple speech features can improve the accuracy of the generated animation.

[0045] In related technologies, a blendshape (BS) deformer can deform a base shape into a target shape by applying deformable targets (also called shape keys) of different shapes to a base shape. For example, the base shape may be an expressionless face, and the deformable targets may include a face with raised eyebrows, an open jaw, closed eyes, and upturned corners of the mouth, etc. The deformable targets applied to the base shape can be set with intensity coefficients, and interpolation operations between the deformable targets and the base shape are performed based on these intensity coefficients to obtain the target shape. The blend shape parameters can be constituted by a set of intensity coefficients of the deformable targets applied to the base shape. For example, the blend shape parameters may include 51-dimensional intensity coefficients. The blend shape parameters can be used on any two-dimensional or three-dimensional object model with a predefined blend shape deformer.

[0046] In this embodiment, the fusion deformation parameter sequence may include a sequence composed of multiple fusion deformation parameters. The first fusion deformation parameter sequence is semantically independent of the speech data. This can be understood as the lip movements in the animation generated based on the first fusion deformation parameter sequence not corresponding to the lip movements in the speech data; that is, the animation generated based on the first fusion deformation parameter sequence does not express the semantics corresponding to the speech data. By obtaining a first fusion deformation parameter sequence that is semantically independent of the speech data, the speech data and facial expression methods can be decoupled, which is beneficial for achieving diverse animation generation.

[0047] Although the first fusion deformation parameter sequence is semantically unrelated to the speech data, it can be related to the object from which the speech data was collected. For example, assuming speech data A and speech data B of user A were collected, speech data A can be used as the speech data in this embodiment, and the fusion deformation parameter sequence corresponding to speech data B can be used as the first fusion deformation parameter sequence. Thus, the generated second fusion deformation parameter sequence not only corresponds to the speech data but also maintains user A's facial expression, achieving a consistent presentation effect in the animation where both the speech and facial expression belong to user A. In this case, the fusion deformation parameter sequence of another segment of speech data from the same object can be used as the first fusion deformation parameter sequence.

[0048] Furthermore, the first fusion deformation parameter sequence can also be independent of the object from which the voice data was collected. For example, assuming voice data A from user A and voice data B from user B were collected, voice data A can be used as the voice data in this embodiment, and the fusion deformation parameter sequence corresponding to user B's voice data B can be used as the first fusion deformation parameter sequence. Thus, the generated second fusion deformation parameter sequence, while corresponding to the voice data, can mimic user B's facial expressions, achieving a diverse presentation effect where the voice in the animation belongs to user A, while the facial expressions belong to user B. In this case, the fusion deformation parameter sequence of another voice data segment from an object different from the object from which the voice data was collected can be used as the first fusion deformation parameter sequence.

[0049] In this embodiment, voice data from different objects can be pre-collected, and the corresponding first fusion deformation parameter sequence can be determined and stored based on the collected voice data. Accordingly, based on the user's selection operation, the desired first fusion deformation parameter can be selected from the pre-stored first fusion deformation parameter sequences as the generation condition for generating the second fusion deformation parameter corresponding to the voice data.

[0050] S120. Generate an encoding sequence based on speech features and the first fusion deformation parameter sequence using a preset generation model.

[0051] In this embodiment of the disclosure, the encoded sequence can be understood as an encoded form of the fused deformation parameter sequence. The preset generation model can include an existing neural network model with encoding generation capabilities, such as a generative pre-trained transformer model. The preset generation model can be pre-built and can have the ability to generate an encoded sequence based on the input speech features and the first fused deformation parameter sequence.

[0052] In some optional implementations, generating an encoding sequence based on speech features and a first fusion deformation parameter sequence may include: extracting features from the first fusion deformation parameter sequence to obtain a style vector; wherein the style vector represents the facial expression of the object model; and generating an encoding sequence based on the speech features and the style vector.

[0053] In this embodiment of the disclosure, the object model may include at least one of the following: a two-dimensional object model and a three-dimensional object model. The two-dimensional object model may include a realistic facial model and a simulated facial model; the three-dimensional object model may be a pre-built or real-time built stereoscopic head model.

[0054] The object model can include a model with a similar shape to the object from which the voice data is collected, or a model with an arbitrary shape determined from multiple preset models based on user input. Alternatively, an object model with a similar shape can be generated based on the shape of the object from which the voice data is collected, using existing object model generation methods.

[0055] In this embodiment of the disclosure, facial expression methods may include the way the object model presents itself, such as facial expressions and lip movements, when speaking. Specifically, features can be extracted from the first fused deformation parameter sequence based on existing sequence feature extraction methods, and the resulting feature vector can be used as a style vector; this style vector can then be used to characterize the facial expression method. For example, see [link to example]. Figure 2 The first fused deformation parameter sequence can be compressed into a vector using stacked Multilayer Perceptron (MLP) modules in the preset generative model, thereby enabling feature extraction from the first fused deformation parameter sequence. The stacking of MLP modules can be understood as multiple MLP modules being set sequentially. The style vector can be used as the initial output, or starting value, of the preset generative model.

[0056] For example, see Figure 2 The speech features can be compressed to a suitable size using the MLPs module in the preset generative model. The compressed speech features and style vectors can then be input into the transformer decoder module of the preset generative model to output an encoded sequence. The transformer decoder module can encode and generate encoded sequences that are adapted to lip movements and conform to style vectors.

[0057] S130. The encoded sequence is decoded into a second fused deformation parameter sequence using a preset decoder.

[0058] In this embodiment, the second fused deformation parameter sequence is semantically related to the speech data; that is, the lip movements generated in the animation based on the first fused deformation parameter sequence correspond to the lip movements in the speech data. The preset decoder may include an existing neural network model with decoding capabilities. See also... Figure 2 The preset decoder can be pre-built and has the ability to decode the input encoded sequence into a second fused deformation parameter sequence.

[0059] S140. Drive the object model according to the second fusion deformation parameter sequence to generate facial animation corresponding to the speech data.

[0060] In this embodiment, an object model with a predefined fusion deformation deformer can be driven based on a second fusion deformation parameter sequence. This allows for the generation of facial animation from an object model driven by speech data, where the object model in the facial animation displays lip movements consistent with the speech. For example, if the speech data is "Hello," the object model in the output facial animation can display the lip movements corresponding to "Hello." The technical solution of this embodiment involves acquiring the speech features of the speech data and a first fusion deformation parameter sequence; wherein the first fusion deformation parameter sequence is semantically independent of the speech data; generating an encoding sequence based on the speech features and the first fusion deformation parameter sequence using a preset generation model; decoding the encoding sequence into a second fusion deformation parameter sequence using a preset decoder; and driving the object model according to the second fusion deformation parameter sequence to generate a facial animation corresponding to the speech data.

[0061] By utilizing a pre-built generative model, an encoded sequence is generated based on speech features and a first fusion deformation parameter sequence. This encoded sequence is then decoded into a second fusion deformation parameter sequence using a pre-built decoder, enabling the creation of an object model driven by the fusion deformation parameter sequence. Because the pre-built generative model possesses strong generative capabilities, it can generate encoded sequences that adapt to lip movements and conform to the style of the first fusion deformation parameter sequence. This ensures that while lip-syncing animation matches the speech data, it also guarantees the richness and diversity of facial expression animations in dimensions other than lip movements.

[0062] This disclosure embodiment can be combined with various optional schemes in the animation generation method provided in the above embodiments. The animation generation method provided in this embodiment details the construction process of the preset decoder. In the animation generation method provided in this disclosure embodiment, the preset decoder is included in a preset vector quantization (VQ) model, and the preset vector quantization model is constructed based on a third fusion deformation parameter sequence semantically related to the sample speech data. By encoding, feature replacement, and reconstruction of the third fusion deformation parameter sequence, and making the reconstructed fusion deformation parameter sequence approximate the third fusion deformation parameter sequence, the preset vector quantization model can be constructed; that is, the preset decoder in the preset vector quantization model can be constructed simultaneously. Based on the constructed preset decoder, the encoded sequence can be decoded into a fusion deformation parameter sequence.

[0063] Figure 3 This is a schematic diagram illustrating the process of constructing a preset vector quantization model in an animation generation method provided by an embodiment of this disclosure. Figure 3 As shown, the animation generation method provided in this embodiment further includes a preset encoder and a codebook in the preset vector quantization model; the construction process of the preset vector quantization model may include:

[0064] S310. The third fused deformation parameter sequence is encoded by a preset encoder to obtain the predicted value of the first encoded sequence.

[0065] In this embodiment, the third fusion deformation parameter sequence is semantically related to the sample speech data and can be considered as the ground truth of the fusion deformation parameter sequence corresponding to the sample speech data. This can be achieved by acquiring the user's facial shape in real time during the sample speech data acquisition process and determining the third fusion deformation parameter sequence based on the acquired facial shape. Alternatively, other methods, such as manually adjusting model vertices, can also be used to obtain the third fusion deformation parameter sequence; these are not exhaustive examples here. The acquisition of sample speech data and facial shapes should comply with relevant laws, regulations, and related provisions.

[0066] For example, Figure 4 This is a schematic block diagram illustrating the data flow of a preset vector quantization model in an animation generation method provided in this embodiment of the disclosure. See also... Figure 4 Given a third fusion deformation parameter sequence Y∈R containing fusion deformation parameters of T frames. T,51 It can be encoded by the preset encoder in the VQ model to obtain the first coded sequence prediction value c∈R. T .

[0067] Because the sample speech data is temporally sequential, the corresponding fusion deformation parameter sequence also exhibits temporal continuity. The process of encoding the fusion deformation parameter sequence using a pre-defined encoder can be considered as classifying the fusion deformation parameters of each frame, thus obtaining discrete classifications of the fusion deformation parameters for each frame. The target shape corresponding to each class of fusion deformation parameters can be considered to possess homogeneity in shape features. Therefore, the predicted value of the first encoded sequence can be called the discrete predicted value of the first encoded sequence.

[0068] S320. From the codebook, determine the feature vectors that are similar to each code in the predicted value of the first coding sequence, and determine the coding features based on each similar feature vector.

[0069] In this embodiment, the codebook can be viewed as an M×N dimension matrix, and the codebook can contain m (e.g., 0≤m≤1024) preset feature vectors, each of which can have a dimension of N (e.g., 512). Each preset feature vector can represent a category of the fused deformation parameters.

[0070] Specifically, based on existing vector similarity calculation methods, the predicted value c∈R of the first encoded sequence can be determined from the codebook. T For each encoded feature vector, the most similar feature vectors can be obtained, and these most similar feature vectors can be concatted to obtain the encoded features C∈R. T,512 .

[0071] S330. The encoded features are decoded by a preset decoder to obtain the predicted value of the third fused deformation parameter sequence.

[0072] See you again Figure 4 Given the encoded features C∈R T,512 The fused deformation parameter sequence can be reconstructed using the preset decoder in the VQ model to obtain the predicted value of the third fused deformation parameter sequence.

[0073] S340. Determine the reconstruction loss based on the third fused deformation parameter sequence and the predicted value of the third fused deformation parameter sequence.

[0074] In this embodiment, the third fused deformation parameter sequence Y∈R can be determined based on at least one existing loss function. T,51 and the predicted values ​​of the third fused deformation parameter sequence The reconstruction loss between [the two]. For example, it can be determined through a loss function. Determine Y and The loss. For example, it can be determined through a loss function. Determine Y and The loss; where V can characterize the computation speed of Y, Characterizable The calculation speed; where the calculation speed can represent the change of the fused deformation parameters in the next frame compared to the fused deformation parameters in the previous frame.

[0075] S350. Adjust the parameters of the preset encoder, preset decoder and codebook according to the reconstruction loss.

[0076] In this embodiment, the reconstruction loss can be backpropagated to adjust the parameters of the preset encoder, the parameters of the preset decoder, and the feature vectors in the codebook, thereby constructing the VQ model. In this embodiment, an existing optimizer can be used for iterative optimization, with the learning rate, training batch size, and number of training iterations set based on the actual application scenario.

[0077] In some optional implementations, the process of constructing the preset vector quantization model may also include: constructing a first coding loss based on the predicted value of the first coding sequence, and adjusting the parameters of the preset encoder based on the first coding loss.

[0078] Among them, the value c∈R can be predicted based on the first encoded sequence using the loss function CL(c). T The first encoding loss. Here, CL(·) can represent commitment loss, which can be used to increase the difference between encodings.

[0079] In this case, see Figure 4 The VQ model is constructed based on the following optimization objectives:

[0080]

[0081] in, CL(c) can represent the reconstruction loss; CL(c) can represent the first encoding loss.

[0082] The technical solution of this disclosure describes in detail the construction process of the preset decoder. In the animation generation method provided by this disclosure, the preset decoder is included in a preset vector quantization model, and the preset vector quantization model is constructed based on a third fusion deformation parameter sequence semantically related to the sample speech data. By encoding, feature replacement, and reconstruction of the third fusion deformation parameter sequence, and by making the reconstructed fusion deformation sequence approximate the third fusion deformation parameter sequence, the preset vector quantization model can be constructed; that is, the preset decoder in the preset vector quantization model can be constructed simultaneously. Based on the constructed preset decoder, the encoded sequence can be decoded into a fusion deformation parameter sequence.

[0083] Furthermore, the animation generation method provided in this embodiment belongs to the same disclosed concept as the animation generation method provided in the above embodiments. Technical details not described in detail in this embodiment can be found in the above embodiments, and the same technical features have the same beneficial effects in this embodiment and the above embodiments.

[0084] This embodiment can be combined with various optional schemes in the animation generation method provided in the above embodiments. The animation generation method provided in this embodiment details the construction process of the preset generation model. In the animation generation method provided in this embodiment, the preset generation model is constructed based on sample speech features of sample speech data, a fourth fusion deformation parameter sequence, and the ground truth value of the sample encoding sequence; wherein, the fourth fusion deformation parameter sequence is semantically independent of the sample speech data; wherein, the ground truth value of the sample encoding sequence is obtained by encoding the third fusion deformation parameter sequence through a preset encoder in the constructed vector quantization model.

[0085] In this embodiment, the construction process of the preset vector quantization model can be referred to as the first-stage construction process, and the construction process of the preset generation model can be referred to as the second-stage construction process. After the preset vector quantization model is constructed, the third fusion deformation parameter sequence can be encoded according to the preset encoder in the preset vector quantization model to obtain the true value of the sample encoded sequence. The sample speech features of the sample speech data can be obtained based on the speech feature acquisition method.

[0086] By constructing a preset generation model based on sample speech features, the fourth fusion deformation parameter sequence, and the true value of the sample coding sequence, the preset generation model can generate a coding sequence based on speech features and the fusion deformation sequence.

[0087] Figure 5 This is a schematic diagram illustrating the process of constructing a preset generation model in an animation generation method provided in this embodiment. Figure 5 As shown, the animation generation method provided in this embodiment, the preset model construction process, may include:

[0088] S510. Using a preset generation model, generate the second coding sequence prediction value based on the sample speech features and the fourth fusion deformation parameter sequence.

[0089] Because facial expressions vary significantly among different users, this embodiment incorporates a fourth fusion deformation parameter sequence as a style condition for generating the fusion deformation parameter sequence, in order to simultaneously utilize sample speech data from multiple users to construct a preset generation model. This fourth fusion deformation parameter sequence is semantically independent of the sample speech data; however, to extract facial expressions from the sample speech data, it can belong to a fusion deformation parameter sequence corresponding to another segment of speech data from the user who collected the sample speech data.

[0090] In some optional implementations, generating a second coding sequence prediction value based on the sample speech features and the fourth fusion deformation parameter sequence may include: extracting features from the fourth fusion deformation parameter sequence to obtain a sample style vector; and generating a second coding sequence prediction value based on the sample speech features and the sample style vector.

[0091] For example, Figure 6 This is a schematic block diagram illustrating the data flow for constructing a preset generation model in an animation generation method provided in this embodiment of the disclosure. See also... Figure 6 The fourth fusion deformation parameter sequence can be compressed into a vector using the MLPs module in the preset generative model to obtain the sample style vector. The sample speech features are then compressed to a suitable size using the MLPs module in the preset generative model. The compressed sample speech features and sample style vector can then be input into the transformer decoder module in the preset generative model to output the predicted value of the second encoded sequence.

[0092] S520. Determine the second coding loss based on the true value of the sample coding sequence and the predicted value of the second coding sequence.

[0093] In this embodiment, the second encoding loss can be determined based on an existing classification loss function. For example, a preset generative model can be constructed based on the following optimization objective: in, c can represent the predicted value of the second encoded sequence; c can represent the true value of the sample encoded sequence.

[0094] S530. Adjust the parameters of the preset generation model according to the second encoding loss.

[0095] In this embodiment, the second encoding loss can be backpropagated to adjust the parameters of various modules such as MLPs and transformer decoder in the preset generative model, thereby realizing the construction of the preset generative model. In this embodiment, an existing optimizer can be used for iterative optimization, and the learning rate, training batch, and training iterations can be set based on the actual application scenario.

[0096] The technical solution of this disclosure provides a detailed description of the construction process of the preset generation model. By constructing the preset generation model based on sample speech features, the fourth fusion deformation parameter sequence, and the true value of the sample coding sequence, the preset generation model can generate a coding sequence based on speech features and the fusion deformation sequence. Furthermore, the animation generation method provided in this disclosure belongs to the same disclosed concept as the animation generation method provided in the above embodiments. Technical details not described in detail in this embodiment can be found in the above embodiments, and the same technical features have the same beneficial effects in this embodiment and the above embodiments.

[0097] Figure 7 This is a schematic diagram of an animation generation apparatus provided in an embodiment of the present disclosure. The animation generation apparatus provided in this embodiment is suitable for generating facial animations that correspond to lip movements in speech data.

[0098] like Figure 7 As shown, the animation generation apparatus provided in this embodiment may include:

[0099] The acquisition module 710 is used to acquire the speech features and the first fusion deformation parameter sequence of the speech data; wherein, the first fusion deformation parameter sequence is independent of the semantics of the speech data;

[0100] The encoding generation module 720 is used to generate an encoding sequence based on speech features and a first fusion deformation parameter sequence using a preset generation model;

[0101] The decoding module 730 is used to decode the encoded sequence into a second fused deformation parameter sequence through a preset decoder;

[0102] Animation generation module 740 is used to drive the object model according to the second fusion deformation parameter sequence to generate facial animation corresponding to the voice data.

[0103] In some alternative implementations, the encoding generation module can be used for:

[0104] Features of the first fused deformation parameter sequence are extracted to obtain a style vector; whereby the style vector represents the facial expression of the object model.

[0105] Generate a coding sequence based on speech features and style vectors.

[0106] In some alternative implementations, the preset decoder is contained in a preset vector quantization model, which is constructed based on a third fusion deformation parameter sequence that is semantically related to the sample speech data.

[0107] In some optional implementations, the preset generation model is constructed based on the sample speech features of the sample speech data, the fourth fusion deformation parameter sequence, and the ground truth of the sample coding sequence; wherein, the fourth fusion deformation parameter sequence is semantically independent of the sample speech data;

[0108] The true value of the sample encoding sequence is obtained by encoding the third fusion deformation parameter sequence through the preset encoder in the constructed vector quantization model.

[0109] In some optional implementations, the preset vector quantization model also includes a preset encoder and a codebook; the animation generation device may further include:

[0110] The first construction module is used to construct a pre-defined vector quantization model based on the following process:

[0111] The third fused deformation parameter sequence is encoded by a preset encoder to obtain the predicted value of the first encoded sequence;

[0112] From the codebook, identify feature vectors that are similar to each code in the predicted value of the first coding sequence, and determine the coding features based on each similar feature vector;

[0113] The coded features are decoded by a preset decoder to obtain the predicted value of the third fused deformation parameter sequence;

[0114] The reconstruction loss is determined based on the third fused deformation parameter sequence and the predicted value of the third fused deformation parameter sequence;

[0115] Based on the reconstruction loss, the parameters of the preset encoder, preset decoder, and codebook are adjusted.

[0116] In some alternative implementations, the first building block can also be used for:

[0117] A first coding loss is constructed based on the predicted value of the first coding sequence, and the parameters of the preset encoder are adjusted based on the first coding loss.

[0118] In some alternative implementations, the animation generation device may also include:

[0119] The second construction module is used to build a pre-defined generative model based on the following process:

[0120] By using a pre-defined generation model, a second coding sequence prediction value is generated based on the sample speech features and the fourth fusion deformation parameter sequence.

[0121] The second coding loss is determined based on the true value of the sample coding sequence and the predicted value of the second coding sequence;

[0122] The parameters of the preset generative model are adjusted based on the second encoding loss.

[0123] In some alternative implementations, the object model includes at least one of the following: a two-dimensional object model and a three-dimensional object model.

[0124] The animation generation apparatus provided in this disclosure can execute the animation generation method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of the execution method.

[0125] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.

[0126] The following is for reference. Figure 8 It illustrates an electronic device suitable for implementing embodiments of the present disclosure (e.g., Figure 8 The diagram below shows the structure of the terminal device or server 800. The terminal device in this embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and vehicle terminals (e.g., vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 8 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0127] like Figure 8 As shown, the electronic device 800 may include a processing unit (e.g., a central processing unit, a graphics processor, etc.) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803. The RAM 803 also stores various programs and data required for the operation of the electronic device 800. The processing unit 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0128] Typically, the following devices can be connected to I / O interface 805: input devices 806 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 807 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 808 including, for example, magnetic tapes, hard disks, etc.; and communication devices 809. Communication device 809 allows electronic device 800 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 8 An electronic device 800 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0129] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 809, or installed from a storage device 808, or installed from a ROM 802. When the computer program is executed by the processing device 801, it performs the functions defined in the animation generation method of embodiments of this disclosure.

[0130] The electronic device provided in this embodiment and the animation generation method provided in the above embodiments belong to the same disclosed concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.

[0131] This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the animation generation method provided in the above embodiments.

[0132] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory (FLASH), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. The transmitted data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0133] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0134] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0135] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to:

[0136] The process involves acquiring speech features and a first fusion deformation parameter sequence from the speech data; wherein the first fusion deformation parameter sequence is semantically independent of the speech data; generating an encoding sequence based on the speech features and the first fusion deformation parameter sequence using a preset generation model; decoding the encoding sequence into a second fusion deformation parameter sequence using a preset decoder; and driving an object model based on the second fusion deformation parameter sequence to generate a facial animation corresponding to the speech data.

[0137] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0138] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0139] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units and modules do not, in certain circumstances, constitute a limitation on the unit or module itself.

[0140] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Parts (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on.

[0141] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0142] According to one or more embodiments of this disclosure, an animation generation method is provided, the method comprising:

[0143] Acquire speech features and a first fusion deformation parameter sequence from the speech data; wherein the first fusion deformation parameter sequence is semantically independent of the speech data;

[0144] A coding sequence is generated based on the speech features and the first fusion deformation parameter sequence using a preset generation model.

[0145] The encoded sequence is decoded into a second fused deformation parameter sequence using a preset decoder;

[0146] The object model is driven according to the second fusion deformation parameter sequence to generate a facial animation corresponding to the speech data.

[0147] According to one or more embodiments of this disclosure, an animation generation method is provided, further comprising:

[0148] In some optional implementations, generating the encoding sequence based on the speech features and the first fusion deformation parameter sequence includes:

[0149] Features of the first fused deformation parameter sequence are extracted to obtain a style vector; wherein the style vector represents the facial expression of the object model;

[0150] A coding sequence is generated based on the speech features and the style vector.

[0151] According to one or more embodiments of this disclosure, an animation generation method is provided, further comprising:

[0152] In some alternative implementations, the preset decoder is contained in a preset vector quantization model, and the preset vector quantization model is constructed based on a third fusion deformation parameter sequence that is semantically related to the sample speech data.

[0153] According to one or more embodiments of this disclosure, an animation generation method is provided, further comprising:

[0154] In some optional implementations, the preset generation model is constructed based on the sample speech features of the sample speech data, the fourth fusion deformation parameter sequence, and the ground truth of the sample encoding sequence; wherein, the fourth fusion deformation parameter sequence is semantically independent of the sample speech data;

[0155] The true value of the sample encoding sequence is obtained by encoding the third fusion deformation parameter sequence through a preset encoder in the constructed vector quantization model.

[0156] According to one or more embodiments of this disclosure, an animation generation method is provided, further comprising:

[0157] In some optional implementations, the preset vector quantization model further includes a preset encoder and a codebook; the construction process of the preset vector quantization model includes:

[0158] The third fused deformation parameter sequence is encoded using the preset encoder to obtain the predicted value of the first encoded sequence;

[0159] From the codebook, feature vectors similar to each code in the predicted value of the first coding sequence are determined, and coding features are determined based on each similar feature vector;

[0160] The encoded features are decoded using the preset decoder to obtain the predicted value of the third fused deformation parameter sequence;

[0161] The reconstruction loss is determined based on the third fused deformation parameter sequence and the predicted value of the third fused deformation parameter sequence;

[0162] Based on the reconstruction loss, the parameters of the preset encoder, the preset decoder, and the codebook are adjusted.

[0163] According to one or more embodiments of this disclosure, an animation generation method is provided, further comprising:

[0164] In some alternative implementations, a first coding loss is constructed based on the predicted value of the first coding sequence, and the parameters of the preset encoder are adjusted based on the first coding loss.

[0165] According to one or more embodiments of this disclosure, an animation generation method is provided, further comprising:

[0166] In some optional implementations, the construction process of the preset generative model includes:

[0167] The second coding sequence prediction value is generated by the preset generation model based on the sample speech features and the fourth fusion deformation parameter sequence.

[0168] The second coding loss is determined based on the true value of the sample coding sequence and the predicted value of the second coding sequence;

[0169] The parameters of the preset generation model are adjusted based on the second encoding loss.

[0170] According to one or more embodiments of this disclosure, an animation generation method is provided, further comprising:

[0171] In some alternative implementations, the object model includes at least one of the following: a two-dimensional object model and a three-dimensional object model.

[0172] According to one or more embodiments of this disclosure, an animation generation apparatus is provided, the apparatus comprising:

[0173] The acquisition module is used to acquire the speech features and a first fusion deformation parameter sequence of the speech data; wherein, the first fusion deformation parameter sequence is semantically independent of the speech data;

[0174] The encoding generation module is used to generate an encoding sequence based on the speech features and the first fusion deformation parameter sequence using a preset generation model;

[0175] The decoding module is used to decode the encoded sequence into a second fused deformation parameter sequence using a preset decoder;

[0176] An animation generation module is used to drive the object model according to the second fusion deformation parameter sequence to generate a facial animation corresponding to the voice data.

[0177] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0178] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0179] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. An animation generation method, characterized in that, include: Acquire speech features and a first fusion deformation parameter sequence from the speech data; wherein the first fusion deformation parameter sequence is semantically independent of the speech data; A coding sequence is generated based on the speech features and the first fusion deformation parameter sequence using a preset generation model. The encoded sequence is decoded into a second fused deformation parameter sequence using a preset decoder; The object model is driven by the second fusion deformation parameter sequence to generate a facial animation corresponding to the speech data.

2. The method according to claim 1, characterized in that, The step of generating an encoding sequence based on the speech features and the first fused deformation parameter sequence includes: Features of the first fused deformation parameter sequence are extracted to obtain a style vector; wherein the style vector represents the facial expression of the object model; A coding sequence is generated based on the speech features and the style vector.

3. The method according to claim 1, characterized in that, The preset decoder is contained in the preset vector quantization model, and the preset vector quantization model is constructed based on a third fusion deformation parameter sequence that is semantically related to the sample speech data.

4. The method according to claim 3, characterized in that, The preset generation model is constructed based on the sample speech features, the fourth fusion deformation parameter sequence, and the ground truth of the sample encoding sequence of the sample speech data; wherein, the fourth fusion deformation parameter sequence is semantically independent of the sample speech data; The true value of the sample encoding sequence is obtained by encoding the third fusion deformation parameter sequence through a preset encoder in the constructed vector quantization model.

5. The method according to claim 3, characterized in that, The preset vector quantization model further includes a preset encoder and a codebook; the construction process of the preset vector quantization model includes: The third fused deformation parameter sequence is encoded using the preset encoder to obtain the predicted value of the first encoded sequence; From the codebook, feature vectors similar to each code in the predicted value of the first coding sequence are determined, and coding features are determined based on each similar feature vector; The encoded features are decoded using the preset decoder to obtain the predicted value of the third fused deformation parameter sequence; The reconstruction loss is determined based on the third fused deformation parameter sequence and the predicted value of the third fused deformation parameter sequence; Based on the reconstruction loss, the parameters of the preset encoder, the preset decoder, and the codebook are adjusted.

6. The method according to claim 5, characterized in that, Also includes: A first coding loss is constructed based on the predicted value of the first coding sequence, and the parameters of the preset encoder are adjusted based on the first coding loss.

7. The method according to claim 4, characterized in that, The construction process of the preset generative model includes: The second coding sequence prediction value is generated by the preset generation model based on the sample speech features and the fourth fusion deformation parameter sequence. The second coding loss is determined based on the true value of the sample coding sequence and the predicted value of the second coding sequence; The parameters of the preset generation model are adjusted based on the second encoding loss.

8. The method according to any one of claims 1-7, characterized in that, The object model includes at least one of the following: a two-dimensional object model and a three-dimensional object model.

9. An animation generation device, characterized in that, include: The acquisition module is used to acquire the speech features and a first fusion deformation parameter sequence of the speech data; wherein, the first fusion deformation parameter sequence is semantically independent of the speech data; The encoding generation module is used to generate an encoding sequence based on the speech features and the first fusion deformation parameter sequence using a preset generation model; The decoding module is used to decode the encoded sequence into a second fused deformation parameter sequence using a preset decoder; An animation generation module is used to drive an object model based on the second fusion deformation parameter sequence to generate a facial animation corresponding to the voice data.

10. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the animation generation method as described in any one of claims 1-8.

11. A storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the animation generation method as described in any one of claims 1-8.