A speech-driven three-dimensional face animation generation method, system and terminal based on double-time encoding and geometric perception mean flow correction

CN122820933APending Publication Date: 2026-09-25PENG CHENG LAB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611092427.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2026-04-22
Filing Date
2026-07-22
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0011]本发明的主要目的在于提供一种基于双时刻编码与几何感知均值流校正的语音驱动三维面部动画生成方法、系统及终端,旨在解决现有技术中现有语音驱动三维面部动画方法中存在的确定性回归结果过平滑、生成式方法推理效率较低以及一步生成过程中轨迹易偏离显式几何流形的问题

Benefits of technology

[0022]本发明中,输入参考视频和驱动语音,对参考视频和驱动语音进行数据预处理,提取3DMM参数、静态形状系数、风格编码和音频特征;构建仿真时间与参考时间,形成双时刻时间嵌入,并结合静态形状系数、风格编码、历史动作与音频特征预测速度场;利用平均速度场恢复目标动作端点,并通过FLAME解码器映射到显式三维顶点空间;通过雅可比向量积构造校正目标流,并结合几何损失和区域语义约束,对潜空间中的均值流轨迹进行修正,得到经过轨迹校正的平均速度场及对应的区域语义约束结果;采用分阶段递进式训练策略,在训练阶段先学习瞬时速度再学习平均速度,在推理阶段,将参考时间固定为 1,利用网络预测得到的平均速度场恢复目标动作端点,并输出语音驱动的三维面部动画结果。本发明通过双时刻去噪Transformer对速度场进行稳定建模,并结合几何感知轨迹学习对一步生成路径进行显式流形约束和区域感知校正,从而在保证较高动作表达性和唇形同步精度的同时,实现低时延、高效率且几何合理的语音驱动三维面部动画生成。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820933A_ABST
    Figure CN122820933A_ABST
Patent Text Reader

Abstract

The application discloses a speech-driven three-dimensional face animation generation method, system and terminal based on double-time coding and geometric perception mean flow correction, and the method comprises the following steps: data preprocessing is performed on a reference video and driving speech to extract a plurality of parameters; simulation time and reference time are constructed to form double-time time embedding, and static shape coefficients, style coding, historical action and audio feature prediction speed field are combined; a target action endpoint is recovered by using a mean speed field and is mapped to an explicit three-dimensional vertex space; a correction target flow is constructed, and a latent space trajectory is modified in combination with geometric loss and regional semantic constraint; in the training stage, instantaneous speed is learned first and then mean speed is learned, in the inference stage, the reference time is fixed as 1, the target action endpoint is recovered by using the mean speed field predicted by the network, and a speech-driven three-dimensional face animation result is output. The application realizes low-latency, high-efficiency and geometrically reasonable speech-driven three-dimensional face animation generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of voice-driven 3D facial animation generation technology, and in particular to a voice-driven 3D facial animation generation method, system and terminal based on dual-moment encoding and geometrically perceptual mean flow correction. Background Technology

[0002] Voice-driven 3D facial animation generation technology, as a fundamental core technology in fields such as virtual reality or augmented reality, digital human interaction, game animation, and film and television production, aims to automatically generate 3D facial motion sequences with synchronized lip movements, natural expressions, and a certain degree of stylistic control based on arbitrary input audio. With the development of digital content production and real-time interactive applications, users have placed higher demands on voice-driven 3D facial animation in terms of lip-sync accuracy, richness of action expression, and real-time inference efficiency. While traditional deterministic regression methods can output 3D facial movements relatively quickly, they often struggle to fully express the diverse facial dynamics under complex speech conditions; while generative methods can improve action expressiveness, they are usually accompanied by high inference latency. Therefore, how to achieve highly expressive, low-latency, and explicitly geometrically constrained voice-driven 3D facial animation generation while ensuring action accuracy and natural timing has become an important research direction in the field of computer vision and graphics.

[0003] Previous voice-driven 3D facial animation technologies mainly fall into four categories, each with its own characteristics: (1) Methods based on rules or procedural mapping drive facial movements by predefined phoneme-lip shape correspondences, which have a certain degree of control, but limited naturalness and generalization ability.

[0004] (2) Based on deterministic regression, the speech signal is directly mapped to vertex displacement or 3DMM (3DMorphable Model) parameterized face model (such as FLAME, Faces Learned with an Articulated Model and Expressions) coefficients, which can generate facial movements quickly, but usually tends to produce oversmoothing.

[0005] (3) Generation methods based on autoregressive or discrete action codebooks enhance action diversity by progressively predicting action sequences.

[0006] (4) Methods based on diffusion models or probabilistic generation models can generate three-dimensional facial movements by gradually denoising from noise, which can better model complex movement distributions and style changes, but usually require multiple iterative sampling steps and inference time is high.

[0007] The main drawbacks of existing technologies are as follows: (1) Facial animation methods based on deterministic regression models are prone to oversmoothing. These methods typically treat the mapping from speech to 3D facial movements as a single-value regression task. However, there is often a one-to-many correspondence between actual speech and facial movements, meaning that the same speech segment may correspond to multiple reasonable facial expressions and movement styles. Therefore, after training with deterministic optimization, the model tends to learn the average result of the conditional distribution, resulting in facial animations that lack subtle facial expressions and style variations, manifested as insufficient lip opening and closing, stiff local movements, and weak overall expressiveness.

[0008] (2) Although probabilistic generation methods based on diffusion models and autoregressive models can improve action diversity and expressive ability, they have large inference delays. These methods usually require multi-step iterative sampling or frame-by-frame sequence prediction, which results in large computational load and high latency during inference, making it difficult to meet the efficiency requirements of real-time driven and interactive applications.

[0009] (3) While existing one-step generation or accelerated generation methods improve inference efficiency, they are prone to problems such as trajectory deviation from explicit geometric manifold. Since voice-driven 3D facial animation essentially requires the generation of 3D geometric motion that satisfies a clear topological structure and physical rationality, if there is a lack of explicit geometric constraints on the generated trajectory, problems such as inaccurate local lip movements, distorted facial movements, unnatural geometric surfaces, and unstable motion in key areas are likely to occur, affecting the realism and synchronization accuracy of the final 3D animation.

[0010] Therefore, existing technologies still need to be improved and developed. Summary of the Invention

[0011] The main objective of this invention is to provide a speech-driven 3D facial animation generation method, system, and terminal based on dual-time encoding and geometrically perceptual mean flow correction. This invention aims to solve the problems of overly smooth deterministic regression results, low inference efficiency of generative methods, and easy deviation of trajectories from explicit geometric manifolds in existing speech-driven 3D facial animation methods.

[0012] To achieve the above objectives, this invention provides a speech-driven 3D facial animation generation method based on dual-time-mapping encoding and geometrically perceptual mean flow correction. The speech-driven 3D facial animation generation method based on dual-time-mapping encoding and geometrically perceptual mean flow correction includes the following steps: Input reference video and driving speech, perform data preprocessing on the reference video and driving speech, and extract 3DMM parameters, static shape coefficients, style codes and audio features; Simulation time and reference time are constructed to form a dual-time embedding, and the velocity field is predicted by combining static shape coefficients, style coding, historical actions and audio features. The target motion endpoints are recovered using the average velocity field and mapped to an explicit 3D vertex space using the FLAME decoder; The target flow is corrected by constructing the Jacobi vector product and combining geometric loss and regional semantic constraints to correct the mean flow trajectory in the latent space, thus obtaining the mean velocity field after trajectory correction and the corresponding regional semantic constraint results. A phased progressive training strategy is adopted. In the training phase, instantaneous velocity is learned first and then average velocity is learned. In the inference phase, the reference time is fixed at 1, and the target action endpoint is recovered by using the average velocity field predicted by the network, and the voice-driven 3D facial animation result is output.

[0013] Optionally, the speech-driven 3D facial animation generation method based on dual-moment coding and geometrically perceptual mean stream correction, wherein the input reference video and driving speech are preprocessed to extract 3DMM parameters, static shape coefficients, style codes, and audio features, specifically including: For the input reference video, parametric face reconstruction and tracking methods are used to extract the 3DMM parameters corresponding to each frame. The FLAME model is used to obtain the explicit parameter representation of the facial action sequence, and static shape coefficients are extracted from the reference video to characterize the speaker's static identity geometry. Simultaneously, a temporal style encoder is used to encode the reference action sequence corresponding to the reference video to obtain the style code. S ; For the input driving speech, a pre-trained HuBERT audio encoder is used to extract audio features corresponding to the time series. .

[0014] Optionally, the speech-driven 3D facial animation generation method based on dual-moment coding and geometrically perceptual mean flow correction, wherein the input reference video and driving speech are preprocessed to extract 3DMM parameters, static shape coefficients, style codes, and audio features, and then further includes: Record the actual facial movement sequence as the target movement. And sample noise actions from a standard Gaussian distribution. Then, based on the simulation time t Constructing noisy action states ; Extract a certain length of historical action sequence from a known action sequence. It is used to provide temporal context constraints for the generation of actions at the current moment.

[0015] Optionally, in the speech-driven 3D facial animation generation method based on dual-time encoding and geometrically perceptual mean flow correction, the velocity field includes an instantaneous velocity field or an average velocity field.

[0016] Optionally, the speech-driven 3D facial animation generation method based on dual-moment encoding and geometrically perceptual mean flow correction, wherein constructing simulation time and reference time to form a dual-moment temporal embedding, and combining static shape coefficients, style encoding, historical actions, and audio features to predict the velocity field, specifically includes: From the interval Internal sampling simulation time t Used to represent the current noisy action state The corresponding generation progress is recorded, and a random time interval is sampled independently. u ~ As a construction reference time r Candidate time variables; Design an interleaving switching mechanism for the reference time. r The assignment is performed in stages. In the first training stage, the reference time is set to... This causes the network to degenerate into learning instantaneous velocity fields. In the second training phase, the reference time is set to random time. This relates the average speed of network learning from any reference time to the current time, and the reference time. r Represented as: ; Simulation time t and reference time r Sine wave position encoding is performed separately, and then each is input into the corresponding multilayer perceptron. and The mapping process yields two temporal embeddings, which are then added together to form a composite temporal embedding. : ; in, This represents the sinusoidal position coding function. and These represent the multilayer perceptrons that process the simulation time and the reference time, respectively. static shape factor and style coding S The parts are combined to form an identity code. Composite time is embedded by adding elements one by one. Injected into the identity encoding to construct a time-aware identity prefix. ; Time-aware identity prefix Historical action sequence and noisy action state The sequences are concatenated into a composite input sequence and then input into a two-time-phase denoising Transformer for modeling. The dual-time denoising Transformer utilizes bidirectional self-attention to establish global dependencies between input segments, and then uses masked cross-attention to extract audio features. An action generation process is introduced, and an alignment mask is used to limit the local audio window corresponding to each action position. After processing by the Transformer decoder module and the linear projection layer, the velocity field prediction result corresponding to the current state is output: when At that time, the instantaneous velocity field is output. ; when At that time, the average velocity field is output. .

[0017] Optionally, the speech-driven 3D facial animation generation method based on dual-time-time encoding and geometrically perceptual mean flow correction, wherein the step of recovering the target motion endpoint using the average velocity field and mapping it to the explicit 3D vertex space through the FLAME decoder specifically includes: For any simulation time t Noisy action state The target motion is implicitly recovered using the average velocity field. During the inference phase, with the reference time fixed at 1, the displacement estimate from the current state to the target motion endpoint is directly obtained. The recovered target motion endpoint is expressed as: ; in, This represents the target action endpoint recovered by the network prediction. This represents the average velocity field output when the reference time is fixed at 1. This represents the remaining time span from the current simulation moment to the target endpoint; The recovered motion parameters are decoded into corresponding 3D facial mesh vertices using the FLAME decoder to obtain the predicted mesh, while the real motion parameters are decoded into the real mesh, and the two are compared in the explicit vertex space. Constructing geometric loss : ; in, This represents the vertex constraint loss, used to constrain the difference between the predicted mesh and the true mesh in the explicit vertex space. This represents the velocity constraint loss, used to suppress abrupt changes in vertex motion between adjacent time steps, thus reducing temporal jitter in facial movements. This represents the coefficient constraint loss, used to maintain the continuity of motion parameters in the boundary region and structural connection region, avoiding unnatural local structures or geometric distortions. , and These are the weight coefficients for vertex constraint loss, velocity constraint loss, and coefficient constraint loss, respectively.

[0018] Optionally, the speech-driven 3D facial animation generation method based on dual-time encoding and geometrically perceptual mean flow correction, wherein the step of constructing a correction target flow through the Jacobian vector product and correcting the mean flow trajectory in the latent space by combining geometric loss and regional semantic constraints specifically includes: Constructing a corrected target flow for monitoring the mean velocity field : ; in, This represents the target action corresponding to the actual action sequence. This represents noise action sampled from a standard Gaussian distribution. This represents the average velocity field predicted by the two-time-time denoised Transformer, and the derivative term is calculated using the Jacobian vector product; Construct the mean flow loss and apply it to the predicted mean velocity field. and correct target flow Constrain the differences between them; The mean flow error is constructed based on the difference between the predicted mean velocity field and the corrected target flow. And the mean stream error is obtained through the FLAME decoder in a local approximation sense. Mapped to explicit mesh vertex deformation : ; in, Indicates FLAME decoder, Indicates the current baseline action state. These are the small perturbation coefficients used for local approximation; Constructing semantically aware manifold constraint loss : ; in, This represents the constraint terms that apply to the dynamic region of the lips. This represents the constraint terms that act on the structural context region. and They are respectively and The weighting coefficients.

[0019] Optionally, the speech-driven 3D facial animation generation method based on dual-moment encoding and geometrically perceptual mean flow correction, wherein the method employs a phased progressive training strategy, first learning instantaneous velocity and then average velocity during the training phase, and fixing the reference time to 1 during the inference phase, using the average velocity field predicted by the network to recover the target action endpoint, and outputting the speech-driven 3D facial animation result, specifically includes: A phased, progressive training strategy is adopted. The first training phase focuses on instantaneous speed learning, with the reference time set to [missing information]. r = t This enables the dual-time denoising Transformer to learn the instantaneous velocity field under the standard flow matching framework, establish a stable basic cross-modal correspondence between speech signals and three-dimensional facial movements, and form an initial motion manifold under explicit geometric constraints. After completing the first phase of training, the second phase begins. This phase continues training based on the initial weights, using an alternating switching mechanism to transition between instantaneous velocity learning and average velocity learning: a portion of the training samples are... r = t To maintain the model's ability to stably model local instantaneous dynamics, another part of the training samples uses a random reference time. r = u , used to learn the average velocity field; During the inference phase, the reference time is fixed at 1, and the target action endpoint is recovered directly using the average velocity field predicted by the network, and the corresponding 3D facial animation result is output.

[0020] Furthermore, to achieve the above objectives, the present invention also provides a speech-driven 3D facial animation generation system based on dual-time-matter encoding and geometrically perceptual mean flow correction, wherein the speech-driven 3D facial animation generation system based on dual-time-matter encoding and geometrically perceptual mean flow correction includes: The data preprocessing module is used to input reference video and driving speech, perform data preprocessing on the reference video and driving speech, and extract 3DMM parameters, static shape coefficients, style codes and audio features; The dual-moment encoding construction module is used to construct the simulation time and reference time, forming a dual-moment time embedding, and combining static shape coefficients, style encoding, historical action and audio features to predict the velocity field; The implicit recovery-based manifold projection module is used to recover the target action endpoints using the average velocity field and map them to the explicit 3D vertex space through the FLAME decoder. The geometrically perceptual mean flow trajectory correction module is used to construct the target flow through the Jacobian vector product and, combined with geometric loss and regional semantic constraints, correct the mean flow trajectory in the latent space to obtain the mean velocity field after trajectory correction and the corresponding regional semantic constraint results. The progressive training and single-step inference module is used to adopt a phased progressive training strategy. In the training phase, instantaneous velocity is learned first and then average velocity is learned. In the inference phase, the reference time is fixed to 1, the average velocity field predicted by the network is used to recover the target action endpoint, and the voice-driven 3D facial animation result is output.

[0021] Furthermore, to achieve the above objectives, the present invention also provides a terminal, wherein the terminal includes: a memory, a processor, and a voice-driven three-dimensional facial animation generation program based on dual-moment encoding and geometrically perceptual mean flow correction stored in the memory and executable on the processor. When the voice-driven three-dimensional facial animation generation program based on dual-moment encoding and geometrically perceptual mean flow correction is executed by the processor, it implements the steps of the voice-driven three-dimensional facial animation generation method based on dual-moment encoding and geometrically perceptual mean flow correction as described above.

[0022] In this invention, a reference video and driving speech are input, and the data of the reference video and driving speech are preprocessed to extract 3DMM parameters, static shape coefficients, style codes, and audio features. A simulation time and a reference time are constructed to form a dual-moment time embedding, and the velocity field is predicted by combining static shape coefficients, style codes, historical actions, and audio features. The target action endpoint is recovered using the average velocity field and mapped to the explicit 3D vertex space through the FLAME decoder. The corrected target flow is constructed by the Jacobian vector product, and the mean flow trajectory in the latent space is corrected by combining geometric loss and regional semantic constraints to obtain the trajectory-corrected average velocity field and the corresponding regional semantic constraint results. A phased progressive training strategy is adopted. In the training phase, the instantaneous velocity is learned first and then the average velocity is learned. In the inference phase, the reference time is fixed to 1, the target action endpoint is recovered using the average velocity field predicted by the network, and the voice-driven 3D facial animation result is output. This invention uses a dual-time denoising Transformer to stably model the velocity field and combines geometrically perceptual trajectory learning to explicitly constrain the one-step generated path and perform region-aware correction. This achieves low-latency, high-efficiency and geometrically reasonable voice-driven 3D facial animation generation while ensuring high action expressiveness and lip-sync accuracy. Attached Figure Description

[0023] Figure 1 This is a schematic diagram illustrating the process of outputting the three-dimensional facial animation result after dual-time denoising Transformer and geometric perception trajectory learning in a preferred embodiment of the speech-driven three-dimensional facial animation generation method based on dual-time encoding and geometric perception mean flow correction of the present invention. Figure 2 This is a flowchart of the entire implementation process in a preferred embodiment of the speech-driven 3D facial animation generation method based on dual-time encoding and geometrically perceptual mean flow correction of the present invention. Figure 3 This is a flowchart of a preferred embodiment of the speech-driven 3D facial animation generation method based on dual-time coding and geometrically perceptual mean flow correction of the present invention; Figure 4 This is a structural diagram of a preferred embodiment of the speech-driven 3D facial animation generation system based on dual-time encoding and geometrically perceptual mean flow correction of the present invention; Figure 5 This is a structural diagram of a preferred embodiment of the terminal of the present invention. Detailed Implementation

[0024] To make the objectives, technical solutions, and advantages of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0025] To address the problems in existing speech-driven 3D facial animation methods, such as overly smooth deterministic regression results, low inference efficiency of generative methods, and easy deviation of trajectories from explicit geometric manifolds during one-step generation, a speech-driven 3D facial animation generation method based on dual-moment encoding and geometrically perceptual mean flow correction is proposed. This method uses a dual-moment denoising Transformer to stably model the velocity field and combines geometrically perceptual trajectory learning to apply explicit manifold constraints and region-aware correction to the one-step generation path. Thus, while ensuring high motion expressiveness and lip-sync accuracy, it achieves low-latency, high-efficiency, and geometrically reasonable speech-driven 3D facial animation generation.

[0026] This invention further introduces the existing mean flow generation mechanism and FLAME-based explicit 3D face geometric representation into the speech-driven 3D facial animation generation process. Unlike previous methods that use a single deterministic regression network or rely on multi-step diffusion / autoregressive sampling to model the speech-to-facial motion mapping, this invention proposes a speech-driven 3D facial animation generation method based on dual-moment encoding and geometrically perceptive mean flow correction. This method uses a dual-moment denoising Transformer to stably model the velocity field and combines geometrically perceptive trajectory learning to apply explicit manifold constraints and region-aware correction to the one-step generation path. As a result, it outperforms existing methods in terms of lip-sync, motion expressiveness, geometric rationality, and real-time inference efficiency of facial animation.

[0027] The purpose of this invention is to generate a 3D facial animation result that is synchronized with the speech content, exhibits natural movements, and possesses high expressiveness, given a reference video and arbitrary input speech, while achieving low-latency one-step inference while ensuring lip-sync accuracy and geometric plausibility. This method uses the parametric 3D face model FLAME as the explicit geometric representation basis, combining a mean-flow one-step generation mechanism, a dual-time denoising Transformer (DTDT), and geometric-aware trajectory learning (GATL) to achieve efficient and stable generation of 3D facial animation results from input speech. Figure 1 As shown, this invention consists of two parts: "Module 1: Dual-Time Denoising Transformer (DTDT)" and "Module 2: Geometric Aware Trajectory Learning (GATL)". Module 1 is used to predict the instantaneous or average velocity field under dual-time encoding conditions, while Module 2 is used to perform explicit manifold projection and region semantic constraint correction on the one-step generated trajectory. Figure 2 As shown, the method flow of the present invention includes steps such as data preprocessing, dual-time-time encoding construction, manifold projection based on implicit recovery, geometrically perceptual mean flow trajectory correction, progressive training and single-step inference, and finally outputting three-dimensional facial animation results.

[0028] like Figure 2 As shown, the inputs are: reference video and driving speech; Step 1: Data preprocessing, extracting 3DMM parameters, static shape coefficients, style codes, and audio features; Step 2: Dual-time encoding construction, constructing simulation time... t Reference time r Step 1: Form a dual-time embedding and combine static shape coefficients, style encoding, historical actions, and audio features to predict the velocity field; Step 2: Based on implicit recovery, manifold projection is used to recover the target action endpoints using the average velocity field and map them to the explicit 3D vertex space through the FLAME decoder; Step 3: Geometrically perceptual mean flow trajectory correction is performed by constructing a corrected target flow through the Jacobian-Vector Product (JVP) and combining geometric loss and regional semantic constraints to correct the mean flow trajectory in the latent space; Step 4: Progressive training and single-step inference are performed by first learning the instantaneous velocity and then the average velocity during the training phase, and fixing the reference time to 1 during the inference phase to achieve one-step generation of the target action sequence; Output: Voice-driven 3D facial animation results.

[0029] The preferred embodiment of the speech-driven 3D facial animation generation method based on dual-time coding and geometrically perceptual mean flow correction described in this invention, such as... Figure 3As shown, the speech-driven 3D facial animation generation method based on dual-time-time encoding and geometrically perceptual mean flow correction includes the following steps: Step S10: Input reference video and driving speech, perform data preprocessing on reference video and driving speech, and extract 3DMM parameters, static shape coefficients, style codes and audio features.

[0030] Specifically, this invention trains on a large-scale public dataset containing numerous videos of individuals with different identities and corresponding 3D facial parameters to learn a speech-to-3D facial motion mapping relationship with cross-identity generalization ability. Before training, the reference videos and corresponding audio data are preprocessed.

[0031] For the input reference video, existing parametric face reconstruction and tracking methods are used to extract the 3DMM parameters (3D deformable face model parameters) corresponding to each frame. For example, the FLAME model can be used to obtain the explicit parameter representation of the facial action sequence, and further static shape coefficients used to characterize the static identity geometry of the speaker are extracted from the reference video. Simultaneously, to characterize the speech style and facial expressions of the reference person changing over time, a temporal style encoder is used to encode the reference action sequence corresponding to the reference video, resulting in style codes. S For the input driving speech, a pre-trained HuBERT (Hidden-Unit BERT, a pre-trained speech representation model) audio encoder is used to extract audio features corresponding to the time series. .

[0032] Furthermore, during the training phase, to construct the generation path based on flow matching, it is also necessary to establish an intermediate state representation between the target action data and the standard Gaussian distribution. Specifically, the real facial action sequence is denoted as the target action. And sample noise actions from a standard Gaussian distribution. Then, based on the simulation time t Constructing noisy action states This serves as one of the inputs to the subsequent dual-timetime denoising Transformer. Simultaneously, a certain length of historical action sequence is extracted from the known action sequence. This provides temporal context constraints for generating actions at the current moment. After the above data preprocessing, the 3DMM parameters and static shape coefficients used in subsequent modules can be obtained. Style coding S Audio features Historical action sequence and noisy action state This provides a unified data foundation for dual-time-matter encoding construction and geometrically perceptual trajectory learning.

[0033] Step S20: Construct simulation time and reference time to form a dual-time embedding, and combine static shape coefficients, style coding, historical actions and audio features to predict the velocity field.

[0034] The purpose of this step is to establish a unified temporal condition that can simultaneously represent the "current generation time" and the "target reference time" for subsequent voice-driven 3D facial motion generation, and on this basis, to complete velocity field prediction, enabling the network to learn both local instantaneous motion patterns and, further, a global average velocity field suitable for single-step inference. For example... Figure 1 As shown in Module 1, this invention introduces simulation time, building upon the traditional flow matching method which uses only a single time variable. t and reference time r Two time anchor points are used, and temporal condition information is constructed through a dual-moment encoding method. This information is then combined with shape encoding, style encoding, historical action, and audio features, and input into a dual-moment denoising Transformer to output an instantaneous velocity field. or average velocity field ;like Figure 2 As shown, this step is located after data preprocessing and provides a unified time condition representation and velocity field input for subsequent manifold projection and trajectory correction.

[0035] Specifically, during the training phase, the first step is to start from the interval Internal sampling simulation time t Used to represent the current noisy action state The corresponding generation progress; at the same time, a random time is sampled independently. u ~ As a construction reference time r Candidate time variables; to enable the model to gradually transition from instantaneous velocity learning to average velocity learning in an increasingly difficult manner, this invention designs an interleaved switching mechanism for the reference time. r The assignment is performed in stages. In the first training stage, the reference time is set to... This causes the network to degenerate into learning instantaneous velocity fields; in the second training phase, the reference time is set to random time. This allows the network to learn the average velocity relationship from any reference time to the current time. Correspondingly, the reference time *r* can be expressed as: ; In this process, Stage I establishes a stable, basic cross-modal correspondence between speech and facial movements, enabling the network to have good local motion modeling capabilities. Stage II builds upon this foundation by further learning the average motion patterns across time spans, allowing the model to directly generate data in one step during inference. Through this staggered switching mechanism, this invention avoids the training instability problem caused by directly learning the average velocity field from scratch, achieving a smooth transition from local dynamic modeling to global average velocity modeling.

[0036] After sampling the time variable, in order to transform the scalar time information into a high-dimensional feature representation suitable for neural network processing, this invention modifies the simulation time... t and reference time r Sine wave position encoding is performed separately, and then each is input into the corresponding multilayer perceptron. and The mapping process yields two temporal embeddings, which are then added together to form a composite temporal embedding. : ; in, This represents the sinusoidal position coding function. and These represent the multilayer perceptrons that process the simulation time and the reference time, respectively.

[0037] Using the above method, two time variables can be mapped to a unified feature space, preserving their temporal semantics in terms of numerical magnitude and relative interval, thereby enabling the network to explicitly perceive "which generation moment it is currently in" and "which reference moment it needs to model the velocity towards".

[0038] Furthermore, to enable the temporal condition to work synergistically with the speaker's identity geometry and style information, this invention embeds the composite time. This is incorporated into the identity criteria extracted from the reference video. Specifically, the static shape coefficients are... and style coding S The parts are combined to form an identity code. Then, the composite time is embedded by adding elements one by one. Injected into the identity encoding to construct a time-aware identity prefix. Subsequently, the time-aware identity prefix was added. Historical action sequence and noisy action state The sequences are concatenated into a composite input sequence and then fed into a two-time denoising Transformer for modeling. The two-time denoising Transformer utilizes bidirectional self-attention to establish global dependencies between each input segment, and then uses masked cross-attention to extract audio features. An action generation process is introduced, and an alignment mask is used to limit the local audio window corresponding to each action position, thereby enhancing the local correspondence between speech content and mouth movements. After processing by the Transformer decoder module and the linear projection layer, the predicted velocity field for the current state is output. when At that time, the instantaneous velocity field is output. ; when At that time, the average velocity field is output. .

[0039] Through the aforementioned dual-moment encoding construction steps, this invention completes the mapping from scalar time variables to high-dimensional temporal conditional representations and establishes a joint representation mechanism among simulation time, reference time, identity geometric information, and style dynamic information. After this step, a reference time can be obtained for subsequent dual-moment denoising Transformer operations. r Simulation time t Composite temporal embedding Time-aware identity prefix and predicted velocity field or This provides the input basis for the implicit recovery-based manifold projection in the next step.

[0040] Step S30: Recover the target motion endpoint using the average velocity field and map it to the explicit three-dimensional vertex space using the FLAME decoder.

[0041] The purpose of this step is to, based on the predicted results of the average velocity field, no longer rely on traditional multi-step numerical integration to gradually recover the target motion, but instead use the average velocity field to directly and implicitly recover the endpoints of the target motion and map them into the explicit three-dimensional vertex space. This imposes constraints on the generated results at the physical geometry level, ensuring the geometric rationality and local lip-sync accuracy of the one-step generated results. For example... Figure 1 As shown in the upper part of Module 2, this invention denoises the Transformer output average velocity field at two time points. Then, construct a recovery operator to restore the current noisy action state. Directly restore to the target action endpoint Then, it is mapped to an explicit 3D mesh space via the FLAME decoder; such as Figure 2 As shown, this step occurs after the two-time-time encoding construction and is used to provide an explicit geometric reference for subsequent mean flow trajectory correction.

[0042] Specifically, for any simulation time t Noisy action state This invention utilizes an average velocity field to implicitly recover the target motion. Since the average velocity field describes the average evolution direction from the reference time to the current time, when the reference time is fixed to 1 during the inference phase, the displacement estimate from the current state to the target motion endpoint can be directly obtained. The recovered target motion endpoint is expressed as: ; in, This represents the target action endpoint recovered by the network prediction. This represents the average velocity field output when the reference time is fixed at 1. This represents the remaining time span from the current simulation moment to the target endpoint. Through the above recovery operator, the present invention can directly recover the target action result from the current noisy action state without relying on multi-step iterative integration, thereby providing a clear action recovery path for one-step inference.

[0043] At the target action endpoint that has been restored Subsequently, this invention further maps the results to an explicit 3D mesh vertex space to provide geometric supervision. Specifically, the FLAME decoder is used to decode the recovered motion parameters into corresponding 3D facial mesh vertices to obtain the predicted mesh; simultaneously, the real motion parameters are decoded into the real mesh, and the two are compared in the explicit vertex space.

[0044] Compared to supervision only in the latent space or parameter space, this approach directly reflects the errors in the generated results at the 3D geometric level, making it more suitable for constraining lip opening and closing, local motion continuity, and overall facial structure stability. To ensure that the recovered 3D facial animation meets requirements in terms of local lip shape, temporal continuity, and structural boundaries, this invention constructs a geometric loss function in this step. : ; in, This represents the vertex constraint loss, used to constrain the difference between the predicted mesh and the true mesh in the explicit vertex space. This represents the velocity constraint loss, used to suppress abrupt changes in vertex motion between adjacent time steps, thus reducing temporal jitter in facial movements. This represents the coefficient constraint loss, used to maintain the continuity of motion parameters in the boundary region and structural connection region, avoiding unnatural local structures or geometric distortions. , and These are the weight coefficients for vertex constraint loss, velocity constraint loss, and coefficient constraint loss, respectively.

[0045] By introducing the aforementioned constraints into the explicit 3D geometric space, this step not only ensures that the restored result maintains global shape consistency with the actual facial movement, but also more precisely constrains the local geometric changes of the lips, perioral region, and other key areas. As a result, the average velocity field prediction is no longer merely an abstract displacement direction in latent space, but corresponds to 3D mesh changes with a clear topological structure and physical meaning, thus providing a reliable geometric basis for subsequent trajectory correction.

[0046] Through the aforementioned manifold projection step based on implicit restoration, this invention completes the direct mapping from the average velocity field to the explicit 3D facial mesh and establishes an explicit geometric supervision relationship between the restored motion endpoints and the real geometric target. After this step, the target motion endpoints can be obtained. Predicted mesh and corresponding geometric loss This provides an explicit manifold basis and geometric constraint information for the geometrically sensed mean flow trajectory correction in the next step.

[0047] Step S40: Construct the corrected target flow through the Jacobian vector product, and combine geometric loss and regional semantic constraints to correct the mean flow trajectory in the latent space, thereby obtaining the mean velocity field after trajectory correction and the corresponding regional semantic constraint results.

[0048] The purpose of this step is to further correct the latent space trajectory in the one-step generation process, based on the aforementioned implicit recovery-based manifold projection, to reduce problems such as local lip-sync distortion, unnatural facial movements, and unstable structural region motion caused by trajectory curvature and manifold deviation, thereby simultaneously improving the lip-sync accuracy, expressiveness, and geometric rationality of speech-driven 3D facial animation. Figure 1 As shown in the lower half of Module 2, this invention, based on obtaining the average velocity field and explicit geometric constraints, corrects the mean flow trajectory in the latent space by constructing a corrected target flow and introducing regional semantic constraints; as shown... Figure 2 As shown, this step follows the implicit recovery-based manifold projection and is used to further optimize the potential generation path corresponding to the mean velocity field.

[0049] Specifically, to eliminate latent space trajectory curvature errors that may occur during one-step generation, this invention first constructs a corrected target flow for supervising the average velocity field. Based on the actual displacement direction, this target flow further considers the derivative of the average velocity field changing with time, thus more accurately characterizing the correction direction required for single-step generation. Its expression can be written as: ; in, This represents the target action corresponding to the actual action sequence. This represents noise action sampled from a standard Gaussian distribution. This represents the average velocity field predicted by the denoised Transformer at two time points. This represents its total derivative with respect to time.

[0050] In implementation, the derivative term is calculated using the Jacobian-Vector Product (JVP), thus avoiding the computational overhead of explicitly calculating the high-dimensional Jacobian matrix. This is achieved by introducing a corrected target flow. This invention can explicitly correct the predicted direction of the average velocity field, making it closer to the straight path corresponding to the actual motion distribution. After obtaining the corrected target flow, this invention further constructs a mean flow loss to improve the predicted average velocity field. and correct target flow The differences between them are constrained. To improve training stability and reduce the impact of outliers or local noise, this invention employs a weighted adaptive mean flow loss. By applying weights to the error term to suppress outliers, the average velocity field learning process becomes more stable. In this way, the model not only learns the average evolution direction from noisy states to the target action, but also gradually corrects the trajectory curvature in the latent space during training, making the generated path in one step closer to the ideal straight trajectory.

[0051] Furthermore, correcting the mean velocity field solely in the latent space is insufficient to guarantee that the generated result exhibits reasonable motion patterns across different facial regions. Therefore, this invention introduces a semantically aware manifold (GVC) constraint in this step, projecting the mean flow error in the latent space onto the explicit 3D facial surface manifold, thus providing explicit geometric interpretability for the trajectory correction process. Specifically, the mean flow error is first constructed based on the difference between the predicted mean velocity field and the target flow being corrected. The error is then mapped to explicit mesh vertex deformation using the FLAME decoder in a local approximation sense. This establishes a correspondence between latent space updates and explicit 3D surface changes. The process can be represented as: ; in, Indicates FLAME decoder, Indicates the current baseline action state. These are small perturbation coefficients used for local approximation; through the above mapping, the present invention can transform the velocity error in the original latent space into an observable explicit mesh deformation, thereby analyzing and constraining the trajectory correction effect on a specific facial region.

[0052] Considering that voice-driven 3D facial animation exhibits different motion attributes across different facial regions, this invention further employs a region-based semantic constraint strategy, applying constraints of varying intensities to key pronunciation regions and structural context regions. Specifically, for dynamic regions directly related to pronunciation, such as the lips and perioral area, stronger region constraints are applied to ensure the accuracy of mouth opening and closing, and local lip shape changes. For regions that prioritize style expression and overall structural stability, such as the forehead, eyes, nose, and facial contours, relatively gentler structural context constraints are applied to maintain the speaker's overall style dynamics and the naturalness of facial structure. Correspondingly, this invention constructs a semantic-aware manifold constraint loss. : ; in, This represents the constraint terms that apply to the dynamic region of the lips. This represents the constraint terms that act on the structural context region. and They are respectively and The weighting coefficients.

[0053] By using the aforementioned regionally differentiated constraint method, this invention can ensure the synchronization accuracy of key pronunciation areas while avoiding interference from random identity dynamics or style changes on key lip signals, thereby balancing local mouth shape accuracy and overall action expressiveness.

[0054] Through the aforementioned geometrically aware mean flow trajectory correction step, this invention achieves joint optimization at the latent space trajectory level and the explicit geometric manifold level: on the one hand, by correcting the target flow and adaptive mean flow loss, the curvature error of the one-step generated trajectory in the latent space is reduced; on the other hand, through semantically aware manifold-GVC constraints, the trajectory correction result is projected onto the explicit 3D surface and subjected to regional differential constraints, thereby improving the overall performance of the generated result in terms of lip-sync, facial expression, and geometric rationality. After this step, the average velocity field after trajectory correction and the corresponding regional semantic constraint results can be obtained, providing an optimization foundation for the progressive training and single-step inference in the next step.

[0055] Step S50: Adopt a phased progressive training strategy. In the training phase, instantaneous velocity is learned first and then average velocity is learned. In the inference phase, the reference time is fixed to 1. The target action endpoint is recovered by using the average velocity field predicted by the network and the voice-driven three-dimensional facial animation result is output.

[0056] The purpose of this step is to smoothly transition the model from instantaneous velocity learning to average velocity learning through a progressive training strategy, and to achieve one-step generation using the average velocity field during the inference phase, thereby balancing training stability and inference efficiency. Figure 2As shown, this step follows the geometric perception mean flow trajectory correction and is used to define the training method and final inference method of the overall model.

[0057] Specifically, this invention employs a phased, progressive training strategy. The first training phase focuses on instantaneous speed learning, setting the reference time to... r = t This approach enables the dual-time denoising Transformer to learn the instantaneous velocity field within the standard flow matching framework. This establishes a stable, fundamental cross-modal correspondence between the speech signal and 3D facial movements, and forms a better initial motion manifold under explicit geometric constraints. The training objective in this stage primarily consists of the instantaneous velocity learning loss and the aforementioned geometric loss. After completing the first stage of training, the second stage begins. This stage continues training based on the initial weights, transitioning between instantaneous velocity learning and average velocity learning through an interleaved switching mechanism: a portion of the training samples still use... r = t To maintain the model's ability to stably model local instantaneous dynamics, another part of the training samples uses a random reference time. r = u The average velocity field is learned. Simultaneously, by combining the aforementioned corrected target flow, adaptive mean flow loss, and semantically aware manifold-GVC constraints, the latent space trajectory and explicit geometric manifold are jointly optimized, thereby enabling the model to gradually acquire the global trajectory correction capability required for one-step generation.

[0058] During the inference phase, this invention no longer employs multi-step iterative sampling or stepwise numerical integration. Instead, it fixes the reference time at 1 and directly utilizes the average velocity field predicted by the network to recover the target action endpoints, outputting the corresponding 3D facial animation results. In other words, for any input speech and reference video, this invention can complete the generation of the target action sequence through a single forward computation, thereby significantly reducing inference latency and improving the real-time performance and deployment efficiency of speech-driven 3D facial animation.

[0059] Through the aforementioned progressive training and single-step inference steps, this invention achieves a smooth transition from instantaneous dynamics modeling to average dynamics modeling, and directly applies the average velocity field prediction capability obtained in the training phase to the one-step generation in the inference phase. After this process, a voice-driven 3D facial animation result with high lip-sync accuracy, expressiveness, geometric rationality, and inference efficiency can be obtained.

[0060] The key innovations of this invention are as follows: (1) Dual-time encoding driven velocity field modeling method: A method that simultaneously incorporates simulation time is proposed. t and reference time rThe dual-moment encoding mechanism establishes a unified temporal conditional representation between instantaneous velocity learning and average velocity learning through an interleaved switching strategy. It combines shape encoding, style encoding, historical action and audio feature inputs into the dual-moment denoising Transformer to achieve stable prediction of instantaneous velocity field and average velocity field.

[0061] (2) Geometrically perceptive mean flow trajectory correction method: A method is proposed to construct the target flow through Jacobi vector product and use adaptive mean flow loss to correct the trajectory of the mean velocity field. At the same time, the mean flow error in the latent space is projected onto the explicit three-dimensional surface manifold. Combined with semantic constraints divided by region, different strengths of restrictions are applied to the dynamic region of the lip and the structural context region, thereby reducing trajectory curvature and manifold deviation in one-step generation and improving the lip synchronization accuracy, action expressiveness and geometric rationality.

[0062] (3) Progressive training and single-step inference generation process: A progressive training process is proposed, which gradually transitions from instantaneous velocity learning to average velocity learning. In the inference stage, the target action endpoint is recovered directly using the average velocity field through a fixed reference time, thereby realizing a single forward generation of the target action sequence. This approach balances training stability, generative expressiveness, and real-time inference efficiency.

[0063] Compared to the current voice-driven 3D facial animation generation method MSMD, the voice-driven 3D facial animation generation model proposed in this invention, based on dual-moment encoding and geometrically perceptual mean flow correction, no longer relies on the multi-step iterative sampling process of the diffusion model. Instead, it uses a dual-moment denoising Transformer to uniformly model the instantaneous velocity field and the average velocity field, and directly uses the average velocity field to recover the target action endpoints during the inference stage, achieving single-step generation of the target action sequence. This method significantly reduces inference latency while maintaining the expressive power of facial movements, making it more suitable for applications such as real-time driving, online interaction, and engineering deployment.

[0064] Furthermore, unlike existing methods that primarily rely on latent space generation and lack explicit geometric trajectory correction, this invention introduces a manifold projection based on implicit recovery and a geometrically aware mean flow trajectory correction mechanism. This maps the generated average velocity field to the vertex space of the FLAME explicit 3D mesh, and combines a correction target flow constructed using the Jacobian vector product with semantic constraints based on region partitioning to apply differentiated constraints to the dynamic lip region and the structural context region. Therefore, this invention not only maintains high lip-sync accuracy and geometric plausibility under one-step inference conditions, but also balances the expressiveness of facial movements with overall structural stability, outperforming existing solutions in both generation quality and inference efficiency.

[0065] Furthermore, the present invention can be replaced or updated in some model details, specifically including: (1) Parametric face model and action representation: This method currently uses FLAME as the explicit 3D face geometric representation, and its parameter space and decoded mesh vertices as the basis for action modeling and geometric constraints. This part can also be replaced with other face models with explicit topological structure and expression parameter representation capabilities, such as BFM, FaceWarehouse, ICT-FaceKit, or other parametric face models or bound mesh systems that can provide 3D facial parameter sequences and mesh vertex outputs. As long as it can realize voice-driven 3D facial action representation and explicit geometric supervision, it can achieve the purpose of this invention.

[0066] (2) Audio Feature Extraction and Style Encoding: This method currently uses HuBERT to extract speech features and extracts style encoding and static shape information from the reference video through a temporal style encoder. The audio encoder can be replaced with Wav2Vec 2.0, Whisper encoder, AV-HuBERT or other speech representation models; the style encoder can also be replaced with a temporal feature extraction structure based on Transformer, temporal convolutional network, recurrent neural network or variational autoencoder, as long as it can extract identity geometric information and style dynamic information from the reference video, the same invention purpose can be achieved.

[0067] (3) Dual-moment velocity field prediction network: This method currently uses a dual-moment denoising Transformer to uniformly model the instantaneous velocity field and the average velocity field. This prediction network can also be replaced by other temporal neural network structures that can simultaneously receive dual-moment conditions, historical action context, and audio feature inputs, such as temporal convolution-based networks, Long Short-Term Memory (LSTM) networks, Gated Recurrent Units (GRUs), or other attention networks. As long as it can complete the velocity field prediction under dual-moment conditions, it can be used as an alternative implementation of this invention.

[0068] Furthermore, such as Figure 4 As shown, based on the above-mentioned speech-driven 3D facial animation generation method based on dual-time-mapping encoding and geometrically perceptual mean flow correction, the present invention also provides a speech-driven 3D facial animation generation system based on dual-time-mapping encoding and geometrically perceptual mean flow correction, wherein the speech-driven 3D facial animation generation system based on dual-time-mapping encoding and geometrically perceptual mean flow correction includes: Data preprocessing module 51 is used to input reference video and driving speech, perform data preprocessing on reference video and driving speech, and extract 3DMM parameters, static shape coefficients, style codes and audio features; The dual-moment encoding construction module 52 is used to construct the simulation time and the reference time, form a dual-moment time embedding, and combine static shape coefficients, style encoding, historical action and audio features to predict the velocity field; The implicit recovery-based manifold projection module 53 is used to recover the target action endpoint using the average velocity field and map it to the explicit three-dimensional vertex space through the FLAME decoder. The geometrically perceptual mean flow trajectory correction module 54 is used to construct the correction target flow through the Jacobian vector product, and combine geometric loss and regional semantic constraints to correct the mean flow trajectory in the latent space, so as to obtain the average velocity field after trajectory correction and the corresponding regional semantic constraint results. The progressive training and single-step inference module 55 is used to adopt a phased progressive training strategy. In the training phase, instantaneous velocity is learned first and then average velocity is learned. In the inference phase, the reference time is fixed to 1, the average velocity field predicted by the network is used to recover the target action endpoint, and the voice-driven 3D facial animation result is output.

[0069] Furthermore, such as Figure 5 As shown, based on the above-mentioned voice-driven 3D facial animation generation method and system based on dual-time encoding and geometric perception mean flow correction, the present invention also provides a terminal, which includes a processor 10, a memory 20 and a display 30. Figure 5 Only some of the terminal components are shown; however, it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.

[0070] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory. In other embodiments, the memory 20 may be an external storage device of the terminal, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc. Further, the memory 20 may include both internal and external storage devices. The memory 20 is used to store application software and various types of data installed on the terminal, such as program code installed on the terminal. The memory 20 can also be used to temporarily store data that has been output or will be output. In one embodiment, the memory 20 stores a voice-driven 3D facial animation generation program 40 based on dual-moment encoding and geometrically perceptual mean flow correction. This voice-driven 3D facial animation generation program 40 based on dual-moment encoding and geometrically perceptual mean flow correction can be executed by the processor 10, thereby implementing the voice-driven 3D facial animation generation method based on dual-moment encoding and geometrically perceptual mean flow correction in this application.

[0071] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chip, used to run program code stored in the memory 20 or process data, such as executing the speech-driven three-dimensional facial animation generation method based on dual-time encoding and geometric perception mean stream correction.

[0072] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display 30 is used to display information on the terminal and to display a visual user interface. The terminal's processor 10, memory 20, and display 30 communicate with each other via a system bus.

[0073] In one embodiment, when the processor 10 executes the speech-driven 3D facial animation generation program 40 based on dual-time encoding and geometrically perceptual mean flow correction stored in the memory 20, it implements the steps of the speech-driven 3D facial animation generation method based on dual-time encoding and geometrically perceptual mean flow correction as described above.

[0074] Furthermore, the present invention can also provide a computer-readable storage medium, wherein the computer-readable storage medium stores a speech-driven three-dimensional facial animation generation program based on dual-time encoding and geometrically perceptual mean flow correction, wherein when the speech-driven three-dimensional facial animation generation program based on dual-time encoding and geometrically perceptual mean flow correction is executed by a processor, it implements the steps of the speech-driven three-dimensional facial animation generation method based on dual-time encoding and geometrically perceptual mean flow correction as described above.

[0075] In summary, this invention provides a method, system, and terminal for voice-driven 3D facial animation generation based on dual-moment encoding and geometrically perceptual mean flow correction. The method includes: inputting a reference video and driving speech; preprocessing the reference video and driving speech to extract 3DMM parameters, static shape coefficients, style codes, and audio features; constructing a simulation time and a reference time to form a dual-moment temporal embedding, and predicting the velocity field by combining static shape coefficients, style codes, historical actions, and audio features; recovering the target action endpoint using the average velocity field and mapping it to the explicit 3D vertex space through a FLAME decoder; constructing a corrected target flow through the Jacobian vector product, and correcting the mean flow trajectory in the latent space by combining geometric loss and regional semantic constraints to obtain the trajectory-corrected average velocity field and the corresponding regional semantic constraint results; and adopting a phased progressive training strategy, first learning the instantaneous velocity and then learning the average velocity during the training phase, and fixing the reference time to 1 during the inference phase, recovering the target action endpoint using the average velocity field predicted by the network, and outputting the voice-driven 3D facial animation result. This invention uses a dual-time denoising Transformer to stably model the velocity field and combines geometrically perceptual trajectory learning to explicitly constrain the one-step generated path and perform region-aware correction. This achieves low-latency, high-efficiency and geometrically reasonable voice-driven 3D facial animation generation while ensuring high action expressiveness and lip-sync accuracy.

[0076] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal that includes that element.

[0077] Of course, those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware (such as a processor, controller, etc.). The program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The computer-readable storage medium can be a memory, magnetic disk, optical disk, etc.

[0078] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.

Claims

1. A speech-driven 3D facial animation generation method based on dual-time-time encoding and geometrically perceptual mean flow correction, characterized in that, The speech-driven 3D facial animation generation method based on dual-time-time encoding and geometrically perceptual mean flow correction includes: Input reference video and driving speech, perform data preprocessing on the reference video and driving speech, and extract 3DMM parameters, static shape coefficients, style codes and audio features; Simulation time and reference time are constructed to form a dual-time embedding, and the velocity field is predicted by combining static shape coefficients, style coding, historical actions and audio features. The target motion endpoints are recovered using the average velocity field and mapped to an explicit 3D vertex space using the FLAME decoder; The target flow is corrected by constructing the Jacobi vector product and combining geometric loss and regional semantic constraints to correct the mean flow trajectory in the latent space, thus obtaining the mean velocity field after trajectory correction and the corresponding regional semantic constraint results. A phased progressive training strategy is adopted. In the training phase, instantaneous velocity is learned first and then average velocity is learned. In the inference phase, the reference time is fixed at 1, and the target action endpoint is recovered by using the average velocity field predicted by the network, and the voice-driven 3D facial animation result is output.

2. The speech-driven 3D facial animation generation method based on dual-time-time encoding and geometrically perceptual mean flow correction according to claim 1, characterized in that, The input reference video and driving speech are preprocessed to extract 3DMM parameters, static shape coefficients, style codes, and audio features, specifically including: For the input reference video, parametric face reconstruction and tracking methods are used to extract the 3DMM parameters corresponding to each frame. The FLAME model is used to obtain the explicit parameter representation of the facial action sequence, and static shape coefficients are extracted from the reference video to characterize the speaker's static identity geometry. Simultaneously, a temporal style encoder is used to encode the reference action sequence corresponding to the reference video to obtain the style code. S ; For the input driving speech, a pre-trained HuBERT audio encoder is used to extract audio features corresponding to the time series. .

3. The speech-driven 3D facial animation generation method based on dual-time-time encoding and geometrically perceptual mean flow correction according to claim 2, characterized in that, The input reference video and driving speech are preprocessed to extract 3DMM parameters, static shape coefficients, style codes, and audio features. This process also includes: Record the actual facial movement sequence as the target movement. And sample noise actions from a standard Gaussian distribution. Then, based on the simulation time t Constructing noisy action states ; Extract a certain length of historical action sequence from a known action sequence. It is used to provide temporal context constraints for the generation of actions at the current moment.

4. The speech-driven 3D facial animation generation method based on dual-time-time encoding and geometrically perceptual mean flow correction according to claim 3, characterized in that, The velocity field includes an instantaneous velocity field or an average velocity field.

5. The speech-driven 3D facial animation generation method based on dual-time-time encoding and geometrically perceptual mean flow correction according to claim 4, characterized in that, The construction of the simulation time and reference time forms a dual-time embedding, which is then combined with static shape coefficients, style coding, historical actions, and audio features to predict the velocity field. Specifically, this includes: From the interval Internal sampling simulation time t Used to represent the current noisy action state The corresponding generation progress is recorded, and a random time interval is sampled independently. u ~ As a construction reference time r Candidate time variables; Design an interleaving switching mechanism for the reference time. r The assignment is performed in stages. In the first training stage, the reference time is set to... This causes the network to degenerate into learning instantaneous velocity fields. In the second training phase, the reference time is set to random time. This relates the average speed of network learning from any reference time to the current time, and the reference time. r Represented as: ; Simulation time t and reference time r Sine wave position encoding is performed separately, and then each is input into the corresponding multilayer perceptron. and The mapping process yields two temporal embeddings, which are then added together to form a composite temporal embedding. : ; in, This represents the sinusoidal position coding function. and These represent the multilayer perceptrons that process the simulation time and the reference time, respectively. static shape factor and style coding S The parts are combined to form an identity code. Composite time is embedded by adding elements one by one. Injected into the identity encoding to construct a time-aware identity prefix. ; Time-aware identity prefix Historical action sequence and noisy action state The sequences are concatenated into a composite input sequence and then input into a two-time-phase denoising Transformer for modeling. The dual-time denoising Transformer utilizes bidirectional self-attention to establish global dependencies between input segments, and then uses masked cross-attention to extract audio features. An action generation process is introduced, and an alignment mask is used to limit the local audio window corresponding to each action position. After processing by the Transformer decoder module and the linear projection layer, the velocity field prediction result corresponding to the current state is output: when At that time, the instantaneous velocity field is output. ; when At that time, the average velocity field is output. .

6. The speech-driven 3D facial animation generation method based on dual-time coding and geometrically perceptual mean flow correction according to claim 5, characterized in that, The process of recovering the target action endpoint using the average velocity field and mapping it to an explicit 3D vertex space via the FLAME decoder specifically includes: For any simulation time t Noisy action state The target motion is implicitly recovered using the average velocity field. During the inference phase, with the reference time fixed at 1, the displacement estimate from the current state to the target motion endpoint is directly obtained. The recovered target motion endpoint is expressed as: ; in, This represents the target action endpoint recovered by the network prediction. This represents the average velocity field output when the reference time is fixed at 1. This represents the remaining time span from the current simulation moment to the target endpoint; The recovered motion parameters are decoded into corresponding 3D facial mesh vertices using the FLAME decoder to obtain the predicted mesh, while the real motion parameters are decoded into the real mesh, and the two are compared in the explicit vertex space. Constructing geometric loss : ; in, This represents the vertex constraint loss, used to constrain the difference between the predicted mesh and the true mesh in the explicit vertex space. This represents the velocity constraint loss, used to suppress abrupt changes in vertex motion between adjacent time steps and reduce temporal jitter in facial movements. This represents the coefficient constraint loss, used to maintain the continuity of motion parameters in the boundary region and structural connection region, avoiding unnatural local structures or geometric distortions. , and These are the weight coefficients for vertex constraint loss, velocity constraint loss, and coefficient constraint loss, respectively.

7. The speech-driven 3D facial animation generation method based on dual-time-time encoding and geometrically perceptual mean flow correction according to claim 6, characterized in that, The method of constructing a corrected target flow through the Jacobian vector product and modifying the mean flow trajectory in the latent space by combining geometric loss and regional semantic constraints specifically includes: Constructing a corrected target flow for monitoring the mean velocity field : ; in, This represents the target action corresponding to the actual action sequence. This represents noise action sampled from a standard Gaussian distribution. This represents the average velocity field predicted by the denoised Transformer at two time points, and the derivative term is calculated using the Jacobian vector product; Construct the mean flow loss and apply it to the predicted mean velocity field. and correct target flow Constrain the differences between them; The mean flow error is constructed based on the difference between the predicted mean velocity field and the corrected target flow. And the mean stream error is obtained through the FLAME decoder in a local approximation sense. Mapped to explicit mesh vertex deformation : ; in, Indicates FLAME decoder, Indicates the current baseline action state. These are the small perturbation coefficients used for local approximation; Constructing semantically aware manifold constraint loss : ; in, This represents the constraint terms that apply to the dynamic region of the lips. This represents the constraint terms that act on the structural context region. and They are respectively and The weighting coefficients.

8. The speech-driven 3D facial animation generation method based on dual-time-time encoding and geometrically perceptual mean flow correction according to claim 7, characterized in that, The method employs a phased, progressive training strategy. During the training phase, instantaneous velocity is learned first, followed by average velocity. In the inference phase, the reference time is fixed at 1, and the target action endpoint is recovered using the average velocity field predicted by the network. The result is then output as a voice-driven 3D facial animation. Specifically, this includes: A phased, progressive training strategy is adopted. The first training phase focuses on instantaneous speed learning, with the reference time set to [missing information]. r = t This enables the dual-time denoising Transformer to learn the instantaneous velocity field under the standard flow matching framework, establish a stable basic cross-modal correspondence between speech signals and three-dimensional facial movements, and form an initial motion manifold under explicit geometric constraints. After completing the first phase of training, the second phase begins. This phase continues training based on the initial weights, using an alternating switching mechanism to transition between instantaneous velocity learning and average velocity learning: a portion of the training samples are... r = t To maintain the model's ability to stably model local instantaneous dynamics, another part of the training samples uses a random reference time. r = u , used to learn the average velocity field; During the inference phase, the reference time is fixed at 1, and the target action endpoint is recovered directly using the average velocity field predicted by the network, and the corresponding 3D facial animation result is output.

9. A speech-driven 3D facial animation generation system based on dual-time-time encoding and geometrically perceptual mean flow correction, characterized in that, The speech-driven 3D facial animation generation system based on dual-time-time encoding and geometrically perceptual mean flow correction includes: The data preprocessing module is used to input reference video and driving speech, perform data preprocessing on the reference video and driving speech, and extract 3DMM parameters, static shape coefficients, style codes and audio features; The dual-moment encoding construction module is used to construct the simulation time and reference time, forming a dual-moment time embedding, and combining static shape coefficients, style encoding, historical action and audio features to predict the velocity field; The implicit recovery-based manifold projection module is used to recover the target action endpoints using the average velocity field and map them to the explicit 3D vertex space through the FLAME decoder. The geometrically perceptual mean flow trajectory correction module is used to construct the target flow through the Jacobian vector product and, combined with geometric loss and regional semantic constraints, correct the mean flow trajectory in the latent space to obtain the mean velocity field after trajectory correction and the corresponding regional semantic constraint results. The progressive training and single-step inference module is used to adopt a phased progressive training strategy. In the training phase, instantaneous velocity is learned first and then average velocity is learned. In the inference phase, the reference time is fixed to 1, the average velocity field predicted by the network is used to recover the target action endpoint, and the voice-driven 3D facial animation result is output.

10. A terminal, characterized in that, The terminal includes: a memory, a processor, and a voice-driven 3D facial animation generation program based on dual-time encoding and geometrically perceptual mean flow correction, which is stored in the memory and can run on the processor. When the voice-driven 3D facial animation generation program based on dual-time encoding and geometrically perceptual mean flow correction is executed by the processor, it implements the steps of the voice-driven 3D facial animation generation method based on dual-time encoding and geometrically perceptual mean flow correction as described in any one of claims 1-8.