A real-time digital human generation system and method based on low-computation voice driving

By decoupling end-to-end generation technology into low-dimensional lip-move parameter prediction and local area gradual rendering, and introducing PID feedback mechanism and heterogeneous hardware acceleration, the contradiction between computing efficiency and real-time and insufficient audio and video synchronization accuracy in the existing technology is solved, and an efficient and low-latency digital human generation system is realized.

CN119991893BActive Publication Date: 2025-06-17LIANGSHENG DIGITAL ARTIFICIAL INTELLIGENCE (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510432794.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-06-17
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

There are problems such as the contradiction between computing efficiency and real-time, imbalance in resource occupation and generalization capabilities, insufficient audio and video synchronization accuracy in existing voice-driven digital life technology, and it is difficult to achieve low-latency response and high-precision synchronization on the mobile terminal.

Method used

By decoupling end-to-end generation into low-dimensional lip-move parameter prediction and local area progressive rendering, LSTM prediction parameters are used for linear interpolation to generate the target mouth shape, introducing a PID feedback mechanism to calibrate the audio and video timestamp deviation in real time, and integrating NPU accelerated inference on the mobile side, the desktop side uses CUDA parallel computing, and embedded devices deploy a binary convolutional network to maximize resource efficiency.

Benefits of technology

It realizes the model volume is greatly compressed while maintaining high fidelity, adapting to the real-time interaction requirements of multiple platforms, reducing latency and computing complexity, improving audio and video synchronization accuracy and system adaptability to complex voice inputs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991893B_ABST
    Figure CN119991893B_ABST
Patent Text Reader

Abstract

The present invention discloses a real-time digital human generation system and method based on low-computation speech driving, which relates to the field of digital human interaction technologies. The system includes: an audio processing module configured to receive speech input in real time and extract audio feature vectors; a driving and rendering module for mapping the audio feature vectors into parameters representing mouth movements, generating dynamic mouth images based on preprocessed static face reference data and the parameters, and fusing them with reference images; a synchronization control module for ensuring the synchronization of audio features and rendered video frames according to a timestamp alignment mechanism and a PID feedback control algorithm; and a dynamic scheduling module for monitoring the hardware resource load in real time and achieving dynamic allocation of computing resources through multi-thread parallelism and task priority adjustment. According to the technical solution of the present application, breakthroughs in low latency, high fidelity, and low energy consumption of digital humans can be achieved in mobile, embedded, and multi-platform scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of digital human interaction technologies, and particularly to a real-time digital human generation system and method based on low-computation voice driving. Background Art

[0002] As a cutting-edge field integrating artificial intelligence and computer vision, digital human interaction technologies aim to generate virtual avatars with high-fidelity expressions and lip synchronization in real time through voice or text input. Traditional technical solutions are mainly divided into two categories: one is a rule-driven animation system that drives a 3D model through a predefined phoneme mouth shape mapping table, and its lip movement generation depends on a limited number of manual animation sequences, resulting in rigid expressions and difficulty in covering the subtle changes of continuous speech; the other is an end-to-end generation model based on deep learning that directly synthesizes face videos by training a neural network with a large-scale dataset. Although the naturalness has been significantly improved, the number of model parameters is huge, and the inference latency is as high as hundreds of milliseconds, making it difficult to meet the real-time interaction requirements of mobile devices.

[0003] In addition, although existing lightweight solutions reduce the computational amount through model compression, they still face two major bottlenecks: 1. The accuracy loss is significant after model pruning and quantization, especially in fast speech scenarios, where lip jitter or mouth shape lag is likely to occur; 2. The rendering process is not optimized for local regions, and still generates pixels one by one for the full-resolution image, resulting in excessive GPU video memory occupancy and difficulty in maintaining a stable frame rate on low-end devices. Therefore, based on the above problems, the present invention proposes a real-time digital human generation system and method based on low-computation voice driving. Summary of the Invention

[0004] Technical Objectives

[0005] In order to solve the above problems, the objective of the present invention is to provide a real-time digital human generation system and method based on low-computation voice driving, aiming to solve the core problems existing in the existing voice-driven digital human generation technologies, such as the contradiction between computational efficiency and real-time performance, the imbalance between resource occupancy and generalization ability, and the insufficient accuracy of audio-visual synchronization. It gets rid of the dependence on high-complexity models in traditional solutions, enables mobile devices to achieve low-latency response, avoids lip jitter and accuracy loss, and solves the problem that the fixed-delay synchronization mechanism cannot dynamically compensate for hardware fluctuations or network jitters.

[0006] Technical Solutions

[0007] To achieve the above object, the present invention provides a real-time digital human generation system and method based on low-computation speech driving. The system and method decouple the traditional end-to-end generation into a mapping from speech features to low-dimensional lip motion parameters and local mouth shape rendering based on pre-stored static face reference data; pre-store multi-scale mouth deformation templates, and linearly interpolate to generate the target mouth shape through LSTM prediction parameters; introduce a PID feedback mechanism to calibrate the audio-visual timestamp deviation in real time; integrate NPU on the mobile side to accelerate LSTM inference, utilize CUDA for parallel interpolation on the desktop side, and deploy a binary convolutional network on embedded devices to maximize resource efficiency; extract static face features from a single video sample and generate personalized digital humans in combination with dynamic lip motion parameters. This solution greatly compresses the model volume while maintaining high fidelity and adapts to the real-time interaction requirements of multiple platforms.

[0008] In a first aspect, the present invention provides a real-time digital human generation system based on low-computation speech driving, including:

[0009] An audio processing module configured to receive speech input in real time and extract low-dimensional audio feature vectors;

[0010] A driving and rendering module, including a lip motion parameter prediction unit and a region rendering unit based on a lightweight deep learning model; the lip motion parameter prediction unit maps the audio feature vectors to low-dimensional parameters representing mouth movements, and the region rendering unit generates dynamic mouth images based on the preprocessed static face reference data and the low-dimensional parameters, and fuses them with the reference images through Alpha blending technology;

[0011] A synchronization control module for ensuring millisecond-level synchronization of audio features and rendered video frames according to the timestamp alignment mechanism and the PID feedback control algorithm;

[0012] A dynamic scheduling module for real-time monitoring of the hardware resource load and realizing dynamic allocation of computing resources through multi-thread parallelism and task priority adjustment;

[0013] Among them, the single-frame inference computation of the system does not exceed 50 MFlops, and the end-to-end response delay is less than 100 milliseconds.

[0014] Further, the audio processing module frames the 16 kHz sampled speech with a frame length of 25 milliseconds and a frame shift of 10 milliseconds, multiplexes the Mel spectrum features of the overlapping regions of adjacent frames, and dynamically estimates the background noise and generates a mask matrix using a lightweight algorithm based on spectral subtraction;

[0015] Among them, the audio feature extraction uses a two-layer LSTM network, and the parameter quantity is less than 300 KB after 8-bit quantization.

[0016] Further, the preprocessing process of the static face reference data includes:

[0017] Extract high-dimensional face features using a teacher network and train a lightweight student network through knowledge distillation;

[0018] Concatenate the feature vectors of the 128×128 pixel image of the mouth region and the 512×512 pixel global reference image to enhance local details and global consistency.

[0019] Furthermore, the driving and rendering module pre-stores multiple groups of mouth deformation base meshes, generates a target mouth shape deformation field by interpolating the base meshes with the low-dimensional parameters, uses a lightweight convolutional network to perform pixel-level repair on the gaps and edges of the deformed image. The repair network includes a 3-layer 3×3 convolutional kernel structure, the output layer uses a Sigmoid activation function, and calls the GPU skeletal animation pipeline or Metal / OpenGL ES graphics API to achieve real-time rendering of the deformed mesh.

[0020] Furthermore, the area rendering unit divides the mouth region into 8×8 sub-blocks, only triggers redrawing for sub-blocks whose deformation amplitude exceeds the threshold, and the background region is compressed using differential coding and updated in full every 10 frames to reduce the pixel filling amount.

[0021] Furthermore, the PID feedback control algorithm of the synchronization control module satisfies:

[0022]

[0023] where is the control signal for fine-tuning the video frame generation interval; , is the audio-visual timestamp error; is the independent variable in the continuous time domain; ; ; .

[0024] Furthermore, the dynamic scheduling module realizes dynamic allocation of computing resources based on a load-aware strategy, including:

[0025] When the CPU utilization rate exceeds 70%, automatically reduce the execution frequency of non-critical tasks and enable GPU hardware acceleration;

[0026] When it is detected that the rendering times out for 5 consecutive frames, dynamically reduce the output frame rate from 30 FPS to 25 FPS;

[0027] The computing resource scheduling priority order is: audio feature extraction, lip movement parameter prediction, background texture cache update.

[0028] Furthermore, when the input signal is lost, the system predicts subsequent lip movement parameters based on Kalman filtering and smoothly transitions to the default closed mouth shape; when the GPU video memory overflows, it automatically switches to the CPU software rendering mode and reduces the resolution to 640×480 pixels; in the case of network jitter, the look-ahead buffer pool is used to dynamically adjust the buffer depth, discard expired frames, and insert motion compensation frames.

[0029] Furthermore, the system integrates the ARM NEON instruction set on the mobile side to optimize matrix operations and calls the NPU to accelerate LSTM inference; on the desktop side, CUDA parallel computing is used to implement deformation field interpolation, and multi-vendor GPU heterogeneous scheduling is supported through OpenCL; on embedded devices, a binary convolutional network is deployed to compress the rendering model parameter quantity to less than 1 MB.

[0030] Furthermore, it also includes a joint interpolation module, which is used to adaptively adjust the deformation field interpolation weight and define the spatio-temporal energy factor by dynamically analyzing the spectral energy distribution of continuous audio frames and the lip movement parameter change rate. Its value is determined by the audio Mel spectral energy of the current frame and the lip movement parameter change rate of adjacent frames jointly:

[0031]

[0032] In the formula, is the mixed weight of spectral energy and motion rate; is the audio Mel spectral energy of the current frame; is the lip movement parameter change rate of adjacent frames; and are normalization factors;

[0033] Based on select the optimal interpolation template from the pre-stored base mesh library:

[0034]

[0035] In the formula, is the spatio-temporal joint interpolation output; is the high-frequency deformation base mesh; is the index of the discrete time step, which is used to ensure temporal continuity and avoid frame skipping; is the low-frequency deformation base mesh.

[0036] By fusing the spectral energy of speech and the dynamic change characteristics of lip movement parameters, a spatio-temporal joint interpolation weight function is constructed, which optimizes the generation process of the deformation field. This mechanism significantly improves the smoothness of mouth shape transition in fast speech scenarios, eliminates the lip jitter or lag caused by a fixed interpolation step size, reduces the redundant overhead of deformation field calculation, and enhances the adaptability of the system to complex speech inputs.

[0037] Furthermore, it also includes a rendering optimization module for dynamically evaluating the lighting consistency, texture continuity, and motion naturalness between the generated mouth region and the reference face image, and adjusting the rendering parameters in real time through backpropagation of gradients, and defining a fusion quality score in the adversarial loss function for optimizing the rendering parameters :

[0038]

[0039] In the formula, is the mean square error of brightness between the generated region and the reference image; is the LBP texture similarity; is the difference in the motion amplitude of optical flow between adjacent frames; , , , with asymmetric allocation of weight coefficients;

[0040] According to Backpropagate to optimize the Alpha blending coefficient and the interpolation weight of the deformation field:

[0041]

[0042] In the formula, is the output of parameter adjustment; is the learning rate; is the Alpha blending transparency parameter.

[0043] By introducing a multi-modal discriminator network, jointly evaluating the lighting consistency, texture continuity, and motion naturalness of the generated region, and back-optimizing the rendering parameters, effectively solves the problems of skin color deviation and edge artifacts that may be caused by local rendering. It improves the seamless fusion quality between the mouth region and the reference face image, reduces the number of iterative optimizations, and realizes the collaborative optimization of high-fidelity rendering and computational efficiency.

[0044] In the second aspect, the present invention also provides a real-time digital human generation method based on low-computation speech driving, and the method is based on the system described in the first aspect above, including:

[0045] Convert the speech signal into lip movement parameters through a lightweight network to describe the dynamic characteristics such as mouth opening and closing, displacement, etc.;

[0046] Based on pre-stored static face reference data, generate a target mouth shape through a dynamic deformation interpolation base mesh library, and combine Alpha blending technology to achieve seamless fusion with the reference image;

[0047] Calculate the interpolation weight according to the speech spectrum energy and the change rate of lip movement parameters, and select the base network template according to the interpolation weight function to generate a smooth deformation field;

[0048] Real-time adjust the audio-visual synchronization error through the PID algorithm, and dynamically allocate task priorities according to the CPU / GPU utilization rate.

[0049] In a third aspect, the present invention also provides a computer device, including a management platform and a memory. The management platform is connected to the memory. The memory is used to store a computer program, and the management platform is used to execute the computer program stored in the memory, so that the computer device executes and implements the aforementioned real-time digital human generation method based on low-computation speech driving.

[0050] In a fourth aspect, the present invention also provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium. When the computer program is executed by the management platform, it implements the aforementioned real-time digital human generation method based on low-computation speech driving.

[0051] The present invention decouples speech feature mapping and local mouth shape rendering through a two-stage lightweight inference pipeline, combines a dynamic deformation interpolation base mesh library with a PID adaptive synchronization control mechanism to achieve efficient conversion and precise synchronization from speech to lip shape; adopts a cross-platform heterogeneous hardware acceleration strategy, integrates NPU / CUDA / embedded binary network to adapt to multi-terminal computing power, and reduces the deployment complexity through a training-free general generation framework. The exponential improvement of its computing efficiency and the extreme compression of end-to-end latency, while ensuring high-fidelity lip synchronization and visual consistency, break through the bottlenecks of traditional solutions in terms of real-time performance, resource occupancy, and long-term operation stability on mobile devices, providing a systematic solution with both efficiency and universality for lightweight digital human interaction scenarios.

[0052] Beneficial Effects

[0053] By implementing the real-time digital human generation system and method based on low-computation speech driving provided by the present invention, the following technical effects are achieved:

[0054] (1) By decoupling end-to-end generation into low-dimensional lip movement parameter prediction and local area progressive rendering, the present invention breaks through the high-computation bottleneck of traditional full-image generation. Reducing the computational complexity to one-tenth of the traditional solution while retaining personalized face features provides an extensible architecture basis for real-time interaction on low-computation devices.

[0055] (2) Dynamically calibrate the audio-visual timestamp deviation based on the PID feedback mechanism, overcoming the cumulative error defect of the fixed-delay synchronization strategy. This enables the system to maintain millisecond-level audio-visual synchronization accuracy even in scenarios with fluctuating hardware resources or unstable inputs, ensuring the consistency of the user experience during long-term operation.

[0056] (3) Construct a spatio-temporal joint interpolation weight function by fusing the dynamic change characteristics of speech spectral energy and lip movement parameters, optimizing the generation process of the deformation field. This mechanism significantly improves the smoothness of mouth shape transitions in fast speech scenarios, eliminates lip jitter or lag caused by fixed interpolation steps, reduces redundant overhead in deformation field calculations, and enhances the system's adaptability to complex speech inputs.

[0057] (4) By introducing a multi-modal discriminator network, jointly evaluate the lighting consistency, texture continuity, and motion naturalness of the generated region, and reverse-optimize the rendering parameters to effectively solve the skin color deviation and edge artifact problems that may be caused by local rendering. This improves the seamless fusion quality between the mouth region and the reference face image, reduces the number of iterative optimizations, and achieves the collaborative optimization of high-fidelity rendering and computational efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] To make the above-mentioned real-time digital human generation system and method based on low computational cost speech driving of the present invention more clearly understandable, the drawings required for the specific implementation manners of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings without creative efforts.

[0059] Figure 1 represents the architecture diagram of the real-time digital human generation system;

[0060] Figure 2 represents the schematic diagram of the system data flow;

[0061] Figure 3 represents the schematic diagram of the optimization strategy. DETAILED DESCRIPTION OF THE INVENTION

[0062] Example 1:

[0063] A real-time digital human generation system and method based on low computational cost speech driving are provided, and the system architecture is as shown in Figure 1As shown in the figure, it includes: an audio processing module configured to receive voice input in real time and extract low-dimensional audio feature vectors; a driving and rendering module including a lip movement parameter prediction unit and a region rendering unit based on a lightweight deep learning model; the lip movement parameter prediction unit maps the audio feature vectors into low-dimensional parameters representing mouth movements, and the region rendering unit generates dynamic mouth images based on preprocessed static face reference data and the low-dimensional parameters, and fuses them with the reference images through Alpha blending technology; a synchronization control module for ensuring millisecond-level synchronization of audio features and rendered video frames according to a timestamp alignment mechanism and a PID feedback control algorithm; a dynamic scheduling module for real-time monitoring of hardware resource loads and dynamically allocating computing resources through multi-thread parallelism and task priority adjustment; where the single-frame inference calculation amount of the system does not exceed 50 MFlops, and the end-to-end response delay is less than 100 milliseconds. Details are as follows.

[0064] The system data flow is as Figure 2 shown. Starting from voice input, after audio preprocessing and feature extraction, it is sent to a sequence model to generate low-dimensional mouth movement parameters, and combined with offline preprocessed reference face data to finally generate video frames.

[0065] In the offline stage, the system extracts the static face and key part features of the target person by pre-collecting video samples of the target person, specifically including:

[0066] Using mature face detection algorithms, such as those based on MediaPipe or Dlib, to detect 68 or more key points from video frames; among them, the mouth area is extracted with emphasis, for example, selecting key points numbered 48 to 68 as the core data for subsequent lip movement generation;

[0067] Calculating the cropping area containing the mouth according to the detected key points; usually, the minimum bounding rectangle based on the mouth key points is used to expand the boundary, for example, each side is expanded by 10% to ensure complete inclusion of the mouth contour; subsequently, the cropping area is scaled to a fixed size, such as 128×128 pixels, specifically determined according to subsequent rendering requirements;

[0068] Using a pre-trained lightweight convolutional neural network to extract low-dimensional feature vectors of the target face, such as 256 dimensions, as reference information for the subsequent generation stage, so as to dynamically update only the mouth area while maintaining the personality characteristics of the face.

[0069] After the preprocessing is completed, the reference image of the target person, the normalized mouth area and its corresponding feature vectors are stored in local or embedded memory, and the total data volume is controlled within 3MB; through memory mapping and streaming loading technology, resources are loaded in batches to ensure extremely low startup latency and fast access to necessary data during operation.

[0070] The online stage is the core process of real-time voice input driving the lip movement generation of a digital human. The audio processing module frames the 16 kHz sampled speech with a frame length of 25 milliseconds and a frame shift of 10 milliseconds, extracts 13-dimensional basic features using the MFCC algorithm, supplements them with energy, first-order and second-order differences to form a total of 39-dimensional feature vectors, and uses a lightweight LSTM or CNN model for real-time processing to ensure low processing latency. It also reuses the Mel spectrum features in the overlapping area of adjacent frames, and uses a lightweight algorithm based on spectral subtraction to dynamically estimate background noise and generate a mask matrix.

[0071] Use a lightweight temporal neural network to map audio features to a parameter vector representing the mouth movement state. The network adopts a two-layer LSTM structure with 64 hidden units in each layer; the input is the audio features of several consecutive frames; the output is a 20-dimensional lip movement parameter vector describing the shape change of the mouth; after pruning and quantization optimization, the network parameters are controlled within hundreds of thousands, making the computational amount per frame inference about 39 MFlops.

[0072] The driving and rendering module pre-stores multiple groups of mouth deformation base meshes, generates the target mouth shape deformation field by interpolating the base meshes with the low-dimensional parameters, and uses a lightweight convolutional network to perform pixel-level repair on the gaps and edges of the deformed image. The repair network contains a 3-layer 3×3 convolution kernel structure, the activation function is ReLU, and the output layer uses the Sigmoid activation function; the input includes the target human face reference image, the preprocessed mouth area and the 20-dimensional lip movement parameters; the output is a new mouth image, which is fused with the mouth area of the reference image using the Alpha blending technique, and the edges are smoothed to eliminate obvious transitions. The entire image generation process is completed at a resolution of 128×128 pixels, with a latency lower than 30ms, ensuring that the system reaches a frame rate of 30fps or higher. And call the GPU skeletal animation pipeline or Metal / OpenGL ES graphics API to achieve real-time rendering of the deformed mesh.

[0073] The area rendering unit divides the mouth area into 8×8 sub-blocks, and only triggers redrawing for sub-blocks whose deformation amplitude exceeds the threshold. The background area is compressed using differential coding and updated in full every 10 frames to reduce the pixel filling amount.

[0074] The details of the low-computation rendering method of the driving and rendering module are as follows:

[0075] Only redraw the area around the mouth frequently, and use the previous frame buffer for unchanged areas such as the background.

[0076] Prepare multiple mouth shape deformation meshes in advance and quickly generate mouth animations by interpolation.

[0077] Bind the facial blendshape parameters in the 3D digital human scene, and perform interpolation and rendering by the GPU. The CPU only needs to provide the driving parameters.

[0078] The system ensures the consistency of audio and video through a synchronization control mechanism, and each audio frame and video frame are both attached with timestamps and ;

[0079] Theoretically, the video frame playback time satisfies:

[0080]

[0081] wherein, is the theoretical playback timestamp of the video frame, indicating when the frame should be displayed; is the serial number of the video frame; is the video frame rate, that is, the number of video frames rendered per second;

[0082] The calculation formula for the time difference is:

[0083]

[0084] wherein, is the audio-visual synchronization error, indicating the deviation between the actual playback time of the video frame and the timestamp of the corresponding audio frame; is the actual playback timestamp of the video frame.

[0085] For example , then the generation delay of the next frame needs to be adjusted:

[0086]

[0087] wherein, is the generation delay of the adjusted next frame, used to dynamically control the rendering rhythm; is the generation delay of the previous frame, that is, the actual processing time of the current frame; is the adjustment factor; is the reference time interval.

[0088] The PID feedback control algorithm of the synchronization control module satisfies:

[0089]

[0090] wherein, is the control signal, used to finely adjust the video frame generation interval; , is the audio-visual timestamp error; is the independent variable in the continuous time domain; ; ; .

[0091] The synchronization control module adopts unified timestamps and buffer control, attaches timestamps when generating audio frames and performs precise alignment during the rendering phase, and introduces an adaptive synchronization adjustment algorithm. If lip movement lags or leads, the system fine-tunes the rhythm of subsequent frames, such as slightly accelerating rendering or discarding expired frames, so as to maintain audio-visual synchronization even during long-term operation. It is more flexible than fixed-delay buffering and can better adapt to differences in device and network environments.

[0092] The optimization strategies adopted by the system are as Figure 3 shown, which demonstrates three major strategies of the system in real-time optimization: dynamic scheduling, cache optimization, and model pruning and quantization. By these strategies working together, it can help the system achieve the goals of reducing latency, computational load, and power consumption.

[0093] The dynamic scheduling module monitors the CPU, GPU load, and memory utilization of the device in real-time, dynamically adjusts the task priorities and time slice allocations of each module according to the load. It uses multi-threading or asynchronous tasks to process audio, expression generation, and rendering respectively without blocking. When the CPU load exceeds the preset threshold, it reduces the frequency of non-critical tasks to ensure that voice processing and lip movement rendering run preferentially. Among them, the scheduling formula is:

[0094]

[0095] In the formula, is the scheduling output; is the delay of the previous frame; is the adjustment factor; is the current CPU occupancy rate; is the preset threshold; is the load threshold.

[0096] The dynamic scheduling module ensures the real-time parallel and efficient operation of each module through a multi-threading scheduling strategy. The multi-threading scheduling strategy includes:

[0097] Audio acquisition and feature extraction continuously run in a separate thread;

[0098] Lip movement prediction and image generation are processed in another thread, and audio frame data is buffered through a queue to achieve asynchronous parallelism;

[0099] Dynamically adjust the priorities and execution times of each module according to the current device resources to ensure the smooth operation of the system even under limited hardware resources.

[0100] The system stores all model parameters, texture materials, and intermediate cache data, supports efficient loading on demand, and loads deep learning models, reference images, and animation materials in batches, using memory mapping and streaming loading to reduce the startup latency.

[0101] The system reuses consecutive frame data through caching technology to reduce repeated calculations, specifically including: the audio module caches the overlapping parts of consecutive frames; the rendering module caches the background and static areas and only updates the dynamically changing areas; when the characteristics of consecutive audio frames change slightly, the generation result of the previous frame is reused, and interpolation is used for smooth transition.

[0102] The system uses pruning and quantization technologies to reduce model parameters and computational complexity, specifically including: pruning the expression generation network to remove low-contribution weights, with a pruning ratio of 30% - 50%; converting the weights from 32-bit floating-point numbers to fixed-point integers through 8-bit quantization technology; using knowledge distillation to train lightweight models and providing multi-level model versions to adapt to different hardware environments. After optimization by pruning and quantization technologies, the computational amount per frame of model inference is reduced to about 39 MFlops.

[0103] To reduce the computational consumption of the model, the system uses a large model to guide the training of lightweight models, ensuring that the output accuracy is not significantly reduced, and dynamically selects high-precision or streamlined models according to device performance to ensure smooth operation on low-end devices.

[0104] To ensure the stable operation of the system in various environments, the system is designed with a multi-level exception handling strategy, specifically including: if a microphone interruption or abnormal data is detected, the current mouth shape is maintained or switched to the default static expression, while continuously monitoring the signal recovery; when abnormal device load or frequency reduction due to high temperature is detected, the rendering details or frame rate are automatically reduced, such as from 30 FPS to 25 FPS, and non-critical tasks are paused; a jitter buffer mechanism is used to sort out-of-order data, and the buffer length is dynamically adjusted according to timestamps, and expired data is discarded to ensure synchronization; timeout monitoring and error capture mechanisms are embedded in each module, and if an exception occurs, it will be automatically restarted and the error log will be recorded.

[0105] The system integrates the ARM NEON instruction set in the mobile terminal to optimize matrix operations and calls the NPU to accelerate LSTM inference; on the desktop, CUDA parallel computing is used to implement deformation field interpolation, and multi-vendor GPU heterogeneous scheduling is supported through OpenCL; in embedded devices, a binary convolutional network is deployed to compress the rendering model parameters to less than 1 MB.

[0106] The system is developed using C++ combined with cross-platform libraries and generates Android, iOS, Windows, and Linux versions through CMake, specifically including:

[0107] On the mobile terminal, local libraries are called through JNI and Swift / Objective-C interfaces, and ARM NEON and iOS Metal are fully utilized for acceleration;

[0108] On Windows and Linux, multi-threading and GPU acceleration technologies are used to ensure high-frame-rate output;

[0109] After testing, the system can stably output 25 - 30 FPS on each platform, and the response latency is less than 100 ms.

[0110] The comparison of the CPU / GPU occupancy rates of the system before and after optimization on each platform is shown in Table 1.

[0111] Table 1. CPU / GPU Occupancy Rate Comparison Chart

[0112] Platform Before Optimization After Optimization Android 75% 45% iOS 70% 40% Windows 50% 30% Linux 55% 35%

[0113] The comparison of the frame rates on each platform before and after optimization is shown in Table 2.

[0114] Table 2. Frame Rate Comparison Chart

[0115] Platform Before Optimization After Optimization Android 20FPS 30FPS iOS 24FPS 30FPS Windows 30FPS 30FPS Linux 28FPS 30FPS

[0116] The comparison of the audio-visual response latencies on each platform before and after optimization is shown in Table 3.

[0117] Table 3. Response Latency Comparison Chart

[0118] Platform Before Optimization After Optimization Android 200ms 80ms iOS 180ms 70ms Windows 100ms 50ms Linux 110ms 60ms

[0119] The comparison of the power consumptions on the mobile phone before and after optimization is shown in Table 4.

[0120] Table 4. Power Consumption Comparison Chart

[0121] Platform Before Optimization After Optimization Android 20% 15% iOS 18% 12%

[0122] Example 2:

[0123] On the basis of the foregoing embodiment, a user uploads a video of a target anchor on the mobile phone. The video is preprocessed offline to extract the static data of the target face and the mouth movement benchmark. The real-time voice extracts the MFCC feature sequence through the audio processing module and sends it to the lightweight LSTM model to generate 20-dimensional mouth movement parameters; the rendering module uses the preprocessed data and the predicted parameters to perform deformation and texture fusion only on the mouth area to generate new frames. Finally, after dynamic scheduling and cache optimization, the system stably reaches 30 FPS on mid- to high-end Android mobile phones, and the response latency is controlled within 80 ms, providing a smooth and natural live broadcast experience.

[0124] Example 3:

[0125] Based on the foregoing embodiments, the core module is compiled into a WebAssembly version and deployed on the computer web page. A certain user inputs voice through the web page microphone. The audio processing module extracts features from the user's voice in real time, generates lip movement parameters, and after aligning the timestamps through the synchronization module, transmits them to the rendering module to realize the real-time generation of the digital human mouth shape animation. This system can ensure audio-video synchronization and smooth interaction on Windows, Linux, iOS, and Android devices, and the response latency is less than 100 ms.

[0126] Embodiment 4:

[0127] Based on the foregoing embodiments, in view of the problem that the traditional deformation interpolation base mesh library has a rigid deformation transition in fast speech scenarios, a joint interpolation module is added. By dynamically analyzing the spectral energy distribution of continuous audio frames and the change rate of lip movement parameters, the interpolation weight of the deformation field is adaptively adjusted, so that the mouth shape change not only conforms to the current phoneme characteristics but also maintains the motion smoothness between adjacent frames. By introducing a spatio-temporal consistency constraint term, the interpolation process is optimized in both the frequency domain and the time domain to avoid lip shape jitter or lag caused by a fixed interpolation step size.

[0128] The dynamic interpolation weight calculation includes:

[0129] Define the spatio-temporal energy factor , whose value is jointly determined by the current frame audio Mel spectral energy and the change rate of lip movement parameters of adjacent frames :

[0130]

[0131] In the formula, is the mixed weight of spectral energy and motion rate; is the current frame audio Mel spectral energy; is the change rate of lip movement parameters of adjacent frames; and are normalization factors;

[0132] The spatio-temporal joint interpolation includes:

[0133] Based on select the optimal interpolation template from the pre-stored base mesh library:

[0134]

[0135] In the formula, is the spatio-temporal joint interpolation output; is the high-frequency deformation base mesh, suitable for frames with sudden energy increase; is the index of discrete time steps, used to ensure temporal continuity and avoid frame jumps; It is a low-frequency deformation base grid, suitable for smooth transition frames.

[0136] Verification shows that when obtaining an average error similar to that of the above embodiment, the lip jitter rate is reduced by 52%, from 3.2 frames per second to 1.5 frames per second; the matching accuracy of high-frequency phoneme mouth shapes is improved by 37%, and the error is reduced from 1.8 mm to 1.1 mm; the interpolation calculation amount is reduced by 28%, and the time-consuming for generating the deformation field per frame is reduced from 4.2 ms to 3.0 ms. The results show that by dynamically fusing the speech spectrum energy and the change rate of lip movement parameters, constructing a spatio-temporal joint interpolation weight function, and optimizing the generation process of the deformation field, this module can significantly suppress the lip jitter phenomenon in fast speech scenarios, improve the matching accuracy of high-frequency phoneme mouth shapes, make the mouth deformation transition more natural, reduce the calculation redundancy of the deformation field interpolation at the same time, avoid mutations or lags caused by fixed interpolation step sizes, reduce the dependence of the deformation field generation process on hardware resources, and meet the real-time requirements of low-computing-power devices.

[0137] Embodiment 5:

[0138] On the basis of the foregoing embodiment, in view of the problem that local rendering may cause inconsistent skin colors between the mouth and the face, a rendering optimization module is added. A lightweight discriminator network is introduced in the Alpha fusion stage to dynamically evaluate the lighting consistency, texture continuity, and motion naturalness between the generated mouth region and the reference face image, and the rendering parameters are adjusted in real time through gradient backpropagation to ensure seamless fusion of the synthesized region.

[0139] The discriminator network adopts a multi-branch convolution structure to extract the following features respectively:

[0140] The lighting branch, with 3 layers of convolution to analyze the brightness distribution in the HSV space;

[0141] The texture branch, with local binary pattern feature matching;

[0142] The motion branch, with optical flow field estimation for the motion smoothness of consecutive frames.

[0143] Define the fusion quality score in the adversarial loss function for optimizing the rendering parameters :

[0144]

[0145] In the formula, is the mean square error of the brightness between the generated region and the reference image; is the LBP texture similarity; is the difference in the motion amplitude of the optical flow between adjacent frames; , , , with asymmetric allocation of weight coefficients;

[0146] According to the Alpha blending coefficient of the backpropagation optimization rendering module and the deformation field interpolation weight:

[0147]

[0148] wherein is the parameter adjustment output; is the learning rate; is the Alpha blending transparency parameter.

[0149] Suppose 10 test videos containing rapid speech changes are selected, with a resolution of 1080p and 30fps;

[0150] The discriminator network's illumination branch is 3 layers of Conv+ReLU, the texture branch is LBP feature extraction, and the motion branch is PWC-Net lightweight optical flow;

[0151] The training parameters include the Adam optimizer, with a learning rate of 0.01, a batch size of 8, and 50 training epochs.

[0152] The effect of the rendering optimization module is shown in Table 5.

[0153] Table 5. Summary of the effects of the rendering optimization module

[0154] Evaluation Metrics Traditional Alpha Blending Adversarial Rendering Improvement Rate PSNR (dB) 32.1 40.3 +8.2dB Proportion of Skin Tone Inconsistent Areas (%) 15.4 3.7 -76% Number of Rendering Iterations per Frame 5 3 -40% Optical Flow Motion Error (pixels) 2.8 1.2 -57% Single Frame Processing Time (ms) 12.5 9.1 -27%

[0155] According to the experimental table, the PSNR value calculated based on the luminance component of the generated mouth region and the real human face is increased by 8.2dB, indicating that the luminance consistency between the synthesized region and the reference image is significantly enhanced; in the Lab color space, the pixel proportion of the color difference is reduced from 15.4% to 3.7%, proving that adversarial rendering effectively eliminates color deviation; the average displacement deviation of the optical flow vector in the mouth region between consecutive frames is reduced from 2.8 pixels to 1.2 pixels, indicating a significant improvement in motion smoothness. The results show that this module dynamically evaluates the fusion quality of the generated region through a multi-modal discriminator network and reversely optimizes the rendering parameters, effectively solving the problems of skin color deviation and texture discontinuity in local rendering. The illumination, texture of the mouth region are seamlessly fused with the reference human face, eliminating the artificial synthesis trace. The adaptive adjustment of the rendering parameters reduces redundant optimization steps, shortens the single-frame processing time, significantly enhances the smoothness of the mouth movement between consecutive frames, and avoids frame jumps caused by local rendering.

[0156] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable non-transitory storage media containing computer-usable program code.

[0157] The present invention can provide computer program instructions to the management platform of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing devices to generate a machine, so that the instructions executed through the management platform of the computer or other programmable data processing devices generate a device for implementing the system.

[0158] These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions of the system.

[0159] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, so that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable devices provide steps for implementing the functions of the system.

Claims

1. A real-time digital human generation system based on low-computation voice drive, characterized in that: include: An audio processing module configured to receive voice input in real time and extract audio feature vectors; A driving and rendering module, used for mapping the audio feature vector into parameters representing mouth movement, generating a dynamic mouth image based on preprocessed static face reference data and the parameters, and fusing the dynamic mouth image with the reference image; A synchronization control module, used to ensure the synchronization of audio features and rendered video frames based on a timestamp alignment mechanism and a PID feedback control algorithm; Dynamic scheduling module, which is used to monitor hardware resource load in real time and realize dynamic allocation of computing resources through multi-threaded parallelism and task priority adjustment; The joint interpolation module is used to dynamically analyze the spectral energy distribution and lip movement parameter change rate of continuous audio frames, adaptively adjust the deformation field interpolation weight, and define the spatiotemporal energy factor , whose value is determined by the Mel spectrum energy of the current frame audio The rate of change of lip movement parameters in adjacent frames Joint decision: In the formula, is the mixed weight of spectrum energy and motion rate; The Mel spectrum energy of the audio in the current frame; is the rate of change of lip movement parameters in adjacent frames; and is the normalization factor; based on Select the optimal interpolation template from the pre-stored base grid library: In the formula, It is the joint interpolation output of time and space; is the high-frequency deformation base grid; is the index of the discrete time step; is the low-frequency deformation base grid.

2. The system according to claim 1, characterized in that: The audio processing module reuses the Mel spectrum features of the overlapping areas of adjacent frames, adopts a lightweight algorithm based on spectrum subtraction to dynamically estimate the background noise and generate a mask matrix; Among them, audio feature extraction adopts a two-layer LSTM network.

3. The system according to claim 1, characterized in that: The driving and rendering module pre-stores multiple groups of mouth deformation base grids, interpolates the base grids through the parameters to generate a target mouth shape deformation field, and uses a lightweight convolutional network to perform pixel-level repair on the gaps and edges of the deformed image, and the output layer uses a Sigmoid activation function.

4. The system according to claim 1 or 3, characterized in that: The driving and rendering module includes a regional rendering unit, which divides the mouth area into multiple sub-blocks, triggers redrawing only for the sub-blocks whose deformation amplitude exceeds a threshold, and uses differential coding compression for the background area, triggering full update for multiple frames.

5. The system according to claim 1, characterized in that: The dynamic scheduling module implements dynamic allocation of computing resources based on a load-aware strategy, including: When the CPU utilization exceeds the threshold, the execution frequency of non-critical tasks is automatically reduced and GPU hardware acceleration is enabled; When a continuous multi-frame rendering timeout is detected, the output frame rate is dynamically reduced; Computing resources are scheduled in priority order.

6. The system according to claim 1, characterized in that: When the input signal is lost, the system predicts subsequent lip movement parameters based on Kalman filtering and smoothly transitions to the default closed mouth shape; when the GPU video memory overflows, it automatically switches to CPU soft rendering mode and reduces the resolution; In the network jitter scenario, a forward-looking buffer pool is used to dynamically adjust the cache depth, discard expired frames and insert motion compensated frames.

7. The system according to claim 1, characterized in that: It also includes a rendering optimization module for dynamically evaluating the illumination consistency, texture continuity, and motion naturalness of the generated mouth area and the reference face image, and adjusting the rendering parameters in real time through gradient backpropagation, defining the fusion quality score in the adversarial loss function. , used to optimize rendering parameters : In the formula, is the mean square error of brightness between the generated area and the reference image; is the LBP texture similarity; is the difference in the optical flow motion amplitude of adjacent frames; , , ; according to Backpropagation optimizes the alpha blending coefficient of the rendering module Interpolation weights with deformation field: In the formula, Adjust output for parameters; is the learning rate; Alpha blending transparency parameter.

8. A real-time digital human generation method based on low-computation voice drive, characterized in that: The method is implemented based on the system according to any one of claims 1 to 7: The method comprises: Convert speech signals into lip movement parameters through a lightweight network; Based on the pre-stored static face reference data, the target mouth shape is generated through the dynamic deformation interpolation base grid library, and combined with Alpha blending technology to achieve seamless fusion with the reference image; The interpolation weight is calculated according to the speech spectrum energy and the rate of change of the lip movement parameters, and the base network template is selected according to the interpolation weight function to generate a smooth deformation field; The audio and video synchronization error is adjusted in real time through the PID algorithm, and task priorities are dynamically allocated according to CPU utilization.

9. A computer-readable storage medium having a computer program stored therein, characterized in that: The computer program executes the method of claim 8 when executed by a processor.

Citation Information

Patent Citations

  • Generator training method of digital human generation model and digital human generation method and device

    CN117456062A

  • Audio-driven digital human generation system and method based on user prompt

    CN118968579A