Real-time digital human generation system and method based on low-calculation-amount voice driving
By decoupling end-to-end generation into speech features to low-dimensional lip-motion parameter mapping and local mouth rendering, combined with LSTM prediction and PID feedback mechanism, the problems of the contradiction between computing efficiency and real-time, imbalance in resource occupation and generalization capabilities, and insufficient audio and video synchronization accuracy in the existing technology are solved, and low-latency response and high-precision synchronization on the mobile terminal are achieved, adapting to the real-time interaction requirements of multi-platforms.
Patent Information
- Application Number
- CN202510432794.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-08
AI Technical Summary
The existing voice-driven digital life generation technology has problems such as the contradiction between computing efficiency and real-time, imbalance in resource occupation and generalization capabilities, and insufficient audio and video synchronization accuracy, making it difficult to achieve low-latency response and high-precision synchronization on the mobile terminal.
By decoupling end-to-end generation into speech features to low-dimensional lip movement parameter mapping and local mouth rendering, combined with LSTM prediction and PID feedback mechanism, efficient speech-to-lip shape conversion and precise synchronization are achieved. It adopts cross-platform heterogeneous hardware acceleration strategy, integrates NPU/CUDA/embedded binary networks, adapts to multi-end computing power, and reduces deployment complexity through a training-free general generation framework.
It achieves the real-time bottleneck of traditional solutions in terms of real-time mobile, resource occupation and long-term operation stability while maintaining high fidelity, and provides a systematic solution that is both efficient and universal.
Smart Images

Figure CN119991893A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of digital human interaction technology, and in particular to a real-time digital human generation system and method based on low-computation voice drive. Background Art
[0002] As a frontier field integrating artificial intelligence and computer vision, digital human interaction technology aims to generate virtual images with high-fidelity expressions and lip synchronization in real time through voice or text input. Traditional technical solutions are mainly divided into two categories: one is a rule-driven animation system that drives the 3D model through a predefined phoneme mouth shape mapping table. Its lip movement generation relies on a limited number of manual animation sequences, resulting in stiff expressions and difficulty in covering the subtle changes of continuous speech; the other is an end-to-end generative model based on deep learning, which uses a large-scale data set to train a neural network to directly synthesize face videos. Although it significantly improves the naturalness, the model has a large number of parameters and the inference delay is as high as hundreds of milliseconds, which is difficult to meet the real-time interaction needs of mobile terminals.
[0003] In addition, although the existing lightweight solutions reduce the amount of calculation by compressing the model, they still face two major bottlenecks: 1. The accuracy loss after model pruning and quantization is significant, especially in fast speech scenarios, lip jitter or lip lag is prone to occur; 2. The rendering process is not optimized for local areas, and the full-resolution image is still generated pixel by pixel, resulting in excessive GPU memory usage and difficulty in maintaining a stable frame rate on low-end devices. Therefore, based on the above problems, the present invention proposes a real-time digital human generation system and method based on low-computation voice drive. Summary of the invention
[0004] Technical Purpose In order to solve the above problems, the purpose of the present invention is to provide a real-time digital human generation system and method based on low-computational voice drive, aiming to solve the core problems existing in the existing voice-driven digital human generation technology, such as the contradiction between computational efficiency and real-time performance, the imbalance between resource occupation and generalization capability, and the insufficient audio and video synchronization accuracy. It gets rid of the traditional solution's reliance on high-complexity models, enables mobile devices to achieve low-latency response, avoids lip jitter and loss of accuracy, and solves the problem that the fixed delay synchronization mechanism cannot dynamically compensate for hardware fluctuations or network jitter.
[0005] Technical Solution In order to achieve the above objectives, the present invention provides a real-time digital human generation system and method based on low-computation voice drive, which decouples the traditional end-to-end generation into voice feature to low-dimensional lip movement parameter mapping and local mouth shape rendering based on pre-stored static face reference data; pre-store multi-scale mouth deformation templates, and generate target mouth shapes through linear interpolation of LSTM prediction parameters; introduce PID feedback mechanism to calibrate audio and video timestamp deviations in real time; integrate NPU on the mobile end to accelerate LSTM reasoning, use CUDA parallel interpolation on the desktop end, and deploy binary convolutional networks on embedded devices to maximize resource efficiency; extract static face features through a single video sample, and generate personalized digital humans in combination with dynamic lip movement parameters. While maintaining high fidelity, this solution greatly compresses the model size and adapts to the real-time interaction requirements of multiple platforms.
[0006] In a first aspect, the present invention provides a real-time digital human generation system based on low-computation voice-driven, comprising: An audio processing module configured to receive speech input in real time and extract a low-dimensional audio feature vector; A driving and rendering module, comprising a lip movement parameter prediction unit and a region rendering unit based on a lightweight deep learning model; the lip movement parameter prediction unit maps an audio feature vector into a low-dimensional parameter representing mouth movement, and the region rendering unit generates a dynamic mouth image based on preprocessed static face reference data and the low-dimensional parameter, and fuses it with the reference image through Alpha blending technology; A synchronization control module is used to ensure millisecond synchronization of audio features and rendered video frames based on the timestamp alignment mechanism and PID feedback control algorithm; Dynamic scheduling module, which is used to monitor hardware resource load in real time and realize dynamic allocation of computing resources through multi-threaded parallelism and task priority adjustment; Among them, the system's single-frame inference computing power does not exceed 50MFlops, and the end-to-end response delay is less than 100 milliseconds.
[0007] Furthermore, the audio processing module divides the 16 kHz sampled speech into frames with a frame length of 25 milliseconds and a frame shift of 10 milliseconds, reuses the Mel spectrum features of the overlapping areas of adjacent frames, and uses a lightweight algorithm based on spectrum subtraction to dynamically estimate the background noise and generate a mask matrix; Among them, the audio feature extraction adopts a two-layer LSTM network, and the parameter size is less than 300KB after 8-bit quantization.
[0008] Furthermore, the preprocessing process of the static face reference data includes: Use the teacher network to extract high-dimensional facial features and train a lightweight student network through knowledge distillation; The feature vectors of the 128×128 pixel image of the mouth region and the global 512×512 pixel reference image are concatenated to enhance local details and global consistency.
[0009] Furthermore, the driving and rendering module pre-stores multiple sets of mouth deformation base grids, interpolates the base grids through the low-dimensional parameters to generate a target mouth shape deformation field, and uses a lightweight convolutional network to perform pixel-level repair on the gaps and edges of the deformed image. The repair network includes a 3-layer 3×3 convolution kernel structure, the output layer uses a Sigmoid activation function, and calls the GPU skeletal animation pipeline or Metal / OpenGL ES graphics API to achieve real-time rendering of the deformed grid.
[0010] Furthermore, the regional rendering unit divides the mouth area into 8×8 sub-blocks, and triggers redrawing only for the sub-blocks whose deformation amplitude exceeds the threshold. The background area is compressed using differential coding and fully updated every 10 frames to reduce the amount of pixel filling.
[0011] Furthermore, the PID feedback control algorithm of the synchronous control module satisfies:
[0012] In the formula, is a control signal used to fine-tune the video frame generation interval; , is the audio and video timestamp error; is the independent variable in the continuous time domain; ; ; .
[0013] Furthermore, the dynamic scheduling module implements dynamic allocation of computing resources based on a load-aware strategy, including: When the CPU utilization rate exceeds 70%, the execution frequency of non-critical tasks is automatically reduced and GPU hardware acceleration is enabled; When 5 consecutive frames of rendering timeout are detected, the output frame rate will be dynamically reduced from 30 FPS to 25 FPS; The priority order of computing resource scheduling is: audio feature extraction, lip movement parameter prediction, and background texture cache update.
[0014] Furthermore, when the input signal is lost, the system predicts subsequent lip movement parameters based on Kalman filtering and smoothly transitions to the default closed mouth shape; when the GPU video memory overflows, it automatically switches to CPU soft rendering mode and reduces the resolution to 640×480 pixels; in the network jitter scenario, the forward-looking buffer pool is used to dynamically adjust the cache depth, discard expired frames and insert motion compensation frames.
[0015] Furthermore, the system integrates the ARM NEON instruction set on the mobile side to optimize matrix operations, and calls the NPU to accelerate LSTM reasoning; uses CUDA parallel computing on the desktop side to implement deformation field interpolation, and supports heterogeneous scheduling of GPUs from multiple vendors through OpenCL; and deploys a binary convolutional network on embedded devices to compress the amount of rendering model parameters to less than 1 MB.
[0016] Furthermore, it also includes a joint interpolation module for adaptively adjusting the deformation field interpolation weight by dynamically analyzing the spectrum energy distribution and lip movement parameter change rate of continuous audio frames, and defining the spatiotemporal energy factor , whose value is determined by the Mel spectrum energy of the current frame audio The rate of change of lip movement parameters in adjacent frames Joint decision:
[0017] In the formula, is the mixed weight of spectral energy and motion rate; The Mel spectrum energy of the audio in the current frame; is the rate of change of lip movement parameters in adjacent frames; and is the normalization factor; based on Select the optimal interpolation template from the pre-stored base grid library:
[0018] In the formula, It is the joint interpolation output of time and space; is the high-frequency deformation base grid; It is the index of discrete time step, which is used to ensure the continuity of timing and avoid frame jumps; is the low-frequency deformation base grid.
[0019] By integrating the speech spectrum energy and the dynamic change characteristics of lip movement parameters to construct a spatiotemporal joint interpolation weight function, the generation process of the deformation field is optimized. This mechanism significantly improves the smoothness of the mouth shape transition in fast speech scenes, eliminates the lip shape jitter or lag caused by the fixed interpolation step size, and reduces the redundant overhead of deformation field calculation, enhancing the system's adaptability to complex speech input.
[0020] Furthermore, it also includes a rendering optimization module, which is used to dynamically evaluate the illumination consistency, texture continuity and motion naturalness of the generated mouth area and the reference face image, and adjust the rendering parameters in real time through gradient backpropagation, and define the fusion quality score in the adversarial loss function. , used to optimize rendering parameters :
[0021] In the formula, is the mean square error of brightness between the generated area and the reference image; is the LBP texture similarity; is the difference in the optical flow motion amplitude of adjacent frames; , , , weight coefficients are asymmetrically distributed; according to Backpropagation optimizes the alpha blending coefficient of the rendering module Interpolation weights with deformation field:
[0022] In the formula, Adjust output for parameters; is the learning rate; Alpha blending transparency parameter.
[0023] By introducing a multimodal discriminator network, the illumination consistency, texture continuity and naturalness of the generated area are jointly evaluated, and the rendering parameters are reversely optimized, which effectively solves the skin color deviation and edge artifact problems that may be caused by local rendering. It improves the seamless fusion quality of the mouth area and the reference face image, reduces the number of iterative optimizations, and achieves the coordinated optimization of high-fidelity rendering and computational efficiency.
[0024] In a second aspect, the present invention further provides a method for generating a real-time digital human based on low-computation voice-driven, the method being based on the system described in the first aspect, comprising: The speech signal is converted into lip movement parameters through a lightweight network to describe the dynamic characteristics of the mouth opening, closing, displacement, etc. Based on the pre-stored static face reference data, the target mouth shape is generated through the dynamic deformation interpolation base grid library, and combined with Alpha blending technology to achieve seamless fusion with the reference image; The interpolation weight is calculated according to the speech spectrum energy and the rate of change of the lip movement parameters, and the base network template is selected according to the interpolation weight function to generate a smooth deformation field; The audio and video synchronization error is adjusted in real time through the PID algorithm, and task priorities are dynamically allocated according to CPU / GPU utilization.
[0025] In the third aspect, the present invention also provides a computer device, including a management platform and a memory, wherein the management platform is connected to the memory, the memory is used to store computer programs, and the management platform is used to execute the computer programs stored in the memory, so that the computer device executes the aforementioned real-time digital human generation method based on low-computational voice-driven.
[0026] In a fourth aspect, the present invention further provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a management platform, the method for generating a real-time digital human based on low-computational voice-driven is implemented.
[0027] The present invention decouples speech feature mapping and local lip rendering through a two-stage lightweight inference pipeline, combines a dynamic deformation interpolation base grid library with a PID adaptive synchronization control mechanism, and achieves efficient conversion and precise synchronization of speech to lip shape; adopts a cross-platform heterogeneous hardware acceleration strategy, integrates NPU / CUDA / embedded binary network adaptation to multi-terminal computing power, and reduces deployment complexity through a training-free general generation framework. Its exponential improvement in computing efficiency and extreme compression of end-to-end latency, while ensuring high-fidelity lip synchronization and visual consistency, breaks through the bottlenecks of traditional solutions in terms of real-time performance, resource occupation, and long-term operation stability on mobile terminals, and provides a systematic solution with both efficiency and universality for lightweight digital human interaction scenarios.
[0028] Beneficial Effects By implementing the above-mentioned low-computation voice-driven real-time digital human generation system and method provided by the present invention, the following technical effects are achieved: (1) This invention breaks through the high computational bottleneck of traditional full-image generation by decoupling end-to-end generation into low-dimensional lip movement parameter prediction and local area progressive rendering. It reduces the computational complexity to one-tenth of that of traditional solutions while retaining personalized facial features, providing a scalable architectural foundation for real-time interaction on low-computing power devices.
[0029] (2) Based on the PID feedback mechanism, the audio and video timestamp deviation is dynamically calibrated to overcome the cumulative error defect of the fixed delay synchronization strategy. This enables the system to maintain millisecond-level audio and video synchronization accuracy in scenarios with fluctuating hardware resources or unstable input, ensuring the consistency of user experience during long-term operation.
[0030] (3) By integrating the speech spectrum energy and the dynamic change characteristics of lip movement parameters to construct a spatiotemporal joint interpolation weight function, the generation process of the deformation field is optimized. This mechanism significantly improves the smoothness of the mouth shape transition in fast speech scenes, eliminates the lip shape jitter or lag caused by the fixed interpolation step size, and reduces the redundant overhead of deformation field calculation, enhancing the system's adaptability to complex speech input.
[0031] (4) By introducing a multimodal discriminator network, the illumination consistency, texture continuity, and naturalness of the generated area are jointly evaluated, and the rendering parameters are reversely optimized, which effectively solves the skin color deviation and edge artifact problems that may be caused by local rendering. It improves the seamless fusion quality of the mouth area and the reference face image, reduces the number of iterative optimizations, and achieves the coordinated optimization of high-fidelity rendering and computational efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to make the above-mentioned real-time digital human generation system and method based on low-computational voice-driven of the present invention more obvious and easy to understand, the following is a brief introduction to the drawings required for use in the specific implementation methods of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without paying creative labor.
[0033] Figure 1 The architecture diagram of the real-time digital human generation system is shown; Figure 2 A schematic diagram showing the system data flow; Figure 3 Schematic diagram of the optimization strategy. DETAILED DESCRIPTION
[0034] Embodiment 1: A real-time digital human generation system and method based on low-computation voice drive is provided, and the system architecture is as follows Figure 1 As shown, it includes: an audio processing module configured to receive voice input in real time and extract low-dimensional audio feature vectors; a driving and rendering module, including a lip movement parameter prediction unit and a regional rendering unit based on a lightweight deep learning model; the lip movement parameter prediction unit maps the audio feature vector to a low-dimensional parameter representing the mouth movement, and the regional rendering unit generates a dynamic mouth image based on the pre-processed static face reference data and the low-dimensional parameter, and fuses it with the reference image through Alpha blending technology; a synchronization control module, which is used to ensure millisecond synchronization of audio features and rendered video frames according to the timestamp alignment mechanism and PID feedback control algorithm; a dynamic scheduling module, which is used to monitor the hardware resource load in real time, and realize dynamic allocation of computing resources through multi-threaded parallelism and task priority adjustment; wherein, the system single-frame inference calculation amount does not exceed 50MFlops, and the end-to-end response delay is less than 100 milliseconds. The details are as follows.
[0035] System data flow Figure 2 As shown, starting from speech input, after audio preprocessing and feature extraction, it is sent to the sequence model to generate low-dimensional mouth movement parameters, and combined with offline preprocessing to obtain reference face data, and finally generate video frames.
[0036] In the offline stage, the system collects video samples of the target person in advance to extract the static face and key features of the target person, including: Using a mature facial detection algorithm, such as one based on MediaPipe or Dlib, 68 or more key points are detected from the video frame. The mouth area is extracted with emphasis, for example, key points 48 to 68 are selected as the core data for subsequent lip movement generation. The cropping area including the mouth is calculated based on the detected key points. The minimum bounding rectangle based on the mouth key points is usually used to expand the boundary, for example, by 10% on each side to ensure that the mouth outline is completely included. Subsequently, the cropping area is scaled to a fixed size, for example, 128×128 pixels, which is determined according to subsequent rendering requirements. A pre-trained lightweight convolutional neural network is used to extract a low-dimensional feature vector of the target face, such as 256 dimensions, as reference information for the subsequent generation stage, so that only the mouth area can be dynamically updated while maintaining the individual characteristics of the face.
[0037] After preprocessing is completed, the reference image of the target person, the normalized mouth area and its corresponding feature vector are stored in the local or embedded memory, and the overall data volume is controlled within 3MB; resources are loaded in batches through memory mapping and streaming loading technology to ensure extremely low startup delay and quick call of necessary data during runtime.
[0038] The online stage is the core process of real-time speech input driving the generation of digital human lip movements. The audio processing module divides the 16 kHz sampled speech into frames with a frame length of 25 milliseconds and a frame shift of 10 milliseconds, and uses the MFCC algorithm to extract 13-dimensional basic features, supplemented by energy, first-order and second-order differences to form a total of 39-dimensional feature vectors. A lightweight LSTM or CNN model is used for real-time processing to ensure low processing latency, and the Mel spectrum features of the overlapping areas of adjacent frames are reused. A lightweight algorithm based on spectral subtraction is used to dynamically estimate the background noise and generate a mask matrix. A lightweight temporal neural network is used to map audio features into parameter vectors representing the state of mouth movement. The network adopts a two-layer LSTM structure with 64 hidden units in each layer. The input is the audio features of several consecutive frames. The output is a 20-dimensional lip movement parameter vector that describes the changes in mouth shape. After pruning and quantization optimization, the network parameters are controlled within hundreds of thousands, making the computational complexity of each frame inference about 39 MFlops.
[0039] The driving and rendering module pre-stores multiple groups of mouth deformation base grids, interpolates the base grids through the low-dimensional parameters to generate the target mouth deformation field, and uses a lightweight convolutional network to perform pixel-level repair on the gaps and edges of the deformed image. The repair network contains a 3-layer 3×3 convolution kernel structure, the activation function is ReLU, and the output layer uses a Sigmoid activation function; the input includes the target face reference image, the preprocessed mouth area and 20-dimensional lip movement parameters; the output is a new mouth image, which is fused with the mouth area of the reference image using Alpha blending technology, and the edges are smoothed to eliminate obvious transitions. The entire image generation process is completed at a resolution of 128×128 pixels, with a delay of less than 30ms, ensuring that the system reaches a frame rate of 30fps or higher. The GPU skeletal animation pipeline or Metal / OpenGL ES graphics API is called to achieve real-time rendering of the deformed grid.
[0040] The regional rendering unit divides the mouth area into 8×8 sub-blocks, and triggers redrawing only for the sub-blocks whose deformation amplitude exceeds the threshold. The background area is compressed by differential coding and fully updated every 10 frames to reduce the pixel filling amount.
[0041] The details of the low computational rendering method of the driver and rendering module are: Only the area around the mouth is frequently redrawn, and the background and other unchanged areas use the previous frame cache; Prepare multiple mouth shape meshes in advance and quickly generate mouth shape animations through interpolation; In the 3D digital human scene, the facial blendshape parameters are bound, and the GPU performs interpolation and rendering, while the CPU only needs to provide driving parameters.
[0042] The system ensures the consistency of audio and video through a synchronization control mechanism. and video frames All with timestamp and ; In theory, the video frame playback time satisfies:
[0043] In the formula, The theoretical playback timestamp of the video frame, indicating when the frame should be displayed; is the sequence number of the video frame; is the video frame rate, that is, the number of video frames rendered per second; The time difference is calculated as:
[0044] In the formula, The audio and video synchronization error indicates the deviation between the actual playback time of the video frame and the timestamp of the corresponding audio frame. The actual playback timestamp of the video frame.
[0045] For example , you need to adjust the next frame delay:
[0046] In the formula, Generates a delay for the adjusted next frame, which is used to dynamically control the rendering rhythm; The generation delay of the previous frame, that is, the actual processing time of the current frame; is the regulating factor; is the reference time interval.
[0047] The PID feedback control algorithm of the synchronous control module satisfies:
[0048] In the formula, is a control signal used to fine-tune the video frame generation interval; , is the audio and video timestamp error; is the independent variable in the continuous time domain; ; ; .
[0049] The synchronization control module adopts a unified timestamp and buffer control, attaches a timestamp when the audio frame is generated and performs precise alignment in the rendering stage, and introduces an adaptive synchronization adjustment algorithm. If the lip movement lags behind or ahead, the system fine-tunes the rhythm of subsequent frames, such as slightly speeding up the rendering or discarding expired frames, so that the audio and video can be synchronized even in long-term operation. Compared with fixed delay buffering, it is more flexible and can better adapt to differences in equipment and network environment.
[0050] The optimization strategy adopted by the system is as follows Figure 3 As shown in the figure, the three major strategies of this system in real-time optimization are shown: dynamic scheduling, cache optimization, and model pruning and quantization. These strategies work together to help the system achieve the goal of reducing latency, computing load, and power consumption.
[0051] The dynamic scheduling module monitors the device CPU, GPU load and memory utilization in real time, and dynamically adjusts the task priority and time slice allocation of each module according to the load. It uses multi-threading or asynchronous tasks to process audio, expression generation and rendering respectively without blocking each other. When the CPU load exceeds the preset threshold, the frequency of non-critical tasks is reduced to ensure that voice processing and lip movement rendering run first. The scheduling formula is:
[0052] In the formula, For scheduling output; Delayed by the previous frame; is the regulating factor; is the current CPU usage; is the preset threshold; is the load threshold.
[0053] The dynamic scheduling module ensures that each module runs in real time and in parallel efficiently through a multi-thread scheduling strategy, and the multi-thread scheduling strategy includes: Audio acquisition and feature extraction run continuously in separate threads; Lip movement prediction and image generation are processed in another thread, and audio frame data is buffered by queues to achieve asynchronous parallelism; Dynamically adjust the priority and execution time of each module according to the current device resources to ensure smooth operation of the system even when hardware resources are limited.
[0054] The system stores all model parameters, texture materials, and intermediate cache data, supports efficient on-demand loading, and loads deep learning models, reference images, and animation materials in batches, using memory mapping and streaming loading to reduce startup latency.
[0055] The system reuses continuous frame data through caching technology to reduce repeated calculations, specifically including: the audio module caches the overlapping parts of continuous frames; the rendering module caches the background and static areas and only updates the dynamically changing areas; when the characteristics of continuous audio frames change slightly, the generation results of the previous frame are reused and interpolation is used for smooth transition.
[0056] The system uses pruning and quantization techniques to reduce model parameters and computational complexity, including: pruning the expression generation network, removing low-contribution weights, and pruning 30% to 50%; converting weights from 32-bit floating-point numbers to fixed-point integers through 8-bit quantization technology; using knowledge distillation to train lightweight models, and providing multi-level model versions to adapt to different hardware environments. After optimization by pruning and quantization technology, the amount of computation per frame of model inference is reduced to about 39 MFlops.
[0057] To reduce model computing consumption, the system uses large models to guide lightweight model training to ensure that output accuracy is not significantly reduced, and dynamically selects high-precision or streamlined models based on device performance to ensure smooth operation on low-end devices.
[0058] To ensure the stable operation of the system in various environments, the system is designed with a multi-level exception handling strategy, including: if microphone interruption or abnormal data is detected, maintain the current mouth shape or switch to the default static expression, and continuously monitor signal recovery; when abnormal device load is detected or frequency reduction due to high temperature, automatically reduce rendering details or frame rate, such as from 30FPS to 25 FPS, and suspend non-critical tasks; use a jitter buffer mechanism to sort out of order data, dynamically adjust the buffer length according to the timestamp, discard expired data, and ensure synchronization; each module is embedded with a timeout monitoring and error capture mechanism, and automatically restarts and records an error log if an abnormality occurs.
[0059] The system integrates the ARM NEON instruction set to optimize matrix operations on the mobile side, and calls the NPU to accelerate LSTM reasoning; uses CUDA parallel computing to implement deformation field interpolation on the desktop side, and supports heterogeneous scheduling of GPUs from multiple manufacturers through OpenCL; and deploys a binary convolutional network in embedded devices to compress the rendering model parameters to less than 1 MB.
[0060] The system is developed using C++ combined with a cross-platform library, and generates Android, iOS, Windows, and Linux versions through CMake, including: On the mobile side, native libraries are called through JNI and Swift / Objective-C interfaces, making full use of ARM NEON and iOS Metal for acceleration; On Windows and Linux, use multi-threading and GPU acceleration technology to ensure high frame rate output; After testing, the system can stably output 25-30FPS on all platforms, and the response delay is less than 100ms.
[0061] The comparison of system CPU / GPU occupancy before and after optimization on each platform is shown in Table 1.
[0062] Table 1. CPU / GPU occupancy rate comparison chart platform Before optimization After optimization Android 75% 45% iOS 70% 40% Windows 50% 30% Linux 55% 35% The comparison of frame rates of each platform before and after optimization is shown in Table 2.
[0063] Table 2, frame rate comparison chart platform Before optimization After optimization Android 20FPS 30FPS iOS 24FPS 30FPS Windows 30FPS 30FPS Linux 28FPS 30FPS The comparison of audio and video response delay before and after optimization on each platform is shown in Table 3.
[0064] Table 3. Response delay comparison chart platform Before optimization After optimization Android 200ms 80ms iOS 180ms 70ms Windows 100ms 50ms Linux 110ms 60ms The comparison of mobile phone power consumption before and after optimization is shown in Table 4.
[0065] Table 4. Power consumption comparison chart platform Before optimization After optimization Android 20% 15% iOS 18% 12% Embodiment 2: Based on the above-mentioned embodiment, a user uploads a video of a target anchor on the mobile phone. The video is pre-processed offline to extract the static data of the target face and the mouth movement benchmark. The real-time voice is processed by the audio processing module to extract the MFCC feature sequence and sent to the lightweight LSTM model to generate 20-dimensional mouth movement parameters. The rendering module uses the pre-processed data and predicted parameters to deform and fuse the texture of the mouth area to generate new frames. Finally, after dynamic scheduling and cache optimization, the system can stably reach 30 FPS on mid-to-high-end Android phones, and the response delay is controlled within 80ms, which can provide a smooth and natural live broadcast experience.
[0066] Embodiment 3: Based on the above embodiments, the core module is compiled into a WebAssembly version and deployed on the computer web page. A user inputs voice through the web microphone, and the audio processing module extracts features of the user's voice in real time, generates lip movement parameters, and passes them to the rendering module after aligning the timestamps through the synchronization module, so as to realize the real-time generation of digital human mouth animation. The system can ensure audio and video synchronization and smooth interaction on Windows, Linux, iOS and Android devices, with a response delay of less than 100ms.
[0067] Embodiment 4: On the basis of the above-mentioned embodiments, in order to solve the problem of abrupt deformation transition of the traditional deformation interpolation base grid library in fast speech scenes, a joint interpolation module is added. By dynamically analyzing the spectrum energy distribution and lip movement parameter change rate of continuous audio frames, the deformation field interpolation weight is adaptively adjusted, so that the mouth shape changes are consistent with the current phoneme characteristics and maintain the smoothness of the movement between adjacent frames. By introducing the spatiotemporal consistency constraint term, the interpolation process is optimized in both the frequency domain and the time domain to avoid lip shape jitter or lag caused by the fixed interpolation step size.
[0068] Dynamic interpolation weight calculation includes: Defining the space-time energy factor , whose value is determined by the Mel spectrum energy of the current frame audio The rate of change of lip movement parameters in adjacent frames Joint decision:
[0069] In the formula, is the mixed weight of spectral energy and motion rate; The Mel spectrum energy of the audio in the current frame; is the rate of change of lip movement parameters in adjacent frames; and is the normalization factor; Joint space-time interpolation includes: based on Select the optimal interpolation template from the pre-stored base grid library:
[0070] In the formula, It is the joint interpolation output of time and space; It is a high-frequency deformable base grid, suitable for energy surge frames; It is the index of discrete time step, which is used to ensure the continuity of timing and avoid frame jumps; It is a low-frequency deformation base mesh, suitable for smooth transition frames.
[0071] Verification shows that while obtaining an average error similar to that of the above embodiment, the lip jitter rate is reduced by 52%, from 3.2 frames / second to 1.5 frames / second; the high-frequency phoneme lip shape matching accuracy is improved by 37%, and the error is reduced from 1.8mm to 1.1mm; the interpolation calculation amount is reduced by 28%, and the time consumption for generating a single-frame deformation field is reduced from 4.2ms to 3.0ms. The results show that this module can significantly suppress the lip jitter phenomenon in fast speech scenarios by dynamically fusing the speech spectrum energy and the lip movement parameter change rate, constructing a spatiotemporal joint interpolation weight function, and optimizing the deformation field generation process, thereby improving the lip shape matching accuracy of high-frequency phonemes, making the mouth deformation transition more natural, and reducing the calculation redundancy of deformation field interpolation, avoiding mutations or lags caused by fixed interpolation steps, and reducing the dependence of the deformation field generation process on hardware resources, so as to meet the real-time requirements of low-computing power devices.
[0072] Embodiment 5: On the basis of the above-mentioned embodiments, in order to address the problem of inconsistent skin color between the mouth and the face due to local rendering, a rendering optimization module is added. A lightweight discriminator network is introduced in the Alpha fusion stage to dynamically evaluate the lighting consistency, texture continuity and naturalness of the generated mouth area and the reference face image, and adjust the rendering parameters in real time through gradient backpropagation to ensure seamless fusion of the synthesized area.
[0073] The discriminator network adopts a multi-branch convolution structure to extract the following features: Illumination branch: 3-layer convolution to analyze HSV spatial brightness distribution; Texture branch, local binary pattern feature matching; In the motion branch, the optical flow field estimates the smoothness of motion in consecutive frames.
[0074] Defining the fusion quality score in the adversarial loss function , used to optimize rendering parameters :
[0075] In the formula, is the mean square error of brightness between the generated area and the reference image; is the LBP texture similarity; is the difference in the optical flow motion amplitude of adjacent frames; , , , weight coefficients are asymmetrically distributed; according to Backpropagation optimizes the alpha blending coefficient of the rendering module Interpolation weights with deformation field:
[0076] In the formula, Adjust output for parameters; is the learning rate; Alpha blending transparency parameter.
[0077] Assume that 10 test videos with fast speech changes are selected, with a resolution of 1080p and 30fps; The illumination branch of the discriminator network is 3-layer Conv+ReLU, the texture branch is LBP feature extraction, and the motion branch is PWC-Net lightweight optical flow; The training parameters include Adam optimizer, learning rate 0.01, batch size 8, and training epochs 50.
[0078] The effect of the rendering optimization module is shown in Table 5.
[0079] Table 5. Summary of rendering optimization module effects Evaluation Metrics Traditional Alpha Blending Adversarial Rendering Improvement PSNR (dB) 32.1 40.3 +8.2dB Proportion of areas with inconsistent skin color (%) 15.4 3.7 -76% Rendering Iterations / Frame 5 3 -40% Optical flow motion error (pixels) 2.8 1.2 -57% Single frame processing time (ms) 12.5 9.1 -27% According to the experimental table, the PSNR value calculated based on the brightness component of the generated mouth area and the real face increased by 8.2dB, indicating that the brightness consistency between the synthetic area and the reference image is significantly enhanced; in the Lab color space, the color difference between the generated mouth and the reference face skin color points is statistically The pixel ratio of the mouth area is reduced from 15.4% to 3.7%, proving that adversarial rendering effectively eliminates color deviation; the average displacement deviation of the optical flow vector of the mouth area between consecutive frames is reduced from 2.8 pixels to 1.2 pixels, indicating that the motion smoothness is greatly improved. The results show that this module dynamically evaluates the fusion quality of the generated area through a multimodal discriminator network and reversely optimizes the rendering parameters. It effectively solves the skin color deviation and texture discontinuity problems of local rendering, and the illumination and texture of the mouth area are seamlessly integrated with the reference face, eliminating traces of artificial synthesis. The adaptive adjustment of rendering parameters reduces redundant optimization steps and shortens the processing time of a single frame. The smoothness of the mouth movement between consecutive frames is significantly enhanced, avoiding the picture jump caused by local rendering.
[0080] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable non-transient storage media containing computer-usable program code.
[0081] The present invention can provide computer program instructions to a management platform of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the management platform of the computer or other programmable data processing device produce a device for implementing the system.
[0082] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions of the system.
[0083] These computer program instructions may also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide steps for implementing the functions of the system.
Claims
1. A real-time digital human generation system based on low-computation voice drive, characterized in that: include: An audio processing module configured to receive voice input in real time and extract audio feature vectors; A driving and rendering module, used for mapping the audio feature vector into parameters representing mouth movement, generating a dynamic mouth image based on preprocessed static face reference data and the parameters, and fusing the dynamic mouth image with the reference image; A synchronization control module, used to ensure the synchronization of audio features and rendered video frames based on a timestamp alignment mechanism and a PID feedback control algorithm; Dynamic scheduling module, which is used to monitor hardware resource load in real time and realize dynamic allocation of computing resources through multi-threaded parallelism and task priority adjustment; The joint interpolation module is used to dynamically analyze the spectral energy distribution and lip movement parameter change rate of continuous audio frames, adaptively adjust the deformation field interpolation weight, and define the spatiotemporal energy factor , whose value is determined by the Mel spectrum energy of the current frame audio The rate of change of lip movement parameters in adjacent frames Joint decision: In the formula, is the mixed weight of spectrum energy and motion rate; The Mel spectrum energy of the audio in the current frame; is the rate of change of lip movement parameters in adjacent frames; and is the normalization factor; based on Select the optimal interpolation template from the pre-stored base grid library: In the formula, It is the joint interpolation output of time and space; is the high-frequency deformation base grid; is the index of the discrete time step; is the low-frequency deformation base grid.
2. The system according to claim 1, characterized in that: The audio processing module reuses the Mel spectrum features of the overlapping areas of adjacent frames, adopts a lightweight algorithm based on spectrum subtraction to dynamically estimate the background noise and generate a mask matrix; Among them, audio feature extraction adopts a two-layer LSTM network.
3. The system according to claim 1, characterized in that: The driving and rendering module pre-stores multiple groups of mouth deformation base grids, interpolates the base grids through the parameters to generate a target mouth shape deformation field, and uses a lightweight convolutional network to perform pixel-level repair on the gaps and edges of the deformed image, and the output layer uses a Sigmoid activation function.
4. The system according to claim 1 or 3, characterized in that: The driving and rendering module includes a regional rendering unit, which divides the mouth area into multiple sub-blocks, triggers redrawing only for the sub-blocks whose deformation amplitude exceeds a threshold, and uses differential coding compression for the background area, triggering full update for multiple frames.
5. The system according to claim 1, characterized in that: The dynamic scheduling module implements dynamic allocation of computing resources based on a load-aware strategy, including: When the CPU utilization exceeds the threshold, the execution frequency of non-critical tasks is automatically reduced and GPU hardware acceleration is enabled; When a continuous multi-frame rendering timeout is detected, the output frame rate is dynamically reduced; Computing resources are scheduled in priority order.
6. The system according to claim 1, characterized in that: When the input signal is lost, the system predicts subsequent lip movement parameters based on Kalman filtering and smoothly transitions to the default closed mouth shape; when the GPU video memory overflows, it automatically switches to CPU soft rendering mode and reduces the resolution; In the network jitter scenario, a forward-looking buffer pool is used to dynamically adjust the cache depth, discard expired frames and insert motion compensated frames.
7. The system according to any one of claims 1 to 6, characterized in that: It also includes a rendering optimization module for dynamically evaluating the illumination consistency, texture continuity, and motion naturalness of the generated mouth area and the reference face image, and adjusting the rendering parameters in real time through gradient backpropagation, defining the fusion quality score in the adversarial loss function. , used to optimize rendering parameters : In the formula, is the mean square error of brightness between the generated area and the reference image; is the LBP texture similarity; is the difference in the optical flow motion amplitude of adjacent frames; , , ; according to Backpropagation optimizes the alpha blending coefficient of the rendering module Interpolation weights with deformation field: In the formula, Adjust output for parameters; is the learning rate; Alpha blending transparency parameter.
8. A real-time digital human generation method based on low-computation voice drive, characterized in that: The method is implemented based on the system according to any one of claims 1 to 7: The method comprises: Convert speech signals into lip movement parameters through a lightweight network; Based on the pre-stored static face reference data, the target mouth shape is generated through the dynamic deformation interpolation base grid library, and combined with Alpha blending technology to achieve seamless fusion with the reference image; The interpolation weight is calculated according to the speech spectrum energy and the rate of change of the lip movement parameters, and the base network template is selected according to the interpolation weight function to generate a smooth deformation field; The audio and video synchronization error is adjusted in real time through the PID algorithm, and task priorities are dynamically allocated according to CPU utilization.
9. A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, characterized in that: The computer program executes the method of claim 8 when executed.
Citation Information
Patent Citations
Speech synthesis system based on face grid
CN116825083A
Voice-driven face generation model construction method and target person speaking video generation method
CN117237521A
Generator training method of digital human generation model and digital human generation method and device
CN117456062A
Voiceprint recognition method and device based on face lip movement voice separation
CN117877482A
Audio-driven digital human generation system and method based on user prompt
CN118968579A
Cited By
Personalized digital human generation method based on single video
CN120298557A
Digital human rendering method and device, storage medium and program product
CN120766706A
Multi-terminal and multi-scene adaptive digital image generation method and device
CN121120886A
Digital human expression driving method and device and electronic equipment
CN121545196A
Digital human display and interaction system based on multiple clients
CN121711503A