Cross-platform real-time digital human rendering system and method without being supported by GPU (Graphics Processing Unit)

By building a low-dimensional representation and lightweight mapping architecture, using quantitative compression dual-layer LSTM network and alpha hybrid technology, the problems of strong computing power dependence, poor cross-platform adaptability and inexpensive deployment in digital human rendering technology are solved, and efficient real-time digital human rendering is achieved in an environment without GPU support.

CN119941959AActive Publication Date: 2025-05-06LIANGSHENG DIGITAL ARTIFICIAL INTELLIGENCE (SHENZHEN) CO LTD

Patent Information

Application Number
CN202510432797.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-05-06
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

The existing digital human rendering technology has problems such as strong computing power dependence, poor cross-platform adaptability and inefficient personalized deployment, especially in an environment without GPU support, which is difficult to achieve efficient real-time rendering.

Method used

By constructing a low-dimensional representation and lightweight mapping architecture, the mouth area characteristics of the target person are extracted, and the low-dimensional orthogonal basis space model is constructed. The speech features are mapped into dynamic PCA coefficients using a quantitative compression dual-layer LSTM network, and the mouth image reconstruction and reference face fusion are achieved by combining alpha mixing and gradient masking technology.

Benefits of technology

Efficiently generate real-time voice-synced digital human animations on general computing devices without GPU support, reducing model size and training costs, improving the robustness and scalability of multi-scene applications, and achieving stable cross-platform operation and high frame rate rendering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941959A_ABST
    Figure CN119941959A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-platform GPU support-free real-time digital human rendering system and method, and relates to the technical field of digital human rendering, and the system comprises an offline modeling module which constructs a mouth base space model through principal component analysis based on video data of a target person; the real-time driving module is used for mapping real-time voice features into dynamic coefficients of the mouth base space model; the image synthesis module is used for reconstructing a mouth image according to the dynamic coefficient and fusing the mouth image with a pre-stored reference face image; the cross-platform adaptation layer is configured to realize real-time rendering without GPU dependence through compiling, memory mapping and multi-thread scheduling technologies; and the dynamic resource manager is used for optimizing resource occupation and ensuring stable operation in a multi-platform environment. According to the technical scheme, dependence on high-computing-power hardware and an ideal acoustic environment can be broken through, and a complete theoretical framework and an engineering implementation path are provided for large-scale popularization of a digital human technology in a lightweight terminal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of digital human rendering, and in particular to a cross-platform real-time digital human rendering system and method without GPU support. Background Art

[0002] Digital human technology generates realistic portraits through computer graphics and artificial intelligence algorithms, and realizes real-time expression and lip-syncing, showing broad application prospects in the fields of virtual anchors, smart assistants, and film and television production. The current mainstream technologies mainly include the following three categories: 2D photo-level animation based on static image deformation technology to simulate dynamic expressions. This method is limited by the limitations of pixel-by-pixel operations and is difficult to capture the complex movements of the lips during pronunciation, resulting in insufficient dynamic coherence and poor visual realism; 2.5D voice-driven lip-syncing that uses a small amount of video data to learn dynamic features and combines voice input to drive lip shape changes. Although this method strikes a balance between computational efficiency and realism, its deformation model based on fixed rules is difficult to adapt to individual mouth shape differences and has limited ability to model the transition state of continuous phonemes; 3D hyper-realistic modeling technology that generates high-fidelity digital humans through fine 3D models and motion capture. This method relies on professional modeling processes and high-end graphics hardware, making it difficult to deploy directly on mobile or web platforms, and real-time rendering requires extremely high GPU computing power.

[0003] The core bottlenecks of the existing technology are concentrated in the following aspects: the voice-driven model based on deep learning needs to rely on GPU accelerated reasoning, the frame rate drops significantly in the ordinary CPU environment, and the 3D solution cannot be adapted to lightweight terminals due to the complex rendering pipeline; the deep learning model has a large number of parameters, resulting in high loading delays on mobile and Web terminals, and the memory usage is difficult to meet the requirements of low-power devices; discrete lip frame switching or key point driving solutions cannot generate continuous intermediate states, resulting in stiff animations and limited voice-lip synchronization accuracy. Therefore, based on the above difficulties, the present invention proposes a cross-platform real-time digital human rendering system and method that does not require GPU support. Summary of the invention

[0004] Technical Purpose In order to solve the above problems, the purpose of the present invention is to provide a cross-platform real-time digital human rendering system and method that does not require GPU support, aiming to solve the problems of strong computing power dependence, poor cross-platform adaptability and low efficiency of personalized deployment in the existing digital human rendering technology. By constructing a low-dimensional representation and lightweight mapping architecture, it is possible to efficiently generate real-time voice-synchronized digital human animation on general-purpose computing devices without GPU support, while reducing the model size and training costs, and improving the robustness and scalability of multi-scenario applications.

[0005] Technical Solution In order to achieve the above-mentioned purpose, the present invention provides a cross-platform real-time digital human rendering system and method without GPU support. The system and method extract the mouth area features of the target person through video data, construct a low-dimensional orthogonal basis space model to compress the mouth shape data dimension; use a quantized and compressed double-layer LSTM network to map the speech features into dynamic PCA coefficients, and combine alpha blending and gradient masking technology to achieve mouth image reconstruction and reference face fusion; optimize cross-platform deployment through WebAssembly compilation, memory mapping and multi-threaded scheduling, and introduce dynamic degradation strategy and exception handling mechanism to ensure stable output of high-synchronization and high-continuity digital human animation in a resource-constrained environment.

[0006] In a first aspect, the present invention provides a cross-platform real-time digital human rendering system that does not require GPU support, comprising: The offline modeling module builds a low-dimensional mouth base space model through principal component analysis based on the video data of the target person; A real-time driving module, used for mapping real-time speech features into dynamic coefficients of the mouth basis space model through a lightweight mapping network; An image synthesis module, used for reconstructing a mouth image according to the dynamic coefficients, and fusing it with a pre-stored reference face image through an alpha blending technique; A cross-platform adaptation layer, configured to achieve GPU-free real-time rendering on browsers, mobile devices, and desktops through WebAssembly compilation, memory mapping, and multi-threaded scheduling technology; Dynamic resource manager, including on-demand loading strategy, shared memory pool and exception handling mechanism, is used to optimize resource usage and ensure stable operation in multi-platform environment.

[0007] Furthermore, the video data covers complete viseme changes.

[0008] Furthermore, MediaPipe FaceMesh is used to detect the key points of the face in each frame of the video data, crop and align the mouth area to a fixed size, such as 15×30 pixels, and grayscale, center-align and size-standardize the cropped image to eliminate illumination and position deviations. If it is three-dimensional data, the lip mesh vertices are extracted and the coordinate range is unified.

[0009] Furthermore, the normalized mouth image is flattened into a one-dimensional vector, such as 450 dimensions, and a sample matrix is ​​constructed in chronological order. The covariance matrix of the centralized sample matrix is ​​calculated, and the first k principal component vectors are obtained by eigenvalue decomposition. k is determined by the cumulative variance ≥ 95%, forming a low-dimensional mouth basis space model. The mouth key points are mirrored and the pixel mean of the left and right regions is calculated to eliminate redundant information and improve the efficiency of principal component representation.

[0010] Furthermore, a neutral closed mouth frame is selected as a reference face image, and its texture and key point coordinates are stored.

[0011] Furthermore, by collecting audio clips in real time and extracting Mel spectrum or MFCC features, a continuous speech feature sequence is formed, and the dynamic PCA coefficients are output by inputting the audio features into a lightweight mapping network.

[0012] Furthermore, the lightweight mapping network is a two-layer LSTM or Transformer architecture, and its training loss function constrains the deviation between the predicted coefficient and the true PCA coefficient through the mean square error term, and constrains the animation continuity through the difference of adjacent frame coefficients. The expression is: Mean squared error loss:

[0013] In the formula, is the mean square error; is the number of principal components; is the number of frames; is the prediction coefficient of the current frame, representing the projection component of the mouth shape of the current frame in the basis space; is the true coefficient of the current frame, which serves as a supervisory signal; Time series smoothing loss:

[0014] In the formula, is the time series smoothing term; is the prediction coefficient of the previous frame, reflecting the mouth shape state of the previous frame; Provides a benchmark for timing changes for the true coefficients of the previous frame; The total loss is:

[0015] In the formula, The weight of the time series smoothing term, the default value is 0.1.

[0016] Furthermore, the mouth image vector is generated by the following formula:

[0017] In the formula, is the reconstructed mouth shape vector; is the average mouth shape vector; is the i-th principal component vector.

[0018] Furthermore, the alpha blending process of the image synthesis module adopts a dynamically generated gradient mask, the generation of the mask adjusts the transition strength of the mask edge based on the deformation field of the facial key points, and eliminates the color difference and brightness inconsistency of the fusion area by matching the local brightness histogram of the reference facial image.

[0019] Furthermore, the gradient mask , the inner part of the lip is 1, and the edge transitions to 0. The fusion formula is:

[0020] In the formula, is the output image pixel value; is the pixel value of the mouth image; is the pixel value of the benchmark face.

[0021] Furthermore, the reference face image supports dynamic replacement with a composite image containing expressions, and achieves coordinated rendering of expressions and mouth shapes through local deformation driven by key points.

[0022] Furthermore, the implementation of the cross-platform adaptation layer includes: The browser loads the core algorithm module through WebAssembly and uses WebWorker multi-threaded asynchronous execution of reasoning and rendering. The rendering is adapted to WebGL and Canvas 2D dual modes, and compatibility is ensured through automatic downgrade; The mobile end converts the speech-to-lip mapping network into TensorFlow Lite or Core ML format, uses memory mapping technology to accelerate model loading, loads the model accuracy according to device performance, and uses a shared memory pool to reuse high-frequency resources; The desktop quickly loads model files through mmap system calls, assigns background threads to handle computationally intensive tasks, and caches recently used mouth shape coefficients and fusion results to speed up responses to similar voice inputs.

[0023] Furthermore, the exception handling and degradation mechanism of the dynamic resource manager includes: When an exception occurs in the graphics interface, the WebGL support status is detected and the mode is automatically switched to Canvas 2D mode. When a driver error occurs, the mode is returned to the pre-recorded lip sequence. Monitor CPU usage and frame processing delay in real time, and reduce rendering resolution or switch to a simplified model when the threshold is exceeded; The noise reduction pre-processing module is enabled when the audio input is abnormal, and the transition mouth shape is interpolated to maintain the continuity of the animation when the microphone permission is abnormal; Dynamically release low-frequency cache data and limit the number of concurrent digital human instances to avoid memory overflow.

[0024] Furthermore, it also includes a weight adjustment module for dynamically adjusting the spectral feature weight according to the audio band signal-to-noise ratio, wherein the weight generation function is defined as:

[0025] In the formula, Output of nonlinear weight function; Used to control the steepness of the weight curve; is the frequency band signal-to-noise ratio; is the signal-to-noise ratio offset threshold.

[0026] Through the dynamic signal-to-noise ratio weighting mechanism, the speech feature contribution of the high signal-to-noise ratio frequency band is adaptively enhanced, and the influence of the noise interference frequency band is suppressed. This mechanism significantly improves the lip synchronization accuracy of the system in complex acoustic environments, while avoiding the excessive loss of high-frequency speech details by traditional noise reduction algorithms. Without increasing the computing power burden, it expands the applicability of the system in non-ideal scenarios such as noise and reverberation, and enhances robustness and environmental adaptability.

[0027] Furthermore, a spatiotemporal constraint module is included, which is used to suppress the inter-frame jitter of the PCA coefficient through second-order difference constraints and dynamic smoothing weights. The dynamic smoothing weights adjust the smoothing strength based on the rate of change of speech features, and the expression is:

[0028] In the formula, For dynamically generated smoothing weights; is the benchmark smoothing weight; is the attenuation coefficient; is the first-order derivative modulus of the speech feature.

[0029] By introducing dynamic smoothing constraints based on the rate of speech change, the inter-frame jitter phenomenon in the PCA coefficient reconstruction process is effectively suppressed. The algorithm balances the rapid response of speech drive and the visual coherence of animation. Through second-order difference constraints and dynamic weight adjustment, it solves the "over-smoothing" or "lag" problem caused by fixed smoothing intensity. Without affecting real-time performance, it significantly improves the naturalness and fluency of digital human animation, which is consistent with the dynamic characteristics of human pronunciation.

[0030] Furthermore, the system has a computational load of ≤40MFLOPs per frame, a total model size of ≤3MB, a cross-platform rendering frame rate of ≥25FPS, and a speech-lip synchronization error of ≤40ms in a GPU-free environment.

[0031] In a second aspect, the present invention further provides a cross-platform real-time digital human rendering method without GPU support, the method is based on the system described in the first aspect, and comprises: Based on the target person’s video data, a low-dimensional mouth base space model is constructed through principal component analysis; The audio features of the input speech are read in real time, and the dynamic coefficients of the low-dimensional mouth basis space model are predicted by a lightweight neural network; The mouth image is reconstructed in combination with the low-dimensional mouth basis space model, and alpha blended with the pre-stored reference face image to generate a real-time digital human output with synchronized speech.

[0032] In a third aspect, the present invention further provides a computer device, comprising a management platform and a memory, wherein the management platform is connected to the memory, the memory is used to store computer programs, and the management platform is used to execute the computer programs stored in the memory, so that the computer device executes the aforementioned cross-platform real-time digital human rendering method that does not require GPU support.

[0033] In a fourth aspect, the present invention further provides a computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by a management platform, the computer program implements the aforementioned cross-platform real-time digital human rendering method that does not require GPU support.

[0034] The present invention constructs a low-dimensional mouth base space model in the offline stage, compresses high-dimensional mouth shape data into 6-10 dimensional dynamic coefficients, combines a double-layer LSTM network in the online stage to realize accurate mapping of speech features to low-dimensional parameters, and adopts gradient mask alpha blending technology to complete the dynamic fusion of mouth image and reference face; optimizes cross-platform deployment through WebAssembly compilation, memory mapping and multi-thread scheduling, supplemented by dynamic degradation strategy and exception handling mechanism, significantly reduces computing power requirements and model size, realizes 25FPS smooth rendering of browser, mobile terminal and desktop terminal in a GPU-free environment, and the synchronization error is ≤40ms, and supports single video acquisition to complete personalized modeling, which solves the core problems of high computing power dependence, poor cross-platform adaptability and low deployment efficiency in the prior art, and provides an efficient technical path for the popularization of digital human applications on lightweight terminals.

[0035] Beneficial Effects By implementing the cross-platform real-time digital human rendering system and method provided by the present invention, the following technical effects are achieved: (1) The present invention maps high-dimensional mouth shape data to a low-dimensional orthogonal basis space through principal component analysis to construct a personalized mouth representation model. This method significantly reduces the computational complexity of real-time rendering, and at the same time, through single video acquisition and offline modeling, it achieves rapid personalized deployment without repeated training; the introduction of low-dimensional coefficients makes the description of dynamic changes in mouth shape more efficient, solves the contradiction between high computing power requirements and poor cross-platform adaptability in traditional solutions, and provides theoretical support for real-time rendering of lightweight terminals.

[0036] (2) A lightweight neural network with a two-layer LSTM or Transformer architecture is used, combined with quantization compression and temporal smoothing loss functions, to achieve accurate mapping of speech features to low-dimensional mouth shape coefficients. Through model volume compression and computational optimization, it can run efficiently on general-purpose computing devices without GPU support, significantly reducing resource usage and power consumption. The core advantage lies in balancing model complexity and prediction accuracy, laying the algorithmic foundation for cross-platform real-time speech driving.

[0037] (3) Through the dynamic signal-to-noise ratio weighting mechanism, the speech feature contribution of the high signal-to-noise ratio frequency band is adaptively enhanced, and the influence of the noise interference frequency band is suppressed. This mechanism significantly improves the lip synchronization accuracy of the system in complex acoustic environments, while avoiding the excessive loss of high-frequency speech details by traditional noise reduction algorithms. Without increasing the computing power burden, the system's applicability in non-ideal scenarios such as noise and reverberation is expanded, and its robustness and environmental adaptability are enhanced.

[0038] (4) By introducing a dynamic smoothing constraint based on the speech change rate, the inter-frame jitter phenomenon in the PCA coefficient reconstruction process is effectively suppressed. This algorithm balances the rapid response of speech drive and the visual coherence of animation. Through second-order difference constraints and dynamic weight adjustment, it solves the "over-smoothing" or "lag" problem caused by fixed smoothing intensity. Without affecting real-time performance, it significantly improves the naturalness and fluency of digital human animation, which is consistent with the dynamic characteristics of human pronunciation. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to make the above-mentioned cross-platform real-time digital human rendering system and method without GPU support of the present invention more obvious and easy to understand, the following is a brief introduction to the drawings required for use in the specific implementation of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0040] Figure 1 The architecture diagram of the real-time digital human rendering system that is cross-platform and does not require GPU support; Figure 2 A schematic diagram showing the PCA mouth shape modeling process; Figure 3 A schematic diagram showing real-time lip coefficient prediction and image synthesis in the online rendering stage; Figure 4 A schematic diagram showing a cross-platform deployment structure. DETAILED DESCRIPTION

[0041] Embodiment 1: Provides a cross-platform real-time digital human rendering system and method without GPU support. The system includes: an offline modeling module, which is used to construct a low-dimensional mouth base space model through principal component analysis based on the video data of the target person; a real-time driving module, which is used to map the real-time speech features into the dynamic coefficients of the mouth base space model through a lightweight mapping network; an image synthesis module, which is used to reconstruct the mouth image according to the dynamic coefficients and fuse it with the pre-stored reference face image through alpha blending technology; a cross-platform adaptation layer, which is configured to achieve real-time rendering without GPU dependence on the browser, mobile terminal and desktop terminal through WebAssembly compilation, memory mapping and multi-thread scheduling technology; a dynamic resource manager, including an on-demand loading strategy, a shared memory pool and an exception handling mechanism, which is used to optimize resource usage and ensure stable operation in a multi-platform environment. The details are as follows.

[0042] The system architecture is as follows Figure 1 As shown in the figure, the upper part is the offline preparation stage: from video acquisition, audio extraction, facial key point detection, to normalized alignment and PCA modeling of the mouth area, output of personalized mouth base space model and reference face image; the lower part is the online rendering stage: user voice input is subjected to audio feature extraction and lightweight voice mapping model to generate mouth shape parameters, which are reconstructed by PCA to generate mouth image, and then fused with the reference face image to output the final digital human frame image; the arrows in the figure indicate the data flow, and it includes a cross-platform adaptation layer and anomaly monitoring mechanism.

[0043] The offline preparation stage performs statistical modeling on a large number of human face mouth area images or three-dimensional vertex data through principal component analysis to reduce the dimension of mouth shape data and reduce the computational complexity during real-time rendering, specifically including: 1. Data collection and preprocessing Collect a video containing speech from the target person, which can be from a few seconds to one minute in length. The video should include continuous changes in the person's lip area to capture sufficient mouth shape changes; Perform facial landmark detection on each frame of the video, using MediaPipe FaceMesh or other face detection algorithms to accurately locate the mouth area.

[0044] Based on the detected lip key points, each frame is cropped and aligned, for example, the center of the lip is placed in a fixed position and the upper and lower lips occupy similar scales on the image; If it contains three-dimensional data, such as face mesh vertices, the mouth mesh coordinates are extracted based on the lip vertices instead of the image; The aligned mouth image or grid is scaled to a uniform specification, such as fixed to 15×30 pixels, or a uniform grid point coordinate range.

[0045] The purpose of the data collection and preprocessing is to ensure that subsequent steps can be operated within the same vector dimension.

[0046] 2. PAC Modeling Perform an expansion operation on the normalized mouth image or three-dimensional vertex coordinates to convert it into a one-dimensional vector form. For example, if the size of the mouth image is 15×30 pixels and grayscale is used, each frame of the mouth image can be flattened into a vector with a length of 450. ; Collect the mouth vectors of all frames in the video time sequence and form a sample matrix, where D represents the dimension of the mouth vector of each frame and N represents the total number of frames; Average lip shape vector Perform calculations , each frame vector Subtract the mean to get the centered vector , and concatenate the centered vectors into a centered data matrix .

[0047] Calculate the covariance matrix:

[0048] right Perform eigenvalue decomposition to obtain the orthogonal eigenvectors corresponding to the first k largest eigenvalues ; When selecting k, the cumulative variance ratio is usually used as the criterion. For example, selecting the first 6 principal components can explain 95% to 98% of the variance.

[0049] These feature vectors are called "feature mouth shapes" or "principal component mouth shapes". A main mode that represents the deformation of the mouth; For any mouth shape vector , its projection coefficient in the PCA subspace The calculation process is as follows: , so we can reconstruct: .

[0050] If only the first k principal components are selected, the mouth shape can be represented by a coefficient vector of length k, which greatly reduces the data dimension from hundreds or thousands of dimensions to a few dimensions.

[0051] The left-right symmetry of the face is used to impose additional symmetry constraints or average the mouth image or mesh, thereby further reducing noise and redundant dimensions. For example, after mirroring the left and right mouth corner key points, averaging the same pixel positions can reduce redundant information by half, making it easier for PCA to extract the real main deformation.

[0052] For each target person, the offline stage finally outputs the mouth base space model, that is, the average mouth shape vector and k characteristic mouth shape vectors , and the reference face image, a frame with the target person’s mouth in a neutral closed state is selected, and the reconstructed mouth area is fused into the image during subsequent rendering.

[0053] At this point, the entire mouth deformation can be described by a small number of principal component coefficients, and subsequent online rendering only needs to operate these small-dimensional coefficients, greatly reducing the amount of real-time calculations.

[0054] The PCA mouth shape modeling process is as follows Figure 2 As shown in the figure, the workflow of the offline stage is described. The left side shows, for example, video acquisition and audio extraction, the middle shows facial key point detection and mouth area capture and alignment, and the right side shows the PCA modeling process, including data vectorization, covariance matrix calculation, feature decomposition, and the final output of the average mouth shape and principal component vector, and explains the reconstruction principle.

[0055] The online rendering stage is used for real-time prediction and synthesis of lip coefficients, specifically including: 1. The user provides real-time voice input through a microphone, and the audio data is sampled in fixed time slices, such as 40ms, corresponding to 25fps; The features of each frame are extracted through the audio processing module, such as Mel spectrum, MFCC or potential features extracted by self-supervision, to form a continuous audio feature sequence, providing data for subsequent model input.

[0056] 2. Use a pre-trained lightweight sequence neural network, such as a two-layer LSTM or a small Transformer, whose input is a sequence of continuous audio feature vectors and output is a vector of PCA coefficients of the mouth shape corresponding to each frame; Through training, the model prediction value is made to match the actual mouth shape parameters obtained by offline PCA. To ensure smooth prediction, the following loss function is used: Mean squared error loss:

[0057] In the formula, is the mean square error; is the number of principal components; is the number of frames; is the prediction coefficient of the current frame, representing the projection component of the mouth shape of the current frame in the basis space; is the true coefficient of the current frame, which serves as a supervisory signal; Time series smoothing loss:

[0058] In the formula, is the time series smoothing term; is the prediction coefficient of the previous frame, reflecting the mouth shape state of the previous frame; Provides a benchmark for timing changes for the true coefficients of the previous frame; The total loss is:

[0059] In the formula, The weight of the time series smoothing term, the default value is 0.1.

[0060] After training is completed, the model can map the input audio features into mouth shape parameters in real time. There is no need to repeat training for new characters. Just load the corresponding offline PCA basis space model.

[0061] Real-time lip coefficient prediction and image synthesis in the online rendering stage Figure 3 As shown, Figure 3 The entire data flow from real-time speech input to the final output of synchronized mouth shape images is shown. First, real-time speech is collected through a microphone, and the audio is converted into feature vectors through the audio feature extraction module; then, these continuous audio features are input into a lightweight speech-to-mouth mapping model to map speech features into low-dimensional mouth shape parameters, namely PCA coefficients; then, the PCA mouth shape model obtained offline is used to generate the mouth shape parameters according to the formula The predicted coefficients are reconstructed into a mouth image; finally, the generated mouth image is alpha-blended with the pre-stored reference face image through the image synthesis module to obtain the final digital human output. The data flow between the modules in the figure clearly shows how the speech signal drives the entire process of mouth shape reconstruction.

[0062] 3. Reconstruct the mouth image of the current frame using the predicted mouth shape coefficients and the PCA model obtained offline. The specific calculation formula is:

[0063] In the formula, is the reconstructed mouth shape vector; is the average mouth shape vector; is the i-th principal component vector; The reconstructed mouth image is merged with the reference face image, and the alpha blending method is used to achieve seamless connection and construct the alpha mask. , the value inside the lips is 1, the edge gradually transitions to 0, and then each pixel is fused:

[0064] In the formula, is the output image pixel value; is the pixel value of the mouth image; is the pixel value of the benchmark face; This process is executed cyclically for each frame to achieve real-time speech-synchronized digital mouth animation output.

[0065] 4. In order to achieve fast startup and low memory usage across platforms, the system adopts a variety of optimization strategies in resource management, including: Key model files, such as PCA model parameters and neural network weights, are compressed and loaded on demand. For example, the PCA average mouth shape and feature vector data are very small and can be directly packaged in the application or accessed in the form of local cache. The neural network model weights are stored in an efficient format. When loading, an asynchronous non-blocking method is preferred to initialize the model in the background to avoid blocking the main thread rendering.

[0066] Through the delayed loading strategy, a resource is loaded only when it is needed. For example, only basic digital human static resources are loaded when the application starts, and the voice-mouth mapping model and related data are loaded / initialized when the user triggers the voice-driven function. This can significantly shorten the initial startup time and reduce unnecessary memory usage.

[0067] For resources that occupy a large amount of memory, such as mouth texture maps and teeth images, shared memory and reuse technology are used; the same buffer is reused between multiple rendering frames to avoid frequent allocation and release; for static and unchanging resources, such as background images and head models, read-only memory mapping or GPU video memory are used as much as possible on each platform to reduce copy overhead; on the mobile side, the memory optimization methods provided by the platform are used, such as Ashmem shared memory on Android or memory mapped files on iOS to load large resources; the above methods can effectively reduce the overhead of model loading and running, and achieve smooth digital human-driven rendering in an environment without GPU acceleration.

[0068] In order to achieve cross-platform applications, the system has designed loading and deployment strategies for different platforms to ensure that digital human rendering runs efficiently in various environments, including: 1. Browser platform: In the Web environment, WebAssembly technology is preferably used to load and execute the core model algorithm to achieve end-side computing for real-time rendering of digital humans. Specifically, the following key steps are included: The speech-to-lip mapping neural network, audio feature extraction module, and image fusion algorithm are pre-compiled into WebAssembly modules and <script>标签或动态加载方式引入浏览器;采用流式编译和异步初始化策略,在后台下载和编译WebAssembly字节码,以缩短等待时间;一旦WebAssembly模块就绪,JavaScript主线程即可调用其导出函数,实现语音特征处理、PCA系数计算以及嘴型渲染等核心运算。

[0069] 对于PCA模型参数及其他小型数据,直接以JSON或二进制格式嵌入WebAssembly模块,或通过XHR / Fetch接口获取数据后反序列化存储于内存中;此方式能够确保数据加载高效且不会阻塞主线程。

[0070] 充分利用浏览器的WebWorker功能,将音频特征提取和模型推理等计算密集型任务移至后台线程执行,从而减少UI线程阻塞,提高整体响应速度。此外,针对浏览器的沙箱安全限制,系统在加载前通过功能检测优先使用SharedArrayBuffer和多线程WebWorker;若浏览器不支持,则自动降级为单线程执行,确保兼容性。

[0071] 在渲染输出方面,系统首先检测目标浏览器是否支持WebGL;若支持,则创建WebGL上下文,通过快速纹理更新将处理后的嘴部图像贴图到数字人脸部模型上;若WebGL不可用或被禁用,则自动切换到Canvas 2D绘图模式,通过Canvas的putImageData接口将合成的嘴部区域图像叠加到基准人脸画布上;这种双通道渲染逻辑确保无论浏览器环境如何,都能正确加载、显示并实时更新数字人嘴型动画。

[0072] 综上所述,浏览器平台方案通过WebAssembly模块加载、数据异步加载、多线程任务调度以及灵活的渲染策略,有效降低了对硬件的要求,确保在资源受限的Web环境中也能实现实时、高效的数字人渲染。

[0073] 2、移动端平台,在移动设备上,系统对模型加载和运行进行了针对性优化,以适应移动设备计算资源有限和对应用包体积敏感的特点,具体包括:将语音-嘴型映射模型转换为适合移动平台运行的格式,如TensorFlow Lite、Core ML或ONNX格式;模型文件可以在应用安装时随包提供,或在首次运行时从服务器动态下载,以保证资源获取的灵活性和更新速度。

[0074] 内存映射,利用Android平台的JNI接口调用本地C / C++代码,或在iOS平台利用Foundation框架的NSData映射或Core ML加载API,将大型模型文件采用内存映射的方式加载,从而减少加载延时和内存拷贝开销;异步加载,在App启动时预先加载关键数据,例如PCA模型、基准人脸图像,而将计算密集型部分,如神经网络权重,延迟加载至实际使用时启动,确保启动速度更快且初始内存占用较低。

[0075] 根据设备性能自动调整加载内容;例如,高端设备可以一次性加载完整精度模型,而低端设备则加载经过降采样或减少参数量的精简版本,以确保低功耗和稳定运行;针对移动端内存紧张情况,采用按需加载策略,仅在检测到用户开始语音输入时才加载用于渲染嘴部的纹理或模板数据,若用户长时间不讲话,系统会释放相关缓存数据以节约内存资源。

[0076] 通过上述策略,移动端应用在保持完整功能的同时,将加载耗时和内存占用降到最低,确保数字人渲染在Android和iOS设备上均能以低延迟、低功耗、高稳定性的方式运行。

[0077] 3、桌面端平台,在PC或Mac等桌面环境中,由于通常具备更高的CPU性能和较充裕的内存,系统模型加载和运行实现得以更加直接高效,具体包括:直接读取本地模型文件,桌面端应用可以直接从硬盘或SSD中读取预先保存的模型参数文件,并在应用启动时进行初始化;为提高加载速度,系统采用内存映射或顺序读取的方式加载大型模型文件;在Windows平台上,通过内存映射文件技术将模型权重映射到进程地址空间,实现即时访问;在Linux和macOS平台上,利用mmap系统调用实现相似效果,充分利用操作系统文件缓存机制,加快文件读取速度。

[0078] 桌面端应用在后台线程中完成模型文件的加载和初始化,并通过信号或回调通知主线程,实时向用户展示加载进度或提示信息,从而避免主界面阻塞;同时,利用多线程或事件驱动机制,将计算密集型任务,例如数据预处理、模型推理,分配到后台并行处理,确保实时渲染的平稳运行。

[0079] 由于桌面端计算资源相对充裕,系统还提供了资源管理策略,以充分利用性能冗余;例如,系统将最近使用的嘴型系数序列和渲染帧进行缓存,以便在用户发出相似语音时快速复用之前计算的结果,而无需每次从头计算;该缓存策略采用按需清理机制,当缓存数量超过预设阈值或内存占用过大时,自动释放最久未使用的缓存数据,确保系统资源得到合理利用。

[0080] 桌面端还支持开发者模式,允许开发人员通过调试接口监控加载进度、帧率及异常处理情况;开发者模式下,用户可按需切换不同精度的模型,例如选择加载更高维度的PCA组件或更复杂的网络结构,以验证系统性能和效果,而不会影响普通用户使用的精简模型。

[0081] 总体而言,通过在桌面端采用内存映射、异步多线程、动态资源管理以及开发者调试支持,系统实现了模型在PC或Mac环境下的快速、高效加载与实时渲染,保证了系统在高性能设备上以更低延迟、更高帧率稳定运行,同时为开发人员提供了灵活的调试和测试手段。

[0082] 跨平台部署结构示意图如图4所示,系统核心包括PCA嘴型模型和语音-嘴型映射模型,实现为"跨平台渲染核心”,各平台均通过适配层调用这一核心模块。

[0083] 本方法已在不同平台环境下进行验证,具体结果如下:在Chrome、Firefox、Safari等浏览器上,通过WebAssembly模块和WebGL / Canvas渲染,数字人口型同步系统平均帧率为18~25FPS,CPU占用率在普通PC上约50%,内存占用稳定在几十MB以内,长时间运行测试无内存泄露,系统稳定性高。

[0084] 在Android和iOS设备上,通过TensorFlow Lite或Core ML加载轻量级语音映射模型,在中端设备上实现15~22FPS,高端设备上超过25FPS,语音响应延迟控制在0.2秒以内,且系统支持动态降级以适应低性能设备,能耗和温度均在安全范围内。

[0085] 在无独立显卡的办公笔记本和台式机上,系统通过文件映射和多线程机制实现模型加载和实时渲染,平均帧率约25FPS以上;在低性能双核PC上,适用降级策略仍可保证15FPS左右的基本实时性;在macOS系统上可达到近30FPS,整体兼容性和实时性均得到验证。

[0086] 为提高系统鲁棒性,本系统在各模块引入了多级异常检测和处理策略,具体包括:检测浏览器是否支持WebGL,若不支持,则自动切换到Canvas 2D模式;在图形驱动错误或上下文创建失败时,系统降级输出预录制嘴型序列或静态闭嘴状态,并在UI上提示用户。

[0087] 对CPU占用率和帧处理时间进行实时监测,当检测到连续帧处理时间超过40ms时,自动降低帧率和分辨率;降低语音映射模型的复杂度以降低计算量,并使用插值算法平滑过渡。

[0088] 如果麦克风权限拒绝或音频信号异常,则系统保留最近有效嘴型或自动切换到预设的闭嘴状态,同时提示用户;针对短暂噪声干扰,系统内置降噪预处理模块,确保后续特征提取稳定。

[0089] 内存使用情况进行动态监控,当检测到内存紧张时,自动释放不必要的缓存数据,采用延迟加载策略重新加载大文件;限制同时运行的数字人实例数量,确保系统在多任务环境下仍能稳定运行。

[0090] 所有异常处理机制均通过实际测试验证,能够在各种异常和受限环境下保证系统不崩溃、性能平稳、用户体验尽可能保持连贯。

[0091] 本系统及方法提供的实时语音驱动数字人技术可广泛应用于教育、娱乐、客服等领域并展现出独特优势,具体包括:在线教育平台上,可通过创建虚拟教师形象,其嘴型与讲解语音实时同步;相比传统视频课程,虚拟教师的嘴唇准确同步语音能增强学习者对发音口型的观察,有助于语言学习;对于听障人士,系统可将语音转化为可见的嘴型运动,辅助手语或唇语教学,提升信息获取效率。

[0092] 在电影配音和动画制作中,本系统可用于自动口型对齐;例如,将名演员的配音通过本系统驱动角色模型的嘴部,使动画角色的口型与台词完美匹配,减少手工逐帧调整的成本;在游戏中,可赋予NPC角色实时对话的能力,当有语音对白时,角色脸部模型由本系统驱动产生精确的唇动,使互动更逼真,相比纯音频对话,有表情的虚拟角色更具沉浸感;此外,在社交娱乐应用中,用户可通过上传自己的照片和录制的语音,生成专属的说话表情包或短视频,在社交平台分享,增强趣味性和个性表达。

[0093] 许多公司使用智能客服机器人,本系统可为客服的头像添加实时语音驱动的嘴型,使其在回答用户提问时"开口说话”,提供类人化的交流体验;在远程视频会议中,如果网络带宽有限导致摄像头视频模糊或关闭,本系统可根据语音在接收端合成高清的说话人脸,保证与会者依然能够看到清晰的唇动,提升交流准确性;这对于需要读取唇语或表情辅助理解的会议尤其有用。

[0094] 综上所述,本发明通过离线建立个性化嘴部基空间和在线实时语音驱动映射,结合高效的资源加载、跨平台适配和完善的异常处理机制,实现了一种在无GPU条件下也能稳定运行的实时数字人口型同步系统及方法。该技术方案不仅大幅降低了计算复杂度和资源占用,而且适用于多种终端平台,为"普及型实时数字人”的广泛应用提供了切实可行的技术途径。

[0095] 实施例2:以人物A为例,使用一段3秒长的视频构建人物A的数字人模型。在离线阶段,对该视频提取约75帧的嘴部图像,分辨率统一到15×30像素,对应每帧音频窗口为40ms;PCA分析选取6个主成分,累计保留98%以上的嘴部像素方差。在在线阶段,将任意语音输入轻量模型,所述轻量模型为双层LSTM,隐藏单元各192维,输出每帧6个系数。利用这些系数与人物A的嘴部PCA模型,可重建出128×128分辨率人脸中嘴部区域的图像,并与预存的人物A脸部其余部分合成,得到人物A的说话视频。测试表明,即使在无GPU的普通笔记本电脑CPU上,该方法也能以每秒25帧以上速度实时生成嘴型同步的视频画面。

[0096] 实施例3:在前述实施例的基础上,针对复杂声学环境下语音特征易受噪声干扰的问题,提出一种基于信噪比的动态频带权重调整机制。该机制通过实时分析音频信号中各频带的信噪比,自适应增强高信噪比频段的特征贡献,抑制低信噪比频段的噪声干扰,从而提升语音-嘴型映射模型在噪声环境下的鲁棒性与预测精度。

[0097] 将输入音频分帧后,通过短时傅里叶变换提取各频带能量,并分别计算语音段与静默段的能量均值,得到各频带信噪比:

[0098] 式中,为频带信噪比;为第f个频带的语音;为第f个频带的噪声能量;为防除零常数;通过非线性权重函数,将信噪比映射至[0,1]区间:

[0099] 式中,为非线性权重函数输出;用于控制权重曲线陡峭度,默认为2.5;为频带信噪比;为信噪比偏移阈值,默认为8dB;对Mel频谱特征按频带进行加权:

[0100] 式中,为加权输出;为原始Mel频谱特征向量;为频带权重向量。

[0101] 将加权后的特征输入至语音-嘴型映射网络,优化模型对清晰语音频段的敏感性。

[0102] 验证表明在获得与上述实施例相近平均误差的情况下,在信噪比≤10dB的噪声环境下,嘴型同步误差降低约30%;额外计算开销仅为0.8ms / 帧,实时性影响可忽略。结果表明,该机制通过动态分析音频频带信噪比,自适应调整特征权重,显著提升了系统在复杂声学环境下的鲁棒性,在噪声干扰场景中,语音特征提取的稳定性得到增强,有效抑制了低频噪声对嘴型预测的干扰,使得生成的嘴型动画与语音信号的同步精度显著提高;同时,特征加权的引入未对实时性产生可感知的影响,系统仍能在无GPU环境下保持流畅的渲染帧率,确保了跨平台部署的普适性。

[0103] 实施例4:在前述实施例的基础上,针对PCA系数重构可能导致的帧间抖动问题,提出一种时空一致性约束算法。该算法通过分析连续帧PCA系数的二阶差分与语音特征的时序变化率,动态调整平滑强度,抑制非自然突变的同时保留语音驱动的快速响应能力,从而提升动画流畅性。

[0104] 对预测的PCA系数序列,计算其加速度项:

[0105] 式中,为PCA系数的二阶差分,用于量化嘴型变化的突变程度;为当前帧的预测系数,表征当前帧嘴型在基空间中的投影分量;为前一帧的预测系数,反映前一帧的嘴型状态;为前两帧的预测系数。

[0106] 根据语音特征变化率,即MFCC一阶导数模长,调整平滑强度:

[0107] 式中,为动态生成的平滑权重,值越小表示允许更大幅度的系数变化;为基准平滑权重,默认为0.3,用于控制整体平滑强度;为衰减系数,默认为1.2,用于调节语音变化率对权重的影响灵敏度;为语音特征的一阶导数模长,表征当前帧语音变化的剧烈程度。

[0108] 在总损失中增加时空平滑项:

[0109] 式中,为时空平滑项;为PCA主成分数量。

[0110] 总损失更新为:

[0111] 式中,为时空平滑项权重,默认为0.2,用于控制其对总损失的贡献比例。

[0112] 假设连续5帧的语音特征变化率与PCA系数如表1所示,其他参数固定为:,,,基准抖动幅度为12像素。

[0113] 表1、语音特征变化率与PCA系数

[0114] 对二阶差分计算进行计算,以为例:

[0115] t为3时,

[0116] t为4时,

[0117] t为5时,

[0118] 对动态平滑权重进行计算:

[0119] t为1时,

[0120] t为2时,

[0121] t为3时,

[0122] 对时空平滑损失项进行计算:

[0123] t为3时,

[0124] t为4时,

[0125] t为5时,

[0126] 总计,。

[0127] 时空约束模块效果如表2所示。

[0128] 表2、时空约束模块效果汇总评估指标应用前应用后提升幅度帧间抖动幅度(像素)12.07.2-40%流畅度评分(MOS)3.84.4+15.8%峰值位移(t=3帧)12.06.5-45.8%CPU额外占用率-+3%可忽略根据实验表格可知,在快速发音切换时,动态平滑权重降至0.02,弱化平滑约束,允许PCA系数快速响应语音变化;而在平缓段,升高至0.20,抑制抖动;时空一致性约束使帧间过渡更自然,额外计算延迟仅1.2ms / 帧,CPU占用率增加≤3%,完全满足实时性要求。结果表明,该算法通过动态调节平滑强度与时空一致性约束,解决了因PCA系数突变导致的动画帧间抖动问题,能够在不牺牲语音驱动实时性的前提下,显著改善嘴型动画的视觉连贯性;尤其在快速发音切换场景中,系统既能快速响应语音变化,又能平滑过渡相邻帧的嘴型状态,避免了传统方案中因固定平滑强度导致的"滞后”或"过平滑”现象;此外,算法的计算开销被严格控制在极低范围内,未对系统整体资源占用产生显著负担。

[0129] 本领域内的技术人员应明白,本发明的实施例可提供为方法、系统或计算机程序产品。因此,本发明可采用完全硬件实施例、完全软件实施例、或结合软件和硬件方面的实施例的形式。而且,本发明可采用在一个或多个其中包含有计算机可用程序代码的计算机可用非瞬时性存储介质上实施的计算机程序产品的形式。

[0130] 本发明可提供计算机程序指令到通用计算机、专用计算机、嵌入式处理机或其他可编程数据处理设备的管理平台以产生一个机器,使得通过计算机或其他可编程数据处理设备的管理平台执行的指令产生用于实现所述系统的装置。

[0131] 这些计算机程序指令也可存储在能引导计算机或其他可编程数据处理设备以特定方式工作的计算机可读存储器中,使得存储在该计算机可读存储器中的指令产生包括指令装置的制造品,该指令装置实现所述系统的功能。

[0132] 这些计算机程序指令也可装载到计算机或其他可编程数据处理设备上,使得在计算机或其他可编程设备上执行一系列操作步骤以产生计算机实现的处理,从而在计算机或其他可编程设备上执行的指令提供用于实现所述系统功能的步骤。< / script>

Claims

1. A cross-platform real-time digital human rendering system that does not require GPU support, characterized in that: include: The offline modeling module builds a mouth-based space model through principal component analysis based on the target person’s video data; A real-time driving module, used for mapping the real-time speech features into dynamic coefficients of the mouth basis space model through a mapping network; An image synthesis module, used for reconstructing a mouth image according to the dynamic coefficients, and fusing it with a pre-stored reference face image through a blending technique; A cross-platform adaptation layer, configured to achieve real-time rendering without GPU dependency through compilation, memory mapping, and multi-threaded scheduling techniques; Dynamic resource manager, used to optimize resource usage and ensure stable operation in multi-platform environments; The spatiotemporal constraint module is used to suppress the inter-frame jitter of the PCA coefficients through second-order difference constraints and dynamic smoothing weights. The dynamic smoothing weights adjust the smoothing strength based on the rate of change of speech features. The expression is: , In the formula, For dynamically generated smoothing weights; is the benchmark smoothing weight; is the attenuation coefficient; is the first-order derivative modulus of the speech feature.

2. The system according to claim 1, characterized in that: The offline modeling module processes the mouth key points by mirroring, eliminates data redundancy by calculating the pixel means of the left and right regions, and adaptively determines the number of principal components according to the cumulative variance ratio threshold.

3. The system according to claim 1, characterized in that: The training loss function of the mapping network constrains the deviation between the predicted coefficient and the true PCA coefficient through the mean square error term, and constrains the animation continuity through the difference of adjacent frame coefficients. The expression is: , In the formula, is the time series smoothing term; is the number of principal components; is the number of frames; is the prediction coefficient of the current frame; is the prediction coefficient of the previous frame; is the true coefficient of the current frame; is the true coefficient of the previous frame.

4. The system according to claim 1, characterized in that: The mixing process of the image synthesis module adopts a dynamically generated gradient mask, the generation of the mask is based on the deformation field of the key points of the face to adjust the transition strength of the mask edge, and the color difference and brightness inconsistency of the fusion area are eliminated by matching the local brightness histogram of the reference face image.

5. The system according to claim 4, characterized in that: The reference face image supports dynamic replacement with a composite image containing expressions, and realizes coordinated rendering of expressions and mouth shapes through local deformation driven by key points.

6. The system according to claim 1, characterized in that: The dynamic resource manager detects the CPU occupancy rate and frame processing delay in real time, triggers a dynamic degradation strategy, and generates a mouth shape based on historical coefficient interpolation and loads a static reference image when the audio input is abnormal.

7. The system according to any one of claims 1 to 6, characterized in that: It also includes a weight adjustment module for dynamically adjusting the spectral feature weight according to the audio band signal-to-noise ratio, wherein the weight generation function is defined as: , In the formula, Output of nonlinear weight function; Used to control the steepness of the weight curve; is the frequency band signal-to-noise ratio; is the signal-to-noise ratio offset threshold.

8. The system according to claim 1, characterized in that: Add a spatiotemporal smoothing term to the total loss according to the dynamic smoothing weight: , In the formula, Output of spatiotemporal smoothing term; is the number of PCA principal components; , In the formula, is the mean square error; is the weight of the time series smoothing term; is the weight of the spatiotemporal smoothing term.

9. A cross-platform real-time digital human rendering method without GPU support, characterized in that: The method is implemented based on the system according to any one of claims 1 to 8: The method comprises: Based on the video data of the target person, the mouth base space model is constructed through principal component analysis; The audio features of the input speech are read in real time, and the dynamic coefficients of the mouth basis space model are predicted through a neural network; The mouth image is reconstructed in combination with the mouth base space model, and alpha blended with the pre-stored reference face image to generate a real-time digital human output with synchronized speech.

10. A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, characterized in that: The computer program executes the method of claim 9 when executed.

Citation Information

Patent Citations

  • Method for rendering 3D digital virtual human real-time behavior based on WebGL

    CN111882628A

  • Digital human video generation method and device, electronic equipment and storage medium

    CN116385629A

  • High-performance cross-platform three-dimensional rendering engine based on C + + development and physical materials

    CN117218256A

  • Expression-editable voice-driven face reconstruction method based on three-dimensional Gaussian sputtering technology

    CN118762133A

  • Digital human image fusion method, device and equipment and readable storage medium

    CN118799439A

Cited By

  • Real-time interactive digital human system supporting high concurrency and implementation method thereof

    CN120179081A

  • A real-time interactive digital human system supporting high concurrency and its implementation method

    CN120179081B

  • Two-dimensional code real-time digital human interaction system based on deep learning

    CN120234088A

  • Deep Learning-Based Real-Time QR Code Interactive System for Digital Humans

    CN120234088B

  • Rendering node video stream low-delay synthesis output system

    CN120935379A