A cross-platform real-time digital human rendering system and method without GPU support
By building a low-dimensional mouth-based space model and a lightweight mapping network, combined with dynamic resource management, real-time digital human rendering without GPU support is achieved, solving the problems of strong computing power dependence and poor cross-platform adaptability in the existing technology, and achieving efficient real-time voice synchronization and smooth rendering.
Patent Information
- Application Number
- CN202510432797.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-08
AI Technical Summary
The existing digital human rendering technology has significantly reduced frame rates in ordinary CPU environments, making it difficult to adapt to lightweight terminals, and has poor cross-platform adaptability, low personalized deployment efficiency, and cannot achieve efficient real-time voice synchronization.
By building a low-dimensional mouth-based space model, combining a lightweight mapping network and a dynamic resource manager, WebAssembly compilation, memory mapping and multi-threaded scheduling are adopted to realize real-time rendering without GPU support, and introduce dynamic downgrade strategies and exception handling mechanisms to ensure stable output of high synchronization and high continuity digital human animations in resource-constrained environments.
Achieve efficient real-time digital rendering in a GPU-free environment, reduces computing complexity and resource usage, supports multi-platform applications, improves robustness and environmental adaptability, and ensures smooth rendering of 25FPS and voice-mouth synchronization error of ≤40ms.
Smart Images

Figure CN119941959B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital human rendering, and particularly to a real-time digital human rendering system and method that is cross-platform and does not require GPU support. Background Art
[0002] Digital human technology generates lifelike portraits through computer graphics and artificial intelligence algorithms, and realizes real-time expression and lip-sync driving, showing broad application prospects in the fields of virtual anchors, intelligent assistants, and film and television production. The current mainstream technologies mainly include the following three categories: 2D photo-realistic animation that simulates dynamic expressions based on static image deformation technology. This method is limited by the limitations of per-pixel operations and is difficult to capture the complex movements inside the lips during pronunciation, resulting in insufficient dynamic coherence and poor visual realism; 2.5D voice-driven lip-sync that uses a small amount of video data to learn dynamic features and combines voice input to drive lip shape changes. Although this method achieves a balance between computational efficiency and realism, its deformation model based on fixed rules is difficult to adapt to individual lip shape differences, and its ability to model the transition states of continuous phonemes is limited; 3D hyper-realistic modeling technology that generates high-fidelity digital humans through fine 3D models and motion capture. This method relies on professional modeling processes and high-end graphics hardware, is difficult to directly deploy on mobile or web platforms, and has extremely high requirements for GPU computing power for real-time rendering.
[0003] The core bottlenecks of the existing technologies are concentrated in the following aspects: The speech-driven model based on deep learning requires GPU acceleration for inference, and the frame rate drops significantly in a common CPU environment. The 3D solution cannot be adapted to lightweight terminals due to the complex rendering pipeline; The deep learning model has a large number of parameters, resulting in high loading latency on mobile and Web sides, and the memory occupancy is difficult to meet the requirements of low-power devices; The discrete lip shape frame switching or key point-driven scheme cannot generate continuous intermediate states, resulting in rigid animations and limited speech-lip sync accuracy. Therefore, based on the above problems, the present invention proposes a real-time digital human rendering system and method that is cross-platform and does not require GPU support. Summary of the Invention
[0004] Technical Objectives
[0005] In order to solve the above problems, the objective of the present invention is to provide a real-time digital human rendering system and method that is cross-platform and does not require GPU support, aiming to solve the problems of strong computing power dependence, poor cross-platform adaptability, and low personalized deployment efficiency in the existing digital human rendering technology. By constructing a low-dimensional representation and lightweight mapping architecture, it is possible to efficiently generate real-time speech-synchronized digital human animations on general computing devices without GPU support, while reducing the model size and training cost, and improving the robustness and scalability of multi-scenario applications.
[0006] Technical Solutions
[0007] To achieve the above object, the present invention provides a real-time digital human rendering system and method that is cross-platform and does not require GPU support. The system and method extract the mouth region features of the target person from video data, construct a low-dimensional orthogonal basis space model to compress the dimension of the mouth shape data; use a quantized compressed double-layer LSTM network to map speech features into dynamic PCA coefficients, and combine alpha blending and gradient mask technology to realize mouth image reconstruction and fusion with the reference face; optimize cross-platform deployment through WebAssembly compilation, memory mapping and multi-threaded scheduling, and introduce a dynamic degradation strategy and an exception handling mechanism to ensure stable output of high-synchronization and high-continuity digital human animations in resource-constrained environments.
[0008] In a first aspect, the present invention provides a real-time digital human rendering system that is cross-platform and does not require GPU support, including:
[0009] An offline modeling module that constructs a low-dimensional mouth basis space model based on the video data of the target person through principal component analysis;
[0010] A real-time driving module for mapping real-time speech features into dynamic coefficients of the mouth basis space model through a lightweight mapping network;
[0011] An image synthesis module for reconstructing a mouth image according to the dynamic coefficients and fusing it with a pre-stored reference face image through alpha blending technology;
[0012] A cross-platform adaptation layer configured to achieve GPU-independent real-time rendering on browsers, mobile devices and desktops through WebAssembly compilation, memory mapping and multi-threaded scheduling technologies;
[0013] A dynamic resource manager, including an on-demand loading strategy, a shared memory pool and an exception handling mechanism, for optimizing resource occupancy and ensuring stable operation in a multi-platform environment.
[0014] Furthermore, the video data covers complete pronunciation mouth shape changes.
[0015] Furthermore, the key points of each frame of the face in the video data are detected by MediaPipe FaceMesh, the mouth region is cropped and aligned to a fixed size, such as 15×30 pixels, and the cropped image is grayscale, centered and size-normalized to eliminate illumination and position deviations. If it is three-dimensional data, the lip mesh vertices are extracted and the coordinate range is unified.
[0016] Further, the normalized mouth image is flattened into a one-dimensional vector, such as 450 dimensions, and a sample matrix is constructed in chronological order. The covariance matrix of the centralized sample matrix is calculated, and the first k principal component vectors are obtained through eigenvalue decomposition, where k is determined by the cumulative variance ≥ 95%, forming a low-dimensional mouth basis space model. By mirroring the mouth key points and calculating the pixel mean values of the left and right regions, redundant information is eliminated to improve the efficiency of the principal component representation.
[0017] Further, a neutral closed mouth shape frame is selected as the reference face image, and its texture and key point coordinates are stored.
[0018] Further, by collecting audio segments in real time and extracting Mel spectrogram or MFCC features, a continuous speech feature sequence is formed, and the audio features are input into a lightweight mapping network to output dynamic PCA coefficients.
[0019] Further, the lightweight mapping network is a two-layer LSTM or Transformer architecture, and its training loss function constrains the deviation between the predicted coefficients and the true PCA coefficients through the mean square error term, and constrains the animation continuity through the difference of the coefficients of adjacent frames. The expression is:
[0020] Mean square error loss:
[0021]
[0022] In the formula, is the mean square error; is the number of principal components; is the number of frames; is the predicted coefficient of the current frame, representing the projection component of the current frame's mouth shape in the basis space; is the true coefficient of the current frame, serving as a supervision signal;
[0023] Temporal smoothing loss:
[0024]
[0025] In the formula, is the temporal smoothing term; is the predicted coefficient of the previous frame, reflecting the mouth shape state of the previous frame; is the true coefficient of the previous frame, providing a benchmark for temporal changes;
[0026] The total loss is:
[0027]
[0028] In the formula, is the weight of the temporal smoothing term, defaulting to 0.1.
[0029] Further, a mouth image vector is generated through the following formula:
[0030]
[0031] In the formula, is the reconstructed mouth shape vector; is the average mouth shape vector; is the i-th principal component vector.
[0032] Further, the alpha blending process of the image synthesis module adopts a dynamically generated gradient mask. The generation of the mask adjusts the edge transition intensity of the mask based on the face key point deformation field, and eliminates the color difference and brightness inconsistency in the fusion area through the local brightness histogram matching of the reference face image.
[0033] Further, for the gradient mask , the value inside the lips is 1, and it transitions to 0 at the edge. The fusion formula is:
[0034]
[0035] In the formula, is the output image pixel value; is the mouth image pixel value; is the reference face pixel value.
[0036] Further, the reference face image supports dynamic replacement with a composite image containing expressions, and realizes the collaborative rendering of expressions and mouth shapes through key point-driven local deformation.
[0037] Further, the implementation methods of the cross-platform adaptation layer include:
[0038] On the browser side, the core algorithm module is loaded through WebAssembly, and WebWorker multi-threading is used to asynchronously execute inference and rendering. The rendering adapts to both WebGL and Canvas 2D modes, and ensures compatibility through automatic degradation;
[0039] On the mobile side, the voice-mouth mapping network is converted into the TensorFlow Lite or Core ML format, and the memory mapping technology is adopted to accelerate the model loading. The model accuracy is loaded according to the device performance, and the shared memory pool is used to reuse high-frequency resources;
[0040] On the desktop side, the model file is quickly loaded through the mmap system call, a background thread is allocated to handle computationally intensive tasks, and the recently used mouth shape coefficients and fusion results are cached to accelerate the response to similar voice inputs.
[0041] Further, the exception handling and degradation mechanism of the dynamic resource manager include:
[0042] Detect the WebGL support status when a graphics interface exception occurs, automatically switch to the Canvas 2D mode, and fallback to the pre-recorded lip sequence when a driver error occurs;
[0043] Monitor the CPU occupancy rate and frame processing latency in real time, and reduce the rendering resolution or switch to a simplified model when the threshold is exceeded;
[0044] Enable the noise reduction preprocessing module when the audio input is abnormal, and interpolate to generate transitional lip shapes when the microphone permission is abnormal to maintain animation continuity;
[0045] Dynamically release low-frequency cache data and limit the number of concurrent digital human instances to avoid memory overflow.
[0046] Furthermore, it also includes a weight adjustment module for dynamically adjusting the spectral feature weights according to the signal-to-noise ratio of the audio frequency band. Among them, the weight generation function is defined as:
[0047]
[0048] In the formula, is the output of the non-linear weight function; is used to control the steepness of the weight curve; is the signal-to-noise ratio of the frequency band; is the signal-to-noise ratio offset threshold.
[0049] Through the signal-to-noise ratio dynamic weighting mechanism, the contribution of speech features in high signal-to-noise ratio frequency bands is adaptively enhanced, and the influence of noise interference frequency bands is suppressed. This mechanism significantly improves the lip synchronization accuracy of the system in complex acoustic environments, while avoiding the excessive loss of high-frequency speech details by traditional noise reduction algorithms. Without increasing the computing power burden, it expands the applicability of the system in non-ideal scenarios such as noise and reverberation, and enhances the robustness and environmental adaptability.
[0050] Furthermore, it also includes a spatio-temporal constraint module for suppressing the inter-frame jitter of PCA coefficients through second-order difference constraints and dynamic smoothing weights. The dynamic smoothing weights adjust the smoothing intensity based on the speech feature change rate, and the expression is:
[0051]
[0052] In the formula, is the dynamically generated smoothing weight; is the reference smoothing weight; is the attenuation coefficient; is the magnitude of the first-order derivative of the speech feature.
[0053] By introducing dynamic smoothing constraints based on the speech change rate, the inter-frame jitter phenomenon in the PCA coefficient reconstruction process is effectively suppressed. This algorithm balances the fast response driven by speech and the visual coherence of the animation. Through second-order difference constraints and dynamic weight adjustment, it solves the "over-smoothing" or "lag" problems caused by fixed smoothing intensity. Without affecting real-time performance, it significantly improves the naturalness and fluency of the digital human animation, conforming to the dynamic characteristics of human pronunciation.
[0054] Furthermore, in an environment without a GPU, the calculation amount per frame of the system is ≤40MFLOPs, the overall model volume is ≤3MB, the cross-platform rendering frame rate is ≥25FPS, and the speech-lip synchronization error is ≤40ms.
[0055] In a second aspect, the present invention also provides a real-time digital human rendering method that is cross-platform and does not require GPU support. The method is based on the system described in the first aspect above and includes:
[0056] Based on the video data of the target person, a low-dimensional mouth base space model is constructed through principal component analysis;
[0057] The audio features of the input speech are read in real time, and the dynamic coefficients of the low-dimensional mouth base space model are predicted through a lightweight neural network;
[0058] The mouth image is reconstructed by combining the low-dimensional mouth base space model and alpha-blended with the pre-stored reference face image to generate a real-time digital human output synchronized with the speech.
[0059] In a third aspect, the present invention also provides a computer device, including a management platform and a memory. The management platform is connected to the memory. The memory is used to store a computer program, and the management platform is used to execute the computer program stored in the memory so that the computer device executes and implements the above-mentioned real-time digital human rendering method that is cross-platform and does not require GPU support.
[0060] In a fourth aspect, the present invention also provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium. When the computer program is executed by the management platform, the above-mentioned real-time digital human rendering method that is cross-platform and does not require GPU support is implemented.
[0061] Through the construction of a low-dimensional mouth base space model in the offline stage, the present invention compresses high-dimensional mouth shape data into 6-10 dimensional dynamic coefficients, combines a double-layer LSTM network in the online stage to achieve an accurate mapping from speech features to low-dimensional parameters, and uses a gradient mask alpha blending technique to complete the dynamic fusion of the mouth image and the reference face; through WebAssembly compilation, memory mapping and multi-threaded scheduling optimization for cross-platform deployment, supplemented by a dynamic degradation strategy and an exception handling mechanism, the computing power requirement and model volume are significantly reduced, and smooth rendering at 25FPS is achieved on browsers, mobile devices and desktops without a GPU environment, with a synchronization error ≤ 40ms. At the same time, it supports personalized modeling with a single video capture, solving the core problems of high computing power dependence, poor cross-platform adaptability and low deployment efficiency in the prior art, and providing an efficient technical path for the popularization of digital humans on lightweight terminals.
[0062] Beneficial effects
[0063] By implementing the real-time digital human rendering system and method provided by the present invention without GPU support across platforms, the following technical effects are achieved:
[0064] (1) The present invention maps high-dimensional mouth shape data to a low-dimensional orthogonal basis space through principal component analysis, and constructs a personalized mouth representation model. This method significantly reduces the computational complexity of real-time rendering, and at the same time, through single video capture and offline modeling, rapid personalized deployment without repeated training is achieved; the introduction of low-dimensional coefficients makes the description of mouth shape dynamic changes more efficient, solves the contradiction between high computing power requirements and poor cross-platform adaptability in traditional solutions, and provides a theoretical support for real-time rendering on lightweight terminals.
[0065] (2) A lightweight neural network using a double-layer LSTM or Transformer architecture, combined with quantization compression and a temporal smoothing loss function, realizes an accurate mapping from speech features to low-dimensional mouth shape coefficients. Through model volume compression and computing optimization, it can run efficiently on general computing devices without GPU support, significantly reducing resource occupancy and power consumption. The core advantage lies in balancing model complexity and prediction accuracy, laying an algorithmic foundation for cross-platform real-time speech driving.
[0066] (3) Through a signal-to-noise ratio dynamic weighting mechanism, the contribution of speech features in high signal-to-noise ratio frequency bands is adaptively enhanced, and the influence of noise interference frequency bands is suppressed. This mechanism significantly improves the mouth shape synchronization accuracy of the system in complex acoustic environments, and at the same time avoids excessive loss of high-frequency speech details by traditional noise reduction algorithms. Without increasing the computing power burden, the applicability of the system in non-ideal scenarios such as noise and reverberation is extended, and the robustness and environmental adaptability are enhanced.
[0067] (4) By introducing a dynamic smoothing constraint based on the speech change rate, the inter-frame jitter phenomenon in the PCA coefficient reconstruction process is effectively suppressed. This algorithm balances the fast response driven by speech and the visual coherence of the animation. By means of second-order difference constraints and dynamic weight adjustment, it solves the problems of "over-smoothing" or "lag" caused by a fixed smoothing intensity, and significantly improves the naturalness and fluency of the digital human animation without affecting real-time performance, which conforms to the dynamic characteristics of human pronunciation. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] To make the above-mentioned cross-platform real-time digital human rendering system and method without GPU support of the present invention more obvious and understandable, the drawings required for the specific implementation manners of the present invention will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings without creative efforts.
[0069] Figure 1 It represents the architecture diagram of the cross-platform real-time digital human rendering system without GPU support;
[0070] Figure 2 It represents the schematic diagram of the PCA mouth shape modeling process;
[0071] Figure 3 It represents the schematic diagram of real-time mouth shape coefficient prediction and image synthesis in the online rendering stage;
[0072] Figure 4 It represents the schematic diagram of the cross-platform deployment structure. DETAILED DESCRIPTION OF THE INVENTION
[0073] Embodiment 1:
[0074] A cross-platform real-time digital human rendering system and method without GPU support are provided.
[0075] The system includes: an offline modeling module for constructing a low-dimensional mouth base space model based on the video data of the target person through principal component analysis; a real-time driving module for mapping real-time speech features to dynamic coefficients of the mouth base space model through a lightweight mapping network; an image synthesis module for reconstructing a mouth image according to the dynamic coefficients and fusing it with a pre-stored reference face image through alpha blending technology; a cross-platform adaptation layer configured to achieve GPU-independent real-time rendering on browsers, mobile devices, and desktops through WebAssembly compilation, memory mapping, and multi-threaded scheduling technologies; a dynamic resource manager including an on-demand loading strategy, a shared memory pool, and an exception handling mechanism for optimizing resource occupancy and ensuring stable operation in a multi-platform environment. Specifically, it is described as follows.
[0076] The system architecture is asFigure 1 As shown, the upper part is the offline preparation stage: from video acquisition, audio extraction, face key point detection, to the normalization alignment and PCA modeling of the mouth region, outputting a personalized mouth base space model and a reference face image; the lower part is the online rendering stage: the user's voice input generates mouth shape parameters through audio feature extraction and a lightweight speech mapping model, generates a mouth image through PCA reconstruction, and then fuses with the reference face image to output the final digital human frame image; the arrows in the figure indicate the data flow direction and include a cross-platform adaptation layer and an anomaly monitoring mechanism.
[0077] The offline preparation stage statistically models a large number of face mouth region images or three-dimensional vertex data through principal component analysis to reduce the dimension of the mouth shape data and reduce the computational complexity during real-time rendering. Specifically, it includes:
[0078] 1. Data acquisition and preprocessing
[0079] Collect a video containing speech from the target person, with a duration ranging from several seconds to one minute. This video should include continuous changes in the person's lip region to facilitate capturing sufficient mouth shape deformations.
[0080] Perform face key point detection on each frame of the video, using MediaPipe FaceMesh or other face detection algorithms to accurately locate the mouth region.
[0081] According to the detected lip key points, crop and align each frame. For example, place the center of the lips at a fixed position and make the upper and lower lips occupy a similar scale in the image.
[0082] If three-dimensional data is included, such as face mesh vertices, extract the mouth mesh coordinates based on the lip vertices to replace the image.
[0083] Scale the aligned mouth images or meshes according to a unified specification, such as fixing them to 15×30 pixels, or unifying the range of mesh point coordinates.
[0084] The purpose of the data acquisition and preprocessing is to ensure that subsequent steps can perform operations within the same vector dimension.
[0085] 2. PAC modeling
[0086] Perform an unfolding operation on the normalized mouth images or three-dimensional vertex coordinates to convert them into a one-dimensional vector form. For example, if the size of the mouth image is 15×30 pixels and grayscale representation is used, each frame of the mouth image can be flattened into a vector with a length of 450. ; Collect the mouth vectors of all frames in chronological order of the video and form a sample matrix, where D represents the dimension of each frame of the mouth vector and N represents the total number of frames;
[0087] Calculate the average mouth shape vector and subtract this average value from each frame vector to obtain a centered vector . Then concatenate the centered vectors into a centered data matrix .
[0088] Calculate the covariance matrix:
[0089]
[0090] Perform eigenvalue decomposition on to obtain the orthogonal eigenvectors corresponding to the top k largest eigenvalues ;
[0091] When selecting k, the cumulative variance ratio is usually used as the criterion. For example, when selecting the first 6 principal components, 95% - 98% of the variance can be explained.
[0092] These eigenvectors are called "feature mouth shapes" or "principal component mouth shapes", and each vector represents a main mode of mouth deformation;
[0093] For any mouth shape vector , the calculation process of its projection coefficient in this PCA subspace is as follows: , and thus it can be reconstructed as: .
[0094] If only the first k principal components are selected, the mouth shape can be represented by a coefficient vector of length k, greatly reducing the data dimension, from hundreds or thousands of dimensions to several dimensions.
[0095] Utilize the left - right symmetry of the human face to impose additional symmetry constraints or averaging on the mouth image or mesh, thereby further reducing noise and redundant dimensions. For example, after mirroring the key points of the left and right mouth corners, taking the average of the same pixel positions can reduce half of the redundant information, making it easier for PCA to extract the true main deformations.
[0096] For each target person, the offline stage finally outputs a mouth base space model, namely the average mouth shape vector and k feature mouth shape vectors , as well as a reference face image. Select a frame when the target person's mouth is in a neutral closed state, and during subsequent rendering, fuse the reconstructed mouth region into this image.
[0097] At this time, the entire mouth deformation can be described by a small number of principal component coefficients, and subsequent online rendering only needs to operate on these small - dimension coefficients, greatly reducing the real - time calculation amount.
[0098] The PCA mouth shape modeling process is as follows Figure 2 shown in the figure, which describes the workflow in the offline stage. The left side shows, for example, video capture and audio extraction. The middle shows face key point detection, mouth region cropping and alignment. The right side shows the PCA modeling process, including data vectorization, covariance matrix calculation, eigen decomposition, and the process of finally outputting the average mouth shape and principal component vectors, and explains the reconstruction principle.
[0099] In the online rendering stage described above, it is used for real-time prediction and synthesis of mouth shape coefficients, specifically including:
[0100] 1. The user provides real-time voice input through the microphone, and the audio data is sampled at a fixed time slice, for example, 40 ms, corresponding to 25 fps;
[0101] Extract the features of each frame through the audio processing module, such as Mel spectrogram, MFCC or self-supervised extracted latent features, to form a continuous audio feature sequence, providing data for the subsequent model input.
[0102] 2. Adopt a pre-trained lightweight sequence neural network, such as a two-layer LSTM or a small Transformer, whose input is a sequence of continuous audio feature vectors, and the output is a vector of mouth shape PCA coefficients corresponding to each frame;
[0103] Through training, the predicted values of the model are made to match the true mouth shape parameters obtained by offline PCA. To ensure smooth prediction, the following loss function is adopted:
[0104] Mean squared error loss:
[0105]
[0106] In the formula, is the mean squared error; is the number of principal components; is the number of frames; is the predicted coefficient of the current frame, representing the projection component of the mouth shape of the current frame in the basis space; is the true coefficient of the current frame, serving as a supervision signal;
[0107] Temporal smoothing loss:
[0108]
[0109] In the formula, is the temporal smoothing term; is the predicted coefficient of the previous frame, reflecting the mouth shape state of the previous frame; is the true coefficient of the previous frame, providing a benchmark for temporal changes;
[0110] The total loss is:
[0111]
[0112] Wherein, is the weight of the temporal smoothing term, and the default value is 0.1.
[0113] After the training is completed, the model can map the input audio features to lip parameters in real time. There is no need to repeat the training for new characters. Just load the corresponding offline PCA basis space model.
[0114] The real-time lip coefficient prediction and image synthesis in the online rendering stage are as Figure 3 shown, Figure 3 which shows the entire data stream from the real-time voice input to the final output of the synchronized lip image. First, the real-time voice is collected through the microphone, and the audio is converted into feature vectors by the audio feature extraction module; then, these continuous audio features are input into the lightweight voice-lip mapping model to map the voice features to low-dimensional lip parameters, that is, PCA coefficients; subsequently, using the offline obtained PCA lip model, according to the formula the predicted coefficients are reconstructed into a lip image; finally, the generated lip image is alpha-blended with the pre-stored reference face image through the image synthesis module to obtain the final digital human output. The data stream between the modules in the figure clearly shows the whole process of how the voice signal drives the lip reconstruction.
[0115] 3. Reconstruct the lip image of the current frame using the predicted lip coefficients and the offline obtained PCA model. The specific calculation formula is:
[0116]
[0117] Wherein, is the reconstructed lip vector; is the average lip vector; is the i-th principal component vector;
[0118] Fuse the reconstructed lip image with the reference face image, and use the alpha-blending method to achieve seamless connection, constructing an alpha mask , with a value of 1 inside the lips, gradually transitioning to 0 at the edges, and then perform the fusion calculation for each pixel:
[0119]
[0120] Wherein, is the pixel value of the output image; is the pixel value of the lip image; is the pixel value of the reference face;
[0121] This process is executed in a loop for each frame to achieve the real-time voice-synchronized digital human lip animation output.
[0122] 4. To achieve fast startup across platforms and low memory occupancy, the system adopts various optimization strategies in resource management, specifically including:
[0123] Compress and load key model files, such as PCA model parameters, neural network weights, etc., on demand; for example, the PCA average mouth shape and eigenvector data volume is very small and can be directly packaged in the application or stored and accessed in the form of local cache, while the neural network model weights are stored in an efficient format; when loading, preferentially use the asynchronous non-blocking method to initialize the model in the background to avoid blocking the main thread rendering.
[0124] Adopt the lazy loading strategy to load a certain resource only when it is needed; for example, only load the basic digital human static resources when the application starts, and wait for the user to trigger the voice-driven function to load / initialize the voice-mouth mapping model and related data; this can significantly shorten the initial startup time and reduce unnecessary memory occupancy.
[0125] For resources with large memory occupancy, such as mouth texture maps, tooth images, etc., adopt shared memory and reuse technology; reuse the same buffer between multiple rendering frames to avoid frequent allocation and release; for static and unchanging resources, such as background images, head models, try to use read-only memory mapping or save them in the GPU video memory on each platform to reduce the copy overhead; utilize the memory optimization means provided by the platform on mobile devices, such as Ashmem shared memory on Android or memory mapping files on iOS to load large resources; through the above methods, the overhead of model loading and operation can be effectively reduced, and smooth digital human driving rendering can be achieved even in an environment without GPU acceleration.
[0126] To achieve cross-platform applications, the system designs loading and deployment strategies for different platforms to ensure efficient operation of digital human rendering in various environments, specifically including:
[0127] 1. Browser platform, preferably adopt WebAssembly technology to load and execute the core model algorithm in the Web environment to achieve end-side computing for real-time digital human rendering, specifically including the following key steps:
[0128] Pre-compile the voice-mouth mapping neural network, audio feature extraction module, and image fusion algorithm into WebAssembly modules, and through <script>标签或动态加载方式引入浏览器;采用流式编译和异步初始化策略,在后台下载和编译WebAssembly字节码,以缩短等待时间;一旦WebAssembly模块就绪,JavaScript主线程即可调用其导出函数,实现语音特征处理、PCA系数计算以及嘴型渲染等核心运算。
[0129] 对于PCA模型参数及其他小型数据,直接以JSON或二进制格式嵌入WebAssembly模块,或通过XHR / Fetch接口获取数据后反序列化存储于内存中;此方式能够确保数据加载高效且不会阻塞主线程。
[0130] 充分利用浏览器的WebWorker功能,将音频特征提取和模型推理等计算密集型任务移至后台线程执行,从而减少UI线程阻塞,提高整体响应速度。此外,针对浏览器的沙箱安全限制,系统在加载前通过功能检测优先使用SharedArrayBuffer和多线程WebWorker;若浏览器不支持,则自动降级为单线程执行,确保兼容性。
[0131] 在渲染输出方面,系统首先检测目标浏览器是否支持WebGL;若支持,则创建WebGL上下文,通过快速纹理更新将处理后的嘴部图像贴图到数字人脸部模型上;若WebGL不可用或被禁用,则自动切换到Canvas 2D绘图模式,通过Canvas的putImageData接口将合成的嘴部区域图像叠加到基准人脸画布上;这种双通道渲染逻辑确保无论浏览器环境如何,都能正确加载、显示并实时更新数字人嘴型动画。
[0132] 综上所述,浏览器平台方案通过WebAssembly模块加载、数据异步加载、多线程任务调度以及灵活的渲染策略,有效降低了对硬件的要求,确保在资源受限的Web环境中也能实现实时、高效的数字人渲染。
[0133] 2、移动端平台,在移动设备上,系统对模型加载和运行进行了针对性优化,以适应移动设备计算资源有限和对应用包体积敏感的特点,具体包括:
[0134] 将语音-嘴型映射模型转换为适合移动平台运行的格式,如TensorFlow Lite、Core ML或ONNX格式;模型文件可以在应用安装时随包提供,或在首次运行时从服务器动态下载,以保证资源获取的灵活性和更新速度。
[0135] 内存映射,利用Android平台的JNI接口调用本地C / C++代码,或在iOS平台利用Foundation框架的NSData映射或Core ML加载API,将大型模型文件采用内存映射的方式加载,从而减少加载延时和内存拷贝开销;异步加载,在App启动时预先加载关键数据,例如PCA模型、基准人脸图像,而将计算密集型部分,如神经网络权重,延迟加载至实际使用时启动,确保启动速度更快且初始内存占用较低。
[0136] 根据设备性能自动调整加载内容;例如,高端设备可以一次性加载完整精度模型,而低端设备则加载经过降采样或减少参数量的精简版本,以确保低功耗和稳定运行;针对移动端内存紧张情况,采用按需加载策略,仅在检测到用户开始语音输入时才加载用于渲染嘴部的纹理或模板数据,若用户长时间不讲话,系统会释放相关缓存数据以节约内存资源。
[0137] 通过上述策略,移动端应用在保持完整功能的同时,将加载耗时和内存占用降到最低,确保数字人渲染在Android和iOS设备上均能以低延迟、低功耗、高稳定性的方式运行。
[0138] 3、桌面端平台,在PC或Mac等桌面环境中,由于通常具备更高的CPU性能和较充裕的内存,系统模型加载和运行实现得以更加直接高效,具体包括:
[0139] 直接读取本地模型文件,桌面端应用可以直接从硬盘或SSD中读取预先保存的模型参数文件,并在应用启动时进行初始化;为提高加载速度,系统采用内存映射或顺序读取的方式加载大型模型文件;在Windows平台上,通过内存映射文件技术将模型权重映射到进程地址空间,实现即时访问;在Linux和macOS平台上,利用mmap系统调用实现相似效果,充分利用操作系统文件缓存机制,加快文件读取速度。
[0140] 桌面端应用在后台线程中完成模型文件的加载和初始化,并通过信号或回调通知主线程,实时向用户展示加载进度或提示信息,从而避免主界面阻塞;同时,利用多线程或事件驱动机制,将计算密集型任务,例如数据预处理、模型推理,分配到后台并行处理,确保实时渲染的平稳运行。
[0141] 由于桌面端计算资源相对充裕,系统还提供了资源管理策略,以充分利用性能冗余;例如,系统将最近使用的嘴型系数序列和渲染帧进行缓存,以便在用户发出相似语音时快速复用之前计算的结果,而无需每次从头计算;该缓存策略采用按需清理机制,当缓存数量超过预设阈值或内存占用过大时,自动释放最久未使用的缓存数据,确保系统资源得到合理利用。
[0142] 桌面端还支持开发者模式,允许开发人员通过调试接口监控加载进度、帧率及异常处理情况;开发者模式下,用户可按需切换不同精度的模型,例如选择加载更高维度的PCA组件或更复杂的网络结构,以验证系统性能和效果,而不会影响普通用户使用的精简模型。
[0143] 总体而言,通过在桌面端采用内存映射、异步多线程、动态资源管理以及开发者调试支持,系统实现了模型在PC或Mac环境下的快速、高效加载与实时渲染,保证了系统在高性能设备上以更低延迟、更高帧率稳定运行,同时为开发人员提供了灵活的调试和测试手段。
[0144] 跨平台部署结构示意图如图4所示,系统核心包括PCA嘴型模型和语音-嘴型映射模型,实现为"跨平台渲染核心”,各平台均通过适配层调用这一核心模块。
[0145] 本方法已在不同平台环境下进行验证,具体结果如下:
[0146] 在Chrome、Firefox、Safari等浏览器上,通过WebAssembly模块和WebGL / Canvas渲染,数字人口型同步系统平均帧率为18~25FPS,CPU占用率在普通PC上约50%,内存占用稳定在几十MB以内,长时间运行测试无内存泄露,系统稳定性高。
[0147] 在Android和iOS设备上,通过TensorFlow Lite或Core ML加载轻量级语音映射模型,在中端设备上实现15~22FPS,高端设备上超过25FPS,语音响应延迟控制在0.2秒以内,且系统支持动态降级以适应低性能设备,能耗和温度均在安全范围内。
[0148] 在无独立显卡的办公笔记本和台式机上,系统通过文件映射和多线程机制实现模型加载和实时渲染,平均帧率约25FPS以上;在低性能双核PC上,适用降级策略仍可保证15FPS左右的基本实时性;在macOS系统上可达到近30FPS,整体兼容性和实时性均得到验证。
[0149] 为提高系统鲁棒性,本系统在各模块引入了多级异常检测和处理策略,具体包括:
[0150] 检测浏览器是否支持WebGL,若不支持,则自动切换到Canvas 2D模式;在图形驱动错误或上下文创建失败时,系统降级输出预录制嘴型序列或静态闭嘴状态,并在UI上提示用户。
[0151] 对CPU占用率和帧处理时间进行实时监测,当检测到连续帧处理时间超过40ms时,自动降低帧率和分辨率;降低语音映射模型的复杂度以降低计算量,并使用插值算法平滑过渡。
[0152] 如果麦克风权限拒绝或音频信号异常,则系统保留最近有效嘴型或自动切换到预设的闭嘴状态,同时提示用户;针对短暂噪声干扰,系统内置降噪预处理模块,确保后续特征提取稳定。
[0153] 内存使用情况进行动态监控,当检测到内存紧张时,自动释放不必要的缓存数据,采用延迟加载策略重新加载大文件;限制同时运行的数字人实例数量,确保系统在多任务环境下仍能稳定运行。
[0154] 所有异常处理机制均通过实际测试验证,能够在各种异常和受限环境下保证系统不崩溃、性能平稳、用户体验尽可能保持连贯。
[0155] 本系统及方法提供的实时语音驱动数字人技术可广泛应用于教育、娱乐、客服等领域并展现出独特优势,具体包括:
[0156] 在线教育平台上,可通过创建虚拟教师形象,其嘴型与讲解语音实时同步;相比传统视频课程,虚拟教师的嘴唇准确同步语音能增强学习者对发音口型的观察,有助于语言学习;对于听障人士,系统可将语音转化为可见的嘴型运动,辅助手语或唇语教学,提升信息获取效率。
[0157] 在电影配音和动画制作中,本系统可用于自动口型对齐;例如,将名演员的配音通过本系统驱动角色模型的嘴部,使动画角色的口型与台词完美匹配,减少手工逐帧调整的成本;在游戏中,可赋予NPC角色实时对话的能力,当有语音对白时,角色脸部模型由本系统驱动产生精确的唇动,使互动更逼真,相比纯音频对话,有表情的虚拟角色更具沉浸感;此外,在社交娱乐应用中,用户可通过上传自己的照片和录制的语音,生成专属的说话表情包或短视频,在社交平台分享,增强趣味性和个性表达。
[0158] 许多公司使用智能客服机器人,本系统可为客服的头像添加实时语音驱动的嘴型,使其在回答用户提问时"开口说话”,提供类人化的交流体验;在远程视频会议中,如果网络带宽有限导致摄像头视频模糊或关闭,本系统可根据语音在接收端合成高清的说话人脸,保证与会者依然能够看到清晰的唇动,提升交流准确性;这对于需要读取唇语或表情辅助理解的会议尤其有用。
[0159] 综上所述,本发明通过离线建立个性化嘴部基空间和在线实时语音驱动映射,结合高效的资源加载、跨平台适配和完善的异常处理机制,实现了一种在无GPU条件下也能稳定运行的实时数字人口型同步系统及方法。该技术方案不仅大幅降低了计算复杂度和资源占用,而且适用于多种终端平台,为"普及型实时数字人”的广泛应用提供了切实可行的技术途径。
[0160] 实施例2:
[0161] 以人物A为例,使用一段3秒长的视频构建人物A的数字人模型。在离线阶段,对该视频提取约75帧的嘴部图像,分辨率统一到15×30像素,对应每帧音频窗口为40ms;PCA分析选取6个主成分,累计保留98%以上的嘴部像素方差。在在线阶段,将任意语音输入轻量模型,所述轻量模型为双层LSTM,隐藏单元各192维,输出每帧6个系数。利用这些系数与人物A的嘴部PCA模型,可重建出128×128分辨率人脸中嘴部区域的图像,并与预存的人物A脸部其余部分合成,得到人物A的说话视频。测试表明,即使在无GPU的普通笔记本电脑CPU上,该方法也能以每秒25帧以上速度实时生成嘴型同步的视频画面。
[0162] 实施例3:
[0163] 在前述实施例的基础上,针对复杂声学环境下语音特征易受噪声干扰的问题,提出一种基于信噪比的动态频带权重调整机制。该机制通过实时分析音频信号中各频带的信噪比,自适应增强高信噪比频段的特征贡献,抑制低信噪比频段的噪声干扰,从而提升语音-嘴型映射模型在噪声环境下的鲁棒性与预测精度。
[0164] 将输入音频分帧后,通过短时傅里叶变换提取各频带能量,并分别计算语音段与静默段的能量均值,得到各频带信噪比:
[0165]
[0166] 式中,为频带信噪比;为第f个频带的语音;为第f个频带的噪声能量;为防除零常数;
[0167] 通过非线性权重函数,将信噪比映射至[0,1]区间:
[0168]
[0169] 式中,为非线性权重函数输出;用于控制权重曲线陡峭度,默认为2.5;为频带信噪比;为信噪比偏移阈值,默认为8dB;
[0170] 对Mel频谱特征按频带进行加权:
[0171]
[0172] 式中,为加权输出;为原始Mel频谱特征向量;为频带权重向量。
[0173] 将加权后的特征输入至语音-嘴型映射网络,优化模型对清晰语音频段的敏感性。
[0174] 验证表明在获得与上述实施例相近平均误差的情况下,在信噪比≤10dB的噪声环境下,嘴型同步误差降低约30%;额外计算开销仅为0.8ms / 帧,实时性影响可忽略。结果表明,该机制通过动态分析音频频带信噪比,自适应调整特征权重,显著提升了系统在复杂声学环境下的鲁棒性,在噪声干扰场景中,语音特征提取的稳定性得到增强,有效抑制了低频噪声对嘴型预测的干扰,使得生成的嘴型动画与语音信号的同步精度显著提高;同时,特征加权的引入未对实时性产生可感知的影响,系统仍能在无GPU环境下保持流畅的渲染帧率,确保了跨平台部署的普适性。
[0175] 实施例4:
[0176] 在前述实施例的基础上,针对PCA系数重构可能导致的帧间抖动问题,提出一种时空一致性约束算法。该算法通过分析连续帧PCA系数的二阶差分与语音特征的时序变化率,动态调整平滑强度,抑制非自然突变的同时保留语音驱动的快速响应能力,从而提升动画流畅性。
[0177] 对预测的PCA系数序列,计算其加速度项:
[0178]
[0179] 式中,为PCA系数的二阶差分,用于量化嘴型变化的突变程度;为当前帧的预测系数,表征当前帧嘴型在基空间中的投影分量;为前一帧的预测系数,反映前一帧的嘴型状态;为前两帧的预测系数。
[0180] 根据语音特征变化率,即MFCC一阶导数模长,调整平滑强度:
[0181]
[0182] 式中,为动态生成的平滑权重,值越小表示允许更大幅度的系数变化;为基准平滑权重,默认为0.3,用于控制整体平滑强度;为衰减系数,默认为1.2,用于调节语音变化率对权重的影响灵敏度;为语音特征的一阶导数模长,表征当前帧语音变化的剧烈程度。
[0183] 在总损失中增加时空平滑项:
[0184]
[0185] 式中,为时空平滑项;为PCA主成分数量。
[0186] 总损失更新为:
[0187]
[0188] 式中,为时空平滑项权重,默认为0.2,用于控制其对总损失的贡献比例。
[0189] 假设连续5帧的语音特征变化率与PCA系数如表1所示,其他参数固定为:,,,基准抖动幅度为12像素。
[0190] 表1、语音特征变化率与PCA系数
[0191]
[0192] 对二阶差分计算进行计算,以为例:
[0193]
[0194] t为3时,
[0195] t为4时,
[0196] t为5时,
[0197] 对动态平滑权重进行计算:
[0198]
[0199] t为1时,
[0200] t为2时,
[0201] t为3时,
[0202] 对时空平滑损失项进行计算:
[0203]
[0204] t为3时,
[0205] t为4时,
[0206] t为5时,
[0207] 总计,。
[0208] 时空约束模块效果如表2所示。
[0209] 表2、时空约束模块效果汇总
[0210] 评估指标应用前应用后提升幅度帧间抖动幅度(像素)12.07.2-40%流畅度评分(MOS)3.84.4+15.8%峰值位移(t=3帧)12.06.5-45.8%CPU额外占用率-+3%可忽略
[0211] 根据实验表格可知,在快速发音切换时,动态平滑权重降至0.02,弱化平滑约束,允许PCA系数快速响应语音变化;而在平缓段,升高至0.20,抑制抖动;时空一致性约束使帧间过渡更自然,额外计算延迟仅1.2ms / 帧,CPU占用率增加≤3%,完全满足实时性要求。结果表明,该算法通过动态调节平滑强度与时空一致性约束,解决了因PCA系数突变导致的动画帧间抖动问题,能够在不牺牲语音驱动实时性的前提下,显著改善嘴型动画的视觉连贯性;尤其在快速发音切换场景中,系统既能快速响应语音变化,又能平滑过渡相邻帧的嘴型状态,避免了传统方案中因固定平滑强度导致的"滞后”或"过平滑”现象;此外,算法的计算开销被严格控制在极低范围内,未对系统整体资源占用产生显著负担。
[0212] 本领域内的技术人员应明白,本发明的实施例可提供为方法、系统或计算机程序产品。因此,本发明可采用完全硬件实施例、完全软件实施例、或结合软件和硬件方面的实施例的形式。而且,本发明可采用在一个或多个其中包含有计算机可用程序代码的计算机可用非瞬时性存储介质上实施的计算机程序产品的形式。
[0213] 本发明可提供计算机程序指令到通用计算机、专用计算机、嵌入式处理机或其他可编程数据处理设备的管理平台以产生一个机器,使得通过计算机或其他可编程数据处理设备的管理平台执行的指令产生用于实现所述系统的装置。
[0214] 这些计算机程序指令也可存储在能引导计算机或其他可编程数据处理设备以特定方式工作的计算机可读存储器中,使得存储在该计算机可读存储器中的指令产生包括指令装置的制造品,该指令装置实现所述系统的功能。
[0215] 这些计算机程序指令也可装载到计算机或其他可编程数据处理设备上,使得在计算机或其他可编程设备上执行一系列操作步骤以产生计算机实现的处理,从而在计算机或其他可编程设备上执行的指令提供用于实现所述系统功能的步骤。< / script>
Claims
1. A real-time digital human rendering system that is cross-platform and does not require GPU support, characterized in that, It includes: An offline modeling module that constructs a mouth base space model through principal component analysis based on the video data of the target person; A real-time driving module for mapping real-time speech features to dynamic coefficients of the mouth base space model through a mapping network; An image synthesis module for reconstructing a mouth image according to the dynamic coefficients and fusing it with a pre-stored reference face image through a blending technique; A cross-platform adaptation layer configured to achieve GPU-independent real-time rendering through compilation, memory mapping, and multi-threaded scheduling technologies; A dynamic resource manager for optimizing resource occupancy and ensuring stable operation in a multi-platform environment through an on-demand loading strategy, a shared memory pool, and an exception handling mechanism; A spatio-temporal constraint module for suppressing the inter-frame jitter of PCA coefficients through second-order difference constraints and dynamic smoothing weights, where the dynamic smoothing weights adjust the smoothing intensity based on the speech feature change rate, and the expression is: ; wherein, is a dynamically generated smoothing weight; is a reference smoothing weight; is an attenuation coefficient; is the magnitude of the first-order derivative of the speech feature; A weight adjustment module, which is used to dynamically adjust the weights of spectral features according to the signal-to-noise ratio of audio frequency bands. The weight generation function is defined as: ; where is the output of the non-linear weight function; is used to control the steepness of the weight curve; is the signal-to-noise ratio of the frequency band; is the signal-to-noise ratio offset threshold.
2. The system according to claim 1, wherein: The offline modeling module processes mouth key points through mirroring, eliminates data redundancy by calculating the pixel mean of the left and right regions, and adaptively determines the number of principal components according to the cumulative variance ratio threshold.
3. The system according to claim 1, wherein: The training loss function of the mapping network constrains the deviation between the predicted coefficients and the true PCA coefficients through a mean square error term, and constrains the animation continuity through the difference between adjacent frame coefficients. The expression is: , In the formula, is the temporal smoothing term; is the number of principal components; is the number of frames; is the prediction coefficient of the current frame; is the prediction coefficient of the previous frame; is the true coefficient of the current frame; is the true coefficient of the previous frame.
4. The system according to claim 1, wherein: The blending process of the image synthesis module uses a dynamically generated gradient mask. The generation of the mask adjusts the mask edge transition intensity based on the deformation field of the face key points, and eliminates the color difference and brightness inconsistency in the fusion area through local brightness histogram matching of the reference face image.
5. The system according to claim 4, wherein: The reference face image supports being dynamically replaced with a composite image containing expressions, and realizes the collaborative rendering of expressions and mouth shapes through key point-driven local deformation.
6. The system according to claim 1, wherein: The dynamic resource manager real-time detects the CPU occupancy rate and frame processing delay, triggers a dynamic degradation strategy, and when the audio input is abnormal, generates a mouth shape based on historical coefficient interpolation and loads a static reference image.
7. The system according to claim 1, wherein: Add a spatio-temporal smoothing term to the total loss according to the dynamic smoothing weights: , In the formula, is the output of the spatio-temporal smoothing term; is the number of PCA principal components; , In the formula, is the mean square error; is the weight of the temporal smoothing term; is the weight of the spatio-temporal smoothing term.
8. A real-time digital human rendering method that is cross-platform and does not require GPU support, wherein: The implementation of the method is based on the system according to any one of claims 1-7: The method includes: Constructing a mouth base space model through principal component analysis based on the video data of the target person; Realtime reading the audio features of the input speech and predicting the dynamic coefficients of the mouth base space model through a neural network; Reconstructing a mouth image in combination with the mouth base space model and performing alpha blending with a pre-stored reference face image to generate a real-time digital human output synchronized with the speech.
9. A computer-readable storage medium storing a computer program therein, characterized in that: The computer program, when run, executes the method according to claim 8.
Citation Information
Patent Citations
Method for rendering 3D digital virtual human real-time behavior based on WebGL
CN111882628A
Photo-realistic synthesis of three dimensional animation with facial features synchronized with speech
US20120280974A1