Method for quickly switching and multiplexing digital human images

By building a digital human basic template and multi-layer layer architecture with a unified skeletal structure, combined with dynamic fault tolerance strategies, the problems of low generation efficiency, high resource consumption and difficult real-time interaction in digital human technology are solved, and efficient and stable digital human image switching and reuse are achieved, and multi-role parallel driving and cross-platform adaptation are supported.

CN120298558AActive Publication Date: 2025-07-11LIANGSHENG DIGITAL CREATIVE DESIGN (HANGZHOU) CO LTD

Patent Information

Application Number
CN202510774721.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-07-11
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

The existing digital human technology has bottlenecks in multimodal data processing, real-time optimization and privacy protection, resulting in inefficient generation, huge consumption of computing resources, and difficulty in real-time interaction and personalized performance.

Method used

By building a digital human basic template with a unified skeletal structure, using a multi-layer layer architecture and dynamic fault tolerance strategy, the rapid switching and reuse of the digital human image is achieved, the resource caching mechanism is used to reduce duplicate data processing, and combining affine transformation matrix and key point mapping to ensure the stability and continuity of the switching process.

Benefits of technology

It significantly improves the efficiency and reliability of digital person image switching, reduces resource consumption, realizes high fluency real-time performance and multi-role parallel driving, and supports cross-platform adaptation and voice synchronization lip generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298558A_ABST
    Figure CN120298558A_ABST
Patent Text Reader

Abstract

The invention discloses a digital human image rapid switching and multiplexing method, and relates to the technical field of digital humans, and the method comprises the steps: pre-constructing a digital human image template parameter file, and aligning key points of different digital human images to a standard reference coordinate system through an affine transformation matrix; mapping the key point data of the driving source to a coordinate system of the target digital human image to generate a target key point set; a resource caching mechanism is used for managing texture and template data of the digital human image, and resource loading delay during switching is reduced through a caching multiplexing strategy; and based on the target key point set, synthesizing a digital human image frame through a hierarchical rendering engine, mapping key point data of the same driving source to a plurality of digital human images through differential parameter fine tuning, and independently controlling the skeleton deformation proportion and expression amplitude of each image. According to the scheme, the switching delay can be remarkably reduced, the resource reuse rate is improved, and multi-role parallel driving and voice synchronization mouth shape generation are supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of digital humans, and particularly to a method for rapid switching and reuse of digital human images. Background Art

[0002] With the breakthroughs in generative artificial intelligence and real-time rendering technologies, digital human technology has gradually penetrated into fields such as virtual live streaming, online education, and customer service, becoming an important carrier for human-computer interaction. Its core goal is to achieve high-fidelity image generation, low-latency interaction, and personalized adaptation. However, existing technologies still face significant bottlenecks in multi-modal data processing, real-time optimization, and privacy protection. For example, traditional digital human systems rely on high-precision 3D modeling and motion capture, resulting in low generation efficiency and huge consumption of computing resources; when switching each time, the system needs to repeatedly load large-volume models and textures, and the data volume of a single high-precision digital human model can reach several GB, resulting in dropped frame rates, excessive memory occupancy, and switching stuttering, making it difficult to achieve real-time interaction; at the same time, the fixed-template-based driving mechanism cannot meet the diverse scenario requirements, and the personalized performance is limited. Therefore, based on the above problems, the present invention proposes a method for rapid switching and reuse of digital human images. Summary of the Invention

[0003] Object of the Invention To solve the above problems, the object of the present invention is to provide a method for rapid switching and reuse of digital human images, which realizes modular reuse of digital human resources through template standardization and multi-layer architecture, greatly reduces repeated data processing, and improves the efficiency of switching between different character images; at the same time, the introduced dynamic fault tolerance strategy can ensure the continuity and stability of the switching process, avoid problems such as stuttering and interruption that may occur during switching, and significantly improve the reliability and real-time performance of digital human image switching.

[0004] Technical Solution To achieve the above object, the present invention provides a method for rapid switching and reuse of digital human images, which realizes rapid replacement and component reuse of digital human images through standardization of digital human image templates, pre-configuration of multi-role mapping parameters, and application of a multi-layer mechanism, and ensures stable continuity of the switching process by introducing a dynamic fault tolerance strategy. Specifically, the method establishes a digital human basic template with a unified bone structure, and each preset digital human role corresponds to a set of mapping parameters; uses the digital human basic template and mapping parameters to construct a digital human image with a multi-layer architecture; when switching roles, calls the mapping parameters of the target role to update the relevant layers of the basic template, so as to quickly switch the current digital human image to the target role image; and monitors the resource loading situation during the switching process, and automatically executes a fault tolerance strategy if a loading failure or delay occurs to ensure the smooth completion of the switching.

[0005] In the first aspect, the present invention provides a method for rapid switching and reuse of digital human images, including: Pre - construct a digital human image template parameter file in a unified format. The template parameter file includes a set of facial key feature point coordinates, basic texture map data, and local material layer information, and align the key points of different digital human images to a standard reference coordinate system through an affine transformation matrix; Map the key point data of the driving source to the coordinate system of the target digital human image to generate a set of target key points; Use a resource caching mechanism to manage the texture and template data of digital human images, and reduce the resource loading delay during switching through a cache reuse strategy; Based on the set of target key points, synthesize digital human image frames through a hierarchical rendering engine, and ensure the continuity and stability of the animation by combining key point sequence smoothing processing and anomaly compensation strategies; the hierarchical rendering engine divides the digital human face into multiple independent layers, and each layer synthesizes a complete image according to a preset occlusion order and transparency blending rule, and supports the rapid replacement and reuse of local textures; The method supports parallel driving of multiple characters. By fine - tuning differential parameters, map the key point data of the same driving source to multiple digital human images, and independently control the bone deformation ratio and expression amplitude of each image.

[0006] Further, the digital human image template parameter file is a general digital human model with a unified bone structure and interface definition. The unified bone structure includes the positions and names of predefined main human bone nodes, and the interface definition includes standardized parameter interfaces for appearance features.

[0007] Further, the generation method of the affine transformation matrix is: Solve the optimal transformation matrix by the least - squares method to minimize the mapping error between the key point set of the source digital human image and the corresponding key points of the standard reference coordinate system.

[0008] Further, the resource caching mechanism manages cached data through a unified resource ID index, supports asynchronous loading and frame - by - frame batch loading to achieve smooth transition during switching, and maintains a dynamic resource cache pool, and adopts the LRU strategy to eliminate unused resources.

[0009] Further, the key point sequence smoothing processing adopts an exponential moving average algorithm, and its recurrence formula is: (3); In the formula, is the key point position after smoothing processing; is the frame index; is the smoothing coefficient; is the original key point position of the current frame; It is the position of the key points after smoothing the previous frame.

[0010] Furthermore, the anomaly compensation strategy includes: When key points are missing, the coordinates after smoothing the previous frame are preferentially used for filling; If multiple consecutive frames are missing, the preset static neutral expression coordinates in the template file are called as replacement values.

[0011] Furthermore, the hierarchical rendering engine splits the digital human image into at least a basic image layer, a clothing layer, and an expression layer. Each layer is superimposed through a preset alignment method to form a complete image, and the switching of different digital human roles is achieved by replacing one or more of the layers.

[0012] Furthermore, it also includes a lip synchronization step based on voice driving. Through phoneme recognition and predefined lip shape mapping relationships, the audio input is converted into a sequence of mouth key points of the target digital human image, and coordinated facial movements are generated in combination with the default expression curve.

[0013] Furthermore, the method supports parallel driving of multiple roles. Through differential parameter fine-tuning, the key point data of the same driving source is mapped to multiple digital human images, and the bone deformation ratio and expression amplitude of each image are independently controlled.

[0014] Furthermore, the implementation methods for adapting to multiple platforms of the method include: Adopt a cross-platform graphics rendering framework to dynamically load rendering pipelines adapted to different terminals; Through the interface adaptation middle platform module, the coordinate system and resolution are uniformly converted to ensure the consistency of cross-platform animation output.

[0015] In a second aspect, the present invention also provides a system for rapid switching and reuse of digital human images. The system is based on the method described in the first aspect above and includes: A template preprocessing module for parsing each digital human image into a template parameter file; A key point alignment module for calculating the affine transformation matrix of each digital human image relative to the standard reference coordinate system based on the template parameter file to achieve the standardized alignment of the key point set; A drive mapping module for mapping the key point data of the drive source to the coordinate system of the target digital human image through affine transformation to generate a target key point set; A resource cache management module for maintaining a dynamic resource cache pool, managing texture and template data through the LRU strategy, and supporting asynchronous loading and reuse; The hierarchical rendering engine module is used to divide the digital human face into multiple independent layers according to the target key point set and reusable resources, synthesize images in a preset order, and output continuous animation frames by combining smoothing processing and an anomaly compensation mechanism; The interface adaptation middleware module is used to shield the underlying platform differences and achieve unified scheduling of cross-terminal resolutions, coordinate systems, and rendering pipelines by dynamically loading adaptation sub-modules.

[0016] Furthermore, it also includes a voice analysis module, which is used to perform phoneme recognition on the audio data and generate a sequence of mouth key points corresponding to the target digital human image when the driving source is audio input.

[0017] In a third aspect, the present invention also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is run by a processor, it executes the method for rapid switching and reuse of the digital human image described above.

[0018] The present invention constructs a standardized template file containing a set of facial key feature point coordinates, a base texture map, and hierarchical materials, and combines an affine transformation matrix to align the key point coordinate systems of multiple digital human images; in the real-time driving stage, it uses a dynamic resource cache pool and an LRU strategy to manage texture and template data, divides the human face into independent layers for synthesis through a hierarchical rendering engine, and introduces an exponential moving average algorithm to smooth the key point sequence, and combines a priority compensation strategy to solve the problem of abnormal data. In addition, the system dynamically loads a cross-platform rendering pipeline through the interface adaptation middleware module to achieve unified multi-terminal resolution and coordinate system. This solution significantly reduces the switching latency, improves the resource reuse rate to over 90%, and supports multi-role parallel driving and voice-synchronized lip generation, solving the core defects of high latency, resource redundancy, and poor cross-platform adaptability in traditional technologies, and achieving the technical effects of high smoothness, low resource consumption, and multi-terminal consistency in scenarios such as virtual live streaming and online education.

[0019] Beneficial effects By implementing the method for rapid switching and reuse of a digital human image provided by the present invention, the following technical effects are achieved: (1) By constructing a standardized template parameter file and combining an affine transformation matrix, the present application realizes the unified alignment of the key point coordinate systems of different digital human images. This mechanism breaks through the constraints of the model topological structure differences in traditional multi-image switching, enables the driving data to migrate seamlessly between heterogeneous digital human images, solves the problem of cross-image continuity of actions and expressions, and lays a foundation for the generalization of the core driving process.

[0020] (2) Based on the LRU strategy and asynchronous loading dynamic resource management mechanism, a multi-level cache pool is constructed and the on-demand scheduling of textures and template data is realized. This architecture significantly reduces the memory occupancy and IO load during switching by intelligently eliminating redundant resources and reusing high-frequency data, eliminating the latency and resource waste problems caused by repeated loading of large models in traditional technologies at the system level.

[0021] (3) The digital human face is divided into independent functional layers such as the skin layer, eye layer, and mouth layer, and is rendered and synthesized layer by layer through preset occlusion order and transparency blending rules. This technology realizes the decoupled control of the facial area and the rapid replacement of local resources, greatly improving the efficiency of texture update and local feature switching while ensuring a high-fidelity visual effect, providing a refined operation space for multi-style image reuse.

[0022] (4) The exponential moving average algorithm is used to perform low-pass filtering on the key point trajectory, and an abnormal data processing mechanism is combined with a priority compensation strategy. This mechanism effectively suppresses the interference of high-frequency noise and sudden loss of the driving signal on the animation coherence, ensuring the stable output of the digital human's expression actions in complex interaction scenarios through dynamic deviation correction and fault tolerance.

[0023] (5) Compared with the end-to-end image generation method, such as deep learning face swapping, this method ensures the controllability of details and the smoothness of frame transition during switching through preprocessing and layer rendering, avoiding the distortion and latency that may occur in the end-to-end method, and having better real-time performance and stability.

[0024] (6) By configuring multiple sets of differentiated mapping parameters for a single driving source, multiple digital human characters can be driven in parallel to achieve synchronous animation of "one person driving multiple people". This mechanism ensures that the bone deformation ratio and expression amplitude of each character conform to its own image characteristics, expanding the application scenarios of digital humans. For example, in a virtual meeting, the facial expressions of a single user can simultaneously drive multiple digital human images with different styles to achieve collaborative performances.

[0025] (7) A general lightweight cross-platform architecture is designed, and the underlying platform differences are shielded through the interface adaptation middleware, ensuring that the real-time driving and image switching processes of this method can be smoothly executed on various terminal devices such as mobile terminals, web terminals, and desktop terminals. The animation performance and effects on different platforms are consistent, and the development cost of multi-platform adaptation is reduced. The animation performance and effects on different platforms are consistent, and the development cost of multi-platform adaptation is reduced.

[0026] (8)Introduce a pure voice-driven mechanism. Through audio analysis and phoneme-lip mapping methods, even with only microphone frequency input, it can drive the mouth movements of the digital human, and supplement with the default expression curve to ensure the overall facial expression is natural and coordinated. Therefore, in scenarios without a camera such as telephone conferences and voice live broadcasts, this method can still achieve real-time synchronization of the digital human's lip shape and voice content, expanding the applicability of this method in various input modes.

[0027] (9) The basic template predefines the positions and names of the main skeletal nodes of the human body, as well as the parameter interfaces for appearance features, enabling different digital human roles to be constructed based on this template. By adopting a standardized basic template, the differences between different roles only need to be described by parameterized data, realizing the unified standardization of digital human image construction, facilitating subsequent rapid switching and reuse, and enabling the system to manage and apply various role features centered on the template. Brief Description of the Drawings

[0028] To make the above method for rapid switching and reuse of digital human images of the present invention more clearly understandable, the drawings required for the specific implementation manners of the present invention will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings without creative efforts.

[0029] Figure 1 Indicates the flowchart of the method of the present application; Figure 2 Indicates the system architecture diagram of the present application. Detailed Embodiments

[0030] Without departing from the core idea of the present invention, those skilled in the art can also conceive of various deformation schemes. For example, for the mapping of driving data between different digital human images, key point alignment can also be achieved by establishing a unified bone binding relationship or a standardized expression parameter system (such as FACS) to achieve a similar cross-image driving effect. Another example is in the generation of digital human images, a pre-trained generative adversarial network model can be used to directly map the driving source to the animation frames of the target image, realizing real-time "face swapping" without explicit key point conversion. The present invention is equally applicable to the above alternative implementation methods, and the technical effects obtained are also within the coverage of the present invention.

[0031] Embodiment 1: Provides a method for rapid switching and reuse of digital human images, and the method flow is as Figure 1As shown, it includes: pre-building a digital human image template parameter file in a unified format, where the template parameter file includes a set of facial key feature point coordinates, basic texture map data, and local material layering information, and aligning the key points of different digital human images to a standard reference coordinate system through an affine transformation matrix; mapping the key point data of the driving source to the coordinate system of the target digital human image to generate a target key point set; using a resource caching mechanism to manage the texture and template data of the digital human image, and reducing the resource loading delay during switching through a cache reuse strategy; based on the target key point set, synthesizing digital human image frames through a hierarchical rendering engine, and ensuring the continuity and stability of the animation by combining key point sequence smoothing processing and anomaly compensation strategies.

[0032] A system for rapid switching and reuse of digital human images is provided. The system is based on the aforementioned method, and its architecture is as Figure 2 shown, including: an interface adaptation middle platform module that shields the differences of the underlying platforms (PC, mobile, Web, etc.), provides a unified API interface, and is responsible for coordinate system and resolution conversion; an image processing and key point extraction module that is responsible for preprocessing the key points of the digital human image template in the offline stage and performing face key point detection, preprocessing, and normalization in the real-time stage; a key point alignment module that aligns the key points of each digital human image template to a standard reference coordinate system based on an affine transformation matrix; a driving mapping module that maps the real-time driving source key points to the coordinate system of the target digital human image through an affine transformation to generate a target key point set; a cache and resource management module that maintains a dynamic resource cache pool, adopts the LRU strategy and asynchronous loading to manage texture and template data, realizes resource reuse, and reduces the loading delay; a hierarchical rendering engine module that splits the digital human face into multiple layers, synthesizes images in a preset order, and combines key point sequence smoothing processing and anomaly compensation strategies to output continuous animation frames; outputs animation frames / video streams, provides the continuous image frames generated by the rendering engine to the upper-layer application or encoder to form a real-time video stream for display.

[0033] Specifically, it is as described below.

[0034] 1. Technical solution The core idea of the method is to introduce a unified image template parameter structure and key point mapping driving mechanism, enabling different digital human images to be quickly interchanged under the same set of driving data while maintaining consistent actions and expressions. The method includes the following steps: Step 1. Image template preprocessing: First, a standardized digital human basic template is established. This basic template is a general human body model with a unified bone skeleton and interface definition. For example, the basic template predefines the positions and names of the main bone nodes of the human body, as well as the parameter interfaces for appearance features, enabling different digital human characters to be constructed based on this template. By adopting a standardized basic template, the differences between different characters only need to be described by parameterized data, realizing the unified standardization of digital human image construction. This measure provides a basis for subsequent rapid switching and reuse, enabling the system to manage and apply various character features centered around the template. Specifically, it includes: In the offline stage, each digital human image is parsed into a template parameter file in a unified format. The template content includes: Coordinate set of facial key feature points: Extract the coordinates of the main feature points on the face of this image as the reference skeleton marks representing the face shape.

[0035] Basic texture mapping data: The skin / facial texture resources of this image are used to present the details of the character's appearance.

[0036] Local mask or layered material: The material layering or mask information for special parts, such as the transparent mask in the mouth area, the eye highlight layer, the semi-transparent texture layer of the hair, etc.

[0037] The template parameters are saved using a unified data structure to ensure the consistent correspondence of key parameters between different images. After preprocessing, each image generates a standard template file, recording the positions of its feature points in the standard reference coordinate system and the associated texture resource indices, facilitating subsequent rapid loading and mapping.

[0038] Step 2: Affine alignment of key points: Calculate the coordinate transformation relationship of each image relative to the standard reference coordinate system, specifically including: Using the set of key points of this image extracted from the template as the input, solve an affine transformation matrix that maps it to the corresponding key points in the reference coordinate system. Preferably, the least squares method is used to solve the transformation matrix that minimizes the error. For example, solve the matrix To minimize , where is the frame index; is the affine transformation matrix to be solved; is the set of key points of the source digital human image; is the set of key points in the standard reference coordinate system to obtain the optimal affine matrix. This matrix is used to align the coordinate system of the source image to the unified reference system, thereby realizing the standardization of the initial postures of different images. The calculated affine matrix of each image is stored in its template file for later use.

[0039] Step 3: Driving key point mapping: During the real-time running stage, the current set of facial key points is obtained from the driving source for each frame, and its format is the same as that of the key points in the template. To apply this driving action to a new target image, it is necessary to convert the source key point coordinates to the coordinate system of the target image. Assume that the affine matrix of the target digital human image relative to the reference coordinate system is, and the matrix of the corresponding image of the current driving source is. Then the relationship between the source key points of the current frame mapped to the coordinate system of the target image can be expressed as: (1); The above formula means: First, use the transformation matrix of the source image to map the set of driving key points captured in the current frame to the standard reference coordinate system, and then through the matrix of the target image convert the reference coordinate system coordinates to the coordinate space of the target image itself to obtain the corresponding set of key points of the target image . After such a coordinate system transformation, the new digital human image can accurately reproduce the facial expressions and action postures of the driving source at this moment, laying a foundation for subsequent rendering.

[0040] Step Four: Resource Loading and Cache Reuse: To achieve fast image switching and avoid repeated loading of a large number of resources, the system designs an efficient texture / model cache mechanism. A cache pool for image resources is maintained during runtime. The loading logic is as follows: When it is necessary to render a certain image, first check whether there are already resources of this image in the cache. If the cache hits, directly reuse the existing data without reloading; if the cache misses, read the texture and template data of this image from the storage medium, calculate its affine alignment matrix, and then add these data to the cache. To prevent the cache from growing infinitely, an LRU strategy is adopted to manage the cache, and image resources that have not been used for a long time are eliminated when the memory occupancy approaches the upper limit. In addition, each image is managed through a unified resource ID index to ensure efficient and reliable searching and recycling. Thanks to the above strategies, frequently used digital human images do not need to be reloaded, greatly reducing the switching latency and memory overhead. For extremely large texture or model data, background asynchronous loading or frame-by-frame batch loading is used for smooth transition to avoid stuttering during the switching instant.

[0041] Step Five: Image Composition and Rendering: The rendering engine draws the digital human based on the set of key points of the target image and the corresponding texture resources to generate the current frame image. Denote the drawing function of the rendering engine as , then the output image for each frame is expressed as: (2); In the formula, is the output digital human image frame; The texture resources for the target digital human image; That is, the texture resources are positioned and deformed according to the key point set and then attached to the digital human model to generate the image of the current frame. During the rendering process, the texture of the facial area is affine deformed according to the key point positions to synchronize the expression actions with the texture deformation, and the corresponding lighting effects and local mask processing are superimposed to ensure a realistic visual effect. This process is repeated for consecutive frames to form an animation sequence of the target digital human changing with the actions of the driving source.

[0042] Step Six: Smoothing Processing of the Key Point Sequence: Since the key point sequence output by the actual driving source contains tiny jitter noises, directly using it will cause slight jitters in the rendering results. Therefore, the method introduces a low-pass filtering smoothing strategy for the key points of consecutive frames, preferably using the exponential moving average method. Its recurrence formula is: (3); In the formula, is the position of the key point after smoothing processing; is the smoothing coefficient, and its value range is from 0 to 1; is the position of the original key point of the current frame; is the position of the key point after smoothing processing of the previous frame. When is small, it emphasizes the position of the previous frame, achieving a strong smoothing filter; When

[0043] is close to 1, it follows the changes of the current frame more, and can respond to large actions in a timely manner. For the initial frame, can be directly set. After the EMA smoothing processing, the position changes of the key points in each frame are more continuous, eliminating high-frequency jitters, making the animation generated by rendering transition smoothly without jitters. at the th frame is : If the current frame is detected as valid, that is, the indicator function is 。

[0044] If the current frame is missing and the previous frame has a valid value, that is and , then the smoothed value of the previous frame is adopted: 。

[0045] If the current frame is missing and there is no previous frame reference, that is and there is no valid value before, then the template default value is adopted: 。

[0046] According to the above rules, reasonable alternative coordinates can be found for each abnormal key point, avoiding local distortion of the output portrait caused by the missing of individual points. The set of key points after inter-frame smoothing and missing compensation processing will be used to replace the original input for subsequent rendering, so as to ensure the stability and coherence of the animation sequence during the switching process.

[0047] In addition, it should be noted that the smoothing processing and compensation strategy of the present invention are not limited to the methods adopted above. For example, to smooth the driving key point sequence to avoid high-frequency jitter, in addition to using the exponential moving average algorithm, methods such as Kalman filtering, bilateral filtering, or noise reduction filtering based on frequency domain analysis can also be used to achieve a similar smoothing effect. Similarly, for the compensation of missing key points, in addition to preferentially filling with the smoothed value of the previous frame or the preset template default value, other alternative solutions can also be adopted according to needs, such as interpolating and predicting using the positional relationship of adjacent feature points, or calculating the missing point value from the corresponding point on the other side using the symmetry of the human face. No matter which specific smoothing algorithm or compensation method is adopted, as long as it can achieve the continuous stability of the change of the driving key point sequence, thus ensuring the smooth and coherent switching process of the digital human animation, it belongs to the protection scope of the present invention.

[0048] In addition, to improve the robustness of the switching process, the system introduces a dynamic fault tolerance strategy to handle abnormal situations that may occur during the switching process. Specifically, when performing the digital human image switching, the system continuously monitors the loading status of each layer resource. If it is detected that the resource loading of a certain layer fails (such as file missing or damaged) or has not been loaded successfully after exceeding the predetermined time, the system will automatically trigger the fault tolerance mechanism: for the situation where the key layer resource loading fails, the pre-set default resource is used for substitution, such as using the default clothing to replace the specific clothing model with loading failure; for non-critical layers or resources that can be temporarily omitted, the system can choose to skip the loading of this resource to give priority to completing the switching and presentation of the main image. Through the above dynamic fault tolerance strategy, even when some resources are abnormal, the switching of the digital human image can still be completed continuously and stably, without interrupting the entire switching process due to a single module failure. This fault tolerance design further ensures the reliability of the method of the present invention in practical applications.

[0049] 2. System Architecture and Collaboration Process This application adopts a loosely coupled system architecture, where multiple functional modules collaborate to achieve real-time driving and image switching of digital humans. The responsibilities of each module are clearly defined, and they interact through standard interfaces. The overall workflow is as follows: Rendering Engine Module: Responsible for the final rendering of the digital human graphics. It receives the key point positions and texture resources of the target image for the current frame, deforms and fits the texture according to the key point driving, and superimposes the corresponding mask and lighting effects to output the digital human image frame. The rendering engine is built with efficient texture mapping and mesh deformation shaders, which can perform real-time deformation and smooth transition processing on the model mesh of the local facial area to ensure natural connection of facial expressions and details at the moment of image switching, and generate highly realistic animation frames.

[0050] Image Processing and Key Point Extraction Module: It includes two parts of functions: offline template preprocessing and real-time key point detection. In the offline stage, this module uses a face key point detection algorithm to extract the feature key points of each digital human image template, aligns their reference coordinates, and generates a standardized image parameter file for the system to load and use. In the real-time stage, this module obtains the original driving data of each frame from the driving source, performs face key point recognition or expression parameter calculation to obtain the key point set of the driving person for the current frame. Subsequently, necessary normalization processing and coordinate transformation are performed on the key point data, and then it is passed to the subsequent key point mapping algorithm through the interface. When necessary, this module also preprocesses the input image to improve the accuracy and robustness of key point detection.

[0051] Interface Adaptation Middleware Module: As a bridge between the system and upper-layer applications, it provides unified API interfaces for front-ends on different platforms to call the digital human driving rendering and image switching functions. Whether the front-end is a web page, a mobile App, or a desktop application, the middleware module shields the underlying platform differences and provides consistent services for them. This module is responsible for the conversion of different platform coordinate systems, resolution adaptation, and data synchronization between the UI interface layer and the rendering engine to ensure consistent effects during the switching process on various devices. By dynamically loading adaptation sub-modules, the middleware applies corresponding optimization measures for different operating environments to achieve smooth cross-platform docking.

[0052] Cache and resource management module: responsible for loading, caching and managing digital human image template data and texture resources. This module works in conjunction with the rendering engine and image processing module: when receiving a request to switch images, it first queries whether the target image resource is already in the cache. If so, it directly returns the resource handle to avoid repeated loading; if not, it triggers the asynchronous loading process, reads the texture and template data required for the new image into the memory, and updates the cache for the rendering engine to use after completion. Cache management uses the aforementioned unified index and LRU strategy to control memory usage, ensuring that the required resources can still be quickly retrieved and scheduled when multiple image resources coexist. Resources are delivered through standardized data interfaces, and this module is decoupled from other modules. It does not rely on specific implementation details, but can efficiently support resource reuse and switching.

[0053] The above modules work together through interfaces, and the overall operation process is as follows: First, the image processing module continuously obtains the facial key point data of the driving character from the camera video or other inputs, and transmits it to the rendering engine through the interface middle platform after preprocessing; the rendering engine uses the currently activated digital human image template to perform drawing, generate and output continuous animation frames. When it is necessary to switch the image, the upper-layer application issues a switching instruction, which is captured by the interface middle platform and notifies the cache management module to prepare the target image resources. The cache module loads the new image texture and template data in the background, and the key point mapping algorithm converts the key points of the current driving frame to the new image coordinate system according to the aforementioned step three. Then, starting from the next frame, the rendering engine uses the texture resources of the new image and the mapped key points for drawing. Due to the use of strategies such as key point smoothing / compensation and rendering transition, the frames are still closely connected during the switching process: on the switching frame, the action positions of the new and old images remain continuous, and the scene background is not suddenly cleared or redrawn, thereby avoiding any flickering or pauses. The entire switch does not affect users and upper-level applications. The system still maintains real-time performance under the module decoupling architecture, and the collaboration of various units ensures smooth image replacement.

[0054] 4. Extended functionality Based on the above solutions, extended functions are also provided to enhance the flexibility and scope of application of digital human driving and rendering.

[0055] Multi-role parallel driving: The system supports driving inputs to simultaneously control multiple digital human avatars, enabling parallel driving of multiple roles. In scenarios such as virtual meetings and group performances, the expressions or actions of a single user can be synchronously mapped to multiple virtual characters with different styles. To ensure that the performances of each role conform to their respective image characteristics, the system introduces a differential parameter fine-tuning strategy: A set of mapping parameters is pre-configured for each target virtual character, including the face shape ratio adjustment coefficient, expression amplitude scaling factor, and offset correction value of the character. When a series of facial key points are generated by the driving source, the system sends this key point sequence in parallel to multiple role mapping modules. Each role module applies its specific parameter set to transform and fine-tune the input key points to obtain an animation control signal adapted to the skeletal / mesh structure of the role. For example, for a character with a larger facial skeleton, the key point displacement is amplified to present the same amplitude of expression; for a cartoon-style character, an exaggeration coefficient is set to enhance its expression changes, making its movements have an exaggerated cartoon effect. The driving processes of each role are independent and parallel, without interfering with each other. This parallel driving ability meets the requirement of synchronous performances of multiple digital humans in virtual scenarios, enabling them to maintain synchronized actions and distinct characteristics under the same input even if the characters have different appearances.

[0056] Layered rendering mechanism: The system realizes the composition of digital human avatars by splitting them into multiple layers (hierarchies). Preferably, a digital human avatar includes at least multiple layers such as a basic image layer, a clothing layer, and an expression layer, with each layer carrying different types of content: The basic image layer is used to present the basic physical appearance of the digital human, such as body shape and basic features; the clothing layer is used to overlay the clothing and accessories worn by the character; the expression layer is used to overlay facial expressions or dynamic effects. The layers are superimposed on each other through a predefined alignment method to form a complete digital human avatar. When it is necessary to switch digital human roles, the system gradually replaces the resources of the corresponding layers for the changed parts, such as replacing the data of the clothing layer and the expression layer, while the basic image layer remains unchanged, thereby quickly generating the image of the new role. Such a multi-layer mechanism ensures the maximum reuse of resources, enabling only the necessary parts to be updated during the switching process and significantly improving the switching efficiency.

[0057] During the process of splitting the human face into multiple layers according to functional regions and then drawing and synthesizing them separately, the system divides the digital human face into a bottom skin layer and several independent upper layers such as lips, teeth, eyes, eyebrows, and hair according to the facial structure. Each layer maintains the texture and geometric information of the corresponding region. During rendering, each layer is drawn in sequence according to the preset occlusion relationship and order, and transparency is used for blending and overlaying to finally generate a complete human face image. For example, the hair is drawn as an independent layer after the skin layer, and the semi-transparent effect of the hair strands is presented through blending to ensure that the hair is correctly occluded behind the face; another example is in the treatment of the mouth, where the lips and teeth are respectively used as different sub-layers: when the character opens its mouth, the opacity of the teeth / tongue layer is increased to make it clearly visible, and when the mouth is closed, its opacity is decreased so that the teeth are occluded by the lip layer. Through layer-based drawing, decoupled control of each facial region is achieved: modifying or replacing a certain layer will not affect the display of other layers. This greatly improves the efficiency of local adjustment and switching of the image, supports quickly replacing local features as needed to reuse image resources. In addition, each layer adopts a unified blending rule to ensure that the final synthesized image is natural and smooth at the edge transitions and the connection between different regions is realistic. The layered rendering mechanism enables the system to flexibly adjust the facial detail effects, correctly handle the hierarchical occlusion in complex scenarios, enhance the authenticity of the digital human expression, and provide a smoother visual transition during image switching.

[0058] Speech-based Lip Sync: Further provides a method to achieve digital human lip sync when only driven by speech. Even without camera video input, the system can synchronize the digital human's lip movements with the spoken content in real time based on pure audio. When the driving source is audio, the speech analysis sub-module is started: First, speech recognition or feature extraction is performed on the input speech to obtain the time axis at the phoneme level, and the continuous speech is divided into a series of phoneme segments annotated by time. Next, the system maps the phoneme sequence to a predefined standard lip animation sequence. Specifically, the system pre-establishes a speech expression library to store the key frame mapping relationships between the main phonemes / syllables and the corresponding lip shapes. For example, vowel phonemes correspond to open lip shapes, bilabial phonemes correspond to pursed or closed lip shapes, and fricative phonemes correspond to tooth-revealing lip shapes, etc. The speech driving module sequentially extracts the lip key frames corresponding to each phoneme in the speech according to the audio time axis, and interpolates and smooths these key frames according to the phoneme duration to generate a continuous and natural lip movement sequence. At the same time, to avoid only lip movements while the rest of the face appears rigid, the system combines the default facial expression curves in the expression library to apply basic coordinated actions to other facial areas. Finally, the lip sync module outputs a sequence of mouth key points that are strictly aligned with the input speech, driving the digital human to produce accurate speaking lip shapes. For example, when an introductory greeting is input, the system extracts the continuous phoneme sequence therein and drives the digital human to perform actions such as opening the mouth, closing the mouth, and showing a toothy smile in sequence, corresponding precisely to the speech content. In this way, even in a scenario without video, the digital human's lip movements can be consistent with the speech, presenting a natural and smooth speaking effect. This speech driving mechanism significantly enhances the digital human's expressiveness in a pure audio scenario, enabling the digital human image to achieve realistic lip sync in various input modes.

[0059] 5. Cross-platform Applicability and Exception Handling To enable this method to run stably on various hardware platforms and complex environments, the system focuses on cross-platform compatibility in design and is equipped with a complete exception handling mechanism.

[0060] The system uses a highly portable graphics rendering framework and standardized interface programming, enabling the core rendering and driving codes to maintain consistent running performance on different operating systems such as Windows, Linux, macOS, Android, and iOS. The interface adaptation middle platform dynamically loads the corresponding adaptation modules according to the current platform environment during initialization. For example, a streamlined rendering pipeline is enabled on mobile devices to reduce power consumption, while high-performance shaders are called on desktop devices to fully utilize the GPU performance. For graphics hardware from different manufacturers, the system uses common shader language features and standard extensions, avoiding relying on private functions of specific devices to enhance cross-hardware compatibility. Through the above measures, the entire system has good compatibility and real-time performance on various platforms and devices, ensuring the wide applicability of the digital human switching function.

[0061] For abnormal situations that may occur during operation, design an exception handling strategy to ensure stable and continuous operation of the system: If network delays cause failure to obtain remote texture resources or local files are missing, the system will detect this and promptly enable alternative solutions, such as switching to low-resolution alternative textures or using preset placeholder textures instead, to avoid interruptions in the rendering of the screen, and asynchronously retry loading the original resources in the background.

[0062] If the key point data of real-time driving is abnormal, the system will start the fault tolerance strategy. For slight fluctuations, the system will combine the smooth data of the previous frame to make a transition, or temporarily reduce the driving frame rate to wait for data recovery; for severely distorted frames, the frame data will be directly discarded to prevent exaggerated and distorted expression output.

[0063] If some rendering effects are not supported on a specific device, the system presets multiple rendering schemes for switching. When a compatibility error occurs, it will automatically downgrade to a conservative compatibility mode, such as turning off advanced lighting effects or using software rendering paths to ensure that basic functions operate normally.

[0064] The system allows you to set a maximum switching delay threshold. If an image switching operation exceeds the preset time and is still not completed, the system will automatically terminate the current switching process and roll back to the state before the switch to maintain the continuity of the user's visual experience and prevent the interface from staying in an unfinished transition for a long time.

[0065] Through the above exception handling strategies, the system can maintain continuous and stable operation even in the case of unstable network, large hardware differences or abnormal sensor data. Users will hardly notice the internal failure and recovery process, and always see the smooth and natural digital human presentation effect.

[0066] In the event of an abnormality in resource loading or real-time driving, the present invention can also take other equivalent fault-tolerant measures. For example, when remote texture resource acquisition fails or local files are missing, in addition to enabling low-definition textures or placeholder textures as substitutes, the system can also call the local cached version of the last valid resource screen if conditions permit, or directly skip the rendering output of the frame in extreme cases to avoid screen freeze interruption. For another example, for frames with severely distorted camera acquisition data, in addition to directly discarding them, an expression action database can also be established in advance to replace the erroneous key points identified in the abnormal frames with the most similar legal expression gesture data. All these modifications are intended to ensure that the system can still run continuously and stably under abnormal circumstances without affecting the switching effect of the digital human. These equivalent deformations based on the concept of the present invention also fall within the scope of protection required by the present invention.

[0067] Embodiment 2: In response to the possible abnormal fluctuations or missing of key-point data during real-time acquisition, a key-point protection mechanism is added to ensure the stable continuity of the digital human animation. The protection point refers to the safety alternative value or compensation measure adopted when the key-point detection result is abnormal, that is, a kind of fault tolerance protection for key points. The system first performs abnormal detection on the set of key points input for each frame to determine whether there are any invalid or abnormal key points. For example, when the camera capture quality is poor or the user's action is too fast, abnormal values or direct missing may occur in the coordinates of individual key points. Once abnormal key-point data is detected, the system activates the protection mechanism and corrects the key points through filtering and compensation strategies: slight jitter noise is buffered through sequence smoothing, and severe missing is filled with historical data or template data. The entire protection process is completed within milliseconds, effectively preventing the digital human expression from being distorted or discontinuous due to abnormal key points.

[0068] For the detected slight abnormalities, the system introduces low-pass filtering to smooth the key-point sequence, preferably using the exponential moving average algorithm. By smoothing and filtering high-frequency noise and weakening overly sensitive tiny jitters, the position change of key points becomes more continuous. Its recurrence formula is expressed as: (4); In the formula, is the key-point coordinate after smoothing processing; is the smoothing coefficient, and its value range is from 0 to 1; is the original key-point coordinate detected in the th frame; is the key-point coordinate after smoothing processing of the previous frame. A smaller value makes the smoothing result more biased towards the data of the previous frame, thus filtering out high-frequency jitters with very small amplitudes;

[0069] If abnormal detection finds that there are obvious errors or missing in the key-point data of some frames, the system adopts a priority compensation strategy to provide alternative coordinates for abnormal points, that is, using historical valid data or template reference values as "protection points" to avoid abnormal distortion in the rendering output. The compensation mechanism searches for the best alternative value in the order from new to old, from dynamic data to static templates. The specific rules are as follows: Previous frame reuse: If a key point in the current frame is missing or invalid, but the previous frame has valid coordinates for that point, then the smoothed position of that key point from the previous frame is preferentially used to replace the current frame value to ensure the continuity of the motion trajectory. This allows for a smooth transition during brief detection failures, and users will not notice the loss of single-frame data.

[0070] Template default value: If a key point in the current frame is missing and there is no valid reference in the previous frame, then the predefined template default coordinate value is used. This default value comes from the key point template of the target digital human image and represents a reasonable static position. Filling in with the template value can ensure the integrity of the basic structure of the digital human face and prevent mesh stretching or distortion due to the missing coordinate of a certain point.

[0071] The above compensation rules can also be described in mathematical form: Let the original key point coordinates detected in the th frame be , its validity indicator function be , the smoothed key point coordinates be , and the template default value be . Then the final coordinates of this key point output by the protection mechanism are expressed as: (5); When the detection value of the current frame is valid, the smoothed result of the current frame is directly adopted; when the current frame is missing and there is a reference in the previous frame, the protected coordinate of the previous frame is reused; if there is a continuous missing, the static coordinate in the template is used for replacement. Through this strategy, whenever an abnormal key point appears, the system can always select a reasonable replacement position for it, thus ensuring the integrity and stability of the digital human face animation.

[0072] The above smoothing and compensation mechanisms together constitute the key point protection strategy. In practical applications, different types of key point anomalies will be corrected in time to avoid negative impact on the animation. For example, when the user blinks quickly, the eye features change dramatically, which may cause the eye corner key points of the current frame to be unable to be recognized normally. At this time, the system will detect the "eye corner key point missing" anomaly and automatically enable the protection mechanism: if the eye corner coordinates of the previous frame exist, the smoothing value of the previous frame will be used to maintain the continuous transition of the eye position; if the eye corners cannot be detected in consecutive frames, the reference coordinates of the eye corners in the template will be temporarily used to keep the digital human eyes in a natural open state. In this way, no matter whether the user's eyes close or open quickly, the eye movements of the digital human image are still smoothly connected, and there will be no abnormal phenomena such as collapse and jumping. For another example, for the key points of the corners of the mouth, when the user makes exaggerated laughter or speaks, the position of the corners of the mouth is sometimes not detected accurately due to facial angles or occlusion. Assuming that the coordinates of the right corner of the mouth are abnormally offset or even missing in a certain frame, the system will give priority to using the smooth coordinates of the corner of the mouth in the previous frame to continue to drive the mouth shape; if the abnormality persists for multiple frames, the system will place the right corner of the mouth in the neutral position predetermined in the template. As a result, the deformation of the digital human character's mouth will pause slightly but maintain a normal shape, and will not cause mouth deformity due to single-frame data errors. Through the above strategy, even if individual feature points are temporarily lost or abnormal, the various parts of the digital human face can still maintain their proper expression and relative position, and the entire animation effect is still natural and credible to the viewer.

[0073] In summary, the system drives the digital human animation through the precise extraction and mapping of key points, and combines the protection point mechanism to ensure the stability and reliability of the key point data. On the one hand, the standardized key points serve as a unified driving medium to achieve a smooth transition of expressions and movements between multiple digital human images; on the other hand, for random jitter and abnormal missing in the key point sequence, the system uses protection strategies such as exponential smoothing and graded compensation to correct in real time, so that the animation frames in the switching process can remain continuous and complete in any case, and will not produce distortion and jitter due to abnormal sensor data. The key points and protection points cooperate with each other to play the dual role of driving animation and fault-tolerant voltage stabilization: it not only ensures that the movements are highly restored when switching between different images, but also provides a safety bottom line for each key feature through abnormal detection and compensation mechanisms, and ultimately ensures the smoothness, stability, realism and reliability of the digital human image switching process.

[0074] Embodiment 3: Based on the above-mentioned embodiment, the specific process of the method for fast switching of digital human images is described in combination with a virtual live broadcast scene: Step 1: Preload image template: The host prepares two digital human images A and B before the live broadcast begins. The system performs offline preprocessing on image A and image B respectively, generates standardized key point templates and texture data files, and loads these two sets of template resources into the system cache in advance for backup.

[0075] Step 2: Start live broadcast and drive animation: At the beginning of the live broadcast, the host selects image A as the initial digital human image. The system activates the template parameters of image A, starts to capture the host's facial expressions and movements through the camera, calculates the key points of the face in real time and drives image A to move synchronously. The rendering engine continuously outputs animated images of image A changing with the host's expression, and the host's digital human image A follows the host's movements in real time and responds smoothly.

[0076] Step 3: Trigger image switching: During the live broadcast, the anchor decides to change the digital human image from A to B and sends a switching command through the operation interface. After the interface adaptation center captures the command, it immediately notifies the system to prepare to switch the current digital human image from A to B and dispatches related resources.

[0077] Step 4: Load and map the new image: After receiving the switching instruction, the system quickly obtains the texture and template data of image B from the cache module. At the same time, the image processing module calls the key point mapping algorithm to convert the current key point set of the anchor's face from image A to the coordinate system of image B. In the next frame, the rendering engine starts to draw using the texture and mapped key points of image B. Because the driving data of the previous image A is seamlessly mapped to image B, the digital human image B immediately takes over and appears on the screen.

[0078] Step 5: Complete the switch presentation: At the moment of switching, the system uses a smooth transition strategy to ensure the continuity of the picture: when the rendering engine switches frames to draw image B, it uses the background and scene status of the previous frame without suddenly clearing or resetting the picture, thus avoiding any screen flickering or frame skipping. After the switch is completed, the digital image of the anchor in the live broadcast screen has seamlessly changed from A to B, and the facial expressions and body movements remain consistent with before the switch, and only the appearance has changed.

[0079] The audience will see the anchor's digital human image switch from A to B in a very short time. The whole process is smooth and the switching delay is almost imperceptible to the naked eye. The switched image B can immediately continue the anchor's expression and action state before the switch, without any pause or lag. This proves that this solution is efficient and stable. Even in scenarios with extremely high real-time requirements such as live broadcasts, the rapid switching of digital human images can still be completed smoothly without affecting the user experience.

[0080] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable non-transitory storage media containing computer-usable program code.

[0081] The present invention can provide computer program instructions to a management platform of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed through the management platform of the computer or other programmable data processing devices generate a device for implementing the system.

[0082] These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions of the system.

[0083] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, so that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable devices provide steps for implementing the functions of the system.

Claims

1. A method for rapid switching and reuse of digital human images, characterized in that, Including: Pre - construct the digital human image template parameter file, and align the key points of different digital human images to the standard reference coordinate system through the affine transformation matrix; Map the key point data of the driving source to the coordinate system of the target digital human image to generate a target key point set; Use the resource caching mechanism to manage the texture and template data of the digital human image, and reduce the resource loading delay during switching through the cache reuse strategy; Based on the target key point set, synthesize digital human image frames through a hierarchical rendering engine, and combine key point sequence smoothing processing and anomaly compensation strategies to ensure the continuity and stability of the animation; The hierarchical rendering engine divides the digital human face into multiple independent layers, and each layer synthesizes a complete image according to the preset occlusion order and transparency blending rules, and supports the rapid replacement and reuse of local textures; The method supports parallel driving of multiple characters. Through differential parameter fine - tuning, map the key point data of the same driving source to multiple digital human images, and independently control the bone deformation ratio and expression amplitude of each image.

2. The method according to claim 1, wherein: The digital human image template parameter file is a general digital human model with a unified bone structure and interface definition. The unified bone structure includes the predefined positions and names of the main human bone nodes, and the interface definition includes the standardized parameter interfaces for appearance features.

3. The method according to claim 1, wherein The method for generating the affine transformation matrix is: Solve the optimal transformation matrix by the least - squares method to minimize the mapping error between the key point set of the source digital human image and the corresponding key points of the standard reference coordinate system.

4. The method according to claim 1, wherein: The resource caching mechanism manages cached data through a unified resource ID index, supports asynchronous loading and frame - by - frame batch loading, and maintains a dynamic resource cache pool, and adopts the LRU strategy to eliminate unused resources.

5. The method according to claim 1, wherein: The key point sequence smoothing processing adopts the exponential moving average algorithm, and its recurrence formula is: (3); Wherein, is the position of the key point after smoothing; is the frame index; is the smoothing coefficient; is the position of the original key point in the current frame; is the position of the key point after smoothing in the previous frame.

6. The method according to claim 1, wherein The anomaly compensation strategy includes: When a key point is missing, preferentially fill it with the coordinates smoothed in the previous frame; If multiple consecutive frames are missing, call the preset static neutral expression coordinates in the template file as replacement values.

7. The method according to claim 1, wherein: The hierarchical rendering engine splits the digital human image into at least a basic image layer, a clothing layer, and an expression layer. Each layer is superimposed through a preset alignment method to form a complete image, and the switching of different digital human characters is realized by replacing one or more of the layers.

8. The method according to claim 1, wherein: It further includes a lip - sync step based on voice driving. Through phoneme recognition and predefined lip - shape mapping relationships, convert the audio input into the key point sequence of the mouth of the target digital human image, and generate coordinated facial movements in combination with the default expression curve.

9. A system for rapid switching and reuse of digital human images, wherein: The system runs to execute the method according to any one of claims 1 - 8: The system includes: A template pre - processing module for parsing each digital human image into a template parameter file; The key point alignment module is used to calculate the affine transformation matrix of each digital human image relative to the standard reference coordinate system based on the template parameter file, so as to realize the standardized alignment of the key point set; The driving mapping module is used to map the key point data of the driving source to the coordinate system of the target digital human image through affine transformation to generate a target key point set; The resource cache management module is used to maintain a dynamic resource cache pool, manage texture and template data through the LRU policy, and support asynchronous loading and reuse; The hierarchical rendering engine module is used to divide the digital human face into multiple independent layers according to the target key point set and the reused resources, synthesize images in a preset order, and output continuous animation frames in combination with the smoothing process and the anomaly compensation mechanism; The interface adaptation middleware module is used to shield the underlying platform differences and realize the unified scheduling of cross-terminal resolutions, coordinate systems and rendering pipelines through dynamically loading the adaptation sub-module.

10. The system according to claim 9, wherein: It further includes a voice analysis module, which is used to perform phoneme recognition on the audio data and generate a mouth key point sequence corresponding to the target digital human image when the driving source is an audio input.

Citation Information

Patent Citations

  • Dynamic picture loading method and device, storage medium and terminal equipment

    CN111292387A

  • Digital human image design method based on human body posture consistency and texture mapping

    CN116704097A

  • Digital human video generation method and device, electronic equipment and storage medium

    CN119729145A

  • User operation guiding method based on digital human and intelligent terminal

    CN120010664A

  • Methods and systems for forming personalized 3D head and facial models

    US11417053B1

Cited By

  • Digital human rendering method and device, storage medium and program product

    CN120766707A

  • Digital character animation control method, modeling method, equipment and storage medium

    CN121746555A