A digital human real-time interaction method based on lip synchronization
Patent Information
- Application Number
- CN202610958030.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-09-25
AI Technical Summary
在政务云部署环境中,服务器通常配置有数十个虚拟网络接口,同时现有方案通常在用户发起连接请求后才启动ICE流程,每次连接均需要承受完整的流程耗时,严重影响用户交互体验
[0027]由于采用了上述的技术方案,本发明与现有技术相比,具有以下的优点和积极效果:
Smart Images

Figure CN122816459A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of real-time digital human interaction technology, and in particular to a real-time digital human interaction method based on lip-syncing. Background Technology
[0002] With the development of artificial intelligence and real-time communication technologies, the application scope of virtual digital humans in scenarios such as government services, public service halls, and exhibition hall guidance is constantly expanding. Introducing virtual digital human technology into a comprehensive elderly care service platform can effectively improve the intuitiveness of data broadcasting and the naturalness of human-computer interaction, adapting to the needs of multiple scenarios such as data broadcasting in the supervision hall and policy consultation at service windows. However, existing real-time digital human interaction systems have the following shortcomings in practical deployment and application:
[0003] Firstly, real-time audio and video transmission solutions based on the WebRTC protocol require an ICE candidate address collection process when establishing a peer-to-peer connection. This process traverses all network interfaces of the server. In government cloud deployment environments, servers are typically configured with dozens of virtual network interfaces. Furthermore, existing solutions usually only initiate the ICE process after a user initiates a connection request. Each connection incurs the full process time, severely impacting the user experience.
[0004] Secondly, most existing digital human-driven solutions only support a single speaking mode, that is, only converting text into speech and then driving lip-syncing animation. They do not provide a marking and arrangement mechanism for inserting custom actions during speech playback. When the digital human needs to perform body movements such as nodding or waving in coordination with the broadcast content, it cannot achieve precise timing synchronization between the action and the speech content, resulting in insufficient realism of the interaction.
[0005] Third, when the deep learning-based lip-sync model performs inference for the first time, it needs to complete initialization operations such as CUDA kernel compilation and memory allocation, which results in the inference latency of the first frame being much higher than that of subsequent frames, causing the first screen to stutter. At the same time, when the digital human is in a silent waiting state, the existing solution still performs complete model inference frame by frame to maintain the smoothness of the screen, causing unnecessary waste of computing resources. Summary of the Invention
[0006] The technical problem to be solved by this invention is to provide a real-time digital human interaction method based on lip-sync, which can improve the digital human interaction response and the synchronization between voice and action while reducing resource consumption, and effectively adapt to the high naturalness real-time digital human interaction needs of scenarios such as government service and elderly care service integrated platforms.
[0007] The technical solution adopted by this invention to solve its technical problem is: to provide a real-time digital human interaction method based on lip-sync, applied to a server, including the following steps:
[0008] In response to a client's connection request, a pre-created WebRTC peer connection is invoked to establish a session channel with the client;
[0009] Receive the broadcast text sent by the client through the session channel, and split the received broadcast text into broadcast text fragments bound to action type identifiers;
[0010] Speech synthesis is performed based on the broadcast text segments to form an audio frame sequence, and an action switching signal is generated after the speech synthesis of each broadcast text segment is completed.
[0011] Feature extraction is performed on the audio frame sequence, and the extracted audio features are fed together with the face reference frame image into the lip-sync model for inference to obtain face region image frames containing the dynamic lip movements of the digital human.
[0012] In response to the action switching signal, the action frame data corresponding to the action type identifier after the switch is obtained, wherein the action frame data includes the background frame image and face coordinates;
[0013] The face region image frames are composited into the corresponding positions of the face coordinates in the background image to form an image frame sequence;
[0014] The generated audio and image frame sequences are aligned to form a digital human video stream, which is then sent to the client via a session channel.
[0015] Furthermore, WebRTC peering connections are created when the server starts up, and each WebRTC peering connection includes pre-collected ICE candidate addresses and pre-generated SDP offers.
[0016] Furthermore, when collecting ICE candidate addresses, the server's main network interface address is obtained by probing external addresses via UDP sockets as the ICE candidate address, and mDNS type ICE candidate addresses are skipped.
[0017] Furthermore, when the number of unused WebRTC peer connections is less than a preset threshold, new WebRTC peer connections are created to supplement them.
[0018] Furthermore, the broadcast text sent by the client contains action instructions; the received broadcast text is split into broadcast text segments bound to action type identifiers, including the following steps:
[0019] Extract action instructions from the received broadcast text;
[0020] The broadcast text is divided into multiple broadcast text segments and each segment is associated with an action type identifier. Broadcast text segments marked with action instructions are associated with the action type identifier corresponding to the action instructions, while broadcast text segments without action instructions are associated with the default action type identifier.
[0021] Furthermore, the broadcast text fragments are mapped to predefined action configuration information through action type identifiers, whereby the action configuration information includes action name, action type number, and action data storage path.
[0022] Furthermore, when no broadcast request is received or the action type is identified as silent, a pre-generated sequence of digital human closed-mouth face frames is retrieved to generate a digital human video stream, which is then sent to the client via the session channel.
[0023] Furthermore, the pre-created WebRTC peer connections, face reference frame images, and pre-generated digital human closed-mouth face frame sequences are stored in their respective static caches. When multiple clients are connected, each session channel shares the static cache.
[0024] Furthermore, a TTS engine is used to synthesize the spoken text segments into an audio stream. Each TTS engine adapts its speech synthesis interface by inheriting a uniformly defined abstract base class. The abstract base class includes the conversion interface from spoken text segments to audio streams, audio frame sequence management, and event callback mechanism.
[0025] Furthermore, when extracting features from the audio stream, the corresponding feature extraction algorithm is selected based on the lip-sync model.
[0026] Beneficial effects
[0027] By adopting the above-mentioned technical solution, the present invention has the following advantages and positive effects compared with the prior art:
[0028] This invention reduces connection establishment latency from seconds to milliseconds through WebRTC connection pool pre-creation mechanism and network interface filtering optimization. Users can see the digital human screen almost instantly after initiating an interaction request, significantly improving the smoothness of the interaction experience.
[0029] This invention enables the insertion of action commands at any position in the text through an action mark parsing mechanism for the broadcast text and an action signal triggering mechanism associated with the speech synthesis queue. The digital human can perform custom actions such as nodding and waving at precise times during the speech broadcast, making action switching more timely and smooth.
[0030] This invention utilizes a class-level static caching resource-sharing mechanism, allowing multiple concurrent sessions to share the same digital human image data, action frame sequences, and silent frame caches. This optimizes memory consumption from linearly increasing with the number of sessions to a constant level, significantly reducing memory usage in multi-concurrent session scenarios.
[0031] This invention eliminates the latency spike during the first inference of the lip-sync model through silent frame pre-computation and model warm-up mechanisms, effectively reducing the first frame output latency. In the silent state of the digital human, unnecessary GPU inference computations are avoided by reading the pre-computed silent frame buffer, significantly reducing GPU utilization in the silent state.
[0032] This invention, through a unified speech synthesis abstraction layer design, supports plug-and-play access to multiple TTS engines. During runtime, the TTS engine can be dynamically switched according to scenario requirements without modifying the core lip-sync driver code, enabling rapid engine switching. Attached Figure Description
[0033] Figure 1 This is a flowchart illustrating an embodiment of the present invention;
[0034] Figure 2 This is a schematic diagram of the real-time digital human interaction system architecture according to an embodiment of the present invention;
[0035] Figure 3 This is a schematic diagram of the overall process of the real-time digital human interaction method according to an embodiment of the present invention;
[0036] Figure 4 This is a schematic diagram of the WebRTC connection pool pre-creation and connection acquisition process according to an embodiment of the present invention;
[0037] Figure 5 This is a schematic diagram illustrating the process of action marker parsing and voice action synchronization according to an embodiment of the present invention. Detailed Implementation
[0038] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Furthermore, it should be understood that after reading the teachings of this invention, those skilled in the art can make various alterations or modifications to the invention, and these equivalent forms also fall within the scope defined by the appended claims.
[0039] The abbreviations used in this implementation are explained below:
[0040] WebRTC (Web Real-Time Communication) – Real-time communication for web pages
[0041] ICE (Interactive Connectivity Establishment) – Interactive connection establishment
[0042] SDP (Session Description Protocol)
[0043] TTS (Text-To-Speech) – Text to Speech
[0044] PCM (Pulse Code Modulation)
[0045] CUDA (Compute Unified Device Architecture) – A unified computing device architecture
[0046] mDNS (Multicast DNS) – Multicast Domain Name System
[0047] GPU (Graphics Processing Unit)
[0048] UDP (User Datagram Protocol)
[0049] The embodiments of the present invention relate to a real-time digital human interaction method based on lip-sync, applied to a server, such as... Figure 1 As shown, it includes the following steps:
[0050] In response to a client's connection request, a pre-created WebRTC peer connection is invoked to establish a session channel with the client;
[0051] Receive the broadcast text sent by the client through the session channel, and split the received broadcast text into broadcast text fragments bound to action type identifiers;
[0052] Speech synthesis is performed based on the broadcast text segments to form an audio frame sequence, and an action switching signal is generated after the speech synthesis of each broadcast text segment is completed.
[0053] Feature extraction is performed on the audio frame sequence, and the extracted audio features are fed together with the face reference frame image into the lip-sync model for inference to obtain face region image frames containing the dynamic lip movements of the digital human.
[0054] In response to the action switching signal, the action frame data corresponding to the action type identifier after the switch is obtained, wherein the action frame data includes the background frame image and face coordinates;
[0055] The face region image frames are composited into the corresponding positions of the face coordinates in the background image to form an image frame sequence;
[0056] The generated audio and image frame sequences are aligned to form a digital human video stream, which is then sent to the client via a session channel.
[0057] WebRTC peer connections are created when the server starts. Upon system startup, the connection pool management module is initialized, pre-creating a preset number of WebRTC peer connection objects. The number of pre-created connections can be dynamically adjusted based on server performance and expected concurrency, ranging from 1 to 10.
[0058] Each connection object completes ICE candidate address collection and SDP (Session Description Protocol) offer generation. Simultaneously, the ICE candidate collection process is optimized for network interface filtering: the server's main network interface IP address is determined by creating a UDP (User Datagram Protocol) socket connection to the external address; the ICE candidate collection function is modified to return only the main network interface address, skipping mDNS (Multicast DNS) type candidate addresses. The connection pool asynchronously monitors the number of available connections in the background and automatically replenishes new pre-built connections when the number falls below a preset threshold.
[0059] When a client initiates a connection request, a pre-built connection object that has completed ICE candidate collection is retrieved from the connection pool, and its associated SDP Offer is returned directly to the client without waiting for the ICE candidate collection process. Upon receiving the SDP Offer, the client generates an SDP Answer and sends it back. The server then sets the remote description information to complete the establishment of the WebRTC peer connection. After the connection is established, the connection pool asynchronously creates new pre-built connections in the background to replenish the available number.
[0060] The system receives and parses text messages to be played from clients via data channels or HTTP interfaces. Action instructions can be marked using double brackets within the text. The parsing result is an ordered list of text-action type pairs. An action marker parser uses regular expressions to match double bracket markers (formatted as "[[Action Name]]") in the text, segmenting it into multiple fragments. Each fragment is associated with an action type identifier. Action type identifiers are mapped to a predefined action configuration table, which contains the correspondence between action names, action type numbers, and action data paths. Text fragments without action markers use the default speaking action type. Each action marker is bound to the following text fragment, and the system synchronously triggers the corresponding action when the corresponding audio for that fragment begins output or reaches its corresponding timestamp.
[0061] The parsed text segments are sequentially fed into the TTS engine for streaming speech synthesis. The speech synthesis process outputs PCM (Pulse Code Modulation) audio frames in a streaming manner, with each audio frame carrying a corresponding action type identifier and event marker. Once the speech synthesis of a text segment is complete, an action switching signal associated with that segment is triggered. The TTS engine can be replaced with other speech synthesis services that support streaming output.
[0062] The system supports multiple TTS engine implementations, all of which can be based on a unified abstract base class. This abstract base class defines the text-to-audio stream conversion interface, audio frame queue management, and event callback mechanism. Each engine adapts by inheriting from the abstract base class and implementing a unified speech synthesis interface. The event callback mechanism, in particular, is designed to unify event notification rules across different TTS engines, decouple TTS synthesis logic from upper-layer business logic, and pre-define standardized event triggering interfaces within the abstract base class. In this solution, this callback mechanism primarily includes:
[0063] (1) Single text fragment synthesis completion callback: Triggered when all text fragments bound to an action type are synthesized, the action type identifier corresponding to the fragment is passed as a parameter;
[0064] (2) Single audio frame generation callback: Triggered when a PCM audio frame is synthesized, the audio frame data and the corresponding bound action type identifier are synchronously passed to the audio feature extraction module;
[0065] (3) Error and full synthesis completion callback: Triggered when TTS synthesis fails or all text to be broadcast is synthesized, notifying the upper module to perform error downgrade or switch to digital human silent state.
[0066] PCM audio frames are obtained from the synthesized audio frame queue and fed into the audio feature extraction module. The feature extraction module selects the appropriate feature extraction algorithm based on the type of lip-sync model used; different lip-sync models correspond to different audio feature representations. Feature extraction employs a sliding window mechanism to preserve the audio context of consecutive frames to improve feature continuity. The extracted audio features, along with pre-loaded face reference frame images, are fed into the lip-sync model for inference. The model outputs a face region image containing accurate lip movements. The lip-sync model can be replaced with other deep learning-based lip-driven models.
[0067] When the action switching signal arrives, if the switched action type is identified as the default speaking action, the face region image frame obtained from lip-sync inference is composited onto the corresponding coordinate position of the default background frame; if the switched action type is identified as a custom action, the rendering module reads the corresponding custom action frame sequence from the cache. The custom action frame sequence data includes the background frame image, face region image, face coordinates, and accompanying audio. If there is corresponding audio playback content, the face region image containing dynamic lip movements obtained from the lip-sync model is composited onto the corresponding coordinate position of the background frame. Simultaneously, the playback audio and accompanying audio are mixed or synthesized as needed to form a video stream; if there is no corresponding audio playback content, the pre-loaded corresponding action frame sequence is retrieved and output directly. After the action playback is complete, it automatically switches back to the default speaking or silent state.
[0068] Custom action frame sequences are preloaded into a class-level static cache from the path specified in the configuration file during system startup. The following class-level cache resources are shared by all concurrent sessions: (a) a pre-computed silent frame sequence cache; (b) a cache of custom action frame sequences and audio data; and (c) a list of face reference frames and coordinate data for the digital human's basic image. Each session maintains independent runtime state information (including current speaking state, action state, TTS queue pointer, and WebRTC connection object), and resource data is not repeatedly loaded.
[0069] Example 1: Digital Human Broadcasting Scenario in Elderly Care Service Data Cockpit
[0070] This embodiment uses the data dashboard of a municipal civil affairs bureau's comprehensive supervision platform for elderly care services as an example to illustrate the complete execution process of the method of the present invention. This dashboard is deployed in the lobby of the civil affairs bureau's elderly care service supervision center and is used to display the operational indicators of elderly care service institutions within its jurisdiction in real time.
[0071] The system is deployed on a government cloud server, with the front end being a large-screen display terminal in the supervision center lobby, accessed via a browser to a digital human interactive page. For example... Figure 2 As shown, it includes:
[0072] Connection Pool Management Module 100: Corresponds to the pre-built connection optimization logic. It is responsible for pre-creating a specified number of WebRTC peer connections during the system startup phase, completing ICE candidate filtering and collection, SDP Offer generation, and audio / video track initialization. During runtime, it dynamically monitors the number of available connections and automatically supplements new pre-built connections when the number is lower than the preset threshold. When a client initiates a connection request, it directly allocates an available pre-built connection to achieve millisecond-level session establishment.
[0073] Action tag parsing module 200: Corresponds to action orchestration logic. It receives the broadcast text input by the client and the interactive text output by the large language model. It extracts the double bracket action tags in the text through regular expression matching, splits the text into an ordered list of text-action type pairs, automatically associates unmarked fragments with default speaking actions, and automatically matches the corresponding parameters in the predefined action configuration table.
[0074] Speech synthesis module 300: Corresponding to the TTS abstraction layer design, it encapsulates multiple TTS engines based on a unified abstract base class (including the first speech synthesis engine 301, the second speech synthesis engine 302, and the third speech synthesis engine 303, which correspond to different timbres and streaming TTS services from different manufacturers). It receives a list of text-action type pairs and performs streaming speech synthesis. It binds a corresponding action identifier to each output PCM audio frame and triggers audio frame output events and text segment synthesis completion events (i.e., action switching signals) through a preset callback mechanism. It supports dynamic switching of TTS engines at runtime without modifying the core process.
[0075] Audio feature extraction module 400: Corresponds to audio feature adaptation logic, receives PCM audio frames output by the speech synthesis module, automatically matches the corresponding feature extraction algorithm according to the lip-sync model type configured in the current system, uses a sliding window to retain contextual audio information, and outputs audio features that meet the model input requirements, which can adapt to the input specifications of different deep learning lip-sync models.
[0076] The lip-sync rendering module 500 is the core rendering execution module, corresponding to the audio and video generation and output logic. It has three core functions: ① In speaking mode, it receives audio features, combines them with cached face reference frames to perform lip-sync inference, generates a matching lip-sync face region, and synthesizes it into a background frame to generate a complete video frame; ② When receiving an action switching signal, it retrieves a pre-stored custom action frame sequence from the cache for playback, achieving temporal synchronization between action and speech; ③ In silent mode, it directly retrieves a pre-stored silent frame sequence and outputs it in a loop, without performing GPU inference. Finally, the audio and video frames are aligned by timestamp and sent to the client via the WebRTC media track.
[0077] Cache management module 600: Corresponds to multi-session resource sharing logic, managing three types of class-level static caches, which are shared by all concurrent sessions without repeated loading: ① Silent frame 601, which stores the pre-calculated closed state frame sequence; ② Action frame 602, which stores the frame sequence, accompanying audio, and coordinate parameters of each custom action; ③ Image data cache 603, which stores basic image data such as digital human face reference frames, default background, and coordinate parameters, optimizing memory consumption to a constant level and improving the system's concurrent carrying capacity.
[0078] When interacting with digital humans, the system processing procedure is as follows: Figure 3As shown, it mainly includes the following steps.
[0079] 1. System startup phase
[0080] When the server starts, it performs the following initialization operations:
[0081] (a) The connection pool management module 100 is initialized, pre-creating two WebRTC connection objects. The creation process for each connection object includes: creating a WebRTC peer connection instance and configuring the signaling server address; performing network interface filtering, determining the server's main network interface address through socket probing, and modifying the ICE candidate collection function to only return this address; creating audio and video media tracks, completing ICE candidate collection and SDP Offer generation. The entire pre-creation process is completed asynchronously in the background and does not affect system services.
[0082] (b) Cache management module 600 performs resource preloading: loads the face reference frame sequence of the digital human image and stores it in the image data cache 603; performs a complete lip-sync model inference using silent audio input, generates a silent frame sequence of closed mouth state and stores it in silent frame cache 601, and completes model warm-up to eliminate subsequent first frame inference stuttering; reads the custom action list (such as "nodding" and "waving") from the action configuration file, loads the frame sequence and audio data of each action and stores it in the action frame cache 602.
[0083] (c) The speech synthesis module 300 initializes the default TTS engine and establishes a connection with the TTS service.
[0084] 2. User Connection Phase
[0085] like Figure 4 As shown, the browser on the large-screen terminal in the monitoring center lobby opens the data dashboard page and initiates a WebRTC connection request. The connection pool management module 100 retrieves an available connection object (containing a completed SDP Offer) from the pre-built connection pool and returns the SDP Offer to the browser. The browser generates an SDP Answer and sends it back to the server, completing the establishment of the WebRTC peer connection. The entire connection establishment process takes approximately 150 milliseconds, with virtually no noticeable delay for the user. The connection pool asynchronously creates new pre-built connections in the background to replenish the pool.
[0086] 3. Digital Human Broadcasting Stage
[0087] like Figure 5As shown, the comprehensive supervision platform for elderly care services sends a broadcast text to the digital human via an HTTP interface: "[[Nodding]] Hello leaders, welcome to view today's elderly care service operation status. [[Waving]] Now broadcasting the elderly care service data for your jurisdiction: There are currently 236 registered elderly care institutions, with a total of 18,520 beds, 15,680 elderly residents, and an overall occupancy rate of 84.7%. Two new institutions were registered today. [[Nodding]] All service indicators are operating normally, and there have been no major safety incidents."
[0088] Action tag parsing module 200 uses regular expressions to match the [[]] tag, parsing the text into the following fragments:
[0089] Segment 1: "Good morning, leaders. Welcome to today's report on the operation of elderly care services." Related action type: Nodding (No. 3)
[0090] Segment 2: "Now broadcasting the elderly care service data for our jurisdiction: There are currently 236 registered elderly care institutions with a total of 18,520 beds, 15,680 elderly residents, and an overall occupancy rate of 84.7%. Two new institutions were registered today." Related action type: Waving (No. 4)
[0091] Segment 3: "All service indicators are operating normally, and there have been no major safety incidents." Related action type: Nodding (Number 3)
[0092] Each segment is sequentially fed into the speech synthesis module 300 for streaming speech synthesis. The TTS engine streams PCM audio frames. After the audio frames enter the queue, the audio feature extraction module 400 extracts audio features and sends them to the lip-sync rendering module 500 for lip-sync inference, generating a face region image containing lip shape changes. After being synthesized into the background frame, the image is output to the browser via the WebRTC video track.
[0093] When the speech synthesis of segment 1 is completed and the speech output of segment 2 begins, the system triggers the action switching signal bound to segment 2. The rendering module reads the frame sequence of the "waving" action from the action frame buffer 602 and plays it, so that the waving action and the speech of segment 2 are presented synchronously. After the waving action is played, it automatically switches back to the speaking state to continue lip-syncing.
[0094] 4. Silent waiting phase
[0095] After the broadcast is completed, the digital human enters a silent state. The lip-sync rendering module 500 reads the pre-computed closure frames sequentially from the silent frame buffer 601, does not perform GPU inference calculations, and continuously outputs silent video frames to maintain smooth video playback.
[0096] Example 2: Dialogue and Interaction Scenario for Elderly Care Policy Consultation
[0097] This embodiment, based on Embodiment 1, illustrates the dialogue and interaction process between the digital human and visiting members of the public. This scenario is applicable to policy consultations at civil affairs service windows or elderly care service hotlines.
[0098] The system integrates an LLM (Large Language Model) dialogue engine. When visitors ask the digital human via voice or text, "My father is 80 years old this year, what kind of pension subsidies can he apply for?"
[0099] 1. The client sends text to the server via the WebSocket data channel.
[0100] 2. The server sends the question to the LLM dialogue engine, which combines the policy knowledge base and subsidy database query results of the comprehensive supervision platform for elderly care services to generate the answer text.
[0101] 3. The streaming output text of the LLM is parsed by the action tag parsing module 200 and then sent to the speech synthesis module 300 for streaming speech synthesis. In this embodiment, to obtain a more natural speech effect, the TTS engine is switched from the default engine to a speech synthesis engine with a more natural tone during runtime. The switching process only requires modifying the TTS engine type parameter; the core lip-syncing process does not need to be changed.
[0102] 4. Following the process in step S5, the audio feature extraction module 400 and the lip-sync rendering module 500 extract audio features from the PCM audio frame output by speech synthesis and send them into the lip-sync model for inference to generate video frames containing accurate lip shapes, which are then output to the client in real time via WebRTC.
[0103] 5. The digital avatar, using natural voice and lip-syncing, answers the public: "According to the city's elderly care service subsidy policy, seniors aged 80 and above can apply for a high-age allowance. If, after a capacity assessment, they are identified as disabled, they can also apply for an elderly care subsidy. We suggest you bring the senior's ID card to the elderly care service center in your community to apply."
[0104] In this embodiment, the runtime switching of the TTS engine demonstrates the advantages of the unified abstraction layer design of this invention. Different TTS engines with different timbres and expressiveness can be selected for different scenarios, while the system's audio feature extraction, lip-sync inference, and WebRTC transmission processes remain unchanged.
[0105] Example 3: Multi-terminal concurrent deployment scenario
[0106] This embodiment illustrates the resource optimization effect of the present invention when multiple terminals of the comprehensive supervision platform for elderly care services are deployed simultaneously.
[0107] A municipal civil affairs bureau deployed three data dashboard terminals in the lobby of the elderly care service supervision center, its offices, and the district / county civil affairs bureaus, all connected to a digital human system.
[0108] 1. The connection pool management module 100 pre-creates two WebRTC connections when the system starts. When the first terminal connects, it obtains a connection from the pool (within 200ms), and the pool automatically replenishes it in the background. When the second and third terminals connect, if there is an available connection in the pool, it directly obtains the connection; otherwise, it waits for the background to complete the creation.
[0109] 2. All resource data in the three session-shared cache management module 600: image data cache 603, silent frame cache 601, and action frame cache 602. Without the sharing mechanism, each of the three sessions needs to load all resource data, and memory consumption increases linearly with the number of sessions; with the sharing mechanism, resource data is loaded only once, resulting in significant memory savings.
[0110] 3. The runtime state information maintained independently by each session only includes the TTS queue pointer, the current action state, and the WebRTC connection object, and the memory overhead of the independent state of each session is very small.
[0111] 4. When all three terminals are in a silent state, the lip-sync rendering module 500 reads the pre-computed frame from the silent frame buffer 601. The GPU inference computation is zero. The lip-sync model inference is only started when a terminal triggers voice broadcast.
Claims
1. A real-time digital human interaction method based on lip-sync, applied to a server, characterized in that, Includes the following steps: In response to a client's connection request, a pre-created WebRTC peer connection is invoked to establish a session channel with the client; Receive the broadcast text sent by the client through the session channel, and split the received broadcast text into broadcast text fragments bound to action type identifiers; Speech synthesis is performed based on the broadcast text segments to form an audio frame sequence, and an action switching signal is generated after the speech synthesis of each broadcast text segment is completed. Feature extraction is performed on the audio frame sequence, and the extracted audio features are fed together with the face reference frame image into the lip-sync model for inference to obtain face region image frames containing the dynamic lip movements of the digital human. In response to the action switching signal, the action frame data corresponding to the action type identifier after the switch is obtained, wherein the action frame data includes the background frame image and face coordinates; The face region image frames are composited into the corresponding positions of the face coordinates in the background image to form an image frame sequence; The generated audio and image frame sequences are aligned to form a digital human video stream, which is then sent to the client via a session channel.
2. The real-time interaction method according to claim 1, characterized in that, WebRTC peering connections are created when the server starts up, and each WebRTC peering connection includes a pre-collected ICE candidate address and a pre-generated SDPOffer.
3. The real-time interaction method according to claim 2, characterized in that, When collecting ICE candidate addresses, the server's main network interface address is obtained by probing external addresses via UDP sockets as the ICE candidate address, and ICE candidate addresses of type mDNS are skipped.
4. The real-time interaction method according to claim 2, characterized in that, When the number of unused WebRTC peer connections falls below a preset threshold, new WebRTC peer connections are created to supplement them.
5. The real-time interaction method according to claim 1, characterized in that, The broadcast text sent by the client contains action instructions; the received broadcast text is split into broadcast text segments bound to action type identifiers, including the following steps: Extract action instructions from the received broadcast text; The broadcast text is divided into multiple broadcast text segments and each segment is associated with an action type identifier. Broadcast text segments marked with action instructions are associated with the action type identifier corresponding to the action instructions, while broadcast text segments without action instructions are associated with the default action type identifier.
6. The real-time interaction method according to claim 5, characterized in that, The broadcast text snippets are mapped to predefined action configuration information through action type identifiers. The action configuration information includes action name, action type number, and action data storage path.
7. The real-time interaction method according to claim 5, characterized in that, When no broadcast request is received or the action type is identified as silent, a pre-generated sequence of digital human closed-mouth face frames is retrieved to generate a digital human video stream, which is then sent to the client via the session channel.
8. The real-time interaction method according to claim 7, characterized in that, The pre-created WebRTC peer connections, face reference frame images, and pre-generated digital human closed-mouth face frame sequences are stored in their respective static caches. When multiple clients are connected, each session channel shares the static cache.
9. The real-time interaction method according to claim 1, characterized in that, A TTS engine is used to synthesize spoken text segments into an audio stream. Each TTS engine adapts its speech synthesis interface by inheriting a uniformly defined abstract base class. The abstract base class includes the conversion interface from spoken text segments to audio stream, audio frame sequence management, and event callback mechanism.
10. The real-time interaction method according to claim 1, characterized in that, When extracting features from an audio stream, the corresponding feature extraction algorithm is selected based on the lip-sync model.