Real-time interactive digital human system supporting high concurrency and implementation method thereof
Through innovative means such as memory mapping technology and dynamic priority algorithms, a real-time interactive digital human system that supports high concurrency is built, which solves the problems of high response latency, low resource utilization and insufficient scalability, and achieves low latency, high throughput and strong robust interaction capabilities.
Patent Information
- Application Number
- CN202510670640.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-23
AI Technical Summary
In the high concurrency scenarios, existing digital human interaction systems have problems such as high response latency, low resource utilization and insufficient scalability, which is difficult to meet the performance needs of large-scale users' concurrency.
Through memory mapping technology, lightweight model instance pool is built to realize multi-thread sharing of the same model weight; dynamic priority algorithm and lock-free queue scheduling tasks are adopted, and streaming processing mechanism of asynchronous pipeline microservices is combined; Kubernetes elastic scaling capacity and edge node deployment strategies are integrated, service instances are dynamically adjusted and cross-region request routing is optimized; clients adapt different terminal performance and network conditions through adaptive rendering strategies and adaptive bit rate algorithms; audio and video hierarchical synchronization calibration mechanism is introduced to ensure interaction consistency.
It realizes low latency, high throughput and strong robust interaction capabilities, reduces resource waste caused by repeated loading of models, optimizes the end-to-end delay of multi-module collaborative processing, improves the system's elastic expansion and cross-region deployment capabilities, and supports the rendering needs of diversified terminal devices.
Smart Images

Figure CN120179081A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital human interaction, and particularly relates to a real-time interactive digital human system supporting high concurrency and its implementation method. Background Art
[0002] With the rapid development of artificial intelligence and computer graphics technologies, virtual digital humans have become an important research direction in the field of human-computer interaction. Its core goal is to build a highly realistic digital human image through technologies such as speech synthesis, natural language processing, and real-time animation rendering to support real-time interactions in scenarios such as intelligent customer service, online education, and virtual live streaming. However, large-scale user concurrency scenarios pose severe challenges to the system architecture: on the one hand, users expect digital humans to have low-latency real-time response capabilities; on the other hand, traditional architectures have significant bottlenecks in resource management, multi-module collaboration, and scalability.
[0003] Existing technologies mainly achieve interactions through centralized cloud processing and streaming media transmission. However, limited by problems such as redundant model loading, low asynchronous scheduling efficiency, and cumulative network latency, it is difficult to meet the performance requirements in high-concurrency scenarios, specifically including: existing solutions usually independently load a complete AI model for each user session, resulting in linear growth of memory and computing resource consumption with the number of concurrent users; modules such as speech recognition, semantic understanding, speech synthesis, and animation rendering use independent asynchronous processing, lacking a global timestamp synchronization mechanism, which easily leads to audio-video desynchronization and cumulative latency; existing systems mostly rely on centralized cloud processing, making it difficult to reduce cross-regional latency through edge computing and lacking an elastic scaling mechanism. Therefore, based on the above problems, the present invention proposes a real-time interactive digital human system supporting high concurrency and its implementation method. Summary of the Invention
[0004] Technical Objectives In order to solve the above problems, the objective of the present invention is to provide a real-time interactive digital human system supporting high concurrency and its implementation method, aiming to solve problems such as high response latency, low resource utilization, and insufficient scalability in the real-time interactive digital human system in high-concurrency scenarios, and to achieve interactive capabilities with low latency, high throughput, and strong robustness through innovative architecture design. Its core objectives include reducing resource waste caused by repeated model loading, optimizing the end-to-end latency of multi-module collaborative processing, enhancing the system's elastic expansion and cross-regional deployment capabilities, and adapting to the rendering requirements of diverse terminal devices to support the efficient operation of large-scale real-time interactive scenarios such as online education and cloud customer service.
[0005] Technical Solutions To achieve the above object, the present invention provides a high-concurrency supported real-time interactive digital human system and its implementation method. The system constructs a lightweight model instance pool through memory mapping technology to enable multiple threads to share the same model weights, thereby reducing memory occupancy and initialization overhead; adopts a dynamic priority algorithm and a lock-free queue to schedule tasks, and combines the streaming processing mechanism of asynchronous pipeline microservices to shorten the end-to-end latency; integrates Kubernetes elastic scaling and edge node deployment strategies to dynamically adjust service instances and optimize cross-regional request routing; the client adapts to different terminal performances and network conditions through an adaptive rendering strategy and an adaptive bitrate algorithm; and also introduces an audio-visual hierarchical synchronization calibration mechanism to ensure interactive coherence. It systematically solves the resource redundancy, latency accumulation, and scalability bottlenecks of the prior art, and realizes efficient real-time interaction under high concurrency.
[0006] In a first aspect, the present invention provides a high-concurrency supported real-time interactive digital human system, including: A model instance pool module for loading deep learning model weights into a shared memory area through memory mapping technology for multiple parallel inference threads to share the same model weights; A multi-thread scheduling module that schedules user requests using a lock-free queue and a dynamic priority algorithm, and the dynamic priority algorithm calculates the priority based on the task waiting time and urgency; An asynchronous pipeline processing module composed of multiple decoupled microservices, including speech recognition, semantic understanding, speech synthesis, animation rendering, and streaming media push modules, which are connected in series through an asynchronous message queue to achieve end-to-end streaming processing; A client SDK that supports an adaptive rendering strategy and dynamically switches between local rendering mode or remote rendering mode according to the terminal hardware performance and network conditions; An elastic scaling module that, based on containerized deployment and Kubernetes orchestration, monitors resource loads in real time and automatically adjusts the number of service instances; An audio-visual synchronization calibration module for ensuring that the synchronization error between audio and video frames is lower than a preset threshold through a timestamp alignment algorithm.
[0007] Further, the model instance pool module supports hot-updating the model weights in the shared memory without interrupting the current session. The update process includes the following steps: loading the new model weights into the spare memory area through double buffering technology; switching to the new weight area during the model inference gap and releasing the memory occupied by the old weights; the hot-update process ensures weight consistency through version number verification.
[0008] Further, the dynamic priority algorithm of the multi-thread scheduling module satisfies the following formula: (1); In the formula, is the scheduling priority; and are adjustable parameters; is the waiting time of task i in the queue; is the urgency of session i; The lock-free queue implements task distribution using atomic operations based on a circular buffer.
[0009] This method significantly reduces memory occupancy and initialization latency, solves the bottleneck of linear growth of resource consumption with the number of concurrencies in traditional architectures, and improves system throughput and high-concurrency support capabilities.
[0010] Furthermore, the asynchronous pipeline processing module adopts a cross-model collaborative inference mechanism, combines the output of the large language model with the rules of the pre-set knowledge base to generate conversation responses, and filters the optimal results through a confidence threshold.
[0011] Furthermore, the adaptive rendering strategy of the client SDK includes: in the local rendering mode, the client receives the parameterized animation instructions sent by the server and generates the digital human image in real time in combination with the local personalized configuration; in the remote rendering mode, the client receives the complete audio-visual stream synthesized by the server and dynamically adjusts the bit rate through the adaptive bit rate algorithm; the adaptive rendering strategy also integrates a network jitter prediction model to predict the network state based on historical latency data and switch the rendering mode in advance.
[0012] Furthermore, the elastic scaling module distributes user requests through the weighted round-robin algorithm based on the geographical location and real-time load of the edge nodes, preferentially expands the edge node instances deployed in the same geographical area when the GPU utilization rate exceeds the threshold, and retains the minimum instance pool during scaling to avoid cold start latency.
[0013] This method optimizes resource utilization and network transmission efficiency, reduces cross-regional interaction latency, and at the same time enhances the robustness of the system in scenarios of high load fluctuations and node failures, providing a scalable underlying architecture support for large-scale real-time interactions.
[0014] Furthermore, the audio-visual synchronization calibration module adopts a hierarchical calibration strategy. At the first layer, the audio and video timestamps are synchronized through the global clock. At the second layer, the frame interval is dynamically adjusted based on the audio features and the key points of the video mouth shape. When the synchronization error exceeds the threshold, the interpolation frame filling or frame dropping mechanism is triggered.
[0015] Furthermore, when the WebRTC protocol is unavailable, the streaming media pushing module automatically degrades to a low-latency transmission protocol based on QUIC; the switching process ensures data integrity through forward error correction and redundant coding.
[0016] Further, it further includes a model weight optimization module for dynamically adjusting local parameters of the shared model weight by analyzing user behavior data in real time during the inference process. The updated model weight is: (2); In the formula, is the updated model weight; is the shared weight; is the learning rate decay factor; is the weight correction function; is the Sigmoid function; is the L2 norm of the feature vector F.
[0017] Without increasing additional resource overhead, this method significantly improves the semantic matching accuracy of the digital human response and user interaction satisfaction, while maintaining the global sharing feature of the model weight and avoiding the disadvantages of redundant model loading in traditional personalized solutions.
[0018] Further, it further includes an audio-video synchronization enhancement module for training a lightweight temporal convolutional network using historical audio-video frame sequences to predict audio features and mouth key points of future frames. The future k-frame features are predicted through the following formula: (3); In the formula, is the predicted future k-frame audio-video features; is the temporal convolutional network; is the input historical frame sequence; The dynamic synchronization correction amount is defined as: (4); In the formula, is the dynamic synchronization correction amount; is the correction intensity coefficient; is the index variable; is the predicted feature at time; is the actual feature at time; is the current time L2 norm of the feature vector.
[0019] This method significantly improves the audio-video synchronization accuracy and the anti-interference ability of the system through the predicted frame buffer and dynamic fusion strategy, while ensuring the stability of the end-to-end delay and overcoming the response lag defect of traditional posterior calibration methods in high-concurrency scenarios.
[0020] In the second aspect, the present invention also provides a method for implementing a real-time interactive digital human system supporting high concurrency. The method is based on the system described in the first aspect above and includes: On the server side, the deep learning model weights are loaded into the shared memory area through memory mapping technology for multi-threaded parallel calls. Each thread independently maintains the session state and shares the same model weights; Adopt a dynamic priority algorithm based on a lock-free queue to schedule user requests, and allocate tasks to idle model instances through atomic operations; Decouple the speech recognition, semantic understanding, speech synthesis, and animation rendering modules into independent microservices, and realize streaming processing through an asynchronous message queue. Each module triggers downstream processing based on segmented output to shorten the end-to-end latency; Dynamically switch the rendering mode according to the terminal hardware performance and network conditions. High-performance terminals generate digital human images through local parameterized instructions, and low-performance terminals receive the complete audio and video streams synthesized by the cloud and adjust the transmission bit rate based on the adaptive bit rate algorithm; Based on containerized orchestration and edge node deployment strategies, monitor the server resource load in real time, dynamically scale the service instances, and allocate user requests to the nearest edge node through geographical location routing; Ensure that the synchronization error between audio and video frames is lower than the preset threshold through the timestamp alignment algorithm.
[0021] In a third aspect, the present invention also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by the management platform, it implements the implementation method of the real-time interactive digital human system supporting high concurrency described above.
[0022] The present invention realizes multi-threaded parallel inference and end-to-end streaming processing through a model instance pool shared memory mapping, a dynamic priority scheduling algorithm, and an asynchronous pipeline microservice module; combines an elastic scaling mechanism and edge node deployment to optimize resource allocation, and ensures cross-terminal interaction fluency through a client-side adaptive rendering strategy and a spatio-temporal prediction synchronization algorithm. This solution systematically solves the core problems in traditional technologies such as redundant model loading, audio and video synchronization errors, low resource utilization, and insufficient scalability, significantly reduces the interaction latency, improves the system throughput and robustness, and at the same time supports personalized semantic adaptation and high-precision synchronization in a dynamic network environment, providing an efficient, stable, and economical technical implementation path for large-scale real-time interaction scenarios.
[0023] Beneficial effects By implementing the real-time interactive digital human system supporting high concurrency and its implementation method provided by the present invention, the following technical effects are achieved: (1) This application loads model weights into the shared memory area through memory mapping technology, combines a lock-free queue with a dynamic priority scheduling algorithm, and realizes the efficient reuse of the same model weights by multiple parallel inference instances. This method significantly reduces memory occupancy and initialization latency, solves the bottleneck of linear growth of resource consumption with the number of concurrencies in traditional architectures, and improves the system throughput and high-concurrency support capabilities.
[0024] (2) Through Kubernetes container orchestration and edge node dynamic routing, it realizes the automatic scaling of service instances and the intelligent distribution of cross-regional requests. This method optimizes resource utilization and network transmission efficiency, reduces cross-regional interaction latency, and enhances the robustness of the system in scenarios of high-load fluctuations and node failures, providing a scalable underlying architecture support for large-scale real-time interactions.
[0025] (3) By analyzing user interaction characteristics in real time and dynamically adjusting the local parameters of the shared model weights, it realizes session-level personalized inference. This method significantly improves the semantic matching accuracy of the digital human response and user interaction satisfaction without additional resource overhead, while maintaining the global sharing characteristics of the model weights and avoiding the drawbacks of redundant model loading in traditional personalized solutions.
[0026] (4) By predicting future audio-visual features through a lightweight temporal convolutional network and combining a feedforward error compensation mechanism, it effectively reduces the audio-visual asynchrony problem caused by network jitter or calculation fluctuations. This method significantly improves the audio-visual synchronization accuracy and the anti-interference ability of the system through prediction frame buffering and dynamic fusion strategies, while ensuring the stability of the end-to-end latency and overcoming the response lag defect of traditional posterior calibration methods in high-concurrency scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] To make the above-mentioned high-concurrency supported real-time interaction digital human system and its implementation method of the present invention more obvious and understandable, the drawings required for the specific implementation manners of the present invention will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings without creative efforts.
[0028] Figure 1 It shows the schematic diagram of the system architecture of this application; Figure 2 It shows the schematic diagram of the backend microservice and asynchronous pipeline architecture; Figure 3 It shows the schematic diagram of elastic scaling and edge node scheduling; Figure 4 It shows the sequence diagram of the interaction process; Figure 5Schematic diagram showing the logic of the client SDK rendering strategy selection. Detailed implementation
[0029] Example 1: Provided is a real-time interactive digital human system supporting high concurrency and its implementation method. The system architecture is as Figure 1 shown, specifically including: a model instance pool module for loading deep learning model weights into a shared memory area through memory mapping technology for multiple parallel inference threads to share the same model weights; a multi-threaded scheduling module that uses a lock-free queue and a dynamic priority algorithm to schedule user requests, and the dynamic priority algorithm calculates the priority based on the task waiting time and urgency; an asynchronous pipeline processing module composed of multiple decoupled microservices, including speech recognition, semantic understanding, speech synthesis, animation rendering, and live streaming push modules, which are connected in series through an asynchronous message queue to achieve end-to-end streaming processing; a client SDK that supports an adaptive rendering strategy and dynamically switches between local rendering mode and remote rendering mode according to the terminal hardware performance and network conditions; an elastic scaling module based on containerized deployment and Kubernetes orchestration, which monitors the resource load in real time and automatically adjusts the number of service instances; an audio-video synchronization calibration module for ensuring that the synchronization error between audio and video frames is lower than a preset threshold through a timestamp alignment algorithm. Specifically, it is described as follows.
[0030] I. Model loading logic and multi-threaded scheduling algorithm To achieve the goal of real-time interaction of digital humans in a high-concurrency environment, the system adopts a lightweight inference model instance pool and multi-threaded shared scheduling technology on the server side to efficiently share AI model weights and support parallel processing of multiple user requests. Specifically, it includes: 1. During the system startup process, the model instance pool module deployed at the backend pre-loads lightweight AI models for each key function of the digital human. The loading process uses memory mapping technology, that is, mapping the model weight file to the shared memory area in read-only mode. This method can ensure that all subsequent created model instances reference the same memory area, avoid repeated loading of model weights, thereby reducing memory consumption and initialization latency. During the loading process, the system also performs integrity verification to ensure the correctness of the data. After loading is completed, each model instance logically generates a reference handle, which is used for subsequent inference calculations, and each instance only maintains its own dynamic state, thus achieving state isolation.
[0031] 2. After a user request arrives, the system encapsulates the request data into a task and enters the global asynchronous task queue. To ensure low-latency distribution in a high-concurrency environment, the scheduler uses a lock-free queue to queue tasks. For task i, its scheduling priority is calculated according to the following formula: (1); wherein is the scheduling priority; and are adjustable parameters for balancing waiting time and task urgency; is the waiting time of task i in the queue; is the urgency or historical completion rate of session i, such as the ratio of unresponsive tasks in continuous interactions; The scheduler selects an idle model instance for task processing according to the value, ensuring that each thread calls the shared model function under an independent session state , and its reasoning process is expressed as: (5); wherein is the reasoning output result; is the shared model function; is the independent session state; is the user input data.
[0032] To ensure multi-thread safety, all model instances share this read-only memory area through memory mapping after loading weights. Each thread only maintains an independent session state to ensure data isolation. At the same time, the multi-thread scheduling module uses a lock-free queue to queue tasks, determines the task scheduling priority according to the scheduling priority formula, and selects an idle model instance to allocate a thread for the task, thereby realizing multi-thread high-concurrency reasoning under state isolation.
[0033] Through the above model loading and multi-thread shared scheduling design, the system can achieve efficient reuse of the same AI model weights under high-concurrency conditions, making the memory consumption not increase linearly with the number of concurrencies, and greatly improving the throughput and response speed of the system.
[0034] II. Asynchronous Pipeline and Microservice Module Design The backend microservice and asynchronous pipeline architecture are as Figure 2 shown. The backend system adopts a microservice architecture, and the core functional modules are decoupled through standard interfaces and asynchronous message queues, forming a pipelined parallel processing process. Each module can run independently and can also form an end-to-end data pipeline through a unified session scheduling and task distribution mechanism to achieve low-latency and high-concurrency real-time interaction of digital humans. Therefore, this module mainly includes the following modules and their collaborative working mechanisms: 1. WebSocket access module, responsible for establishing and maintaining a long connection between the client and the server. This module generates a unique session ID during session initialization, ensures the stability of the connection through heartbeat detection, and forwards the voice or text data uploaded by the user to the backend message queue in a pre-defined format for subsequent asynchronous processing by other modules.
[0035] 2. The ASR recognition module is used to implement automatic speech recognition and real-time transcribe the voice data uploaded by the user into text. This module adopts a streaming algorithm and combines short-time energy and zero-crossing rate to detect endpoints. The basic formula is: (6); In the formula, is the short-time energy of the nth frame, which is used to measure the intensity of the voice signal in this frame; is the window length; is the sampling point index within the window; is the amplitude of the voice signal at the mth sampling point within the nth frame; (7); In the formula, is the zero-crossing rate of the nth frame, which reflects the frequency at which the signal passes through zero within this frame; is the sign function, which returns the sign of the sampling point.
[0036] The ASR module gradually outputs the recognition results in a streaming manner and immediately pushes part of the transcribed text to the semantic understanding module through an asynchronous message queue.
[0037] 3. The semantic understanding and dialogue generation module performs natural language understanding on the text output by ASR based on a large language model or a pre-set knowledge base, and generates the response text of the digital human. To reduce latency, this module supports streaming output, that is, the model outputs sentences one by one when generating responses. Its generation process is described by a recursive function: (8); In the formula, is the generation step size, which represents the time interval or text span from the current moment to the next generation moment; is the generation function, which represents the core logic of the semantic understanding and dialogue generation model; is the text sequence generated at the current t moment; is the dialogue context information; is the input text.
[0038] The generated response is immediately passed to the TTS speech synthesis module through the message queue after being segmented.
[0039] 4、 The TTS speech synthesis module is used to convert the response text into the voice of the digital human speaking using a neural network TTS model. The speech synthesis adopts a streaming synthesis strategy, and its process is expressed as: (9); In the formula, is the generated voice signal; Standard naming prefix for the generation function; Is the input text sequence; Are the TTS model parameters.
[0040] During the synthesis process, the TTS voice synthesis module generates audio data while outputting it to shorten the overall response delay.
[0041] 5. Animation driving and rendering module, which is used to generate the digital human facial expressions and mouth movements according to the audio output by the TTS voice synthesis module and its related features using the speaker generation model. This module uses a temporal convolutional neural network or an LSTM architecture, and the timestamp of the output video frame is calibrated with the audio playback time to ensure lip-sync. The calibration formula is: (10); In the formula, Is the calibration output; Is the audio playback time; Is the video frame timestamp; Is the absolute value symbol. When Is less than the preset threshold, it is regarded as synchronized, otherwise buffer correction is triggered.
[0042] 6. WebRTC streaming module, which is used to encode and encapsulate the audio and video data output by the TTS voice synthesis module and the animation driving and rendering module, and use WebRTC to push it to the client in real time through the RTP protocol. This module supports dynamic adjustment of encoding parameters to adapt to different network conditions and achieve end-to-end delay transmission in milliseconds.
[0043] 7. Task distribution and coordination module, which is responsible for scheduling and distributing the tasks of each module and coordinating the intermediate results. This module uses a unified timestamp and a synchronization signal to ensure that the output data sequence of each stage within the same session is consistent, avoiding lip-sync problems caused by differences in processing speeds. At the same time, it uses an asynchronous message queue to achieve non-blocking transmission and improve the overall throughput.
[0044] 8. Model state synchronization and user data isolation module, which is used to transfer the necessary historical states in continuous sessions to ensure the coherence of the digital human dialogue and actions; at the same time, the user data of each user is completely isolated logically to ensure that the data is only used within the normal range, protecting user privacy and preventing data confusion between different sessions.
[0045] The system adopts a pipeline parallel design, and non-blocking transmission is achieved between modules through an asynchronous message queue, enabling the processing of each stage to overlap. The overall response time is close to the sum of the average processing times of each module, rather than strict accumulation, thus achieving an end-to-end response of less than 1.5 seconds. Specifically, it includes: Step 1: After the user's voice data is uploaded through the WebSocket access module, it immediately enters the ASR recognition module for processing. The ASR recognition module uses a streaming recognition algorithm to output partial text in real time; Step 2: The recognition result is transmitted to the semantic understanding and dialogue generation module through the message queue. This module uses a streaming large language model to generate response text based on the context and transmits it to the TTS speech synthesis module sentence by sentence; Step 3: Upon receiving the text, the TTS speech synthesis module immediately starts streaming speech synthesis, synthesizing and outputting the audio stream simultaneously, while notifying the animation driving and rendering module to generate corresponding video frames; Step 4: The model state synchronization module corrects the timestamps of the audio and video frames to ensure strict alignment between the two, and then transmits them to the client in real time via the WebRTC streaming module; Step 5: Throughout the process, the task distribution and coordination module manages through a unified timestamp and asynchronous queue to ensure that data within the same session is transmitted sequentially and orderly. Each module can process tasks of different sessions simultaneously, achieving a highly parallel pipeline processing.
[0046] Through this asynchronous pipeline and module collaboration mechanism, each microservice module can make full use of server resources in a high-concurrency scenario, achieving real-time interaction effects with low latency and high throughput.
[0047] III. Cross-platform Deployment and Verification The backend of this system adopts containerization technology and Kubernetes orchestration to achieve unified deployment across platforms and regions. When the container starts, the model instance pool module uses a preheating technology to load the model weight file into the shared memory area through memory mapping. During the loading process, the system performs file integrity self-check to ensure the accuracy of the model data. By using memory mapping and sharing technology, each microservice module can achieve consistent behavior on operating systems such as Linux, Windows, and macOS without having to reload the model weights repeatedly, thereby reducing memory consumption and improving startup efficiency. Each module communicates through standard interfaces to ensure interoperability and cross-platform consistency between services.
[0048] In addition, the system supports multi-region distributed deployment. The backend service cluster can be deployed simultaneously in multiple data centers and edge nodes globally. The intelligent routing mechanism automatically distributes user requests to the nearest node to reduce network latency and improve the user experience. After strict verification in cloud environments, edge nodes, and different terminal environments, this system can still operate stably in high-concurrency scenarios and maintain millisecond-level response and smooth adaptive video stream transmission effects in actual interactions.
[0049] The client SDK has also been fully cross - platform tested and verified on PC web pages, Android, iOS, and AR / VR devices. It realizes data interaction and media transmission through standard WebSocket and WebRTC protocols, ensuring that low - latency and real - time interaction performance requirements can be met in various terminal and network environments.
[0050] IV. Exception Handling and Fault - Tolerance Mechanism To ensure the stable operation of this system in high - concurrency and network - fluctuation environments, an exception - handling and fault - tolerance mechanism is added. Strict exception monitoring, error capture, retry, and automatic scaling - in and scaling - out strategies are set for each module during operation, specifically including: 1. Each module sets a clear processing timeout. For example, if the ASR recognition module does not return a recognition result within 500ms, it will automatically trigger a retry or call an alternative algorithm; similarly, if the TTS speech synthesis module detects a long - time stagnation during speech synthesis, it will immediately switch to an alternative model or return a degraded voice prompt. The retry mechanism adopts an exponential back - off algorithm, and its formula is: (11); In the formula, is the exponential back - off output; is the basic retry delay; is the current retry count; Through this mechanism, the system can automatically remedy when some modules encounter short - term failures, ensuring that critical processes are not blocked for a long time.
[0051] 2. To ensure the stability of the WebSocket and WebRTC channels in long - connection environments, the system sets periodic heartbeat packet detection on both sides. If multiple consecutive heartbeat packets do not receive responses, the disconnection reconnection mechanism will be automatically triggered, and the session scheduling module will be notified to rebuild the connection, thus ensuring that each session is not interrupted. This mechanism can effectively handle network fluctuations and short - term connection interruptions, guaranteeing the continuity and real - time nature of data transmission.
[0052] 3. Each module adopts a try - catch exception - capture mechanism during implementation to comprehensively capture exceptions generated during runtime. The captured exception information will be detailedly recorded and reported through a unified monitoring platform. The scheduling module takes corresponding measures according to the exception type, such as redistributing tasks, switching to an alternative instance, or starting degraded processing. This mechanism ensures that local errors do not spread to the entire system and can provide detailed basis for subsequent fault troubleshooting.
[0053] 4. When the system detects that a certain module is overloaded or has frequent anomalies, it will automatically activate the fault tolerance and degradation strategy. For example, in a high-concurrency scenario, the core functions are prioritized. The animation rendering module can simplify some detailed processing. When a node fails, the system reassigns the unfinished tasks to other normal nodes through the task queue for processing. In this way, even if some functions are temporarily degraded, the entire system can still maintain the coherence and stability of the core interaction process.
[0054] 5. The system is built with an elastic scaling module. Elastic scaling and edge node scheduling are as Figure 3 shown. It monitors the CPU, GPU, and memory usage rates of each container instance in real time. When the resource utilization rate reaches the preset threshold, the system automatically triggers the scaling-up operation, quickly starts new instances through the container orchestration platform, and registers them in the service registry for the scheduling module to assign tasks. After the load decreases, it automatically releases the idle instances to ensure efficient resource utilization and reduce operating costs. The scaling operation not only ensures the stable response of the system in a high-concurrency environment but also improves the overall scalability and flexibility of the system.
[0055] The interaction process is as Figure 4 shown. This system adopts an end-to-end streaming processing timing sequence to achieve real-time interaction between the user and the digital human. The overall system process combines a cloud-edge-end collaborative architecture and an asynchronous pipeline processing mechanism to ensure that each processing link can run overlappingly in a high-concurrency environment, and the total response time is less than 1.5 seconds. The specific process is as follows: First, the user inputs voice through the microphone on the terminal device, and the data is uploaded to the access layer in real time via a WebSocket long connection. The access layer intelligently routes the user request to the idle model instance pool in the backend service cluster. The model instance pool uses a multi-threaded shared scheduling mechanism to achieve efficient sharing of the weights of the same deep learning model, so that the model weights are only loaded once, and each parallel thread calls the shared model for inference processing in an independent session state.
[0056] Subsequently, the ASR recognition module uses a segmented recognition algorithm to transcribe and output partial text in real time for the arriving voice data. The session management module passes the recognized text to the semantic understanding and dialogue generation module in sequence. This module uses a streaming large language model to combine the dialogue context and the input text, and generates the answer text sentence by sentence in the form of a recursive function, and starts subsequent processing immediately after the first sentence is generated to achieve streaming output.
[0057] The generated response text is then passed to the TTS speech synthesis module, which is based on a neural network TTS model and adopts a streaming synthesis strategy to generate speech data segment by segment. During the synthesis process of the TTS speech synthesis module, the output audio features will synchronously drive the animation driving and rendering module, which generates digital human facial expression and mouth movement video frames according to the audio features using a temporal convolutional neural network or an LSTM architecture. To ensure strict synchronization of audio and video data, the system uses a timestamp calibration algorithm.
[0058] Finally, the integrated audio and video data is encoded by the WebRTC streaming module and pushed to the client in real time via the RTP protocol. The client uses an adaptive algorithm to dynamically adjust the video bitrate according to the network conditions to ensure smooth playback. The entire processing pipeline adopts an asynchronous parallel mechanism. Each module starts the subsequent processing immediately after completing part of the output, and ensures the correct sequential transfer of session data through unified session management and task scheduling, realizing closed-loop real-time interaction.
[0059] V. Practical Application Scenarios and Industry Implementation Capabilities This system has a wide range of practical application scenarios and industry implementation value, and can provide technical support for the large-scale user interaction needs in various fields, including but not limited to the following examples: In the field of online customer service, this system can be deployed as a cloud customer service agent to provide 7×24-hour uninterrupted intelligent interaction services for a large number of customers. Through the high concurrency support of this system, the customer service centers of e-commerce platforms or operators can provide real-time voice Q&A and problem handling for hundreds or even thousands of users simultaneously. The digital human customer service image can communicate with customers through fluent speech and expressions, answer common questions, guide business handling, or accept complaints and suggestions. With edge node deployment, user requests from different regions will be routed to the nearest server node for processing, reducing network latency and improving the interaction experience. In the cloud customer service scenario, this system can also reduce the server resource occupancy under large-scale deployment through the model instance pool mechanism, save operation costs for enterprises, and provide a consistent and high-quality service experience at the same time.
[0060] In the field of distance education and training, this system can be used as a virtual teacher or intelligent teaching assistant in large-scale online classrooms. Hundreds of students can interact with the digital human teacher in real time to ask questions and get answers, and the high-concurrency architecture of the system ensures that each student can get a timely response and the interactions are independent of each other without interference. For student devices with better terminal performance, the client SDK can select the local rendering mode to achieve a high-resolution presentation of the teacher's image; for mobile devices or low-performance terminals, cloud rendering is adopted to ensure smoothness. The digital human teacher can combine speech recognition and knowledge base to answer students' questions in real time, or actively explain teaching content according to the course progress. By deploying edge computing nodes close to campuses or educational institutions, this system can also reduce the network latency of distance education. In the online education scenario, this system provides an extensible way of teacher-student interaction, which supports a large number of students to be online at the same time while improving the teaching quality, and helps the large-scale dissemination of high-quality educational resources.
[0061] In summary, this system can adapt to the large-scale user interaction needs in multiple fields and has excellent scenario implementation capabilities. Through cloud deployment and edge collaboration, this system can be flexibly applied to the above various scenarios and customized and extended according to actual business needs, thus greatly improving the service automation and intelligence levels in various industries.
[0062] VI. Client Adaptation Solution The rendering strategy selection logic of the client SDK is as Figure 5 shown. To fully meet the computing power, display capabilities of different terminal devices and the real-time interaction needs in various network environments, this system provides a set of flexible and efficient client SDK rendering modes to ensure that the digital human interaction experience reaches the optimal effect on each platform. Specifically, it includes the following aspects: The system supports two rendering modes: Local rendering mode: On high-performance terminals, the server only sends down the processed instructions and control parameters, including audio streams, lip movement parameters, expression curves, and animation compensation information. The client SDK uses the local GPU for real-time graphics rendering to generate digital human videos. This mode can not only give full play to the rendering capabilities of the terminal hardware but also support personalized customization. For example, users can locally customize the skin, clothing, etc. of the digital human, thus achieving high-quality and personalized digital human display effects.
[0063] Remote rendering mode: For terminals with weak performance or scenarios requiring unified output, the server-side completes the entire digital human audio and video synthesis process, and pushes the generated complete video stream to the client in real time in a low-latency manner through the WebRTC protocol. This mode can ensure that a consistent high-quality viewing experience can be obtained regardless of the terminal computing power, and it is suitable for low-power devices or browser-side applications.
[0064] In addition, the client SDK has an adaptive network bandwidth adjustment function. The system uses an adaptive bitrate algorithm to monitor the current network latency, packet loss rate, and bandwidth status in real time, and dynamically adjusts the video resolution, frame rate, and bitrate to ensure high-definition video presentation when the broadband is sufficient, and automatically reduces the bitrate and resolution when the network conditions are poor to avoid lags and maintain smooth interaction. This adaptive algorithm is based on historical network statistics and real-time feedback, and adopts the following algorithm model: (12); In the formula, is the adjusted bitrate; is the base bitrate; is the currently measured bandwidth; is the target bandwidth threshold.
[0065] Based on the above design, the client adaptation solution can not only make full use of the computing power of high-performance terminals to achieve personalized and high-quality rendering, but also provide stable cloud rendering output for low-performance terminals. Coupled with the adaptive network adjustment strategy, this system maintains a smooth and realistic real-time interaction experience of digital humans under various terminal and network conditions.
[0066] In summary, this system not only achieves low latency, low resource consumption, and high scalability of the digital human system in a high-concurrency environment technically, but also proves its cross-platform deployment, stability, and economy in practice. This system architecture provides a mature, economical, and efficient real-time interaction solution for scenarios such as online education, cloud customer service, and large-scale event live broadcasts, and has significant market application prospects and technological innovation value.
[0067] Embodiment 2: On the basis of the foregoing embodiment, a dynamic model weight optimization mechanism based on user behavior feedback is added. By analyzing user behavior data in real time during the inference process, the local parameters of the shared model weight are dynamically adjusted, so as to improve the personalization and accuracy of digital human responses without increasing the additional model loading overhead. The core lies in encoding user behavior characteristics as weight correction factors and fusing them with the basic model weights to achieve session-level adaptive inference.
[0068] During the conversation, user interaction data is collected in real time, including response latency tolerance, semantic keyword matching degree, and emotional tendency score, and a feature vector F is generated through normalization processing; According to the weight correction function , the updated model weight is: (2); In the formula, is the updated model weight; is the shared weight; is the corrected matrix for pre-training; is the learning rate decay factor; is the weight correction function; is the Sigmoid function, used to control the correction amplitude; is the L2 norm of the feature vector F; In the model instance pool, each session thread independently calculates based on this and performs inference. After the session ends, the weight correction value is automatically recycled, and the shared weight remains unchanged.
[0069] Verification shows that in the case of obtaining an average error similar to that of the above embodiment, the dynamic weight optimization mechanism improves the semantic matching accuracy by 15%-20%. The weight correction only increases the single inference time by 3ms, and the memory occupancy increase is less than 1%. The results show that this mechanism can significantly improve the semantic accuracy and personalized adaptation ability of the digital human response by dynamically adjusting the shared model weights. The semantic matching accuracy is significantly optimized during the user interaction process, and the user satisfaction score shows a systematic increase. At the same time, since the weight correction process only acts on the session-level local parameters and the global model weights remain in a shared state, the growth of system memory occupancy and computational overhead is effectively controlled at a very low level, avoiding the resource waste problem caused by multi-model instantiation in traditional personalized solutions.
[0070] Embodiment 3: On the basis of the foregoing embodiment, a spatio-temporal prediction-based audio-visual synchronization enhancement algorithm is added. A lightweight temporal convolutional network is trained using historical audio-visual frame sequences to predict the audio features and mouth key points of future frames, and buffer frames are generated in advance to reduce synchronization errors. By fusing the prediction results with the actual output, feed-forward error compensation is achieved.
[0071] Extract the audio fundamental frequency , energy spectrum and video mouth key point coordinates to construct a spatio-temporal sequence ; Use the TCN network to predict the features of the next k frames: (3); In the formula, is the predicted audio-visual features of the next k frames; is the temporal convolutional network; is the input historical frame sequence; Define the dynamic synchronization correction amount: (4); In the formula, is the dynamic synchronization correction amount; is the correction intensity coefficient, used to control the replacement ratio of the predicted frames; is an index variable; is the predicted feature at the th moment; is the actual feature at the th moment; is the current moment L2 norm of the feature vector.
[0072] When , enable the predicted frame to replace the actual frame; otherwise, fuse the predicted and actual frames according to the weight ; In the formula, is the threshold, which determines whether to enable the predicted frame, usually 0.1; is the fusion weight, the proportion of the actual frame is , and the proportion of the predicted frame is .
[0073] Suppose that the digital human system processes a 300ms voice input in real-time interaction and generates corresponding video frames. Network jitter causes the 6th - 8th frame audio to arrive late. Buffer frames are generated in advance through the spatio-temporal prediction algorithm to compensate for the synchronization error.
[0074] The input data is the audio fundamental frequency, energy spectrum, and mouth key point coordinates of historical frames (t = 1 to t = 10), and the prediction target is the audio-visual features of the next 3 frames; Based on the historical data from t = 1 to t = 10, predict the audio fundamental frequency, energy spectrum, and mouth key point coordinates from t = 11 to t = 13. The actual t = 11 frame arrives late, and the system enables the predicted frame to replace; After the actual t = 11 frame arrives, calculate the prediction error:
[0075] Therefore, use only the predicted frame for rendering.
[0076] The effects of the audio-visual synchronization enhancement algorithm are shown in Table 1.
[0077] Table 1. Summary of the effects of the audio-visual synchronization enhancement algorithm
[0078] According to the experimental table, when the packet loss rate is 10%, the synchronization error drops from 55 ms to 15 ms, a decrease of 72.7%; the incidence of audio-video out-of-sync drops from 18% to 4%, a reduction of 77.8%; the end-to-end delay only increases by 0.2% - 0.3%, the single-frame prediction time of the TCN model is 0.5 ms, and the memory occupancy increases by 50 KB. The results show that the algorithm can significantly reduce the audio-video out-of-sync phenomenon caused by network jitter or calculation delay by predicting future audio-video features and implementing feed-forward error compensation. In the test of simulating a complex network environment, the system's tolerance to sudden delays and packet losses is significantly enhanced, the audio-video synchronization error is compressed to a negligible range, and the overall end-to-end delay does not introduce an additional burden due to prediction calculations. In addition, the design of the lightweight temporal convolutional network ensures the efficiency of the prediction process, and its computational overhead has little impact on the real-time performance of the system.
[0079] Example 4: Based on the foregoing embodiments, this embodiment deploys and tests a high-concurrency digital human real-time interaction system.
[0080] The server side is deployed in the cloud cluster using a containerized microservices architecture and is orchestrated and managed through Kubernetes. First, when the cluster starts, various deep learning models of the digital human are pre-loaded, and the model weights are loaded into the shared memory area at once using memory mapping technology. In this way, all model instances in the entire cluster reference the weight data at the same memory address, avoiding repeated loading of model files by each service node. Subsequently, a multi-threaded scheduling service is started to construct a global asynchronous task queue for receiving user session requests. The scheduler uses a lock-free queue and atomic operations to ensure efficient concurrent access and dynamically allocates tasks to idle model instance threads according to a predetermined priority algorithm. Several computing nodes are deployed at the edge as extended server-side instances, and when users connect nearby, the intelligent routing distributes requests to the nearest edge node, thereby reducing regional network latency. After the above cloud-edge collaborative deployment is completed, the client accesses the server through the distributed SDK and is ready to start the digital human interaction service. The entire system maintains a WebSocket long connection and session routing through the access layer gateway to ensure that user requests can be quickly transmitted to the idle nodes at the back end. In case of high concurrency, more container instances can be automatically expanded through Kubernetes to carry new sessions. Through the above deployment process, end-to-end collaboration from the cloud cluster, edge nodes to the client terminal is achieved, providing low-latency real-time digital human interaction services for high-concurrency users.
[0081] The client SDK adopts a modular design and internally includes a network communication module, a rendering engine module, and an adaptive control module. The network communication module sends the user's voice or text input through the WebSocket protocol and receives audio and video streams or rendering instructions from the server through the WebRTC protocol; the rendering engine module is responsible for generating digital human images locally or remotely according to different modes. Specifically, this system supports two modes: local rendering and remote rendering. When it is detected that the client device has sufficient computing power, the system selects the local rendering mode. The server only sends control information such as processed audio data, lip shapes, and expression parameters. The client SDK calls the local rendering engine to synthesize the digital human's picture in real time and synchronously play the voice. When the client is a low-performance device or in a scenario that requires unified output, the system switches to the remote rendering mode. The server completes the complete synthesis of the digital human's audio and video and pushes the encoded audio and video stream to the client through WebRTC with low latency. The client only needs to decode and play it. The adaptive control module of the client SDK monitors the terminal network status in real time and dynamically adjusts the rendering strategy: in the local rendering mode, if the network condition is poor, it reduces the texture resolution or frame rate; in the remote rendering mode, it automatically adjusts the video bit rate and resolution through the adaptive bit rate algorithm to ensure smooth interaction under any network conditions. Through the logical judgment and strategy switching inside the SDK, it ensures that various terminal users can obtain the best digital human interaction effect.
[0082] Under the above deployment architecture, this embodiment tests the system performance. The test environment includes 5 backend GPU servers and 2 geographically distributed edge nodes. By simulating a high-concurrency user voice session scenario, under the load of approximately 1000 users online and interacting simultaneously, the system runs stably as a whole. The measured results show that the average end-to-end response time per user is about 1.2 seconds, that is, the delay from when the user starts speaking to when the digital human video and voice feedback start playing is about 1.2 seconds, and the response delay of 95% of user sessions is less than 1.5 seconds, meeting the requirements of real-time interaction. When 500 users are concurrent, the CPU utilization rate of a single server is about 65%, the GPU utilization rate is about 75%, and the memory occupancy is about 60% of the total capacity. Compared with the traditional architecture where the memory and GPU occupancy increase linearly with concurrency due to loading the complete model in each session, this system greatly slows down the growth of resource occupancy through the sharing of the model instance pool. For example, when 100 users are concurrent, the total memory occupancy only increases by about 20% compared to when there is a single user, proving the effectiveness of model weight reuse. At the same time, by utilizing the elastic scaling ability of container orchestration, when the concurrent users surge to 1000, the system timely expands and adds container instances, and the overall CPU / GPU utilization rate still remains within a safe range, without any server overload or crash phenomenon, verifying the availability and stability of this system architecture in an actual high-concurrency scenario.
[0083] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable non-transitory storage media containing computer-usable program code.
[0084] The present invention can provide computer program instructions to a management platform of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed through the management platform of the computer or other programmable data processing devices generate a device for implementing the system.
[0085] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions of the system.
[0086] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions of the system.
Claims
1. A real-time interactive digital human system supporting high concurrency, characterized in that, Including: A model instance pool module for loading deep learning model weights into a shared memory area through memory mapping technology; A multi-threaded scheduling module that schedules user requests using a lock-free queue and a dynamic priority algorithm; An asynchronous pipeline processing module composed of decoupled microservices, with the modules connected in series through an asynchronous message queue; A client SDK for dynamically switching the rendering mode according to the terminal hardware performance and network conditions; An elastic scaling module for real-time monitoring of resource loads and automatically adjusting the number of service instances; An audio-video synchronization calibration module for ensuring that the synchronization error between audio and video frames is lower than a preset threshold through a timestamp alignment algorithm.
2. The system according to claim 1, characterized in that: The model instance pool module supports hot-updating the model weights in the shared memory without interrupting the current session.
3. The system according to claim 1, characterized in that: The dynamic priority algorithm of the multi-threaded scheduling module satisfies the following formula: (1); In the formula, is the scheduling priority; and are adjustable parameters; is the waiting time of task i in the queue; is the urgency of session i; The lock-free queue uses atomic operations based on a circular buffer to implement task distribution.
4. The system according to claim 1, characterized in that: The asynchronous pipeline processing module adopts a cross-model collaborative inference mechanism, combines the output of a large language model with preset knowledge base rules to generate conversation responses, and filters the optimal results through a confidence threshold.
5. The system according to claim 1, characterized in that: The elastic scaling module distributes user requests through a weighted round-robin algorithm based on the geographical location and real-time load of edge nodes. When the GPU utilization rate exceeds the threshold, it preferentially expands the edge node instances deployed in the same geographical area, and retains the minimum instance pool during scaling.
6. The system according to claim 1, characterized in that: The audio-video synchronization calibration module adopts a hierarchical correction strategy. In the first layer, the audio and video timestamps are synchronized through a global clock. In the second layer, the frame interval is dynamically adjusted based on audio features and video mouth key points. When the synchronization error exceeds the threshold, an interpolation frame filling or frame dropping mechanism is triggered.
7. The system according to any one of claims 1-6, characterized in that: It also includes a model weight optimization module for dynamically adjusting the local parameters of the shared model weights by real-time analyzing user behavior data during the inference process. The updated model weights are: (2); Wherein, is the updated model weight; is the shared weight; is the learning rate decay factor; is the weight correction function; is the Sigmoid function; is the L2 norm of the feature vector F.
8. The system according to any one of claims 1-6, characterized in that: It also includes an audio-video synchronization enhancement module for training a lightweight temporal convolutional network using historical audio-video frame sequences to predict the audio features and mouth key points of future frames. The features of the future k frames are predicted through the following formula: (3); In the formula, is the predicted audio-visual feature of the future k frames; is the temporal convolutional network; is the input historical frame sequence; The dynamic synchronization correction amount is defined as: (4); Wherein, is the dynamic synchronization correction amount; is the correction intensity coefficient; is the index variable; is the predicted feature at the th time; is the L2 norm of the feature vector at the current time.
9. An implementation method of a real-time interactive digital human system supporting high concurrency, characterized in that: The implementation of the method is based on the system described in any one of claims 1-8. The method includes: Loading deep learning model weights into a shared memory area through memory mapping technology, with each thread independently maintaining the session state and sharing the same model weights; Scheduling user requests using a dynamic priority algorithm based on a lock-free queue, and distributing tasks to idle model instances through atomic operations; Decoupling different modules into independent microservices and implementing streaming processing through connection in series with an asynchronous message queue; Dynamically switching the rendering mode according to the terminal hardware performance and network conditions, and adjusting the transmission bitrate based on an adaptive bitrate algorithm; Based on containerized orchestration and edge node deployment strategies, real-time monitoring the server resource load, dynamically scaling the service instances, and distributing user requests to the nearest edge node through geographical location routing; Ensuring that the synchronization error between audio and video frames is lower than a preset threshold through a timestamp alignment algorithm.
10. A computer-readable storage medium storing a computer program therein, characterized in that: When the computer program is run, it implements the method described in claim 9.
Citation Information
Patent Citations
Message processing method and device
CN105975433A
Distributed and container virtualization-based elastic micro-service system and implementation method
CN114422371A
Neural network model scheduling deployment method and device, equipment, medium and product
CN114489746A
Intelligent environmental adaptation animation rendering optimization system
CN118644588A
Station area intelligent fusion terminal data processing system based on edge calculation
CN119440800A
Cited By
Intelligent reply system suitable for live broadcast matrix scene
CN120378701A
An intelligent reply system suitable for live matrix scenarios
CN120378701B
Audio transmission control system and method and audio equipment
CN120639718A
Audio transmission control system, method and audio device
CN120639718B
All-media all-signal integrated management platform
CN120935094A