A collaborative interaction system and method for a computing power limited environment

CN122816802APending Publication Date: 2026-09-25CHINA UNIV OF GEOSCIENCES (WUHAN) +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610976021.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-02
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

传统方案在处理长度不足的残余音频时,往往采用直接截断丢弃的粗暴策略,这会导致输出语音高频出现“劈啪”爆音,并随着时间推移引发严重的音视频画面(唇形)错位与撕裂

Benefits of technology

[0011]本发明的有益效果为:能够基于软硬件底层的跨模态精细化调度策略,高效解决多模型并发引发的系统崩溃与唇形撕裂等痛点,实现毫秒级打断响应、不宕机的高可用流式交互,为单卡或低配物理机上的本地化数字人应用提供坚实的底层技术支撑。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122816802A_ABST
    Figure CN122816802A_ABST
Patent Text Reader

Abstract

The application provides a kind of collaborative interaction system and method for computing power limited environment, front-end interaction and state control module, cross-modal streaming data hub module, visual driving and dynamic bypass rendering module, can be based on the cross-modal fine scheduling strategy of hardware bottom, efficiently solve the pain points such as system collapse and lip tearing caused by multi-model concurrent, realize millisecond level interruption response, non-downtime high-availability streaming interaction, provide solid underlying technology support for localized digital human application on single card or low-config physical machine;Compared with the traditional normally open type monitoring and full-time strong binding rendering digital human architecture, the method is more efficient, energy-saving and robust in reducing invalid consumption of end-side resources, squeezing peak video memory elastic margin and improving system anti-downtime capability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a collaborative interaction system and method for computing-constrained environments. Background Technology

[0002] With the explosion of Large Language Models (LLM) and generative artificial intelligence, digital human interaction systems are evolving from traditional "pre-set video playback" to "real-time streaming generation." Mainstream digital human architectures typically include multiple heavyweight AI models such as Automatic Speech Recognition (ASR), streaming output from Large Language Models, Text-to-Speech (TTS), and facial visual lip-sync rendering (e.g., the Wav2Lip model). However, in computing-constrained, single-node localized deployment scenarios, existing solutions for consumer hardware suffer from the following fatal flaws: Memory / GPU memory overflow (OOM) and system crash risks under extreme concurrency: Existing open-source components are often designed to run independently, lacking global system-level computing power coordination. In physical machines with limited computing power, to ensure the basic survival of multi-container deployments (such as Docker), stringent storage quotas and hard memory limits are usually imposed. In traditional digital human systems, during "standby listening" or "silent pause" phases, their visual rendering models continue to perform ineffective GPU forward inference calculations. This continuous high energy consumption can easily break through the system's limited memory ceiling, causing the process to be forcibly terminated by the operating system or the entire machine to crash.

[0003] "Lip tearing" and audio popping caused by modal transition rate mismatch: In real-world environments, LLM streaming output exhibits extremely random rate distribution and is often accompanied by Markdown formatting symbols or complex mathematical symbols. If fed directly to TTS without deep cleaning and scientific chunking, it can lead to extreme synthesis delays or stuttering in pronunciation. More seriously, the audio packet size in network streaming is random, while the underlying face rendering model (Wav2Lip) is forced to rely on fixed-length audio feature frames. Traditional solutions often employ a crude strategy of directly truncating and discarding insufficient audio fragments, resulting in high-frequency "crackling" popping in the output speech, and over time, causing severe audio-visual (lip-sync) misalignment and tearing.

[0004] The "always-on" front-end interaction leads to ineffective back-end computing power consumption: Traditional WebRTC voice interaction front-ends heavily rely on back-end VAD (Voice Activity Detection) to continuously listen to and determine user intent. On low-end hardware, this constant, continuous listening and audio uplinking will chronically consume already limited CPU and memory resources, preventing the back-end from releasing peak computing power during peak lip-sync rendering of the digital human, thus significantly slowing down the entire system's first frame response speed. Summary of the Invention

[0005] The purpose of this invention is to address the shortcomings of the prior art by providing a collaborative interaction system and method for computing-constrained environments, efficiently solving the pain points of system crashes and lip tearing caused by multi-model concurrency, and achieving highly available streaming interaction with millisecond-level interruption response and no downtime.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: This invention provides a collaborative interaction system for computing-constrained environments, comprising: Front-end interaction and status control module: It is connected to the back-end via a bidirectional network communication link of HTTP and WebSocket, and is responsible for encapsulating the collected user voice stream, intent truncation command and single interaction status signal into a standard network request and uplinking it to the back-end hub. Cross-modal streaming data hub module: After receiving the front-end signal, it continuously sends the aligned standard fixed-length multimodal data frames to the input buffer pool of the back-end dynamic bypass rendering module through the asynchronous event-driven and memory queue mechanism. The visual-driven and dynamic bypass rendering module extracts data from the buffer pool and performs scheduling based on audio energy detection. According to the detection results, it generates a sequence of audio-visual synchronized streaming video frames through dynamic memory pointer switching or video memory bus data loading. Finally, it is sent back to the front-end module for real-time rendering and playback through the streaming media communication protocol.

[0007] Furthermore, a collaborative interaction system designed for computing-constrained environments is adopted, which also includes: At the interaction layer, a state machine based on a single interaction and a global intent truncation mechanism are used, specifically as follows: During initial wake-up and accidental touch prevention, the front-end audio acquisition of the collaborative interaction system is in a low-power weak listening mode by default. Only when the local lightweight engine detects the preset wake-up entity will the state machine unlock and enter a high-priority active state. At the same time, in order to prevent invalid audio concurrency during network latency, the front-end synchronously enables request locks to physically block noise input from non-main threads within a specified time window. During the instruction flow and active sleep phase, when a user's valid voice command is transmitted to the backend interface, the frontend does not wait for the LLM to return the result, but immediately triggers a sleep command locally to actively suspend continuous voice recognition polling.

[0008] Furthermore, the intermediate streaming layer constructs a dual-buffered hub that includes dynamic throttling and feature reconstruction, specifically as follows: In the pre-cleaning and dynamic throttling stages of the text stream, an asymmetric segmentation pipeline is constructed; To eliminate the first-packet delay of the digital human's opening, the collaborative interaction system sets an adaptive segmentation threshold; The segmented text blocks are fed into a deep regularization cleaner, which forcibly strips Markdown tags using a preset rule engine and strictly converts special symbols into standard Chinese colloquial pronunciation.

[0009] Furthermore, in the residual compensation and fixed-length reconstruction stage of network audio features, the reconstruction logic enforced by the collaborative interaction system is as follows: Seamless pre-stitching: The newly arrived audio array is stitched together with the incomplete tail notes left over from the previous truncation in the buffer pool; On-demand precise cutting: Based on the strictly fixed frame length required for inference from the underlying visual model, the spliced ​​streaming data is cyclically cut into equal lengths and then sequentially delivered to the underlying rendering driver queue. Residual retention and silence alignment: After the loop is cut, audio fragments that are less than a standard frame length are stored back in the residual buffer pool to wait for the next network packet; when the end of the entire interactive voice stream is determined, the system uses silent zero data to fill in the last missing feature in the buffer pool and sends it as the final frame.

[0010] Furthermore, a dynamic bypass scheduling mechanism was embedded in the backend rendering layer. The specific implementation steps are as follows: During the pre-detection phase of batch rendering, the collaborative interaction system engine retrieves multimodal data from the pre-buffer pool at a specific time step to construct the computation batch; before the data is officially sent into the visual model, the collaborative interaction system forcibly starts the audio energy discriminator to perform pure silence detection on the audio sequence extract of the current batch. During the dynamic rendering power distribution phase, the collaborative interaction system triggers a hardware-level distribution switch based on the energy detection results: if the audio clip contains effective acoustic energy, the collaborative interaction system executes the standard rendering link, wakes up the Wav2Lip model, loads the Mel spectrum and facial image array into the GPU memory to perform high-intensity matrix operations, and generates lip-synced dynamic video frames. If the discriminator confirms that the current audio batch is in a completely silent state, the collaborative interaction system will immediately trigger a bypass interception action. At this time, the scheduler forcibly blocks the aforementioned high-energy-consuming GPU forward inference function and instead uses lightweight memory pointer scheduling to directly extract the pre-cached digital human static facial original image and push it directly into the output queue as the video stream of the current time slice.

[0011] The beneficial effects of this invention are: it can efficiently solve the pain points such as system crashes and lip tearing caused by multi-model concurrency based on the cross-modal fine-grained scheduling strategy at the software and hardware level, achieve millisecond-level interruption response and high-availability streaming interaction without downtime, and provide solid underlying technical support for localized digital human applications on single-card or low-configuration physical machines.

[0012] Compared to the traditional always-on monitoring and all-time strongly bound rendering digital human architecture, this method is more efficient, energy-saving and robust in reducing the ineffective waste of edge resources, squeezing out the peak video memory elasticity margin and improving the system's anti-downtime capability.

[0013] Compared to conventional multimodal streaming alignment schemes that rely on simple data truncation or direct discarding of incomplete audio frames, this method's dynamic throttling and residual reconstruction strategy is more refined, systematic, and adaptive. Furthermore, the final rendered streaming digital human video has a clear advantage in terms of audio-visual synchronization accuracy, picture continuity, and robustness against network jitter. Attached Figure Description

[0014] Figure 1 This is a link architecture diagram of a collaborative interaction system designed for computing-constrained environments. Figure 2 This is an interactive control flowchart based on post-attack sleep and intent interception; Figure 3 This is a flowchart of the multimodal data alignment process based on dynamic segmentation and residual reconstruction. Figure 4 This is a flowchart of GPU dynamic bypass rendering based on audio energy detection. Detailed Implementation

[0015] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0016] like Figure 1 As shown, a collaborative interaction system for computing-constrained environments includes: Front-end interaction and status control module: As the entry point for preventing accidental touches and protecting computing power by immediately going to sleep after activation, it is connected to the back-end via a bidirectional network communication link of HTTP and WebSocket. It is responsible for encapsulating the collected user voice stream, intent interception instructions, and single interaction status signals into standard network requests and uplinking them to the back-end hub. Please see Figures 3 to 4 Cross-modal streaming data hub module: As the core of data scheduling, after receiving the front-end signal, it completes the dynamic segmentation of large model text and the seamless compensation and reconstruction of residual audio feature fragments through asynchronous event driving and memory queue mechanism. Then, the aligned standard fixed-length multimodal data frames (such as 20ms audio slices) are continuously delivered to the input buffer pool of the back-end dynamic bypass rendering module through high-speed shared memory or inter-process communication (IPC) link. The visual-driven and dynamic bypass rendering module extracts data from the buffer pool and performs scheduling based on audio energy detection. According to the detection results, it generates a streaming video frame sequence that is synchronized with audiovisuals by switching dynamic memory pointers (triggering silent bypass direct output of static image) or loading video memory bus data (waking up the GPU to perform inference). Finally, it is sent back to the front-end module for real-time rendering and playback through streaming media communication protocols (such as WebRTC / WebSocket video stream delivery).

[0017] This is achieved using a collaborative interaction system designed for computing-constrained environments, and also includes: Interaction layer: A sleep control module that activates immediately after activation based on intent interception and computing power to prevent accidental touches; As the first data entry point for a multimodal digital human system, the interaction layer aims to address the issue of prolonged invalid CPU and memory usage on low-end computing nodes due to constant microphone monitoring (such as WebRTC's default continuous VAD detection). To avoid resource congestion during concurrent execution of large language model inference and visual rendering, the collaborative interaction system restructures the traditional voice assistant's monitoring sequence and introduces a state machine based on one-shot interaction and a global intent truncation mechanism.

[0018] At the interaction layer, a state machine based on a single interaction and a global intent truncation mechanism are used, specifically as follows: During initial wake-up and accidental touch prevention, the front-end audio acquisition of the collaborative interaction system is in a low-power weak listening mode by default. Only when the local lightweight engine detects the preset wake-up entity will the state machine unlock and enter a high-priority active state. At the same time, in order to prevent invalid audio concurrency during network latency, the front-end synchronously enables request locks to physically block noise input from non-main threads within a specified time window. During the command flow and active sleep phase, once a user's valid voice command is transmitted to the backend interface, the frontend does not wait for the LLM's return result but immediately triggers a sleep command locally to actively suspend continuous speech recognition polling. This interrupted sleep mechanism frees up a critical CPU processing queue for high-concurrency computation of large-scale models and TTS voice stream framing at the physical machine level.

[0019] During the abnormal interruption and multimodal queue clearing phases, and considering the high-frequency interruption scenarios during long text broadcasts by the digital human, this collaborative interaction system breaks away from traditional single software process control. It constructs a penetrating deep truncation architecture that collaboratively works with a front-end interactive control module, a network communication module, a central scheduling module, and a multimodal generation engine module. The specific module connection relationships and technical transmission scheme of the penetrating deep truncation architecture are as follows: First, the front-end interaction control module is responsible for real-time acquisition of local audio and wake-up word recognition. When the front-end interaction control module detects a user interrupt command containing a wake-up word during the sleep period or during the digital human's broadcast, it immediately encapsulates a global interruption request with a timestamp. The global interruption request relies on the network communication routing module to establish an uplink network transmission link via the HTTP POST protocol (API route: / interrupt_talk) and send the interruption signal to the back-end server.

[0020] Secondly, after receiving the interrupt signal, the network communication module transmits it through the asynchronous event bus to the backend central scheduling module. The central scheduling module, as the control core, immediately issues hardware-level clearing instructions in parallel to each component of the multimodal generation engine module: Queue cache clearing and transmission: The central scheduling module uses memory address mapping and pointer reset mechanism to call the underlying clearing function to instantly erase the large model text waiting in the memory queue and the audio feature frames that have not yet been sent for rendering; Thread blocking control: The central scheduling module directly cuts off the underlying data stream connection with the TTS synthesis engine through inter-process communication or thread control signals, forcibly terminating the currently executing streaming speech synthesis thread; Finally, after completing the dual blocking of queues and threads, the central scheduling module forcibly resets the status flags in the visual rendering pipeline and sends a confirmation signal of successful interruption to the front-end interactive control module via the established WebSocket full-duplex communication link, completing the closed-loop synchronization of the front-end and back-end state machines. Through the aforementioned cross-module network protocol communication and underlying memory manipulation, the collaborative interaction system achieves millisecond-level response interruption and complete release of limited computing resources.

[0021] In multimodal concurrent environments, the streaming text output rate of Large Language Models (LLMs) exhibits strong randomness, and the data packets of audio streams transmitted over the network vary in size. Directly rendering the raw multimodal data using a low-level, fixed-length vision-driven model results in severe lip tearing, image twitching, and audio popping due to feature step size mismatch. To completely solve the problem of cross-modal alignment in streaming, this collaborative interaction system constructs a dual-buffered hub in the intermediate layer, incorporating dynamic throttling and feature reconstruction.

[0022] The intermediate streaming layer constructs a dual-buffered hub that includes dynamic throttling and feature reconstruction, specifically: In the pre-cleaning and dynamic throttling stages of the text stream, an asymmetric segmentation pipeline is constructed; To eliminate the first-packet delay of the digital human's opening, the collaborative interaction system sets an adaptive segmentation threshold; The segmented text blocks are fed into a deep regularization cleaner, which forcibly strips Markdown tags using a preset rule engine and strictly converts special symbols into standard Chinese colloquial pronunciation.

[0023] During the residual compensation and fixed-length reconstruction stages of network audio features, the reconstruction logic enforced by the collaborative interaction system is as follows: Seamless pre-stitching: The newly arrived audio array is stitched together with the incomplete tail notes left over from the previous truncation in the buffer pool; On-demand precise cutting: Based on the strictly fixed frame length required for inference from the underlying visual model, the spliced ​​streaming data is cyclically cut into equal lengths and then sequentially delivered to the underlying rendering driver queue. Residual retention and silence alignment: After the loop is cut, audio fragments that are less than a standard frame length are stored back in the residual buffer pool to wait for the next network packet; when the end of the entire interactive voice stream is determined, the system uses silent zero data to fill in the last missing feature in the buffer pool and sends it as the final frame.

[0024] This mechanism completely eliminates the popping sounds caused by data stream interruption at the physical level through extremely lightweight memory operations, ensuring millisecond-level high-precision audio-visual synchronization for digital humans even with very low hardware configurations.

[0025] In computationally constrained deployment environments, large vision-driven models (such as Wav2Lip) are a core bottleneck consuming GPU memory and computing power in collaborative interaction systems. Traditional digital human systems employ a "strongly bound rendering" mode, meaning that regardless of whether the digital human is actively speaking or in a silent, idle state, the underlying pipeline continuously sends visual features to the GPU for undifferentiated forward inference. When a single consumer-grade GPU concurrently processes LLM inference and video rendering, this ineffective, constant computational waste can easily trigger system-level memory overflow (OOM). To overcome this energy-consuming blind spot, the collaborative interaction system innovatively embeds a dynamic bypass scheduling mechanism at the underlying layer of the rendering-driven pipeline.

[0026] A dynamic bypass scheduling mechanism is embedded in the backend rendering layer. The specific implementation steps are as follows: During the pre-detection phase of batch rendering, the collaborative interaction system engine retrieves multimodal data from the pre-buffer pool at a specific time step to construct the computation batch; before the data is officially sent into the visual model, the collaborative interaction system forcibly starts the audio energy discriminator to perform pure silence detection on the audio sequence extract of the current batch. During the dynamic rendering power distribution phase, the collaborative interaction system triggers a hardware-level distribution switch based on the energy detection results: if the audio clip contains effective acoustic energy, the collaborative interaction system executes the standard rendering link, wakes up the Wav2Lip model, loads the Mel spectrum and facial image array into the GPU memory to perform high-intensity matrix operations, and generates lip-synced dynamic video frames. If the discriminator confirms that the current audio batch is in a completely silent state, the collaborative interaction system will immediately trigger a bypass interception action. At this time, the scheduler forcibly blocks the aforementioned high-energy-consuming GPU forward inference function and instead uses lightweight memory pointer scheduling to directly extract the pre-cached digital human static facial original image and push it directly into the output queue as the video stream of the current time slice.

[0027] The "static-instead-of-dynamic" bypass control mechanism, while ensuring the absolute continuity of the video stream frame rate, forces the GPU computing requests during the nearly 50% "non-speaking" time of the interaction cycle to zero. This squeezes out crucial video memory elasticity margin for the 32GB memory and the multimodal ultra-fast switching in the restricted Docker container environment, building a physical-level defense against system crashes.

[0028] Please see Figure 2 To address the issues of ineffective resource consumption and accidental touch prevention during multimodal concurrency caused by routine monitoring on low-end computing nodes, the first key aspect of this invention is the design of a front-end interactive state machine scheduling strategy based on intent truncation and "sleep immediately after sending". This strategy can proactively suspend the front-end recognition engine after voice uplink and issue a deep clearing command that penetrates the underlying layer when interruption intent is detected. This completely cuts off the continuous encroachment of environmental noise on the back-end's extreme computing power from the data source, achieving precise tilting of system computing power and millisecond-level response interruption. To address the issues of audio dropouts, crashes, and screen tearing caused by the mismatch between the random output rate of large language models and the fixed frame length of the underlying visual rendering, the second key aspect of this invention is the construction of a fixed-length audio residue compensation mechanism adapted to cross-modal streaming rates. By constructing a residue buffer pool in memory, performing seamless pre-splitting, precise fixed-length cutting on demand, and using silence data to fill in the tail sounds for non-fixed-length network audio packets, a set of audio and video timeline strong alignment algorithms with extremely low memory overhead is constructed, ensuring the smooth reconstruction of streaming multimodal data. To overcome the bottleneck of large visual-driven models constantly occupying GPU memory under limited hardware conditions, thus causing system-level overflow (OOM) crashes, the third key point of this invention is the embedding of a GPU memory dynamic bypass control mechanism based on audio energy detection. This mechanism can accurately detect completely silent features (such as listening or standby periods) at the very bottom of the rendering pipeline, forcibly block ineffective and energy-intensive GPU forward inference calculations, and instead directly send the static original image through memory pointer scheduling, thus achieving a "static-instead-of-dynamic" elastic release of GPU memory usage under multimodal concurrency.

[0029] The embodiments described above are merely illustrative of implementation methods of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be defined by the appended claims.

Claims

1. A collaborative interaction system for computing-constrained environments, characterized in that, include: Front-end interaction and status control module: It is connected to the back-end via a bidirectional network communication link of HTTP and WebSocket, and is responsible for encapsulating the collected user voice stream, intent truncation command and single interaction status signal into a standard network request and uplinking it to the back-end hub. Cross-modal streaming data hub module: After receiving the front-end signal, it continuously sends the aligned standard fixed-length multimodal data frames to the input buffer pool of the back-end dynamic bypass rendering module through the asynchronous event-driven and memory queue mechanism. The visual-driven and dynamic bypass rendering module extracts data from the buffer pool and performs scheduling based on audio energy detection. According to the detection results, it generates a sequence of audio-visual synchronized streaming video frames through dynamic memory pointer switching or video memory bus data loading. Finally, it is sent back to the front-end module for real-time rendering and playback through the streaming media communication protocol.

2. The collaborative interaction method for computing-constrained environments according to claim 1, characterized in that: This is implemented using a collaborative interaction system for computing-constrained environments as described in claim 1, and further includes: At the interaction layer, a state machine based on a single interaction and a global intent truncation mechanism are used, specifically as follows: During initial wake-up and accidental touch prevention, the front-end audio acquisition of the collaborative interaction system is in a low-power weak listening mode by default. Only when the local lightweight engine detects the preset wake-up entity will the state machine unlock and enter a high-priority active state. At the same time, in order to prevent invalid audio concurrency during network latency, the front-end synchronously enables request locks to physically block noise input from non-main threads within a specified time window. During the instruction flow and active sleep phase, when a user's valid voice command is transmitted to the backend interface, the frontend does not wait for the LLM to return the result, but immediately triggers a sleep command locally to actively suspend continuous voice recognition polling.

3. The collaborative interaction method for computing-constrained environments according to claim 2, characterized in that: The intermediate streaming layer constructs a dual-buffered hub that includes dynamic throttling and feature reconstruction, specifically: In the pre-cleaning and dynamic throttling stages of the text stream, an asymmetric segmentation pipeline is constructed; To eliminate the first-packet delay of the digital human's opening, the collaborative interaction system sets an adaptive segmentation threshold; The segmented text blocks are fed into a deep regularization cleaner, which forcibly strips Markdown tags using a preset rule engine and strictly converts special symbols into standard Chinese colloquial pronunciation.

4. The collaborative interaction method for computing-constrained environments according to claim 3, characterized in that: During the residual compensation and fixed-length reconstruction stages of network audio features, the reconstruction logic enforced by the collaborative interaction system is as follows: Seamless pre-stitching: The newly arrived audio array is stitched together with the incomplete tail notes left over from the previous truncation in the buffer pool; On-demand precise cutting: Based on the strictly fixed frame length required for inference from the underlying visual model, the spliced ​​streaming data is cyclically cut into equal lengths and then sequentially delivered to the underlying rendering driver queue. Residual retention and silence alignment: After the loop is cut, audio fragments that are less than a standard frame length are stored back in the residual buffer pool to wait for the next network packet; when the end of the entire interactive voice stream is determined, the system uses silent zero data to fill in the last missing feature in the buffer pool and sends it as the final frame.

5. A collaborative interaction method for computing-constrained environments according to claim 3, characterized in that: A dynamic bypass scheduling mechanism is embedded in the backend rendering layer. The specific implementation steps are as follows: During the pre-detection phase of batch rendering, the collaborative interaction system engine retrieves multimodal data from the pre-buffer pool at a specific time step to construct the computation batch; before the data is officially sent into the visual model, the collaborative interaction system forcibly starts the audio energy discriminator to perform pure silence detection on the audio sequence extract of the current batch. During the dynamic rendering power distribution phase, the collaborative interaction system triggers a hardware-level distribution switch based on the energy detection results: if the audio clip contains effective acoustic energy, the collaborative interaction system executes the standard rendering link, wakes up the Wav2Lip model, loads the Mel spectrum and facial image array into the GPU memory to perform high-intensity matrix operations, and generates lip-synced dynamic video frames. If the discriminator confirms that the current audio batch is in a completely silent state, the collaborative interaction system will immediately trigger a bypass interception action. At this time, the scheduler forcibly blocks the aforementioned high-energy-consuming GPU forward inference function and instead uses lightweight memory pointer scheduling to directly extract the pre-cached digital human static facial original image and push it directly into the output queue as the video stream of the current time slice.