Python-golang dual-architecture multi-modal real-time digital human customer service system based on micro-service deployment

CN122824922APending Publication Date: 2026-09-25FUJIAN GONGTIAN SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611282589.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-24
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]然而,现有的数字人系统在处理这一状态切换时,大多采用直接替换当前播放帧的方式,例如,在某一已知的实现方案中,服务端在收到推理结果后立即停止静默视频循环,并将推理首帧作为当前显示帧推送给前端;推理结束后,再直接切回静默视频的某一固定起始帧;这种硬切换机制在实际运行中容易产生视觉不连续问题,一种常见的情形是用户提问后,数字人从静默姿态直接跳变为开口说话的口型,且面部位置、头部角度或表情与切换前的静默帧存在明显差异,导致前端画面出现明显的闪烁或突变感;推理结束后,从说话状态切回静默状态时,同样可能因为静默帧与推理末帧的不对齐而产生类似跳变,这种突变不仅降低了数字人的逼真度,也削弱了用户的沉浸式交互体验

Benefits of technology

[0017]通过Python侧与Golang侧的明确职责划分,将AI推理密集型任务与高并发网络控制任务分离,并结合gRPC双向流式传输与统一时间戳对齐机制,实现了音频、视频与字幕的多模态同步;同时,系统利用轮询算法进行推理进程池负载均衡,配合节流算法与自适应缓冲机制,有效保障多用户并发场景下的资源稳定性和低延迟响应能力;此外,Golang控制层通过对静默视频与推理视频的锚点帧定位及平滑过渡处理,避免了状态切换时的画面闪烁和面部跳变,最终经由WebRTC协议将时序一致的多模态流同步推送至WEB前端,提升实时数字人交互的流畅性、逼真度及系统整体的可扩展性和可靠性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122824922A_ABST
    Figure CN122824922A_ABST
Patent Text Reader

Abstract

The application provides a Python-Golang dual-architecture multi-modal real-time digital human customer service system based on micro-service deployment, relates to the technical field of data processing, and comprises a generation module, a scheduling module and the like, wherein the generation module is used for acquiring a multi-modal interaction request of a user, performing speech recognition, agent reasoning, text-to-speech and digital human animation generation processing in sequence on the Python side, and obtaining a multi-modal reasoning result; the scheduling module is used for synthesizing video image frames into a basic video sequence on the Python side according to the video image frames in the multi-modal reasoning result, performing frame-by-frame face recognition and face semantic segmentation on the basic video, and distributing a preprocessing task to a stateless AI working node in an inference process pool according to a polling algorithm to obtain a preprocessing video frame sequence with a face semantic segmentation mask and a bounding box coordinate. The application reduces the end-to-end interaction delay of the multi-modal digital human customer service system and improves the customer response speed and interaction naturalness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a Python-Golang dual-architecture multimodal real-time digital human customer service system based on microservice deployment. Background Technology

[0002] With the widespread application of virtual digital humans in scenarios such as intelligent customer service and online education, users have placed higher demands on the fluency and naturalness of real-time interaction between digital humans. In a typical real-time digital human customer service system, the digital human usually needs to frequently switch between two states: silent waiting and inference-based response. When the user does not speak, the system pushes a silent video that plays in a loop to the front end, such as a picture of the digital human standing still or breathing slightly. When the user asks a question and triggers background inference, after speech recognition, semantic understanding, speech synthesis and lip-reading, the system needs to replace the silent video with the speaking video frame generated by the inference to present the effect of the digital human speaking and responding.

[0003] However, most existing digital human systems handle this state transition by directly replacing the currently playing frame. For example, in a known implementation, the server immediately stops the silent video loop after receiving the inference result and pushes the first frame of the inference as the current display frame to the front end. After the inference ends, it directly switches back to a fixed starting frame of the silent video. This hard switching mechanism is prone to visual discontinuity problems in actual operation. A common scenario is that after a user asks a question, the digital human jumps directly from a silent posture to lip-syncing, and the facial position, head angle, or expression is significantly different from the silent frame before the switch, causing obvious flickering or abrupt changes in the front end screen. After the inference ends, when switching back from the speaking state to the silent state, a similar jump may occur due to the misalignment between the silent frame and the last frame of the inference. This abrupt change not only reduces the realism of the digital human but also weakens the user's immersive interactive experience. Summary of the Invention

[0004] This invention provides a Python-Golang dual-architecture multimodal real-time digital human customer service system based on microservice deployment, which reduces the end-to-end interaction latency of the multimodal digital human customer service system and improves the customer service response speed and the naturalness of the interaction.

[0005] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:

[0006] The first aspect is a multimodal real-time digital human customer service system based on a microservice-oriented Python-Golang dual-architecture system, including:

[0007] The generation module is used to obtain user multimodal interaction requests and sequentially perform speech recognition, intelligent agent reasoning, text-to-speech and digital human animation generation processing on the Python side to obtain multimodal reasoning results;

[0008] The scheduling module is used to synthesize video image frames into a basic video sequence on the Python side based on the video image frames in the multimodal inference results, perform frame-by-frame face recognition and facial semantic segmentation on the basic video, and allocate the preprocessing tasks to the stateless AI worker nodes in the inference process pool according to the polling algorithm to obtain a preprocessed video frame sequence with facial semantic segmentation mask and bounding box coordinates.

[0009] The control module is used to perform sound activity detection and recognition of the user's speaking state in the Golang side control layer according to the preprocessed video frame sequence, and to perform H.264 compression encoding and Opus compression encoding on the preprocessed video image frames and audio data blocks respectively. At the same time, it performs throttling algorithm control according to the current session connection state to obtain the encoded audio and video streams and session control instructions.

[0010] The alignment module is used to establish a bidirectional streaming transmission channel between the Python side and the Golang side based on the gRPC protocol according to the encoded audio and video streams, session control instructions and subtitle text, and to perform unified timestamp alignment and packaging of the encoded audio and video streams and subtitle text to obtain time-consistent multimodal streaming data packets;

[0011] The presentation module is used to establish a real-time peer-to-peer connection through the WebRTC protocol based on multimodal streaming data packets, and synchronously write audio streams, video streams and subtitle streams into the WebRTC media track to obtain a multimodal real-time digital human interactive screen that is synchronously pushed to the web front end.

[0012] In a second aspect, a computing device includes:

[0013] One or more processors;

[0014] A storage device for storing one or more programs that, when executed by one or more processors, enable the one or more processors to implement the system.

[0015] Thirdly, a computer-readable storage medium storing a program that, when executed by a processor, implements the system.

[0016] The above-described solution of the present invention has at least the following beneficial effects:

[0017] By clearly defining the responsibilities of the Python and Golang sides, AI inference-intensive tasks are separated from high-concurrency network control tasks. Combined with gRPC bidirectional streaming and a unified timestamp alignment mechanism, multimodal synchronization of audio, video, and subtitles is achieved. Simultaneously, the system utilizes a polling algorithm for load balancing of the inference process pool, along with throttling algorithms and adaptive buffering mechanisms, effectively ensuring resource stability and low-latency response capabilities in multi-user concurrent scenarios. Furthermore, the Golang control layer avoids screen flickering and facial abrupt changes during state transitions by locating anchor frames and smoothly transitioning between silent and inference videos. Finally, the time-consistent multimodal streams are synchronously pushed to the web frontend via the WebRTC protocol, improving the fluency, realism, and overall scalability and reliability of real-time digital human interaction. Attached Figure Description

[0018] Figure 1 This is a flowchart illustrating a Python-Golang dual-architecture multimodal real-time digital human customer service system based on microservice deployment, as provided in an embodiment of the present invention.

[0019] Figure 2 This is a schematic diagram of the control module provided in an embodiment of the present invention. Detailed Implementation

[0020] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0021] like Figure 1 As shown, embodiments of the present invention propose a Python-Golang dual-architecture multimodal real-time digital human customer service system based on microservice deployment, including:

[0022] The generation module is used to obtain user multimodal interaction requests and sequentially perform speech recognition, intelligent agent reasoning, text-to-speech and digital human animation generation processing on the Python side to obtain multimodal reasoning results;

[0023] The scheduling module is used to synthesize video image frames into a basic video sequence on the Python side based on the video image frames in the multimodal inference results, perform frame-by-frame face recognition and facial semantic segmentation on the basic video, and allocate the preprocessing tasks to the stateless AI worker nodes in the inference process pool according to the polling algorithm to obtain a preprocessed video frame sequence with facial semantic segmentation mask and bounding box coordinates.

[0024] The control module is used to perform sound activity detection and recognition of the user's speaking state in the Golang side control layer according to the preprocessed video frame sequence, and to perform H.264 compression encoding and Opus compression encoding on the preprocessed video image frames and audio data blocks respectively. At the same time, it performs throttling algorithm control according to the current session connection state to obtain the encoded audio and video streams and session control instructions.

[0025] The alignment module is used to establish a bidirectional streaming transmission channel between the Python side and the Golang side based on the gRPC protocol according to the encoded audio and video streams, session control instructions and subtitle text, and to perform unified timestamp alignment and packaging of the encoded audio and video streams and subtitle text to obtain time-consistent multimodal streaming data packets;

[0026] The presentation module is used to establish a real-time peer-to-peer connection through the WebRTC protocol based on multimodal streaming data packets, and synchronously write audio streams, video streams and subtitle streams into the WebRTC media track to obtain a multimodal real-time digital human interactive screen that is synchronously pushed to the web front end.

[0027] In this embodiment of the invention, by clearly defining the responsibilities of the Python and Golang sides, AI inference-intensive tasks are separated from high-concurrency network control tasks. Combined with gRPC bidirectional streaming transmission and a unified timestamp alignment mechanism, multimodal synchronization of audio, video, and subtitles is achieved. At the same time, the system uses a polling algorithm to balance the load of the inference process pool, and with the help of a throttling algorithm and an adaptive buffering mechanism, it effectively ensures resource stability and low-latency response capabilities in multi-user concurrent scenarios. In addition, the Golang control layer avoids screen flickering and facial abrupt changes during state switching by anchor frame positioning and smooth transition processing of silent video and inference video. Finally, the time-consistent multimodal stream is synchronously pushed to the WEB front end via the WebRTC protocol, improving the fluency, realism, and overall scalability and reliability of real-time digital human interaction.

[0028] In a preferred embodiment of the present invention, the generation module may include:

[0029] In this embodiment of the invention, user audio streams are extracted from user multimodal interaction requests. Acoustic feature extraction and speech endpoint detection are performed on the user audio streams. After dividing the speech activity intervals, the acoustic features are input into a pre-trained acoustic feature mapping network to obtain the user text sequence corresponding to the user audio stream. Specifically, the system monitors the front-end user multimodal interaction request link in real time. When a user interaction request data packet is captured, the protocol encapsulation header, tail check bit, and redundant padding bits inside the data packet are stripped. The original time-domain audio data stream collected in real time by the user terminal microphone is extracted without loss. To unify the audio operation standards of the entire system, the original audio stream is subjected to standardization processing, and all irregular audio standards are forcibly normalized to a 16kHz sampling rate, mono, and 16 The bit-quantized PCM raw audio format eliminates the inconsistency issues caused by differences in acquisition equipment. Based on this, DC component elimination is performed. All audio sampling points are traversed frame by frame, and the arithmetic mean of the amplitudes of all samples in the entire audio segment is used as the DC offset. This DC offset is subtracted from the original amplitude of each sampling point to eliminate the inherent DC noise of the hardware acquisition equipment, resulting in a zero-offset, high-purity standardized time-domain speech signal. Natural speech signals are non-stationary time-domain signals with irregular global statistical characteristics, remaining stable only for short periods. Therefore, a short-time framing mechanism must be used for decomposition and calculation. The system is fixed with a speech frame duration of 25ms and a frame shift interval of 10ms. The number of fixed sampling points per frame is calculated based on the sampling rate, using the following formula: ,in The total number of discrete sampling points in a single frame of audio is the benchmark for the total amount of data processed in a single frame of speech. The audio sampling rate is fixed at 16000Hz, representing the number of discrete speech samples collected per second. The duration of a single frame of audio is fixed at 0.025s. After calculation with fixed parameters, the number of sampling points per frame is constant at 400.

[0030] The continuous audio stream is truncated using a 10ms frame shift interval, with a 15ms overlap between adjacent audio frames to ensure uninterrupted and complete speech temporal features. To address spectral leakage caused by abrupt changes in signal amplitude at the beginning and end of a single audio frame, a Hanning window weighted smoothing operation is applied to each truncated audio frame. The Hanning window calculation formula is as follows: ,in For the first frame The smoothing weight coefficient of each sampling point is used to weaken the abrupt change signal at the beginning and end of the frame and strengthen the effective signal in the center of the frame. Its value is fixed and ranges from 0 to 1. The weight of the sampling points at the beginning and end of the frame approaches 0, and the weight of the sampling points in the center of the frame approaches 1, which can realize the smooth transition constraint of the audio frame signal. For traversing the discrete sampling points within the frame; The total number of discrete sampling points for a single frame of audio is used to constrain the window function's operation range. This weighted operation ensures a smooth transition in the amplitude of the first and last frames of audio signal.

[0031] For each standard audio frame that has undergone windowing and smoothing, a Fast Fourier Transform (FFT) is performed to convert the time-domain speech signal into a frequency-domain energy distribution signal, unlocking the timbre, pitch, and frequency characteristics of the speech. The FFT calculation formula is as follows: ,in For the first The complex spectral values ​​corresponding to each frequency component simultaneously carry frequency energy and phase information; It is a frequency index in the frequency domain, used to distinguish different frequency levels; It is the traversal index of discrete temporal domain sampling points within a single frame of audio; The original time-domain sampling point represents the speech vibration amplitude. Smoothing weights for the Hanning window; The imaginary unit is used to construct a frequency domain complex number operation system to complete the orthogonal transformation from the time domain to the frequency domain. The power spectrum of each frequency component is obtained by taking the square of the modulus of the transformed complex spectrum, characterizing the energy intensity of speech at different frequencies. This is then connected to 24 sets of stepped triangular Mel filter banks to rigorously simulate the nonlinear characteristics of human hearing, refining the preservation of low-frequency speech details and compressing and filtering high-frequency redundant noise, retaining only the energy of the effective speech segments perceptible to the human ear. Logarithmic energy is calculated filter by filter to complete the nonlinear mapping of linear energy characteristics. The calculation formula is as follows: ,in For the first Logarithmic energy characteristics of the output of a Mel filter; For filter bank index; For the first Each filter corresponds to a specific frequency range; This represents the voice power value at the corresponding frequency point.

[0032] A discrete cosine transform is performed on the 24-dimensional logarithmic energy feature to remove linear redundancy between feature dimensions and compress the feature volume. A 13-dimensional core Mel-frequency cepstral coefficient is extracted as the basic acoustic feature. To compensate for the lack of temporal dynamic features in speech, first-order and second-order difference features are additionally calculated to characterize the short-time change rate and acceleration of speech, respectively. Finally, these are concatenated to form a complete, highly refined, and temporally complete 39-dimensional acoustic feature vector, serving as the core data basis for speech judgment and recognition. A dual-dimensional joint judgment mechanism of short-time energy and short-time zero-crossing rate is adopted. All judgment thresholds have been calibrated through tens of thousands of tests in indoor office customer service scenarios, possessing a fixed numerical range and clear applicable basis. It can distinguish between silent low-noise, environmental noise, and effective human voice. The system solidifies the three types of judgment thresholds. The numerical range and judgment criteria are as follows: First, the low-energy threshold for silence, with a normalized numerical range of 0.02-0.05. This range matches the energy distribution range of indoor equipment background noise and static air noise, serving as the benchmark for determining the silence range. Second, the high-energy threshold for human voice, with a normalized numerical range of 0.15-0.40. This range fully covers the effective speech energy of close-range human speech and can distinguish between human voice and environmental noise. Third, the zero-crossing rate threshold for human voice, with a single-frame numerical range of 35-60 times. This range represents the inherent frequency of the natural human voice waveform crossing the zero point. The zero-crossing rate for silence noise is consistently below 30 times / frame, while the zero-crossing rate for high-frequency sharp noise is consistently above 65 times / frame, enabling the non-overlapping distinction between noise and human voice. The short-time energy of the audio is calculated frame by frame to determine the speech energy level. The calculation formula is as follows: .

[0033] in For the first The short-time total energy of a frame audio represents the overall vibration intensity of a single frame of speech. For audio frame timing index; For the first Frame number Amplitude data at each sampling point; calculate the short-time zero-crossing rate of the audio frame by frame to determine the frequency characteristics of the speech; the calculation formula is as follows. ,in For the first Number of zero-crossings of the frame audio waveform; The amplitude of the previous adjacent sampling point is used for comparison to determine the waveform flipping state. The system executes continuous frame linkage judgment logic to prevent single-frame misjudgment. Five consecutive audio frames simultaneously satisfy the condition that the short-time energy is between 0.15 and 0.40 and the zero-crossing rate is between 35 and 60 times / frame, which is determined to be the speech start point and effective speech truncation is started. Eight consecutive audio frames simultaneously satisfy the condition that the short-time energy is below 0.05 and the zero-crossing rate is below 30 times / frame, which is determined to be the speech end point and speech truncation is terminated. Through this mechanism, a complete, noise-free, and frame-free effective speech temporal feature sequence is extracted. The extracted clean and effective 39-dimensional temporal feature sequence is completely input into the acoustic feature mapping operation architecture with solidified parameters. The input layer of this architecture is adapted to 39-dimensional feature input and supports a maximum of 1000 frames of long temporal sequence. The system employs a 3×1 one-dimensional convolutional computation unit to fuse local features and increase dimensionality, enhancing local speech details. It stacks eight layers of temporal coding computation units, each configured with a 16-head attention computation structure and a 2048-dimensional forward computation structure, capturing long-distance temporal dependencies in speech. This adapts to the speech features of customer service scenarios with uneven speech rates, connected speech, soft tones, and irregular phrasing. Layer normalization and random deactivation regularization operations are configured at the end of the computation to avoid computational redundancy and overfitting. Finally, probability normalization operations output the probability distribution sequence of Chinese characters and punctuation marks. A beam search decoding logic with a beam width of 5 is used to traverse all character combination probabilities, filter the global final semantic combination, and automatically correct speech recognition errors such as phrasing errors, missed soft tones, and misidentification of connected speech, ultimately outputting a standardized user text sequence.

[0034] The user text sequence is input into an agent reasoning service instance deployed on the Python side to obtain customer service response text matching the user's intent. Specifically, this includes: fully normalizing the input raw user text sequence, automatically filtering illegal special characters, garbled characters, extra spaces, and line breaks, unifying full-width and half-width character formats, splitting extremely long and complex sentences into standard short sentences conforming to Chinese grammar, correcting word order errors and redundant words, ensuring the input text is semantically pure, formatted uniformly, and grammatically correct, and adapting to semantic operation input requirements; the agent semantic operation system deployed on the Python side microservice consists of an embedding operation layer, a bidirectional encoding operation layer, a one-way decoding operation layer, and a post-processing operation layer. The embedding operation layer converts each word in the normalized text into a 768-dimensional high-density semantic vector, simultaneously carrying word semantics, positional temporal sequence, and sentence structure features; stacking 12 layers of bidirectional temporal encoding operation units performs global semantic parsing of the entire sentence text, deeply... The system captures user inquiry intent, core demands, question keywords, contextual logic, and business scenario attributes; it stacks 12 layers of unidirectional decoding operation units, combines them with the weight parameters of the customer service business knowledge base, and iteratively generates a probability sequence of reply text word by word; the operation system has a built-in intelligent customer service business knowledge base that covers all five scenarios: business inquiry, process query, problem handling, complaint feedback, and casual conversation. During the operation, it prioritizes classifying user intent, determines the user's interaction scenario, and then calls the corresponding scenario's dedicated response knowledge base to match the final response logic and script template, generating the original customer service reply text sequence word by word; the original generated text undergoes grammatical correction, semantic completion, redundancy removal, and tone standardization to unify the gentle, standardized, and rigorous language style of human customer service, eliminate colloquial, fragmented, and non-standard expressions, correct word order defects, and complete sentence integrity, ultimately generating standardized customer service reply text.

[0035] The customer service response text is input into the speech synthesis processing unit. Text regularization and phoneme conversion are performed on the response text. After extracting prosodic features, a neural vocoder synthesizes a speech waveform corresponding to the phoneme sequence, resulting in a customer service audio data block carrying the customer service voice data. Specifically, this includes: for non-standard text content such as numbers, letters, abbreviations, times, amounts, and special symbols, forced rule conversion is performed to replace them with standard Chinese characters. For example, Arabic numerals are converted to Chinese uppercase numerals, time symbols are converted to Chinese time period descriptions, and amount symbols are converted to Chinese amount descriptions. Simultaneously, word segmentation, semantic sentence segmentation, and punctuation are performed. The system distinguishes semantic pauses at the beginning, middle, and end of sentences, marks stressed words, and generates a standardized text sequence adapted for speech synthesis. Through a Mandarin phoneme conversion operation system with parameter convergence optimization, the system breaks down the standardized text character by character, matching the standard initial consonant, final vowel, tone, and unstressed syllable attributes of each Chinese character. Combined with Chinese semantic logic, it marks short pauses within sentences, long pauses between sentences, and stressed syllable positions, generating a complete phoneme temporal sequence with temporal, pause, and stress attributes, conforming to the pronunciation rules of natural Mandarin speakers. Based on the phoneme temporal sequence, it quantitatively calculates the four core prosodic parameters: pitch, intensity, duration, and pause.

[0036] The formula for fine-tuning pitch (fundamental frequency) is as follows: ,in for The instantaneous fundamental frequency of speech at any given moment represents the pitch of the speech at that moment. The base frequency for customer service voice is fixed at 180Hz, which is the standard base tone for adult female customer service representatives to ensure a consistent and stable overall tone. This is a word weighting coefficient used to quantify the semantic importance level of words. Core content words are assigned a value of 0.7-1.0, while ordinary function words are assigned a value of 0.2-0.6. The importance score for words is dynamically generated by business semantic rules, and the score is output in the range of 0-1 according to the proportion of words in the core semantics of the whole sentence. This is a sentence structure correction coefficient, adapted to the intonation fluctuations of different sentences, with a fixed subdivision value range: 0.8-1.0 for declarative sentences, 1.1-1.3 for interrogative sentences, and 1.2-1.4 for exclamatory sentences. The emotional smoothness coefficient is fixed at 1.0 for customer service scenarios to ensure a gentle tone without excessive fluctuations in voice; the formula for fine-tuning sound intensity (volume) is as follows: ,in for The instantaneous amplitude of speech volume at any given moment represents the loudness of speech; The normalized volume is fixed at 1.0, serving as a unified volume benchmark. This is the volume adjustment coefficient, with a fixed value of 0.1-0.3, which controls the dynamic range of volume changes and avoids sudden volume changes. The parameters for phoneme pronunciation fullness are 0.8-1.0 for fully pronounced phonemes and 0.2-0.5 for unstressed or weakly pronounced phonemes. The rules for adaptive adjustment of sound length and pause duration are as follows: vowel phonemes are extended by 20% by default, consonant phonemes are shortened by 10% by default, core content words are extended, and function words are shortened. Punctuation pauses are strictly graded: commas are paused for 80ms, semicolons for 120ms, and periods and question marks for 200ms, which conforms to the rhythm of natural human speech. All temporal prosodic parameters are spliced ​​frame by frame to form a full-dimensional prosodic feature vector sequence.

[0037] The HiFi-GAN high-precision neural vocoder architecture, optimized through iterative convergence, is progressively composed of five parts: a feature fusion preprocessing unit, a multi-scale residual generation unit, a multi-level adversarial discrimination unit, a waveform refinement and correction unit, and an audio normalization encapsulation unit. This architecture maps abstract prosodic features to high-fidelity natural human voice time-domain waveforms. It performs time-sequence alignment and fusion of the standard phoneme time sequence output from the second step with the multi-dimensional continuous prosodic feature vector sequence generated in the third step. Using the phoneme timestamp as a reference, the pitch, intensity, duration, pauses, and stress features at the corresponding moments are embedded into the phonemes. Within the feature dimension, a fusion feature matrix with unified dimensions, one-to-one temporal correspondence, and complete feature coupling is constructed. At the same time, the fusion feature matrix is ​​normalized and smoothed to eliminate feature mutation points, dimensional amplitude deviations, and temporal jitter issues, ensuring that the feature sequence input to the vocoder is continuous, stable, and free of abnormal noise. The generation operation unit is the core generation structure of the vocoder. A multi-level progressive one-dimensional convolutional residual upsampling architecture is adopted to avoid the problem of missing details in single-level upsampling. Four-level progressive upsampling layers are configured to improve the temporal resolution step by step, and finally the low-dimensional semantic prosodic features are mapped to a 16kHz high-resolution temporal domain original speech waveform.

[0038] Each upsampling level is configured with a residual convolution structure and normalized activation operation. Residual jump connections preserve the detailed features of previous layers. Simultaneously, multi-scale convolutional kernels are configured for parallel operation to capture the global prosodic contour, mid-frequency pronunciation details, and high-frequency timbre texture of the speech, balancing the overall fluency of the sentence with the precision of individual word pronunciation. Structurally, this ensures the naturalness and fidelity of the generated waveform. The discrimination unit adopts a multi-scale parallel discrimination architecture, including three independent verification dimensions: a global waveform discrimination branch, a spectral detail discrimination branch, and a prosodic temporal discrimination branch. This comprehensively verifies the compliance and authenticity of the generated speech waveform. The global waveform discrimination branch is responsible for verifying the overall... The waveform profile conforms to the distribution pattern of natural human voice, eliminating waveform distortion, discontinuity, and abnormal fluctuations. The spectral detail discrimination branch converts the generated waveform into a frequency domain spectrum, compares it with the spectral characteristics of real customer service voices, and corrects issues such as missing high-frequency details, low-frequency timbre distortion, and frequency band energy imbalance. The prosodic timing discrimination branch verifies the matching degree between waveform timing changes and input prosodic parameters, phoneme duration, and pause nodes frame by frame, correcting issues such as intonation misalignment, uneven speech rate, and stress shift. Through adversarial iterative operations between the generation unit and the discrimination unit, the waveform details are continuously optimized, gradually reducing the feature deviation between the generated speech and real human customer service speech, and completing the refined waveform correction.

[0039] For the original speech waveform generated by adversarial iteration, although its overall contour is normal, it has subtle defects such as high-frequency fine spikes, instantaneous impulse noise, abnormal local amplitude jumps, and abrupt waveform transitions between frames. A customized adaptive threshold filtering algorithm and a frame-by-frame temporal smoothing algorithm are used to optimize and correct the waveform layer by layer and in partitions. The specific execution logic is as follows: First, adaptive threshold high-frequency noise reduction and abnormal noise removal operations are performed. Unlike the one-size-fits-all defect of fixed threshold filtering, a temporal local statistical adaptive threshold judgment mechanism is adopted. Taking the current waveform sampling point as the center, 10 sampling points before and after are extracted to construct a local waveform statistical window. The mean and standard deviation of the waveform amplitude within the window are calculated in real time to dynamically generate the instantaneous noise judgment threshold. The amplitude fluctuation range of normal human voice waveform is controlled within the range of local mean ± 1.5 times the standard deviation. Instantaneous amplitude points that exceed this range are judged as impulse noise, abnormal amplitude noise, and high-frequency spike noise. For abnormal noise points that have been judged, they are not directly set to zero, which would cause waveform distortion. The first method is to use local neighborhood interpolation correction, which replaces the amplitude of abnormal points by weighted averaging of the amplitudes of the effective normal sampling points before and after. The second method is frame-by-frame temporal smoothing and waveform connection correction. Neural vocoders are prone to problems such as discontinuous amplitude at frame boundaries, abrupt changes in local waveforms, and gaps in connection when generating waveforms frame by frame, resulting in harsh pronunciation and a stuttering sound. The sliding window temporal smoothing algorithm is used, with a 25ms smoothing window and a 5ms window sliding step size. The entire speech waveform is traversed segment by segment. For problems such as abrupt changes in amplitude, sudden changes in waveform slope, and trajectory breaks at the connection position of adjacent audio frames, linear interpolation combined with cosine smoothing correction logic is used to reconstruct the inter-frame transition waveform trajectory, transforming abrupt amplitude jumps between frames into continuous and smooth gradual curves. At the same time, the smoothing operation is strictly bound to the phoneme pronunciation logic, only correcting invalid inter-frame abrupt noise, without modifying the intonation fluctuations, volume changes, and stress amplitude characteristics of normal pronunciation, ensuring that the original rhythm, timbre, and speech rate of the speech remain unchanged, and only optimizing the temporal continuity and smoothness of the waveform.

[0040] After noise reduction and smoothing correction, the system trims invalid and redundant silent intervals, strictly retains the preset standard pause duration, and eliminates problems such as excessively long silences and abrupt sound cutoffs. It unifies the dynamic range of waveform amplitude, limits the maximum and minimum amplitude ranges, avoids sudden changes in local volume, and ensures that the entire audio segment has a stable volume, uniform timbre, no distortion, and no noise. The optimized complete continuous speech waveform is evenly sliced ​​into blocks with a fixed 20ms block duration to ensure that the audio timing length of each block is uniform and the frame structure is regular, adapting to the timing alignment requirements of subsequent digital human video frames. All audio blocks are uniformly packaged in a standard PCM raw audio format with a 16kHz sampling rate, mono, and 16-bit quantization. Abnormal data blocks, incomplete frames, and noisy frames are removed, and finally, standardized customer service audio data blocks are output.

[0041] The customer service audio data block and customer service reply text are synchronously input into the digital human image synthesis processing unit. The displacement of key points on the digital human's face is controlled based on the speech amplitude and phoneme duration of the customer service audio data block, resulting in video image frames synchronized with the lip movements of the customer service speech data. The customer service reply text is then converted into subtitle text by aligning it word-by-word according to the audio timestamp, yielding multimodal inference results. Specifically, this includes: unpacking the 16kHz, mono, 16-bit standard PCM customer service audio data block frame by frame; performing time-sequential frame segmentation of the continuous audio stream according to a system-preset fixed frame length of 20ms; and ensuring each audio frame contains 320 discrete sampling points. To ensure uniform frame length, consistent sampling point count, and uniform temporal granularity, all audio frame data is traversed frame by frame to achieve high-precision extraction of four core temporal driving parameters: First, instantaneous normalized amplitude, obtained by calculating the root mean square of the amplitudes of all sampling points in a single frame and normalizing it to the 0-1 range, used to characterize the intensity of sound emission in a single frame; second, phoneme duration, bound to the previous phoneme conversion result, recording the start and end millisecond timestamps and effective duration of the phoneme in the current audio frame, anchoring the pronunciation cycle of single words and single sounds; third, pronunciation type label, classifying and marking the phonemes in the current frame to distinguish between different phonemes. Six major pronunciation types—open vowels, closed vowels, rounded vowels, flattened vowels, neutral vowels, and pauses / silences—provide a type basis for subsequent lip shape weighting matching; fourth, a global time-series timestamp is used, with the acquisition time of the first valid sampling point of the audio data stream as the absolute time zero point (t=0ms). The frame duration is accumulated frame by frame, assigning a unique, incremental, and non-repeating global absolute timestamp to each audio frame, marking the temporal position of each frame; based on the temporal parameters extracted frame by frame, a continuous and uninterrupted global temporal scale system is constructed to fill the micro-time-series gaps between frames, eliminate the temporal discrete errors caused by frame segmentation, and address audio silence transition frames and phoneme segmentation. Transition frames and inter-sentence pause frames are used for temporal interpolation to complete the audio timeline, ensuring that the entire audio timeline is continuous, complete, uninterrupted, time-skipping, and deviation-free from the start to the end of the speech. At the same time, a two-layer temporal mapping system is established, including an absolute global timeline and a relative frame temporal index. The absolute timeline is used for millisecond-level alignment of subtitles, and the relative frame index is used for frame-by-frame driving rendering of digital human animation. The completed global audio temporal reference axis is set as the unique temporal scale for the entire link, forcing the subsequent digital human facial animation rendering frame rate, video frame temporal arrangement, subtitle refresh timing, and image composition timing to all be synchronously calibrated by this timeline.

[0042] The system predefines 468 refined core key points for the digital human face, covering the entire area of ​​the lips, corners of the mouth, eyes, jaw, and cheeks. The original baseline 2D coordinates of all key points in a static state are pre-calibrated. Based on audio driving parameters, the system calculates the real-time displacement coordinates of the key points frame by frame using dynamic formulas, matching lip shape changes during pronunciation. The core calculation formula is as follows: ,in for Real-time two-dimensional coordinates of key facial features; The key point is the silent reference coordinate, which is the fixed point coordinate of the digital human when it is stationary and does not make a sound; The scaling factor for facial movements is set to match the gentle persona of a customer service representative, with a fixed value of 0.4-0.6 to ensure that the movements are natural, not exaggerated, and not stiff. This is the normalized amplitude of the audio in the current frame, ranging from 0 to 1. The larger the amplitude, the greater the lip opening and closing amplitude. The duration of the current phoneme is used to control the duration and rate of change of facial movements; For each phoneme, a unique matching coefficient is used, with different initials and finals corresponding to specific displacement weights, accurately matching various pronunciation patterns such as open mouth, closed mouth, rounded lips, and flat lips. The system traverses all 468 facial key points frame by frame, updating their coordinates in real time, and continuously driving the digital human to complete human-like dynamic expressions such as lip opening and closing, slight mouth movements, jaw rise and fall, and slight blinking. This ensures that the lip shape and facial movements corresponding to each phoneme are synchronized with the rhythm of speech pronunciation, without any leading, lagging, or misalignment, ultimately generating a continuous video image frame sequence with a standard frame rate of 25 FPS, smooth movements, and a high degree of accent matching. Based on the global audio timeline, the customer service reply text is broken down word by word and bound to the start and end timestamps of the corresponding phonemes to construct the text. The system maps the audio to a temporal sequence, sets differentiated display rules for different sentence forms, refreshes the display of single characters instantly following single sounds, continuously displays words following the pronunciation range of complete words, uniformly adds punctuation marks at the end of sentences, calibrates the timing of subtitle display frame by frame, automatically corrects subtitles that are ahead, behind, or flickering, and generates temporal subtitle text corresponding to each frame of video; it unifies and integrates the data from the entire processing chain, including the retained user's original audio stream, standardized customer service audio data blocks, lip-synced video image frame sequences, and temporally aligned subtitle text sequences, performs temporal consistency verification, format uniformity standardization, and content matching verification on all data, automatically removes abnormal frames, empty data frames, and temporally disordered data, and finally generates standardized multimodal inference results.

[0043] It avoids screen flickering, jumps, and gaps when switching digital human states, improves the human-like naturalness and visual smoothness of digital human interaction, and adapts to high-concurrency, high-traffic intelligent customer service scenarios.

[0044] In a preferred embodiment of the present invention, the scheduling module may include:

[0045] In this embodiment of the invention, video image frames are extracted from the multimodal inference results. The generation timestamp of each video image frame is read, and the video image frames are arranged into a frame sequence in ascending order of generation timestamps. The frame sequences are then concatenated and synthesized into a continuous basic video sequence using a preset synthesis frame rate. Specifically, this includes: performing layered parsing on the encapsulated data of the multimodal inference results, stripping away irrelevant content such as audio data streams, subtitle timing data, and multimodal verification redundant data, and separating all original data of the video image frames. Each video image frame is bound to a globally unique generation timestamp during the generation stage. This generation timestamp inherits from the aforementioned globally unified audio timing reference axis and is a millisecond-level high-precision absolute timestamp, which can uniquely identify the generation timing position of a single video frame. It does not have repetitive, missing, or abrupt timing characteristics and is the sole reference for video frame timing sorting. An insertion-based stable sorting algorithm is used to complete the global frame sequence rearrangement, and the discrete video frame set is set as... ,in The total number of video frames extracted so far, for each frame Corresponding to a unique globally generated timestamp The formula for calculating the core alignment in the sorting process is as follows: ,in This is a frame position shift determination flag, used to control frame sequence insertion and shift operations; For the sorted interval, the first... Global generation timestamp of the frame; A global timestamp is generated for the current frame to be inserted. When the judgment result is 1, a frame shifting and insertion operation is performed; when the judgment result is 0, the original ordered position is retained. The complete sorting iteration calculation process is as follows: the first video frame is used as the initial ordered reference sequence to construct an initial ordered frame set. Starting from the second video frame, the frames to be sorted are taken out one by one and compared with the timestamps of all the preceding sorted frames. If the timestamp of the current frame to be sorted is less than the timestamp of the already sorted frames, the already sorted frames are shifted one position to the right, and the comparison is continued until a valid insertion position is found. The current frame to be sorted is inserted into the corresponding time sequence position, ensuring that the timestamps of all preceding frames are less than or equal to the timestamp of the current frame. The process is repeated until all discrete video frames have been inserted and sorted.

[0046] The preset composite frame rate is fixed at 25 FPS. This frame rate parameter has been determined through extensive engineering testing and calibration in customer service digital human interaction scenarios. It is adapted to the 16kHz standard customer service audio timing granularity, the persistence of vision in the human eye, and the computing power threshold of real-time inference devices. It can balance video smoothness, multimodal timing synchronization accuracy, and real-time inference efficiency of the system. It is the final fixed frame rate parameter adapted to the entire chain operation of this system. The theoretical display duration of a single frame corresponding to the 25 FPS standard frame rate is fixed at 40ms. The formula for calculating the duration of a single frame is as follows: ,in The standard display duration for a single frame is in milliseconds (ms). The system presets a composite frame rate of 25 FPS, which, when calculated, yields a standard frame duration of 40 ms, providing a timing benchmark for timing regularization. A unified timing normalization calibration is performed to address the minute time intervals, sampling time differences, and generation time differences between frames. Pixel similarity comparison is used to remove a very small number of redundant and duplicate frames. For timing gaps, weighted interpolation between adjacent frames is used to fill in the gaps. The weighted interpolation calculation formula is as follows: ,in The pixel matrix of the interpolated frame image; This is the pixel matrix of the preceding valid video frame at the time-series gap position; This represents the pixel matrix of the subsequent valid video frames at the time-series gap position; , The weighting coefficients for the preceding and following frames satisfy the following conditions: Unified value =0.5、 =0.5, ensuring a natural transition between interpolated frames without abrupt distortion; through frame rate unification constraints and temporal completion regularization, the discrete and ordered frame sequences are ultimately integrated into a continuous basic video sequence.

[0047] Face region localization and detection are performed frame-by-frame on a continuous basic video sequence. In each video frame, the bounding rectangle of the face is determined, and the coordinates of the bounding box of the face region are obtained. Pixel-level facial semantic segmentation is then performed within the face region defined by the bounding box coordinates. Each pixel within the face region is assigned a semantic category to the left eye region, right eye region, nose region, mouth region, facial contour region, and non-face background region, respectively. A semantic category label is assigned to each pixel, resulting in a facial semantic segmentation mask corresponding to the current frame. Specifically, this includes configuring a lightweight face detection computing architecture adapted to digital human face features. The overall computing architecture consists of five parts: an input adaptation layer, a backbone feature extraction layer, a feature fusion and aggregation layer, a detection and regression computing layer, and a result output layer. The input adaptation layer uniformly normalizes the size of single-frame images in the continuous basic video sequence, uniformly outputting a three-channel RGB normalized image of 640×640 pixels to ensure... Subsequent feature operations are scaled uniformly. The backbone feature extraction layer adopts a multi-layer lightweight residual convolution stacked operation structure, relying on the combination of convolution kernel sliding convolution operation, pooling downsampling operation, and residual jump connection operation to complete the layered extraction of shallow texture details, mid-level facial contour structure features, and high-level global semantic features of the face. This can completely preserve the detailed features of the digital face's facial structure, while effectively suppressing invalid feature interference caused by image background interference, light and shadow differences, and image quality noise, thus purifying the effective feature information of the face. The feature fusion and aggregation layer performs feature splicing and weighted fusion processing on the multi-scale layered extracted face features, strengthening the weight of the core features of the face target and weakening the weight of invalid background features, thereby achieving effective differentiation between face features and background features. The detection and regression operation layer, based on the fused high-precision face features, completes the effective confidence determination of the face target and the regression solution of the bounding box coordinates of the face.

[0048] During the overall computation, the system presets a face detection confidence threshold of 0.85. This threshold is designed for scenarios with stable image quality in digital human videos, minor dynamic changes in facial pose, simple and regular backgrounds, and uniform facial lighting. The final threshold was determined through multiple sets of gradient threshold comparison experiments. A threshold that is too low can lead to misidentification of image textures, background color blocks, and lighting edges as face targets, resulting in a large number of invalid detection areas and increasing the computational cost and risk of incorrect segmentation in subsequent semantic segmentation. A threshold that is too high will mask effective facial features under subtle changes in facial pose and slight lighting occlusion, leading to missed detections and misclassifications. After practical testing with gradients ranging from 0.70 to 0.95, a threshold of 0.85 balances the false positive and false negative rates, adapts to the standardized distribution of digital human video features, maximizes the retention of effective face detection areas, and filters out all low-confidence suspected interference areas, fake face areas, and background texture interference areas, ensuring the accuracy and stability of face localization. The bounding box regression operation outputs the coordinates of the face's bounding rectangle quadruple, calculated using the following formula: ,in This is the set of coordinates of the bounding rectangle of the face corresponding to the current video frame, used to uniquely define the effective pixel region of the face in the image; This represents the minimum horizontal pixel coordinates of the face region, corresponding to the pixel position of the leftmost boundary of the face. It represents the minimum vertical pixel coordinates of the face region, corresponding to the pixel position of the uppermost boundary of the face; This represents the maximum horizontal pixel coordinates of the face region, corresponding to the pixel position of the rightmost boundary of the face. The coordinates of the maximum vertical pixel in the face region correspond to the pixel position of the lowest edge of the face. This quadruple coordinates can be used to lock the effective pixel range of the face within a single frame image, dividing the entire frame image into the effective face processing area and the invalid background area.

[0049] The system employs an encoder-decoder symmetrical pixel-level facial semantic segmentation architecture. This architecture has been optimized through convergence with massive amounts of standardized digital human face image data, specifically adapted to the distribution patterns of facial features, facial texture features, and the positional structure of facial features in digital humans. It possesses high-precision and highly adaptable facial pixel classification capabilities. The architecture has six fixed semantic classification standards, corresponding to the left eye region, right eye region, nose region, mouth region, facial contour region, and non-face background region, respectively. The encoding operation performs multi-level deep feature extraction and feature dimension compression processing on the local facial image obtained by bounding box cropping. It deeply condenses refined semantic information such as facial feature position features, contour morphology features, surface texture features, and region boundary features, and eliminates redundant and invalid features to form a facial semantic feature matrix. The decoding operation performs progressive upsampling convolution restoration operations to progressively map and restore the compressed low-dimensional semantic feature matrix to the original image pixel size. During the semantic segmentation operation, the system performs a full-coverage and comprehensive traversal analysis operation on all pixels inside the bounding box. For each independent pixel, it outputs the probability distribution vector of its corresponding six semantic regions.

[0050] By using a maximum probability filtering rule, the semantic category with the highest probability value is selected as the unique label for the current pixel, achieving high-precision semantic classification for a single pixel. The overall classification operation strictly follows the spatial distribution rules of facial features in natural and digital faces, constraining the inherent semantic distribution logic of symmetrical distribution of the left and right eyes, centered distribution of the nose, distribution of the lower middle region of the mouth, and outer wrapping of the facial contour. This eliminates cross-regional misclassification and pixel label confusion at the operational rule level, ultimately achieving full semantic coverage labeling of all pixels within the effective area of ​​the face. Based on the semantic category labeling results of all pixels, the system performs connected component aggregation operations of pixels with the same semantic label, integrating and clustering adjacent pixels with the same semantic label to form structurally complete, well-defined, independent, and non-overlapping semantic regions for the left eye, right eye, nose, mouth, and facial contours. Pixels remaining within the face bounding box that cannot match the semantic features of facial features are uniformly classified as non-face background regions. Based on the global pixel label clustering results, a binary semantic image with the same frame size as the original video image and a one-to-one correspondence between pixel positions is generated, which is the facial semantic segmentation mask.

[0051] According to the polling algorithm, the facial semantic segmentation mask and bounding box coordinates of the current frame are packaged into preprocessing task data units. The current task queue length of each stateless AI worker node in the inference process pool is queried, and the preprocessing task data units are assigned to the stateless AI worker node with the shortest current task queue length. The stateless AI worker node performs morphological edge smoothing on the facial semantic segmentation mask to eliminate jagged discontinuous pixels on the mask edges, and performs pixel filling on connected component holes caused by semantic classification ambiguity inside the mask to obtain the processed facial semantic segmentation mask. Specifically, this includes: using a single frame of video image as the smallest processing unit, processing the facial semantic segmentation mask data, face bounding box quadruple coordinate data, and frame temporal coordinates obtained from the current frame. Data and image resolution parameters are integrated and structured to form independent and complete preprocessing task data units. The system constructs a distributed inference process pool, which deploys multiple sets of stateless AI worker nodes with equivalent performance, consistent parameters, and identical architecture. All worker nodes have the ability to independently receive tasks, perform calculations, and output results. There is no resource contention or data interaction between nodes. During scheduling, the scheduler polls and collects the current task queue length of all stateless AI worker nodes in real time. The task queue length is defined as the cumulative number of preprocessing task data units currently pending processing by the node. The queue length values ​​of all nodes are compared in real time, and the target worker node with the shortest queue length, the lowest load pressure, and the most idle resources is dynamically selected.

[0052] The preprocessing task data units that have been packaged are sent to the target worker nodes. A shortest queue-first polling scheduling logic is used to achieve dynamic and balanced allocation of computing resources in the process pool. After receiving the preprocessing task data units, the stateless AI worker nodes perform edge morphological smoothing and internal connected component hole filling processing. The first step is mask edge morphological smoothing. Due to pixel-level classification discreteness, the original semantic segmentation mask is prone to jagged pixels, breakpoints, concave-convex distortion, and local discrete noise at the edges. A 3×3 standard rectangular structural element is used to perform morphological opening operations on the mask. The opening operation is a combination of erosion and dilation operations. The erosion operation is used to remove small jagged edges, discrete and abrupt pixels, and edge breakpoint noise. The core calculation formula for the erosion operation is... ,in Coordinates after erosion calculation The pixel output value of the position; The pixel input value is the corresponding neighborhood position of the original mask image; It is the set of neighboring pixel coordinates corresponding to a 3×3 rectangular structural element, covering the eight surrounding neighboring pixels centered on the current pixel and the pixel itself, forming a complete nine-pixel neighborhood traversal window; , The horizontal and vertical relative offsets of the neighborhood are represented, with values ​​ranging from -1, 0, to +1. The dilation operation is used to repair the contour loss of the effective region caused by the erosion operation and to compensate for the shrinkage deviation at the edge of the effective region. The core calculation formula for the dilation operation is... ,in Coordinates after dilation operation The pixel output value of the position; This represents the pixel input value at the corresponding neighborhood position after the erosion operation.

[0053] Secondly, the mask's internal connected component hole filling process addresses the issue that, due to factors such as high pixel texture similarity in the face, semantic classification ambiguity, and light and shadow interference, closed, small blank connected components (pixel holes) are easily generated within the facial feature area of ​​the original mask. The preset hole area screening threshold is fixed at 10-50 pixels. This range was determined through multiple rounds of statistical experiments using digital facial semantic mask samples. Blank areas with a hole area less than 10 pixels are invalid, tiny noise holes caused by pixel classification discrepancies and must be filled. Blank areas with a hole area greater than 50 pixels are mostly natural hollow areas of facial features, areas obscured by real light and shadow, or gaps in normal structures; these are legal and valid blank areas and must be retained. The 10-50 pixel threshold range can distinguish between semantic classification defect holes and real facial structure holes. To construct blank areas while ensuring both the integrity of defect repair and the realism of facial structure, the working node employs an eight-connected domain traversal algorithm to scan and identify the entire mask area, locating all isolated and closed hole regions, and calculating the hole pixel area. For all invalid hole regions within a preset threshold range of 10-50 pixels, corresponding semantic pixel values ​​are filled according to the semantic category of the facial features surrounding the hole, ultimately obtaining the processed facial semantic segmentation mask. The eight-connected domain traversal algorithm is the core operational logic for hole recognition. Compared to the four-connected domain method, which only determines the four directional neighbors (up, down, left, right), the eight-connected domain adds four diagonal neighbor determination dimensions (upper left, upper right, lower left, lower right), which can completely cover all adjacent pixel relationships and identify various irregular closed holes. The eight-connected domain neighborhood offset set is defined as the center pixel. The set of eight neighboring coordinates is , , , , , , , It can completely identify all closed tiny holes inside the mask without omission or misjudgment, ensuring that the internal area of ​​the mask is dense and intact.

[0054] The processed facial semantic segmentation mask is reverse-mapped to the original pixel coordinate system of the corresponding video image frame according to the bounding box coordinates. The processed facial semantic segmentation mask and bounding box coordinates are then appended to the video image frame as metadata, resulting in a preprocessed video frame sequence with facial semantic segmentation mask and bounding box coordinates. Specifically, since the aforementioned face detection, semantic segmentation, and mask optimization are all based on the local image coordinate system after bounding box cropping, there is a fixed offset from the global pixel coordinate system of the original video image frame. Pixel alignment must be achieved through coordinate reverse mapping compensation. The specific formula is as follows: ,in The global horizontal pixel coordinates of the original video frame after mapping; The global vertical pixel coordinates of the original video frame after mapping; To optimize the local horizontal pixel coordinates corresponding to the masked image; To optimize the local vertical pixel coordinates of the masked image; Use the minimum horizontal offset coordinates of the face bounding box as a horizontal compensation constant; The minimum vertical offset coordinates of the face bounding box are used as vertical compensation constants.

[0055] Through the aforementioned binary coordinate compensation operation, each semantic pixel of the local mask is mapped to the global pixel position of the original video frame, offsetting the coordinate offset error caused by local cropping. This ensures that the optimized semantic mask region corresponds one-to-one with the pixels of the facial features region in the original video frame, with overlapping positions and matching boundary heights, without misalignment, offset, or deviation. Without altering the pixel color, image quality details, or image content of the original video image frame, the optimized facial semantic segmentation mask and aligned facial bounding box coordinate data, after coordinate mapping, are attached to the data structure of the corresponding video image frame in the form of an independent metadata channel. This forms a multi-layered fused data structure consisting of an image visibility layer, a semantic mask layer, and a coordinate metadata layer, achieving integrated storage and binding of image content and semantic information. The coordinate mapping and metadata attachment processing of all video image frames are completed frame by frame, and the frames are strictly rearranged according to the original temporal order to form a preprocessed video frame sequence.

[0056] Improve the parallel efficiency and real-time performance of batch video frame preprocessing, and repair mask morphological defects through morphological edge smoothing and hole filling to achieve alignment between semantic data and image data.

[0057] like Figure 2 As shown, in a preferred embodiment of the present invention, the control module may include:

[0058] In this embodiment of the invention, the user's original audio stream is extracted from the multimodal inference results and input to the sound activity detection module deployed in the Golang side control layer. Short-time energy calculation and zero-crossing rate analysis are performed on the user's original audio stream to determine whether the current audio frame is in a user speaking state, thus obtaining a sound activity status flag. Specifically, this includes: performing layered parsing on the encapsulated data of the multimodal inference results, stripping away irrelevant information such as video frame sequences, subtitle data, and synthesized customer service audio data, and extracting a complete, temporally continuous user's original audio stream. This audio stream is the original audio data received during the user's real-time interaction, retaining complete temporal features, noise features, human voice features, and silence gap features. It undergoes no post-processing noise reduction or waveform correction, and can accurately reflect the user's real-time speaking and pause states, providing original and authentic audio input data for sound activity detection. According to the source; to adapt to the requirements of short-time audio feature analysis, the continuous user raw audio stream is processed by short-time uniform framing, with a fixed single-frame audio duration of 20ms. There is no overlap or gap between frames, ensuring uniform temporal length and regular feature dimensions for each frame. The user raw audio stream uniformly adopts a 16kHz sampling rate and 16-bit sampling precision. Based on the fixed frame length, the fixed number of sampling points per frame is calculated to be 320. After framing, the audio temporal sampling amplitude is read and the structured matrix is ​​constructed frame by frame, fully constructing the single-frame audio temporal sampling matrix. The specific implementation process is as follows: for each independent audio segment after framing, all temporal sampling amplitudes are collected point by point in chronological order. The one-dimensional discrete temporal sampling data is then restructured to construct a dimensionally regular and temporally ordered single-frame audio temporal sampling matrix. The expression for the single-frame audio temporal sampling matrix is ​​as follows: ,in For the first The temporal sampling matrix corresponding to the frame audio; This represents the total number of audio sampling points per frame, with a fixed value of 320. For the first Frame audio number The original sampling amplitude at each time position.

[0059] Short-time energy is used to characterize the loudness intensity and energy concentration of a single frame of audio, and is a core feature parameter that distinguishes human voice from silent background. The formula for calculating short-time energy is: ,in For the first Short-time energy of frame audio; This represents the total number of sampling points in a single audio frame. For the first The first frame of audio The temporal amplitude of each sampling point is measured. By traversing all sampling point amplitudes frame by frame and performing square accumulation, the precise energy value of each audio frame is obtained. Valid human voice frames have significantly higher short-time energy values ​​due to the presence of effective speech amplitudes, while silent and background noise frames have amplitudes close to zero and extremely low short-time energy values, allowing for a preliminary distinction between valid human voice periods and silent periods. Next, the short-time zero-crossing rate of the audio frame is quantized. The short-time zero-crossing rate characterizes the frequency at which the audio temporal waveform crosses the zero amplitude axis, effectively distinguishing low-frequency background noise from human voice oscillations and compensating for the misjudgment defects of relying solely on short-time energy determination. The formula for calculating the short-time zero-crossing rate is as follows: ,in For the first Short-time zero-crossing rate of frame audio; , These are the temporal amplitudes of the current sampling point and the previous adjacent sampling point, respectively; the human voice waveform exhibits regular zero-crossing oscillation characteristics, with a stable zero-crossing rate value within a fixed range.

[0060] For digital human customer service human-computer interaction scenarios, gradient comparison experiments were conducted using tens of thousands of on-site human voice samples, environmental silence samples, and background noise samples to calibrate a dedicated fixed threshold for the audio acquisition parameters. This provides sufficient engineering basis and is suitable for most customer service interaction scenarios, including regular indoor sound recording, low-noise office environments, and slightly noisy environments. The preset short-time energy threshold for human voice is fixed at 800, and the preset effective short-time zero-crossing rate range for human voice is fixed at 25-80. The short-time energy threshold of 800 effectively filters low-energy interference from equipment background noise and weak environmental noise, preventing silent background noise from being misinterpreted as human voice; the 25-80 range... The zero-crossing rate range of 0 can match the waveform oscillation pattern of natural human voice, filtering out irregular low-frequency noise and pulse interference noise. When the short-time energy of a single frame audio is greater than the preset energy threshold of 800 and the short-time zero-crossing rate is within the effective range of human voice (25-80), the current audio frame is determined to be in the user speaking state, and a high-level effective sound activity status flag is generated. When the short-time energy of a single frame audio is lower than the threshold and the zero-crossing rate exceeds the effective range of human voice, the current audio frame is determined to be in the non-user speaking silent state, and a low-level invalid sound activity status flag is generated. After completing the determination frame by frame, a global sound activity status flag sequence with time alignment and continuous status is generated.

[0061] Based on the sound activity status flags, adaptive parameter adjustments are performed on the preprocessed video image frames in the preprocessed video frame sequence to obtain the target video frame sequence. Specifically, when the sound activity status flag indicates a user speaking state, the original resolution and frame rate of the preprocessed video image frames are maintained as the target video frame sequence; when the sound activity status flag indicates a non-user speaking state, the encoded resolution and frame rate of the preprocessed video image frames are reduced to obtain the target video frame sequence. This includes: using the frame time sequence index of the preprocessed video frame sequence as a basis, binding each preprocessed video image frame to a frame-by-frame generated sound activity status flag to ensure that each frame of video image... The status parameters are synchronized with the audio voice status of the corresponding time period, eliminating the problems of video parameter adjustment being ahead, behind, or misaligned, and providing a time-series matching basis for adaptive parameter adjustment; when the bound sound activity status flag indicates that the current time interval is in the user's speaking state, the current time period is determined to be a valid interaction time period. The user's interaction with the voice is the core observation content. Without compressing the image quality and frame rate, the system directly retains the original encoding resolution and original synthesized frame rate of the pre-processed video image frames, without performing any downgrading processing on the image pixel size or frame refresh rate, fully preserving the digital human facial details, facial feature dynamics, image clarity, and playback smoothness, and keeping the original parameters unchanged. Preprocessed video frames are directly incorporated into the target video frame sequence. When the bound sound activity status flag indicates that the current time interval is in a non-user-speaking, silent state, the current time period is determined to be an interactive idle period with no effective human voice interaction content. This reduces encoding computing power and transmission bandwidth usage without affecting visual experience. For silent interaction scenarios of digital human customer service, after multiple sets of subjective image quality evaluations, bandwidth usage comparisons, and smoothness stability comparison experiments, fixed degradation preset parameters adapted to idle period optimization were calibrated. The parameters take into account the engineering requirements of no image distortion residue, natural dynamic transitions, and extremely low bandwidth consumption. The specific preset standard parameters are based on the original video encoding resolution being uniform. The original composite frame rate is 25FPS and the encoding resolution is downgraded to 720P (1280×720) when the user is not speaking. At the same time, the playback frame rate is downgraded to 15FPS. According to the above preset standard parameters, the encoding resolution scaling down and frame rate extraction downgrade processing is uniformly performed on the preprocessed video image frames corresponding to the current time sequence. While ensuring that the digital human picture is smooth and distortion-free, the amount of encoded data per frame and the number of frames per unit time are reduced. The video frames with downgraded parameters are integrated into the target video frame sequence, and finally a complete target video frame sequence with adaptive state and dynamic parameter matching is generated.

[0062] The process involves performing H.264 compression encoding on the target video frame sequence to obtain an H.264 encoded video stream, and simultaneously performing Opus compression encoding on the customer service audio data block and segmenting it into audio data packets according to a preset duration to obtain an Opus encoded audio stream. The H.264 encoded video stream and the Opus encoded audio stream are then timestamped and encapsulated to obtain the initial encoded audio and video stream. Specifically, this includes: for the adaptively adjusted target video frame sequence, strictly adhering to the layered coding architecture of the H.264 video coding international standard, performing frame-by-frame, full-process standardized compression encoding operations; intra-frame prediction utilizing the strong correlation characteristics of adjacent pixels within a single frame video image, using already encoded pixels in the surrounding image to predict the current pixel block value, eliminating the need to store all original pixel data, achieving single-frame spatial redundancy compression; dividing the single-frame video image into fixed-size 4×4 pixel macroblocks; traversing multiple intra-frame prediction modes for each macroblock, calculating the pixel prediction residual under each mode, and selecting the final prediction mode with the smallest residual to complete pixel prediction; the formula for calculating the intra-frame prediction residual is as follows. ,in coordinates Intra-frame prediction residual value of the location pixel; This is the original pixel brightness value for the current pixel; The pixel prediction values ​​are obtained from intra-frame prediction operations. Considering the temporal similarity between adjacent video frames in the target video frame sequence, motion estimation and motion compensation operations are used to remove inter-frame redundancy. Using the current coded frame as the frame to be processed and the previously coded reference frame as a benchmark, the optimal matching pixel region is searched block by block. Motion vectors are calculated, and the predicted pixel values ​​for the current macroblock are obtained through motion vector mapping. The formula for calculating the inter-frame prediction residual is as follows: ,in These are the pixel residual values ​​after inter-frame motion compensation; The original pixel values ​​of the current encoded frame; The matching pixel values ​​at the offset positions corresponding to the reference frame are used. The pixel residual matrices predicted intra-frame and inter-frame are subjected to discrete cosine transform to convert spatial domain pixel data into frequency domain coefficient data, concentrating effective image energy and eliminating high-frequency subtle redundant information. The transformed and quantized coefficients, prediction mode parameters, motion vector parameters and other encoding information are losslessly compressed and encoded to convert the structured encoded data into a continuous binary bitstream, finally completing the standardized encoding of a single frame image. The encoding process dynamically matches the encoding bitrate according to the resolution and frame rate parameters of the target video frame sequence. For high-definition, high-frame-rate interactive video, a high-fidelity encoding strategy is adopted to reduce the quantization coefficient, retain high-frequency details and ensure high-definition image quality. For idle video after parameter reduction, an efficient compression encoding strategy is adopted to appropriately increase the quantization compression ratio and maximize the compressed data volume, finally generating an H.264 encoded video stream.

[0063] Global temporal analysis is performed on standardized customer service audio data blocks, strictly adhering to the hybrid coding architecture of the Opus audio coding standard. This integrates two core operational logics: linear predictive coding and transform coding. It adapts to human voice audio characteristics to complete full-domain fine-grained compression coding operations. Temporal feature detection is performed on continuous customer service audio data blocks to identify steady-state and transient characteristics of the human voice. The audio signal is adaptively segmented into blocks; long-block coding is used for steady-state human voices to improve compression efficiency, while short-block coding is used for transient human voices to avoid sound quality distortion, adapting to the dynamic changes in human voice characteristics. A linear prediction algorithm is used to fit the changing patterns of the audio temporal waveform, using historical sampling points to predict the current sampling point value, extracting the effective residual signal of the audio, and eliminating audio temporal redundancy. The linear prediction calculation formula is... ,in For the first Predicted audio amplitude at each sampling point; The linear prediction order; For the first Prediction coefficients; For the preceding number The original amplitude values ​​of each historical sampling point are used to calculate the audio residual signal based on the predicted values. The residual calculation formula is as follows: ,in This represents the audio residual value at the current sampling point. The original audio amplitude at the current sampling point is used; the extracted audio residual signal is subjected to discrete Fourier transform to convert the time-domain residual signal into a frequency-domain signal. Based on the human hearing masking effect, high-precision coefficients are retained for the human ear sensitive frequency band, and the human ear insensitive frequency band is appropriately quantized and compressed. Under the premise of ensuring lossless human voice hearing, the audio data volume is compressed to the maximum extent.

[0064] Entropy encoding is performed on the quantized frequency domain coefficients, prediction parameters, and block parameters to generate a continuous and regular Opus binary audio bitstream, completing full-domain audio compression encoding. After encoding, the continuous audio bitstream is divided into independent audio data packets of equal duration according to the system's preset fixed duration of 20ms. Each audio data packet carries independent timing information, encoding parameters, and verification information, ultimately generating a structured, well-organized, and sound-stable Opus encoded audio stream. Using a globally unified audio timing reference axis as the sole synchronization scale, the timing timestamps of each frame of the H.264 encoded video stream and each audio data packet of the Opus encoded audio stream are extracted. The audio and video timing points are traversed and matched, and the video encoded data and audio encoded data that are synchronized are bundled and encapsulated to correct audio and video timing offset and fast / slow step issues. The audio and video encoded data are integrated through a unified encapsulation format to finally generate the initial encoded audio and video stream.

[0065] The system acquires the estimated network bandwidth and round-trip latency of the current session connection, calculates the available transmission rate of the current session, and when the available transmission rate is lower than a preset minimum transmission threshold, it generates session control commands including bitrate reduction and frame rate reduction commands. Based on these commands, it downgrades and adjusts the encoding parameters of the initial encoded audio and video stream to obtain the final encoded audio and video stream and session control commands. Otherwise, it generates session control commands including normal bitrate maintenance and normal frame rate maintenance commands, and directly uses the initial encoded audio and video stream as the final encoded audio and video stream. Specifically, the system continuously detects the network connection status of the current customer service interaction session, continuously collects network bandwidth estimates and network round-trip latency data within a fixed detection period. The network bandwidth estimate represents the maximum data transmission capacity that the current link can carry, and the network round-trip latency represents the link latency of data transmission. These two parameters together reflect the real-time network transmission quality. Combining the real-time collected network bandwidth estimate and round-trip latency parameters, the system calculates the actual available transmission rate of the current session using the network transmission rate calculation formula. The calculation formula is as follows: ,in The current available transmission rate for the session, in bps; Real-time bandwidth estimate for the current network connection, in bps; is the network delay attenuation coefficient, and is a fixed calibration constant of the system, used to characterize the impact of round-trip delay on transmission rate loss; This represents the current network round-trip time, measured in seconds (s).

[0066] The system presets a minimum stable transmission threshold of 800kbps for sessions. This value is a specific critical threshold for the audio and video interaction scenario of this digital human customer service system, based on sufficient engineering test calibration. Through multiple rounds of comparative tests under different network conditions, 800kbps has been verified as the minimum critical bandwidth standard that accommodates the digital human's 720P-1080P dynamic resolution, 15FPS-25FPS dynamic frame rate, and Opus high-definition audio synchronous transmission. When the transmission rate is higher than this value, the audio and video stream transmission is smooth, without packet loss or audio-visual tearing, ensuring a normal interactive experience. When the transmission rate is lower than this value, the data volume of the original encoded audio and video stream exceeds the link's carrying capacity, easily leading to problems such as bitstream accumulation, video stuttering, audio interruption, and audio-visual asynchrony. This threshold can distinguish between stable network transmission conditions and weak network restricted conditions, adapting to the dynamic fluctuation scenarios of most office networks and mobile networks, and can stably guarantee the minimum image and audio quality requirements of the digital human's audio and video interaction. The calculated real-time available transmission rate is compared with the preset minimum transmission threshold of 800kbps. The system compares and determines the appropriate control logic. When the real-time available transmission rate is lower than the preset minimum transmission threshold, it is determined that the current network link bandwidth is insufficient and the transmission quality is poor, unable to support the original transmission parameters of the initial encoded audio and video stream. The system automatically generates a session control command containing bitrate reduction and frame rate reduction instructions. Based on the command requirements, the overall encoding bitrate, video playback frame rate, and video encoding resolution of the initial encoded audio and video stream are adjusted in a coordinated manner to reduce the overall data transmission volume of the audio and video stream, adapting to low-bandwidth and weak network environments. The audio and video stream after parameter downgrading and optimization is the final encoded audio and video stream, and the corresponding session control command is output synchronously. When the real-time available transmission rate is greater than or equal to the preset minimum transmission threshold, it is determined that the current network link transmission quality is good and the bandwidth is sufficient to stably support the normal transmission of the initial encoded audio and video stream. The system generates a session control command containing normal bitrate maintenance and normal frame rate maintenance instructions, without modifying or downgrading any parameters of the initial encoded audio and video stream, and directly determines the initial encoded audio and video stream as the final encoded audio and video stream.

[0067] It improves the smoothness and experience of user interaction in different network environments, while reducing system computing power consumption and network resource occupation, and adapts to real-time human-computer interaction scenarios under various complex network conditions.

[0068] In a preferred embodiment of the present invention, the alignment module may include:

[0069] In this embodiment of the invention, a bidirectional streaming transmission channel is established between the Python side and the Golang side based on the gRPC protocol. Specifically, this includes: uniformly adopting the gRPC remote procedure call protocol as the only communication protocol between the Python side and the Golang side; defining a dedicated bidirectional streaming service file based on Protocol Buffers serialization syntax; unifying the data serialization format, field type, transmission byte order, and protocol version at both ends to prevent data parsing compatibility issues in heterogeneous language services; declaring a bidirectional streaming service method in the service file, which differs from ordinary one-way client-side streaming and server-side streaming transmission modes. This bidirectional streaming method supports independent, continuous, and asynchronous sending and receiving of streaming data between the client and server, with no limit on the length of data transmitted in a single transmission or the number of interactions. It is suitable for continuous interactive scenarios such as long-streamed audio-visual data, real-time subtitle data, and dynamic control commands. At the same time, the basic communication parameters at both ends are uniformly configured, with fixed transmission timeout thresholds, heartbeat detection cycles, and maximum single-frame transmission byte limits to complete the alignment of communication parameters between the two ends.

[0070] Using the Golang-side control service as the communication server, it completes local port binding, service instance registration, route method mounting, and monitor startup. The Golang side continuously monitors the specified communication port and monitors connection handshake requests from Python-side services within the local area network in real time. It also initializes an independent session resource pool, allocating dedicated thread resources, data cache queues, and log recording units for each new connection. Using the Python-side multimodal inference service as the communication client, it actively initiates a gRPC bidirectional streaming connection handshake request based on the preset Golang-side server IP address and monitoring port. Before initiating the request, the client completes local communication parameter verification, protocol adaptation verification, and permission verification. After confirming parameter matching and legal permissions, it communicates with the server. The system completes three-way handshake, protocol negotiation, and identity authentication processes. During the handshake, both ends synchronize heartbeat mechanisms, data retransmission mechanisms, and abnormal reconnection mechanisms. Fault tolerance logic is preset for abnormal scenarios such as network jitter, momentary disconnection, and data packet loss. After both ends pass the handshake authentication and parameter negotiation, a bidirectional full-duplex streaming transmission channel is officially established between the Python side and the Golang side. This channel runs continuously throughout the entire digital human customer service interaction session lifecycle without the need for repeated connection establishment. It supports the Golang side to send session control commands, network status adjustment commands, and encoding parameter adjustment commands to the Python side in real time. At the same time, it supports the Python side to continuously push encoded audio and video streams, multimodal inference data, and subtitle timing data to the Golang side.

[0071] Encoded audio and video streams are received via a bidirectional streaming channel. Frame timestamps of each audio and video frame in the stream are obtained. Subtitle text is extracted from the multimodal inference results, and subtitle timestamps of each subtitle unit are obtained. Using the frame timestamps as the baseline timeline, the subtitle timestamps are matched with the frame timestamps. When the difference between the subtitle timestamps and frame timestamps is less than a preset alignment tolerance, the current subtitle unit is bound to the current audio and video frame to obtain an alignment data pair. Specifically, this includes continuously and in real-time receiving Python side-transmissions via a pre-established resident bidirectional streaming channel. The final encoded audio and video stream is processed frame-by-frame, segmented, and its frame structure verified for the continuous binary streaming bitstream. Abnormal, incomplete, duplicate, redundant, and invalid noise frames are removed, and complete, valid standardized video and audio frames are selected. During frame parsing, the global frame timestamp, uniformly marked during the encoding and encapsulation stage of each audio and video frame, is read synchronously. All frame timestamps originate from the same global timing clock, possessing temporal uniqueness, continuity, and incrementing properties. The generation, encoding, and output times of each audio and video frame are recorded, with the frame timestamps of all audio and video frames serving as the core. A benchmark is constructed to establish a globally unified, linearly increasing baseline timeline, serving as the sole reference standard for all subtitle timing matching. The complete multimodal inference results output by the system in real-time are subjected to layered and field-based structured analysis, stripping away irrelevant and redundant information such as video feature data, audio feature data, and semantic inference intermediate data to extract independent subtitle semantic units. These units are segmented according to human-computer dialogue semantic logic, with each unit corresponding to a complete and independent interactive semantic segment, without semantic truncation or overlap. Simultaneously, a standard subtitle timestamp is extracted for each unit; this timestamp represents the standard moment for subtitle semantic inference generation and output display, originating from the global baseline timeline clock to ensure temporal dimension uniformity. This ultimately forms an ordered, complete sequence of subtitle units carrying temporal information. Based on the globally unified baseline timeline, homologous matching of subtitle timestamps and audio / video frame timestamps is achieved. Subtitle units and audio / video frames within the same time interval on the timeline are paired one-to-one. The offset of the two types of temporal parameters is quantified using an absolute difference calculation formula, characterizing the degree of temporal synchronization between subtitles and audio / video. The specific calculation formula is as follows: ,in This represents the absolute timing offset difference between the subtitle unit and the audio / video frame, expressed in milliseconds (ms). The timestamp of the standard subtitle for the current subtitle unit to be matched, in milliseconds (ms). This is the global frame timestamp of the current audio / video frame to be matched, in milliseconds (ms).

[0072] The preset timing alignment tolerance is 15ms. This value has been comprehensively calibrated through extensive digital human interaction audiovisual testing, cross-module latency statistics, and network jitter scenario simulation experiments, providing a rigorous engineering basis. The 15ms tolerance threshold is far lower than the critical value of audiovisual synchronization deviation perceptible to the human eye and ear. The human body cannot perceive subtitle advance or lag issues within this range. At the same time, it can effectively accommodate objective errors such as system microsecond-level computing latency, minor streaming jitter, and minor device clock offsets. It will not cause normal matching data loss or matching failure due to excessively small tolerance, nor will it cause problems due to excessively large tolerance. The problem of cross-frame mismatch, semantic misalignment, and timing disorder is addressed by comparing the absolute timing offset difference obtained from quantization with a preset alignment tolerance of 15ms. When the absolute timing offset difference is less than 15ms, it is determined that the current subtitle unit is in sync with the current audio and video frame, the match is valid, and there is no perceptible deviation. Data binding is then performed to establish a unique corresponding relationship between the two, generating standardized alignment data pairs that are structurally complete, timing-matched, and semantically corresponding. Each alignment data pair contains only one set of audio and video frames and one set of subtitle units, without any many-to-many or one-to-many disordered matching relationships.

[0073] The subtitle units and audio / video frames bound in the aligned data pairs are encapsulated into multimodal data segments of a unified format. A unified timestamp on the reference timeline is added to the multimodal data segments, and the corresponding session control commands are combined with the multimodal data segments, arranged in ascending order according to the unified timestamp, and then packaged to obtain a time-consistent multimodal streaming data packet. Specifically, this includes: for each independent aligned data pair, extracting the internally bound audio / video encoded data and subtitle text data; standardizing and encapsulating the structural differences between the two types of heterogeneous data; based on the multimodal streaming transmission scenario of digital human customer service, and combining the data parsing specifications of the gRPC bidirectional streaming transmission protocol, audio / video encoding format characteristics, and subtitle text storage characteristics, after multiple rounds of compatibility adaptation tests, a fixed standardized multimodal data segment structure is pre-defined. This structure is specifically used to uniformly carry the time-matched audio / video binary stream data and subtitle text data, and can adapt to the full-process business needs of front-end parsing and rendering, network fragmentation transmission, and back-end time verification, providing sufficient basis for scenario adaptation. Based on the engineering implementation, a pre-defined unified multimodal data segment structure includes four fixed core fields. The field type, length, data storage format, and arrangement order of all fields are fixed system configurations, with no dynamic changes or random arrangements. The specific structure consists of: First, a video encoding data field, used to store the binary bitstream data of a single frame of H.264 encoded video, using a fixed binary byte storage format to fully carry all the encoding information of a single frame of video; Second, an audio encoding data field, used to store the binary data of a corresponding time-series single group of Opus encoded audio data packets, strictly corresponding to the video frame time sequence, using a standard audio bitstream storage format; Third, a subtitle text field, used to store the complete subtitle unit text content matched in the current time sequence, using the UTF-8 standard character encoding format to ensure that multilingual and multi-character text is free of garbled characters and distortion; Fourth, a basic verification field, used to store the data length verification value and data type identifier of the current data segment, used for subsequent data integrity verification and type identification.

[0074] Based on the defined structure, the extracted audio and video encoded binary data and subtitle text character data are sequentially filled into the corresponding fixed fields of the structure. This completes the unified format conversion and structured inclusion of different types of heterogeneous data. During the encapsulation process, an automatic data integrity verification mechanism is simultaneously activated. By comparing the actual filled data length with the preset length of the standard field, it determines whether the current data segment has problems such as missing fields, incomplete data, abnormal truncation, or disordered format. Invalid data segments that do not meet the requirements are automatically eliminated, and only valid multimodal data segments with complete structure, complete fields, compliant format, and unique parsing rules are retained. This achieves the standardization and structured integration of discrete and disordered multimodal raw data. Using the constructed global reference timeline as the sole time-series tracing basis, the reference timestamp of the audio and video frame matched by the current multimodal data segment is extracted. This timestamp is used as the global unified timestamp and is fixed and attached to the header time sequence identifier field of the multimodal data segment. Each multimodal data segment corresponds to only one global unified timestamp.

[0075] Based on the time-series correspondence of globally unified timestamps, a linkage matching mechanism between session control commands and multimodal data segments is established. Session control commands corresponding to the time-series interval of the current multimodal data segment are selected. These commands are adjustment commands for audio and video encoding parameters and network transmission status at the corresponding time. The selected session control commands are embedded as independent control fields into fixed command positions in the corresponding multimodal data segments to achieve time-series binding of business data and control commands. This ensures that the generation, output, and transmission of each set of multimodal data are accompanied by corresponding adjustment commands, achieving synchronous linkage between data and control. For all multimodal data segments that have completed command combinations, the globally unified timestamp appended to the header is used as the core sorting criterion. The global traversal and rearrangement are strictly performed according to the ascending order rule of timestamp values ​​from smallest to largest. After the time-series sorting is completed, all ordered multimodal data segments are encrypted, packaged, and format-encapsulated as a whole, and the header, trailer, and checksum structures of the data packets are unified to generate the final multimodal streaming data packets.

[0076] Achieve unified timing integration of audio and video data, subtitle text, and conversation control commands to ensure the synchronization of audio and video playback, subtitle display, and network control during digital human customer service interaction.

[0077] In a preferred embodiment of the present invention, the presentation module may include:

[0078] In this embodiment of the invention, multimodal streaming data packets are unpacked in a unified timestamp order to extract encoded audio and video streams. These encoded audio and video streams are then decoded into original audio frame sequences and original video frame sequences, respectively. A WebRTC signaling server facilitates session description protocol negotiation and interactive connection establishment signaling exchange between the Golang side and the web frontend, establishing a real-time point-to-point connection. Specifically, to address the issues of asynchronous transmission of multimodal streaming data and data disorder caused by network jitter, a unified global timestamp ascending order rule is pre-set. This rule is generated based on the system's globally consistent clock system, possessing rigorous engineering timing standards. The global clock maintains a unique benchmark throughout the entire multimodal data generation, encoding and encapsulation, timing alignment, and packet output chain. The timestamps of all audio and video frames, subtitle units, and data packets originate from the same clock beat, eliminating cross-module clock offsets and inconsistent timing standards. Furthermore, considering the output characteristics of real-time interactive streaming data for digital humans, the multimodal... All multimodal data segments within a streaming data packet must be encapsulated and sorted strictly according to a linearly increasing timestamp value. This sorting rule is a system-fixed, enforced standard that does not dynamically change with transmission order, reception order, or network status, ensuring the temporal order of the data packet from the data source encapsulation stage. Based on this fixed, preset sorting rule, continuously received multimodal streaming data packets are decapsulated and parsed packet by packet in an orderly manner, strictly following the temporal sequence to complete data decomposition, avoiding data out-of-order, timing reversal, and partial misalignment issues caused by asynchronous transmission. During the decapsulation process, according to the preset structure field rules of the multimodal data segments, session control instructions, timing verification information, and redundant information in the encapsulation header and tail are stripped from the data packet, and the standardized H.264 encoded video stream and Opus encoded audio stream are completely extracted, retaining the complete original audio and video encoded bitstream data and corresponding globally unified timestamp information, ensuring that the extracted encoded audio and video streams are temporally continuous, data complete, without missing or disordered data.

[0079] For the H.264 standard encoded video stream extracted from the unpacking, a full decoding and restoration operation is performed by calling the standardized H.264 video decoder. Following the inverse operation logic of encoding, inverse operations such as entropy decoding, dequantization, inverse discrete cosine transform, and inter-frame and intra-frame prediction compensation are completed in sequence to restore the pixel spatial information and temporal dynamic information of the video bitstream step by step. Redundant compression parameters in the encoding and compression process are removed to restore the uncompressed original video pixel data. Frame by frame, a continuous and complete sequence of original video frames is generated. Each original video frame retains the resolution and frame rate parameters after adaptive adjustment of encoding, and each video frame is bound to the original globally unified timestamp to ensure that the video frame sequence is temporally aligned, the picture is complete, and the clarity is lossless.

[0080] For the unpacked Opus standard encoded audio stream, a global decoding and restoration operation is performed using an Opus audio decoder adapted for human voice scenarios. Following the inverse operation logic of the Opus encoding hybrid architecture, audio entropy decoding, inverse quantization, inverse Fourier transform, and linear prediction compensation are completed to restore the temporal waveform, frequency characteristics, and timbre prosody information of the human voice audio, eliminating data loss during audio compression. The decoded audio data is then time-series reassembled according to a fixed 20ms packet duration, splicing together a continuous, stable, distortion-free, and uninterrupted original audio frame sequence. Each group of audio frames is synchronously bound to a corresponding globally unified timestamp, achieving temporal consistency between the audio and video frame sequences. Based on the completed audio and video frame sequence decoding and restoration, a WebRTC bidirectional interactive connection process is initiated, using an independently deployed WebRTC signaling server as the intermediate interactive hub to achieve… The cross-platform signaling interaction and protocol negotiation between the Golang side and the web frontend is as follows: First, the Golang side acts as the media data sender, generating session description protocol information that includes audio and video encoding parameters, media transmission port, network adaptation parameters, and synchronization timing parameters. This session description protocol information is then pushed to the WebRTC signaling server, which handles protocol data forwarding and synchronization. After receiving the session description protocol information, the web frontend performs protocol parameter parsing, local media adaptation verification, and network port matching verification. It then generates corresponding response session description protocol information and sends it back to the Golang side. After both ends complete session description protocol parameter matching, media capability negotiation, network parameter adaptation, interactive connection establishment signaling exchange, and network penetration verification, a real-time point-to-point transmission connection based on the WebRTC protocol is formally established between the Golang side and the web frontend.

[0081] The process involves writing the original audio frame sequence into the first media track of the WebRTC real-time peer-to-peer connection, writing the original video frame sequence into the second media track of the WebRTC real-time peer-to-peer connection, and writing the subtitle text obtained from the unpacking of multimodal streaming data packets into the third media track of the WebRTC real-time peer-to-peer connection according to a unified timestamp. Specifically, this includes: writing the decoded, sequentially continuous original audio frame sequence frame by frame into the first media track of the WebRTC real-time peer-to-peer connection, following the ascending order of the globally unified timestamp. The first media track is a system-defined audio transmission track used only to carry human voice audio frame data. Audio frames are cached strictly according to the timestamp order within the track to maintain the temporal continuity and rhythmic stability of audio playback. During the writing process, the ascending integrity of the audio frame timestamps is simultaneously verified, and abnormal audio frames with disordered, repeated, or missing timestamps are removed to ensure that the audio frame sequence within the first media track is orderly, complete, free of noise, and without broken frames; and writing the decoded, sequentially complete original video frame sequence frame by frame into the WebRTC real-time peer-to-peer connection according to a globally unified timestamp reference. The system establishes a point-to-point connection to a pre-defined second media track. This second media track is a system-defined video transmission track used solely to carry digital human video frame data. It independently isolates video and audio data, preventing interference between heterogeneous data. The writing process strictly matches the audio-video timing correspondence, ensuring that video and audio frames corresponding to the same timestamp are written to the track synchronously, maintaining the original audio-video timing pairing and guaranteeing the synchronization of subsequent video and audio. Timing-matched subtitle text units are extracted from the unpacked multimodal streaming data packets. Using a globally unified timestamp as the unique matching benchmark, each subtitle text segment is mapped to its corresponding timing node and written sequentially to a pre-defined third media track in the WebRTC real-time point-to-point connection. This third media track is a system-defined subtitle text transmission track specifically designed to carry real-time subtitle text data, achieving track isolation between subtitle data and audio-video media data. During the writing process, subtitle text is strictly bound according to the timestamp, ensuring that the display timing of the subtitle text corresponds to the audio-video playback timing, preventing subtitles from being ahead, behind, or misaligned, and completing the track-based regularization and mounting of the three types of multimodal data.

[0082] Through real-time peer-to-peer connection using WebRTC, the audio stream from the first media track, the video stream from the second media track, and the subtitle stream from the third media track are synchronously pushed to the web front-end with a unified timestamp as the synchronization benchmark. The web front-end renders the digital human avatar, plays synchronized customer service voice, and displays synchronized subtitle text in the browser environment, resulting in a multimodal real-time digital human interactive screen synchronously pushed to the web front-end. Specifically, it includes: using a globally unified timestamp as the core synchronization benchmark, performing time-series linkage scheduling on the audio stream of the first media track, the video stream of the second media track, and the subtitle stream of the third media track; synchronously outputting the three types of media data at each time-series node according to the timestamp; during the push process, the system verifies the consistency of the timestamps of the three tracks in real time; and performs microsecond-level time-series calibration compensation for track data with microsecond-level time-series deviations to ensure that audio playback, video rendering, and subtitle display at the same time-series node are triggered and take effect synchronously. Relying on the low-latency transmission characteristics of WebRTC peer-to-peer direct connection, the calibrated multimodal streaming data is continuously and in real-time pushed to the web front-end without intermediate server forwarding delay, ensuring the real-time and synchronous nature of data push.

[0083] The web front-end receives multimodal data synchronously pushed from three tracks in real time through the browser's native WebRTC interface. It independently parses and processes each of the three types of data. For the video stream data pushed from the second media track, the browser's native video rendering engine parses the original video frames frame by frame to complete the real-time rendering, layer rendering, and dynamic refresh of the digital human's 3D avatar, outputting a smooth and stable interactive digital human screen. For the audio stream data pushed from the first media track, the browser's native audio playback interface decodes and plays the original audio frames in real time, restoring the standard customer service voice and ensuring consistency in voice timbre, speech rate, and rhythm with the original output. For the subtitle stream data pushed from the third media track, the front-end loads the subtitle text in real time according to a unified timestamp sequence, completing the page mounting, position adaptation, and dynamic refresh of the subtitle text, achieving synchronous display of subtitles with the audio and video sequence. Through the integrated collaborative processing of front-end audio-visual synchronous rendering, real-time voice playback, and dynamic subtitle display, millisecond-level synchronous output of the digital human avatar, synchronized customer service voice, and real-time interactive subtitles is achieved, eliminating issues such as timing offsets, screen stuttering, voice delays, and subtitle misalignment during multimodal display, ultimately resulting in a real-time multimodal digital human interactive screen.

[0084] Improve the real-time performance, fluency, and audiovisual consistency of human-computer interaction in digital human customer service, and adapt to real-time interaction scenarios on various terminal browsers.

[0085] Embodiments of the present invention also provide a computing device, including: a processor and a memory storing a computer program, wherein the computer program, when executed by the processor, performs the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.

[0086] Embodiments of the present invention also provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.

[0087] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A multimodal real-time digital human customer service system based on a microservice-deployed Python-Golang dual-architecture system, characterized in that: include: The generation module is used to obtain user multimodal interaction requests and sequentially perform speech recognition, intelligent agent reasoning, text-to-speech and digital human animation generation processing on the Python side to obtain multimodal reasoning results; The scheduling module is used to synthesize video image frames into a basic video sequence on the Python side based on the video image frames in the multimodal inference results, perform frame-by-frame face recognition and facial semantic segmentation on the basic video, and allocate the preprocessing tasks to the stateless AI worker nodes in the inference process pool according to the polling algorithm to obtain a preprocessed video frame sequence with facial semantic segmentation mask and bounding box coordinates. The control module is used to perform sound activity detection and recognition of the user's speaking state in the Golang side control layer according to the preprocessed video frame sequence, and to perform H.264 compression encoding and Opus compression encoding on the preprocessed video image frames and audio data blocks respectively. At the same time, it performs throttling algorithm control according to the current session connection state to obtain the encoded audio and video streams and session control instructions. The alignment module is used to establish a bidirectional streaming transmission channel between the Python side and the Golang side based on the gRPC protocol according to the encoded audio and video streams, session control instructions and subtitle text, and to perform unified timestamp alignment and packaging of the encoded audio and video streams and subtitle text to obtain time-consistent multimodal streaming data packets; The presentation module is used to establish a real-time peer-to-peer connection through the WebRTC protocol based on multimodal streaming data packets, and synchronously write audio streams, video streams and subtitle streams into the WebRTC media track to obtain a multimodal real-time digital human interactive screen that is synchronously pushed to the web front end.

2. The Python-Golang dual-architecture multimodal real-time digital human customer service system based on microservice deployment as described in claim 1, characterized in that, The generation module includes: The user audio stream is extracted from the user's multimodal interaction request. Acoustic feature extraction and speech endpoint detection are performed on the user audio stream. After dividing the speech activity range, the acoustic features are input into a pre-trained acoustic feature mapping network to obtain the user text sequence corresponding to the user audio stream. Input the user's text sequence into an agent reasoning service instance deployed on the Python side to obtain customer service response text that matches the user's intent; The customer service reply text is input into the speech synthesis processing unit. The text is regularized and phoneme converted. After extracting prosodic features, the speech waveform corresponding to the phoneme sequence is synthesized by the neural vocoder to obtain the customer service audio data block carrying the customer service voice data. The customer service audio data block and the customer service reply text are synchronously input into the digital human image synthesis processing unit. The displacement of the digital human's facial key points is controlled according to the speech amplitude and phoneme duration of the customer service audio data block to obtain video image frames that are synchronized with the lip movements of the customer service voice data. The customer service reply text is then converted into subtitle text by aligning it word by word according to the audio timestamp, and the multimodal inference result is obtained.

3. The Python-Golang dual-architecture multimodal real-time digital human customer service system based on microservice deployment as described in claim 2, characterized in that, The multimodal inference results include the user's original audio stream, customer service audio data blocks, video image frames, and subtitle text.

4. The Python-Golang dual-architecture multimodal real-time digital human customer service system based on microservice deployment as described in claim 3, characterized in that, The scheduling module includes: Video image frames are extracted from the multimodal inference results. The generation timestamp of each video image frame is read. The video image frames are arranged into a frame sequence in ascending order of generation timestamps. The frame sequence is then concatenated and synthesized into a continuous basic video sequence at a preset synthesis frame rate. Face region localization and detection are performed frame by frame for a continuous basic video sequence. In each video image frame, the bounding rectangle of the face is determined to obtain the bounding box coordinates of the face region. Pixel-level facial semantic segmentation is performed within the face region defined by the bounding box coordinates. Each pixel in the face region is classified into the left eye region, right eye region, nose region, mouth region, facial contour region and non-face background region according to semantic category. Semantic category labels are assigned to each pixel to obtain the facial semantic segmentation mask corresponding to the current frame. According to the polling algorithm, the facial semantic segmentation mask and bounding box coordinates of the current frame are packaged into preprocessing task data units. The current task queue length of each stateless AI worker node in the inference process pool is queried. The preprocessing task data units are assigned to the stateless AI worker node with the shortest current task queue length. The stateless AI worker node performs morphological edge smoothing on the facial semantic segmentation mask to eliminate jagged discontinuous pixels on the mask edge. Pixel filling is also performed on the connected component holes inside the mask caused by semantic classification ambiguity to obtain the processed facial semantic segmentation mask. The processed facial semantic segmentation mask is reverse-mapped to the original pixel coordinate system of the corresponding video image frame according to the bounding box coordinates. The processed facial semantic segmentation mask and the bounding box coordinates are used together as metadata channels and appended to the video image frame to obtain a preprocessed video frame sequence with facial semantic segmentation mask and bounding box coordinates.

5. The Python-Golang dual-architecture multimodal real-time digital human customer service system based on microservice deployment as described in claim 4, characterized in that, The control module includes: Extract the user's original audio stream from the multimodal inference results, input the user's original audio stream into the sound activity detection module deployed in the Golang side control layer, perform short-time energy calculation and zero-crossing rate analysis on the user's original audio stream, determine whether the current audio frame is in the user's speaking state, and obtain the sound activity status flag; Based on the sound activity status flag, adaptive parameter adjustments are made to the preprocessed video image frames in the preprocessed video frame sequence to obtain the target video frame sequence. When the sound activity status flag indicates that the user is speaking, the original resolution and frame rate of the preprocessed video image frames are maintained as the target video frame sequence. When the sound activity status flag indicates that the user is not speaking, the encoding resolution and frame rate of the preprocessed video image frames are reduced and then used as the target video frame sequence. The target video frame sequence is subjected to H.264 compression encoding to obtain an H.264 encoded video stream. At the same time, the customer service audio data block is subjected to Opus compression encoding and divided into audio data packets according to a preset duration to obtain an Opus encoded audio stream. The H.264 encoded video stream and the Opus encoded audio stream are timestamped and encapsulated to obtain the initial encoded audio and video stream. Obtain the estimated network bandwidth and round-trip latency of the current session connection, calculate the available transmission rate of the current session, and when the available transmission rate is lower than the preset minimum transmission threshold, obtain session control instructions including bitrate reduction instructions and frame rate reduction instructions, and downgrade the encoding parameters of the initial encoded audio and video stream according to the session control instructions to obtain the final encoded audio and video stream and session control instructions; otherwise, generate session control instructions including normal bitrate maintenance instructions and normal frame rate maintenance instructions, and directly use the initial encoded audio and video stream as the final encoded audio and video stream.

6. The Python-Golang dual-architecture multimodal real-time digital human customer service system based on microservice deployment as described in claim 5, characterized in that, Alignment module, including: A bidirectional streaming channel is established between the Python and Golang sides based on the gRPC protocol; The system receives encoded audio and video streams through a bidirectional streaming transmission channel, obtains the frame timestamps of each audio and video frame in the encoded audio and video stream, extracts subtitle text from the multimodal inference results and obtains the subtitle timestamps of each subtitle unit, uses the frame timestamps as the reference time axis, performs difference matching between the subtitle timestamps and the frame timestamps, and when the difference between the subtitle timestamps and the frame timestamps is less than the preset alignment tolerance, the current subtitle unit is bound to the current audio and video frame to obtain an alignment data pair. The subtitle units and audio / video frames bound in the alignment data pair are encapsulated into a unified format multimodal data segment. A unified timestamp on the reference time axis is added to the multimodal data segment, and the corresponding session control instructions are combined with the multimodal data segment. After being arranged in ascending order according to the unified timestamp, they are packaged to obtain a time-consistent multimodal streaming data packet.

7. The Python-Golang dual-architecture multimodal real-time digital human customer service system based on microservice deployment according to claim 6, characterized in that, The bidirectional streaming channel includes a first unidirectional stream from the Golang side to the Python side, which sends encoded audio and video streams, and a second unidirectional stream from the Python side to the Golang side, which sends session control commands and subtitle text.

8. The Python-Golang dual-architecture multimodal real-time digital human customer service system based on microservice deployment according to claim 7, characterized in that, The presentation module includes: Unpack the multimodal streaming data packets in a unified timestamp order, extract the encoded audio and video streams, decode the encoded audio and video streams into the original audio frame sequence and the original video frame sequence respectively, and complete the session description protocol negotiation and interactive connection establishment signaling exchange between the Golang side and the WEB front end through the WebRTC signaling server to establish a real-time point-to-point connection; The original audio frame sequence is written to the first media track of the WebRTC real-time peer-to-peer connection, the original video frame sequence is written to the second media track of the WebRTC real-time peer-to-peer connection, and the subtitle text obtained by unpacking the multimodal streaming data packet is written to the third media track of the WebRTC real-time peer-to-peer connection according to a unified timestamp. Through real-time peer-to-peer connection via WebRTC, the audio stream from the first media track, the video stream from the second media track, and the subtitle stream from the third media track are synchronously pushed to the web front end with a unified timestamp as the synchronization benchmark. The web front end then renders the digital human image, plays synchronized customer service voice, and displays synchronized subtitle text in the browser environment, resulting in a multimodal real-time digital human interactive screen that is synchronously pushed to the web front end.

9. A computing device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the system as described in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when executed by a processor, implements the system as described in any one of claims 1 to 8.