Digital human live voice interaction system fused with affective computing

Through real-time emotion computing and refined time window management, the digital human live voice interaction system solves the problems of inappropriate response timing and insufficient emotion adaptability, achieving a natural and smooth interactive experience.

CN121393436BActive Publication Date: 2026-04-07BEIJING ZHONGSHENGSHENG DIGITAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-29
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing digital human live voice interaction systems lack refined management in response generation and timing, resulting in inappropriate response timing, affecting user experience, and failing to capture user emotional state in real time, leading to a lack of emotional adaptability and naturalness in the interaction.

Method used

By integrating affective computing, the system acquires user voice input data streams in real time, determines the emotional response time window, and judges whether to start response generation based on real-time emotional state vectors. It predicts the total generation time of the digital human's voice response and dynamically adjusts the tone and rhythm of the voice response by combining affective adaptation and content optimization response models.

Benefits of technology

It improves the naturalness and adaptability of live voice interaction for digital humans, ensures precise matching of response generation and time window, enhances the relevance of interaction and user experience, and reduces mechanical feeling and timing disorder.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121393436B_ABST
    Figure CN121393436B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of digital human voice interaction, and discloses a digital human live broadcast voice interaction system fusing emotional computing. The system acquires user voice input data streams, constructs an emotional response time window, the starting point of which is the end time of the user voice input data streams, and the ending point of which is the preset maximum response cutoff time minus the necessary duration of voice response synthesis. The system acquires the emotional state vector of the user in real time, and judges whether the emotional state vector reaches an emotional intensity threshold value; if the emotional state vector reaches the threshold value, the total generation duration of the digital human voice response is predicted. When the residual duration of the emotional response time window is equal to the total generation duration, the starting point of the window is taken as the starting time of the voice response, and the digital human is controlled to start voice response generation. Through fine time window management and emotional state perception, the system generates voice responses that are adapted to the emotional state of the user, reduces invalid responses and interaction delays, and improves the naturalness of digital human live broadcast voice interaction and user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of digital human voice interaction technology, specifically to a digital human live voice interaction system that integrates emotion computing. Background Technology

[0002] With the deep integration of live streaming technology and artificial intelligence, digital human live streaming has become an emerging form of interaction, widely used in e-commerce shopping guides, online education, entertainment interaction and other scenarios. In the process of digital human live streaming, the naturalness and timeliness of voice interaction directly affect the user experience, while emotional factors, as the core element of interpersonal communication, are often overlooked or insufficiently handled in digital human interaction.

[0003] Traditional digital human voice interaction systems are mostly based on semantic understanding, focusing on recognizing and parsing the user's voice content to generate corresponding responses. These systems are relatively simple in controlling the timing of responses, typically initiating the response generation process immediately after detecting the end of the user's voice input, without considering the dynamic changes in the user's emotional state or the time cost required for response generation. This leads to situations where, in actual interaction, the digital human may abruptly insert a response before the user's emotional expression is fully developed or sufficiently intense, interrupting the user's emotional flow; or the response generation may take too long, exceeding the user's acceptable waiting threshold, causing interaction delays and disrupting the continuity of communication.

[0004] Existing systems lack real-time and accurate capture of users' emotional states. Most systems rely solely on emotional words in the speech text for simple emotion judgment, failing to capture real-time changes in the intensity of the user's emotions during voice input and making it difficult to distinguish between genuine emotional expressions and ordinary semantic statements. This results in digital human-generated voice responses often lacking emotional adaptability. For example, when users exhibit clearly positive or negative emotions, the digital human may still respond in a flat tone, failing to create effective emotional resonance and reducing user immersion and engagement.

[0005] Traditional systems lack a sophisticated time window management mechanism for combining response generation and time scheduling. The initiation and execution of response generation are not correlated with the end time of the user's voice input or the preset maximum response deadline, resulting in a mismatch between the response generation duration and the actual available time window. When response generation takes longer than the user can wait, a "timeout" occurs; conversely, if the response is generated too quickly, it may appear abrupt due to premature output. These issues collectively lead to a mechanical, emotionally resonant, and temporally disordered voice interaction in digital human live streaming, failing to meet users' demands for natural, smooth, and human-like interaction. Summary of the Invention

[0006] The purpose of this invention is to provide a digital human live voice interaction system that integrates emotion computing to solve the problems mentioned in the background art.

[0007] To achieve the above objectives, the present invention provides a digital human live-streaming voice interaction system integrating emotion computing, the system comprising:

[0008] Acquire user voice input data stream;

[0009] Determine the emotional response time window; wherein, the end point of the emotional response time window is the preset maximum response cutoff time minus the necessary duration of speech response synthesis, and the start point of the emotional response time window is the end time of the user's speech input data stream;

[0010] Obtain a real-time emotional state vector and determine whether the real-time emotional state vector reaches an emotional intensity threshold;

[0011] If so, predict the total generation time of the digital human's voice response;

[0012] If the remaining duration of the emotional response time window is equal to the total generation duration, then the starting point of the emotional response time window is taken as the voice response start time, and the digital human voice response generation is controlled by the voice response start time.

[0013] Preferably, obtaining the real-time emotional state vector further includes the following steps:

[0014] Extract the speech emotion feature sequence from the user's speech input data stream;

[0015] Based on the aforementioned speech emotion feature sequence, an initial emotion cluster center is generated using an emotion feature clustering algorithm;

[0016] Calculate the mean feature distance between the speech emotion feature sequence and the initial emotion cluster center;

[0017] The sentiment clustering boundary is adjusted based on the mean of the feature distance, and the real-time sentiment state vector is output; wherein, the sentiment feature clustering algorithm adopts a dynamic weight adjustment mechanism, and the weight coefficients are generated by backfitting historical sentiment interaction data.

[0018] Preferably, the total generation time of the predicted digital human voice response further includes the following steps:

[0019] Obtain the emotional intensity value of the real-time emotional state vector and the matching complexity parameter of the digital human response template library;

[0020] Based on the emotional intensity value and the matching complexity parameter, the baseline generation time for a single response is calculated;

[0021] The response segment interval duration is determined based on the number of semantic segments in the user voice input data stream;

[0022] The total generation time is obtained by multiplying the baseline generation time by the number of semantic segments and adding the cumulative value of the response segment interval time.

[0023] Preferably, after the total generation time of the predicted digital human voice response, the following steps are further included:

[0024] If the remaining duration of the emotional response time window is greater than the total generation duration, then at least one candidate response sub-window that satisfies the total generation duration constraint is selected from the emotional response time window.

[0025] The starting point of the candidate response sub-window is used as the start time of the alternative response, and the generation of the digital human's voice response is controlled by the start time of the alternative response.

[0026] Preferably, after selecting at least one candidate response sub-window that satisfies the total generation duration constraint from the emotion response time window, the method further includes the following steps:

[0027] The candidate response sub-window with the longest interval from the end of the emotional response time window is selected as the first priority sub-window;

[0028] The starting point of the first priority sub-window is taken as the start time of the optimized response, and the digital human voice response generation is controlled by the start time of the optimized response.

[0029] Preferably, the system further includes the following steps:

[0030] Based on the real-time emotional state vector, the target response model in the emotional adaptation response model or the content optimization response model is activated; wherein, the emotional adaptation response model includes a dynamic mapping relationship between emotional intensity and response tone, and the content optimization response model includes a non-linear association rule between semantic content and response rhythm.

[0031] The target response model is used to generate dynamic adjustment instructions for the voice response.

[0032] The voice response dynamic adjustment command is input into the digital human voice synthesis engine to correct the response intonation parameters or response rhythm distribution.

[0033] Preferably, the target response model in the activated emotion-adaptive response model or content-optimized response model further includes the following steps:

[0034] Real-time monitoring of user sentiment fluctuation trends and semantic content complexity;

[0035] When the user's emotional fluctuation trend line exceeds the first emotional threshold and the semantic content complexity is within a preset stable range, the emotional adaptation response model is activated.

[0036] Based on the intonation mapping rules in the emotion-adaptive response model, a dynamic adjustment instruction for the response intonation is generated;

[0037] When the semantic content complexity exceeds the second content threshold and the user emotion fluctuation trend line is in a preset stable range, the content optimization response model is activated.

[0038] Based on the optimized response model, the rhythm distribution rules are used to generate dynamic adjustment instructions for the response rhythm.

[0039] Preferably, after inputting the voice response dynamic adjustment command into the digital human speech synthesis engine, the system further includes the following steps:

[0040] During the speech response generation process, real-time response feedback data is collected, including pitch deviation data and rhythm fluctuation data;

[0041] The intonation deviation data is compared with the predicted value of the emotion adaptation response model to calculate the deviation and generate an intonation error signal.

[0042] The rhythm fluctuation data is compared with the expected value of the content optimization response model to generate a rhythm error signal by offset analysis.

[0043] Based on the systematic deviation component in the intonation error signal, adjust the intonation mapping rules of the emotion adaptation response model;

[0044] Based on the random fluctuation components in the rhythm error signal, the rhythm distribution rule of the content optimization response model is optimized.

[0045] Preferably, after adjusting the intonation mapping rules of the emotion-adaptive response model and optimizing the rhythm distribution rules of the content-optimized response model, the system further includes the following steps:

[0046] Obtain the sentiment intensity benchmark template set and semantic content benchmark template set from the historical sentiment interaction database;

[0047] Perform similarity matching on the emotional intensity benchmark template set and filter candidate emotional templates with a matching degree exceeding a threshold;

[0048] Perform collaborative verification on the semantic content baseline template set and remove conflicting and abnormal content templates;

[0049] Based on the emotional stability index of the candidate emotional templates, the optimal emotional baseline template is selected;

[0050] Based on the optimal emotional baseline template and the verified semantic content template, an updated voice response dynamic adjustment instruction is generated.

[0051] Preferably, the system further includes the following steps:

[0052] After the digital human's voice response is generated, obtain the final response output data;

[0053] The final response output data is compared with a preset response quality standard to generate an overall response error signal;

[0054] Based on the correction component in the overall response error signal, the rules for determining the emotional response time window and the algorithm for generating the real-time emotional state vector are iteratively updated.

[0055] Compared with the prior art, the beneficial effects of the present invention are:

[0056] This system enhances the naturalness and adaptability of live voice interaction for digital humans by integrating affective computing with refined time window management. Regarding response timing control, the affective response time window design enables precise regulation of the response start time. This window starts at the end of the user's voice input data stream and ends at the preset maximum response deadline minus the necessary duration for voice response synthesis. This avoids interrupting the interaction by starting the response prematurely before the user's voice has fully finished, while ensuring that the response generation process is completed within the user's maximum acceptable waiting time. From a temporal perspective, this provides a guarantee for smooth interaction, making the digital human's response timing in live streams closer to the natural pauses in human communication.

[0057] The mechanism for acquiring real-time emotional state vectors and determining emotional intensity thresholds enables the digital human to dynamically perceive the intensity of users' emotional expressions. The system no longer relies on simple text sentiment analysis; instead, it captures emotional features in users' speech in real time to form emotional state vectors and determines whether to initiate a response generation process based on these vectors. This mechanism effectively filters out scenarios where users' emotional expression is insufficient or the emotional intensity is low, avoiding meaningless responses from the digital human, making interactions more targeted, and enabling the digital human to respond when users truly need emotional feedback, thus enhancing the relevance and effectiveness of the interaction.

[0058] In terms of response generation timing, the total generation time of the digital human's voice response is predicted and compared with the remaining time of the emotional response time window, ensuring precise adaptation between the response generation process and the time window. When the remaining time equals the total generation time, the start time of the window is used as the response start time, ensuring that the response can be synthesized and output within the preset time range, while avoiding hasty or delayed responses due to time mismatch. This timing scheduling method makes the digital human's voice output in live broadcasts more organized, makes the timing of the entire interaction process more reasonable, and reduces user discomfort during the waiting process.

[0059] By incorporating the design philosophy of affective computing, the digital human's voice response is no longer limited to semantic matching, but can be adjusted in conjunction with the user's real-time emotional state. When the system detects that the user's emotional state vector has reached a threshold, the generated voice response will naturally incorporate tone and rhythm that match the user's emotions, enabling the digital human to exhibit more human-like emotional expression in live broadcasts. Attached Figure Description

[0060] Figure 1 This is a timing diagram of the digital human live voice interaction system that integrates emotion computing as described in this invention.

[0061] Figure 2 A flowchart for generating real-time sentiment state vectors;

[0062] Figure 3 A graph showing the sequence analysis of multi-channel speech emotion features;

[0063] Figure 4 Flowchart optimized for priority child windows;

[0064] Figure 5 This is a system resource load monitoring and analysis chart. Detailed Implementation

[0065] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0066] Please see Figure 1 This invention provides a digital human live-streaming voice interaction system that integrates emotion computing, the system comprising:

[0067] Based on real-time emotion computing and dynamic response control technology, intelligent interaction is achieved by establishing a closed-loop control mechanism for emotion state vectors and speech response generation. During system operation, the system continuously collects user speech input data streams and processes them in real time using sliding window technology. The emotion response time window is determined using a dynamic boundary calculation method. The window endpoint is determined by subtracting the minimum necessary processing time of the speech synthesis engine from the preset maximum response deadline, while the window start point is strictly aligned with the end time of the user's speech input. When the real-time emotion analysis module detects that the emotion state vector exceeds a preset intensity threshold, it triggers a response duration prediction mechanism. By calculating the non-linear relationship between semantic segmentation and emotion intensity, it accurately predicts the total speech response generation time. The system compares the remaining duration of the emotion response time window with the predicted total duration in real time. When they are equal, it immediately locks the window start point as the response start time, driving the digital human speech synthesis engine to initiate the response generation process.

[0068] Example 1: See Figure 2 In the speech emotion feature processing stage, the system captures the user's speech stream in real time through a multi-channel audio processing unit. The feature extraction engine employs mixed signal processing technology to simultaneously calculate Mel-frequency cepstral coefficients and linear predictive coding parameters. These two parameter sets form a multi-dimensional feature matrix in the temporal dimension, generating a set of 128-dimensional feature vectors every 20 milliseconds. These vectors are concatenated through a sliding window to form a continuous sequence of emotion features. The window length is dynamically adjusted according to the average syllable length of the speech content, while the window step size is fixed at 10 milliseconds.

[0069] In the initialization phase of the sentiment clustering module, the system loads a typical sentiment feature set from the historical interaction database. This dataset has been grouped by sentiment category and contains over 2000 labeled samples. The cluster center generation algorithm first performs principal component analysis on the samples to reduce dimensionality, compressing the dimension to 12 while retaining 90% of the feature quantity. Subsequently, a hierarchical clustering method is used to construct the initial cluster topology, and the optimal number of clusters is automatically determined to be 8 using the silhouette coefficient method. The final generated initial cluster centers contain three-dimensional spatial coordinates and sentiment intensity scalars.

[0070] The dynamic weight adjustment mechanism operates continuously during the clustering process. The weight calculation neural network consists of an input layer, three fully connected hidden layers, and an output layer. The input layer receives statistics from the current feature sequence, including the mean, variance, and autocorrelation coefficient. The hidden layers perform nonlinear transformations, and the neuron activation functions use modified linear units with leakage parameters. The output layer generates weight coefficients for each feature dimension, which are fitted based on the distribution of feature importance in historical interaction samples. The network parameters are retrained every 24 hours based on new interaction data.

[0071] The feature distance calculation employs an improved Mahalanobis distance algorithm. This algorithm introduces a sliding window update mechanism for the covariance matrix, recalculating the covariance structure every 100 feature vectors processed. The mean feature distance is smoothed using a time-series window of length 15 to eliminate transient interference. The sentiment cluster boundary adopts a variable radius sphere model, with the boundary radius directly proportional to the square root of the mean feature distance. When a feature distribution shift is detected, the system automatically expands or contracts the clustering region proportionally, with the adjustment range constrained by the dispersion of the most recent 10 clustering results.

[0072] The construction process of the real-time sentiment state vector includes three sub-steps: distance calculation between feature vectors and cluster centers, weighted fusion, and sentiment intensity quantification. The final output sentiment state vector contains six dimensional parameters, where the first three dimensions represent the projected coordinates of the sentiment type in space, and the last three dimensions correspond to the activation degree, persistence, and fluctuation frequency parameters of the sentiment, respectively. All dimensional parameters are normalized to zero mean and mapped to the [0,1] interval.

[0073] The voice response duration prediction phase begins with emotion intensity threshold detection. When the emotion intensity value exceeds a preset threshold, the system initiates a multi-stage duration estimation algorithm. The matching complexity of the response template library is quantified by multiplying the semantic parse tree depth by the number of keywords, where the semantic tree depth is generated in real-time by a dependency parser. The benchmark generated duration calculation model contains quadratic polynomial terms, with coefficients obtained through regression from historical response data. The model's input parameters include standardized values ​​of the emotion intensity value and complexity exponent, with output in milliseconds.

[0074] Semantic segmentation processing employs a bidirectional recurrent neural network model. This model, based on acoustic feature extraction via convolutional neural networks, identifies natural pauses in the speech stream through a gating mechanism. The number of semantic segments is defined as the statistical value of effective segments containing complete semantic units. The response segmentation interval duration library stores standard interval parameters for 16 typical dialogue scenarios, and the best-matching parameter set is selected using the current scenario classifier.

[0075] The final calculation of the total generation duration incorporates a non-linear correction factor. The basic calculation formula multiplies the baseline duration by the number of segments, then sums the arithmetic sum of the intervals between each segment. The correction phase introduces an exponential compensation term; the compensation factor depends on the gradient of emotional intensity changes and the average response duration of the three most recent interactions. The calculation results are validated using confidence intervals; if they exceed a reasonable range, an alternative model is used for recalculation. All time parameters are synchronized to the system's master clock, with timestamp errors controlled within ±2 milliseconds.

[0076] At the real-time operation level, the emotion state monitoring and duration prediction processes employ a parallel pipeline architecture. The feature extraction and clustering engines are deployed on a dedicated signal processor, while the duration prediction algorithm runs on a general-purpose processor core. The two processors exchange intermediate results via shared memory, with a fixed data exchange cycle of 5 milliseconds. When the remaining duration of the emotion response time window equals the total predicted duration, the system immediately sends a trigger command with a precise timestamp to the speech synthesis engine. The timing control circuit monitors the execution process and automatically activates a compensation protocol when the response delay exceeds a preset threshold.

[0077] The entire implementation process includes 48 real-time parameter monitoring points, with the core algorithm having an iteration cycle of 8 milliseconds. An automatic optimization program is initiated after every 1000 interactions in the historical interaction database, updating the cluster centers and prediction model parameters. The audio processing unit uses a 44.1kHz sampling rate, with a feature extraction frame length of 512 sampling points and a frame shift of 128 sampling points. The emotional feature sequence cache queue has a depth of 60 frames; when the queue is full, an overwrite mechanism is activated to retain later-time audio data.

[0078] The response template library employs a hierarchical index structure, with the top-level index based on sentiment type classification and the second-level index organized according to semantic scenarios. The template retrieval process uses an approximate nearest neighbor search algorithm, returning the top 5 most relevant candidate templates for duration prediction. The precision of all time parameters is recorded to three decimal places, and timing control utilizes a hardware interrupt mechanism to achieve microsecond-level precision response.

[0079] See Figure 3 This diagram illustrates the multidimensional emotional feature sequence extracted from user speech. The five curves of different gray levels represent 128-dimensional feature vectors at five consecutive time points. Changes in feature values ​​reflect emotional information in the speech, such as pitch, rhythm, and intensity. Peak regions of the feature vectors typically correspond to the strong parts of emotional expression, while stable regions represent neutral emotional states. The system analyzes the changing patterns of these feature sequences and combines them with an emotional clustering model to determine the user's emotional state in real time. The accuracy and real-time performance of feature extraction are crucial foundations for the system to achieve emotionally adaptive interaction.

[0080] Example 2: See Figure 4 This involves a dynamic selection and optimization control mechanism for candidate response sub-windows. When the remaining duration of the sentiment response time window exceeds the predicted total generation time, the system initiates a multi-level candidate window filtering process. The time window management system maintains a continuously updated pool of time-series parameters, refreshing the remaining duration data every 50 milliseconds. The sliding scan engine performs a comprehensive traversal within the complete sentiment response time window interval, using a basic step size of 10 milliseconds.

[0081] The identification of candidate response sub-windows employs a dual-constraint verification protocol. The first constraint verifies the candidate duration, requiring the length of the candidate sub-window to be exactly equal to the predicted total generation duration, with an error range limited to ±5 milliseconds. The second constraint analyzes the stability of the sentiment state, determining the level of sentiment persistence by calculating the variance of the sentiment vector within the candidate sub-window. Simultaneously, system resource usage is monitored; when CPU utilization exceeds 85% or memory utilization reaches 90%, the system automatically relaxes the constraint thresholds.

[0082] The candidate window pool is organized and managed using a priority queue data structure. The queue sorting key is defined as the time difference between the candidate's starting point and the sentiment window's ending point; the larger the difference, the higher the priority. The queue capacity is set to a maximum of 10 valid candidates, and dynamic replacement is triggered when a new candidate is detected that is superior to the last item in the queue. The queue management program performs a reordering operation every 20 milliseconds to ensure that the best candidate is always at the head of the queue.

[0083] The process of determining the start time of alternative responses includes a real-time load balancing mechanism. The system deploys a resource monitoring agent to periodically collect computing unit load data, including parameters such as graphics processor utilization and the number of pending tasks for digital signal processors. When the overall load index is lower than a preset threshold, the system selects the candidate solution at the top of the priority queue; when the load pressure increases, it automatically switches to the principle of time proximity to select the feasible candidate solution closest to the current time.

[0084] The first priority sub-window is selected using a backtracking search algorithm. The search starting point is fixed at the end of the emotional response time window, and the entire timeline is scanned in reverse. During the scanning process, a time-stamped index table is built to record all candidate starting points that meet the total duration requirement. The backtracking step speed adopts a variable speed control mechanism: when approaching the starting point of the emotional window, the scanning speed is slowed down to perform fine detection in the early part of the window; in areas where the emotional intensity changes gently, the scanning speed is accelerated to complete the scan at 5 times the standard speed.

[0085] The decision-making system for optimizing the response start time integrates multiple compensation mechanisms. The temporal compensation module continuously measures the actual response latency of the speech synthesis engine; this latency parameter is obtained through a high-frequency clock counter with nanosecond-level accuracy. The dynamic compensation algorithm, based on the moving average of the most recent 50 response latencies, adds a compensation offset to the optimized time value. The emotion persistence detection unit analyzes the gradient change of the emotion state vector on the time axis, automatically adding a trend persistence period parameter when an upward trend in emotion intensity is detected.

[0086] The trigger control for digital human voice generation adopts a hardware-software collaborative architecture. The timing trigger circuit is implemented using a field-programmable gate array (FPGA) with a built-in high-precision quartz crystal oscillator. The trigger signal output port is connected to the hardware start pin of the speech synthesis engine, and the signal transmission path is designed with an optocoupler isolation structure to eliminate electrical interference. The control software maintains a time synchronization state machine, performing a time alignment operation with the system master clock every millisecond.

[0087] The execution status monitoring loop tracks the speech generation process in real time. Within the first 5 milliseconds after response initiation, the execution loop monitoring is initialized, and thereafter, synthesis progress data is collected every 2 milliseconds. Progress monitoring employs a three-stage pipeline structure: the first stage tracks the text processing status, the second stage monitors the speech parameter generation progress, and the third stage tracks the waveform synthesis stage. When the progress lag exceeds a set threshold, the system activates a timeline compression algorithm, adaptively adjusting the speech playback rate to meet the original timing requirements.

[0088] The system is configured with a delay compensation protocol to handle abnormal situations. The standard compensation scheme uses a time prediction model, with input parameters including current system load, historical delay records, and sentiment vector complexity. The prediction model outputs three levels of compensation parameters: baseline compensation value, maximum compensation value, and safety margin. When the actual delay exceeds the maximum compensation value, the system activates an emergency response procedure: interrupting the current speech synthesis process and using pre-generated backup response content.

[0089] The load balancing system implements a distributed task scheduling scheme. The speech generation task is decomposed into three sub-tasks: semantic parsing, emotion mapping, and speech parameter generation. The scheduler dynamically allocates sub-tasks based on the real-time load of each computing unit: the central processing unit performs logic-intensive semantic parsing, the graphics processing unit accelerates emotion mapping calculations, and the digital signal processor handles the speech parameter generation task. The task monitor tracks the execution progress of each sub-task, and when a delay in any stage exceeds expectations, the system automatically initiates task rescheduling.

[0090] The backtracking search algorithm features an adjustable scanning depth control mechanism. The depth control parameters are dynamically adjusted based on the stability of the emotional state: when the emotional fluctuation index is less than 0.1, a deep scanning mode is activated, performing a fine search at half speed in the first half of the window; when the fluctuation index is between 0.1 and 0.3, a standard scanning speed is used; when the emotional fluctuation is severe and exceeds 0.3, the system automatically skips the first half of the window and searches only the last third of the window. The search result verification module performs a duration-based validity check on each candidate point, eliminating abnormal candidate points caused by system instability.

[0091] The optimized decision-making system employs an integrated voting mechanism for final determination. The decision-making voting committee comprises five independent model output schemes: a time-series optimization model output scheme, a sentiment continuity model scheme, a system load model scheme, a resource consumption prediction scheme, and a user experience evaluation scheme. The voting algorithm uses a weighted voting mechanism, assigning different weight values ​​to each scheme based on its historical performance. When significant disagreements arise, a secondary coordination process is triggered, where the system re-evaluates all candidate schemes and generates a compromise optimization scheme. The final decision result is transmitted to the execution control system within 20 milliseconds to complete the synchronization of time-series parameters.

[0092] See Figure 5 This chart illustrates the load changes of the system's main computing resources. Three curves of different grayscale represent the utilization rates of the processor, graphics processor, and memory, respectively. Two dashed lines represent the system's set alarm thresholds. When the processor load exceeds 85% or memory utilization exceeds 90%, the system automatically adjusts its response generation strategy, such as reducing computational precision or enabling a simplified model. The periodic changes in resource load reflect the system's processing demands at different interaction stages, with peaks typically occurring during periods of complex sentiment analysis or high response generation requirements. A stable load curve indicates that the system has good resource management capabilities.

[0093] Example 3: This example focuses on a collaborative working mechanism between two models—emotional adaptation and content optimization—to address the dynamic adjustment issue during the digital human's voice response generation process. The system employs a hierarchical decision-making architecture. The bottom-level signal processing unit continuously monitors changes in emotional features and semantic content within the user interaction stream. The mid-level analysis engine performs real-time feature extraction and pattern recognition, while the upper-level control module issues model selection and parameter adjustment commands.

[0094] The emotion-adaptive response model is built upon a deep temporal neural network architecture. The network input layer receives a preprocessed sequence of emotion state vectors. This sequence is organized in a sliding window format, with a window length of 15 sampling points, corresponding to a 450-millisecond emotion evolution process. The network hidden layer contains bidirectional long short-term memory (LSTM) units and an attention mechanism layer. The LSTM units are responsible for capturing the temporal dependencies of emotion features, while the attention layer automatically identifies key emotion inflection points. The network output layer generates a set of intonation control parameters, including fundamental frequency curve control points, speech rate adjustment coefficients, and stress distribution patterns. The model training process employs an incremental learning strategy, automatically updating the network parameters after every 100 user interactions.

[0095] The content optimization response model employs a hybrid architecture, combining a rule-based semantic analysis engine and a statistical rhythm prediction model. The semantic parser transforms user input into a dependency syntax tree structure, quantifying content difficulty through three dimensions: tree depth, branching factor, and node complexity. The rhythm prediction model utilizes an improved conditional random field algorithm, defining the following state transition probability calculation relationship:

[0096]

[0097] Character definition explanation: Given input semantic features Under the given conditions, output the rhythm state sequence probability distribution Normalization factor (partition function) is used to ensure that the probability value is in the interval [0,1]. Timing position The rhythmic state at that point (a discrete variable, taking values ​​from a predefined set of rhythmic patterns), Pre-drive timing position rhythm state Input semantic feature vector (including dimensions such as syntax tree depth and node complexity), : No. A transition feature function characterizes the adjacent states. and The transfer relationship, : No. Each state feature function describes the state. With input degree of matching, : No. The weight parameters of each transition feature (obtained through training on historical data), : No. The weight parameters of each state feature (learned through maximum likelihood estimation), The natural exponential function maps the result of a linear combination to the space of positive real numbers. : Transfer feature function index (value range from 1 to the total number of manually defined transfer features), : Index of state characteristic function.

[0098] The activation mechanism of the target response model is implemented using a gating decision unit. This unit calculates two key indicators in real time: the sentiment fluctuation index and the content complexity score. The sentiment fluctuation index is obtained by analyzing the second derivative of the sentiment state vector, reflecting the intensity of the user's emotional changes. The content complexity score comprehensively considers three dimensions: sentence length, vocabulary difficulty, and logical structure, and is expressed as a percentage. When the sentiment fluctuation index exceeds a dynamic threshold and the content complexity is within a stable range, the system activates the sentiment adaptation model as the dominant response generator; when the content complexity exceeds a grading threshold and the sentiment fluctuation is moderate, it switches to the content optimization model to control response generation. The gating decision unit re-evaluates the model selection strategy every 50 milliseconds to ensure that the system always operates in the optimal response mode.

[0099] The encoding and transmission of voice response dynamic adjustment commands adopts a layered protocol design. The command message is divided into a control header and a data body. The control header contains a timestamp, model identifier, and parameter version information, while the data body carries the specific set of adjustment parameters. The emotion adaptation command data body uses a three-dimensional curve control point sequence to describe intonation changes; each control point includes a time coordinate, frequency value, and intensity coefficient. The content optimization command data body uses a two-dimensional matrix encoding rhythm pattern, where matrix rows correspond to semantic units, columns represent time segments, and matrix element values ​​define the rhythm intensity level for that time segment. The command transmission channel employs a double-buffered queue structure to ensure real-time command transmission even under high load conditions.

[0100] The digital human speech synthesis engine's parameter interface design supports multi-granularity control. The fundamental frequency adjustment module receives a sequence of curve control points from the emotion adaptation model and generates a smooth fundamental frequency trajectory using a piecewise cubic Hermit interpolation algorithm. The speech rate control unit employs non-linear time warping technology to dynamically adjust the syllable duration ratio based on the rhythm matrix output by the content optimization model. The stress synthesis component combines the output parameters of the two models, superimposing emotion-driven intensity modulation onto the basic stress pattern. The engine internally maintains a parameter priority arbitration mechanism; when different model parameters conflict, a weighted decision is made based on three dimensions: emotion intensity, content importance, and temporal urgency.

[0101] The real-time parameter synchronization system employs a master-slave clock architecture to ensure timing accuracy. The master clock source uses a temperature-compensated crystal oscillator with a frequency stability of ±0.1ppm. Slave clock units are distributed across various processing nodes, achieving sub-millisecond synchronization through a precise time protocol. The parameter update event triggering mechanism includes both hardware interrupts and software polling. Critical control parameters are written directly through dedicated registers, while routine parameter updates utilize memory-mapped I / O channels. The timing monitoring circuit continuously measures the end-to-end delay from command issuance to parameter activation, automatically triggering a compensation process when abnormal delays are detected.

[0102] A smooth transition strategy is implemented for state transitions during model switching. When the system decides to switch from the emotion-adaptive model to the content-optimized model, a 150-millisecond hybrid control phase is set as the transition period. During this period, the output parameters of the two models are merged in a linearly gradual manner, with the emotion model weighting at 100% at the beginning and fully transitioning to content model control at the end. The transition algorithm monitors the parameter change rate and automatically inserts an intermediate transition point when a sharp jump is detected to avoid perceptible abrupt changes in speech quality. The state transition log records the context of each switch for subsequent analysis and optimization of the switching strategy.

[0103] The anomaly handling system employs a multi-level solution to address conflicts during model collaboration. The primary conflict detection module compares the differences in key parameters output by the two models, marking potential conflicts as those where the fundamental frequency suggestion difference exceeds 20Hz or the speech rate difference exceeds 15%. The intermediate solution initiates a parameter negotiation process, automatically selecting the dominant model by analyzing the priority settings of the current dialogue scenario. The advanced conflict arbitration mechanism is activated in cases of severe inconsistency, invoking 22 preset typical interaction scenario templates as decision-making references, and introducing a manual intervention interface for special handling when necessary. All conflict events are recorded in the diagnostic log, serving as the data foundation for model collaborative optimization training.

[0104] The system maintenance module performs periodic model performance evaluations and parameter calibrations. The evaluation metric set includes 37 speech quality dimensions, covering aspects such as intelligibility, naturalness, and accuracy of emotional expression. The calibration process is executed automatically every 8 hours, during which the system switches to safe mode and uses standard test cases to verify the performance metrics of each model component. The calibration results trigger an automatic parameter adjustment mechanism, updating the parameters online for model components whose deviations exceed the allowable range, while retaining the previous stable version as a rollback backup.

[0105] Example 4: Focusing on a real-time monitoring and dynamic adjustment mechanism for emotional fluctuations and semantic content, this example utilizes a dual-channel analysis engine to achieve precise adaptation of voice responses. During system operation, a parallel processing pipeline is established: the left channel is dedicated to emotional state tracking, while the right channel processes semantic content parsing. The data from both channels are then fused and analyzed at the decision-making level.

[0106] The emotional fluctuation trend line monitoring employs a multi-stage filtering architecture. The raw emotional signal first undergoes a moving average filter to eliminate high-frequency noise, with a window width set to 7 sampling points. The secondary processing stage uses a bandpass filter based on physiological response characteristics to retain emotional fluctuation components within the 0.5Hz to 3Hz range. Finally, the trend line generation module incorporates an inertial smoothing algorithm to suppress spurious fluctuations caused by short-term speech characteristics. The system continuously calculates the second derivative of the trend line, marking emotional event points when a sudden change in slope exceeds a threshold.

[0107] The semantic content complexity assessment employs a three-tiered quantitative system. The initial analysis statistically analyzes sentence length and lexical difficulty, using a pre-defined 5000-word tiered vocabulary for matching and scoring. The intermediate processing parses the dependency syntax tree structure, measuring average nesting depth and branch density. The deep analysis stage detects the frequency of logical connectors and the complexity of referential relationships, forming a comprehensive score. The system maintains a dynamic benchmark library, storing typical complexity distribution data for different dialogue scenarios for relative evaluation (see Table 1).

[0108] Table 1: Correspondence between emotional fluctuations and semantic content states.

[0109] Timestamp Emotional fluctuation index Trendline slope semantic complexity Dominant Model 08:15:23.456 0.62 +0.18 43 Emotional compatibility 08:15:24.102 0.71 +0.25 45 Emotional compatibility 08:15:25.337 0.55 -0.12 78 Content optimization 08:15:26.891 0.48 -0.05 82 Content optimization 08:15:27.452 0.67 +0.31 51 Emotional compatibility

[0110] The activation of the emotion-adaptive response model follows a progressive decision-making process. When the emotion fluctuation index exceeds the first threshold for three consecutive sampling periods, and the semantic complexity remains within a stable range (40-60 points), the system initiates the model switching preparation phase. During the preparation period, the parameter influence of the current model is gradually reduced, while the contextual data of the emotion-adaptive model is preloaded. During formal activation, a cross-gradual transition technique is used to complete the transfer of control within a 150-millisecond transition period, avoiding perceptible speech abrupt changes.

[0111] The dynamic adjustment of intonation mapping rules is implemented through closed-loop control. The system collects the generated speech fundamental frequency curve in real time and compares it frame-by-frame with the ideal curve predicted by the emotion model. The deviation detection algorithm identifies three abnormal modes: overall offset, local distortion, and periodic fluctuations. The adjustment strategy is automatically selected according to the type of abnormality: overall offset triggers model parameter recalibration, local distortion initiates segment regeneration, and periodic fluctuations adjust the cutoff frequency of the filter. All correction operations are recorded in the version control log, maintaining a complete debugging traceability chain.

[0112] The rhythm control of the content-optimized response model employs a distributed decision-making mechanism. A semantic unit segmenter divides the input text into several meaningful segments, each assigned an independent rhythm control agent. The agent cluster works collaboratively, with the master node responsible for the overall rhythm framework and the child nodes handling local rhythm details. When the content complexity exceeds a second threshold, the system automatically increases the decision weights of the child nodes, making the response more aligned with the expression needs of complex semantic structures.

[0113] The real-time feedback data processing employs a double-buffered pipeline architecture. The acquisition thread stores the raw audio data in a circular buffer A, while the analysis thread reads data from buffer B for feature extraction. The two buffers switch roles every 100 milliseconds to ensure continuous data processing. The intonation analysis module measures the fundamental frequency, energy, and duration characteristics of the actual output and performs multi-dimensional comparisons with expected parameters. The rhythm monitor tracks syllable boundaries and stress distribution, calculating the deviation between the actual rhythm pattern and the ideal pattern using a dynamic time warping algorithm.

[0114] The error signal classification system implements five levels of fine processing. The first-level classifier distinguishes between system bias and random noise, the second-level classifier locates the source module of the bias, the third-level analysis determines the spatiotemporal distribution characteristics of the bias, the fourth-level assessment quantifies the severity of the bias, and the fifth-level decision generates a specific correction scheme.

[0115] The model parameter update mechanism employs a differential synchronization strategy. When an adjustment to the intonation mapping rule is detected, the system only transmits the subset of parameters that have changed, and the update packet uses binary differential encoding to compress its volume. The optimization of the rhythm distribution rule uses an incremental update approach, adjusting only the local parameters of the affected semantic segments each time. The update verification phase runs in shadow mode, with the old and new models processing the same input in parallel. The actual switch is only performed after comparing the output differences to confirm the update effect.

[0116] The anomaly handling system is designed with a three-tiered defense system. The primary defense detects out-of-bounds errors in routine parameters and automatically restrains them within reasonable ranges. The intermediate defense addresses model output conflicts by initiating an arbitration protocol to select the optimal solution. The advanced defense handles system-level anomalies, such as persistently high errors or model failures, switching to a safe mode and using a simplified rules engine to maintain basic interactive functions. All protection events are logged in detail, including anomaly snapshots, handling measures, and follow-up impact tracking.

[0117] The system maintenance module performs periodic health checks. A self-check program is initiated at a fixed time each day, sequentially verifying the operational status of the sentiment analysis channel, semantic processing channel, and model decision-making component. The check includes 128 preset test cases, covering the entire chain from signal input to speech output. During maintenance, the system enters a degraded operation mode, suspending non-core functions to ensure service continuity. The check results generate a detailed diagnostic report to guide subsequent optimization and adjustments.

[0118] A version rollback mechanism ensures system stability. A restore point is automatically created before each parameter update, saving the complete runtime context. Rollback decisions are based on multi-dimensional evaluation: when the performance of new parameters falls below historical benchmarks for three consecutive sampling periods, an automatic rollback process is triggered. The rollback operation employs a transaction processing mechanism to ensure the system state is completely and consistently restored to the specified point in time. The version management system supports selecting rollback targets by time point, by session ID, or by exception type.

[0119] Example 5: Focusing on in-depth mining and intelligent template matching of historical emotional interaction data, the system continuously optimizes voice response quality by establishing a multi-dimensional benchmark template system. The system adopts a layered knowledge base architecture: the bottom layer stores raw interaction data, the middle layer organizes feature indexes, and the upper layer constructs a template matching engine, forming a complete data processing closed loop.

[0120] The historical sentiment interaction database is organized using a temporal graph structure, with each node recording a sentiment state vector and its context at a specific moment. Directed edges between nodes represent sentiment state transition paths, with edge weights reflecting transition probabilities and typical transition durations. The database maintenance engine continuously performs data cleaning and feature enhancement, including outlier correction, missing value imputation, and feature standardization. The sentiment intensity benchmark template set is automatically generated using a clustering algorithm, and the template update cycle is set to trigger reconstruction every 200 new interactions.

[0121] The construction of the semantic content benchmark template set adopts a hierarchical sampling strategy. The system extracts representative semantic units from historical dialogues, categorized by topic type, sentence structure, and expressive meaning. Figure 3 Templates are categorized and stored according to several dimensions. Each template contains three parts: original text, structural feature vector, and applicable scenario tags. The template quality assessment module periodically checks the usage effectiveness of each template and calculates a health score based on actual interaction feedback data.

[0122] The similarity matching algorithm employs a hybrid distance metric strategy. For sentiment intensity template matching, the system calculates the cosine similarity and Euclidean distance between the current sentiment state vector and each candidate template, and obtains the comprehensive matching score through weighted summation. Semantic content matching combines word vector space distance and syntactic tree edit distance, taking into account semantic relevance while maintaining structural similarity. The matching score threshold setting adopts an adaptive mechanism, dynamically adjusting the critical value according to the complexity of the current dialogue scenario.

[0123] The collaborative verification process employs a multi-stage filtering mechanism. Primary verification checks semantic and logical consistency, eliminating candidate templates with obvious contradictions or duplications. Intermediate verification analyzes sentiment-content fit, ensuring that the selected sentiment expression aligns with the semantic connotation. Advanced verification assesses temporal coherence, checking the naturalness of transitions when switching templates. Conflicting templates discovered during verification are moved to an isolation area, awaiting manual review before a decision is made on whether to reuse them.

[0124] The emotional stability index is calculated using a composite algorithm. The duration factor measures the number of consecutive sampling points maintaining the emotional state, using exponential decay weighting to highlight the importance of recent data. The intensity volatility calculates the coefficient of variation of the emotional value within a time window, eliminating the influence of absolute dimensions. The transition smoothness assesses the gradient change during state transitions, using cubic spline interpolation to fit the emotional curve and then calculating the root mean square value of the second derivative. The three sub-indices are combined into a final stability score after dimensionality reduction through principal component analysis.

[0125] The selection process for the optimal sentiment benchmark template involves multiple rounds of screening. In the initial screening stage, the top 20% of candidate templates are selected based on similarity scores and stability indicators. The refined screening stage incorporates diversity control to ensure representative templates are selected for each sentiment dimension. The final selection stage considers the real-time dialogue context, prioritizing templates highly relevant to the current topic. The system retains a second-best template as a backup option, automatically switching when the optimal template fails to produce satisfactory results.

[0126] The optimization of semantic content templates employs knowledge graph-assisted technology. The system links candidate templates to the domain knowledge graph, identifying key concepts and relationships involved. Semantic completeness is checked through graph reasoning, automatically filling in missing logical elements. The expression optimization module detects redundant expressions and ambiguous structures in the templates, simplifying and clarifying them using preset rewriting rules. The final optimized template must pass readability evaluation and naturalness testing before it can be used.

[0127] The update instruction generation module employs a difference-driven strategy. The system compares the parameter distribution differences between the old and new template sets to identify key dimensions requiring adjustment. For sentiment templates, it generates update instruction packages containing adjustments to the fundamental frequency curve, speech rate variation rate, and stress distribution map. Content template update instructions focus on rhythm pattern adjustments and pause distribution optimization. The instruction packages use versioned encoding and include complete forward compatibility specifications and rollback guidelines.

[0128] The verification of dynamically adjusted instructions employs a shadow testing mechanism. The system creates a parallel testing environment, applying new instructions to the replay processing of recent historical interaction data. The effectiveness of the update is quantitatively evaluated by comparing the differences between the original output and the test output. Verification metrics include 36 voice quality dimensions, each with an acceptable range of variation. An updated instruction is only approved for formal deployment when more than half of the dimensions show positive improvement and no key dimensions degrade.

[0129] Version control systems maintain a complete change history graph. Each template update generates a detailed change log, recording the modifications, decision-making basis, and scope of impact. The system supports precise rollback based on time points, allowing restoration to any historical version state. The change impact tracking function automatically marks in-transit interaction sessions affected by updates, highlighting potential adaptation issues requiring special attention. Version comparison tools visualize the parameter differences between different version template sets, assisting analysts in understanding the system's evolution path.

[0130] The exception handling subsystem incorporates a defensive programming mechanism. Strict data validation is implemented during template loading to detect and filter corrupted templates with format errors or out-of-bounds parameters. The matching engine has a built-in circuit breaker that automatically switches to safe mode when consecutive low-quality matching results occur. A conflict resolution protocol defines clear processing priorities, prioritizing tasks according to preset rules when multiple optimization objectives cannot be simultaneously achieved. All abnormal events trigger diagnostic data collection, generating a complete fault analysis report for subsequent improvement reference.

[0131] The system maintenance interface provides a comprehensive monitoring dashboard. Operators can view the health status of the template library, matching success rate, and usage trend in real time. Interactive analysis tools support template effectiveness analysis by time range, dialogue scenario, or user group. The maintenance task scheduler automatically plans to execute background optimization jobs during fragmented time periods, minimizing the impact on online services. The access control system implements fine-grained access control to ensure the security of sensitive template data operations.

[0132] Example 6: Focusing on quality assessment and iterative optimization after voice response generation, a closed-loop feedback mechanism is established to realize a self-evolving intelligent interactive system. After the system completes the digital human's voice response output, it initiates a multi-dimensional quality assessment process. Through in-depth analysis of the final output data, a correction signal is generated to drive the dynamic update of system parameters.

[0133] The final response output data acquisition utilizes a multi-source heterogeneous sensor array. The core audio acquisition unit employs a professional microphone array with a 48kHz sampling rate, deployed at key locations within the user interaction space. Auxiliary acquisition devices include a high-definition camera array to capture user micro-expression responses, photoelectric sensors to record user posture changes, and an environmental noise monitor to record the background acoustic environment. All data streams are strictly aligned in the time domain, with timestamp synchronization accuracy controlled within ±3 milliseconds.

[0134] The response quality assessment model constructs a multi-dimensional quantitative matrix. The quality dimensions are divided into three categories: acoustic features, including 12 indicators such as fundamental frequency stability, harmonic noise ratio, and spectral tilt; emotional expression, encompassing 9 parameters such as emotional matching degree, naturalness of tone, and smoothness of emotional transition; and interaction effectiveness, including 7 attributes such as response latency, dialogue coherence, and information transmission efficiency. Each dimension's indicators are quantified using specialized algorithms. For example, fundamental frequency stability is represented by the root mean square error of the continuous fundamental frequency trajectory, and emotional matching degree is obtained by comparing the emotional target vector with the actual output vector using a deep neural network.

[0135] The overall response error signal is calculated using a composite weighting formula:

[0136]

[0137] in: To standardize the overall response error value (dimensionless); For quality dimension index (value range 1 to 28); The total number of quality dimensions defined for the system (constant 28). The relative importance weight of the j-th dimension (determined through principal component analysis); The measured error value (original measurement) of the j-th dimension; The acceptable error threshold for the j-th dimension (system preset constant). The aggregated value of the key dimension deviation (calculated from the first 5 principal component dimensions); The system fault tolerance threshold (engineering empirical constant 0.15); This is the system-level deviation correction coefficient (obtained through regression analysis of historical performance data).

[0138] The error signal decomposition system employs a five-stage cascaded filtering architecture. The primary separation module distinguishes between equipment-related errors and environmental interference errors; the secondary processing separates inherent algorithm errors and parameter configuration errors; the tertiary decomposition identifies time-related errors and content-related errors; the quaternary processing focuses on the error characteristics of the sentiment computing subsystem; and the quintile final separation extracts the steady-state error components that can be used for system correction. Each decomposition stage utilizes digital filtering algorithms and statistical classification techniques.

[0139] The update of the emotion response time window rules adopts a reinforcement learning framework. The system defines the state space as the configuration parameters of the current time window, and the action space includes 12 operations such as adjusting the window boundaries and modifying the buffer duration. The reward function design comprehensively considers response delay penalties, emotion matching rewards, and fluency bonuses. The Q-learning algorithm maintains the state-action value matrix and continuously updates the state transition probability table based on actual interaction effects. Update decisions must undergo security verification before execution to prevent system anomalies caused by sudden rule changes.

[0140] The real-time sentiment generation algorithm employs an iterative dual-channel learning mechanism. The online learning channel uses stochastic gradient descent to continuously fine-tune the feature extraction network parameters, calculating the gradient of the loss function and updating it with small steps after each interaction. The offline optimization channel initiates deep training daily at midnight, loading 3000 new interaction samples and retraining the core parameter layers of the convolutional neural network and recurrent neural network. The dual-channel parameters are merged through a weighted fusion module, preserving both the system's long-term memory and short-term adaptive capabilities.

[0141] The timing control system is calibrated to establish a sidereal time reference. The master clock system is equipped with a GPS-disciplined rubidium atomic clock, generating a 10MHz standard frequency signal. Distributed nodes achieve nanosecond-level synchronization via the IEEE 1588 precision time protocol. The firmware of the time window controller includes an automatic drift compensation algorithm to periodically measure and correct accumulated clock errors. Redundant check codes are added to the generation process of key timing parameters, such as the voice response start time, to prevent bit errors during data transmission.

[0142] The anomaly response case analysis module implements in-depth root cause analysis. The system establishes a dedicated analysis sandbox for anomaly cases with error values ​​exceeding twice the standard deviation threshold, replaying intermediate data from the complete interaction process. The automatic diagnostic engine analyzes the probability of 32 preset fault modes and generates a fault tree structure analysis report. Confirmed cases are transformed into penalty samples in the reinforcement learning environment, driving the system to avoid repeating similar errors.

[0143] The system iteration management adopts a rolling version update strategy. Daily incremental update packages are generated, including optimized time window rules, sentiment generation algorithm parameters, and quality assessment model weights. Shadow testing is performed before each update to verify the performance of the new version using historical data. Formal deployment adopts a canary release model, enabling the new version only on 5% of interactive sessions on the first day, and gradually expanding the deployment scope based on actual results. The rollback mechanism retains a complete system image from the most recent 7 days, allowing restoration to a previous stable state within 45 seconds in case of anomalies.

[0144] The continuous monitoring dashboard integrates a multi-dimensional visualization interface. The main monitoring area displays a real-time trend chart of the boundary changes of the emotion response time window, supplemented by an emotion matching degree heatmap and a response delay distribution histogram. Expert mode provides probes into the internal state of the algorithm, allowing observation of in-depth information such as neural network activation patterns and feature space projection. The early warning system sets dynamic thresholds, triggering a three-level alarm mechanism when core indicators continuously deviate from the normal range.

[0145] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0146] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A digital human live-streaming voice interaction system integrating affective computing, characterized in that, Includes the following steps: Acquire user voice input data stream; Determine the emotional response time window; wherein, the end point of the emotional response time window is the preset maximum response cutoff time minus the necessary duration of speech response synthesis, and the start point of the emotional response time window is the end time of the user's speech input data stream; Obtain a real-time emotional state vector and determine whether the real-time emotional state vector reaches an emotional intensity threshold; If so, predict the total generation time of the digital human's voice response; If the remaining duration of the emotional response time window is equal to the total generation duration, then the starting point of the emotional response time window is taken as the voice response start time, and the digital human voice response generation is controlled by the voice response start time.

2. The digital human live-streaming voice interaction system integrating emotion computing according to claim 1, characterized in that, The process of obtaining the real-time emotional state vector also includes the following steps: Extract the speech emotion feature sequence from the user's speech input data stream; Based on the aforementioned speech emotion feature sequence, an initial emotion cluster center is generated using an emotion feature clustering algorithm; Calculate the mean feature distance between the speech emotion feature sequence and the initial emotion cluster center; The sentiment clustering boundary is adjusted based on the mean of the feature distance, and the real-time sentiment state vector is output; wherein, the sentiment feature clustering algorithm adopts a dynamic weight adjustment mechanism, and the weight coefficients are generated by backfitting historical sentiment interaction data.

3. The digital human live-streaming voice interaction system integrating emotion computing according to claim 2, characterized in that, The total generation time for predicting the digital human's voice response also includes the following steps: Obtain the emotional intensity value of the real-time emotional state vector and the matching complexity parameter of the digital human response template library; Based on the emotional intensity value and the matching complexity parameter, the baseline generation time for a single response is calculated; The response segment interval duration is determined based on the number of semantic segments in the user voice input data stream; The total generation time is obtained by multiplying the baseline generation time by the number of semantic segments and adding the cumulative value of the response segment interval time.

4. The digital human live-streaming voice interaction system integrating emotion computing according to claim 1, characterized in that, After the total generation time of the predicted digital human voice response, the following steps are also included: If the remaining duration of the emotional response time window is greater than the total generation duration, then at least one candidate response sub-window that satisfies the total generation duration constraint is selected from the emotional response time window. The starting point of the candidate response sub-window is used as the start time of the alternative response, and the generation of the digital human's voice response is controlled by the start time of the alternative response.

5. The digital human live-streaming voice interaction system integrating emotion computing according to claim 4, characterized in that, After selecting at least one candidate response sub-window that satisfies the total generation duration constraint from the emotion response time window, the method further includes the following steps: The candidate response sub-window with the longest interval from the end of the emotional response time window is selected as the first priority sub-window; The starting point of the first priority sub-window is taken as the start time of the optimized response, and the digital human voice response generation is controlled by the start time of the optimized response.

6. The digital human live-streaming voice interaction system integrating emotion computing according to claim 1, characterized in that, It also includes the following steps: Based on the real-time emotional state vector, the target response model in the emotional adaptation response model or the content optimization response model is activated; wherein, the emotional adaptation response model includes a dynamic mapping relationship between emotional intensity and response tone, and the content optimization response model includes a non-linear association rule between semantic content and response rhythm. The target response model is used to generate dynamic adjustment instructions for the voice response. The voice response dynamic adjustment command is input into the digital human voice synthesis engine to correct the response intonation parameters or response rhythm distribution.

7. The digital human live-streaming voice interaction system integrating emotion computing according to claim 6, characterized in that, The target response model in the activated emotion-adaptive response model or content-optimized response model further includes the following steps: Real-time monitoring of user sentiment fluctuation trends and semantic content complexity; When the user's emotional fluctuation trend line exceeds the first emotional threshold and the semantic content complexity is within a preset stable range, the emotional adaptation response model is activated. Based on the intonation mapping rules in the emotion-adaptive response model, a dynamic adjustment instruction for the response intonation is generated; When the semantic content complexity exceeds the second content threshold and the user emotion fluctuation trend line is in a preset stable range, the content optimization response model is activated. Based on the optimized response model, the rhythm distribution rules are used to generate dynamic adjustment instructions for the response rhythm.

8. The digital human live-streaming voice interaction system integrating emotion computing according to claim 7, characterized in that, After inputting the voice response dynamic adjustment command into the digital human voice synthesis engine, the following steps are also included: During the speech response generation process, real-time response feedback data is collected, including pitch deviation data and rhythm fluctuation data; The intonation deviation data is compared with the predicted value of the emotion adaptation response model to calculate the deviation and generate an intonation error signal. The rhythm fluctuation data is compared with the expected value of the content optimization response model to generate a rhythm error signal by offset analysis. Based on the systematic deviation component in the intonation error signal, adjust the intonation mapping rules of the emotion adaptation response model; Based on the random fluctuation components in the rhythm error signal, the rhythm distribution rule of the content optimization response model is optimized.

9. The digital human live-streaming voice interaction system integrating emotion computing according to claim 8, characterized in that, After adjusting the intonation mapping rules of the emotion-adaptive response model and optimizing the rhythm distribution rules of the content-optimized response model, the following steps are also included: Obtain the sentiment intensity benchmark template set and semantic content benchmark template set from the historical sentiment interaction database; Perform similarity matching on the emotional intensity benchmark template set and filter candidate emotional templates with a matching degree exceeding a threshold; Perform collaborative verification on the semantic content baseline template set and remove conflicting and abnormal content templates; Based on the emotional stability index of the candidate emotional templates, the optimal emotional baseline template is selected; Based on the optimal emotional baseline template and the verified semantic content template, an updated voice response dynamic adjustment instruction is generated.

10. The digital human live-streaming voice interaction system integrating emotion computing according to claim 1, characterized in that, It also includes the following steps: After the digital human's voice response is generated, obtain the final response output data; The final response output data is compared with a preset response quality standard to generate an overall response error signal; Based on the correction component in the overall response error signal, the rules for determining the emotional response time window and the algorithm for generating the real-time emotional state vector are iteratively updated.

Citation Information

Patent Citations

  • Customer service method and system based on AI digital human

    CN120406893A

  • Multi-modal interaction method and system of digital human intelligent agent

    CN120653118A