Intelligent real-time interactive question-answering system based on virtual digital human
The integration of multi-modal sensors and a virtual digital human behavior decision model in the system addresses the lack of real-time responsiveness and emotional adaptability in existing systems, enhancing interaction experience through synchronized feedback.
Patent Information
- Application Number
- CN202510779127.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-06-12
AI Technical Summary
The existing virtual digital human system fails to fully consider non-verbal information such as user's body movements and facial expressions, resulting in a relatively single and mechanical interactive experience, lacking deep fusion and real-time response of multimodal information.
By integrating multimodal sensors with virtual digital human behavior decision model, real-time processing and cross-modal fusion of speech, text, facial expressions and body movements are realized, and accurate interaction feature matrix is generated. Combined with knowledge retrieval modules and speech generation modules, the response strategy of virtual digital humans is dynamically adjusted to ensure the naturalness of feedback and emotional adaptation.
It significantly improves the interactive experience and adaptability of virtual digital people, realizes an interactive question-and-answer system that is more intelligent, real-time and emotionally adaptable, improves the accuracy of emotional recognition and response, and is suitable for a variety of application scenarios such as intelligent customer service and virtual assistants.
Smart Images

Figure CN120318388A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of voice signal processing and speech recognition, and particularly to an intelligent real-time interactive Q&A system based on a virtual digital human. Background Art
[0002] With the rapid development of artificial intelligence technology, virtual digital humans, as a new type of human-computer interaction interface, have been widely used in multiple fields such as customer service, education, and entertainment. Traditional virtual digital human systems mostly rely on speech recognition and text processing, but existing systems usually fail to fully consider non-verbal information such as users' body movements and facial expressions, nor can they dynamically adjust the behavior and responses of virtual digital humans according to users' emotional states and interaction histories, resulting in a relatively single and mechanical interaction experience.
[0003] Existing virtual digital human technologies often focus on single voice or text input and rely on preset rules to react, lacking in-depth integration of multi-modal information and intelligent adjustment of real-time responses. This method not only makes the interaction effect and emotional adaptability of virtual digital humans poor, but also cannot make accurate behavioral decisions according to specific situations, affecting the user's interaction experience.
[0004] Therefore, the present invention provides an intelligent real-time interactive Q&A system based on a virtual digital human. Summary of the Invention
[0005] The present invention provides an intelligent real-time interactive Q&A system based on a virtual digital human, which is used to solve the defects existing in the prior art, especially to improve the interaction intelligence, emotional adaptability, and response real-time of the system. By integrating a multi-modal sensor and a virtual digital human behavior decision model, a more intelligent, real-time, and emotion-adaptive interactive Q&A system is realized, which can simultaneously process voice, text, facial expressions, and body movements, dynamically adjust the response strategy of the virtual digital human, generate an accurate interactive feature matrix through spatio-temporal alignment cross-modal fusion, and improve the accuracy of emotion recognition and response. The knowledge retrieval module optimizes the answer content according to the user's emotional tendency, and the voice generation module realizes the synchronization of lip movements, micro-expressions, and body movements, ensuring that the feedback of the virtual digital human is more natural and meets the user's needs. The overall solution significantly improves the interaction experience and adaptability of the virtual digital human.
[0006] The present invention provides an intelligent real-time interactive Q&A system based on a virtual digital human, including: Data acquisition module: Based on an integrated multi-modal sensor array, it receives the voice data and text data input by the user. At the same time, it real-time collects the sequence data of the user's expression dynamic parameters and body movement sequences, and processes the voice data and text data to obtain standardized voice feature vectors and structured text data; Cross-modal fusion module: Align the standardized speech feature vectors, structured text data with the user's facial expression dynamic parameter sequence and limb movement sequence data in space-time, and construct an interaction feature matrix containing multi-modal temporal and spatial correlation features; Behavior decision-making module: Determine a decision instruction set containing emotional tendency parameters, knowledge graph association indexes, and feedback time series based on historical interaction data and the interaction feature matrix; Knowledge retrieval module: According to the knowledge graph association index in the decision instruction set, retrieve the corresponding knowledge fragment set from a preset distributed heterogeneous database, perform tone word embedding and prosody rhythm annotation on the vocabulary corresponding to the knowledge fragments in the knowledge fragment set according to the emotional tendency parameters, and generate an answer text and speech features with emotional adaptability; Speech generation module: Generate the key frames of the lip animation, micro-expression parameter sequence, and limb movement trajectory of the virtual digital human and generate a voice response based on the feedback time series, answer text, and speech features in the decision instruction set, and push it to the user terminal through the edge computing node.
[0007] Preferably, the data acquisition module includes: Speech processing unit: Perform frequency-domain energy detection on the received speech data, identify and remove invalid speech segments with energy lower than the preset threshold, and obtain continuous effective speech waveform data; Text analysis unit: Perform dependency syntax analysis on the text data, annotate entity types and action predicates, and generate initial text data with semantic labels; Environmental analysis unit: Collect the three-dimensional space information of the interaction environment, track the movement trajectory of the target user in the interaction space, and calculate the spatial position coordinates and orientation angle data of the user; Image acquisition unit: Dynamically adjust the focal length of the visual sensor and the directivity parameters of the sound pickup array according to the spatial position coordinates and orientation angle data, and obtain the facial expression image and limb movement image data of the user; Image analysis unit: Analyze the spatial offset of the key feature points in the facial expression image, measure the deformation amplitude and duration of a specific facial area, and generate an expression dynamic parameter sequence; Motion analysis unit: Extract the spatial coordinates of the key bone nodes in the limb movement image data, determine the relative movement trajectory between the nodes, and generate limb movement sequence data; Data processing unit: Process the continuous effective speech waveform data and the initial text data to obtain standardized speech feature vectors and structured text data.
[0008] Preferably, the data processing unit includes: Speech data processing sub-unit: Preprocess the continuous valid speech waveform data using an adaptive noise cancellation algorithm, extract acoustic features, and generate a standardized speech feature vector containing dynamic time warping information; Text data processing sub-unit: Segment the initial text data through a domain-adaptive semantic segmentation model, and perform intention annotation in combination with a preset intention recognition knowledge base to generate structured text data with context semantic associations.
[0009] Preferably, the cross-modal fusion module includes: Data alignment unit: Perform spatio-temporal alignment on the standardized speech feature vector, structured text data, user's facial expression dynamic parameter sequence, and limb movement sequence data; Weight adjustment unit: Adjust the weights of the standardized speech feature vector, structured text data, facial expression dynamic parameter sequence, and limb movement sequence data after spatio-temporal alignment; Matrix construction unit: Integrate the standardized speech feature vector, structured text data, user's facial expression dynamic parameter sequence, limb movement sequence data, and the adjusted corresponding weights to construct an interaction feature matrix containing multi-modal temporal and sequential correlation features.
[0010] Preferably, the weight adjustment unit includes: Coupling sub-unit: Perform coupling analysis on the aligned standardized speech feature vector, facial expression dynamic parameter sequence, and limb movement sequence data to obtain a speech-expression coupling coefficient and a speech-action coupling coefficient; First adjustment sub-unit: Perform a first adjustment on the weight distribution ratios of the aligned standardized speech feature vector, facial expression dynamic parameter sequence, and limb movement sequence data according to the speech-expression coupling coefficient and the speech-action coupling coefficient; Parameter acquisition sub-unit: Perform micro-expression analysis on the user's facial expression dynamic parameter sequence and limb movement sequence data. At the same time, perform speech fundamental frequency jitter detection on the aligned standardized speech feature vector to determine the user's physiological state parameters, where the user's physiological state parameters include an emotional tension index and an attention concentration score; Second adjustment sub-unit: Perform a second adjustment on the weights of the aligned standardized speech feature vector, facial expression dynamic parameter sequence, and limb movement sequence data based on the user's physiological state parameters; Third adjustment sub-unit: Perform a third adjustment on the preset initial weight of the structured text data based on the weights of the standardized speech feature vector, facial expression dynamic parameter sequence, and limb movement sequence data after the second adjustment.
[0011] Preferably, the behavior decision-making module includes: a historical data weight adjustment unit: determining the multimodal weight distribution of the user's interaction behavior based on historical interaction data and the interaction feature matrix, obtaining the short-term interaction frequency, multimodal signal volatility, and emotional tendency deviation degree, performing multimodal confidence analysis and interaction habit matching, and obtaining the voice weight correction coefficient, text weight correction coefficient, and action weight correction coefficient; A cross-modal association network construction unit: monitoring the signal synchronization between the standardized voice feature vector, structured text data, and facial expression dynamic parameter sequence during the interaction process, determining the voice-text synchronization rate and voice-expression synchronization rate, and generating a multimodal synchronization relationship network; A feedback decision splitting and prediction unit: performing semantic parsing on the structured text data, extracting interaction intention labels, and determining the feedback level division ratio and feedback rhythm adjustment coefficient in combination with the emotional tendency deviation degree; An associated multimodal anomaly analysis unit: based on the multimodal synchronization relationship network, using the feedback level division ratio and feedback rhythm adjustment coefficient to perform cross-modal consistency verification, determining the text-voice anomaly coefficient and expression-action anomaly coefficient, and generating a composite anomaly index in combination with the voice weight correction coefficient, text weight correction coefficient, and action weight correction coefficient; An instruction integration and output unit: generating a decision instruction set according to the composite anomaly index and the preset index-instruction database, where the decision instruction set includes emotional tendency parameters, knowledge graph association indexes, and feedback time series.
[0012] Preferably, the knowledge retrieval module includes: A knowledge fragment search unit: according to the knowledge graph association index in the decision instruction set, traversing the preset distributed heterogeneous database, identifying the set of knowledge fragments matching the current interaction intention, preferentially matching the high-frequency accessed knowledge nodes in the user's historical interaction during the retrieval process, and sorting the association degrees of the knowledge fragments; An emotional adaptability correction unit: based on the emotional tendency parameters in the decision instruction set, performing emotional adaptability matching on the retrieved knowledge fragments, identifying the words and expressions that need to be adjusted, and performing tone word embedding and sentence structure optimization in combination with the preset tone mode and response rhythm preferred by the user to generate a response text that conforms to the current interaction emotional state; A voice feature annotation unit: according to the response text, annotating the corresponding voice feature parameters for different semantic paragraphs, including the intonation rise and fall range, speech rate adjustment ratio, and stress distribution position, and generating voice features that can drive the voice synthesis module.
[0013] Preferably, the voice generation module includes: Mouth shape matching unit: According to the timestamp distribution of the feedback time series, the response text is segmented into corresponding pronunciation syllable segments, the speech features of each syllable segment are extracted, and the preset standard mouth shape animation templates of the virtual digital human are matched to generate the key frames of the mouth shape animation synchronized with the actual pronunciation; Micro-expression regulation unit: Based on the intensity and type of the emotional tendency parameters, select the corresponding level of micro-expression parameter combinations from the preset expression library, including the movement amplitude of the eyebrows, the opening and closing frequency of the eyes, and the upward angle of the corners of the mouth, and adjust the micro-expression change rate according to the rhythm of the feedback time series to form a micro-expression parameter sequence matching the speech emotional expression; Body movement choreography unit: Identify the key semantic nodes in the response text. If the speech features contain emphatic stress or long pauses, insert the preset body movement templates at the corresponding moments, including the gesture swing trajectory and the torso tilt angle, and the movement amplitude is scaled proportionally according to the emotional tendency parameters to obtain the body movement trajectory; Multi-modal synchronization unit: Perform time-axis calibration on the key frames of the mouth shape animation, the micro-expression parameter sequence, and the body movement trajectory to ensure that the start moment of the voice response, the stress trigger point, and the peak moment of the movement trajectory are aligned, generate the voice response, and push it to the user terminal after rendering calculations at the edge computing node.
[0014] Compared with the prior art, the beneficial effects of the present application are as follows: By integrating multi-modal sensors and the virtual digital human behavior decision-making model, a more intelligent, real-time, and emotion-adaptive interactive Q&A system is realized, which can simultaneously process voice, text, facial expressions, and body movements, dynamically adjust the response strategy of the virtual digital human, and generate an accurate interactive feature matrix through spatio-temporal alignment cross-modal fusion, improving the accuracy of emotion recognition and response. The knowledge retrieval module optimizes the response content according to the user's emotional tendency, and the voice generation module realizes the synchronization of mouth shape, micro-expression, and body movements, ensuring that the feedback of the virtual digital human is more natural and meets the user's needs. The overall solution significantly improves the interactive experience and adaptability of the virtual digital human, and is applicable to various application scenarios such as intelligent customer service and virtual assistants. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0016] Figure 1 FIG. is a schematic structural diagram of an intelligent real-time interactive Q&A system based on a virtual digital human provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts fall within the scope of protection of the present invention.
[0018] Embodiment 1: The embodiment of the present invention provides an intelligent real-time interactive Q&A system based on a virtual digital human, as Figure 1 shown, including: Data acquisition module: Based on an integrated multi-modal sensor array, it receives the voice data and text data input by the user. At the same time, it real-time collects the sequence of facial expression dynamic parameters and the sequence of limb movement data of the user, and processes the voice data and text data to obtain standardized voice feature vectors and structured text data; Cross-modal fusion module: Aligns the standardized voice feature vectors, structured text data with the sequence of facial expression dynamic parameters and the sequence of limb movement data of the user in space and time, and constructs an interaction feature matrix containing multi-modal temporal and spatial correlation features; Behavior decision module: Determines a decision instruction set containing emotional tendency parameters, knowledge graph association indexes, and feedback time series based on historical interaction data and the interaction feature matrix; Knowledge retrieval module: According to the knowledge graph association index in the decision instruction set, retrieves the corresponding set of knowledge fragments from a preset distributed heterogeneous database, and performs mood word embedding and prosody rhythm annotation on the vocabulary corresponding to the knowledge fragments in the set of knowledge fragments according to the emotional tendency parameters, generating an answer text and voice features with emotional adaptability; Voice generation module: Generates the key frames of the mouth shape animation, the sequence of micro-expression parameters, and the limb movement trajectories of the virtual digital human based on the feedback time series, answer text, and voice features in the decision instruction set, and generates a voice response, which is pushed to the user terminal through an edge computing node.
[0019] In this embodiment, the knowledge graph association index refers to a semantic retrieval pointer system constructed based on a domain knowledge graph, specifically including: 1) Node encoding rule: Hash encoding (SHA-3 algorithm) is used to map knowledge nodes (such as diseases, drugs, operation steps) to 128-bit unique identifiers, where the first 32 bits represent the domain classification (medicine / finance / education); 2) Relationship weight calculation: The embedding vectors of the relationships between entities (such as "treatment", "taboo") are learned through the TransE model. When the user queries "diabetes complications", the adjacent nodes with a relationship weight > 0.85 are retrieved preferentially; 3) Dynamic path optimization: Combining the term usage frequency in the user's historical interactions, the A* algorithm is used to generate the optimal retrieval path with a maximum of 3 hops in the knowledge graph. This index is different from the traditional database primary key and can achieve cross-modal queries (such as directly associating the content of voice consultation with the image of the drug molecular formula), and the response delay is controlled within 50ms.
[0020] In this embodiment, the distributed heterogeneous database refers to a hybrid architecture storage system designed specifically for multi-modal knowledge storage. Its core technical features include: 1) Hierarchical storage strategy - Structured data (drug instructions) are stored in a PostgreSQL cluster, and unstructured data (CT images) are placed in the IPFS distributed file system, and a unified query interface is implemented through a GraphQL gateway; 2) Semantic caching mechanism: A semantic similarity matching cache based on BERT is deployed at the edge node. When the user continuously asks questions such as "symptoms of angina pectoris" and "manifestations of myocardial ischemia", the answer blocks with a semantic similarity > 92% in the cache are directly returned; 3) Cross-domain consistency protocol: An improved RAFT algorithm is used to coordinate data synchronization in multiple data centers to ensure strong consistency (CPA model) for financial data and eventual consistency (BASE model) for entertainment data. This system supports 100,000-level concurrent queries per second, and the data sharding granularity is refined to a single medical concept (such as "aspirin" corresponding to 3 data shards).
[0021] The beneficial effects of the above technical solutions are as follows: By integrating multi-modal sensors and the virtual digital human behavior decision-making model, a more intelligent, real-time, and emotion-adaptive interactive Q&A system is realized. It can process voice, text, facial expressions, and body movements simultaneously, dynamically adjust the response strategy of the virtual digital human, generate an accurate interactive feature matrix through spatio-temporal alignment cross-modal fusion, and improve the accuracy of emotion recognition and response. The knowledge retrieval module optimizes the answer content according to the user's emotional tendency, and the voice generation module realizes the synchronization of lip movements, micro-expressions, and body movements, ensuring that the feedback of the virtual digital human is more natural and meets the user's needs. The overall solution significantly improves the interactive experience and adaptability of the virtual digital human and is applicable to various application scenarios such as intelligent customer service and virtual assistants.
[0022] Embodiment 2: An embodiment of the present invention provides an intelligent real-time interactive Q&A system based on a virtual digital human, and a data acquisition module, including: A voice processing unit: performs frequency-domain energy detection on the received voice data, identifies and eliminates invalid voice segments with energy lower than a preset threshold, and obtains continuous valid voice waveform data; A text analysis unit: performs dependency syntactic analysis on the text data, annotates entity types and action predicates, and generates initial text data with semantic tags; An environment analysis unit: collects three-dimensional space information of the interaction environment, tracks the movement trajectory of the target user in the interaction space, and calculates the spatial position coordinates and orientation angle data of the user; An image acquisition unit: dynamically adjusts the focal length of the visual sensor and the directivity parameters of the sound pickup array according to the spatial position coordinates and orientation angle data, and obtains the facial expression image and limb movement image data of the user; An image analysis unit: analyzes the spatial offset of key feature points in the facial expression image, measures the deformation amplitude and duration of a specific facial area, and generates an expression dynamic parameter sequence; An action analysis unit: extracts the spatial coordinates of key bone nodes in the limb movement image data, determines the relative movement trajectory between each node, and generates limb movement sequence data; A data processing unit: processes the continuous valid voice waveform data and the initial text data to obtain a normalized voice feature vector and structured text data.
[0023] In this embodiment, frequency-domain energy detection is performed on the received voice data, invalid voice segments with energy lower than a preset threshold are identified and eliminated, continuous valid voice waveform data is obtained, and its spectral peak distribution parameters are calculated to generate a normalized voice feature vector. Specifically, it includes: adopting a spectral analysis method based on short-time Fourier transform, performing frame processing on the voice signal with a frame length of 20 ms, and calculating the total energy value of each frame signal within the effective frequency band of 100 - 4000 Hz; setting a dynamic energy threshold, and when the energy of a certain frame is lower than this threshold and the continuous duration exceeds 300 ms, it is determined as an invalid voice segment and eliminated; for the remaining valid voice segments, extract the 12-dimensional features of its Mel frequency cepstral coefficients (MFCC), and statistically analyze the position distribution and amplitude ratio of the spectral peaks of each frame to generate a 128-dimensional normalized voice feature vector containing frequency-domain energy distribution and time-domain continuity features; at the same time, detect the silent segments and non-human voice noises in the voice, and combine the double-threshold endpoint detection algorithm to optimize the determination accuracy of the voice start and end points. Among them, the voice signal refers to the continuous audio waveform data containing human voice information collected during the interaction between the user and the virtual digital human, and its frequency domain range is usually 80 Hz - 8 kHz (covering the fundamental frequency and main formants of spoken language), and its time domain is an analog or digital signal with amplitude changing over time.
[0024] In this embodiment, dependency syntactic analysis is performed on the text data, entity types and action predicates are annotated, and initial structured text data with semantic labels is generated. Specifically, it includes: using a neural dependency syntactic analysis model based on the attention mechanism to identify the core predicates and their dependency relationships in the input text and construct a syntactic dependency tree; using named entity recognition technology to annotate entity categories such as person names, place names, and organization names in the sentence, and combining with a domain knowledge base to expand the professional term dictionary; establishing a semantic role framework for action predicates and annotating semantic roles such as agent, patient, time, and place and their constraint conditions; finally generating structured text data, which contains four layers of information: the original text, the dependency relationship graph, the entity annotation result, and the semantic role framework. Each layer of information is associated through a unified identifier to support the subsequent intention recognition and knowledge retrieval module calls.
[0025] In this embodiment, three-dimensional spatial information of the interaction environment is collected, the movement trajectory of the target user in the interaction space is tracked, the spatial position coordinates and orientation angle data of the user are calculated, and the visual sensor parameters are dynamically adjusted. Specifically, it includes: deploying a binocular stereo vision camera array to reconstruct the three-dimensional point cloud data of the scene in real time based on the parallax principle; using an improved Kalman filtering algorithm to fuse RGB-D sensor and UWB positioning data to track the spatial coordinates of the key points of the user's head, with a position update frequency of 30Hz; calculating the orientation angle by fitting the spatial vector connecting the user's shoulders, with the accuracy controlled within ±3°; dynamically adjusting the camera focal length according to the user's distance to keep the face area occupying 40%-60% of the screen area within the range of 1.5-5 meters; at the same time, controlling the main lobe of the beamforming microphone array to point to the user's mouth area to suppress the environmental noise by more than 15dB and ensure the voice collection quality.
[0026] In this embodiment, the spatial offset of key feature points in the facial expression image is analyzed, the deformation amplitude and duration of a specific facial area are measured, and a sequence of expression dynamic parameters is generated; at the same time, the spatial coordinates of key bone nodes in the limb movement image are extracted to determine the relative movement trajectory between each node. Specifically, it includes: using a deep learning-based facial feature point detection algorithm to real-time track the two-dimensional coordinates of 68 facial key points and calculate the displacement vectors of the control areas of facial muscles such as the corrugator supercilii and zygomaticus major; establishing a normalized facial action coding system (FACS) to map the displacement vectors to the activation intensity and time curve of AUs (Action Units); for limb analysis, extracting the three-dimensional coordinates of 25 human bone nodes through a lightweight OpenPose model to construct a spatio-temporal graph model of limb movements; calculating kinematic parameters such as joint angle change rate and end trajectory curvature, and combining with a preset action template library to identify typical interaction actions such as waving and nodding, and generating an action description sequence containing spatio-temporal features and semantic labels.
[0027] In this embodiment, speech data and text data are processed to obtain a standardized speech feature vector and structured text data, which are then spatio-temporally aligned and feature-fused with other perception data. Specifically, it includes: using the Dynamic Time Warping (DTW) algorithm to align the time axes of the speech feature sequence and the facial expression parameter sequence to ensure that the time deviation between modalities is less than 80 ms; constructing a cross-modal feature fusion model based on a graph neural network to encode the MFCC features of speech, the dependency graph of text, the activation patterns of facial AUs, and the spatio-temporal graph model of the body into a unified 256-dimensional joint feature vector; calculating the contribution weights of the features of each modality through an attention mechanism, assigning a weight coefficient of 0.6 to the facial expression features in the speech emotion recognition task, and increasing the weight of the text semantic features to 0.8 in the intention understanding task; finally, outputting a multi-level data structure containing the original data, the temporal alignment result, and the fusion features, providing a complete environmental perception input for subsequent interaction decisions.
[0028] The beneficial effects of the above technical solution are as follows: Through multi-dimensional data collection and processing, comprehensive acquisition and intelligent analysis of the user's speech, text, facial expressions, body movements, and interaction environment information are achieved. Speech processing and vector generation improve the effectiveness of speech input and the accuracy of feature extraction; text analysis enhances semantic understanding ability; environmental analysis and image acquisition achieve accurate perception of the user's state; image and motion analysis provide dynamic parameter inputs for expressions and actions. The overall solution provides a comprehensive, dynamic, and accurate interaction data foundation for virtual digital humans, significantly improving the intelligent response ability and the naturalness of human-computer interaction of the system.
[0029] Embodiment 3: The embodiment of the present invention provides an intelligent real-time interactive Q&A system based on a virtual digital human, and a data processing unit, including: Speech data processing sub-unit: Use an adaptive noise cancellation algorithm to preprocess the continuous valid speech waveform data, extract acoustic features, and generate a standardized speech feature vector containing dynamic time warping information; Text data processing sub-unit: Perform word segmentation processing on the initial text data through a domain-adaptive semantic word segmentation model, and combine it with a preset intention recognition knowledge base for intention annotation to generate structured text data with context semantic associations.
[0030] In this embodiment, the adaptive noise cancellation algorithm is used to preprocess the voice request, extract acoustic features, and generate a standardized voice feature vector containing dynamic time warping information, including: using an adaptive filtering algorithm based on frequency domain least mean square (FLMS) to estimate the power spectral density of background noise in real time, constructing a multi-level noise reference model, realizing dynamic noise reduction in the frequency band of 20 - 8000 Hz, and the signal-to-noise ratio improvement amplitude ≥ 12 dB; extracting 26-dimensional acoustic features from the denoised voice signal, including 12-dimensional Mel frequency cepstral coefficients (MFCC), 1-dimensional fundamental frequency (F0), 6-dimensional chromatic features, and 7-dimensional spectral centroid features, to form a time-varying feature matrix; using the dynamic time warping (DTW) algorithm to align voice segments of different durations, calculating the distance metric matrix of frame-level features, and generating a 128-dimensional standardized voice feature vector through Gaussian normalization processing; at the same time, detecting the emotional intonation features in the voice, and based on the fundamental frequency trajectory change rate (ΔF0 / Δt) and the energy envelope fluctuation to mark emotional states such as excitement, anger, or calmness, to form an enhanced voice feature descriptor.
[0031] In this embodiment, the text request is segmented by a domain-adaptive semantic segmentation model, and combined with a preset intention recognition knowledge base for intention annotation, generating structured text data with context semantic association, including: adopting a hybrid segmentation strategy, integrating the BiLSTM-CRF neural network and the domain dictionary matching method. For example, on the basis of general Chinese word segmentation (accuracy ≥ 98%), optimizing the segmentation of professional terms for the intelligent customer service scenario (such as "5G package" cannot be split); constructing a hierarchical intention recognition knowledge base, the first layer matches high-frequency intention templates through a rule engine (such as "query traffic" is associated with the "service handling" category), and the second layer calculates the semantic similarity between the user query and the knowledge base questions based on the BERT model, and conducts fine-grained intention classification for requests that do not hit the template; the output structured text data includes: 1) the word segmentation sequence and part-of-speech tagging of the original text, 2) the predicate-argument structure generated by dependency syntactic analysis, 3) the intention category and confidence score, 4) the associated context conversation state (such as the order number involved in the previous round of conversation), supporting multi-round semantic inheritance and ambiguity resolution.
[0032] The beneficial effects of the above technical solutions are: by introducing the adaptive noise cancellation algorithm in the voice data processing subunit, the clarity of voice input and the accuracy of feature extraction in complex environments are effectively improved; at the same time, the text data processing subunit combines the domain-adaptive semantic segmentation model and the intention recognition knowledge base, which can accurately identify the user's intention and generate structured text data with complete semantics. The overall solution realizes high-quality preprocessing and semantic enhancement of voice and text data, provides an accurate and efficient input basis for subsequent cross-modal fusion and behavior decision-making, and thus significantly improves the interaction understanding ability and response quality of virtual digital humans.
[0033] Example 4: An embodiment of the present invention provides an intelligent real-time interactive Q&A system based on a virtual digital human. The cross-modal fusion module includes: Data alignment unit: Perform spatio-temporal alignment on the standardized speech feature vectors, structured text data, the sequence of the user's facial expression dynamic parameters, and the sequence data of limb movement sequences; Weight adjustment unit: Adjust the weights of the standardized speech feature vectors, structured text data, the sequence of facial expression dynamic parameters, and the sequence data of limb movement sequences after spatio-temporal alignment; Matrix construction unit: Integrate the standardized speech feature vectors, structured text data, the sequence of the user's facial expression dynamic parameters, the sequence data of limb movement sequences, and the corresponding adjusted weights to construct an interaction feature matrix containing multi-modal temporal correlation features.
[0034] In this embodiment, performing spatio-temporal alignment on the standardized speech feature vectors, structured text data, the sequence of facial expression dynamic parameters, and the sequence data of limb movement sequences includes: adopting a multi-modal synchronization algorithm based on dynamic time warping (DTW), using the speech feature sequence as the reference time axis, calculating the relative temporal offset of the facial expression parameter sequence (sampled at 50 Hz) and the limb movement sequence (sampled at 30 Hz), and achieving millisecond-level alignment accuracy through spline interpolation with a preset number of times (usually three times); for the text data, generate time tags according to the word-level timestamps of the speech recognition results to associate with the nodes of the dependency syntax tree to ensure the alignment of semantic units and speech segments; at the same time, establish a cross-modal trigger detection mechanism. When it is recognized that the "eyebrow raising" action in the facial expression parameter sequence and the interrogative intonation (F0 rising by 20%) in the speech feature co-occur within ±200 ms, forcefully synchronize and mark it as an "interrogative state" event; finally, output the multi-modal data stream with time alignment, and each modal data is marked with a unified time code and event trigger flag, providing a standardized input for subsequent fusion.
[0035] In this embodiment, the weights of each modality are dynamically adjusted based on the spatio-temporal aligned multi-modal data, including: constructing a weight allocation network based on the attention mechanism, where the input layer receives cross-modal features such as the aligned speech spectrum energy, expression AU intensity, and limb joint angle change rate; analyzing the temporal correlation of each modality feature through a gated recurrent unit (GRU). When it is detected that the "arm waving" appears in the limb action sequence (spatial trajectory variance > 0.3) and overlaps with the speech energy peak (> 65 dB), the limb modality weight is increased from 0.2 to 0.5; using a differentiable decision tree model to achieve context adaptive adjustment, assigning a base weight of 0.7 to the speech and text modalities in the question-and-answer scenario, and increasing the expression modality weight to 0.6 in the emotional interaction scenario; the output layer generates a dynamic weight matrix including the time dimension, and the weight value of each modality at any time t satisfies the constraint condition of ∑w(t)=1, and the sigmoid function is used to prevent gradient explosion.
[0036] In this embodiment, integrating the multi-modal data and dynamic weights to construct an interaction feature matrix, including: designing a hierarchical fusion architecture, where the bottom layer maps the speech MFCC features (128 dimensions), text BERT embeddings (768 dimensions), expression AU parameters (32 dimensions), and limb bone coordinates (75 dimensions) to a unified 256-dimensional latent space through a fully connected layer; the middle layer adopts a cross-modal attention mechanism to calculate the correlation scores between features guided by the dynamic weight matrix. For example, when the expression weight w > 0.4, the correlation between the mouth AU25 (lip opening) feature and the plosive segment in the speech spectrum is strengthened; the top layer outputs an interaction matrix containing a three-dimensional tensor of time-modal-feature, with dimensions [T×M×256] (T is the time step, M = 4 is the number of modalities), and adds learnable position encoding by time slicing; before the matrix is input into the pre-trained virtual digital human behavior decision model, it undergoes layer normalization and residual connection processing to ensure the stability of gradient flow, and finally supports simultaneously driving the speech synthesis, facial animation, and limb movement generation of the digital human.
[0037] The beneficial effects of the above technical solutions are: through the precise data alignment and weight adjustment of the cross-modal fusion module, it is possible to effectively synchronize the speech, text, expression, and limb movement data, ensuring the efficient fusion and temporal correlation of multi-modal information. The weight adjustment unit dynamically adjusts according to the actual importance of each data source, improving the accuracy and emotional adaptability of the virtual digital human response. By constructing an interaction matrix of multi-modal temporal correlation features, it is possible to generate more natural, intelligent, and emotionally expressive interaction feedback, significantly enhancing the user experience and providing reliable support for the application of intelligent virtual digital human technology.
[0038] Embodiment 5: An embodiment of the present invention provides an intelligent real-time interactive question-and-answer system based on a virtual digital human, a weight adjustment unit, including: Coupling subunit: Perform coupling analysis on the aligned and standardized speech feature vectors, expression dynamic parameter sequences, and limb movement sequence data to obtain the speech-expression coupling coefficient and the speech-movement coupling coefficient; First adjustment subunit: Perform a first adjustment on the weight allocation ratios of the aligned and standardized speech feature vectors, expression dynamic parameter sequences, and limb movement sequence data according to the speech-expression coupling coefficient and the speech-movement coupling coefficient; Parameter acquisition subunit: Perform micro-expression analysis on the expression dynamic parameter sequences and limb movement sequence data of the user. At the same time, perform speech fundamental frequency jitter detection on the aligned and standardized speech feature vectors to determine the user's physiological state parameters, where the user's physiological state parameters include the emotional tension index and the attention concentration score; Second adjustment subunit: Perform a second adjustment on the weights of the aligned and standardized speech feature vectors, expression dynamic parameter sequences, and limb movement sequence data based on the user's physiological state parameters; Third adjustment subunit: Perform a third adjustment on the preset initial weights of the structured text data based on the weights of the second-adjusted standardized speech feature vectors, expression dynamic parameter sequences, and limb movement sequence data.
[0039] In this embodiment, adjusting the weight allocation ratio according to the speech-expression coupling coefficient Kve and the speech-movement coupling coefficient Kvb includes: designing a dual-channel adaptive weighting module. When Kve > 0.6 and Kvb < 0.3, the expression modality weight is increased from 0.3 to 0.55, and the limb modality weight Wbody is decreased to 0.15; adopting a logarithmic ratio allocation strategy, the speech weight Wvoice = 1 - α(Kve + Kvb), where α = 0.35 is an empirical attenuation factor; for coupling conflict scenarios (such as Kve > 0.8 but the limb performance is rigid), start a peak suppression mechanism to forcibly limit the growth of the expression weight not exceeding 120% of the coupling coefficient; the adjustment result is processed by a time smoothing filter to prevent the weight jump from causing the digital human movement to be incoherent, and finally output a weight vector with timestamps [wv(t), wf(t), wb(t)], which are the speech weight vector, the expression weight vector, and the limb weight vector respectively, satisfying the real-time update frequency ≥ 10Hz; In this embodiment, physiological state parameters are determined based on micro-expression analysis and voice fundamental frequency jitter detection, including: using a 3DCNN model to detect minute facial muscle movements (lasting frames < 1 / 3 second), and when the levator labii superioris (AU10) and eyelid tension (AU7) co-occur, marking the emotional tension index T ∈ [0, 1]; calculating jitter (jitter rate) and shimmer (amplitude perturbation) through an improved fundamental frequency trajectory analysis algorithm, and when jitter > 1.2%, the attention concentration score C decreases by 0.2 per second; fusing behavioral features such as the head pose angle (pitch > 15°) and hand contact frequency (touching the face > 3 times per minute), constructing a physiological state evaluation model based on LightGBM, and outputting an emotion-attention binary tuple (Tt, Ct) containing timestamps, where when Tt > 0.7, the "high-pressure state" flag is triggered, and when Ct < 0.4, the "distraction warning" is activated; In this embodiment, a second weight adjustment is performed based on physiological state parameters, including: when T > 0.7, introducing an emergency compensation strategy to increase the voice weight vector by an increase value of Δwv = 0.2(1 - C) to enhance the information transmission efficiency; for the distracted state with C < 0.5, adopting a limb movement reinforcement scheme to make the limb weight vector wb += 0.15×∑(limb joint movement amplitude) / T, and attracting the user's attention again through large movements; designing a weight arbiter based on fuzzy logic, inputting the time difference values ΔT / Δt and ΔC / Δt of T and C, and when a rapid increase in tension (ΔT / Δt > 0.1 / s) is detected, immediately reducing the expression weight by 20% to avoid excessive emotional rendering; finally, the output dynamic weight is superimposed on the first adjustment result to form a weight distribution with physiological state adaptability; In this embodiment, the third adjustment includes: after the second adjustment is completed, the system has obtained the weight allocation ratios of voice, expression, and movement (such as 60% for voice, 25% for expression, and 15% for movement). At this time, structured text data needs to be incorporated into the weight adjustment system. The system adopts semantic-modal association rules to pre-define the mapping relationships between different text types and multi-modal data. For example, interrogative sentences (such as "Why?") rely more on the weight of voice intonation, while descriptive texts (such as "I went to the park yesterday") rely more on the coordination of expression and body movements. The system will parse the syntactic structure, sentiment polarity, and keywords of the current text, match the pre-set modal preference templates, generate the initial weight adjustment direction of the text (such as increasing or decreasing the proportion of the text's influence on the final decision), and evaluate the reliability of the current text data through ASR confidence scoring and semantic coherence analysis. If the text confidence is lower than the threshold (such as semantic ambiguity caused by speech-to-text errors), its weight will be reduced, and dynamic compensation will be performed according to the coupling coefficient of voice-expression-movement. For example, when the user says "Yes" but shakes their head, the weight of the text "Yes" will be lowered, and the conflict between the affirmative voice tone and the shaking head movement will trigger modal arbitration, ultimately giving priority to voice and movement. In addition, the system will detect the semantic consistency between the text and the expression / movement (such as the text "happy" but the expression parameters show frowning), and locally reduce the weight of inconsistent text segments; to avoid unnatural feedback from the digital human caused by frequent weight jumps, the system adopts a time-window recursive smoothing strategy. The final weight of the current text segment not only depends on the real-time analysis results but also refers to the historical average weight of the same type of text within the previous 5 seconds. For example, if the user asks questions in a high tone continuously for multiple times, the text weight of subsequent interrogative sentences will inherit the upward trend of the previous ones instead of being recalculated. At the same time, the system performs low-pass filtering on short-term burst noises (such as text jumps caused by speech misrecognition) to ensure that the weight change curve conforms to the inertia and gradualness of human interaction; for long-term interacting users, the system will record their historical adjustment preferences for text weights in the interaction and build a personalized weight model. For example, some users are used to using body movements to replace verbal expressions (such as nodding instead of saying "Yes"), and the system will gradually reduce the default weight of the text in such scenarios and increase the priority of movement data. The adaptive reinforcement module updates the user profile in real time through Online Learning, and the weight results after each adjustment will be fed back to the user preference database to form a closed-loop optimization; after all adjustment steps are completed, the system normalizes the weights of voice, text, expression, and movement to ensure that the sum of the four is 100%. The normalization process introduces the Softmax function to avoid a single modal weight being suppressed to zero. The final output weight combination will be used as the coefficient of the interaction feature matrix and input into the virtual digital human behavior decision-making model. For example, the output result may be: 40% for voice, 25% for text, 20% for expression, and 15% for movement, and the digital human will comprehensively generate feedback behaviors with both semantic accuracy and emotional expressiveness based on this.
[0040] In this embodiment, coupling analysis is performed on the aligned standardized speech feature vectors, expression dynamic parameter sequences, and limb movement sequence data, including determining the speech-expression coupling coefficient Kve based on the aligned standardized speech feature vectors and expression dynamic parameter sequences, and determining the speech-action coupling coefficient Kvb based on the aligned standardized speech feature vectors and limb movement sequence data; In this embodiment, the speech-expression coupling coefficient Kve is used to evaluate the temporal correlation between the speech energy change and the facial expression intensity change. The speech feature sequence: , each is a feature vector with a dimension of . The expression parameter sequence: , each contains dynamic expression parameters (such as AU intensity). The time window length: represents the number of frames in the sliding window. The calculation process: Extract the change rate (within the sliding window): Speech change rate: Expression change rate: Calculate the Pearson correlation coefficient: where, is the Pearson correlation coefficient of the speech change rate and the expression change rate, measuring the synchronization degree of the two change rates. The result value range is . The closer it is to 1, the more positively correlated it is. Take as the speech-expression coupling coefficient Kve to describe the linkage (coupling strength) between speech and expression. is the speech energy change rate sequence within the sliding window; is the expression intensity change rate sequence representing within the sliding window; is the covariance of the two; and are respectively and standard deviations; In this embodiment, the speech-action coupling coefficient Kvb is a measure of the frequency matching degree between the speech rhythm and the limb movement. The speech-action coupling coefficient measures the dynamic matching degree between the speech intonation and the limb movement. The speech feature sequence: , the limb movement sequence: , each is an -dimensional vector composed of joint angles or displacements. L is the joint angle / displacement dimension. The time window length: , calculation process: Extract the speech rhythm frequency (based on the average period of the fundamental frequency ): Extract the action rhythm frequency (based on the frequency domain analysis of the action energy change rate): Calculate the coupling coefficient (based on the normalization of the frequency difference): Among them, is used as the speech-action coupling coefficient Kvb, and the value range is , the closer to 1 indicates that the rhythm is more coordinated and synchronous. is the action rhythm frequency, that is, the main frequency period of the action in window t. is the speech rhythm frequency, that is, the average rhythm period of the speech in window t. is the change rate of limb movement energy. is the absolute value of the frequency difference (the smaller the value, the more synchronous). is the normalization factor to ensure that the output is within the range of [0, 1]. In this embodiment, the calculation process of the user's physiological state parameters: Step 1: Micro-expression recognition and AU intensity analysis. Identify facial micro-expressions and estimate the degree of emotional tension to obtain the micro-expression frequency , method: Use the AU parameters in the FACS system (such as AU2: eyebrow raising, AU12: lip corner tightening) to calculate the AU intensity difference between consecutive frames. If the change rate of a certain AU intensity exceeds the dynamic threshold (such as > 30%) within a very short time (< 500 ms), it is regarded as a micro-expression event, and the corresponding AU intensity change rate is determined as the micro-expression frequency. , meaning: The frequency and amplitude of micro-expressions are positively correlated with the degree of tension. For example, frequent lip pursing (AU14) or frowning (AU4) often indicate that the user is in a stressed state. Normalize the micro-expression frequency ; Step 2: Limb movement stability assessment. Detect the stability change of the user's body and judge the attention state. Method: Perform variance analysis within a sliding window on the joint angular velocity vector : Among them, is the joint velocity vector of the t-th frame. is the average velocity within the window. is the number of frames in the window. If the limb movement stability score Higher than the set benchmark (the threshold in the relaxed state), indicating abnormal jitter or restlessness. Meaning: High variance usually reflects distracted attention or emotional uneasiness and can be used as an indirect indicator of distraction; Step 3: Analysis of fundamental frequency jitter of speech, evaluating the frequency fluctuation characteristics reflecting emotional fluctuations in the speech signal. Method: Perform dynamic time warping on the fundamental frequency (F0) sequence in the standardized speech feature vector and extract its short-time variation coefficient (jitter): Where is the fundamental frequency jitter rate, is the length of the analysis window, is the fundamental frequency value of the t-th frame, and the change trend of the fundamental frequency mean (such as a sudden increase) is combined as an auxiliary signal for tension. An increase in the jitter rate (such as exceeding 5%) usually corresponds to a state of emotional tension. At the same time, the system further corrects the emotional tension index by combining the offset of the fundamental frequency mean (such as a sudden increase may indicate excitement); Step 4: Fusion modeling and calculation of emotion / attention index, integrating three-modal information to calculate a unified physiological state parameter. Method: Emotional tension index Adopt a weighted fusion formula: Where : Normalized micro-expression frequency; Default modal weight: , attention concentration score : A combination of the reciprocal of action stability and speech pause frequency (the specific expression can be extended to a linear or non-linear mapping). Meaning: The higher the value, the more tense; the lower the value, the more distracted the attention; To prevent instantaneous noise interference (such as a sudden cough by the user resulting in abnormal speech), the system uses a sliding median filter to smooth EEE and AAA, retaining the trend changes. The finally output emotional tension index and attention concentration score are quantified on a 0 - 100 scale and are accompanied by a confidence mark (such as a low-confidence result does not trigger weight adjustment).
[0041] The beneficial effects of the above technical solution are: By introducing a coupling subunit and a three-layer weight adjustment mechanism, the system can accurately capture the coupling relationship between speech, expression, and movement, and dynamically optimize the weights of each modal data in combination with the user's physiological state parameters. This mechanism not only improves the accuracy of multi-modal information fusion and the context adaptation ability, but also enhances the behavior decision-making sensitivity and personalized response ability of the virtual digital human in different user states, thus realizing a more natural and emotion-adaptive human-computer interaction experience.
[0042] Example 6: The embodiment of the present invention provides an intelligent real-time interactive Q&A system based on a virtual digital human. The behavior decision-making module includes: A historical data weight adjustment unit: determining the multimodal weight distribution of user interaction behaviors based on historical interaction data and an interaction feature matrix, obtaining the short-term interaction frequency, multimodal signal volatility, and emotional tendency deviation degree, performing multimodal confidence analysis and interaction habit matching, and obtaining a voice weight correction coefficient, a text weight correction coefficient, and an action weight correction coefficient; A cross-modal correlation network construction unit: monitoring the signal synchronization between the standardized voice feature vector, structured text data, and expression dynamic parameter sequence during the interaction process, determining the voice-text synchronization rate and the voice-expression synchronization rate, and generating a multimodal synchronization relationship network; A feedback decision splitting prediction unit: performing semantic parsing on the structured text data, extracting interaction intention tags, and determining the feedback level division ratio and the feedback rhythm adjustment coefficient in combination with the emotional tendency deviation degree; An associated multimodal anomaly analysis unit: based on the multimodal synchronization relationship network, using the feedback level division ratio and the feedback rhythm adjustment coefficient to perform cross-modal consistency verification, determining the text-voice anomaly coefficient and the expression-action anomaly coefficient, and generating a composite anomaly index in combination with the voice weight correction coefficient, the text weight correction coefficient, and the action weight correction coefficient; An instruction integration output unit: generating a decision instruction set according to the composite anomaly index and a preset index-instruction database, where the decision instruction set includes emotional tendency parameters, knowledge graph association indexes, and feedback time series.
[0043] In this embodiment, the historical data weight adjustment unit dynamically weights the multimodal historical data in the interaction feature matrix through a pre-trained time series analysis model. Specifically, the decay coefficient is calculated based on the short-term interaction frequency of 10 consecutive interactions, and the initial weight is generated in combination with the multimodal signal volatility (such as voice spectrum variance, text word frequency change rate); at the same time, the emotional tendency deviation degree is obtained by analyzing the user's historical emotional label sequence (such as the change in the positive / negative ratio) through an LSTM network. The unit outputs the weight correction coefficients of voice, text, and action (range 0.8 - 1.2), and the correction rule is: when the emotional deviation degree exceeds the threshold of ±0.3, the weight of the corresponding modality is increased by 15%. This process performs incremental updates every 5 seconds to ensure that the weights match the user's real-time interaction habits.
[0044] In this embodiment, the cross-modal correlation network construction unit establishes a synchronization relationship graph between multimodal signals, realizes coordinated modeling and topological analysis between modalities. Key steps: Synchronization rate calculation: Voice-text synchronization rate: ; Voice-expression synchronization rate: ; Graph construction: Create a node set , edge weight , forming a dynamic weighted graph , Graph attention mechanism and modal center activation: If and , then activate the "expression" as the central node to guide the outward propagation of information; Modal clustering recognition: Update the graph structure every 200ms, and identify modal clusters through community discovery algorithms; When the modularity of the "expression-action" > 0.7, it is marked as the non-verbal expression dominant state for reference by the anomaly detection module.
[0045] In this embodiment, the feedback decision splitting prediction unit dynamically constructs a feedback strategy according to semantic content and emotion signals, splitting resource ratio and rhythm control. Key steps: Interaction intention recognition: Based on the dependency syntactic tree + BiLSTM-CRF structure, extract 28 types of interaction intention labels (such as "social" "complaint"), Feedback level allocation: Use the emotional offset Calculate the feedback resource ratio triple: ; If , allocate 70% of the feedback resources to the "emotional response" layer, Rhythm control mechanism: Construct a reinforcement learning rhythm controller, input the speech pause duration and dependency distance, and output the feedback rhythm coefficient ; If the user's speech rate > 5 words / second, then accelerate the response , Output the strategy triple: As the global control parameter for multimodal feedback.
[0046] In this embodiment, the associated multimodal anomaly analysis unit improves the accuracy and interpretability of anomaly detection through the consistency check between multimodal signals and the fusion of cross-anomaly indicators. Key steps: Anomaly indicator calculation: Text-to-speech anomaly coefficient: ; Expression-action anomaly coefficient: ; Composite anomaly index construction: ; Among them, the weights are the text-to-speech modal weight and the expression-action modal weight respectively, dynamically adjusted according to the scene, range [0,1]; is the voice credibility attenuation factor, range [0,0.5], when the signal-to-noise ratio of the voice signal (SNR < 15dB = 0.4), or the fundamental frequency has abnormal jitter (when the jitter rate > 10% = 0.3), is the text credibility attenuation factor, range [0,0.5], text logical contradiction (such as when there are semantic conflict sentences in the same round of conversation = 0.2), or when the keyword confidence < 70% = 0.4, is the background interference penalty term, with a range of [0, 0.3]. When the ambient light intensity suddenly changes (when > 200 lux / ms = 0.2) or there is sudden noise (when the energy suddenly increases by 30 dB = 0.3), dynamic weight constraint: α + β ≤ 1, and λv + λt + μb ≤ 1, ensuring that A ∈ [0, 2], is the compensation term. When the reliability of a single modality decreases (such as when speech is contaminated by noise), the compensation mechanism is used to prevent the A value from being falsely high; three-level response mechanism: level 1 (0.6 < A ≤ 0.8): reduce the weight of the abnormal modality by 30%; level 2 (0.8 <a ≤ 1.0):插入澄清提问;三级(a>1.0): Initiate the emergency dialogue process; Abnormal topology update: Synchronize the analysis results to the abnormal attribute matrix of the multi-modal graph to achieve system-level traceability analysis support.
[0047] In this embodiment, the instruction integration output unit integrates all policy results and outputs a structured multi-modal driving instruction set to achieve joint control of voice, actions, expressions, etc., including: Emotional parameter mapping: If the composite abnormal index , select the "neutral comfort" feedback parameter (intensity 0.7), Knowledge node preference: In cases, preferentially select the explanatory nodes in the knowledge graph with a confidence level > 85%, Time series generator: Split according to the rhythm coefficient into: Voice segment: Last seconds; Action segment: Last seconds; Insert intermediate frames: Insert expression transition frames, Output instruction set: Output a structured control set with 17 fields in JSON format, covering semantic tags, feedback rhythm, modal control signals, action trajectories, etc.; Support millisecond-level response distribution.
[0048] The beneficial effects of the above technical solutions are: The behavior decision-making module dynamically optimizes the multi-modal weights through the fusion of historical interaction data and the interaction feature matrix, accurately capturing the user's interaction habits and emotional changes. The cross-modal association network ensures the high synchronization of voice, text, and expression signals, generates a composite abnormal index by combining semantic parsing and abnormal analysis, and improves the accuracy and context adaptability of feedback. The instruction integration output unit generates a decision instruction set containing emotional tendencies, knowledge indexes, and feedback time series according to the abnormal index, significantly enhancing the personalized response ability and interaction naturalness of the virtual digital human, and providing users with an immersive and emotionally resonant intelligent interaction experience.
[0049] Embodiment 7: The embodiment of the present invention provides an intelligent real-time interactive Q&A system based on a virtual digital human. The knowledge retrieval module includes: Knowledge fragment search unit: According to the knowledge graph association index in the decision instruction set, traverse the preset distributed heterogeneous database, identify the set of knowledge fragments that match the current interaction intention, preferentially match the high-frequency accessed knowledge nodes in the user's historical interactions during the retrieval process, and sort the relevance of the knowledge fragments; Emotional adaptability correction unit: Based on the emotional tendency parameters in the decision instruction set, perform emotional adaptability matching on the retrieved knowledge fragments, identify the words and expressions that need to be adjusted, and combine the preset user-preferred tone pattern and response rhythm to perform tone word embedding and sentence structure optimization to generate an answer text that conforms to the current interaction emotional state; Voice Feature Annotation Unit: According to the response text, annotate corresponding voice feature parameters for different semantic paragraphs, including the intonation rise and fall range, speech rate adjustment ratio, and stress distribution position, and generate voice features that can drive the speech synthesis module.
[0050] In this embodiment, according to the knowledge graph association index in the decision instruction set, traverse the preset distributed heterogeneous database, identify the set of knowledge fragments that match the current interaction intention, and give priority to matching the high-frequency accessed knowledge nodes in the user's historical interactions during the retrieval process, and sort the relevance of the knowledge fragments, including the following steps: (1) Knowledge Graph Index Mapping: Analyze the decision instruction set, extract the key entities of the knowledge graph association index (such as "Product Model_A123"), and match the corresponding storage location in the node index table of the distributed heterogeneous database; (2) Historical Interaction Optimization Retrieval: Call the high-frequency accessed knowledge nodes in the user's recent 100 interaction records (such as "Warranty Policy" is accessed 35 times), and set the weight coefficient +0.5 for the knowledge fragments containing these nodes; (3) Multi-dimensional Relevance Calculation: Sort the retrieval results according to the following priorities: Direct Matching Degree: The fragments that exactly match the user's query keywords (such as "Battery Life of A123") are placed at the top; Logical Relevance: The fragments indirectly associated through the knowledge graph relationship chain (such as "Matching Guide for A123 and Compatible Chargers") are sorted in descending order according to the number of associated edges; (4) Real-time Cache Loading: Pre-load the top 5 knowledge fragments into the memory cache area of the edge computing node, and control the response delay within 200ms.
[0051] In this embodiment, according to the sentiment tendency parameters in the decision instruction set, perform sentiment adaptability matching on the retrieved knowledge fragments, identify the words and expressions that need to be adjusted, and combine the preset user-preferred tone mode and response rhythm to perform modal particle embedding and sentence structure optimization to generate a response text that conforms to the current interaction sentiment state, including the following steps: (1) Sentiment Parameter Mapping and Correction Rules: If the sentiment tendency parameter is "Dissatisfaction_0.6": Replace "We suggest you" in the original text with "We will handle it for you immediately", insert the apologetic modal particle "Sorry" and the commitment phrase "Solve it within 24 hours"; If the sentiment tendency parameter is "Pleasure_0.8": Add the exclamation "Great!" at the beginning of the sentence, and optimize the declarative sentence "The operation is completed" to "It's been easily done for you!"; (2) User Preference Matching: Based on the "Preference for Concise Expression" feature marked in the user profile, delete the redundant modifiers in the knowledge fragments and split the complex sentences into short sentence sequences; (3) Rhythm Adjustment: Add segmented pause marks to the technical description paragraphs (such as parameter lists), and insert a 500ms silent interval every 3 items of data.
[0052] In this embodiment, according to the response text, corresponding voice feature parameters are marked for different semantic paragraphs, including the intonation rise and fall range, the speech rate adjustment ratio, and the stress distribution position, to generate emotional voice feature data that can drive the speech synthesis module, including the following steps: (1) Semantic segmentation and marking: Key conclusion paragraphs (such as "Final solution"): Mark the intonation rise range (the fundamental frequency is increased by 20 Hz), and the speech rate is reduced to 0.8 times the reference value; Warning content (such as "Do not disassemble the device"): Set the stress energy increase by 150% at the prohibition verb "Do not", and extend the syllable duration by 50 ms; (2) Emotional rhythm generation: For the text with the emotional tendency parameter "urgent_0.9", the global speech rate is increased to 1.2 times, and a 0.3-second rapid breathing sound effect is inserted every 5 seconds; For the text with "comfort_0.7", the intonation fluctuation range is compressed to ±5 Hz, and the stress distribution interval is expanded by 1.5 times to simulate a gentle tone; (3) Cross-modal verification: Detect whether there is a conflict between the voice feature data and the pronunciation syllable segmentation group of the lip movement matching unit. If there is a cross-modal timing deviation exceeding 100 ms, the text segmentation is preferentially reconstructed based on the voice features.
[0053] The beneficial effects of the above technical solution are: Through the combination of knowledge fragment search and emotional adaptability correction, the system can accurately retrieve and adjust the response content according to the user's historical interaction and current emotional state, achieving a more personalized and emotional response. The combination of emotional tendency parameters and the user's preferred tone pattern optimizes the tone and sentence pattern, making the response more friendly and natural. At the same time, the voice feature marking unit provides accurate voice parameter support for each semantic paragraph, ensuring that the speech synthesis module can generate voice feedback that conforms to the emotional state, thereby greatly enhancing the interaction experience and emotional resonance ability of the virtual digital human.
[0054] Embodiment 8: An intelligent real-time interactive Q&A system based on a virtual digital human provided by an embodiment of the present invention, a voice generation module, includes: Lip movement matching unit: According to the timestamp distribution of the feedback time series, the response text is segmented into corresponding pronunciation syllable segments, the voice features of each syllable segment are extracted and matched with the preset standard lip animation template of the virtual digital human, and key frames of lip animation synchronized with the actual pronunciation are generated; Micro-expression control unit: Based on the intensity and type of the emotional tendency parameter, select a corresponding level of micro-expression parameter combination from the preset expression library, including the movement amplitude of the eyebrows, the opening and closing frequency of the eyes, and the upward angle of the corners of the mouth, and adjust the micro-expression change rate according to the rhythm of the feedback time series to form a micro-expression parameter sequence matching the voice emotional expression; Body movement arrangement unit: Identify key semantic nodes in the answer text. If the speech features contain emphatic stress or long pauses, insert a preset body movement template at the corresponding moment, including the hand gesture swing trajectory and the torso tilt angle. The movement amplitude is scaled proportionally according to the emotional tendency parameters to obtain the body movement trajectory. Multimodal synchronization unit: performs timeline calibration on lip-sync animation key frames, micro-expression parameter sequences, and body movement trajectories to ensure that the start time of the voice response, the accent trigger point, and the peak time of the movement trajectory are aligned, generates a voice response, and pushes it to the user terminal after rendering and calculation at the edge computing node.
[0055] In this embodiment, the answer text is divided into corresponding pronunciation syllable segments according to the timestamp distribution of the feedback time series, including the following steps: performing phoneme-level analysis on the answer text, extracting the start and end timestamps of each syllable, and generating a pronunciation timing mapping table; matching the lip closure, tongue position and jaw opening parameters of each phoneme based on a standard lip animation template library to generate basic lip key frames; dynamically compressing or stretching the key frame interval duration according to the speech speed changes in the voice features to synchronize the lip animation with the spectral features of the real-time pronunciation; and interpolating and smoothing the transition frames of consecutive syllables to eliminate lip shape jumps.
[0056] In this embodiment, micro-expressions are regulated based on the intensity and type of the emotional tendency parameter, including the following steps: when the emotional tendency parameter is "anger_0.8", a frown template is called from the expression library, the eyebrow depression amplitude is set to 120% of the baseline value, and the eye opening and closing frequency is increased to 1.2 times / second; if the total duration of the feedback time series is compressed by 30%, the maintenance duration of the micro-expression is shortened synchronously, and the upward angle of the mouth corners is attenuated from 15° to 5° to achieve an emotional weakening transition; an instantaneous expression enhancement instruction is inserted at the stressed syllable of the speech feature, such as triggering a rapid eyelid closure action at the explosive consonant syllable.
[0057] In this embodiment, body movements are choreographed according to key semantic nodes, including the following steps: detecting logical conjunctions "therefore" and "however" in the answer text, and inserting an emphatic gesture of right hand abduction 30° at the corresponding timestamp; when the speech feature contains a pause of more than 300ms, triggering a listening posture template of torso leaning forward 10°, and the tilt duration is proportionally matched to the pause duration; if the emotional tendency parameter is "joy_0.7", the amplitude of the gesture swing trajectory is expanded to 1.5 times the baseline value, and a celebratory action of raising both hands is added at the end of the speech.
[0058] In this embodiment, multi-modal timeline calibration is performed, including the following steps: taking the starting moment of the voice response as the reference zero point, triggering the rendering of the first frame of the lip keyframe sequence with a delay of 50 ms to compensate for the device response delay; when the voice spectrum analysis shows the peak stress energy, synchronizing the peak of the eyebrow lift of the micro-expression parameter sequence and the highest point of the fingertips of the body movement; adopting a timestamp rollback mechanism at the edge computing node. If the deviation between the lip animation and the voice is detected to exceed 80 ms, reload the multi-modal data packet for the current 200 ms period.
[0059] The beneficial effects of the above technical solution are as follows: Through the collaborative work of lip matching, micro-expression regulation, body movement choreography, and multi-modal synchronization unit, the system achieves highly synchronous and natural expression of voice, expression, and movement. Lip matching ensures the precise correspondence between pronunciation and animation. Micro-expression regulation generates delicate expression changes based on emotional parameters. Body movement choreography drives dynamic movements through semantics and emotions, enhancing the authenticity of the interaction. The multi-modal synchronization unit ensures the real-time consistency of voice, expression, and movement through timeline calibration and edge rendering, significantly improving the interaction naturalness and emotional expression ability of virtual digital humans, and providing users with an immersive and personalized human-computer interaction experience.
[0060] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An intelligent real-time interactive Q&A system based on a virtual digital human, characterized in that, Including: Data acquisition module: Based on the integrated multi-modal sensor array, it receives the voice data and text data input by the user. At the same time, it real-time collects the sequence of dynamic facial expression parameters and the sequence of limb movement data of the user, and processes the voice data and text data to obtain the standardized voice feature vector and the structured text data; Cross-modal fusion module: Aligns the standardized voice feature vector, structured text data with the sequence of dynamic facial expression parameters and the sequence of limb movement data of the user in space-time, and constructs an interaction feature matrix containing multi-modal time-series correlation features; Behavior decision module: Determines a decision instruction set including emotional tendency parameters, knowledge graph association indexes and feedback time series based on the historical interaction data and the interaction feature matrix; Knowledge retrieval module: According to the knowledge graph association index in the decision instruction set, retrieves the corresponding set of knowledge fragments from the preset distributed heterogeneous database, embeds tone words and annotates the prosody rhythm for the vocabulary corresponding to the knowledge fragments in the set of knowledge fragments according to the emotional tendency parameters, and generates the response text and voice features with emotional adaptability; Voice generation module: Generates the key frames of the lip animation, the sequence of micro-expression parameters and the limb movement trajectory of the virtual digital human and generates a voice response based on the feedback time series, response text and voice features in the decision instruction set, and pushes them to the user terminal through the edge computing node.
2. The intelligent real-time interactive Q&A system based on a virtual digital human according to claim 1, wherein, Data acquisition module, including: Voice processing unit: Conducts frequency-domain energy detection on the received voice data, identifies and eliminates the invalid voice segments with energy lower than the preset threshold, and obtains the continuous valid voice waveform data; Text analysis unit: Conducts dependency syntactic analysis on the text data, annotates the entity types and action predicates, and generates the initial text data with semantic labels; Environment analysis unit: Collects the three-dimensional space information of the interaction environment, tracks the movement trajectory of the target user in the interaction space, and calculates the spatial position coordinates and orientation angle data of the user; Image acquisition unit: Dynamically adjusts the focal length of the visual sensor and the directivity parameters of the pickup array according to the spatial position coordinates and orientation angle data, and obtains the facial expression image and limb movement image data of the user; Image analysis unit: Analyzes the spatial offset of the key feature points in the facial expression image, measures the deformation amplitude and duration of a specific facial area, and generates a sequence of dynamic facial expression parameters; Motion analysis unit: Extracts the spatial coordinates of the key bone nodes in the limb movement image data, determines the relative movement trajectory between the nodes, and generates the sequence of limb movement data; Data processing unit: Processes the continuous valid voice waveform data and the initial text data to obtain the standardized voice feature vector and the structured text data.
3. An intelligent real-time interactive Q&A system based on a virtual digital human according to claim 2, characterized in that, Data processing unit, including: Voice data processing subunit: Preprocesses the continuous valid voice waveform data by using the adaptive noise cancellation algorithm, extracts the acoustic features, and generates the standardized voice feature vector containing dynamic time warping information; Text data processing subunit: Conducts word segmentation processing on the initial text data through the domain-adaptive semantic word segmentation model, and combines the preset intention recognition knowledge base for intention annotation to generate the structured text data with context semantic association.
4. An intelligent real-time interactive Q&A system based on a virtual digital human according to claim 1, characterized in that, Cross-modal fusion module, including: Data alignment unit: Perform spatio-temporal alignment on the normalized speech feature vectors, structured text data, sequence of user's facial expression dynamic parameters, and sequence of body movement data. Weight adjustment unit: Adjust the weights of the normalized speech feature vectors, structured text data, sequence of facial expression dynamic parameters, and sequence of body movement data after spatio-temporal alignment. Matrix construction unit: Integrate the normalized speech feature vectors, structured text data, sequence of user's facial expression dynamic parameters, sequence of body movement data, and the corresponding adjusted weights to construct an interaction feature matrix containing multi-modal temporal correlation features.
5. An intelligent real-time interactive Q&A system based on a virtual digital human according to claim 4, characterized in that, Weight adjustment unit, including: Coupling sub-unit: Perform coupling analysis on the aligned normalized speech feature vectors, sequence of facial expression dynamic parameters, and sequence of body movement data to obtain speech-expression coupling coefficients and speech-action coupling coefficients. First adjustment sub-unit: Perform a first adjustment on the weight distribution ratios of the aligned normalized speech feature vectors, sequence of facial expression dynamic parameters, and sequence of body movement data according to the speech-expression coupling coefficients and speech-action coupling coefficients. Parameter acquisition sub-unit: Perform micro-expression analysis on the sequence of user's facial expression dynamic parameters and sequence of body movement data. Meanwhile, perform voice fundamental frequency jitter detection on the aligned normalized speech feature vectors to determine the user's physiological state parameters, where the user's physiological state parameters include emotional tension index and attention concentration score. Second adjustment sub-unit: Perform a second adjustment on the weights of the aligned normalized speech feature vectors, sequence of facial expression dynamic parameters, and sequence of body movement data based on the user's physiological state parameters. Third adjustment sub-unit: Perform a third adjustment on the preset initial weight of the structured text data based on the weights of the normalized speech feature vectors, sequence of facial expression dynamic parameters, and sequence of body movement data after the second adjustment.
6. The intelligent real-time interactive Q&A system based on a virtual digital human according to claim 1, wherein, Behavior decision module, including: Historical data weight adjustment unit: Determine the multi-modal weight distribution of the user's interaction behavior based on historical interaction data and the interaction feature matrix, obtain the short-term interaction frequency, multi-modal signal volatility, and emotional tendency deviation degree, perform multi-modal confidence analysis and interaction habit matching to obtain the voice weight correction coefficient, text weight correction coefficient, and action weight correction coefficient. Cross-modal correlation network construction unit: Monitor the signal synchronization between the normalized speech feature vectors, structured text data, and sequence of facial expression dynamic parameters during the interaction process, determine the speech-text synchronization rate and speech-expression synchronization rate, and generate a multi-modal synchronization relationship network. Feedback decision splitting prediction unit: Perform semantic parsing on the structured text data, extract the interaction intention label and combine it with the emotional tendency deviation degree to determine the feedback level division ratio and feedback rhythm adjustment coefficient. Associated multi-modal anomaly analysis unit: Based on the multi-modal synchronization relationship network, use the feedback level division ratio and feedback rhythm adjustment coefficient to perform cross-modal consistency verification, determine the text-voice anomaly coefficient and expression-action anomaly coefficient, and combine the voice weight correction coefficient, text weight correction coefficient, and action weight correction coefficient to generate a composite anomaly index. Instruction Integration Output Unit: Generate a decision instruction set according to the composite exception index and a preset index-instruction database. The decision instruction set includes sentiment tendency parameters, knowledge graph association indexes, and feedback time series.
7. An intelligent real-time interactive Q&A system based on a virtual digital human according to claim 1, characterized in that, Knowledge Retrieval Module, including: Knowledge Fragment Search Unit: According to the knowledge graph association indexes in the decision instruction set, traverse the preset distributed heterogeneous database, identify the set of knowledge fragments matching the current interaction intention, preferentially match the high-frequency accessed knowledge nodes in the user's historical interactions during the retrieval process, and sort the relevance of the knowledge fragments; Emotional Adaptability Correction Unit: Based on the sentiment tendency parameters in the decision instruction set, perform emotional adaptability matching on the retrieved knowledge fragments, identify the words and expressions that need to be adjusted, and combine the preset tone pattern and response rhythm preferred by the user to perform modal particle embedding and sentence structure optimization to generate an answer text that conforms to the current interaction emotional state; Voice Feature Annotation Unit: According to the answer text, annotate the corresponding voice feature parameters for different semantic paragraphs, including the intonation rise and fall range, speech rate adjustment ratio, and stress distribution position, to generate voice features that can drive the voice synthesis module.
8. An intelligent real-time interactive Q&A system based on a virtual digital human according to claim 1, characterized in that, Voice Generation Module, including: Lip Sync Unit: According to the timestamp distribution of the feedback time series, segment the answer text into corresponding pronunciation syllable segments, extract the voice features of each syllable segment and match the preset standard lip animation template of the virtual digital human to generate the key frames of the lip animation synchronized with the actual pronunciation; Micro-expression Regulation Unit: Based on the intensity and type of the sentiment tendency parameters, select the corresponding level of micro-expression parameter combinations from the preset expression library, including the movement amplitude of the eyebrows, the opening and closing frequency of the eyes, and the upward angle of the corners of the mouth, and adjust the micro-expression change rate according to the rhythm of the feedback time series to form a micro-expression parameter sequence matching the voice emotional expression; Limb Movement Choreography Unit: Identify the key semantic nodes in the answer text. If the voice features include emphasized stress or long pauses, insert the preset limb movement templates at the corresponding moments, including the gesture swing trajectory and the torso tilt angle, and scale the movement amplitude proportionally according to the sentiment tendency parameters to obtain the limb movement trajectory; Multi-modal Synchronization Unit: Perform time-axis calibration on the key frames of the lip animation, the micro-expression parameter sequence, and the limb movement trajectory to ensure that the starting moment of the voice response, the stress trigger point, and the peak moment of the movement trajectory are aligned, generate the voice response, and push it to the user terminal after rendering calculations at the edge computing node.
Citation Information
Patent Citations
Virtual person-based multi-mode interactive processing method and system
CN107765852A
Virtual doctor system based on knowledge graph driving and operation method thereof
CN117056536A
Multi-modal digital employee reception system and method based on edge calculation
CN119693750A
Digital human interaction method and system based on naked eye 3D visualization
CN119782466A
Virtual digital human interaction system based on AI
CN119902625A
Cited By
Method and system for realizing real-time two-way visual interaction of digital human by calling camera
CN120523334A
Method and system for realizing real-time two-way visual interaction of digital human by calling camera
CN120523334B
Multi-mode fusion intelligent inductive switch control system
CN120568555A
Multi-modal fusion intelligent sensing switch control system
CN120568555B
Teaching interaction-oriented digital human three-dimensional reconstruction system
CN120599155A