An intelligent real-time interactive question and answer system based on a virtual digital person
By integrating multimodal sensors with a virtual digital human behavior decision-making model, a real-time interactive question-and-answer system based on voice, text, facial expressions, and body movements was realized. This solves the problem of limited interactive experience in existing technologies, improves the accuracy of emotion recognition and response, and ensures the naturalness and adaptability of feedback.
Patent Information
- Application Number
- CN202510779127.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-06-12
AI Technical Summary
Existing virtual digital human systems fail to fully consider users' non-verbal information such as body movements and facial expressions, resulting in a relatively simple and mechanical interactive experience, lacking deep integration of multimodal information and intelligent adjustment for real-time response.
It integrates multimodal sensors and a virtual digital human behavior decision-making model. The data acquisition module collects voice, text, facial expressions and body movements data in real time. The cross-modal fusion module performs spatiotemporal alignment to generate an interactive feature matrix. Combined with the knowledge retrieval module and the voice generation module, it realizes emotionally adapted interactive question answering.
It improves the interactive intelligence and emotional adaptability of virtual digital humans, ensuring that feedback is more natural and meets user needs, and significantly enhances the interactive experience.
Smart Images

Figure CN120318388B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech signal processing and speech recognition, and in particular to an intelligent real-time interactive question-answering system based on a virtual digital human. Background Art
[0002] With the rapid development of artificial intelligence technology, virtual humans, as a new type of human-computer interaction interface, have been widely used in various fields such as customer service, education, and entertainment. Traditional virtual human systems rely heavily on speech recognition and text processing. However, existing systems often fail to fully consider non-verbal information such as user body movements and facial expressions, and fail to dynamically adjust the virtual human's behavior and responses based on the user's emotional state and interaction history, resulting in a relatively monotonous and mechanical interactive experience.
[0003] Existing virtual human technologies often focus on single voice or text input and rely on pre-set rules for responses, lacking the deep integration of multimodal information and intelligent adjustment of real-time responses. This approach not only results in poor interactive effects and emotional adaptability of virtual humans, but also in their inability to make precise behavioral decisions based on specific situations, impacting the user's interactive experience.
[0004] Therefore, the present invention provides an intelligent real-time interactive question-answering system based on virtual digital humans. Summary of the Invention
[0005] The present invention provides an intelligent real-time interactive question-and-answer system based on a virtual digital human to address the deficiencies in the prior art, particularly to improve the interactive intelligence, emotional adaptability, and real-time response of the system. By integrating multimodal sensors with a virtual digital human behavioral decision-making model, a more intelligent, real-time, and emotionally adaptive interactive question-and-answer system is realized. It can simultaneously process voice, text, facial expressions, and body movements, dynamically adjust the response strategy of the virtual digital human, and generate an accurate interactive feature matrix through spatiotemporal aligned cross-modal fusion, thereby improving the accuracy of emotion recognition and response. The knowledge retrieval module optimizes the answer content based on the user's emotional tendencies, and the voice generation module synchronizes lip shape, micro-expressions, and body movements, ensuring that the virtual digital human's feedback is more natural and meets user needs. The overall solution significantly improves the interactive experience and adaptability of the virtual digital human.
[0006] The present invention provides an intelligent real-time interactive question-answering system based on a virtual digital human, comprising:
[0007] Data acquisition module: Receives user input voice and text data based on an integrated multimodal sensor array. It also collects the user's facial expression and body movement sequence data in real time, and processes the voice and text data to obtain standardized voice feature vectors and structured text data.
[0008] Cross-modal fusion module: This module performs spatiotemporal alignment of standardized speech feature vectors and structured text data with the user's facial expression dynamic parameter sequence and body movement sequence data, and constructs an interactive feature matrix containing multimodal temporal correlation features.
[0009] Behavior decision module: Determines a decision instruction set containing sentiment tendency parameters, knowledge graph association index, and feedback time series based on historical interaction data and interaction feature matrix;
[0010] Knowledge retrieval module: Based on the knowledge graph association index in the decision instruction set, the corresponding knowledge fragment set is retrieved from the preset distributed heterogeneous database. Based on the sentiment tendency parameters, the corresponding vocabulary of the knowledge fragments in the knowledge fragment set is embedded with modal words and rhythmic rhythm annotation, generating emotionally adaptive answer text and voice features;
[0011] Speech generation module: Based on the feedback time series, answer text and speech features in the decision instruction set, it generates the virtual digital human's lip animation key frames, micro-expression parameter sequence and body movement trajectory, and generates speech responses, which are pushed to the user terminal through the edge computing node.
[0012] Preferably, the data acquisition module includes:
[0013] Speech processing unit: performs frequency domain energy detection on the received speech data, identifies and removes invalid speech segments whose energy is lower than a preset threshold, and obtains continuous valid speech waveform data;
[0014] Text analysis unit: performs dependency syntax analysis on text data, annotates entity types and action predicates, and generates initial text data with semantic labels;
[0015] Environmental analysis unit: collects three-dimensional spatial information of the interactive environment, tracks the movement trajectory of the target user in the interactive space, and calculates the user's spatial position coordinates and orientation angle data;
[0016] Image acquisition unit: Dynamically adjusts the focal length of the visual sensor and the directional parameters of the sound pickup array based on the spatial position coordinates and orientation angle data to obtain the user's facial expression image and body movement image data;
[0017] Image analysis unit: Analyzes the spatial offset of key feature points in facial expression images, measures the deformation amplitude and duration of specific facial areas, and generates a sequence of expression dynamic parameters;
[0018] Motion analysis unit: extracts the spatial coordinates of key skeletal nodes from limb motion image data, determines the relative movement trajectories between nodes, and generates limb motion sequence data;
[0019] Data processing unit: processes continuous valid speech waveform data and initial text data to obtain standardized speech feature vectors and structured text data.
[0020] Preferably, the data processing unit includes:
[0021] Speech data processing subunit: Uses an adaptive noise cancellation algorithm to preprocess continuous valid speech waveform data, extract acoustic features, and generate standardized speech feature vectors containing dynamic time warping information;
[0022] Text data processing subunit: The initial text data is segmented through a domain-adaptive semantic segmentation model, and intent is annotated in combination with a preset intent recognition knowledge base to generate structured text data with contextual semantic associations.
[0023] Preferably, the cross-modal fusion module includes:
[0024] Data alignment unit: aligns the standardized speech feature vector, structured text data, user's expression dynamic parameter sequence and body movement sequence data in time and space;
[0025] Weight adjustment unit: adjusts the weights of the standardized speech feature vector, structured text data, expression dynamic parameter sequence and body movement sequence data after spatiotemporal alignment;
[0026] Matrix construction unit: Integrate standardized speech feature vectors, structured text data, user's expression dynamic parameter sequence and body movement sequence data as well as the adjusted corresponding weights to construct an interactive feature matrix containing multimodal temporal correlation features.
[0027] Preferably, the weight adjustment unit includes:
[0028] Coupling subunit: performs coupling analysis on the aligned standardized speech feature vectors, expression dynamic parameter sequences, and body movement sequence data to obtain the speech-expression coupling coefficient and the speech-movement coupling coefficient;
[0029] A first adjustment subunit: performing a first adjustment on the weight distribution ratio of the aligned standardized speech feature vector, the expression dynamic parameter sequence, and the body movement sequence data according to the speech-expression coupling coefficient and the speech-action coupling coefficient;
[0030] Parameter acquisition subunit: This unit performs micro-expression analysis on the user's facial expression dynamic parameter sequence and body movement sequence data. It also performs voice fundamental frequency jitter detection on the aligned standardized speech feature vectors to determine the user's physiological state parameters, including the emotional tension index and attention concentration score.
[0031] A second adjustment subunit: performing a second adjustment on the weights of the aligned standardized speech feature vector, expression dynamic parameter sequence, and body movement sequence data based on the user's physiological state parameters;
[0032] The third adjustment subunit: performs a third adjustment on the preset initial weights of the structured text data based on the weights of the standardized speech feature vector, expression dynamic parameter sequence and body movement sequence data after the second adjustment.
[0033] Preferably, the behavior decision module includes: a historical data weight adjustment unit: based on historical interaction data and an interaction feature matrix, determining the multimodal weight distribution of the user's interaction behavior, obtaining the short-term interaction frequency, multimodal signal volatility, and emotional tendency deviation, performing multimodal confidence analysis and interaction habit matching, and obtaining a voice weight correction coefficient, a text weight correction coefficient, and an action weight correction coefficient;
[0034] Cross-modal association network construction unit: monitors the signal synchronization between standardized speech feature vectors, structured text data and expression dynamic parameter sequences during the interaction process, determines the speech-text synchronization rate and speech-expression synchronization rate, and generates a multimodal synchronization relationship network;
[0035] Feedback decision splitting and prediction unit: This unit performs semantic analysis based on structured text data, extracts interaction intent labels, and combines them with the sentiment bias to determine the feedback level division ratio and feedback rhythm adjustment coefficient.
[0036] Correlated multimodal anomaly analysis unit: Based on the multimodal synchronization relationship network, it uses the feedback level division ratio and feedback rhythm adjustment coefficient to perform cross-modal consistency verification, determine the text-speech anomaly coefficient and the expression-action anomaly coefficient, and combine the speech weight correction coefficient, text weight correction coefficient, and action weight correction coefficient to generate a composite anomaly index;
[0037] Instruction integration output unit: Generates a decision instruction set based on the composite anomaly index and the preset index-instruction database, where the decision instruction set includes sentiment tendency parameters, knowledge graph association index and feedback time series.
[0038] Preferably, the knowledge retrieval module includes:
[0039] Knowledge fragment search unit: Based on the knowledge graph association index in the decision instruction set, it traverses the preset distributed heterogeneous database to identify the set of knowledge fragments that match the current interaction intention. During the retrieval process, it prioritizes frequently accessed knowledge nodes in the user's historical interactions and sorts the relevance of the knowledge fragments.
[0040] Emotional adaptability correction unit: Based on the emotional tendency parameters in the decision instruction set, it performs emotional adaptability matching on the retrieved knowledge fragments, identifies the vocabulary and expressions that need to be adjusted, and combines the preset user-preferred tone pattern and response rhythm to perform modal word embedding and sentence structure optimization to generate a response text that conforms to the current interactive emotional state;
[0041] Speech feature annotation unit: Based on the answer text, it annotates the corresponding speech feature parameters for different semantic paragraphs, including the intonation rise and fall range, speech speed adjustment ratio and stress distribution position, to generate speech features that can drive the speech synthesis module.
[0042] Preferably, the speech generation module includes:
[0043] Lip-sync matching unit: This unit divides the response text into corresponding pronunciation syllable segments based on the timestamp distribution of the feedback time series. It extracts the speech features of each syllable segment and matches them with the preset standard lip-sync animation template of the virtual digital human to generate lip-sync animation keyframes synchronized with the actual pronunciation.
[0044] Micro-expression control unit: Based on the intensity and type of emotional tendency parameters, it selects a corresponding level of micro-expression parameter combination from the preset expression library, including the amplitude of eyebrow movement, the frequency of eye opening and closing, and the angle of mouth corner upward movement. It then adjusts the rate of micro-expression change according to the rhythm of the feedback time series to form a micro-expression parameter sequence that matches the emotional expression of the voice;
[0045] The body movement choreography unit identifies key semantic nodes in the response text. If the speech features contain emphatic accents or long pauses, a preset body movement template is inserted at the corresponding moment, including the hand gesture trajectory and torso tilt angle. The movement amplitude is proportionally scaled according to the emotional tendency parameter to obtain the body movement trajectory.
[0046] Multimodal synchronization unit: This unit performs timeline calibration on lip-sync animation keyframes, micro-expression parameter sequences, and body movement trajectories to ensure that the starting moment and accent trigger point of the voice response are aligned with the peak moment of the movement trajectory. The voice response is generated and pushed to the user terminal after rendering and calculation at the edge computing node.
[0047] Compared with the prior art, the present invention has the following advantages:
[0048] By integrating multimodal sensors with a virtual human behavioral decision-making model, a more intelligent, real-time, and emotionally adaptive interactive question-and-answer system has been implemented. This system can simultaneously process speech, text, facial expressions, and body movements, dynamically adjust the virtual human's response strategy, and generate a precise interaction feature matrix through spatiotemporal alignment of cross-modal fusion, improving the accuracy of emotion recognition and response. The knowledge retrieval module optimizes responses based on user emotional tendencies, while the speech generation module synchronizes lip movements, micro-expressions, and body movements, ensuring that the virtual human's feedback is more natural and meets user needs. This overall solution significantly improves the interactive experience and adaptability of the virtual human, making it suitable for a variety of application scenarios, including intelligent customer service and virtual assistants. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0050] Figure 1 This is a structural diagram of an intelligent real-time interactive question-answering system based on a virtual digital human provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0051] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0052] Example 1:
[0053] The embodiment of the present invention provides an intelligent real-time interactive question-answering system based on virtual digital human. Figure 1 Shown, including:
[0054] Data acquisition module: Receives user input voice and text data based on an integrated multimodal sensor array. It also collects the user's facial expression and body movement sequence data in real time, and processes the voice and text data to obtain standardized voice feature vectors and structured text data.
[0055] Cross-modal fusion module: This module performs spatiotemporal alignment of standardized speech feature vectors and structured text data with the user's facial expression dynamic parameter sequence and body movement sequence data, and constructs an interactive feature matrix containing multimodal temporal correlation features.
[0056] Behavior decision module: Determines a decision instruction set containing sentiment tendency parameters, knowledge graph association index, and feedback time series based on historical interaction data and interaction feature matrix;
[0057] Knowledge retrieval module: Based on the knowledge graph association index in the decision instruction set, the corresponding knowledge fragment set is retrieved from the preset distributed heterogeneous database. Based on the sentiment tendency parameters, the corresponding vocabulary of the knowledge fragments in the knowledge fragment set is embedded with modal words and rhythmic rhythm annotation, generating emotionally adaptive answer text and voice features;
[0058] Speech generation module: Based on the feedback time series, answer text and speech features in the decision instruction set, it generates the virtual digital human's lip animation key frames, micro-expression parameter sequence and body movement trajectory, and generates speech responses, which are pushed to the user terminal through the edge computing node.
[0059] In this embodiment, the knowledge graph association index refers to a semantic search pointer system built based on the domain knowledge graph. Specifically, it includes: 1) Node encoding rules: Using hash encoding (SHA-3 algorithm), knowledge nodes (such as diseases, drugs, and procedures) are mapped to 128-bit unique identifiers, with the first 32 bits representing the domain classification (medicine / finance / education); 2) Relationship weight calculation: Using the TransE model to learn embedding vectors for inter-entity relationships (such as "treatment" and "contraindications"), when a user queries for "diabetes complications," adjacent nodes with a relationship weight greater than 0.85 are prioritized; 3) Dynamic path optimization: Using the A* algorithm, the optimal search path of up to three hops is generated within the knowledge graph, taking into account the frequency of terms used in historical user interactions. This index, unlike traditional database primary keys, enables cross-modal queries (e.g., directly linking a voice medical consultation to an image of a drug's molecular formula), with response latency kept to under 50ms.
[0060] In this embodiment, the distributed heterogeneous database refers to a hybrid architecture storage system designed specifically for multimodal knowledge storage. Its core technical features include: 1) a tiered storage strategy: structured data (drug instructions) is stored in a PostgreSQL cluster, while unstructured data (CT images) is placed in an IPFS distributed file system, with a unified query interface implemented through a GraphQL gateway; 2) a semantic caching mechanism: a BERT-based semantic similarity matching cache is deployed at edge nodes. When a user asks consecutive questions about "angina symptoms" and "myocardial ischemia symptoms," cached answer blocks with a semantic similarity greater than 92% are directly returned; 3) a cross-domain consistency protocol: an improved RAFT algorithm is used to coordinate data synchronization across multiple data centers, ensuring strong consistency (CPA model) for financial data and eventual consistency (BASE model) for entertainment data. The system supports 100,000 concurrent queries per second, with data sharding down to a single medical concept (e.g., "aspirin" corresponds to three data shards).
[0061] The beneficial effects of the above technical solution are: by integrating multimodal sensors with the virtual digital human behavioral decision-making model, a more intelligent, real-time, and emotionally adaptive interactive question-and-answer system is realized. It can simultaneously process voice, text, facial expressions, and body movements, dynamically adjust the virtual digital human's response strategy, and generate a precise interaction feature matrix through spatiotemporal alignment of cross-modal fusion, thereby improving the accuracy of emotion recognition and response. The knowledge retrieval module optimizes the content of the answer based on the user's emotional tendencies, and the speech generation module synchronizes lip shape, micro-expressions, and body movements, ensuring that the virtual digital human's feedback is more natural and meets user needs. The overall solution significantly improves the interactive experience and adaptability of the virtual digital human, and is suitable for a variety of application scenarios such as intelligent customer service and virtual assistants.
[0062] Example 2:
[0063] The embodiment of the present invention provides an intelligent real-time interactive question-answering system based on a virtual digital human, and a data acquisition module, including:
[0064] Speech processing unit: performs frequency domain energy detection on the received speech data, identifies and removes invalid speech segments whose energy is lower than a preset threshold, and obtains continuous valid speech waveform data;
[0065] Text analysis unit: performs dependency syntax analysis on text data, annotates entity types and action predicates, and generates initial text data with semantic labels;
[0066] Environmental analysis unit: collects three-dimensional spatial information of the interactive environment, tracks the movement trajectory of the target user in the interactive space, and calculates the user's spatial position coordinates and orientation angle data;
[0067] Image acquisition unit: Dynamically adjusts the focal length of the visual sensor and the directional parameters of the sound pickup array based on the spatial position coordinates and orientation angle data to obtain the user's facial expression image and body movement image data;
[0068] Image analysis unit: Analyzes the spatial offset of key feature points in facial expression images, measures the deformation amplitude and duration of specific facial areas, and generates a sequence of expression dynamic parameters;
[0069] Motion analysis unit: extracts the spatial coordinates of key skeletal nodes from limb motion image data, determines the relative movement trajectories between nodes, and generates limb motion sequence data;
[0070] Data processing unit: processes continuous valid speech waveform data and initial text data to obtain standardized speech feature vectors and structured text data.
[0071] In this embodiment, frequency domain energy detection is performed on the received speech data, invalid speech segments with energy below a preset threshold are identified and eliminated, continuous valid speech waveform data is obtained, and its spectrum peak distribution parameters are calculated to generate a standardized speech feature vector. Specifically, the method includes: using a spectrum analysis method based on short-time Fourier transform to divide the speech signal into frames with a frame length of 20ms, and calculating the total energy value of each frame signal within the effective frequency band of 100-4000Hz; setting a dynamic energy threshold, and when the energy of a frame is lower than the threshold and the duration exceeds 300ms, it is judged as an invalid speech segment and eliminated; for the retained valid speech segments, the 12-dimensional features of their Mel-frequency cepstral coefficients (MFCCs) are extracted, and the position distribution and amplitude ratio of the spectrum peaks of each frame are statistically analyzed to generate a 128-dimensional standardized speech feature vector containing frequency domain energy distribution and time domain continuity features; simultaneously detecting silence segments and non-human noise in the speech, and combining a dual-threshold endpoint detection algorithm to optimize the accuracy of determining the start and end points of the speech. The speech signal refers to the continuous audio waveform data containing human speech information collected by the user during the interaction with the virtual digital human. Its frequency domain range is usually 80Hz-8kHz (covering the fundamental frequency and main resonance peaks of spoken language), and its time domain representation is an analog or digital signal with amplitude changing over time.
[0072] In this embodiment, dependency syntactic analysis is performed on text data, entity types and action predicates are annotated, and initial structured text data with semantic labels is generated. Specifically, this includes: using a neural dependency syntactic analysis model based on an attention mechanism to identify core predicates and their dependencies in the input text and construct a syntactic dependency tree; using named entity recognition technology to annotate entity categories such as names of people, places, and institutions in sentences, and expanding the professional terminology dictionary in combination with the domain knowledge base; establishing a semantic role framework for action predicates, annotating semantic roles such as agent, patient, time, and place, and their constraints; and finally generating structured text data, which contains four layers of information: original text, dependency graph, entity annotation results, and semantic role framework. Each layer of information is associated through a unified identifier to support subsequent intent recognition and knowledge retrieval module calls.
[0073] In this embodiment, three-dimensional spatial information of the interactive environment is collected, the movement trajectory of the target user in the interactive space is tracked, the user's spatial position coordinates and orientation angle data are calculated, and the visual sensor parameters are dynamically adjusted, specifically including: deploying a binocular stereo vision camera array to reconstruct the three-dimensional point cloud data of the scene in real time based on the parallax principle; using an improved Kalman filter algorithm to fuse RGB-D sensors and UWB positioning data to track the spatial coordinates of key points on the user's head, with a position update frequency of 30Hz; calculating the orientation angle by fitting the spatial vector of the user's shoulder line, with an accuracy controlled within ±3°; dynamically adjusting the camera focal length according to the user's distance, keeping the face area occupying 40%-60% of the screen area within the range of 1.5-5 meters; and controlling the main lobe of the beamforming microphone array to point to the user's mouth area, suppressing ambient noise by more than 15dB, and ensuring the quality of voice collection.
[0074] In this embodiment, the spatial offsets of key feature points in facial expression images are analyzed, the deformation amplitude and duration of specific facial regions are measured, and a sequence of expression dynamic parameters is generated. Simultaneously, the spatial coordinates of key skeletal nodes in limb motion images are extracted to determine the relative movement trajectories between these nodes. Specifically, this includes: using a deep learning-based facial feature point detection algorithm to track the two-dimensional coordinates of 68 key facial points in real time, calculating the displacement vectors of expression muscle control areas such as the glabellar and zygomatic major muscles; establishing a normalized Facial Action Coding System (FACS) to map the displacement vectors to activation intensity and time curves of AUs (Action Units); for limb analysis, the three-dimensional coordinates of 25 skeletal nodes of the human body are extracted using a lightweight OpenPose model to construct a spatiotemporal graph model of limb motion; calculating kinematic parameters such as joint angle change rate and end-point trajectory curvature, and combining them with a preset action template library to identify typical interactive actions such as waving and nodding, generating an action description sequence containing spatiotemporal features and semantic labels.
[0075] In this embodiment, speech data and text data are processed to obtain standardized speech feature vectors and structured text data, and then spatiotemporally aligned and feature fused with other perceptual data. Specifically, the dynamic time warping (DTW) algorithm is used to align the time axis of the speech feature sequence and the facial expression parameter sequence, ensuring that the time deviation between the modalities is less than 80ms; a cross-modal feature fusion model based on a graph neural network is constructed to encode the MFCC features of speech, the dependency graph of text, the activation pattern of facial AUs, and the spatiotemporal graph model of the body into a unified 256-dimensional joint feature vector; the contribution weight of each modal feature is calculated through the attention mechanism, and a weight coefficient of 0.6 is assigned to facial expression features in the speech emotion recognition task, and the weight of text semantic features is increased to 0.8 in the intention understanding task; the final output is a multi-level data structure containing the original data, the temporal alignment results, and the fused features, providing complete environmental perception input for subsequent interactive decision-making.
[0076] The beneficial effects of this technical solution include: through multi-dimensional data collection and processing, comprehensive acquisition and intelligent analysis of user voice, text, facial expressions, body movements, and interactive environment information are achieved. Voice processing and vector generation improve the effectiveness of voice input and the accuracy of feature extraction; text analysis enhances semantic understanding; environmental analysis and image acquisition achieve precise perception of user status; and image and motion analysis provides dynamic parameter input for expressions and movements. The overall solution provides a comprehensive, dynamic, and accurate interactive data foundation for virtual digital humans, significantly improving the system's intelligent responsiveness and the naturalness of human-computer interaction.
[0077] Example 3:
[0078] The embodiment of the present invention provides an intelligent real-time interactive question-answering system based on a virtual digital human, wherein the data processing unit includes:
[0079] Speech data processing subunit: Uses an adaptive noise cancellation algorithm to preprocess continuous valid speech waveform data, extract acoustic features, and generate standardized speech feature vectors containing dynamic time warping information;
[0080] Text data processing subunit: The initial text data is segmented through a domain-adaptive semantic segmentation model, and intent is annotated in combination with a preset intent recognition knowledge base to generate structured text data with contextual semantic associations.
[0081] In this embodiment, an adaptive noise cancellation algorithm is used to preprocess voice requests, extract acoustic features, and generate a standardized voice feature vector containing dynamic time warping information. The following steps are performed: an adaptive filtering algorithm based on frequency domain least mean squares (FLMS) is used to estimate the power spectral density of background noise in real time, construct a multi-level noise reference model, and achieve dynamic noise reduction in the 20-8000Hz frequency band, with a signal-to-noise ratio improvement of ≥12dB; 26-dimensional acoustic features are extracted from the noise-reduced voice signal, including 12-dimensional Mel-frequency cepstral coefficients (MFCCs), 1-dimensional fundamental frequency (F0), 6-dimensional chromaticity features, and 7-dimensional spectral centroid features, to form a time-varying feature matrix; a dynamic time warping (DTW) algorithm is used to align voice segments of different lengths, calculate a distance metric matrix for frame-level features, and generate a 128-dimensional standardized voice feature vector through Gaussian normalization; and emotional intonation features in the voice are simultaneously detected. Emotional states such as excitement, anger, or calmness are marked based on the fundamental frequency trajectory change rate (ΔF0 / Δt) and energy envelope fluctuations to form an enhanced voice feature descriptor.
[0082] In this embodiment, a domain-adaptive semantic segmentation model is used to segment text requests, and intents are annotated in combination with a preset intent recognition knowledge base to generate structured text data with contextual semantic associations. This includes: adopting a hybrid segmentation strategy that integrates a BiLSTM-CRF neural network with a domain dictionary matching method. For example, based on general Chinese word segmentation (accuracy ≥ 98%), professional term segmentation is optimized for intelligent customer service scenarios (e.g., "5G package" cannot be split); constructing a hierarchical intent recognition knowledge base. The first layer uses a rule engine to match high-frequency intent templates (e.g., "query traffic" is associated with the "business processing" category); the second layer uses the BERT model to calculate the semantic similarity between user queries and knowledge base questions, and performs fine-grained intent classification on requests that do not match the templates. The output structured text data includes: 1) the word segmentation sequence and part-of-speech tagging of the original text, 2) the predicate-argument structure generated by dependency syntactic analysis, 3) the intent category and confidence score, and 4) the associated contextual conversation state (e.g., the order number involved in the previous round of conversation), supporting multi-round semantic inheritance and ambiguity resolution.
[0083] The beneficial effects of this technical solution are as follows: The voice data processing subunit introduces an adaptive noise cancellation algorithm, effectively improving voice input clarity and feature extraction accuracy in complex environments. Simultaneously, the text data processing subunit, combined with a domain-adaptive semantic segmentation model and an intent recognition knowledge base, accurately identifies user intent and generates semantically complete structured text data. The overall solution achieves high-quality preprocessing and semantic enhancement of voice and text data, providing an accurate and efficient input foundation for subsequent cross-modal fusion and behavioral decision-making, significantly enhancing the interactive understanding and response quality of virtual digital humans.
[0084] Example 4:
[0085] The embodiment of the present invention provides an intelligent real-time interactive question-answering system based on a virtual digital human, and a cross-modal fusion module, including:
[0086] Data alignment unit: aligns the standardized speech feature vector, structured text data, user's expression dynamic parameter sequence and body movement sequence data in time and space;
[0087] Weight adjustment unit: adjusts the weights of the standardized speech feature vector, structured text data, expression dynamic parameter sequence and body movement sequence data after spatiotemporal alignment;
[0088] Matrix construction unit: Integrate standardized speech feature vectors, structured text data, user's expression dynamic parameter sequence and body movement sequence data as well as the adjusted corresponding weights to construct an interactive feature matrix containing multimodal temporal correlation features.
[0089] In this embodiment, standardized speech feature vectors, structured text data, expression dynamic parameter sequences, and body movement sequence data are spatiotemporally aligned. This involves employing a multimodal synchronization algorithm based on dynamic time warping (DTW), using the speech feature sequence as the reference time axis. The relative timing offset between the expression parameter sequence (sampled at 50 Hz) and the body movement sequence (sampled at 30 Hz) is calculated, and millisecond-level alignment accuracy is achieved through spline interpolation with a preset number (typically cubic). For text data, time tags are generated for the nodes of the dependency syntax tree based on the word-level timestamps of the speech recognition results to ensure alignment between semantic units and speech segments. A cross-modal trigger detection mechanism is also established. When an "eyebrow raise" action in the expression parameter sequence and an interrogative tone (F0 rise of 20%) in the speech features co-occur within ±200 ms, they are forcibly synchronized and marked as an "interrogative tone" event. Finally, a time-aligned multimodal data stream is output, with each modal data stream annotated with a unified time code and event trigger flag, providing standardized input for subsequent fusion.
[0090] In this embodiment, the weights of each modality are dynamically adjusted based on the spatiotemporally aligned multimodal data, including: constructing a weight allocation network based on the attention mechanism, where the input layer receives aligned cross-modal features such as speech spectrum energy, expression AU intensity, and limb joint angle change rate; analyzing the temporal correlation of each modal feature through a gated recurrent unit (GRU); when "arm waving" (spatial trajectory variance > 0.3) is detected in the limb movement sequence and overlaps with the speech energy peak (> 65dB), the limb modality weight is increased from 0.2 to 0.5; a differentiable decision tree model is used to achieve context-adaptive adjustment, assigning a basic weight of 0.7 to the speech and text modalities in question-and-answer scenarios, and increasing the expression modality weight to 0.6 in emotional interaction scenarios; and generating a dynamic weight matrix containing a time dimension at the output layer, where the weight value of each modality at any time t satisfies the constraint condition ∑w(t) = 1, and a sigmoid function is used to prevent gradient explosion.
[0091] In this embodiment, multimodal data and dynamic weights are integrated to construct an interactive feature matrix, including: designing a hierarchical fusion architecture, in which the bottom layer maps the speech MFCC features (128 dimensions), text BERT embedding (768 dimensions), expression AU parameters (32 dimensions), and limb bone coordinates (75 dimensions) to a unified 256-dimensional latent space through a fully connected layer; the middle layer adopts a cross-modal attention mechanism, and uses the dynamic weight matrix as a guide to calculate the correlation score between features. For example, when the expression weight w>0.4, the association between the mouth AU25 (lip opening) feature and the explosive sound segment in the speech spectrum is strengthened; the top layer outputs an interaction matrix containing a time-modality-feature three-dimensional tensor with a dimension of [T×M×256] (T is the time step, M=4 is the number of modalities), and adds learnable position encoding according to time slices; before being input into the pre-trained virtual digital human behavior decision model, the matrix undergoes layer normalization and residual connection processing to ensure the stability of gradient flow, ultimately supporting the simultaneous driving of the digital human's speech synthesis, facial animation, and body movement generation.
[0092] The beneficial effects of this technical solution are: through precise data alignment and weight adjustment in the cross-modal fusion module, voice, text, facial expressions, and body movement data can be effectively synchronized, ensuring the efficient fusion and temporal correlation of multimodal information. The weight adjustment unit dynamically adjusts based on the actual importance of each data source, improving the accuracy and emotional adaptability of the virtual digital human's responses. By constructing an interaction matrix based on multimodal temporal correlation features, more natural, intelligent, and emotionally expressive interactive feedback can be generated, significantly improving the user experience and providing reliable support for the application of intelligent virtual digital human technology.
[0093] Example 5:
[0094] The embodiment of the present invention provides an intelligent real-time interactive question-answering system based on a virtual digital human, and a weight adjustment unit, comprising:
[0095] Coupling subunit: performs coupling analysis on the aligned standardized speech feature vectors, expression dynamic parameter sequences, and body movement sequence data to obtain the speech-expression coupling coefficient and the speech-movement coupling coefficient;
[0096] A first adjustment subunit: performing a first adjustment on the weight distribution ratio of the aligned standardized speech feature vector, the expression dynamic parameter sequence, and the body movement sequence data according to the speech-expression coupling coefficient and the speech-action coupling coefficient;
[0097] Parameter acquisition subunit: This unit performs micro-expression analysis on the user's facial expression dynamic parameter sequence and body movement sequence data. It also performs voice fundamental frequency jitter detection on the aligned standardized speech feature vectors to determine the user's physiological state parameters, including the emotional tension index and attention concentration score.
[0098] A second adjustment subunit: performing a second adjustment on the weights of the aligned standardized speech feature vector, expression dynamic parameter sequence, and body movement sequence data based on the user's physiological state parameters;
[0099] The third adjustment subunit: performs a third adjustment on the preset initial weights of the structured text data based on the weights of the standardized speech feature vector, expression dynamic parameter sequence and body movement sequence data after the second adjustment.
[0100] In this embodiment, the weight distribution ratio is adjusted according to the voice-expression coupling coefficient Kve and the voice-action coupling coefficient Kvb, including: designing a dual-channel adaptive weighting module, when Kve>0.6 and Kvb<0.3, the expression modality weight is increased from 0.3 to 0.55, and the body modality weight Wbody is reduced to 0.15; adopting a logarithmic proportional distribution strategy, the voice weight Wvoice=1-α(Kve+Kvb), where α=0.35 is the empirical attenuation factor; for coupling conflict scenarios (such as Kve>0.8 but the body is stiff), the peak suppression mechanism is activated to forcibly limit the increase of the expression weight to no more than 120% of the coupling coefficient; the adjustment result is processed by a time smoothing filter to prevent weight jumps from causing discontinuous movements of the digital human, and finally outputting a weight vector with a timestamp [wv(t), wf(t), wb(t)], which is the voice weight vector, the expression weight vector, and the body weight vector, respectively, to meet the real-time update frequency of ≥10Hz;
[0101] In this embodiment, physiological state parameters are determined based on micro-expression analysis and voice fundamental frequency jitter detection, including: using a 3D CNN model to detect small facial muscle movements (duration of frames < 1 / 3 second), marking the emotional tension index T∈[0,1] when the upper lip levator (AU10) and eyelid tension (AU7) co-occur; calculating jitter (jitter rate) and shimmer (amplitude disturbance) using an improved fundamental frequency trajectory analysis algorithm; when jitter>1.2%, the attention concentration score C decreases by 0.2 / second; integrating behavioral features such as head posture angle (pitch>15°) and hand contact frequency (touching the face>3 times / minute), a physiological state assessment model based on LightGBM is constructed, and the output is an emotion-attention binary (Tt, Ct) containing a timestamp, where a "high pressure state" flag is triggered when Tt>0.7, and a "distraction warning" is activated when Ct<0.4;
[0102] In this embodiment, a second weight adjustment is performed based on physiological state parameters, including: when T>0.7, an emergency compensation strategy is introduced to increase the voice weight vector by Δwv=0.2(1-C) to enhance information transmission efficiency; for the distracted state of C<0.5, a body movement reinforcement scheme is adopted to make the body weight vector wb+=0.15×∑(body joint movement amplitude) / T, so as to re-attract the user's attention through large movements; a weight arbitrator based on fuzzy logic is designed, which inputs the time difference values ΔT / Δt and ΔC / Δt of T and C. When a rapid increase in tension is detected (ΔT / Δt>0.1 / s), the expression weight is immediately reduced by 20% to avoid excessive emotional rendering; the final output dynamic weight is superimposed on the first adjustment result to form a weight distribution that is adaptive to the physiological state;
[0103] In this embodiment, the third adjustment includes: After the second adjustment is completed, the system has determined the weight distribution ratios for speech, expression, and gesture (e.g., 60% for speech, 25% for expression, and 15% for gesture). At this point, structured text data needs to be incorporated into the weight adjustment system. The system uses semantic-modality association rules to pre-define the mapping relationship between different text types and multimodal data. For example, interrogative sentences (such as "Why?") rely more on the weight of speech and intonation, while descriptive text (such as "I went to the park yesterday") relies more on the coordination of expression and body language. The system analyzes the syntactic structure, sentiment polarity, and keywords of the current text, matches it to a preset modality preference template, and generates an initial weight adjustment direction for the text (e.g., increasing or decreasing the influence of the text on the final decision). The reliability of the current text data is assessed through ASR confidence scoring and semantic coherence analysis. If the text confidence falls below a threshold (e.g., due to speech-to-text errors resulting in semantic ambiguity), its weight is reduced, and dynamic compensation is performed based on the speech-expression-gesture coupling coefficient. For example, if a user says "yes" but shakes their head, the weight of the text "yes" will be lowered. The conflict between the affirmative tone of voice and the head shake triggers modal arbitration, ultimately giving priority to the voice and gesture. Furthermore, the system checks for semantic consistency between text and expressions / actions (e.g., the text reads "happy" but the expression parameters indicate a frown), and locally downgrades the weight of inconsistent text segments. To prevent frequent weight jumps that could cause unnatural digital human feedback, the system employs a time-windowed recursive smoothing strategy. The final weight of the current text segment is determined not only by real-time analysis results but also by the historical average weight of similar text within the previous five seconds. For example, if a user asks a question repeatedly in a high-pitched voice, the weight of subsequent questions will inherit the upward trend of the previous ones, rather than being recalculated. The system also performs low-pass filtering on short-term noise bursts (such as text jumps caused by speech misrecognition) to ensure that the weight change curve conforms to the inertia and gradual nature of human interaction. For users with long-term interactions, the system records their preferences for adjusting text weights in past interactions to build a personalized weight model. For example, some users prefer using body language instead of verbal expression (e.g., nodding instead of saying "yes"). The system will gradually reduce the default weight of text in such scenarios and increase the priority of motion data. The adaptive reinforcement module uses online learning to update user profiles in real time. Each weight adjustment is fed back into the user preference database, forming a closed-loop optimization loop. After all adjustments are completed, the system normalizes the weights of speech, text, expression, and motion to ensure that the sum of these four is 100%. A softmax function is introduced during this normalization process to prevent the weight of a single modality from being suppressed to zero. The resulting weight combination serves as the coefficient of the interaction feature matrix and is input into the virtual digital human's behavioral decision-making model. For example, the output might be: speech 40%, text 25%, expression 20%, and motion 15%. Based on this, the digital human generates feedback behaviors that are both semantically accurate and emotionally expressive.
[0104] In this embodiment, coupling analysis is performed on the aligned standardized speech feature vector, expression dynamic parameter sequence, and body movement sequence data, including determining a speech-expression coupling coefficient Kve based on the aligned standardized speech feature vector and expression dynamic parameter sequence, and determining a speech-action coupling coefficient Kvb based on the aligned standardized speech feature vector and body movement sequence data;
[0105] In this embodiment, the speech-expression coupling coefficient Kve is used to evaluate the temporal correlation between speech energy changes and facial expression intensity changes. The speech feature sequence is: , each The dimension is Characteristic vector, expression parameter sequence: , each Include Dynamic expression parameters (such as AU intensity), time window length: Indicates the number of sliding window frames. The calculation process is: extract the rate of change (within the sliding window):
[0106] Voice change rate:
[0107]
[0108] Expression change rate:
[0109]
[0110] Calculate the Pearson correlation coefficient:
[0111]
[0112] in, The Pearson correlation coefficient of the voice change rate and the expression change rate is used to measure the degree of synchronization between the two change rates. The result range is , the closer to 1, the more positive correlation As the voice-expression coupling coefficient Kve, it describes the linkage (coupling strength) between voice and expression. is the speech energy change rate sequence within the sliding window; is a sequence representing the rate of change of expression intensity within the sliding window; is the covariance of the two; and They are and The standard deviation of
[0113] In this embodiment, the speech-action coupling coefficient Kvb measures the frequency matching between speech rhythm and body movements. The speech-action coupling coefficient measures the dynamic matching between speech intonation and body movements. The speech feature sequence is: , body movement sequence: , each Composed of joint angles or displacements dimensional vector, L is the joint angle / displacement dimension, and the time window length is: , calculation process:
[0114] Extract speech rhythm frequency (based on fundamental frequency Average period of time):
[0115]
[0116] Extracting movement rhythm frequency (based on frequency domain analysis of movement energy change rate):
[0117]
[0118]
[0119] Calculate the coupling coefficient (normalized based on the frequency difference):
[0120]
[0121] Among them, As the voice-action coupling coefficient Kvb, the value range is , the closer it is to 1, the more coordinated and synchronized the rhythm is. is the action rhythm frequency, that is, the main frequency period of the action in window t, is the speech rhythm frequency, that is, the average rhythm period of speech in window t, is the rate of change of limb movement energy, is the absolute value of the frequency difference (the smaller the value, the more synchronized it is), The normalization factor ensures that the output is in the range [0,1];
[0122] In this embodiment, the calculation process of the user's physiological state parameters is as follows:
[0123] Step 1: Micro-expression recognition and AU intensity analysis, identifying facial micro-expressions and estimating emotional tension to obtain micro-expression frequency Method: AU parameters in the FACS system (such as AU2: eyebrow raising, AU12: mouth corner tightening) are used to calculate the AU intensity difference between consecutive frames. If the AU intensity change rate exceeds the dynamic threshold (such as >30%) in a very short time (<500ms), it is considered a micro-expression event, and the corresponding AU intensity change rate is determined as the micro-expression frequency. , Meaning: The frequency and amplitude of micro-expressions are positively correlated with tension. For example, frequent pursing of lips (AU14) or frowning (AU4) often indicate that the user is under stress. Perform normalization;
[0124] Step 2: Body movement stability assessment, detect the stability changes of the user's body, and judge the attention state. Method: joint angular velocity vector Perform a sliding window ANOVA:
[0125]
[0126] in, is the joint velocity vector of the t-th frame, is the mean velocity in the window, is the window frame number, if the limb movement stability score Higher than the set baseline (the threshold in a relaxed state) indicates abnormal jitter or agitation. Meaning: High variance usually reflects inattention or emotional distress and can be used as an indirect indicator of distraction. Step 3: Speech fundamental frequency jitter analysis evaluates the frequency fluctuation characteristics of the speech signal that reflect emotional fluctuations. Method:
[0127] Dynamically regularize the fundamental frequency (F0) sequence in the standardized speech feature vector and extract its short-term coefficient of variation (jitter):
[0128]
[0129] in, is the fundamental frequency jitter rate, For the analysis window length, The fundamental frequency value of the tth frame is combined with the fundamental frequency mean change trend (e.g., a sudden increase) as an auxiliary signal of tension. An increase in jitter rate (e.g., exceeding 5%) generally corresponds to a state of emotional tension. The system also further adjusts the emotional tension index by combining the fundamental frequency mean deviation (e.g., a sudden increase may indicate excitement).
[0130] Step 4: Fusion modeling and emotion / attention index calculation, integrating the three modal information to calculate unified physiological state parameters. Method:
[0131] Emotional tension index Using weighted fusion formula:
[0132]
[0133] in : Normalized micro-expression frequency; modal weight default: , attention concentration score : Based on the inverse combination of action stability and speech pause frequency (the specific expression can be expanded to linear or nonlinear mapping), meaning: the higher the The lower the value, the more tense it is; A higher value indicates more distracted attention. To prevent transient noise interference (such as a user's sudden cough causing speech abnormalities), the system uses a sliding median filter to smooth EEE and AAA, preserving trend changes. The final output emotional tension index and attention concentration score are quantified on a 0-100 scale and accompanied by a confidence level (e.g., low confidence results do not trigger weight adjustment).
[0134] The beneficial effect of this technical solution is that, by introducing coupling subunits and a three-layer weight adjustment mechanism, the system can accurately capture the coupling relationship between speech, expression, and movement, and dynamically optimize the weights of each modal data based on the user's physiological state parameters. This mechanism not only improves the accuracy and contextual adaptability of multimodal information fusion, but also enhances the behavioral decision-making sensitivity and personalized response capabilities of the virtual digital human in different user states, thereby achieving a more natural and emotionally adaptable human-computer interaction experience.
[0135] Example 6:
[0136] An embodiment of the present invention provides an intelligent real-time interactive question-and-answer system based on a virtual digital human, and a behavior decision module, including: a historical data weight adjustment unit: based on historical interaction data and an interaction feature matrix, determines the multimodal weight distribution of user interaction behavior, obtains short-term interaction frequency, multimodal signal volatility, and emotional tendency deviation, performs multimodal confidence analysis and interaction habit matching, and obtains a voice weight correction coefficient, a text weight correction coefficient, and an action weight correction coefficient;
[0137] Cross-modal association network construction unit: monitors the signal synchronization between standardized speech feature vectors, structured text data and expression dynamic parameter sequences during the interaction process, determines the speech-text synchronization rate and speech-expression synchronization rate, and generates a multimodal synchronization relationship network;
[0138] Feedback decision splitting and prediction unit: This unit performs semantic analysis based on structured text data, extracts interaction intent labels, and combines them with the sentiment bias to determine the feedback level division ratio and feedback rhythm adjustment coefficient.
[0139] Correlated multimodal anomaly analysis unit: Based on the multimodal synchronization relationship network, it uses the feedback level division ratio and feedback rhythm adjustment coefficient to perform cross-modal consistency verification, determine the text-speech anomaly coefficient and the expression-action anomaly coefficient, and combine the speech weight correction coefficient, text weight correction coefficient, and action weight correction coefficient to generate a composite anomaly index;
[0140] Instruction integration output unit: Generates a decision instruction set based on the composite anomaly index and the preset index-instruction database, where the decision instruction set includes sentiment tendency parameters, knowledge graph association index and feedback time series.
[0141] In this embodiment, the historical data weight adjustment unit dynamically weights the multimodal historical data in the interaction feature matrix using a pre-trained time series analysis model. Specifically, the attenuation coefficient is calculated based on the short-term interaction frequency of 10 consecutive interactions, and the initial weight is generated by combining the volatility of the multimodal signal (such as the variance of the speech spectrum and the change rate of text word frequency). At the same time, the emotional tendency deviation is obtained by analyzing the user's historical emotional label sequence (such as the change in the positive / negative ratio) through the LSTM network. The unit outputs the weight correction coefficient for speech, text, and action (range 0.8-1.2). The correction rule is: when the emotional deviation exceeds the threshold of ±0.3, the corresponding modal weight is increased by 15%. This process performs incremental updates every 5 seconds to ensure that the weight matches the user's real-time interaction habits.
[0142] In this embodiment, the cross-modal association network construction unit establishes a synchronization relationship diagram between multimodal signals to achieve inter-modal coordination modeling and topological analysis. The key steps are: synchronization rate calculation: speech-text synchronization rate: ;
[0143] Voice-expression synchronization rate: ; Graph construction: create a node set , edge weight , forming a dynamic weighted graph , graph attention mechanism and modal center activation: If and , then activate "expression" as the central node to guide the outward spread of information; modal clustering recognition: update the graph structure every 200ms, and identify modal clusters through the community discovery algorithm; when the "expression-action" modularity is greater than 0.7, it is marked as the dominant state of non-verbal expression for reference by the anomaly detection module.
[0144] In this embodiment, the feedback decision split prediction unit dynamically builds feedback strategies based on semantic content and emotional signals, splits resource ratios and rhythm control. Key steps: Interaction intent recognition: Based on the dependency syntax tree + BiLSTM-CRF structure, 28 types of interaction intent labels (such as "social" and "complaint") are extracted; feedback level allocation: using the emotional offset degree Calculate the feedback resource ratio triple: ;like , 70% of the feedback resources are allocated to the "emotional response" layer, rhythm control mechanism: build a reinforcement learning rhythm controller, input the voice pause duration and dependency distance, and output the feedback rhythm coefficient ; If the user's speaking speed is > 5 words / second, the response will be accelerated , output strategy triplet: As a global control parameter for multimodal feedback.
[0145] In this embodiment, the associated multimodal anomaly analysis unit improves the accuracy and interpretability of anomaly detection by performing consistency checks between multimodal signals and integrating cross-anomaly indicators. The key steps are: anomaly indicator calculation:
[0146] Text-to-speech anomaly coefficient: ; Expression-action abnormality coefficient: ; Construction of composite anomaly index: ; Among them, the weight They are the text-speech modality weight and the expression-action modality weight, which are dynamically adjusted according to the scene and range [0,1]; is the voice credibility attenuation factor, ranging from [0, 0.5], and the voice signal-to-noise ratio (SNR<15dB) =0.4), or abnormal fundamental frequency jitter (jitter rate>10% =0.3), is the text credibility attenuation factor, ranging from [0, 0.5], when the text logical contradiction (such as when a semantic conflict sentence appears in a round of dialogue) =0.2), or keyword confidence <70% =0.4, is the background interference penalty term, ranging from [0, 0.3], when the ambient light intensity changes suddenly (> 200 lux / ms =0.2) or sudden noise (when the energy suddenly increases by 30dB) =0.3), dynamic weight constraints: α+β≤1, and λv+λt+μb≤1, ensuring A∈[0,2], As a compensation item, when the reliability of a single mode decreases (such as when the voice is contaminated by noise), the compensation mechanism is used to prevent the A value from being too high; three-level response mechanism: Level 1 (0.6 <a ≤ 1.0):插入澄清提问;三级(a>1.0): Start the emergency dialogue process; Abnormal topology update: synchronize the analysis results to the abnormal attribute matrix of the multimodal graph to achieve system-level traceability analysis support.
[0147] In this embodiment, the command integration output unit integrates all strategy results and outputs a structured multimodal driving command set to achieve joint control of voice, action, expression, etc., including: Emotion parameter mapping: If the composite abnormality index , select the "neutral comfort" feedback parameter (intensity 0.7), knowledge node optimization: In this case, explanatory nodes with confidence > 85% in the knowledge graph are preferred. Time series generator: according to the rhythm coefficient Split into: Voice segment: Continuous Seconds; Action period: continuous seconds; intermediate interpolation: insert expression transition frames, output instruction set: output a structured control set of 17 fields in JSON format, covering semantic labels, feedback rhythm, modal control signals, action trajectories, etc.; support millisecond-level response delivery.
[0148] The beneficial effects of the above technical solution are as follows: the behavioral decision-making module dynamically optimizes multimodal weights by fusing historical interaction data with the interaction feature matrix, accurately capturing user interaction habits and emotional changes. The cross-modal association network ensures high synchronization of voice, text, and expression signals, and combines semantic parsing with anomaly analysis to generate a composite anomaly index, improving the accuracy and contextual adaptability of feedback. The instruction integration output unit generates a decision instruction set based on the anomaly index, including emotional tendencies, knowledge indexes, and feedback time series, significantly enhancing the personalized response capabilities and natural interaction of the virtual digital human, providing users with an immersive, emotionally resonant intelligent interactive experience.
[0149] Example 7:
[0150] The embodiment of the present invention provides an intelligent real-time interactive question-answering system based on a virtual digital human, and a knowledge retrieval module, including:
[0151] Knowledge fragment search unit: Based on the knowledge graph association index in the decision instruction set, it traverses the preset distributed heterogeneous database to identify the set of knowledge fragments that match the current interaction intention. During the retrieval process, it prioritizes frequently accessed knowledge nodes in the user's historical interactions and sorts the relevance of the knowledge fragments.
[0152] Emotional adaptability correction unit: Based on the emotional tendency parameters in the decision instruction set, it performs emotional adaptability matching on the retrieved knowledge fragments, identifies the vocabulary and expressions that need to be adjusted, and combines the preset user-preferred tone pattern and response rhythm to perform modal word embedding and sentence structure optimization to generate a response text that conforms to the current interactive emotional state;
[0153] Speech feature annotation unit: Based on the answer text, it annotates the corresponding speech feature parameters for different semantic paragraphs, including the intonation rise and fall range, speech speed adjustment ratio and stress distribution position, to generate speech features that can drive the speech synthesis module.
[0154] In this embodiment, according to the knowledge graph association index in the decision instruction set, the preset distributed heterogeneous database is traversed to identify the set of knowledge fragments that match the current interaction intention. During the retrieval process, the frequently accessed knowledge nodes in the user's historical interactions are prioritized, and the relevance of the knowledge fragments is sorted. The following steps are included: (1) Knowledge graph index mapping: parsing the decision instruction set, extracting the key entities of the knowledge graph association index (such as "product model_A123"), and matching the corresponding storage locations in the node index table of the distributed heterogeneous database; (2) Historical interaction optimization retrieval: calling the frequently accessed knowledge nodes in the user's recent 100 interaction records. For each knowledge node (e.g., "warranty policy" was accessed 35 times), a weight coefficient of +0.5 was set for the knowledge fragments containing these nodes; (3) Multi-dimensional relevance calculation: the search results were sorted according to the following priorities: Direct matching: fragments that fully matched the user's query keywords (e.g., "A123 battery life") were placed at the top; Logical relevance: fragments that were indirectly related through the knowledge graph relationship chain (e.g., "A123 and compatible charger matching guide") were sorted in descending order by the number of related edges; (4) Real-time cache loading: the top 5 knowledge fragments were preloaded into the memory cache of the edge computing node, and the response delay was controlled within 200ms.
[0155] In this embodiment, based on the sentiment tendency parameter in the decision instruction set, sentiment adaptability matching is performed on the retrieved knowledge fragments, the vocabulary and expressions that need to be adjusted are identified, and combined with the preset user preferred tone pattern and response rhythm, modal word embedding and sentence structure optimization are performed to generate a response text that conforms to the current interactive emotional state, including the following steps: (1) Sentiment parameter mapping correction rule: If the sentiment tendency parameter is "dissatisfied_0.6": replace "suggest you" in the original text with "we will handle it for you immediately", insert the apologetic modal word "sorry" and the commitment phrase "solve it within 24 hours"; If the sentiment tendency parameter is "pleased_0.8": add the exclamation word "Great!" at the beginning of the sentence, and optimize the declarative sentence "The operation is completed" to "It has been easily solved for you!"; (2) User preference matching: Based on the "prefer concise expression" feature marked in the user profile, delete the redundant modifiers in the knowledge fragment and split the complex sentence into a sequence of short sentences; (3) Rhythm adjustment: add segment pause markers to technical description paragraphs (such as parameter lists), and insert a 500ms silence interval every three data items.
[0156] In this embodiment, according to the answer text, the corresponding speech feature parameters are annotated for different semantic paragraphs, including the intonation rise and fall interval, the speech speed adjustment ratio and the stress distribution position, to generate emotional speech feature data that can drive the speech synthesis module, including the following steps: (1) semantic segmentation annotation: key conclusion paragraph (such as "final solution"): mark the intonation rise interval (the fundamental frequency is increased by 20Hz), and the speech speed is reduced to 0.8 times the baseline value; warning content (such as "do not disassemble the device"): set the stress energy increase of 150% at the prohibition verb "don't" and prolong the tone. The segment duration is 50ms; (2) Emotional rhythm generation: for texts with the emotional tendency parameter "urgent_0.9", the global speaking speed is increased to 1.2 times, and a 0.3-second rapid breathing sound effect is inserted every 5 seconds; for texts with the emotional tendency parameter "comfort_0.7", the intonation fluctuation range is compressed to ±5Hz, and the stress distribution interval is expanded by 1.5 times to simulate a calm tone; (3) Cross-modal verification: Detect whether the speech feature data conflicts with the pronunciation syllable segmentation grouping of the lip matching unit. If there is a cross-modal timing deviation of more than 100ms, the speech feature is given priority to reconstruct the text segmentation.
[0157] The beneficial effects of the above technical solution are: through the combination of knowledge fragment search and emotional adaptability correction, the system can accurately retrieve and adjust the answer content according to the user's historical interactions and current emotional state, achieving a more personalized and emotional response. The emotional tendency parameters are combined with the user's preferred tone mode to optimize the tone and sentence structure, making the answer more friendly and natural. At the same time, the speech feature annotation unit provides accurate speech parameter support for each semantic paragraph, ensuring that the speech synthesis module can generate speech feedback that is consistent with the emotional state, thereby significantly improving the interactive experience and emotional resonance ability of the virtual digital human.
[0158] Example 8:
[0159] The embodiment of the present invention provides an intelligent real-time interactive question-answering system based on a virtual digital human, and a speech generation module, including:
[0160] Lip-sync matching unit: This unit divides the response text into corresponding pronunciation syllable segments based on the timestamp distribution of the feedback time series. It extracts the speech features of each syllable segment and matches them with the preset standard lip-sync animation template of the virtual digital human to generate lip-sync animation keyframes synchronized with the actual pronunciation.
[0161] Micro-expression control unit: Based on the intensity and type of emotional tendency parameters, it selects a corresponding level of micro-expression parameter combination from the preset expression library, including the amplitude of eyebrow movement, the frequency of eye opening and closing, and the angle of mouth corner upward movement. It then adjusts the rate of micro-expression change according to the rhythm of the feedback time series to form a micro-expression parameter sequence that matches the emotional expression of the voice;
[0162] The body movement choreography unit identifies key semantic nodes in the response text. If the speech features contain emphatic accents or long pauses, a preset body movement template is inserted at the corresponding moment, including the hand gesture trajectory and torso tilt angle. The movement amplitude is proportionally scaled according to the emotional tendency parameter to obtain the body movement trajectory.
[0163] Multimodal synchronization unit: This unit performs timeline calibration on lip-sync animation keyframes, micro-expression parameter sequences, and body movement trajectories to ensure that the starting moment and accent trigger point of the voice response are aligned with the peak moment of the movement trajectory. The voice response is generated and pushed to the user terminal after rendering and calculation at the edge computing node.
[0164] In this embodiment, the answer text is divided into corresponding pronunciation syllable segments according to the timestamp distribution of the feedback time series, including the following steps: performing phoneme-level analysis on the answer text, extracting the start and end timestamps of each syllable, and generating a pronunciation timing mapping table; based on a standard lip animation template library, matching the lip closure, tongue position and mandibular opening and closing parameters of each phoneme to generate basic lip key frames; dynamically compressing or stretching the key frame interval duration according to the change in speech speed in the voice characteristics, so that the lip animation is synchronized with the spectral characteristics of the real-time pronunciation; and interpolating and smoothing the transition frames of consecutive syllables to eliminate lip jumps.
[0165] In this embodiment, micro-expressions are regulated based on the intensity and type of the emotional tendency parameter, including the following steps: when the emotional tendency parameter is "anger_0.8", a frown template is called from the expression library, the eyebrow downward pressure amplitude is set to 120% of the baseline value, and the eye opening and closing frequency is increased to 1.2 times / second; if the total duration of the feedback time series is compressed by 30%, the maintenance duration of the micro-expression is synchronously shortened, and the upward angle of the mouth corners is attenuated from 15° to 5° to achieve an emotional weakening transition; an instantaneous expression enhancement instruction is inserted at the stressed syllable of the speech feature, for example, triggering a rapid eyelid closure action at the explosive consonant syllable.
[0166] In this embodiment, body movements are choreographed according to key semantic nodes, including the following steps: detecting the logical conjunctions "therefore" and "however" in the answer text, and inserting an emphatic gesture of 30° abduction of the right hand at the corresponding timestamp; when the voice feature contains a pause of more than 300ms, triggering a listening posture template of 10° forward leaning of the torso, and the tilt duration is proportionally matched to the duration of the pause; if the emotional tendency parameter is "joy_0.7", the amplitude of the gesture swing trajectory is expanded to 1.5 times the baseline value, and a celebratory action of raising both hands is added at the end of the voice.
[0167] In this embodiment, multimodal timeline calibration is performed, including the following steps: taking the starting moment of the voice response as the reference zero point, delaying 50ms to trigger the rendering of the first frame of the lip-sync key frame sequence to compensate for the device response delay; when the voice spectrum analysis shows the accent energy peak, synchronizing the eyebrow lifting peak of the micro-expression parameter sequence and the highest point of the fingertips of the body movement; using a timestamp rollback mechanism at the edge computing node, if it is detected that the lip animation deviates from the voice by more than 80ms, reloading the multimodal data packet of the current 200ms period.
[0168] The beneficial effects of the above technical solution are: through the collaborative work of lip matching, micro-expression control, body movement choreography, and multimodal synchronization units, the system achieves highly synchronized and natural expression of speech, expression, and movement. Lip matching ensures precise correspondence between pronunciation and animation, micro-expression control generates subtle expression changes based on emotional parameters, and body movement choreography drives dynamic movements through semantics and emotions, enhancing the realism of interaction. The multimodal synchronization unit ensures the real-time consistency of speech, expression, and movement through timeline calibration and edge rendering, significantly improving the naturalness of interaction and emotional expression of virtual digital humans, and providing users with an immersive and personalized human-computer interaction experience.
[0169] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. An intelligent real-time interactive question-answering system based on virtual digital human, characterized by: include: Data acquisition module: Receives user input voice and text data based on an integrated multimodal sensor array. It also collects the user's facial expression and body movement sequence data in real time, and processes the voice and text data to obtain standardized voice feature vectors and structured text data. Cross-modal fusion module: This module performs spatiotemporal alignment of standardized speech feature vectors and structured text data with the user's facial expression dynamic parameter sequence and body movement sequence data, and constructs an interactive feature matrix containing multimodal temporal correlation features. Behavior decision module: Determines a decision instruction set containing sentiment tendency parameters, knowledge graph association index, and feedback time series based on historical interaction data and interaction feature matrix; Knowledge retrieval module: Based on the knowledge graph association index in the decision instruction set, the corresponding knowledge fragment set is retrieved from the preset distributed heterogeneous database. Based on the sentiment tendency parameters, the corresponding vocabulary of the knowledge fragments in the knowledge fragment set is embedded with modal words and rhythmic rhythm annotation, generating emotionally adaptive answer text and voice features; Speech generation module: Based on the feedback time series, answer text, and speech features in the decision instruction set, it generates the lip animation key frames, micro-expression parameter sequence, and body movement trajectory of the virtual digital human, and generates speech responses, which are pushed to the user terminal through the edge computing node; The behavior decision module includes: a historical data weight adjustment unit: based on historical interaction data and the interaction feature matrix, it determines the multimodal weight distribution of user interaction behavior, obtains short-term interaction frequency, multimodal signal volatility, and sentiment tendency deviation, performs multimodal confidence analysis and interaction habit matching, and obtains voice weight correction coefficients, text weight correction coefficients, and action weight correction coefficients; Cross-modal association network construction unit: monitors the signal synchronization between standardized speech feature vectors, structured text data and expression dynamic parameter sequences during the interaction process, determines the speech-text synchronization rate and speech-expression synchronization rate, and generates a multimodal synchronization relationship network; Feedback decision splitting and prediction unit: This unit performs semantic analysis based on structured text data, extracts interaction intent labels, and combines them with the sentiment bias to determine the feedback level division ratio and feedback rhythm adjustment coefficient. Correlated multimodal anomaly analysis unit: Based on the multimodal synchronization relationship network, it uses the feedback level division ratio and feedback rhythm adjustment coefficient to perform cross-modal consistency verification, determine the text-speech anomaly coefficient and the expression-action anomaly coefficient, and combine the speech weight correction coefficient, text weight correction coefficient, and action weight correction coefficient to generate a composite anomaly index; Instruction integration output unit: Generates a decision instruction set based on the composite anomaly index and the preset index-instruction database, where the decision instruction set includes sentiment tendency parameters, knowledge graph association index and feedback time series.
2. The intelligent real-time interactive question-answering system based on virtual digital human according to claim 1, characterized in that: Data acquisition module, including: Speech processing unit: performs frequency domain energy detection on the received speech data, identifies and removes invalid speech segments whose energy is lower than a preset threshold, and obtains continuous valid speech waveform data; Text analysis unit: performs dependency syntax analysis on text data, annotates entity types and action predicates, and generates initial text data with semantic labels; Environmental analysis unit: collects three-dimensional spatial information of the interactive environment, tracks the movement trajectory of the target user in the interactive space, and calculates the user's spatial position coordinates and orientation angle data; Image acquisition unit: Dynamically adjusts the focal length of the visual sensor and the directional parameters of the sound pickup array based on the spatial position coordinates and orientation angle data to obtain the user's facial expression image and body movement image data; Image analysis unit: Analyzes the spatial offset of key feature points in facial expression images, measures the deformation amplitude and duration of specific facial areas, and generates a sequence of expression dynamic parameters; Motion analysis unit: extracts the spatial coordinates of key skeletal nodes from limb motion image data, determines the relative movement trajectories between nodes, and generates limb motion sequence data; Data processing unit: processes continuous valid speech waveform data and initial text data to obtain standardized speech feature vectors and structured text data.
3. The intelligent real-time interactive question-answering system based on virtual digital human according to claim 2, characterized in that: Data processing unit, including: Speech data processing subunit: Uses an adaptive noise cancellation algorithm to preprocess continuous valid speech waveform data, extract acoustic features, and generate standardized speech feature vectors containing dynamic time warping information; Text data processing subunit: The initial text data is segmented through a domain-adaptive semantic segmentation model, and intent is annotated in combination with a preset intent recognition knowledge base to generate structured text data with contextual semantic associations.
4. The intelligent real-time interactive question-answering system based on virtual digital human according to claim 1, characterized in that: Cross-modal fusion module, including: Data alignment unit: aligns the standardized speech feature vector, structured text data, user's expression dynamic parameter sequence and body movement sequence data in time and space; Weight adjustment unit: adjusts the weights of the standardized speech feature vector, structured text data, expression dynamic parameter sequence and body movement sequence data after spatiotemporal alignment; Matrix construction unit: Integrate standardized speech feature vectors, structured text data, user's expression dynamic parameter sequence and body movement sequence data as well as the adjusted corresponding weights to construct an interactive feature matrix containing multimodal temporal correlation features.
5. The intelligent real-time interactive question-answering system based on virtual digital human according to claim 4 is characterized in that: Weight adjustment unit, including: Coupling subunit: performs coupling analysis on the aligned standardized speech feature vectors, expression dynamic parameter sequences, and body movement sequence data to obtain the speech-expression coupling coefficient and the speech-movement coupling coefficient; A first adjustment subunit: performing a first adjustment on the weight distribution ratio of the aligned standardized speech feature vector, the expression dynamic parameter sequence, and the body movement sequence data according to the speech-expression coupling coefficient and the speech-action coupling coefficient; Parameter acquisition subunit: This unit performs micro-expression analysis on the user's facial expression dynamic parameter sequence and body movement sequence data. It also performs voice fundamental frequency jitter detection on the aligned standardized speech feature vectors to determine the user's physiological state parameters, including the emotional tension index and attention concentration score. A second adjustment subunit: performing a second adjustment on the weights of the aligned standardized speech feature vector, expression dynamic parameter sequence, and body movement sequence data based on the user's physiological state parameters; The third adjustment subunit: performs a third adjustment on the preset initial weights of the structured text data based on the weights of the standardized speech feature vector, expression dynamic parameter sequence and body movement sequence data after the second adjustment.
6. The intelligent real-time interactive question-answering system based on virtual digital human according to claim 1, characterized in that: Knowledge retrieval module, including: Knowledge fragment search unit: Based on the knowledge graph association index in the decision instruction set, it traverses the preset distributed heterogeneous database to identify the set of knowledge fragments that match the current interaction intention. During the retrieval process, it prioritizes frequently accessed knowledge nodes in the user's historical interactions and sorts the relevance of the knowledge fragments. Emotional adaptability correction unit: Based on the emotional tendency parameters in the decision instruction set, it performs emotional adaptability matching on the retrieved knowledge fragments, identifies the vocabulary and expressions that need to be adjusted, and combines the preset user-preferred tone pattern and response rhythm to perform modal word embedding and sentence structure optimization to generate a response text that conforms to the current interactive emotional state; Speech feature annotation unit: Based on the answer text, it annotates the corresponding speech feature parameters for different semantic paragraphs, including the intonation rise and fall range, speech speed adjustment ratio and stress distribution position, to generate speech features that can drive the speech synthesis module.
7. The intelligent real-time interactive question-answering system based on virtual digital human according to claim 1, characterized in that: Speech generation module, including: Lip-sync matching unit: This unit divides the response text into corresponding pronunciation syllable segments based on the timestamp distribution of the feedback time series. It extracts the speech features of each syllable segment and matches them with the preset standard lip-sync animation template of the virtual digital human to generate lip-sync animation keyframes synchronized with the actual pronunciation. Micro-expression control unit: Based on the intensity and type of emotional tendency parameters, it selects a corresponding level of micro-expression parameter combination from the preset expression library, including the amplitude of eyebrow movement, the frequency of eye opening and closing, and the angle of mouth corner upward movement. It then adjusts the rate of micro-expression change according to the rhythm of the feedback time series to form a micro-expression parameter sequence that matches the emotional expression of the voice; The body movement choreography unit identifies key semantic nodes in the response text. If the speech features contain emphatic accents or long pauses, a preset body movement template is inserted at the corresponding moment, including the hand gesture trajectory and torso tilt angle. The movement amplitude is proportionally scaled according to the emotional tendency parameter to obtain the body movement trajectory. Multimodal synchronization unit: This unit performs timeline calibration on lip-sync animation keyframes, micro-expression parameter sequences, and body movement trajectories to ensure that the starting moment and accent trigger point of the voice response are aligned with the peak moment of the movement trajectory. The voice response is generated and pushed to the user terminal after rendering and calculation at the edge computing node.
Citation Information
Patent Citations
Virtual person-based multi-mode interactive processing method and system
CN107765852A
Virtual doctor system based on knowledge graph driving and operation method thereof
CN117056536A
Multi-modal digital employee reception system and method based on edge calculation
CN119693750A
Digital human interaction method and system based on naked eye 3D visualization
CN119782466A