Virtual digital human multi-modal collaborative interaction method and system driven by large language model
Patent Information
- Application Number
- CN202610828584.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-09
- Publication Date
- 2026-09-18
AI Technical Summary
[0005]为解决上述技术问题,本发明提供大语言模型驱动的虚拟数字人多模态协同交互方法及系统,用于解决现有虚拟数字人交互技术中语义理解不准确、情感表达能力缺乏、交互输出协同性不足的问题
本发明通过跨模态语义对齐与注意力融合机制生成多模态语义特征向量,解决了现有技术中多模态交互数据集语义错位、语义割裂的问题,大幅提升了复杂交互场景下的语义理解精度与抗干扰能力;本发明通过多模态语义特征向量与历史上下文数据库构建语义提示向量,并引入迭代式推理策略,提升大语言模型的上下文连贯理解与动态意图更新能力,有效解决了语义断层与意图误判的问题。
Smart Images

Figure CN122776977A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of human-computer intelligent interaction technology, and in particular to a method and system for multimodal collaborative interaction of virtual digital humans driven by a large language model. Background Technology
[0002] With the rapid development of artificial intelligence, virtual digital humans have been widely used in fields such as virtual customer service, intelligent education, and entertainment media. However, existing virtual digital human interaction technologies still face problems such as limited interaction methods and insufficient utilization of multimodal information. They rely solely on single-channel interaction methods driven by text or voice, resulting in a stiff interactive experience and insufficient immersion.
[0003] Existing technologies rely on single-modal feature processing mechanisms and lack the ability to jointly model multimodal interaction data such as speech, text, actions, and facial expressions. This leads to problems such as insufficient contextual semantic understanding, weak cross-modal semantic correlation, low emotion recognition accuracy, and poor coordination between actions and speech. Furthermore, the interactive output of existing virtual digital humans is severely disconnected from the user's emotional state and intensity, failing to integrate real-time emotional signals such as user voice intonation and facial micro-expressions, and unable to handle inconsistencies in emotional expression between multimodal interaction data. The interactive output of existing virtual digital humans mainly employs individual synthesis and simple splicing methods, lacking collaborative optimization based on emotion expression strategies. This results in problems such as audio-visual asynchrony and a disconnect between facial expressions and semantics, severely limiting the naturalness and realism of the interaction.
[0004] Therefore, it is necessary to provide a method and system for multimodal collaborative interaction of virtual digital humans driven by a large language model to solve the above-mentioned technical problems. Summary of the Invention
[0005] To address the aforementioned technical problems, this invention provides a method and system for multimodal collaborative interaction of virtual digital humans driven by a large language model, which solves the problems of inaccurate semantic understanding, lack of emotional expression ability, and insufficient collaborative interaction output in existing virtual digital human interaction technologies.
[0006] The present invention provides a method for multimodal collaborative interaction of virtual digital humans driven by a large language model, the method comprising: A multimodal interaction dataset is collected in parallel during the user interaction process. The multimodal interaction dataset includes voice data, text data, gesture data, and facial expression data. The multimodal interaction dataset is input into a multimodal coding network for feature extraction to obtain speech semantic features, text semantic features, action semantic features and facial expression semantic features. A cross-modal feature fusion model is then called to perform semantic alignment and attention fusion to generate a multimodal semantic feature vector. Based on the multimodal semantic feature vector and the historical context database, a semantic prompt vector is constructed and input into the large language model. An iterative reasoning strategy is used to perform contextual understanding and user intent recognition on the semantic prompt vector to generate semantic response content and corresponding action planning instructions. A multimodal emotion fusion network is used to identify the emotional state and emotional intensity of the multimodal semantic feature vectors, and to determine the emotional expression strategy corresponding to the semantic response content; The semantic response content, the emotion expression strategy, and the action planning instructions are converted into a multimodal synchronous expression sequence through a multimodal synthesis engine, and the virtual digital human is driven to perform multimodal interaction.
[0007] Preferably, the step of inputting the multimodal interaction dataset into a multimodal coding network for feature extraction to obtain speech semantic features, text semantic features, action semantic features, and facial expression semantic features specifically includes: The speech data is subjected to sliding frame processing based on a preset time window length and window shift length, and the speech data after sliding frame processing is subjected to spectrum transformation using fast Fourier transform to extract speech spectrum features; the speech spectrum features are input into the speech encoder and local feature extraction and downsampling are performed to generate the speech semantic features; The text data is split into word sequence using the WordPiece word segmentation algorithm and input into the text encoder for feature modeling and nonlinear mapping to generate the text semantic features. The gesture data is used to estimate the pose to obtain a sequence of human joint coordinates. The sequence of human joint coordinates is temporally sampled based on a preset sampling time interval to generate a temporal sampling sequence of human joints and calculate the corresponding motion displacement features and motion velocity features. The sequence of human joint coordinates, the motion displacement features and the motion velocity features are concatenated by channel dimension to generate joint motion trajectory features and input into the motion encoder for graph convolution operation to generate the motion semantic features. The facial expression data is subjected to face region detection, facial key point coordinates are extracted and input into a visual encoder, local expression texture features are extracted using a multi-layer convolutional feature extraction network, global facial structure features are extracted using a spatial pooling layer, and the local expression texture features and the global facial structure features are concatenated by channel dimension to generate the expression semantic features.
[0008] Preferably, the step of calling the cross-modal feature fusion model to perform semantic alignment and attention fusion to generate a multimodal semantic feature vector specifically includes: The speech semantic features, text semantic features, action semantic features and facial expression semantic features are unified in dimension by a unified feature mapping layer, and a multimodal feature set is obtained by summarizing them. Based on the multimodal feature set, the cross-modal feature fusion model is invoked to calculate the semantic association weights, and the corresponding calculation formula is as follows: In the formula, Representing semantic features With semantic features Semantic association weights between them; This represents an exponential function with the natural constant e as the base; n represents the number of multimodal feature sets. Construct a cross-modal attention matrix based on the semantic association weights. The multimodal feature set is then semantically aligned to generate a semantically aligned multimodal feature set. An attention fusion strategy is used to perform weighted fusion on the semantically aligned multimodal feature set to generate the multimodal semantic feature vector.
[0009] Preferably, the step of constructing a semantic cue vector based on the multimodal semantic feature vector and the historical context database specifically includes: The historical context database is invoked, which includes a short-term context database, a medium-term context database, and a long-term context database. A preset semantic similarity threshold is used, and the cosine similarity algorithm is employed to calculate the semantic similarity between the multimodal semantic feature vector and the short-term context data, medium-term context data, and long-term context data, respectively. Extract the historical context data whose semantic similarity is greater than or equal to the semantic similarity threshold, and sort and concatenate the historical context data in descending order of priority to generate a historical context representation. The historical context data in descending order of priority are the short-term context data, the medium-term context data, and the long-term context data. Combine the multimodal semantic feature vector with the historical context representation to generate the semantic cue vector.
[0010] Preferably, the step of inputting the semantic cue vector into the large language model, and using the iterative reasoning strategy to perform contextual understanding and user intent recognition on the semantic cue vector to generate the semantic response content and the action planning instruction specifically includes: The semantic cue vector is input into the large language model, and an initial semantic response sequence is generated through autoregressive iterative inference using a multi-layer self-attention mechanism and a cross-attention mechanism. The initial semantic response sequence is then used to identify the intent using an intent classification head, and the initial intent category is output. confidence level of initial intent The corresponding calculation formula is as follows: In the formula, Indicates the probability distribution of intent categories; Intended classification weight matrix; Indicates the intention to classify the bias vector; A vector representing the initial semantic response sequence; Represents the probability distribution function; A preset intent confidence threshold is set; if the initial intent confidence level is... If the value is less than the intent confidence threshold, the user's intent is determined to be unclear and a clarification and follow-up questioning mechanism is triggered. A clarification and follow-up questioning text is generated and returned to the user's end to wait for the user to supplement the input. If the initial intent confidence level If the user's intent confidence is greater than or equal to the intent confidence threshold, then the user's intent is determined to be clear, and the initial intent category is considered. Perform semantic response sequence on the initial semantic response sequence Figure 1 Consistency check, output the semantic response content; The semantic response content is input into the action planning head. The action planning network is used to extract the action trigger words in the semantic response content and query the action semantic mapping table to determine the gesture type code, gesture intensity parameter and gesture timing parameter corresponding to the action trigger word and summarize them to obtain the action planning instruction.
[0011] Preferably, the step of using a multimodal emotion fusion network to identify the emotional state and emotional intensity of the multimodal semantic feature vector, and to determine the emotional expression strategy corresponding to the semantic response content, specifically includes: The multimodal semantic feature vector is input into the sentiment coding layer, and forward and backward temporal coding are performed based on a bidirectional gated recurrent unit to extract fused sentiment features. The fused emotional features are input into the emotional classification layer. The emotional category probability distribution of the fused emotional features is calculated using the Softmax classification function. The emotional category corresponding to the maximum emotional category probability value is labeled as the emotional state. The maximum emotional category probability value and the second largest emotional category probability value of the emotional category probability distribution are extracted. The difference between the maximum emotional category probability value and the second largest emotional category probability value is calculated to obtain the emotional intensity corresponding to the emotional state. The preset mapping rule library is invoked to map the emotional state and the emotional intensity, generate the speech expression parameters and facial expression driving parameters corresponding to the semantic response content, and summarize and output the emotional expression strategy.
[0012] Preferably, the step of converting the semantic response content, the emotional expression strategy, and the action planning instructions into a multimodal synchronous expression sequence through a multimodal synthesis engine, and driving the virtual digital human to perform multimodal interaction, specifically includes: The semantic response content is input into the speech synthesis engine, and the speech rate, fundamental frequency, volume, and duration are adjusted by a deep neural network according to the speech expression parameters to generate a speech output sequence. The motion planning instruction is input into the motion driving engine, which calls the motion template library and matches the motion template corresponding to the gesture type code. Based on the voice output sequence, the gesture intensity parameter and gesture timing parameter in the motion planning instruction, the motion template is subjected to skeletal trajectory correction and motion interpolation processing to generate a gesture motion sequence. Based on the speech output sequence and the expression driving parameters, the expression driving engine controls the facial key point displacement, mouth corner deformation amplitude, eyebrow raising amplitude and eye opening and closing degree of the virtual digital human to generate an expression animation sequence. Read the timestamp and prosodic feature frames of the speech output sequence. Using the timestamp of the speech output sequence as a reference, use a dynamic time warping algorithm to stretch or compress the gesture sequence and the facial animation sequence on the time axis. Map the gesture sequence, the facial animation sequence and the speech output sequence to a unified time axis to generate the multimodal synchronous expression sequence. The multimodal synchronous expression sequence is input into the real-time rendering engine and drives the virtual digital human to synchronously output speech, actions and expressions.
[0013] Preferably, dynamically updating the historical context database specifically includes: During the interaction between the user and the virtual digital human, the interaction feedback data is recorded and stored in the short-term context database; The short-term context database is maintained in real time using a sliding window mechanism. When the short-term context data reaches the capacity limit of the short-term context database, the short-term context data of the earliest round in the short-term context database is migrated to the medium-term context database. When the intermediate context data reaches the upper limit of the intermediate context database capacity, a hierarchical compression mechanism is used to recursively compress the intermediate context data that exceeds the upper limit of the intermediate context database capacity and store it in the long-term context database.
[0014] A large language model-driven virtual digital human multimodal collaborative interaction system, the system comprising: A multimodal data input module is used to collect multimodal interaction datasets in parallel during user interaction. The multimodal interaction datasets include voice data, text data, gesture data, and facial expression data. The feature extraction and fusion module is used to input the multimodal interaction dataset into a multimodal coding network for feature extraction, obtain speech semantic features, text semantic features, action semantic features and facial expression semantic features, and call a cross-modal feature fusion model to perform semantic alignment and attention fusion to generate a multimodal semantic feature vector; The large language model reasoning module is used to construct semantic prompt vectors based on the multimodal semantic feature vectors and the historical context database and input them into the large language model. It uses an iterative reasoning strategy to perform contextual understanding and user intent recognition on the semantic prompt vectors, and generates semantic response content and corresponding action planning instructions. The multimodal sentiment analysis module is used to identify the sentiment state and sentiment intensity of the multimodal semantic feature vector using a multimodal sentiment fusion network, and to determine the sentiment expression strategy corresponding to the semantic response content; The virtual digital human rendering module is used to convert the semantic response content, the emotion expression strategy and the action planning instructions into a multimodal synchronous expression sequence through a multimodal synthesis engine, and drive the virtual digital human to perform multimodal interaction.
[0015] Compared with existing technologies, the large language model-driven virtual digital human multimodal collaborative interaction method and system provided by this invention have the following beneficial effects: This invention generates multimodal semantic feature vectors through cross-modal semantic alignment and attention fusion mechanisms, solving the problems of semantic misalignment and semantic fragmentation in existing multimodal interaction datasets, and significantly improving the semantic understanding accuracy and anti-interference ability in complex interaction scenarios. This invention constructs semantic prompt vectors by combining multimodal semantic feature vectors with historical context databases, and introduces an iterative reasoning strategy to improve the contextual coherence understanding and dynamic intent update ability of large language models, effectively solving the problems of semantic discontinuity and intent misjudgment.
[0016] This invention identifies users' emotional states and intensity through a multimodal emotion fusion network and generates multimodal synchronous expression sequences based on emotion expression strategies. This solves the problems of mechanical and lacking in hierarchy in existing virtual digital human interaction technologies, achieving dynamic adaptation of emotional expression to users' real-time emotions and improving the empathetic interaction capabilities of virtual digital humans. This invention performs frame-level time alignment and collaborative control on gesture action sequences, facial expression animation sequences, and voice output sequences, achieving synchronous coordination and rhythmic consistency of interactive outputs such as voice, actions, and facial expressions, significantly improving the naturalness and user experience of virtual digital human interaction outputs.
[0017] This invention constructs a three-tiered historical context database (short-term, medium-term, and long-term), employing a sliding window mechanism and a hierarchical compression mechanism to achieve dynamic migration and recursive compression of interactive feedback data. This solves the problems of fixed historical context database capacity and loss of interactive feedback data in existing technologies. Furthermore, this invention filters historical context data through a hierarchical retrieval strategy based on priority levels and semantic similarity thresholds, addressing the issues of blind historical information retrieval and irrelevant information interfering with reasoning in existing technologies. This improves the coherence of multi-turn dialogues and the accuracy of context awareness. Attached Figure Description
[0018] Figure 1 A flowchart of a large language model-driven multimodal collaborative interaction method for virtual digital humans provided in an embodiment of the present invention; Figure 2 A system block diagram of a large language model-driven virtual digital human multimodal collaborative interaction system provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] like Figure 1 The diagram shown is a flowchart of a multimodal collaborative interaction method for virtual digital humans driven by a large language model, provided in an embodiment of the present invention. Figure 1 The execution entity of the method shown can be a software and / or hardware device. The execution entity of this application can include, but is not limited to, at least one of the following: user equipment, network equipment, etc. User equipment can include, but is not limited to, computers, smartphones, personal digital assistants (PDAs), and the aforementioned electronic devices. Network equipment can include, but is not limited to, a single network server, a server group consisting of multiple network servers, or a cloud based on cloud computing consisting of a large number of computers or network servers. Cloud computing is a type of distributed computing, consisting of a super virtual computer composed of a group of loosely coupled computers. This embodiment does not limit this. Steps S1 to S5 are detailed as follows: S1, Parallel acquisition of multimodal interaction datasets during user interaction, the multimodal interaction datasets including voice data, text data, gesture data and facial expression data; A multimodal interaction data acquisition system was built to collect multimodal interaction datasets during user interactions in parallel. User voice signals are acquired in real time via a microphone array, and analog-to-digital conversion is performed based on a preset sampling frequency to generate voice data. The preset sampling frequency covers an effective frequency band of 16kHz-48kHz, capable of completely capturing human voice and emotional prosody information. Voice data is transcribed into text in real time using a speech recognition engine, or text data can be obtained through user terminal input. A depth camera acquires images of the user's limbs, and a human detection algorithm is used to locate the user's body regions, identifying joints such as shoulders, elbows, wrists, and finger joints to generate gesture data. A high-definition RGB camera acquires user facial images, and a face localization algorithm extracts the user's facial image sequence to generate facial expression data, including complete texture information of the user's face.
[0021] It is understandable that multimodal interaction datasets have problems such as inconsistent sampling frequency, inconsistent data dimensions, and timestamp drift during the collection process. Therefore, after the collection is completed, the voice data, text data, gesture data, and facial expression data are timestamp aligned.
[0022] S2, the multimodal interaction dataset is input into a multimodal coding network for feature extraction to obtain speech semantic features, text semantic features, action semantic features and facial expression semantic features, and a cross-modal feature fusion model is called to perform semantic alignment and attention fusion to generate a multimodal semantic feature vector; The step of inputting the multimodal interaction dataset into a multimodal coding network for feature extraction to obtain speech semantic features, text semantic features, action semantic features, and facial expression semantic features specifically includes: The speech data is subjected to sliding frame processing based on a preset time window length and window shift length, and the speech data after sliding frame processing is subjected to spectrum transformation using fast Fourier transform to extract speech spectrum features; the speech spectrum features are input into the speech encoder and local feature extraction and downsampling are performed to generate the speech semantic features; The text data is split into word sequence using the WordPiece word segmentation algorithm and input into the text encoder for feature modeling and nonlinear mapping to generate the text semantic features. The gesture data is used to estimate the pose to obtain a sequence of human joint coordinates. The sequence of human joint coordinates is temporally sampled based on a preset sampling time interval to generate a temporal sampling sequence of human joints and calculate the corresponding motion displacement features and motion velocity features. The sequence of human joint coordinates, the motion displacement features and the motion velocity features are concatenated by channel dimension to generate joint motion trajectory features and input into the motion encoder for graph convolution operation to generate the motion semantic features. The facial expression data is subjected to face region detection, facial key point coordinates are extracted and input into a visual encoder, local expression texture features are extracted using a multi-layer convolutional feature extraction network, global facial structure features are extracted using a spatial pooling layer, and the local expression texture features and the global facial structure features are concatenated by channel dimension to generate the expression semantic features.
[0023] For speech data, a time window of 25ms and a window shift of 10ms are used, and a sliding frame process is applied using a Hamming window to eliminate spectral leakage. A Fast Fourier Transform is then performed on the framed speech data to extract speech spectral features. These features are input into a Transformer-based speech encoder, where local feature extraction is performed using a multi-head self-attention mechanism, followed by progressive downsampling using a one-dimensional convolutional layer to obtain speech semantic features. These semantic features include semantic information, speech rate, intonation, pauses, and other information from the speech data.
[0024] For text data, the WordPiece segmentation algorithm is used to statistically analyze the frequency of word occurrences and the probability of word combination, splitting the text data into word-level word sequences and inputting them into a pre-trained BERT-base text encoder. This text encoder contains 12 Transformer encoding layers, each with 12 attention heads. Each Transformer encoding layer sequentially captures the contextual semantic dependencies of the word sequences through a multi-layer self-attention mechanism, and performs non-linear feature transformation through a feedforward neural network, outputting the contextual semantic representation layer corresponding to the word sequence and performing normalization processing to generate text semantic features.
[0025] For gesture data, the OpenPose pose estimation algorithm is used to detect human joint coordinates, resulting in a sequence of human joint coordinates, each containing three-dimensional spatial coordinates (x, y, z). The human joint coordinate sequence is temporally sampled according to a preset sampling time interval, generating a temporal sampling sequence of human joints. The Euclidean distance between temporal sampling points of human joints at adjacent sampling times is calculated as the motion displacement feature, and the ratio of the motion displacement feature to the preset sampling time interval is calculated as the motion velocity feature. The human joint coordinate sequence, motion displacement feature, and motion velocity feature are concatenated along the channel dimension to generate joint motion trajectory features. This joint motion trajectory feature is input into the motion encoder, which aggregates the spatial relationships and dynamic information between human joints through a spatiotemporal graph convolutional network to generate action semantic features.
[0026] For facial expression data, an MTCNN multi-task cascaded convolutional network is used for face detection and keypoint alignment. Facial keypoint coordinates are extracted in real time and input into a visual encoder. Local expression texture features are extracted through a multi-layer convolutional feature extraction network, and global facial structure features are extracted through a global average spatial pooling layer. The local expression texture features and global facial structure features are then concatenated along the channel dimension to generate semantic features of the expression.
[0027] The step of calling a cross-modal feature fusion model to perform semantic alignment and attention fusion to generate multimodal semantic feature vectors specifically includes: The speech semantic features, text semantic features, action semantic features and facial expression semantic features are unified in dimension by a unified feature mapping layer, and a multimodal feature set is obtained by summarizing them. Based on the multimodal feature set, the cross-modal feature fusion model is invoked to calculate the semantic association weights, and the corresponding calculation formula is as follows: In the formula, Representing semantic features With semantic features Semantic association weights between them; This represents an exponential function with the natural constant e as the base; n represents the number of multimodal feature sets. Construct a cross-modal attention matrix based on the semantic association weights. The multimodal feature set is then semantically aligned to generate a semantically aligned multimodal feature set. An attention fusion strategy is used to perform weighted fusion on the semantically aligned multimodal feature set to generate the multimodal semantic feature vector.
[0028] Because the dimensions of speech semantic features, text semantic features, action semantic features, and facial expression semantic features are inconsistent, cross-modal attention computation is impossible. Therefore, a unified feature mapping layer is used to unify the dimensions. Specifically, a fully connected layer is used to map speech semantic features, text semantic features, action semantic features, and facial expression semantic features to a unified dimension and summarize them to obtain a multimodal feature set, thereby eliminating dimensional differences and achieving feature space matching.
[0029] Based on a multimodal feature set, a cross-modal feature fusion model is invoked to calculate semantic association weights. Taking the calculation of semantic association weights between speech semantic features and text semantic features as an example, firstly, the dot product of speech semantic features and text semantic features is calculated. This dot product reflects the directional consistency of the two types of modal features in the semantic space. The dot product result is input into an exponential function with the natural constant e as the base to obtain a similarity measure. Simultaneously, the dot product of speech semantic features and the multimodal feature set is calculated separately, and the exponent is taken and then summed to obtain a normalized denominator. Based on this dot product and the normalized denominator, the semantic association weights between speech semantic features and text semantic features are calculated. These semantic association weights represent the contribution weight of text semantic features to speech semantic features from the perspective of speech semantic features. For example, if a user says "hello" while waving, the semantic association weight between speech semantic features and text semantic features is relatively high, while the semantic association weight between speech semantic features and action semantic features is lower, indicating that action information plays an auxiliary role in semantic understanding.
[0030] A 4×4 cross-modal attention matrix A is constructed based on the calculated semantic association weights, and the matrix elements are... This represents the attention allocation of the i-th modal feature to the j-th modal feature. The attention matrix A is used to perform semantic alignment on the multimodal feature set. Specifically, for each modal feature, the semantic information of other modal features is fused through a weighted summation method, achieving interactive transmission and complementary enhancement of cross-modal semantic information. For example, when speech semantic features are interfered with by environmental noise, text semantic features and facial expression semantic features can inject supplementary semantic information into the speech semantic features through semantic association weights, improving the robustness of overall semantic understanding.
[0031] An attention fusion strategy is employed to introduce importance weights corresponding to the multimodal feature set. These importance weights are then used to perform weighted fusion of the semantically aligned multimodal feature set, and a gating mechanism dynamically adjusts the contribution ratio of each modality feature within the multimodal feature set. In practical applications, when the interaction scenario is primarily voice-based, the importance weights of voice semantic features and text semantic features are relatively large. When the interaction scenario involves complex gesture commands, the importance weight of action semantic features is automatically increased to highlight action semantic information. The generated multimodal semantic feature vector contains deep semantic association information of four modalities: voice, text, action, and facial expression.
[0032] S3. Based on the multimodal semantic feature vector and the historical context database, construct a semantic prompt vector and input it into the large language model. Use an iterative reasoning strategy to perform contextual understanding and user intent recognition on the semantic prompt vector, and generate semantic response content and corresponding action planning instructions. The construction of semantic cue vectors based on the multimodal semantic feature vectors and the historical context database specifically includes: The historical context database is invoked, which includes a short-term context database, a medium-term context database, and a long-term context database. A preset semantic similarity threshold is used, and the cosine similarity algorithm is employed to calculate the semantic similarity between the multimodal semantic feature vector and the short-term context data, medium-term context data, and long-term context data, respectively. Extract the historical context data whose semantic similarity is greater than or equal to the semantic similarity threshold, and sort and concatenate the historical context data in descending order of priority to generate a historical context representation. The historical context data in descending order of priority are the short-term context data, the medium-term context data, and the long-term context data. Combine the multimodal semantic feature vector with the historical context representation to generate the semantic cue vector.
[0033] Building a three-tiered historical context database supports long-term interactive memory management. The short-term context database uses a circular buffer structure, with its capacity capped at the last 10 rounds of interactive feedback data. The medium-term context database uses a key-value pair storage structure, with its capacity capped at the last 100 rounds of interactive feedback data. A near-nearest neighbor algorithm is used to index the medium-term context data, improving matching efficiency with multimodal semantic feature vectors. The long-term context database uses a distributed file storage structure, storing earlier interactive feedback data to build long-term user profiles across sessions.
[0034] The priority of historical context data, from highest to lowest, is as follows: short-term context data, medium-term context data, and long-term context data. This priority is based on the principle of temporal locality. Recent short-term context data is most relevant to the user's current intent and is given priority in constructing semantic cue vectors; medium-term context data provides session context continuity; and long-term context data can represent user preferences.
[0035] The semantic similarity threshold is calculated offline using a pre-trained validation set, and the semantic similarity between the multimodal semantic feature vectors and historical context data at each level is calculated using the cosine similarity algorithm. The formula for calculating semantic similarity is as follows: In the formula, Represents a multimodal semantic feature vector; This represents the vector corresponding to the historical context data. When the calculated semantic similarity is greater than or equal to the semantic similarity threshold, the historical context data is determined to be semantically relevant to the multimodal semantic feature vector and is extracted; otherwise, the historical context data is filtered to avoid irrelevant information interfering with the reasoning process.
[0036] After extracting historical context data filtered by semantic similarity, the short-term, medium-term, and long-term context data are sorted and concatenated in descending order of priority. It's important to note that during sorting and concatenation, the data is sorted and concatenated in reverse timestamp order, and separators are inserted between each layer of historical context data to allow the large language model to distinguish contextual information at different time scales. The sorted and concatenated historical context representation is then combined with the multimodal semantic feature vector. This combination involves concatenating the historical context representation and the multimodal semantic feature vector and performing dimensionality reduction using a linear projection layer to generate a semantic cue vector. This semantic cue vector incorporates both the multimodal semantic information of the user's current input and highly matched historical interaction memories.
[0037] The step of inputting the semantic cue vector into the large language model, and using the iterative reasoning strategy to perform contextual understanding and user intent recognition on the semantic cue vector to generate the semantic response content and the action planning instruction specifically includes: The semantic cue vector is input into the large language model, and an initial semantic response sequence is generated through autoregressive iterative inference using a multi-layer self-attention mechanism and a cross-attention mechanism. The initial semantic response sequence is then used to identify the intent using an intent classification head, and the initial intent category is output. confidence level of initial intent The corresponding calculation formula is as follows: In the formula, Indicates the probability distribution of intent categories; Intended classification weight matrix; Indicates the intention to classify the bias vector; A vector representing the initial semantic response sequence; Represents the probability distribution function; A preset intent confidence threshold is set; if the initial intent confidence level is... If the value is less than the intent confidence threshold, the user's intent is determined to be unclear and a clarification and follow-up questioning mechanism is triggered. A clarification and follow-up questioning text is generated and returned to the user's end to wait for the user to supplement the input. If the initial intent confidence level If the user's intent confidence is greater than or equal to the intent confidence threshold, then the user's intent is determined to be clear, and the initial intent category is considered. Perform semantic response sequence on the initial semantic response sequence Figure 1 Consistency check, output the semantic response content; The semantic response content is input into the action planning head. The action planning network is used to extract the action trigger words in the semantic response content and query the action semantic mapping table to determine the gesture type code, gesture intensity parameter and gesture timing parameter corresponding to the action trigger word and summarize them to obtain the action planning instruction.
[0038] A large language model with an iterative inference architecture is used as the core inference engine. After the semantic prompt vector is input into the large language model, the semantic prompt vector is mapped to the hidden space through the word embedding layer, and deep semantic inference is performed by multi-layer self-attention mechanism and cross-attention mechanism.
[0039] In the intent recognition stage, the initial semantic response sequence output by the last layer of the large language model's Transformer decoder is extracted. This initial semantic response sequence refers to the semantic information of the entire semantic cue vector. The initial semantic response sequence is input into the intent classification head, where it undergoes a linear transformation using the intent classification weight matrix and intent classification bias vector. Then, it is normalized using the Softmax classification function to obtain the intent category probability distribution. Intent categories include interactive intents such as greetings, inquiries, instructions, confirmations, negations, emotional expressions, and casual conversation.
[0040] The initial intent category is selected based on the maximum probability value in the intent category probability distribution. The initial intent confidence is the maximum sentiment category probability value in the intent category probability distribution, reflecting the large language model's confidence in the intent. If the initial intent confidence is less than the intent confidence threshold, the user's intent is deemed unclear, potentially due to incomplete user input, multimodal signal conflict, or encountering an untrained intent type. In this case, a clarification and follow-up questioning mechanism is triggered, generating clarification and follow-up text and returning it to the user. (For example, "I apologize, I did not fully understand your meaning. Please explain in more detail.") The intent recognition is then restarted after the user provides additional input.
[0041] If the initial intent confidence level is greater than or equal to the intent confidence threshold, the user's intent is determined to be clear, and the initial semantic response sequence is then processed based on the initial intent category. Figure 1 Consistency verification. Specifically, the semantic matching degree between the initial semantic response sequence generated by the large language model and the initial intent category is calculated. The closer the semantic matching degree is to 1, the better the consistency between the initial semantic response sequence and the initial intent category. Figure 1 The initial semantic response sequence is determined as the semantic response content. If the semantic matching degree is closer to 0, it is determined that the initial semantic response sequence is inconsistent with the initial intent category, triggering a response regeneration mechanism to adjust decoding parameters and regenerate the initial semantic response sequence until the intent is correct. Figure 1 Consistency check passed. This iterative reasoning strategy effectively avoids the semantic drift problem between intent recognition and response generation.
[0042] During the action planning phase, the validated semantic response content is input into the action planning head. The action planning head uses an action planning network to segment and tag the semantic response content, extracting action trigger words such as "display," "point," and "welcome." It then queries the action semantic mapping table, which contains the mapping relationships between action trigger words and gesture type codes, gesture intensity parameters, and gesture timing parameters. For example, the action trigger word "welcome" is mapped to the gesture type code WAVE (waving), the gesture intensity parameter 0.8 (relatively large amplitude), and the gesture timing parameter 4 seconds (duration). The determined gesture type codes, gesture intensity parameters, and gesture timing parameters are then summarized to generate a structured action planning instruction, which drives the virtual digital human to perform corresponding limb movements.
[0043] S4, a multimodal emotion fusion network is used to identify the emotional state and emotional intensity of the multimodal semantic feature vector, and to determine the emotional expression strategy corresponding to the semantic response content; The step of employing a multimodal emotion fusion network to identify the emotional state and intensity of the multimodal semantic feature vectors, and determining the emotional expression strategy corresponding to the semantic response content, specifically includes: The multimodal semantic feature vector is input into the sentiment coding layer, and forward and backward temporal coding are performed based on a bidirectional gated recurrent unit to extract fused sentiment features. The fused emotional features are input into the emotional classification layer. The emotional category probability distribution of the fused emotional features is calculated using the Softmax classification function. The emotional category corresponding to the maximum emotional category probability value of the emotional category probability distribution is labeled as the emotional state. The maximum emotional category probability value and the second largest emotional category probability value of the emotional category probability distribution are extracted. The difference between the maximum emotional category probability value and the second largest emotional category probability value is calculated to obtain the emotional intensity corresponding to the emotional state. The preset mapping rule library is invoked to map the emotional state and the emotional intensity, generate the speech expression parameters and facial expression driving parameters corresponding to the semantic response content, and summarize and output the emotional expression strategy.
[0044] Multimodal semantic feature vectors are input into a multimodal sentiment fusion network, which includes a sentiment encoding layer, a sentiment classification layer, and a mapping rule base. The sentiment encoding layer contains a forward temporal encoder and a backward temporal encoder, which operate in parallel using bidirectional gated recurrent units. The forward temporal encoder processes the multimodal semantic feature vectors in forward chronological order to capture the gradual evolution of sentiment states; the backward temporal encoder processes the multimodal semantic feature vectors in reverse chronological order to capture the contextual dependencies of sentiment states. By mining multidimensional emotional cues such as intonation fluctuations in speech semantic features, emotional words in text semantic features, body tension in action semantic features, and facial micro-changes in facial expression semantic features through bidirectional encoding, speech emotional features, text emotional features, action emotional features, and facial expression emotional features are obtained and fused to generate fused emotional features. Among them, speech emotional features include fundamental frequency change rate and energy dynamic range, text emotional features include distribution of emotional polarity words and intensity of negative words, action emotional features include changes in gesture amplitude and head tilt, and facial expression emotional features include the curvature of the corners of the mouth and the depth of the frown lines.
[0045] The fused emotional features are input into the emotional classification layer, which uses the Softmax classification function to calculate the probability distribution of emotional categories. Emotional categories include joy, anger, sadness, fear, surprise, disgust, and neutrality. The emotional category corresponding to the highest probability value in the probability distribution is labeled as the emotional state. Simultaneously, the difference between the highest and second-highest emotional category probability values is calculated to obtain the emotional intensity. The emotional intensity ranges from [0,1]. When the emotional intensity is close to 1, it indicates a clear emotional state and strong, typical user emotional expression; when the emotional intensity is close to 0, it indicates an ambiguous emotional state, and the user may be in a transitional or mixed emotional state.
[0046] The system uses a pre-defined mapping rule library to map emotional states and emotional intensity, generating an emotional expression strategy. The mapping rule library includes an emotional state-voice expression parameter mapping table and an emotional state-facial expression driving parameter mapping table. For example, when the emotional state is joy and the emotional intensity is 0.9 (high-intensity joy), the voice expression parameters are set to a speech rate of 1.2x (relatively fast), a fundamental frequency increase of 15% (high pitch), a volume of 0.9 (louder), and a note length reduction of 10% (light and cheerful); the facial expression driving parameters are set to a mouth corner raise of 0.85 (laughing), an eyebrow raise of 0.6 (raising eyebrows), and an eye opening / closing degree of 0.9 (wide-eyed). When the emotional intensity is 0.4, the voice expression parameters are adjusted to a speech rate of 1.0x (normal), a fundamental frequency increase of 5% (slightly higher pitch), a volume of 0.7 (normal), and the note length remains unchanged; the facial expression driving parameters are adjusted to a mouth corner raise of 0.4 (smiling), an eyebrow raise of 0.2 (slightly raised), and an eye opening / closing degree of 0.7 (normal). By precisely quantifying the intensity of emotions, we can achieve dynamic adaptation between emotional expression strategies and users' real-time emotions.
[0047] S5, the semantic response content, the emotion expression strategy and the action planning instructions are converted into a multimodal synchronous expression sequence through a multimodal synthesis engine, and the virtual digital human is driven to perform multimodal interaction.
[0048] The process of converting the semantic response content, the emotional expression strategy, and the action planning instructions into a multimodal synchronous expression sequence using a multimodal synthesis engine, and driving the virtual digital human to perform multimodal interaction, specifically includes: The semantic response content is input into the speech synthesis engine, and the speech rate, fundamental frequency, volume, and duration are adjusted by a deep neural network according to the speech expression parameters to generate a speech output sequence. The motion planning instruction is input into the motion driving engine, which calls the motion template library and matches the motion template corresponding to the gesture type code. Based on the voice output sequence, the gesture intensity parameter and gesture timing parameter in the motion planning instruction, the motion template is subjected to skeletal trajectory correction and motion interpolation processing to generate a gesture motion sequence. Based on the speech output sequence and the expression driving parameters, the expression driving engine controls the facial key point displacement, mouth corner deformation amplitude, eyebrow raising amplitude and eye opening and closing degree of the virtual digital human to generate an expression animation sequence. Read the timestamp and prosodic feature frames of the speech output sequence. Using the timestamp of the speech output sequence as a reference, use a dynamic time warping algorithm to stretch or compress the gesture sequence and the facial animation sequence on the time axis. Map the gesture sequence, the facial animation sequence and the speech output sequence to a unified time axis to generate the multimodal synchronous expression sequence. The multimodal synchronous expression sequence is input into the real-time rendering engine and drives the virtual digital human to synchronously output speech, actions and expressions.
[0049] During speech synthesis, speech expression parameters in the emotion expression strategy are adjusted in real time. The speech rate is changed by adjusting the weight distribution of the attention mechanism, the fundamental frequency is adjusted by modifying the frequency axis scaling factor of the Mel spectrogram, the volume is adjusted by gain control, and the duration is controlled by the scaling factor of the duration predictor, generating a speech output sequence containing timestamps and prosodic feature frames.
[0050] In the motion-driven phase, the corresponding motion template is matched according to the gesture type encoding. For example, WAVE encoding corresponds to a waving animation template. The gesture intensity parameter is multiplied by the motion amplitude in the motion template to correct the skeletal trajectory of the motion template. The motion speed in the motion template is scaled and adjusted based on the gesture timing parameters. Bezier curves are used for motion interpolation to smooth the transition between keyframes and generate a gesture motion sequence.
[0051] In the expression-driven phase, based on the speech output sequence and expression-driven parameters, the expression-driven engine controls the facial channels of the virtual digital human, including raising or lowering the corners of the mouth to correspond to joy or sadness; raising or furrowing the eyebrows to correspond to surprise or anger; and widening or squinting the eyes to correspond to surprise or disgust. Based on the amplitude of mouth corner deformation, eyebrow raising, and eye opening / closing in the expression-driven parameters, the weight coefficients of each facial channel are calculated in real time to generate an expression animation sequence.
[0052] A dynamic time warping algorithm is used to align the timelines of gesture sequences and facial expression animation sequences. The algorithm constructs a cumulative distance matrix and finds the optimal matching path between the gesture sequence and prosodic feature frames. Based on the optimal matching path, the transition time is dynamically stretched or compressed. When the duration of the gesture sequence is long, but the duration of the corresponding prosodic feature frames is short (i.e., multiple prosodic feature frames are mapped to a single or a few consecutive gesture keyframes in the gesture sequence), the transition time between the gesture keyframes in the gesture sequence is compressed to accelerate the completion of the gesture. When the duration of the gesture sequence is short, but the duration of the corresponding prosodic feature frames is long (i.e., multiple consecutive gesture keyframes in the gesture sequence are mapped to a single or a few consecutive prosodic feature frames in the speech output sequence), the transition time between the gesture keyframes in the gesture sequence is stretched to ensure that the start and end times of the gesture are accurately aligned with the stress and pauses of the speech. Simultaneously, the facial expression animation sequence is time-axis normalized. If the peak expression moment lags behind the prosodic feature frame, the transition time before the peak expression moment is compressed, thus advancing the peak expression moment. Conversely, if the peak expression moment precedes the prosodic feature frame, the transition time before the peak expression moment is stretched to ensure a delay, thereby synchronizing the peak expression moment with the prosodic feature frame. The keyframes of the speech output sequence, gesture sequence, and facial expression animation sequence are then fully aligned on the time axis to obtain a multimodal synchronized expression sequence.
[0053] The multimodal synchronous expression sequence is input into the real-time rendering engine and voice, actions and expressions are output synchronously to achieve natural multimodal interaction between users and virtual digital humans.
[0054] Dynamically updating the historical context database specifically includes: During the interaction between the user and the virtual digital human, the interaction feedback data is recorded and stored in the short-term context database; The short-term context database is maintained in real time using a sliding window mechanism. When the short-term context data reaches the capacity limit of the short-term context database, the short-term context data of the earliest round in the short-term context database is migrated to the medium-term context database. When the intermediate context data reaches the upper limit of the intermediate context database capacity, a hierarchical compression mechanism is used to recursively compress the intermediate context data that exceeds the upper limit of the intermediate context database capacity and store it in the long-term context database.
[0055] Understandably, after each round of interaction between the user and the virtual digital human, the interaction feedback data is automatically recorded and stored in a short-term context database. The interaction feedback data includes multimodal semantic feature vectors, semantic response content, emotional state, interaction content, and interaction timestamps.
[0056] It should be noted that a preset sliding window size is set. When the most recent interaction feedback data arrives in the short-term context database, if the current short-term context data volume has not reached the capacity limit, it is added for storage; if the current short-term context data volume has reached the capacity limit, the short-term context data of the earliest turn is removed and migrated to the medium-term context database. Medium-term context data exceeding the capacity limit of the medium-term context database is recursively compressed to ensure reduced storage overhead while preserving core semantic information. This achieves efficient storage and intelligent retrieval of historical context data, avoids the loss of historical information, helps large language models understand users' long-term preferences and behavioral patterns, and significantly improves the coherence and personalization of multi-turn dialogues.
[0057] like Figure 2 The diagram shown is a system block diagram of a large language model-driven virtual digital human multimodal collaborative interaction system provided in an embodiment of the present invention. The system includes: A multimodal data input module is used to collect multimodal interaction datasets in parallel during user interaction. The multimodal interaction datasets include voice data, text data, gesture data, and facial expression data. The feature extraction and fusion module is used to input the multimodal interaction dataset into a multimodal coding network for feature extraction, obtain speech semantic features, text semantic features, action semantic features and facial expression semantic features, and call a cross-modal feature fusion model to perform semantic alignment and attention fusion to generate a multimodal semantic feature vector; The large language model reasoning module is used to construct semantic prompt vectors based on the multimodal semantic feature vectors and the historical context database and input them into the large language model. It uses an iterative reasoning strategy to perform contextual understanding and user intent recognition on the semantic prompt vectors, and generates semantic response content and corresponding action planning instructions. The multimodal sentiment analysis module is used to identify the sentiment state and sentiment intensity of the multimodal semantic feature vector using a multimodal sentiment fusion network, and to determine the sentiment expression strategy corresponding to the semantic response content; The virtual digital human rendering module is used to convert the semantic response content, the emotion expression strategy and the action planning instructions into a multimodal synchronous expression sequence through a multimodal synthesis engine, and drive the virtual digital human to perform multimodal interaction.
[0058] Figure 2 The system of the illustrated embodiment can be used to perform corresponding operations. Figure 1 The steps in the method embodiments shown are implemented in a similar manner and have similar technical effects, and will not be repeated here.
[0059] An electronic device includes a memory and a processor, wherein the memory stores a computer program, and when the processor runs the computer program stored in the memory, the processor performs the steps of the large language model-driven virtual digital human multimodal collaborative interaction method as described in any of the above.
[0060] like Figure 3 The diagram shown is a hardware structure schematic of an electronic device according to an embodiment of the present invention. The electronic device 30 includes: a processor 31, a memory 32, and a computer program; wherein... The memory 32 is used to store the computer program, and the memory may also be flash memory. The computer program is, for example, an application program or functional module that implements the above method.
[0061] Processor 31 is configured to execute the computer program stored in the memory to implement the various steps performed by the device in the above method. For details, please refer to the relevant descriptions in the preceding method embodiments.
[0062] Alternatively, the memory 32 can be either standalone or integrated with the processor 31.
[0063] When the memory 32 is a device independent of the processor 31, the device may further include: Bus 33 is used to connect the memory 32 and the processor 31.
[0064] A readable storage medium storing a computer program, which, when executed by a processor, is used to implement the steps of the large language model-driven virtual digital human multimodal collaborative interaction method as described above.
[0065] The readable storage medium can be a computer storage medium or a communication medium. A communication medium includes any medium that facilitates the transfer of computer programs from one location to another. A computer storage medium can be any available medium accessible to a general-purpose or special-purpose computer. For example, a readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application-Specific Integrated Circuit (ASIC). Alternatively, the ASIC can be located in a user equipment. Of course, the processor and the readable storage medium can also exist as discrete components in a communication device. The readable storage medium can be a read-only memory (ROM), random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0066] The present invention also provides a program product including executable instructions stored in a readable storage medium. At least one processor of the device can read the executable instructions from the readable storage medium, and the at least one processor executes the executable instructions to cause the device to implement the methods provided in the various embodiments described above.
[0067] In the embodiments of the above-described device, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly manifested as execution by a hardware processor, or execution by a combination of hardware and software modules within the processor.
[0068] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for multi-modal collaborative interaction of a large language model driven virtual digital human, characterized in that, The method includes: A multimodal interaction dataset is collected in parallel during the user interaction process. The multimodal interaction dataset includes voice data, text data, gesture data, and facial expression data. The multimodal interaction dataset is input into a multimodal coding network for feature extraction to obtain speech semantic features, text semantic features, action semantic features and facial expression semantic features. A cross-modal feature fusion model is then called to perform semantic alignment and attention fusion to generate a multimodal semantic feature vector. Based on the multimodal semantic feature vector and the historical context database, a semantic prompt vector is constructed and input into the large language model. An iterative reasoning strategy is used to perform contextual understanding and user intent recognition on the semantic prompt vector to generate semantic response content and corresponding action planning instructions. A multimodal emotion fusion network is used to identify the emotional state and emotional intensity of the multimodal semantic feature vectors, and to determine the emotional expression strategy corresponding to the semantic response content; The semantic response content, the emotion expression strategy, and the action planning instructions are converted into a multimodal synchronous expression sequence through a multimodal synthesis engine, and the virtual digital human is driven to perform multimodal interaction.
2. The method of claim 1, wherein the method further comprises: The step of inputting the multimodal interaction dataset into a multimodal coding network for feature extraction to obtain speech semantic features, text semantic features, action semantic features, and facial expression semantic features specifically includes: The speech data is subjected to sliding frame processing based on a preset time window length and window shift length, and the speech data after sliding frame processing is subjected to spectrum transformation using fast Fourier transform to extract speech spectrum features; the speech spectrum features are input into the speech encoder and local feature extraction and downsampling are performed to generate the speech semantic features; The text data is split into word sequence using the WordPiece word segmentation algorithm and input into the text encoder for feature modeling and nonlinear mapping to generate the text semantic features. The gesture data is used to estimate the pose to obtain a sequence of human joint coordinates. The sequence of human joint coordinates is temporally sampled based on a preset sampling time interval to generate a temporal sampling sequence of human joints and calculate the corresponding motion displacement features and motion velocity features. The sequence of human joint coordinates, the motion displacement features and the motion velocity features are concatenated by channel dimension to generate joint motion trajectory features and input into the motion encoder for graph convolution operation to generate the motion semantic features. The facial expression data is subjected to face region detection, facial key point coordinates are extracted and input into a visual encoder, local expression texture features are extracted using a multi-layer convolutional feature extraction network, global facial structure features are extracted using a spatial pooling layer, and the local expression texture features and the global facial structure features are concatenated by channel dimension to generate the expression semantic features.
3. The method of claim 1, wherein, The step of calling a cross-modal feature fusion model to perform semantic alignment and attention fusion to generate multimodal semantic feature vectors specifically includes: The speech semantic features, text semantic features, action semantic features and facial expression semantic features are unified in dimension by a unified feature mapping layer, and a multimodal feature set is obtained by summarizing them. Based on the multimodal feature set, the cross-modal feature fusion model is invoked to calculate the semantic association weights, and the corresponding calculation formula is as follows: wherein, denotes semantic features denotes semantic association weights between semantic features denotes semantic association weights between semantic features denotes an exponential function with the natural constant e as base; n denotes the number of multi-modal feature sets. constructing a cross-modal attention matrix based on the semantic correlation weight and performing semantic alignment on the multi-modal feature set to generate a semantic-aligned multi-modal feature set An attention fusion strategy is used to perform weighted fusion on the semantically aligned multimodal feature set to generate the multimodal semantic feature vector.
4. The method of claim 1, wherein, The construction of semantic cue vectors based on the multimodal semantic feature vectors and the historical context database specifically includes: The historical context database is invoked, which includes a short-term context database, a medium-term context database, and a long-term context database. A preset semantic similarity threshold is used, and the cosine similarity algorithm is employed to calculate the semantic similarity between the multimodal semantic feature vector and the short-term context data, medium-term context data, and long-term context data, respectively. Extract the historical context data whose semantic similarity is greater than or equal to the semantic similarity threshold, and sort and concatenate the historical context data in descending order of priority to generate a historical context representation. The historical context data in descending order of priority are the short-term context data, the medium-term context data, and the long-term context data. Combine the multimodal semantic feature vector with the historical context representation to generate the semantic cue vector.
5. The large language model driven virtual digital human multi-modal collaborative interaction method according to claim 1, characterized in that, The step of inputting the semantic cue vector into the large language model, and using the iterative reasoning strategy to perform contextual understanding and user intent recognition on the semantic cue vector to generate the semantic response content and the action planning instruction specifically includes: input the semantic prompt vector into the large language model, and generate an initial semantic response sequence through self-attention mechanism and cross-attention mechanism by self-recurrent iterative reasoning, and use an intent classification head to identify the initial semantic response sequence, and output an initial intent category corresponding to the initial intent confidence The corresponding calculation formula is as follows: In the formula, Indicates the probability distribution of intent categories; Intended classification weight matrix; Indicates the intention to classify the bias vector; A vector representing the initial semantic response sequence; Represents the probability distribution function; A preset intent confidence threshold is set; if the initial intent confidence level is... If the value is less than the intent confidence threshold, the user's intent is determined to be unclear and a clarification and follow-up questioning mechanism is triggered. A clarification and follow-up questioning text is generated and returned to the user's end to wait for the user to supplement the input. If the initial intent confidence level If the user's intent confidence is greater than or equal to the intent confidence threshold, then the user's intent is determined to be clear, and the initial intent category is considered. Perform intent consistency verification on the initial semantic response sequence and output the semantic response content; The semantic response content is input into the action planning head. The action planning network is used to extract the action trigger words in the semantic response content and query the action semantic mapping table to determine the gesture type code, gesture intensity parameter and gesture timing parameter corresponding to the action trigger word and summarize them to obtain the action planning instruction.
6. The method for multimodal collaborative interaction of virtual digital humans driven by a large language model according to claim 1, characterized in that, The step of employing a multimodal emotion fusion network to identify the emotional state and intensity of the multimodal semantic feature vectors, and determining the emotional expression strategy corresponding to the semantic response content, specifically includes: The multimodal semantic feature vector is input into the sentiment coding layer, and forward and backward temporal coding are performed based on a bidirectional gated recurrent unit to extract fused sentiment features. The fused emotional features are input into the emotional classification layer. The emotional category probability distribution of the fused emotional features is calculated using the Softmax classification function. The emotional category corresponding to the maximum emotional category probability value is labeled as the emotional state. The maximum emotional category probability value and the second largest emotional category probability value of the emotional category probability distribution are extracted. The difference between the maximum emotional category probability value and the second largest emotional category probability value is calculated to obtain the emotional intensity corresponding to the emotional state. The preset mapping rule library is invoked to map the emotional state and the emotional intensity, generate the speech expression parameters and facial expression driving parameters corresponding to the semantic response content, and summarize and output the emotional expression strategy.
7. The method for multimodal collaborative interaction of virtual digital humans driven by a large language model according to claim 6, characterized in that, The process of converting the semantic response content, the emotional expression strategy, and the action planning instructions into a multimodal synchronous expression sequence using a multimodal synthesis engine, and driving the virtual digital human to perform multimodal interaction, specifically includes: The semantic response content is input into the speech synthesis engine, and the speech rate, fundamental frequency, volume, and duration are adjusted by a deep neural network according to the speech expression parameters to generate a speech output sequence. The motion planning instruction is input into the motion driving engine, which calls the motion template library and matches the motion template corresponding to the gesture type code. Based on the voice output sequence, the gesture intensity parameter and gesture timing parameter in the motion planning instruction, the motion template is subjected to skeletal trajectory correction and motion interpolation processing to generate a gesture motion sequence. Based on the speech output sequence and the expression driving parameters, the expression driving engine controls the facial key point displacement, mouth corner deformation amplitude, eyebrow raising amplitude and eye opening and closing degree of the virtual digital human to generate an expression animation sequence. Read the timestamp and prosodic feature frames of the speech output sequence. Using the timestamp of the speech output sequence as a reference, use a dynamic time warping algorithm to stretch or compress the gesture sequence and the facial animation sequence on the time axis. Map the gesture sequence, the facial animation sequence and the speech output sequence to a unified time axis to generate the multimodal synchronous expression sequence. The multimodal synchronous expression sequence is input into the real-time rendering engine and drives the virtual digital human to synchronously output speech, actions and expressions.
8. The method for multimodal collaborative interaction of virtual digital humans driven by a large language model according to claim 4, characterized in that, Dynamically updating the historical context database specifically includes: During the interaction between the user and the virtual digital human, the interaction feedback data is recorded and stored in the short-term context database; The short-term context database is maintained in real time using a sliding window mechanism. When the short-term context data reaches the capacity limit of the short-term context database, the short-term context data of the earliest round in the short-term context database is migrated to the medium-term context database. When the intermediate context data reaches the upper limit of the intermediate context database capacity, a hierarchical compression mechanism is used to recursively compress the intermediate context data that exceeds the upper limit of the intermediate context database capacity and store it in the long-term context database.
9. A virtual digital human multimodal collaborative interaction system driven by a large language model, characterized in that: The system, applied to the large language model-driven multimodal collaborative interaction method for virtual digital humans as described in any one of claims 1-8, comprises: A multimodal data input module is used to collect multimodal interaction datasets in parallel during user interaction. The multimodal interaction datasets include voice data, text data, gesture data, and facial expression data. The feature extraction and fusion module is used to input the multimodal interaction dataset into a multimodal coding network for feature extraction, obtain speech semantic features, text semantic features, action semantic features and facial expression semantic features, and call a cross-modal feature fusion model to perform semantic alignment and attention fusion to generate a multimodal semantic feature vector; The large language model reasoning module is used to construct semantic prompt vectors based on the multimodal semantic feature vectors and the historical context database and input them into the large language model. It uses an iterative reasoning strategy to perform contextual understanding and user intent recognition on the semantic prompt vectors, and generates semantic response content and corresponding action planning instructions. The multimodal sentiment analysis module is used to identify the sentiment state and sentiment intensity of the multimodal semantic feature vector using a multimodal sentiment fusion network, and to determine the sentiment expression strategy corresponding to the semantic response content; The virtual digital human rendering module is used to convert the semantic response content, the emotion expression strategy and the action planning instructions into a multimodal synchronous expression sequence through a multimodal synthesis engine, and drive the virtual digital human to perform multimodal interaction.