A Digital Human Interaction System and Method Based on Multimodal Perception AI Agent

By employing a multimodal perception and encoding module, a modal graph construction and path optimization module, a fusion expression generation module, and a policy response generation module, the problems of information redundancy and rigid behavior in digital human interaction are solved, achieving natural, low-latency, and privacy-preserving multimodal interaction.

CN120408125BActive Publication Date: 2025-11-14BEIJING INFINITE SMART TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510886441.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-11-14
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

Existing digital human interaction technologies suffer from a lack of selection and dynamic control mechanisms in the multimodal feature fusion process, leading to information redundancy and the masking of key features; the behavioral decision-making process has a simple response mechanism, lacks strategic flexibility, and is difficult to balance between low latency and privacy protection.

Method used

Employing a multimodal perception and encoding module, a modality graph construction and path optimization module, a fusion expression generation module, and a policy response generation module, this system dynamically selects the modality combination and behavioral decision with the most information-rich content through information entropy-driven modality fusion path optimization and the Softmax policy function. Combined with an edge-cloud collaborative strategy, it achieves efficient processing and natural interaction of multimodal information.

Benefits of technology

It achieves automatic selection of modal features and the most comprehensive combination of information, dynamic response to action signals and adjustable distribution, natural and smooth digital human interaction, low latency and privacy protection capabilities, strong adaptability, and excellent retention of key modal information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120408125B_ABST
    Figure CN120408125B_ABST
Patent Text Reader

Abstract

This application relates to the field of artificial intelligence technology and discloses an AI intelligent agent digital human interaction system and method based on multimodal perception. The system includes: a multimodal perception and encoding module, a modality graph construction and path optimization module, a fusion representation generation module, a policy response generation module, and a digital human driving module. The method includes: preprocessing the collected perception information to obtain perception features; performing feature encoding to obtain multiple modal feature vectors; constructing a modality information graph and determining a set of modality feature indices; selecting corresponding modality features for fusion to generate a compressed representation; generating interaction response action signals; receiving and parsing the interaction response action signals; and realizing multimodal interaction based on the parsed control parameters. This invention achieves the effect of automatically selecting the modality combination with the most information content by constructing a modality information graph and dynamically selecting the optimal feature index, solving the problems of high modality redundancy and easy dilution of effective features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, specifically to an AI intelligent agent digital human interaction system and method based on multimodal perception. Background Technology

[0002] In recent years, with the development of technologies such as artificial intelligence, speech recognition, image understanding, and 3D modeling, digital humans are gradually moving from concept to application. Especially in education, services, and healthcare, digital humans are being given functional roles such as "contextual response," "knowledge transfer," and "companionship and guidance." Taking education as an example, students exhibit a variety of sensory signals during the learning process, including language, actions, and emotions; text or voice alone is far from sufficient to cover the expressive dimensions required for interaction.

[0003] Existing digital human interaction technologies typically employ an "encoding-concatenation-decision" process for processing multimodal input. The system independently encodes modal information such as speech and images, then concatenates or averages them before feeding them into the decision network. This type of architecture is already used in voice assistants, virtual customer service, and some voice-guided learning products. For example, Google's Multimodal Transformer uses multimodal channels for feature concatenation, and Microsoft's affective computing framework also attempts to incorporate speech emotion information to modulate the output.

[0004] However, existing digital human interaction technologies lack selection and dynamic control mechanisms in their multimodal feature fusion process. A one-size-fits-all approach can easily lead to information redundancy and the obscuring of key features. Furthermore, the behavioral decision-making process has a simple response mechanism, often triggering actions solely based on maximum confidence, lacking the "strategic flexibility" required in real-world scenarios, resulting in stiff and unnatural digital human performance. Secondly, current systems mostly rely on single-terminal computing or cloud models, lacking effective edge-cloud collaboration strategies and struggling to achieve a balance between low latency and privacy protection. Therefore, this invention provides an AI intelligent agent digital human interaction system and method based on multimodal perception to address the shortcomings of existing technologies. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides an AI-powered intelligent agent digital human interaction system and method based on multimodal perception, which solves the problems of uncontrollable feature redundancy, unnatural interaction actions, high difficulty in terminal deployment, and lack of security guarantees in cross-institutional modeling in existing multimodal interactions.

[0006] To achieve the above objectives, the present invention provides the following technical solution: an AI intelligent agent digital human interaction system based on multimodal perception, comprising:

[0007] The multimodal perception and encoding module is used to collect user perception information and preprocess and encode the collected user perception information to obtain multiple modal feature vectors.

[0008] The modality graph construction and path optimization module is used to construct a modality information graph based on the modality feature vector, determine the modality fusion path with the minimum information entropy, and output the modality feature index set corresponding to the modality fusion path;

[0009] The fusion representation generation module is used to select corresponding modal features from the modal feature vectors according to the modal feature index set, fuse them, and generate a compressed representation.

[0010] The strategy response generation module is used to input the fused expression into the behavior policy network and generate interactive response action signals through the policy optimization mechanism;

[0011] The digital human driving module is used to control the digital human to perform voice, facial expression or action interaction output according to the response action.

[0012] Preferably, the multimodal sensing and coding module includes:

[0013] The sensory information acquisition unit is used to collect user sensory information, including voice signals, image information, tactile pressure signals, physiological signals, and environmental sensory data.

[0014] The modal preprocessing unit is used to perform noise reduction, normalization, temporal alignment and missing information processing on the perceived information.

[0015] The feature encoding unit is used to extract features from the preprocessed perceptual information using a modal-specific encoder, and output a modal feature vector of uniform dimension.

[0016] Preferably, the modality graph construction and path optimization module includes:

[0017] The modal information graph construction unit is used to construct a modal information graph based on the multiple modal feature vectors, where nodes in the graph represent different modal features and edges represent the relationships between modal features;

[0018] The path optimization unit is used to determine the path of modal fusion based on the modal information graph through an information entropy evaluation mechanism, and output the optimal set of modal feature indices.

[0019] Preferably, the path optimization unit performs modal fusion path optimization according to the following formula:

[0020] ;

[0021] in, , , These are the weighting coefficients for speech, emotion, and context modalities, respectively. This is the final optimized modality fusion feature vector; , , These are feature vectors for speech, emotion, and context modalities, respectively.

[0022] Preferably, the fusion expression generation module includes:

[0023] A modal feature selection unit is used to select a corresponding modal feature from multiple modal feature vectors based on the modal feature index set;

[0024] The modal feature fusion unit is used to fuse the selected modal features to generate a modal fusion feature vector with a compressed representation.

[0025] Preferably, the modal feature fusion unit performs modal feature fusion processing according to the following formula:

[0026] ;

[0027] in, The fused modal features; For the first Modal feature vectors; For the first Weight coefficients for each modal feature; This represents the total number of modal features.

[0028] Preferably, the strategy response generation module includes:

[0029] The semantic parsing unit is used to perform semantic decomposition on the fused representation and extract semantic tags and context states;

[0030] The behavior policy network unit is used to input the semantic tags and context state into the policy network and output the probability distribution of interactive response action signals.

[0031] Preferably, the behavior policy network unit optimizes the output of the response action based on the following policy function:

[0032] ;

[0033] in, Indicates the state Take action below The probability of the strategy; For action value functions; Temperature coefficient; The normalization term for all candidate actions; is the base of the natural logarithm.

[0034] Preferably, the digital human driving module includes:

[0035] The signal parsing unit is used to receive the interactive response action signal and parse it into voice control parameters, facial expression control parameters and body movement control parameters.

[0036] The output execution unit is used to drive the speech synthesizer, expression driver and motion actuator respectively according to the control parameters to realize the multimodal interactive output of the digital human.

[0037] It also provides a method for AI-powered digital human interaction based on multimodal perception, including the following steps:

[0038] Collect multimodal perception information from users, preprocess the perception information, and obtain multiple preprocessed perception features;

[0039] The preprocessed perceptual features are encoded to obtain multiple modal feature vectors;

[0040] A modal information graph is constructed based on the modal feature vectors, and the optimal modal fusion path is determined by the information entropy minimization method. The set of modal feature indices corresponding to the modal fusion path is then output.

[0041] Based on the modal feature index set, the corresponding modal features are selected from the modal feature vectors and fused to generate a compressed representation;

[0042] The compressed representation is input into the behavior policy network, and interactive response action signals are generated through the policy optimization mechanism;

[0043] Receive interactive response action signals and parse them into voice control parameters, facial expression control parameters, and body movement control parameters;

[0044] The speech synthesizer, expression driver, and motion actuator are driven by the parsed control parameters to control the digital human to perform corresponding speech, expression, and motion outputs, thereby realizing multimodal interaction.

[0045] This invention provides an AI-powered digital human interaction system and method based on multimodal perception. It offers the following advantages:

[0046] 1. This invention employs an "information entropy-driven modal fusion path optimization" scheme. By constructing a modal information graph and dynamically selecting the optimal feature index, it achieves the effect of automatically selecting the modal combination with the most sufficient information. Compared with the static splicing or average fusion methods commonly used in existing technologies, it solves the problems of high modal redundancy and easy dilution of effective features.

[0047] 2. This invention introduces a behavior decision-making mechanism based on the Softmax policy function, enabling the selection of response action signals to be dynamic and adjustable in distribution. This design is more flexible and produces a more natural response when facing complex states and multiple action options. Existing solutions generally use threshold decisions or fixed mappings, which can easily lead to rigid behavior and a lack of diversity. This solution effectively avoids these limitations.

[0048] 3. This invention utilizes a three-channel parsing output method encompassing voice, facial expressions, and motion to decompose the strategy response signal into controllable driving parameters, thereby achieving overall unity in voice synchronization, facial expression naturalness, and motion continuity for the digital human. Compared to traditional solutions relying on a single audio or motion model, it eliminates output disconnection or "sluggish response," resulting in a natural and smooth interactive experience.

[0049] 4. This invention employs a weighted fusion calculation formula during the modal feature fusion process, allowing for flexible adjustment of weights for different modal features based on scene importance. Compared to the one-size-fits-all fusion mode in existing technologies, this method is more adaptable and better preserves key modal information, exhibiting excellent stability, especially in multi-user and non-ideal environments. Attached Figure Description

[0050] Figure 1 This is a system architecture diagram of the present invention;

[0051] Figure 2 This is a schematic diagram of the multimodal sensing and encoding module of the present invention;

[0052] Figure 3 This is a schematic diagram of the modality graph construction and path optimization module of the present invention;

[0053] Figure 4 This is a schematic diagram of the fusion expression generation module of the present invention;

[0054] Figure 5 This is a schematic diagram of the strategy response generation module of the present invention;

[0055] Figure 6 This is a schematic diagram of the digital human driving module of the present invention;

[0056] Figure 7 This is a flowchart of the method steps of the present invention. Detailed Implementation

[0057] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0058] Please see the appendix Figure 1 -Appendix Figure 6 This invention provides an AI-powered digital human interaction system based on multimodal perception, comprising:

[0059] The multimodal perception and encoding module is used to collect user perception information and preprocess and encode the collected user perception information to obtain multiple modal feature vectors.

[0060] The modality graph construction and path optimization module is used to construct a modality information graph based on the modality feature vector, determine the modality fusion path with the minimum information entropy, and output the modality feature index set corresponding to the modality fusion path;

[0061] The fusion representation generation module is used to select corresponding modal features from the modal feature vectors according to the modal feature index set, fuse them, and generate a compressed representation.

[0062] The strategy response generation module is used to input the fused expression into the behavior policy network and generate interactive response action signals through the policy optimization mechanism;

[0063] The digital human driving module is used to control the digital human to perform voice, facial expression or action interaction output according to the response action.

[0064] In this embodiment, the multimodal perception and encoding module is responsible for collecting and processing multimodal perception information from the user. This multimodal information includes voice signals, image information, tactile pressure signals, physiological signals, and environmental perception data. This information forms the basis for the system's interaction with the user; through effective processing of this information, the system can better understand and respond to the user's needs.

[0065] By transforming multimodal information into a unified feature vector, the integrity and accuracy of the information are ensured, so that subsequent modules can effectively utilize this data.

[0066] Generally, the processing flow of a multimodal sensing and coding module can be divided into three sub-modules: a sensing information acquisition unit, a modal preprocessing unit, and a feature coding unit. The specific functions of each sub-module are as follows:

[0067] In this embodiment, the sensing information acquisition unit is responsible for acquiring multimodal information from different sensors. These sensors include, but are not limited to, voice acquisition devices (such as microphones), visual acquisition devices (such as cameras), tactile sensors (such as pressure sensors), physiological signal acquisition devices (such as electrocardiogram or skin conductance response sensors), and environmental sensing devices (such as temperature sensors, humidity sensors, etc.). Signals from different modalities are acquired in their own specific formats, providing basic data for subsequent processing.

[0068] For example, voice signals can be captured using a microphone, while image signals are collected using a camera. Tactile pressure signals can be acquired through gloves or sensors embedded in clothing, and physiological signals are collected through wearable devices. Environmental perception data, such as temperature, humidity, and air pressure, is obtained from environmental sensors. All of this information provides the raw data source for subsequent analysis of user intent.

[0069] In this embodiment, the modal preprocessing unit is responsible for preprocessing the acquired sensing information. Specifically, this unit is responsible for processing various types of sensing data, including noise reduction, normalization, temporal alignment, and missing data repair.

[0070] In practical applications, noise reduction is particularly important because perceived information can be interfered with by noise, especially in image or speech signals. Noise reduction can be achieved through techniques such as filters or neural networks. Normalization, on the other hand, ensures that the data scale is consistent across different modalities, preventing certain modalities from having excessively large or small data ranges that could negatively impact subsequent processing.

[0071] Temporal alignment is an operation performed on data with temporal characteristics. Data from different modalities may have temporal inconsistencies; temporal alignment ensures that all modal information is aligned at the same point in time, facilitating subsequent feature fusion. Missing data repair addresses data loss caused by acquisition issues, ensuring the system can complete data processing without losing critical information.

[0072] The role of the feature encoding unit is to transform preprocessed perceptual information into modal feature vectors. Each modal information is processed by its own dedicated encoder for feature extraction. Specifically, speech signals can be used to extract their feature vectors using speech recognition models, image data can be processed by models such as convolutional neural networks (CNNs) to extract visual features, and tactile and physiological signals are encoded using specialized sensors and algorithms.

[0073] In this embodiment, the encoded feature vectors are processed by a feature encoding unit, and the feature vectors of all modalities are eventually transformed into vectors of a uniform dimension. The purpose of this step is to ensure that information from different modalities can be effectively fused and used in subsequent processing.

[0074] Specifically, modal feature vectors, after feature encoding, typically have a fixed dimension so that the system can combine and compare features from different modalities. For example, an image feature vector might be a vector with hundreds of dimensions, while a speech feature vector might be a vector with tens of dimensions. In this process, by using deep learning models (such as autoencoders, convolutional neural networks, etc.), the features of each modality are mapped to a unified feature space, thereby ensuring that they have a consistent representation.

[0075] In this embodiment, the modal feature vector after feature encoding This is a multi-dimensional vector representing the information features of each modality. After the above processing, the system can convert the information of each modality into a feature vector of the same dimension, thus making the fusion of information from different modalities possible.

[0076] For example, the encoded modal feature vector Features of speech modality, Features of emotional modality Features are defined for the context modality. Information from each modality is represented within a unified feature space, allowing features from different modalities to be directly compared and fused.

[0077] In one possible implementation, the system uses multiple sensors to perceive the user in real time. For example, a microphone collects the user's voice, a camera captures the user's facial expressions, a pressure sensor obtains tactile feedback, and a physiological signal acquisition device monitors the user's heart rate, skin conductance, etc. Through the perception information acquisition unit, this data is synchronously transmitted to the modality preprocessing unit for processing such as noise reduction, normalization, and temporal alignment.

[0078] The preprocessed data is transformed into feature vectors by a dedicated feature encoder. At this point, speech information is converted into speech feature vectors, image data into image feature vectors, and tactile and physiological signals into their respective feature vectors. Feature vectors from all modalities are integrated into the same feature space, ensuring the system can process and utilize information from different sources in a consistent manner.

[0079] In this embodiment, for the modality graph construction and path optimization module, a multimodal information graph is constructed based on the modal feature vectors generated by the multimodal perception and encoding module. Through this information graph, the system can establish connections between multimodal features and further optimize the modality fusion path, thereby effectively improving the system's response accuracy and efficiency.

[0080] In the processing of multimodal information, due to the differences in data characteristics and processing requirements of each modality, the use of a single modality often cannot provide sufficient contextual information. In this context, the modality graph construction and path optimization module becomes particularly important. It integrates the feature information of multiple modalities to construct a graphical structure that comprehensively represents the relationships between modalities, thereby ensuring that the system can continuously optimize processing paths when facing complex and ever-changing interaction scenarios, improving the system's real-time responsiveness and decision-making accuracy.

[0081] Typically, the modal information graph (MAP) construction unit first builds a graph structure based on feature vectors from different modalities. Nodes in the graph represent features of different modalities, and edges represent the relationships between these features. For example, one node might represent features of the speech modality, while another might represent features of the visual modality. The edges connecting these nodes represent potential associations or mutual influences between speech and vision. In this way, the system can explicitly perceive the connections between modalities and provide data support for path optimization.

[0082] In one possible implementation, the construction of the modal information graph also considers the correlation between modalities. For example, the weights of the graph edges can be adjusted based on the correlation between modalities. High correlation between modalities may result in larger edge weights, thus prioritizing the fusion of these modalities during path optimization.

[0083] In this embodiment, the path optimization unit is responsible for optimizing the modal fusion path based on the constructed modal information graph. Specifically, this unit utilizes an information entropy evaluation mechanism to select the fusion path with the minimum information entropy by calculating the information entropy values ​​corresponding to different paths. Information entropy is used to measure the degree of disorder of information; therefore, the path with the minimum information entropy can effectively achieve accurate fusion of modal information.

[0084] In one possible implementation, the path optimization unit performs modal fusion path optimization according to the following formula:

[0085] ;

[0086] in, This represents the final optimized modality fusion feature vector; , , These are the weighting coefficients for speech, emotion, and contextual modalities, respectively. , , These are feature vectors for speech, emotion, and contextual modalities, respectively.

[0087] The system will obtain a comprehensive feature vector by weighted summation of the feature vectors of speech, emotion, and contextual modalities, which will serve as the final modality fusion result. During the adjustment of weight coefficients, the path optimization unit will dynamically evaluate the importance of each modality in the current context and adjust the weights accordingly.

[0088] In the path optimization process, the information entropy evaluation mechanism plays a crucial role. Specifically, information entropy, as a standard for measuring information uncertainty, can effectively guide the selection of modality fusion paths. By calculating the information entropy of each path, the system can evaluate which path can provide the most accurate and efficient information fusion. For example, between emotion and speech modalities, if the correlation between the two is strong and the information entropy is low, then this path may be selected as the optimal path.

[0089] To further enhance the accuracy of path optimization, the system may also introduce other evaluation criteria, such as intermodal correlation metrics and feature signal-to-noise ratios. These additional evaluation criteria can further improve the selection accuracy of modal fusion paths, thereby enhancing the overall performance of the system.

[0090] In this embodiment, the set of modal feature vectors selected by the optimized path is further semantically aligned, fused, and generated with low-dimensional representation, thereby providing a unified semantic input for the subsequent policy response generation module.

[0091] In general, the modal features extracted during multimodal perception are high-dimensional, numerous, and suffer from inconsistencies in dimensionality and redundancy in representation. To address these issues, the fusion representation generation module employs an index-guided modal feature selection mechanism, combined with a weighted fusion strategy and a compressed representation generation method, effectively avoiding information redundancy and improving computational efficiency. This module achieves a good balance between the continuity of information paths and the semantic preservation of compressed representations.

[0092] In this embodiment, the fusion expression generation module includes a modality feature selection unit and a modality feature fusion unit. The modality feature selection unit selects the corresponding modality channel, such as the feature vectors of speech, visual, or tactile modalities, from the pre-encoded multimodal feature vector library based on the modality feature index set output by the upstream module. This selection operation is dynamically adjusted not only based on path optimization results but also by incorporating the confidence information of each modality feature in the current interaction state.

[0093] In one possible implementation, modality feature selection can be controlled using a gating mechanism, such as using a sigmoid activation function to determine whether each modality is activated.

[0094] ;

[0095] in, Indicates the first The activation state of each modal channel; The control parameters for this modal channel are determined by the fusion context; This represents the sigmoid function.

[0096] After selecting modal features, the modal feature fusion unit performs weighted fusion calculations and outputs a single modal fusion vector representation. Specifically, this module uses a weighted linear combination method to integrate information between modalities, and the fusion operation conforms to the following calculation formula:

[0097] ;

[0098] in, This represents the modal representation features after fusion; For the first Each modal feature vector has a uniform dimension and originates from the selected modal set; For the first The fusion weight coefficients of each modality channel reflect the importance of that modality in the current semantic task; This represents the number of modalities currently participating in the fusion.

[0099] As an option, The determination can be based on the combined effect of multiple factors such as modal confidence, information gain, and temporal context correlation, and can be performed offline training or online optimization and updating through historical interaction data.

[0100] In some embodiments, to further improve the compactness and semantic integrity of the fused expression, this module may add a compression transformation layer after the fusion calculation, such as using PCA, sparse coding or low-rank tensor projection, to reduce redundant feature dimensions while retaining key semantic information.

[0101] Specifically, after the fusion Dimensionality reduction operations can be introduced above:

[0102] ;

[0103] in, It is the fusion of multimodal features or data; This is the fused data after compression and transformation; This represents a compression transformation function, which can take the form of a linear projection matrix, principal component analysis operator, or sparse autoencoder, etc.

[0104] In another implementation, the compressed representation can automatically match the fused output dimension based on the input dimension of the target policy network, thereby avoiding redundant mapping and transformation operations. This mechanism ensures that the fused representation not only has uniform dimensions but also maintains a consistent semantic hierarchy, which is beneficial for the policy decision network to generate the next response.

[0105] In this embodiment, the strategy response generation module is responsible for inputting the generated unified semantic expression into the strategy reasoning path and ultimately outputting an interactive response action signal with timeliness and semantic matching. This module realizes the complete link transformation from "perception—cognition—decision—response" and is the core intermediary connecting user input and digital human execution output.

[0106] Generally, because fused expressions may have high information density but unclear semantic boundaries, the policy response generation module not only needs to have the ability to decode fused semantics, but also needs to support dynamic scene adaptation and multi-round response persistence. In this system, this module uses a two-layer structure of "semantic parsing unit" and "behavioral policy network unit" to achieve dynamic optimization of semantic content decomposition and policy generation.

[0107] In this embodiment, the policy response generation module first includes a semantic parsing unit. This unit receives fused representation features. Or compressed expression Multi-scale semantic parsing is then performed on it. On the one hand, explicit semantic labels are extracted, including intent categories and behavioral goals; on the other hand, the interaction history state vector is extracted using a contextual memory mechanism. This helps in understanding the semantic position of the current expression.

[0108] Specifically, the following mapping function can be used to extract tags:

[0109] ;

[0110] in, For the current set of semantic tags; This represents a label prediction function, which can be constructed based on shallow fully connected networks or multi-task learning models. This is the set of model parameters.

[0111] In one possible implementation, context state It can be constructed using a gated recurrent unit (GRU), which has strong ability to preserve expression across time, as shown below:

[0112] ;

[0113] in, Indicates the state of the previous interaction; This represents the current input fusion feature.

[0114] In this embodiment, the module further includes a behavior policy network unit, which will process semantic tags. With context state Joint input is used, and a policy network is employed to model the probabilistic action response. To achieve more discriminative action output selection, this module introduces a decision-making mechanism based on a soft policy function, combined with temperature regulation parameters. Control the smoothness of the output distribution.

[0115] The policy optimization function is defined as follows:

[0116] ;

[0117] in, Indicates the state Take action below The probability of the strategy; For action value functions; Temperature coefficient; The normalization term for all candidate actions; is the base of the natural logarithm.

[0118] Alternatively, the policy function can be continuously optimized via backpropagation. The function's parameters are adapted to the user's behavioral preferences during continuous interaction.

[0119] In another possible implementation, the policy network employs a multi-channel parallel convolutional structure to process the semantic intent vector and the historical state trajectory vector separately. After fusion, an attention mechanism is used to weight and combine the action channels to output the action control signal. :

[0120] ;

[0121] in, This represents the attention aggregation function; Represents the state Below, all possible actions Value (action value function); For context state; This is the final selected interactive response action signal.

[0122] In this embodiment, the digital human driving module is crucial for converting interactive response action signals into executable physical control commands and driving the digital human to perform multimodal outputs such as voice, facial expressions, and body movements. This module realizes the mapping and transformation from abstract semantic action signals to specific interactive performances. It is the terminal output link of the perception-cognition-response chain in the system and directly determines the interactive experience and immersion level between the digital human and the user.

[0123] Generally, the interactive response action signal output by the strategy response generation module As an abstract vector in a high-dimensional semantic space, it cannot be directly used to drive actual synthesis or execution devices. Therefore, the digital human driving module uses a signal parsing and execution management mechanism to parse the signal into structured control parameters, which are then allocated to the speech, facial expression, and motion driving sub-modules to achieve coordinated, synchronized, and realistic interactive output behaviors.

[0124] In this embodiment, the digital human driving module includes a signal parsing unit and an output execution unit. The signal parsing unit is responsible for receiving policy response action signals. It then uses an action parameter mapping function to decode it into three sets of control parameters:

[0125] ;

[0126] in, It represents the control parameters for speech synthesis, including phoneme sequence, speech rate, intonation, emotion tags, etc. It represents facial expression control parameters, including facial key point displacement, expression labels, muscle drive, etc. It represents body movement control parameters, including joint angle sequence, movement rhythm, and posture target point.

[0127] In one possible implementation, the control parameter parsing process can be based on a multi-head attention mechanism to achieve decoupled representation between submodals, as shown below:

[0128] ;

[0129] in, For the first Attention aggregation module for class control channels; For the corresponding control channel, a multilayer perceptron is used to complete the mapping learning from semantic actions to control parameters; For the final generated first Class control parameters; This is a strategy response action signal.

[0130] As an alternative, to ensure temporal consistency between speech and lip movements, the speech synthesis module in the system adopts a VITS-Wav2Vec2 joint model structure, synchronously generating a time-aligned signal during speech generation. The lip-sync animation generation part then uses the LipSync model to reference this time signal for frame-level driving. Voice control can employ the following synthesis function:

[0131] ;

[0132] in, In time The voice frames output in real time; It represents the control parameters for speech synthesis, including phoneme sequence, speech rate, intonation, emotion tags, etc. This represents a speech synthesis model that performs end-to-end generation by combining speech features and emotion modulation parameters. The corresponding lip-sync function is:

[0133] ;

[0134] in, This refers to the displacement of the lip shape keyframe; It represents facial expression control parameters, including facial key point displacement, expression labels, muscle drive, etc. In time The voice frames output in real time; This represents an audio-driven lip-sync prediction function that combines speech frames with facial control features for regression.

[0135] In some embodiments, the expression-driven component can further integrate emotional semantic vectors to achieve an explicit mapping between emotional consistency and expression synthesis, thereby enhancing the naturalness and emotional realism of the digital human's output. The motion control path relies on a physics simulation engine, such as Bullet, combined with a motion sequence prediction module to form an interpretable output path from intention to physical motion trajectory.

[0136] The output execution unit is responsible for distributing the parsed control parameters to the underlying rendering and execution modules.

[0137] Voice control parameters are used to drive the TTS synthesis engine;

[0138] Facial expression parameters are mapped to a 3D facial model controller;

[0139] The motion parameters are input into the skeletal animation system or the physics simulation module.

[0140] These submodules support heterogeneous platform output, including WebGL front-end rendering, high-resolution rendering on PCs, or playback on browser-based media players.

[0141] Specifically, the output execution process can be represented by the following combined function:

[0142] ;

[0143] in, This represents the final rendered output sequence; This represents the rendering and action execution functions, which complete the mapping from control parameters to the digital human's trimodal actions; It represents the control parameters for speech synthesis, including phoneme sequence, speech rate, intonation, emotion tags, etc. It represents facial expression control parameters, including facial key point displacement, expression labels, muscle drive, etc. It represents body movement control parameters, including joint angle sequence, movement rhythm, and posture target point.

[0144] In practical applications, this module also supports state mechanism management of multimodal output processes, including standby, activation, response, transition and other multi-state switching logic, to ensure the continuity of digital human behavior output and the consistency of state control.

[0145] The AI ​​agent digital human interaction method based on multimodal perception described below can be referred to in correspondence with the AI ​​agent digital human interaction system based on multimodal perception described above.

[0146] Please see the appendix Figure 7 The present invention also provides a digital human interaction method for AI intelligent agents based on multimodal perception, comprising the following steps:

[0147] S1. Collect the user's multimodal perception information, preprocess the perception information, and obtain multiple preprocessed perception features;

[0148] S2. Perform feature encoding on the preprocessed perceptual features to obtain multiple modal feature vectors;

[0149] S3. Construct a modal information graph based on the modal feature vectors, determine the optimal modal fusion path using the information entropy minimization method, and output the set of modal feature indices corresponding to the modal fusion path;

[0150] S4. Based on the modal feature index set, select the corresponding modal features from the modal feature vector and fuse them to generate a compressed representation;

[0151] S5. Input the compressed expression into the behavior policy network, and generate interactive response action signals through the policy optimization mechanism;

[0152] S6. Receive interactive response action signals and parse them into voice control parameters, facial expression control parameters and body movement control parameters;

[0153] S7. Drive the speech synthesizer, expression driver, and motion actuator according to the parsed control parameters to control the digital human to perform corresponding speech, expression, and motion outputs, thereby realizing multimodal interaction.

[0154] For step S1, in this embodiment, the system first collects the user's multimodal perception information, including but not limited to speech signals, facial images, body posture, and contextual information. The collected raw perception data undergoes preprocessing operations, including denoising, normalization, image enhancement, and endpoint detection, to ensure the stability and feature reliability of subsequent encoded inputs. The preprocessed data is converted into multiple perception feature sets, each corresponding to a different modal input.

[0155] In step S2, in this embodiment, the perceptual features of the multiple modalities are input into their respective encoding modules. Each modality is encoded using a differentiated feature extraction network. For example, spatial features of the image modality can be extracted based on a convolutional neural network, while prosodic and semantic features of the speech modality are extracted through temporal modeling. Finally, multiple modal feature vectors are generated, providing a basic representation for subsequent modality fusion.

[0156] For step S3, in this embodiment, the system constructs a modal information graph based on the acquired modal feature vectors, where each node in the graph represents a different modal feature, and edges represent the correlation or complementarity between modalities. To determine the optimal modal fusion path, the system introduces an information entropy minimization method, combined with the information transmission efficiency between nodes, to optimize the path selection logic. The final output fusion path corresponds to a modal feature index set, which is used for subsequent simplification and fusion.

[0157] For step S4, in this embodiment, specified modal features are selected from the original modal feature vectors and weighted fusion is performed. The fusion method can employ attention mechanisms, principal component analysis, or sparse autoencoders to generate compressed representation vectors, which serve as a unified perceptual input representation of the digital human's response behavior.

[0158] For step S5, in this embodiment, the compressed representation vector is input into the behavior policy network. This network structure employs a multilayer perceptron or a Transformer-based decision architecture, and combines policy optimization mechanisms (such as PPO or behavior cloning) to train the model to generate action response signals. The generated signals are used to represent the feedback response behavior that the digital human should exhibit.

[0159] In step S6, in this embodiment, the response action signal output by the behavior strategy is decoded by the parsing module and parsed into three types of specific control information: voice control parameters, facial expression control parameters, and body movement control parameters. The parsing method is based on a predefined action mapping table and parameterized behavior templates to achieve modular understanding of the action signal.

[0160] In step S7, in this embodiment, the speech synthesizer, expression driver, and motion actuator are driven according to the control parameters obtained from the above analysis. The digital human generates speech output based on the speech driver module, the expression driver module adjusts facial muscle movements according to the control parameters, and the motion actuator module controls the body skeletal animation to achieve coordinated output of speech, expression, and motion, thereby completing multimodal interactive behavior.

[0161] The method in this embodiment can be used to execute the above system embodiment, and its principle and technical effect are similar, so it will not be described again here.

[0162] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A multimodal perception-based AI intelligent agent digital human interaction system, characterized in that, include: The multimodal perception and encoding module is used to collect user perception information and preprocess and encode the collected user perception information to obtain multiple modal feature vectors. The modality graph construction and path optimization module is used to construct a modality information graph based on the modality feature vector, determine the modality fusion path with the minimum information entropy, and output the modality feature index set corresponding to the modality fusion path; The fusion representation generation module is used to select corresponding modal features from the modal feature vectors according to the modal feature index set, fuse them, and generate a compressed representation. The strategy response generation module is used to input the fused expression into the behavior policy network and generate interactive response action signals through the policy optimization mechanism; The digital human driving module is used to control the digital human to perform voice, facial expression or action interaction output according to the response action; The modality graph construction and path optimization module includes: The modal information graph construction unit is used to construct a modal information graph based on the multiple modal feature vectors, where nodes in the graph represent different modal features and edges represent the relationships between modal features; The path optimization unit is used to determine the path of modal fusion based on the modal information graph through an information entropy evaluation mechanism, and output the optimal set of modal feature indices. The path optimization unit performs modality fusion path optimization according to the following formula: W total =αW Voice +βW Emotion +γW Context ; Where α, β, and γ are the weighting coefficients for speech, emotion, and contextual modalities, respectively; W total W represents the final optimized modality fusion feature vector. Voice W Emotion W Context These are feature vectors for speech, emotion, and contextual modalities, respectively.

2. The AI ​​intelligent agent digital human interaction system based on multimodal perception according to claim 1, characterized in that, The multimodal sensing and coding module includes: The sensory information acquisition unit is used to collect user sensory information, including voice signals, image information, tactile pressure signals, physiological signals, and environmental sensory data. The modal preprocessing unit is used to perform noise reduction, normalization, temporal alignment and missing information processing on the perceived information. The feature encoding unit is used to extract features from the preprocessed perceptual information using a modal-specific encoder, and output a modal feature vector of uniform dimension.

3. The AI ​​intelligent agent digital human interaction system based on multimodal perception according to claim 1, characterized in that, The fusion expression generation module includes: A modal feature selection unit is used to select a corresponding modal feature from multiple modal feature vectors based on the modal feature index set; The modal feature fusion unit is used to fuse the selected modal features to generate a modal fusion feature vector with a compressed representation.

4. The AI ​​intelligent agent digital human interaction system based on multimodal perception according to claim 3, characterized in that, The modal feature fusion unit performs modal feature fusion processing according to the following formula: Among them, F fusion The fused modal features; W i Let α be the eigenvector of the i-th modality. i is the weight coefficient of the i-th modal feature; n is the total number of modal features.

5. The AI ​​intelligent agent digital human interaction system based on multimodal perception according to claim 1, characterized in that, The policy response generation module includes: The semantic parsing unit is used to perform semantic decomposition on the fused representation and extract semantic tags and context states; The behavior policy network unit is used to input the semantic tags and context state into the policy network and output the probability distribution of interactive response action signals.

6. The AI ​​intelligent agent digital human interaction system based on multimodal perception according to claim 5, characterized in that, The behavioral policy network unit optimizes the output of response actions based on the following policy function: Where π(a|s) represents the policy probability of taking action a in state s; Q(s,a) is the action value function; τ is the temperature coefficient; ∑ a′ e Q(s,a′) / τ is the normalization term for all candidate actions; e is the base of the natural logarithm.

7. The AI ​​intelligent agent digital human interaction system based on multimodal perception according to claim 1, characterized in that, The digital human driving module includes: The signal parsing unit is used to receive the interactive response action signal and parse it into voice control parameters, facial expression control parameters and body movement control parameters. The output execution unit is used to drive the speech synthesizer, expression driver and motion actuator respectively according to the control parameters to realize the multimodal interactive output of the digital human.

8. A multimodal perception-based AI agent digital human interaction method, applied to the multimodal perception-based AI agent digital human interaction system according to any one of claims 1-7, characterized in that, Includes the following steps: Collect multimodal perception information from users, preprocess the perception information, and obtain multiple preprocessed perception features; The preprocessed perceptual features are encoded to obtain multiple modal feature vectors; A modal information graph is constructed based on the modal feature vectors, and the optimal modal fusion path is determined by the information entropy minimization method. The set of modal feature indices corresponding to the modal fusion path is then output. Based on the modal feature index set, the corresponding modal features are selected from the modal feature vectors and fused to generate a compressed representation; The compressed representation is input into the behavior policy network, and interactive response action signals are generated through the policy optimization mechanism; Receive interactive response action signals and parse them into voice control parameters, facial expression control parameters, and body movement control parameters; The speech synthesizer, expression driver, and motion actuator are driven by the parsed control parameters to control the digital human to perform corresponding speech, expression, and motion outputs, thereby realizing multimodal interaction.

Citation Information

Patent Citations

  • Knowledge fusion method and system for multi-source heterogeneous multi-modal data

    CN118690838A

  • Virtual character interaction method and device

    CN120105009A