Multi-user emotion recognition and digital human feedback method based on behavior data
By employing a multi-user emotion recognition and digital human feedback method based on behavioral data, and utilizing identity recognition and emotion analysis technologies, personalized strategies are generated to drive digital human feedback. This solves the problems of privacy protection and personalized interaction, and achieves efficient and secure emotional interaction.
Patent Information
- Application Number
- CN202511500705.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-21
- Publication Date
- 2026-01-16
AI Technical Summary
Existing technologies cannot meet users' core needs for privacy protection and personalized emotional interaction, thus hindering the practical implementation and development of digital human emotional interaction technology.
User identity is confirmed through RFID, facial recognition, and voiceprint recognition. Real-time sequential dialogue behavior data of user interaction with digital human is captured. Emotion analysis is performed using a sequence information processing model based on state space theory to generate multimodal control commands to drive digital human feedback. Personalized strategies are generated by combining user profiles.
It achieves non-intrusive, multi-user emotion recognition and feedback, eliminates privacy concerns, improves the usability of the interaction and user trust, reduces the uncanny valley effect, and enhances interaction stickiness.
Smart Images

Figure CN121350883A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of affective computing of artificial intelligence, and particularly relates to a multi-user emotion recognition and digital human feedback method based on behavior data. BACKGROUND
[0002] With the rise of artificial intelligence, 5G communication and the concept of meta universe, human-computer interaction is evolving from traditional graphical user interface (GUI) to a more immersive, natural and intelligent form. Digital human technology, as a representative of the frontier, has expanded its application boundaries from entertainment and media to key industries such as enterprise services, online education, medical and health care, and virtual companionship. In this process, the effectiveness of interaction depends not only on the accuracy of information transmission, but also on the depth of emotional connection. Interactions with "empathy" capabilities can significantly improve user learning efficiency, problem-solving rate and brand loyalty. Therefore, the ability of digital humans to understand and respond to user emotions has become a core focus of the industry. Current related technical solutions mainly include: traditional Web interaction systems (relying on lagging operation log analysis, single interaction, no multi-user recognition capability), emotion recognition solutions based on sensors (relying on cameras and wearable devices, high invasiveness and high privacy risk, only recognizing basic expressions), emotion analysis solutions based on explicit input content (only relying on text / voice content, ignoring behavioral information, one-sided and lagging understanding), and digital humans / smart speakers integrated with voice recognition (core response to voice commands, lack of deep emotion analysis and multi-user complex scene adaptation capability. To optimize and improve digital human technology, this project is supported by "Sichuan Science and Technology Plan", project number: 2024NSFSC0499.
[0003] Traditional Web systems lack real-time emotion perception and natural interaction capabilities. Sensor solutions have privacy and scene adaptation problems due to reliance on hardware devices. Explicit input analysis solutions ignore emotional information in user behavior. Even more advanced digital human solutions integrated with voice recognition lack multi-user accurate recognition and conversation management capabilities, as well as deep emotion-driven dynamic feedback mechanisms. In summary, existing technologies have not formed a complete and effective technical closed loop in the three key dimensions of non-intrusive deep emotion perception, multi-user accurate recognition in complex scenarios, and dynamic multi-modal emotional feedback. They cannot meet the core needs of users for privacy protection and personalized emotional interaction, and have become a major bottleneck restricting the development of digital human emotional interaction. SUMMARY
[0004] The purpose of the present application is to provide a multi-user emotion recognition and digital human feedback method based on behavior data, to solve the following technical problems raised in the background art: Existing technologies cannot meet the core needs of users for privacy protection and personalized emotional interaction, restricting the actual landing and development of digital human emotional interaction technology.
[0005] To solve the above technical problems, the technical solution adopted by the present application is: A multi-user emotion recognition and digital human feedback method based on behavior data, comprising the following steps: Step one, perceive user proximity, confirm the identity of the registered user and assign a unique identifier for the current session, or assign a temporary session identifier for unregistered visitors, through at least one of RFID, facial recognition and voiceprint recognition; when there are multiple users, distinguish different users and establish independent session channels through sound source positioning and speaker diarization technology; Step two, during the conversation, real-time capture of the time-series dialogue behavior data generated when the user interacts with the digital human, dialogue behavior data including speech prosody features, dialogue interaction patterns and linguistic features; Step three, input the dialogue behavior data into the pre-trained emotion analysis model, the model adopts a sequence information processing model based on state space theory with data-dependent selection mechanism and a static context feature splicing architecture, and outputs the user's current multi-dimensional emotion and cognitive state category emotion state label related to the interaction scene; Step four, according to the emotion state label and the personal profile associated with the user's identity, generate an emotional support strategy through a decision model; Step five, according to the emotional support strategy, generate and execute multi-modal control instructions that drive the digital human to perform emotional performance, which synchronously controls the digital human's expression, action, voice parameters and scene special effects.
[0006] Further, sound source positioning is achieved by calculating the time difference of sound arriving at different microphones or applying beamforming algorithm, and speaker diarization technology is achieved by analyzing the voiceprint features of continuous speech stream, cutting the speech stream into independent speech segments corresponding to different users, to complete the multi-user conversation binding.
[0007] Further, speech prosody features include speech rate, pitch variation range, volume fluctuation amplitude, and pause duration and frequency; dialogue interaction patterns include user response delay, frequency of interrupting digital human speech, turn-taking rules, and frequency of requesting clarification or repetition; linguistic features include filler word usage frequency and user question type and complexity.
[0008] Further, the core model for processing time-series dialogue behavior data is a sequence information processing model based on state space theory with data-dependent selection mechanism; static context features include the difficulty of the current interaction task and the user's knowledge level; the multi-dimensional emotion and cognitive state category corresponding to the emotion state label at least covers the state types related to user task engagement, cognitive load and emotional tendency; the sequence information processing model is a Mamba state space model.
[0009] Furthermore, the personal profile includes the user's historical interaction preferences, personalized characteristics related to the interaction scenario, and past emotional response patterns; the decision-making model is a rule system or reinforcement learning agent, and the emotion support strategy defines the intervention type, speech style, and guidance direction, and the intervention type is adapted to the user's state corresponding to the emotion state label.
[0010] Furthermore, the multimodal control instructions are JSON format objects, containing digital human emotion expression tags, digital human action tags, voice parameters, and scene effect parameters; the digital human rendering adopts rendering technology adapted to 2D or 3D styles.
[0011] A multi-user emotion recognition and digital human feedback device based on behavioral data is used to implement the aforementioned multi-user emotion recognition and digital human feedback method based on behavioral data, including an identity recognition module, a data acquisition module, an emotion analysis module, a strategy generation module, a feedback control module, and a storage module; The identity recognition module is used to sense the approach of users. It uses RFID readers, image sensors and microphone arrays to confirm the identity of registered users in at least one of RFID, face recognition and voiceprint recognition. It assigns temporary session identifiers to visitors and establishes multi-user independent session channels through sound source localization and speaker log technology. The data acquisition module is used to capture the temporal dialogue behavior data stream of the user's interaction with the digital human in real time. The data stream includes speech prosody features, dialogue interaction patterns and linguistic features. The emotion analysis module is pre-set with an emotion analysis model that uses a sequence information processing model based on state space theory and a data dependency selection mechanism and a static context feature splicing architecture. It is used to receive dialogue behavior data and output the emotion state label of the user’s current multi-dimensional emotion and cognitive state category related to the interaction scenario. The strategy generation module is used to call the user's personal profile in the storage module, combine it with the emotion state tags, and generate emotion support strategies through the decision model. The feedback control module is used to generate multimodal control commands based on the emotion support strategy, drive the digital human to perform facial expressions, actions and voice output, and trigger scene effects; The storage module is used to store user profiles, conversation behavior data, and an emotion support strategy library.
[0012] Furthermore, the image sensor of the identity recognition module runs a lightweight face detection and recognition algorithm, and the microphone array runs a voiceprint recognition model to achieve rapid verification of registered users' identities and differentiation of sessions for unregistered users.
[0013] Furthermore, the storage module adopts a hybrid database architecture: a graph database is used to store user knowledge structure and learning path, a time-series database is used to store time-series dialogue behavior data streams, and a relational database or document database is used to store user profiles, emotion support strategy library and conversation records, and all stored data is bound to the user's unique identifier.
[0014] Furthermore, the feedback control module includes a text-to-speech engine, which adjusts the speech rate, pitch, and volume of the synthesized speech according to the speech parameters in the multimodal control instructions to match the emotional tone corresponding to the emotion support strategy; scene effects include switching the background style of the interactive interface and playing background music that matches the emotional atmosphere.
[0015] Compared with the prior art, the present invention has the following beneficial effects: In this invention, non-invasive data acquisition is sensor-free, eliminating user privacy concerns. The multi-user session isolation technology is adapted to public and multi-user interaction scenarios, breaking through the limitations of traditional single-user solutions.
[0016] In this invention, based on temporal behavioral data and a hybrid architecture model, it is possible to identify deep cognitive states that are strongly related to the scene, such as confusion and focus, rather than just basic emotions, thereby improving the practicality of emotion understanding.
[0017] In this invention, a user profile generation strategy is combined to ensure personalized feedback for each individual; multimodal synchronous feedback enhances the warmth of digital human interaction, reduces the uncanny valley effect, and improves user trust and interaction stickiness.
[0018] In this invention, a complete process is formed from identity perception to feedback execution, and the data is bound to the user identifier, providing data support for model iteration, realizing the system's self-learning and adaptive capabilities, and improving the effectiveness of interaction in the long term. Attached Figure Description
[0019] Figure 1 This is a schematic diagram of the overall process of the present invention; Figure 2 This is a schematic diagram of the multi-user identification process of the present invention; Figure 3 This is a schematic diagram of the multi-user differentiation processing flow of the present invention; Figure 4 This is a schematic diagram of the user emotion analysis and digital human feedback control process of the present invention. Detailed Implementation
[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0021] Implementation 1: A method for multi-user emotion recognition and digital human feedback based on behavioral data, such as Figure 1 As shown, it includes the following steps: Step one: Detect user approach and verify the registered user's identity using at least one of RFID, facial recognition, and voiceprint recognition, assigning a unique identifier to the current session, or assigning a temporary session identifier to unregistered visitors, such as... Figure 2 As shown, when there are multiple users, different users are distinguished and independent session channels are established through sound source localization and speaker log technology.
[0022] This step serves as the foundation for interaction, enabling user identity anchoring and session isolation. Multimodal recognition confirms registered user identities and assigns temporary identifiers to visitors. Simultaneously, sound source localization and speaker logging technologies address session attribution issues in multi-user scenarios, ensuring that each user's interaction data is independently linked and avoids confusion.
[0023] Step two: During the conversation, capture the temporal dialogue behavior data stream generated when the user interacts with the digital human in real time. The dialogue behavior data includes speech prosody features, dialogue interaction patterns and linguistic features.
[0024] This step serves as the data input for emotion analysis. It captures temporal behavioral data during the conversation in real time, such as speech prosody, interaction patterns, and linguistic features, providing non-invasive, highly relevant raw data support for subsequent emotion recognition and avoiding the privacy risks associated with traditional sensors.
[0025] It should be noted that the "non-invasive" approach described in this invention primarily differs from technical solutions that require users to wear additional specialized sensors (such as EEG caps, heart rate bracelets, and electrodermal sensors) to collect physiological signals. The technical solution of this invention relies solely on conventionally configured devices in the interactive environment (such as microphone arrays and cameras). During natural dialogue interaction, data collection can be completed without the user's active cooperation or wearing any additional devices, thus achieving non-invasive emotional state recognition that is uninterrupted and burden-free for the user. Step 3: Input the dialogue behavior data into the pre-trained sentiment analysis model. The model adopts an architecture that combines a sequence information processing model based on state space theory and a data-dependent selection mechanism with static context features, and outputs the sentiment state label of the user's current sentiment and cognitive state category related to the interaction scenario.
[0026] In the specific implementation of this step, the real-time captured dialogue behavior data first needs to be preprocessed and vectorized using a systematic approach. Within a preset time window, such as 5 seconds, the system extracts multi-dimensional prosodic features from the user's speech signal, including the number of syllables V reflecting speech rate. rate The fundamental frequency mean μ, which reflects the pitch characteristics F0、 Standard deviation σ F0 and the range of variation R F0 The root mean square energy μ that reflects the change in volume E and standard deviation σ E And the total duration T of pauses, which reflects the fluency of the conversation. pause With frequency F pause At the same time, the system will also quantify dialogue interaction pattern characteristics that reflect user engagement, such as user response latency T. latency The frequency F of interrupting the digital human's speech interrupt Number of turn-taking sessions N turn Combined with filler words, such as the frequency of use of "um" and "that", etc., F filler Linguistic features, etc. All these numerical features, after normalization, are concatenated into a unified temporal feature vector X. t Its continuous sequence {X1,X2,...,X} T This constitutes the dynamic input data stream of the sentiment analysis model.
[0027] Next, the temporal feature vector sequence is input into a pre-trained sentiment analysis model that employs a sequence information processing model based on state space theory and possessing a data dependency selection mechanism, along with a static context feature concatenation architecture. The core of this model is a sequence information processing model based on state space theory and possessing a data dependency selection mechanism. Through its core state space model (SSM) and a data dependency selection mechanism, it efficiently compresses the sequence information into a compact hidden state, thereby effectively capturing long-distance dependencies of data over time with linear time complexity. Specifically, the temporal feature sequence {X1, X2, ..., X...} T First, it is encoded using the Mamba module, which outputs h in the hidden state at the last time step T. TThis is extracted as a generalized representation of the dynamic behavior throughout the entire time series. Simultaneously, static contextual features representing the current interaction scenario, such as the pre-defined task difficulty C, are also extracted. task and the knowledge level C based on user history assessment user This is constructed into a static feature vector C. Subsequently, in the feature concatenation layer, h, representing the dynamic behavior, is... T The vector F is concatenated with C, which represents the static scene, to form a comprehensive feature vector F that integrates dynamic and static information. fused =concat(h T (C). This vector is then subjected to a deeper nonlinear feature transformation through one or more fully connected layers to extract deep abstract features for judging emotional state.
[0028] Finally, the model maps these deep features to specific emotion state labels through the output layer. These labels are not simply single categories, but rather multi-dimensional floating-point vectors E=[e1,e2,...,e...]. N Each dimension corresponds to a cognitive or emotional state closely related to the interaction scenario, such as EngagementLevel, CognitiveLoad, and EmotionalValence. To obtain quantified results, the output layer uses the Sigmoid activation function, normalizing the output value of each node to the (0,1) interval. This value directly reflects the confidence or activation strength of the corresponding state. The complete mapping process from input to output is defined by the state-space model. The core of Mamba, the state-space model SSM, maps the input sequence x(t) to the output y(t) through a hidden state h(t). Its continuous form can be described by the following linear ordinary differential equation (LDE): State equation: h′(t) = Ah(t) + Bx(t) Output equation: y(t) = Ch(t) In practical computation, this continuous system is transformed into a discrete form for efficient computation, where A, B, and C are the state, input, and output matrices, respectively. The Mamba model achieves powerful selectivity by making these matrix parameters functions of the input data, allowing it to dynamically retain or forget information based on context. Through this series of rigorous mathematical transformations, the model transforms complex raw behavioral data into a set of quantifiable and interpretable multi-dimensional emotional state labels, providing accurate and reliable decision-making basis for generating personalized emotion support strategies in subsequent steps.
[0029] Step four: Based on the emotional state labels and the personal profiles associated with the user's identity, generate emotional support strategies through a decision model.
[0030] In the specific implementation of this step, it is first necessary to clarify the origin and composition of the user's personal profile. This profile is initialized and created when the user first interacts with the system, whether through registration to confirm identity or as a visitor assigned a temporary identifier, and is bound to the user's unique identifier. Subsequently, in each interaction session, the system continuously and automatically adds and records information such as the user's dialogue behavior data, the emotional state label sequence generated in step three, and the system's feedback strategy and the user's subsequent reactions after each interaction to this profile. Therefore, the personal profile is a dynamically updated database, and its content mainly includes: (User's historical interaction preferences, such as responses to proactive guidance or passive listening). Personalized features related to the interaction scenario, such as common emotional fluctuation patterns under specific tasks; And past emotional response patterns, such as the average time it takes to recover from a state of confusion to a state of focus.
[0031] This archive provides a crucial data foundation for enabling personalized decision-making for each individual.
[0032] After obtaining the real-time emotion state label output from step three, the system combines it with the user's personal profile retrieved from the storage module, and inputs both into a pre-trained reinforcement learning agent—the core of the decision-making model—to generate the optimal emotion support strategy. The key elements of this reinforcement learning agent are defined as follows: State space: The state st observed by the agent at any decision time t is a high-dimensional vector that integrates current emotion and historical personality. It consists of the multi-dimensional emotion state label E output from step three. t =[e1,e2,...,e N ] and key historical features P extracted from user profile P t It is pieced together, that is, s t =concat(E t ,P t ).
[0033] Action Space: The "actions" that the agent can execute are a predefined set of emotion support strategies throughout the system. Each strategy AI is structured as a specific instruction containing multiple dimensions, such as a triple defined as intervention type, discourse style, and guidance direction, like soothing, gentleness, or proactively offering simplified solutions. The entire action space A = {a1, a2, ..., a...} K It contains all possible strategy combinations.
[0034] • Reward function: Reward function R(s) t,a t ,s t+1 The design of the reward r is the core of agent learning, aiming to quantify the effectiveness of a policy under specific conditions. t It is based on the action a. t Afterwards, the user's emotional state changed from E t To E t+1 The reward function is calculated based on the changes in the emotional state vector. A specific reward function can be designed as a weighted sum of the changes in the emotional state vector: where t represents the amount of change in each emotional dimension. For example, the weights of w1 task engagement and w3 emotional tendency are positive, encouraging strategies to improve user engagement and enthusiasm; while the weight of w2 cognitive load is negative, punishing strategies that increase the user's cognitive load.
[0035] The generation process of the emotion support strategy, i.e., the agent's decision-making process, is based on a learned policy function that gives the probability of choosing action a in state s. In practical applications, this is usually implemented through a deep Q-network, which learns an optimal action value function Q(s,a) to predict the expected future cumulative reward obtained after performing action a in state s. Therefore, at decision time t, the agent calculates the current state s. t Find the Q-values of all possible actions a∈A, and select the action that maximizes the Q-value (as the output). Where θ is the parameter of the deep Q-network. This selected action a t This results in the final generated emotion support strategy, which is output in a specific, structured form, providing clear and quantifiable input instructions for the multimodal feedback execution in step five. Through this closed-loop learning and decision-making mechanism, the system can dynamically and data-drivenly generate emotion support strategies that are highly adapted to the current user's state and personalized needs.
[0036] Step 5: Based on the emotion support strategy, generate and execute multimodal control instructions that drive the digital human to perform emotional expressions. These instructions simultaneously control the digital human's facial expressions, movements, voice parameters, and scene effects.
[0037] In the specific implementation of this step, the core task is to accurately transform the abstract, high-level emotion support strategy generated in step four into a set of structured multimodal control instructions that can be directly executed by the digital human rendering engine, speech synthesis engine, and scene control module. This mapping process is accomplished through a pre-defined policy-instruction mapping rule base, which defines a corresponding multimodal performance paradigm for each emotion support strategy. Specifically, upon receiving the emotion support strategy output by the reinforcement learning agent, the system first dynamically generates specific dialogue content that conforms to the current dialogue context based on all elements of the strategy through a pre-trained text generation module, for example: "It's okay, let's take it step by step. This problem can be understood in a simpler way." While generating text content, the system uses the intervention type in the strategy as the primary key to retrieve a basic multimodal instruction template from the mapping rule base. This template predefines baseline performance parameters that match the intervention type, such as {facial expression label: "smile", action label: "slow nod", scene effect: "warm background"}. Subsequently, the system uses the "speech style" and "guidance direction" in the strategy as modifiers to refine this basic template. For example, a "gentle" style will further specify the speech synthesis parameters as a lower speech rate, a smoother pitch, and a moderate volume, and may modify the action label to a more subtle "slightly reassuring gesture"; while the guidance direction of "providing a simplified solution" may trigger a specific supplementary action label of "pointing to the screen".
[0038] Finally, after the above generation and refinement process, the system encapsulates the dynamically generated text content, finely tuned facial expression tags, action tags, speech synthesis parameters, and scene effect parameters into a standardized, JSON-formatted, multimodal control instruction object. This JSON object contains all the control information needed to drive a complete emotional interaction of the digital human and is synchronously sent to the corresponding execution modules: facial expression and action tags are sent to the digital human rendering engine, text content with emotional parameters is sent to the text-to-speech engine, and scene effect parameters are sent to the user interface rendering module. Through this hierarchical and clearly defined mapping mechanism from abstract strategies to specific instructions, it ensures that every response from the digital human is not only intelligent in content but also highly synchronized and empathetic with the user's emotional state in terms of expressiveness, thus forming a complete emotional interaction loop. Specifically, this solution follows a closed-loop logic of perception, collection, analysis, decision-making, and feedback: First, the identity recognition module detects the user's approach and establishes an independent conversation channel; then, it continuously collects temporal dialogue behavior data during the interaction; it processes the data using a pre-trained hybrid architecture model to output accurate emotion state labels; based on emotion labels and user profiles, it generates personalized emotion support strategies through a decision model; finally, it transforms the strategies into multimodal control commands to drive the digital human to perform emotional feedback, while all interaction data is bound and stored with user identifiers, providing a foundation for subsequent iterations.
[0039] This design enables non-intrusive data acquisition without sensor dependence, eliminating user privacy concerns. The multi-user session isolation technology is suitable for public, multi-user interaction scenarios, breaking through the limitations of traditional single-user solutions.
[0040] This design, based on temporal behavioral data and a hybrid architecture model, can identify deep cognitive states that are strongly related to the context, such as confusion and focus, rather than just basic emotions, thus improving the practicality of emotion understanding.
[0041] This design, combined with a user profile generation strategy, ensures personalized feedback for each individual; multimodal synchronous feedback enhances the warmth of digital human interaction, reduces the uncanny valley effect, and increases user trust and engagement.
[0042] This design creates a complete process from identity perception to feedback execution, and the data is bound to the user's identifier, providing data support for model iteration, enabling the system to learn and adapt, and improving the effectiveness of interaction in the long term.
[0043] Example 2: Sound source localization is achieved by calculating the time difference of sound reaching different microphones or by applying beamforming algorithms. Speaker logging technology analyzes the voiceprint features of a continuous speech stream and segments the speech stream into independent speech segments corresponding to different users in order to complete multi-user session affiliation binding.
[0044] This content serves as the core technical support for session attribution binding in multi-user scenarios, enabling precise association between multi-user voice input and corresponding user identities. It provides clear user attribution criteria for subsequent dialogue behavior data collection, sentiment analysis, and personalized feedback, resolving the session confusion problem when multiple users interact simultaneously.
[0045] Specifically, sound source localization relies on microphone arrays. By calculating the time difference of sound signals arriving at different microphones in the array, or by applying beamforming algorithms to focus sound in a specific direction, the current speaker's position in physical space can be determined.
[0046] Speaker log technology is based on the uniqueness of voiceprint features. It analyzes voiceprint information in a continuous speech stream, identifies differences in voiceprints among different users, and then segments the mixed speech stream into independent speech segments corresponding to different users.
[0047] The attribution binding principle is based on combining the spatial location information of the sound source localization with the voiceprint segmentation results of the speaker log to bind independent speech segments with specific users, thereby completing the attribution matching of multi-user sessions.
[0048] This design allows for the precise differentiation of mixed voice streams generated simultaneously by multiple users, avoiding the confusion of interaction data from different users. It also ensures that subsequent data collection on each user's dialogue behavior, emotional state analysis, and the generation of personalized emotional support strategies can be accurately linked to the corresponding user. Furthermore, it provides technical support for establishing independent conversation channels in multi-user scenarios, enabling the system to achieve personalized interactions tailored to each user, overcoming the limitations of traditional single-user interaction solutions.
[0049] Example 3: Speech prosodic features include speech rate, pitch range, volume fluctuation, and pause duration and frequency; dialogue interaction patterns include user response delay, frequency of interruptions to the digital human, turn-taking patterns, and frequency of requests for clarification or repetition; linguistic features include the frequency of filler words and the type and complexity of user-asked questions.
[0050] Clearly defining the core behavioral data dimensions required for non-invasive emotion recognition provides a standardized data collection scope and definition for the input acquisition module, while also providing quantifiable and highly correlated raw data support for the emotion analysis module. This serves as the fundamental data basis for realizing the transformation from behavioral data to emotional state.
[0051] Specifically, based on the core logic that there is an objective correlation between user interaction behavior and internal emotions and cognitive states, we break down the key data dimensions in the dialogue interaction process: Speech prosodic features focus on the physical attributes of speech, such as speech rate, pitch, volume, and pauses. These attributes change naturally with emotional fluctuations, such as speaking faster when anxious and pausing more when confused.
[0052] Dialogue interaction patterns focus on the rhythm of interaction between users and digital humans, such as response delays and interruption frequency. These patterns directly reflect the user's level of engagement and cognitive load, such as longer response delays when confused.
[0053] Linguistic features focus on details of language expression, such as filler words and question types. These details reflect the user's hesitation and the complexity of their needs. For example, more filler words are used when there is uncertainty. Ultimately, this forms a multi-dimensional data system covering speech, interaction, and language, providing comprehensive input for sentiment analysis.
[0054] To obtain the pre-trained emotion analysis model, a large amount of real or simulated human-computer dialogue data covering the various behavioral features described in this invention can be collected and labeled, and the model can be trained using supervised learning methods so that the model can accurately map behavioral data sequences to corresponding emotion and cognitive state labels.
[0055] This design avoids the randomness of data collection, ensuring that the captured behavioral data comprehensively covers the dimensions of emotion association, providing complete input for emotion analysis. Multi-dimensional data corroborate each other, reducing the risk of misjudgment from a single data dimension, enabling the emotion analysis model to more accurately infer users' deep emotions. All data comes from natural behavior during dialogue interactions, without relying on sensors such as cameras or wearable devices, supporting non-invasive emotion recognition from the data source.
[0056] Example 4: The core model for processing temporal dialogue behavior data is a sequence information processing model based on state space theory and possessing a data dependency selection mechanism; static context features include the difficulty of the current interaction task and the user's knowledge level; the emotional state labels correspond to multi-dimensional emotional and cognitive state categories, covering at least the state types related to the user's task participation, cognitive load, and emotional tendencies.
[0057] Specifically, clarifying the core technical components of the sentiment analysis model, providing standardized technical parameters and boundary definitions for model construction, and ensuring that the model can accurately process temporal data of dialogue behavior to effectively identify users' deep emotions and cognitive states are key technical supports for the sentiment analysis process. Regarding the selection of a temporal data processing model, this invention adopts a sequence information processing model (State Space Model, SSM) based on state space theory and possessing a data dependency selection mechanism. Compared to traditional Long Short-Term Memory (LSTM) networks, the Mamba model has significant advantages in processing long sequence data. Through a selective state space mechanism, it can model sequences with linear time complexity, effectively capturing ultra-long-distance dependencies, while avoiding the gradient vanishing problem that may occur with LSTM, and also solving the bottleneck of the Transformer architecture's quadratic increase in computational cost when processing long sequences. In real-time human-computer interaction scenarios, the dialogue behavior data stream is continuous and infinitely growing; Mamba's efficiency and powerful long-term memory capabilities make it a better choice for achieving accurate and real-time sentiment analysis. For the integration of static contextual features, the difficulty of the current interaction task and the user's knowledge level are key scenario and individual factors influencing the user's emotion and cognitive state. Concatenating these as static features with the output of the Mamba module can supplement scenario and individual information not covered by time-series behavioral data, avoiding the one-sidedness of sentiment analysis relying solely on dynamic behavioral data. Regarding the definition of emotion and cognitive state categories, based on the logic that emotions and cognitive states must align with the actual needs of the interaction scenario, the focus is on user task participation, cognitive load, and emotional tendency, ensuring that the output emotion state labels can directly support the generation of subsequent personalized support strategies.
[0058] Example 5: The personal profile includes the user's historical interaction preferences, personalized characteristics related to the interaction scenario, and past emotional response patterns; the decision-making model is a rule system or reinforcement learning agent, and the emotion support strategy defines the intervention type, speech style, and guidance direction, and the intervention type is adapted to the user's state corresponding to the emotion state label.
[0059] Specifically, it clarifies the core input elements, technical tools, and output standards for generating personalized emotion support strategies, providing a standardized technical framework for the transformation from emotion recognition to strategy generation. This ensures that strategy generation is supported by data, has practical tools, and has an adaptation basis, serving as a key technical link between emotion analysis and feedback execution.
[0060] Among them, the personal profile is based on historical interaction data to map the logic of users' personalized needs. Users' historical interaction preferences, contextual characteristics, and past emotional response patterns can quantify their acceptance and adaptability to different intervention methods, providing a data foundation for personalized strategies.
[0061] In the decision-making model, the rule system quickly outputs basic strategies based on preset emotion and strategy matching rules to ensure response efficiency; the reinforcement learning agent dynamically optimizes the decision-making logic by continuously learning user feedback on the strategy, balancing stability and iteration.
[0062] The strategy adaptation principle follows the logic of precise matching between needs and solutions. Based on emotional state tags, it first determines the corresponding intervention type, and then matches the appropriate communication style and guidance direction to ensure that the strategy is consistent with the user's current state.
[0063] This design allows personal profile support strategies to be tailored to individual user differences, avoiding homogeneous interventions. The rule system meets the need for real-time response, and reinforcement learning agents enable continuous strategy iteration. The strategies are precisely matched to the user's emotional state, reducing ineffective interventions, increasing user acceptance of emotional support, and optimizing the user experience.
[0064] To obtain the reinforcement learning agent, offline reinforcement learning training can be performed by building a simulated interaction environment or by utilizing existing anonymous interaction logs. The goal of training is to enable the agent to learn an optimal policy function that selects actions that maximize long-term positive feedback (such as increased user engagement or relief of negative emotions) based on the input user state; that is, an emotion support strategy.
[0065] Example 6: The multimodal control instructions are JSON format objects, containing digital human emotion expression tags, digital human action tags, voice parameters, and scene effect parameters; digital human rendering uses rendering technology that adapts to 2D or 3D styles.
[0066] Specifically, this design is used to clarify the technical implementation standards of multimodal feedback, define the format and core components of multimodal control commands, and determine the technical adaptation scope of digital human rendering. It provides executable technical specifications for the feedback control module to transform abstract emotion support strategies into concrete digital human interactive performances, and serves as a key technical bridge connecting strategy generation and actual interactive output.
[0067] Among them, the multimodal control instructions adopt JSON format objects. Due to their structured and easy-to-parse characteristics, they can clearly carry multi-dimensional parameters such as the digital human's emotional expression, actions, voice, and scene effects, ensuring that the instructions are efficiently transmitted between system modules and that the parameters are unambiguous.
[0068] The rendering principle of digital humans can select appropriate rendering technologies based on the different needs of different application scenarios to ensure the smoothness and visual adaptability of digital human image and motion rendering. For example, Live2D is used for 2D, and Unity / UnrealEngine is used for 3D.
[0069] This design avoids parameter confusion using the JSON format, ensuring the digital human can accurately receive and parse control requirements such as emotional expressions and actions. The selection of 2D / 3D rendering technologies covers different application scenarios, improving the system's compatibility with various hardware and environments. Multi-dimensional parameter synchronous control enables coordinated performance of the digital human's expressions, actions, voice, and scene effects, avoiding fragmented feedback and enhancing the emotional coherence perceived by the user.
[0070] Example 7: A multi-user emotion recognition and digital human feedback device based on behavioral data includes an identity recognition module, a data acquisition module, an emotion analysis module, a strategy generation module, a feedback control module, and a storage module; The identity recognition module is used to sense the approach of users. It uses RFID readers, image sensors and microphone arrays to confirm the identity of registered users in at least one of RFID, face recognition and voiceprint recognition. It assigns temporary session identifiers to visitors and establishes multi-user independent session channels through sound source localization and speaker log technology. The data acquisition module is used to capture the temporal dialogue behavior data stream of the user's interaction with the digital human in real time. The data stream includes speech prosody features, dialogue interaction patterns and linguistic features. The emotion analysis module is pre-set with an emotion analysis model that uses a sequence information processing model based on state space theory and a data dependency selection mechanism and a static context feature splicing architecture. It is used to receive dialogue behavior data and output the emotion state label of the user’s current multi-dimensional emotion and cognitive state category related to the interaction scenario. The strategy generation module is used to call the user's personal profile in the storage module, combine it with the emotion state tags, and generate emotion support strategies through the decision model. The feedback control module is used to generate multimodal control commands based on the emotion support strategy, drive the digital human to perform facial expressions, actions and voice output, and trigger scene effects; The storage module is used to store user profiles, conversation behavior data, and an emotion support strategy library.
[0071] Among them, the identity recognition module lays the foundation for interaction, and is responsible for sensing the approach of users, confirming the identity of registered users, assigning temporary session identifiers to visitors, and establishing independent session channels for multiple users through sound source localization and speaker log technology to achieve user differentiation and session isolation.
[0072] The data acquisition module is the data source for sentiment analysis, capturing in real time a temporal dialogue behavior data stream containing phonological prosody features, dialogue interaction patterns, and linguistic features.
[0073] The sentiment analysis module is responsible for data and sentiment transformation, outputting multi-dimensional sentiment and cognitive state labels related to the interaction scenario based on a preset model. The strategy generation module is the core of personalized decision-making, calling upon the user's personal profile stored in the module and combining the sentiment state labels to generate sentiment support strategies through the decision model, such as... Figure 4 As shown.
[0074] The feedback control module is key to interactive output. It generates multimodal control commands based on the strategy, drives the digital human to perform facial expressions, actions, and voice output, and triggers scene effects.
[0075] The storage module provides data support for the system, storing user profiles, dialogue behavior data, and an emotion support strategy library to ensure the data needs for strategy generation and subsequent iterations.
[0076] Specifically, the device operates according to a collaborative logic of sensing, data collection, analysis, decision-making, feedback, and storage: First, the identification module utilizes RFID... The system utilizes a reader, image sensor, and microphone array to verify registered user identities through multimodal recognition methods such as RFID, facial recognition, and voiceprint recognition. It assigns temporary IDs to visitors and uses sound source localization and speaker logging technologies to segment multi-user sessions and establish independent channels. Next, the data acquisition module continuously captures temporal dialogue behavior data between users and the digital human within these independent sessions. Then, the emotion analysis module inputs this data into a model based on state-space theory and a data-dependent selection mechanism, combined with a static context feature concatenation architecture. The model outputs emotion and cognitive state labels. Following this, the strategy generation module retrieves the corresponding user profile from the storage module, combines it with emotion labels, and generates appropriate emotion support strategies through a rule system or reinforcement learning agent decision-making model. Finally, the feedback control module translates the strategies into multimodal control commands, driving the digital human to perform emotional expressions and trigger scene effects. Finally, the storage module synchronously stores user profiles, dialogue behavior data, and emotion support strategies from the interaction process, providing a data foundation for subsequent interactions and model optimization.
[0077] This device effectively addresses the shortcomings of existing technologies in multi-user differentiation, non-invasive emotion recognition, and personalized feedback. The multi-user session isolation technology in the identity recognition module overcomes the limitations of traditional single-user interaction, adapting to simultaneous interaction by multiple people in public settings and avoiding session confusion.
[0078] The data collection module focuses on dialogue behavior data, without relying on intrusive sensors such as cameras and wearable devices, thus protecting user privacy from the source and increasing user acceptance.
[0079] The hybrid architecture model and multi-dimensional emotion tags of the emotion analysis module break through the limitations of traditional basic emotion recognition and achieve accurate judgment of the user's deep cognitive state.
[0080] The strategy generation module combines personal profiles to generate personalized strategies, and the multimodal output of the feedback control module enables the digital human to provide personalized feedback, enhancing the warmth of the interaction and reducing the uncanny valley effect.
[0081] The data source support of the storage module provides a foundation for subsequent model iteration and interaction optimization of the system, forming a closed-loop improvement mechanism to ensure the effectiveness and adaptability of the interaction in the long term.
[0082] In a further optimized version, the image sensor of the identity recognition module runs a lightweight face detection and recognition algorithm, and the microphone array runs a voiceprint recognition model to achieve rapid verification of registered user identities and differentiation of sessions for unregistered users.
[0083] Specifically, the technical implementation details of the image sensor and microphone array in the identity recognition module are clarified. By configuring lightweight algorithms and voiceprint models, the efficiency of registered user identity verification and the accuracy of session differentiation for unregistered users are enhanced. This provides technical support for the identity recognition module to complete the core functions of user identity confirmation and session attribution binding, ensuring the orderly development of subsequent multi-user independent interactions.
[0084] The image sensor runs lightweight face detection and recognition algorithms, such as MTCNN detection and MobileFaceNet recognition. These algorithms reduce hardware resource consumption by simplifying model parameters and optimizing the computation process, and can quickly capture facial features and compare them with the feature database of registered users; for example... Figure 3 As shown, the microphone array runs a voiceprint recognition model, such as the x-vector model, to complete identity matching by extracting unique voiceprint features from speech, such as spectrum and formants. The two work together to achieve multimodal identity verification for registered users by quickly comparing facial and voiceprint features. For unregistered users whose registration information is not matched, the differences in voiceprint features help to distinguish different users, assign temporary session identifiers to them and bind session data to ensure clear session ownership.
[0085] This design significantly reduces the time required for user authentication during registration, achieving near-seamless verification, eliminating user waiting time, and improving the convenience of interaction.
[0086] This design improves the accuracy of distinguishing unregistered user sessions, effectively avoids session data confusion in multi-user scenarios, and ensures the user-specificity of subsequent sentiment analysis and strategy generation.
[0087] This design enables lightweight algorithms to reduce the computing power and energy consumption requirements of devices, ensuring stable operation of modules under different hardware configurations and expanding the scope of device adaptability.
[0088] Further optimized, the storage module adopts a hybrid database architecture: a graph database is used to store user knowledge structure and learning path, a time-series database is used to store time-series dialogue behavior data streams, and a relational database or document database is used to store user profiles, emotion support strategy library and conversation records, and all stored data is bound to the user's unique identifier.
[0089] Specifically, the database architecture selection and data storage allocation rules of the storage module are clearly defined. Through a hybrid architecture design that adapts to the characteristics of different types of data, efficient management of multiple types of data such as user knowledge structure, time-series dialogue behavior, and personal profiles can be achieved. At the same time, all data is bound with a unique user identifier to ensure clear data ownership and provide support for accurate data retrieval by modules such as identity recognition and policy generation.
[0090] The logic of matching and adapting storage media to different types of data to optimize management efficiency lies in the fact that graph databases have the ability to efficiently store and query complex relational data, so they are used to store data with related attributes such as user knowledge structure and learning path.
[0091] Time-series databases offer advantages in high-concurrency writing and time-range querying for time-series data, making them suitable for storing time-series dialogue behavior data streams; relational / document databases excel at processing structured / semi-structured data, making them suitable for storing user profiles, emotion support strategy libraries, and other data; at the same time, by using unique user identifiers to associate various types of data, all data can be accurately traced back to the corresponding user, meeting the data access needs of each module.
[0092] This design improves data storage and retrieval efficiency, with each database performing its own function, avoiding the performance bottleneck of a single database handling multiple types of data, and ensuring the response speed when modules call data.
[0093] This design ensures accurate data ownership, and the user unique identifier binding mechanism prevents data confusion between different users, providing a reliable data foundation for personalized strategy generation and historical interaction tracing.
[0094] This design enhances system scalability, and the hybrid architecture can flexibly adapt to the new data type requirements, ensuring that the storage module supports the stable operation of the system in the long term.
[0095] In a preferred embodiment, the feedback control module includes a text-to-speech engine, which adjusts the speech rate, pitch, and volume of the synthesized speech according to the speech parameters in the multimodal control instructions to match the emotional tone corresponding to the emotion support strategy; the scene effects include switching the background style of the interactive interface and playing background music that matches the emotional atmosphere.
[0096] Among them, the details of the optimized technical implementation of the feedback control module in voice output and scene atmosphere creation are clarified. Through the parameter adjustment function of the text-to-speech engine and the specific form of scene effects, the voice expression and interactive environment atmosphere of the digital human are accurately matched with the emotional tone corresponding to the emotion support strategy, filling the gap in the emotional transmission of basic feedback forms and supporting the complete implementation of multimodal emotional feedback.
[0097] Specifically, the text-to-speech engine receives multimodal control commands generated by the feedback control module, extracts speech-related parameters from the commands, such as speech rate, pitch, and volume, and dynamically adjusts the physical attributes of the synthesized speech based on these parameters. For example, a slower speech rate, lower pitch, and softer volume are matched for soothing strategies.
[0098] Scene effects trigger preset background style changes for the interactive interface based on the emotional atmosphere defined by the emotional support strategy; for example, the relaxation strategy switches to a natural landscape background. Appropriate background music is also played, such as soft, instrumental music for the focus strategy. These two elements work together to ensure that the feedback format aligns with the emotional tone.
[0099] This design improves the accuracy of emotional expression in voice, avoids the monotonous expression of synthesized speech, and allows users to intuitively perceive the emotional tendencies of the digital human through voice, enhancing the interactive experience. Furthermore, it strengthens the immersive emotional atmosphere, with scene effects constructing a suitable interactive environment from visual and auditory dimensions to support the effectiveness of emotional support strategies. Finally, it reduces the disconnect between feedback and user emotions, making multimodal feedback more relevant to the user's current state, further increasing user acceptance of emotional support and strengthening the emotional connection.
[0100] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for multi-user emotion recognition and digital human feedback based on behavioral data, characterized in that, Includes the following steps: Step 1: Sensing user approach, confirming the identity of registered users and assigning a unique identifier to the current session through at least one of RFID, facial recognition, and voiceprint recognition, or assigning a temporary session identifier to unregistered visitors; when there are multiple users, distinguishing different users and establishing independent session channels through sound source localization and speaker log technology. Step 2: During the conversation, capture the temporal dialogue behavior data stream generated when the user interacts with the digital human in real time. The dialogue behavior data includes speech prosody features, dialogue interaction patterns, and linguistic features. Step 3: Input the dialogue behavior data into a pre-trained sentiment analysis model. The model adopts an architecture that combines a sequence information processing model based on state space theory and a data-dependent selection mechanism with static context features, and outputs the sentiment state label of the user's current sentiment and cognitive state category related to the interaction scenario. Step four: Based on the emotional state tags and the personal profile associated with the user's identity, generate an emotion support strategy through a decision model; Step five: Based on the emotion support strategy, generate and execute multimodal control instructions that drive the digital human to perform emotional expressions. These instructions simultaneously control the digital human's facial expressions, actions, voice parameters, and scene effects.
2. The method for multi-user emotion recognition and digital human feedback based on behavioral data according to claim 1, characterized in that: Sound source localization is achieved by calculating the time difference of sound reaching different microphones or by applying beamforming algorithms. Speaker logging technology analyzes the voiceprint features of a continuous speech stream and segments the speech stream into independent speech segments corresponding to different users in order to complete multi-user session affiliation binding.
3. The method for multi-user emotion recognition and digital human feedback based on behavioral data according to claim 1, characterized in that: Speech prosodic features include speech rate, pitch range, volume fluctuation, and pause duration and frequency; dialogue interaction patterns include user response delay, frequency of interruptions to the digital human, turn-taking patterns, and frequency of requests for clarification or repetition; linguistic features include the frequency of filler words and the type and complexity of user-asked questions.
4. The method for multi-user emotion recognition and digital human feedback based on behavioral data according to claim 1, characterized in that: The core model for processing temporal dialogue behavior data is a sequence information processing model based on state space theory and equipped with a data dependency selection mechanism; static context features include the difficulty of the current interaction task and the user's knowledge level; the emotional state labels correspond to multi-dimensional emotional and cognitive state categories, covering at least the state types related to the user's task participation, cognitive load, and emotional tendency; the sequence information processing model is the Mamba state space model.
5. The method for multi-user emotion recognition and digital human feedback based on behavioral data according to claim 1, characterized in that: The personal profile includes the user's historical interaction preferences, personalized characteristics related to the interaction scenario, and past emotional response patterns; the decision-making model is a rule system or reinforcement learning agent, and the emotion support strategy defines the intervention type, speech style, and guidance direction, and the intervention type is adapted to the user's state corresponding to the emotion state label.
6. The method for multi-user emotion recognition and digital human feedback based on behavioral data according to claim 1, characterized in that: The multimodal control instructions are JSON format objects, containing digital human emotion expression tags, digital human action tags, voice parameters, and scene effect parameters; digital human rendering uses rendering technology that adapts to 2D or 3D styles.
7. A multi-user emotion recognition and digital human feedback device based on behavioral data, used to implement the multi-user emotion recognition and digital human feedback method based on behavioral data as described in any one of claims 1-6, characterized in that: It includes an identity recognition module, a data acquisition module, a sentiment analysis module, a strategy generation module, a feedback control module, and a storage module; The identity recognition module is used to sense the approach of users. It uses RFID readers, image sensors and microphone arrays to confirm the identity of registered users in at least one of RFID, face recognition and voiceprint recognition. It assigns temporary session identifiers to visitors and establishes multi-user independent session channels through sound source localization and speaker log technology. The data acquisition module is used to capture the temporal dialogue behavior data stream of the user's interaction with the digital human in real time. The data stream includes speech prosody features, dialogue interaction patterns and linguistic features. The emotion analysis module is pre-set with an emotion analysis model that uses a sequence information processing model based on state space theory and a data dependency selection mechanism and a static context feature splicing architecture. It is used to receive dialogue behavior data and output the emotion state label of the user’s current multi-dimensional emotion and cognitive state category related to the interaction scenario. The strategy generation module is used to call the user's personal profile in the storage module, combine it with the emotion state tags, and generate emotion support strategies through the decision model. The feedback control module is used to generate multimodal control commands based on the emotion support strategy, drive the digital human to perform facial expressions, actions and voice output, and trigger scene effects; The storage module is used to store user profiles, conversation behavior data, and an emotion support strategy library.
8. The multi-user emotion recognition and digital human feedback device based on behavioral data according to claim 7, characterized in that: The image sensor in the identity recognition module runs a lightweight face detection and recognition algorithm, while the microphone array runs a voiceprint recognition model to achieve rapid verification of registered users' identities and differentiation of conversations with unregistered users.
9. A multi-user emotion recognition and digital human feedback device based on behavioral data according to claim 7, characterized in that: The storage module adopts a hybrid database architecture: a graph database is used to store user knowledge structure and learning path, a time-series database is used to store time-series dialogue behavior data streams, and a relational database or document database is used to store user profiles, emotion support strategy library and conversation records, and all stored data is bound to the user's unique identifier.
10. A multi-user emotion recognition and digital human feedback device based on behavioral data according to claim 7, characterized in that: The feedback control module includes a text-to-speech engine, which adjusts the speech rate, pitch, and volume of the synthesized speech according to the speech parameters in the multimodal control commands to match the emotional tone corresponding to the emotion support strategy; scene effects include switching the background style of the interactive interface and playing background music that matches the emotional atmosphere.
Citation Information
Cited By
Intelligent interaction system
CN122174843A