Digital human emotion perception and dynamic response method, system, device and storage medium

By combining multimodal affective computing with hotel scenario knowledge graphs, the emotional perception and response capabilities of hotel human-computer interaction systems have been improved, enabling accurate emotion recognition and personalized services while ensuring data security and cross-cultural adaptability.

CN119884327BActive Publication Date: 2025-12-16BEIJING VCONTROL TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510361255.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-12-16
Estimated Expiration
2045-03-26

AI Technical Summary

Technical Problem

Existing hotel human-computer interaction systems have limitations in terms of emotion perception and response, including insufficient emotion perception, weak contextual understanding and emotional association capabilities, rigid emotion response mechanisms, and poor cross-cultural adaptability, resulting in a disconnect between emotion and context and incomplete protection of implicit sensitive information.

Method used

A multimodal emotion computing model is adopted, which combines speech, facial expression and body language analysis. An emotion state classification model is constructed through deep learning and graph neural networks. A human-like response strategy is generated by combining hotel scene knowledge graph. Multimodal features are weighted and fused through attention mechanism to realize emotion intensity judgment and dynamic adjustment.

Benefits of technology

It improves the accuracy of emotion recognition, enhances the ability to capture multi-dimensional emotional information of customers, provides personalized and human-like responses, improves the affinity of interaction, and ensures the secure storage and processing of customer emotional data, meeting privacy protection requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119884327B_ABST
    Figure CN119884327B_ABST
Patent Text Reader

Abstract

The embodiment of the application discloses a kind of digital person's emotional perception and dynamic response method, system, equipment and storage medium, the embodiment of the application can realize accurate emotional perception, improve the identification precision of multi-dimensional emotional information such as customer voice, expression, text, accurately capture complex emotional state;According to customer emotional intensity, cultural background and dialogue context, generate personified voice, expression and solution, enhance interaction affinity;Combined with hotel industry knowledge graph and high-frequency service scene, provide personalized, high-target emotional service;Ensure the safe storage and processing of customer emotional data, meet the privacy protection regulations requirements.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The embodiment of the application relates to the technical field of human-computer interaction, and particularly relates to a digital person emotional perception and dynamic response method, system, device and storage medium. BACKGROUND

[0002] In the hotel service scenario, the traditional human-computer interaction system mainly relies on a rule-based or simple natural language processing (NLP) dialogue system. These systems usually use keyword matching or pre-defined dialogue flows to realize customer service functions such as room booking, information inquiry and common question answering. In recent years, with the development of artificial intelligence technology, some hotels have begun to introduce voice assistants and digital human images to improve customer experience. However, the existing technology has significant limitations in emotional perception and response, and the specific problems and shortcomings are as follows:

[0003] 1) Insufficient emotional perception: existing systems mainly rely on single-modal analysis of text or voice, lacking multi-dimensional perception ability of customer emotions. For example, only through voice tone or text keywords to judge customer emotions, it is difficult to capture non-verbal cues such as facial expressions and body language;

[0004] 2) Weak context understanding and emotional association ability: it is difficult to maintain context consistency in multi-turn dialogue, and it is difficult to dynamically associate customer emotional state with dialogue history and scene background. For example, when a customer repeatedly complains about a room problem, the system may repeatedly provide the same solution, leading to escalation of customer emotions;

[0005] 3) Emotional response mechanism is rigid: the response is usually based on fixed rules or simple templates, lacking dynamic adjustment ability of personification. For example, when a customer expresses dissatisfaction, the system may only provide standardized apology statements, and cannot convey empathy through voice tone, facial expressions, etc.

[0006] At the same time, single-modal emotion recognition has certain limitations. Speech emotion recognition (SER) is susceptible to environmental noise, and text emotion analysis is difficult to handle implicit emotional expressions in high-context cultures. Traditional dialogue management systems (DMS) lack the ability to remember long dialogue context, and it is difficult to capture emotional trends after losing context. The existing technology does not deeply combine emotional analysis with hotel scene knowledge (such as customer preferences, service history), resulting in lack of pertinence of emotional response, causing the problem of emotional-scene disconnection. The existing technology usually only performs simple replacement on explicit sensitive information (such as phone numbers), and does not protect implicit sensitive information (such as emotional state), and the desensitization of data is not thorough enough. SUMMARY

[0007] To this end, the embodiment of the present application provides a digital person emotional perception and dynamic response method, system, device and storage medium to solve the technical problems of insufficient emotional perception ability, rigid response mechanism and poor cross-cultural adaptability of the prior art.

[0008] To achieve the above object, the embodiment of the present application provides the following technical scheme:

[0009] According to a first aspect of the embodiment of the present application, a digital person emotional perception and dynamic response method is provided, which is applied to a hotel digital person and includes:

[0010] S1, receiving interactive information of a user and collecting multi-modal data, extracting a plurality of modal features from the multi-modal data and weighting and fusing each modal feature by using an attention mechanism to generate a comprehensive emotional feature vector;

[0011] S2, constructing a deep learning model and training the deep learning model by using preset emotional state data to generate an emotional state classification model;

[0012] S3, using the emotional state classification model and the comprehensive emotional feature vector to judge the emotional intensity, and when the emotional intensity is a medium or low intensity emotion, generating an emotional response strategy in combination with the dialogue history and the hotel scene knowledge graph;

[0013] S4, generating a personification behavior by using the emotional response strategy, outputting the personification behavior, collecting interactive feedback of the user, detecting emotional change, dynamically adjusting the response strategy, storing the interactive data of this time and generating a service report.

[0014] Further, the multi-modal data includes interactive voice of the user and facial expression data and body language of the user.

[0015] Further, the multi-modal data includes interactive voice of the user and facial expression data and body language of the user.

[0016] A multi-modal emotional perception network is constructed by integrating speech emotion recognition, facial expression analysis, text emotion analysis and body language capture;

[0017] The feature extraction of the speech spectrum is performed by using a deep convolutional neural network model, and the time sequence emotional change is captured by combining a long short-term memory network;

[0018] The facial expression data is acquired, and the facial emotional change of the user is captured by using a micro-expression recognition algorithm;

[0019] The dialogue analysis is performed by using a language model to generate language emotional change;

[0020] The voice spectrum features, the timing emotional changes, the facial emotional changes and the language emotional changes are weighted and fused through an attention mechanism to generate a comprehensive emotional feature vector.

[0021] Further, the emotional intensity is judged using the emotional state classification model and the comprehensive emotional feature vector, and when the emotional intensity is a medium-low intensity emotion, an emotional response strategy is generated in combination with the dialogue history and the hotel scene knowledge graph, including:

[0022] A dialogue context modeling framework based on a graph neural network is constructed to dynamically associate the user's emotional state with the dialogue history and the scene background;

[0023] The multi-round dialogue information is stored through a memory enhancement network for long-time emotional state tracking;

[0024] The hotel scene knowledge graph is obtained and used for scene knowledge enhancement to generate an emotional response strategy;

[0025] In the process of judging the emotional intensity, the emotional state is classified through an emotional intensity quantification model, and the emotional state of the emotional intensity quantification model is divided into a first preset number of large categories and a second preset number of subcategories.

[0026] Further, the emotional response strategy is used to generate a personification behavior, and the personification behavior is output and the user's interactive feedback is collected to dynamically adjust the response strategy according to the emotional changes, including:

[0027] A digital human behavior generation engine driven by emotions is developed to dynamically adjust the voice tone, facial expression and body movement of the digital human, specifically including:

[0028] The synthesized voice is adjusted based on the emotional intensity to adjust the voice synthesis parameters, including pitch and speech rate;

[0029] Through an emotion-action mapping rule base, facial expressions and body movements matching the emotional state are generated, and the facial expressions of the robot are rendered according to the emotional intensity;

[0030] The model parameters are dynamically updated through the gradient descent method, and the real-time emotional state monitoring algorithm is combined with an online learning mechanism to adjust the response strategy according to the user's immediate interactive feedback data;

[0031] The immediate interactive feedback data includes voice interruptions and expression changes.

[0032] Further, the method further includes:

[0033] The emotional intensity is judged using the emotional state classification model and the comprehensive emotional feature vector, and when the emotional intensity is a high-intensity emotion, an emergency response mechanism is triggered to calm the user's emotions.

[0034] Further, the storage of the interaction data this time includes:

[0035] The interaction data is desensitized in real time, and the emotional data and the business data are separated;

[0036] The emotional data and the business data are stored in separate databases, the data is encrypted using a preset encryption algorithm, and all data access behaviors are recorded;

[0037] The desensitization operation is to dynamically filter sensitive information based on role permissions.

[0038] According to a second aspect of an embodiment of the present application, a digital person emotional perception and dynamic response system is provided, the system comprising:

[0039] The acquisition module is configured to receive interaction information of a user and acquire multi-modal data, extract a plurality of modal features from the multi-modal data, and generate a comprehensive emotional feature vector by weighting and fusing each modal feature using an attention mechanism.

[0040] The model construction module is configured to construct a deep learning model and train the deep learning model using preset emotional state data to generate an emotional state classification model.

[0041] The emotional response strategy generation module is configured to use the emotional state classification model and the comprehensive emotional feature vector to determine the emotional intensity, and when the emotional intensity is a medium or low intensity emotion, generate an emotional response strategy in combination with the dialogue history and the hotel scene knowledge graph.

[0042] The interaction feedback module is configured to generate a personification behavior using the emotional response strategy, output the personification behavior, acquire interaction feedback of the user, detect emotional changes, dynamically adjust the response strategy, store the interaction data this time, and generate a service report.

[0043] According to a third aspect of an embodiment of the present application, a digital person emotional perception and dynamic response device is provided, the device comprising a processor and a memory;

[0044] The memory is configured to store one or more program instructions;

[0045] The processor is configured to run one or more program instructions to perform the steps of the digital person emotional perception and dynamic response method according to any one of the above.

[0046] According to a fourth aspect of an embodiment of the present application, a computer readable storage medium is provided, the computer readable storage medium storing a computer program, the computer program being executed by a processor to implement the steps of the digital person emotional perception and dynamic response method according to any one of the above.

[0047] The embodiment of the present application has the following advantages:

[0048] The embodiment of the present application comprises: S1, receiving interactive information of a user and collecting multi-modal data, extracting a plurality of modal features from the multi-modal data and weighting and fusing each modal feature by using an attention mechanism to generate a comprehensive emotional feature vector; S2, constructing a deep learning model and training the deep learning model by using preset emotional state data to generate an emotional state classification model; S3, judging an emotional intensity by using the emotional state classification model and the comprehensive emotional feature vector, and when the emotional intensity is a medium-low intensity emotion, generating an emotional response strategy in combination with a dialogue history and a hotel scene knowledge graph; S4, generating a personification behavior by using the emotional response strategy, outputting the personification behavior, collecting interactive feedback of the user, detecting emotional changes to dynamically adjust the response strategy, storing this time interactive data and generating a service report. The embodiment of the present application can realize accurate emotional perception, improve the recognition accuracy of multi-dimensional emotional information such as customer voice, expression and text, accurately capture complex emotional states; generate personified voice, expression and solutions according to the customer emotional intensity, cultural background and dialogue context, enhance the interaction affinity; provide personalized and high-target emotional services in combination with the hotel industry knowledge graph and high-frequency service scenes; ensure the safe storage and processing of customer emotional data, and meet the requirements of privacy protection regulations. BRIEF DESCRIPTION OF DRAWINGS

[0049] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are only exemplary, and for those skilled in the art, other drawings can be obtained without creative labor on the basis of the provided drawings.

[0050] The structures, proportions, sizes, etc. shown in the specification are only used to cooperate with the content disclosed in the specification, to be understood and read by those skilled in the art, and do not define the limiting conditions for the implementation of the present application, so they do not have technical significance. Any modification of structure, change of proportion relationship or adjustment of size, without affecting the effect and purpose that the present application can produce, should still fall within the scope of the technical content disclosed by the present application.

[0051] Figure 1 A logical structure schematic diagram of a digital person emotional perception and dynamic response system provided by the embodiment of the present application;

[0052] Figure 2 A flowchart schematic diagram of a digital person emotional perception and dynamic response method provided by the embodiment of the present application;

[0053] Figure 3 The overall architecture diagram of a digital human emotional perception and dynamic response method provided for an embodiment of the present application. DETAILED DESCRIPTION

[0054] The embodiments of the present application will be described in detail by specific embodiments, and those skilled in the art can easily understand other advantages and effects of the present application from the disclosed content of the specification. Obviously, the described embodiments are part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0055] In the hotel service scenario, traditional human-computer interaction systems mainly rely on rule-based or simple natural language processing (NLP) dialogue systems. These systems usually use keyword matching or pre-defined dialogue flows to achieve customer service functions such as room booking, information inquiry, and common question answering. In recent years, with the development of artificial intelligence technology, some hotels have begun to introduce voice assistants and digital human images to improve customer experience. However, existing technologies have significant limitations in emotional perception and response, with specific problems and shortcomings as follows:

[0056] 1) Insufficient emotional perception: existing systems mainly rely on single-modal analysis of text or voice, lacking multi-dimensional perception ability of customer emotions. For example, only through voice tone or text keywords to judge customer emotions, unable to capture non-verbal cues such as facial expressions and body language;

[0057] 2) Weak context understanding and emotional association ability: difficult to maintain context consistency in multi-turn dialogue, unable to dynamically associate customer emotional state with dialogue history, scene background. For example, when a customer repeatedly complains about room problems, the system may repeatedly provide the same solution, leading to escalation of customer emotions;

[0058] 3) Emotional response mechanism is rigid: response is usually based on fixed rules or simple templates, lacking dynamic adjustment ability of personification. For example, when a customer expresses dissatisfaction, the system may only provide standardized apology statements, unable to convey empathy through voice tone, facial expressions, etc.

[0059] At the same time, single-modal emotion recognition has certain limitations, speech emotion recognition (SER) is susceptible to environmental noise, and text sentiment analysis is difficult to handle implicit emotional expression in high-context culture. Traditional dialog management systems (DMS) lack the ability to remember long-term dialog context, and it is difficult to capture the trend of emotional changes after losing the context. The existing technology does not deeply combine emotion analysis with hotel scene knowledge (such as customer preferences and service history), resulting in a lack of targeted emotional response and a problem of emotion-scene disconnection. The existing technology usually only performs simple replacement on explicit sensitive information (such as phone numbers), and does not protect implicit sensitive information (such as emotional state), and the desensitization of data is not thorough enough.

[0060] In order to solve the technical problems of insufficient emotion perception ability, rigid response mechanism and poor cross-cultural adaptability.

[0061] Reference Figure 1 The embodiment of the present application discloses a digital person's emotion perception and dynamic response system, which comprises: a collection module 1; a model construction module 2; an emotional response strategy generation module 3; and an interactive feedback module 4.

[0062] The technical category of the embodiment of the present application covers: a multi-modal emotion computing model based on deep learning, a semantic-emotion association analysis algorithm in a cross-cultural context, a real-time emotional feedback driven digital person behavior generation engine, and a privacy protection type dialog data processing architecture. The embodiment of the present application is directly applied to the hotel customer service scene, integrates speech emotion recognition (SER), micro-expression capture, dialog text sentiment analysis, multi-turn dialog management technology with context awareness, and innovatively integrates hotel industry knowledge graph and cross-cultural emotion expression rule library, and is suitable for emotional interaction of intelligent digital person in hotel front desk consultation, room service, complaint handling and other links, realizes the full-link closed loop from emotion recognition (such as anger, satisfaction, anxiety) to personification response (voice tone adjustment, facial expression rendering, solution recommendation), and meets the reliability and affinity requirements of the emotional human-computer interaction system in ISO 9241-210 standard.

[0063] Corresponding to the above disclosed digital person's emotion perception and dynamic response system, the embodiment of the present application also discloses a digital person's emotion perception and dynamic response method. The following will introduce the digital person's emotion perception and dynamic response method disclosed in the embodiment of the present application in detail in combination with the above described digital person's emotion perception and dynamic response system.

[0064] The embodiment of the application aims to solve the problem of insufficient emotional perception and response capability of digital human interaction system in hotel service scenarios, and innovatively combines multi-modal emotion computing, context perception technology and personification behavior generation engine to improve the technical problems of insufficient emotion perception capability and rigid response mechanism of the prior art. It is committed to creating an intelligent, emotional, safe and reliable hotel digital human perception and dynamic response system, which significantly improves customer service experience and hotel operation efficiency.

[0065] Reference Figures 2 to 3 The application discloses a kind of digital human emotional perception and dynamic response method, the method is applied to hotel digital human, it includes:

[0066] S1, receive the interactive information of user and collect multi-modal data, extract multiple modal features from the multi-modal data and utilize attention mechanism to each modal feature Weighted fusion, generate comprehensive emotional feature vector;

[0067] S2, construct deep learning model and utilize preset emotional state data to train the deep learning model, generate emotional state classification model;

[0068] S3, utilize the emotional state classification model and the comprehensive emotional feature vector to judge emotional intensity, when emotional intensity is low intensity emotion, generate emotional response strategy in combination with dialogue history and hotel scene knowledge graph;

[0069] S4, utilize the emotional response strategy to generate personification behavior, output the personification behavior and collect the interactive feedback of user, detect emotional change Dynamic adjustment response strategy, store this interaction data and generate service report.

[0070] Further, the multi-modal data includes the interactive voice of user and the facial expression data and body language of user.

[0071] Further, extract multiple modal features from the multi-modal data and utilize attention mechanism to each modal feature Weighted fusion, generate comprehensive emotional feature vector, including: by integrating speech emotion recognition, facial expression analysis, text sentiment analysis and body language capture Construct multi-modal emotion perception network;Utilize deep convolutional neural network model to carry out feature extraction of voice frequency spectrum and combine long short-term memory network to capture time sequence emotional change;Obtain facial expression data and capture the facial emotional change of user by micro-expression recognition algorithm;Utilize language model to carry out dialogue analysis, generate language emotional change;Voice frequency spectrum features, time sequence emotional change, facial emotional change and language emotional change are weighted fusion through attention mechanism, generate comprehensive emotional feature vector.

[0072] Integrate voice sentiment recognition (SER), facial expression analysis, text sentiment analysis, and body language capture technology to build a multi-modal sentiment perception network.

[0073] Voice sentiment recognition: Use deep convolutional neural network (CNN) based voice spectrum feature extraction combined with long short-term memory network (LSTM) to capture temporal sentiment changes.

[0074] Facial expression analysis: Capture subtle emotional changes of customers through facial action unit (AU) detection and micro-expression recognition algorithms.

[0075] Text sentiment analysis: Context-aware sentiment classification based on pre-trained language models (such as BERT) to support implicit sentiment analysis in high-context cultures.

[0076] The embodiment of the present application proposes a multi-modal feature fusion mechanism, dynamically weights each modal feature through attention mechanism, and improves the sentiment recognition accuracy (>92% F1 value). Introduce a sentiment intensity quantification model to divide the sentiment state into 9 categories and 32 subcategories, supporting more fine-grained sentiment classification.

[0077] Further, the emotion state classification model and the comprehensive emotion feature vector are used to judge the emotion intensity, and when the emotion intensity is a low-intensity emotion, a sentiment response strategy is generated by combining the dialogue history and the hotel scene knowledge graph, including: building a dialogue context modeling framework based on graph neural network, dynamically associating the user's emotional state with the dialogue history and scene background; store multi-round dialogue information through memory enhancement network, and track long-time emotional state; obtain the hotel scene knowledge graph and use the hotel scene knowledge graph for scene knowledge enhancement to generate a sentiment response strategy.

[0078] Build a dialogue context modeling framework based on graph neural network (GNN) to dynamically associate the customer's emotional state with the dialogue history and scene background.

[0079] Dialogue context modeling: Store multi-round dialogue information through memory enhancement network (Memory Network) to support long-time emotional state tracking.

[0080] Scene knowledge enhancement: Combine with hotel industry knowledge graph (such as customer preferences, service history) to provide sentiment-based service suggestions.

[0081] Propose a sentiment-scene correlation rule engine to dynamically adjust the emotional response strategy. For example, for customers who have complained multiple times, automatically trigger the escalation service process.

[0082] Introduce a sentiment decay model to simulate the characteristics of human emotions changing over time, avoiding emotional response lag.

[0083] The emotion intensity quantification model classifies the emotion state into a first preset number of large categories and a second preset number of subcategories. The first preset number is 9, and the second preset number is 32.

[0084] Further, the emotional response strategy is used to generate a personification behavior, the personification behavior is output, the interactive feedback of the user is collected, the emotional change is detected, and the response strategy is dynamically adjusted, including: developing an emotion-driven digital human behavior generation engine, dynamically adjusting the voice tone, facial expression and body movement of the digital human, specifically including: synthesizing voice and adjusting voice synthesis parameters based on emotion intensity, the voice synthesis parameters including pitch and speech rate; generating facial expressions and body movements matched with the emotion state through an emotion-action mapping rule library, rendering the facial expression of the robot according to the emotion intensity; dynamically updating the model parameters through the gradient descent method, combining the real-time emotion state monitoring algorithm with the online learning mechanism, and adjusting the response strategy according to the instant interactive feedback data of the user.

[0085] The instant interactive feedback data includes voice interruption and expression change.

[0086] Developing an emotion-driven digital human behavior generation engine to dynamically adjust voice tone, facial expression and body movement:

[0087] Voice synthesis: adjusting voice synthesis parameters (such as pitch and speech rate) based on emotion intensity to deliver empathy effect.

[0088] Expression and action generation: generate facial expressions and body movements matched with the emotion state through an emotion-action mapping rule library.

[0089] Propose an emotional response diversity mechanism to generate differentiated response content according to emotion intensity and cultural background. Introduce a real-time emotional feedback mechanism to dynamically adjust the response strategy through customer expressions and voice changes.

[0090] Further, the method further includes: using the emotion state classification model and the comprehensive emotion feature vector to judge the emotion intensity, and when the emotion intensity is high-intensity emotion, triggering an emergency response mechanism and calming the user's mood.

[0091] Further, storing the interaction data includes: performing a desensitization operation on the interaction data in real time and separating emotion data and business data; storing the emotion data and business data in separate databases, encrypting the data using a preset encryption algorithm, and recording all data access behaviors.

[0092] The desensitization operation is to filter sensitive information based on role permissions dynamically.

[0093] A layered data desensitization and encrypted storage architecture is designed to ensure the security and compliance of customer sentiment data. An emotion data anonymization processing technology is proposed to ensure that customer privacy is not disclosed during sentiment analysis. A data access audit mechanism is introduced to record all data access behaviors, meeting the requirements of privacy protection regulations such as GDPR.

[0094] Real-time data desensitization: dynamically filter sensitive information (such as identity information, payment information) based on role permissions.

[0095] Sharded storage mechanism: isolate sentiment data and business data for storage, and use AES-256 encryption algorithm to protect data security.

[0096] The embodiment of the application also has cross-cultural support function, adapts to the difference of emotional expression in multi-language and multi-cultural background, and improves the customer satisfaction in global hotel service scene.

[0097] A cross-cultural emotional expression rule library is constructed to support emotional interaction in multi-language and multi-cultural background. A cultural context perception mechanism is proposed to dynamically adjust the sentiment recognition threshold. For example, the implicit sentiment recognition sensitivity is improved for East Asian customers. Cross-cultural sentiment mapping rules are introduced to ensure that the emotional response meets the customer's cultural background.

[0098] Cultural difference compensation: adjust the sentiment recognition and response strategy according to the emotional expression habits of customers with different cultural backgrounds.

[0099] Multi-language support: integrate multi-language sentiment analysis models to support emotional interaction in mainstream languages such as English, Chinese and Japanese.

[0100] In the embodiment of the application, the following advantages are achieved: high-precision emotion perception: through multi-modal fusion and context association, the emotion recognition accuracy and scene adaptation ability are significantly improved; anthropomorphic interaction experience: dynamically generate emotional voice, expression and action to enhance customer interaction experience; safety and compliance: use layered data desensitization and encryption storage technology to ensure customer privacy security; cross-cultural support: adapt to multi-language and multi-cultural background to meet the needs of global hotel service.

[0101] Through the above technical solutions, the application realizes a full-link closed loop from emotion perception to anthropomorphic response, and provides an intelligent, emotional, safe and reliable digital human interaction solution for hotel service scenarios.

[0102] The embodiment of the application has the following advantages:

[0103] 1) The fusion of multi-modal emotion perception is enhanced, which significantly improves the accuracy and scene adaptability of emotion recognition:

[0104] Adopting multi-modal data fusion technology, the user's facial expressions, voice tone, body movements (such as arm posture), and conversation text are synchronously collected through cameras, microphones, and text inputs, and are jointly trained in combination with deep learning models (such as LSTM and attention mechanism). For example, micro-expressions are analyzed through a facial action coding system (FACS), combined with the fundamental frequency (F0) of the voice and the semantic of the text, to form a comprehensive emotional judgment. The problem of single-modal data being easily disturbed by noise is solved, and especially in the hotel service scenario, the implicit emotions of customers (such as dissatisfaction in a neutral state) can be more accurately identified.

[0105] 2) Dynamic real-time emotional response optimization, realizing low-latency personalized feedback and improving the naturalness of interaction:

[0106] The real-time emotional state monitoring algorithm (REM) is introduced, the model parameters are dynamically updated through gradient descent method, and the response strategy is adjusted according to the real-time data of user interaction (such as voice interruption and expression change) in combination with online learning mechanism. For example, when the customer expresses the demand, the system optimizes the feedback content by real-time calculation of emotional label difference (such as the loss function in formula (1)); the response time is shortened to milliseconds, and can adapt to sudden emotional changes (such as the customer from calm to anxious), and provides more scene-adaptive service suggestions.

[0107] 3) Empathy ability improvement based on common sense knowledge graph, generating more logical and context-related conversation content:

[0108] The open-domain dialogue model of the prior art relies on large-scale pre-training language models (such as GPT and PLATO), but lacks common sense reasoning ability, and is prone to logical contradictions (such as "cows going to college"), the embodiment of the present invention constructs a social common sense knowledge graph, including event causal relationship (such as "thirst -> buy water"), social intention (such as "engagement needs ring"), and the like triplets, and fuses the dialogue history and graph knowledge through the BART model. For example, when the customer mentions "birthday gift", the system combines the graph to infer the "expectation" emotion, generates a "the other party will definitely like" empathetic reply, reduces the reply that violates common sense, and enhances the coherence of the dialogue and the user's trust.

[0109] 4) Multi-scene adaptive service improvement suggestion generation, providing practical business optimization solutions, and surpassing basic interaction functions:

[0110] Traditional hotel feedback systems rely on manual analysis of questionnaire data, which is inefficient and subjective. Embodiments of the present application improve the coefficient calculation model, quantify the preference of customers of different age groups for dishes and service experience, and generate improvement suggestions in combination with emotional recognition results (such as high frequency of "neutral" expressions). For example, if it is identified that young customers have a high frequency of "frowning" on a certain dish, the system automatically calculates the "negative experience coefficient" of the dish and recommends adjusting the taste or plating. Converting emotional data into executable operation strategies helps hotels improve customer satisfaction and repeat purchase rate.

[0111] In addition, the embodiment of the present application also provides a digital person emotional perception and dynamic response device, the device comprises: a processor and a memory; the memory is used to store one or more program instructions; the processor is used to run one or more program instructions to execute the steps of the digital person emotional perception and dynamic response method according to any one of the above.

[0112] In addition, the embodiment of the present application also provides a computer readable storage medium, the computer readable storage medium stores a computer program, and the computer program is executed by a processor to realize the steps of the digital person emotional perception and dynamic response method according to any one of the above.

[0113] In the embodiment of the present application, the processor can be an integrated circuit chip with signal processing capability. The processor can be a general processor, a digital signal processor (Digital Signal Processor, DSP), an application specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field programmable gate array (Field Programmable Gate Array, FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component.

[0114] The disclosed methods, steps and logic block diagrams in the embodiments of the present application can be implemented or executed. The general processor can be a microprocessor or the processor can be any conventional processor. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as hardware code processor execution or executed by a combination of hardware and software modules in the code processor. The software module can be located in a random memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register or other mature storage medium in the art. The processor reads the information in the storage medium and combines the hardware to complete the steps of the above method.

[0115] The storage medium can be a memory, for example, can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories.

[0116] The non-volatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically EPROM (EEPROM), or a flash memory.

[0117] The volatile memory can be a Random Access Memory (RAM) used as an external cache memory. By way of example, and not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct Rambus RAM (DRRAM).

[0118] The storage media described in the embodiments of the present application are intended to include, but are not limited to these and any other suitable types of memory.

[0119] Those skilled in the art should be aware that the functions described in the embodiments of the present application can be implemented in combination of hardware and software in one or more of the above examples. When the software is applied, the corresponding functions can be stored in a computer readable medium or transmitted as one or more instructions or codes on the computer readable medium. The computer readable medium includes a computer storage medium and a communication medium, wherein the communication medium includes any medium that facilitates the transfer of computer programs from one place to another. The storage medium can be any available medium accessible by a general or special purpose computer.

[0120] Although the present application has been described in detail with reference to the general description and the specific embodiments above, some modifications or improvements can be made to the present application on the basis of the present application, which is obvious to those skilled in the art. Therefore, these modifications or improvements made on the basis of not deviating from the spirit of the present application, are within the scope of the present application.

Claims

1. A digital human emotional perception and dynamic response method, characterized in that, The method is applied to a hotel digital person, and the method comprises the following steps: S1, receiving interactive information of a user and collecting multi-modal data, extracting a plurality of modal features from the multi-modal data, and weighting and fusing each modal feature by using an attention mechanism to generate a comprehensive emotional feature vector; S2, constructing a deep learning model and training the deep learning model by using preset emotional state data to generate an emotional state classification model; S3, judging an emotional intensity by using the emotional state classification model and the comprehensive emotional feature vector, and when the emotional intensity is a medium or low intensity emotion, generating an emotional response strategy in combination with a dialogue history and a hotel scene knowledge graph; S4, generating a personification behavior by using the emotional response strategy, outputting the personification behavior, collecting interactive feedback of the user, detecting emotional changes to dynamically adjust the response strategy, storing this time of interactive data, and generating a service report; extracting a plurality of modal features from the multi-modal data and weighting and fusing each modal feature by using an attention mechanism to generate a comprehensive emotional feature vector, comprising: constructing a multi-modal emotional perception network by integrating voice emotion recognition, facial expression analysis, text emotion analysis, and body language capture; using a deep convolutional neural network model to extract voice spectrum features and combining a long short-term memory network to capture temporal emotional changes; obtaining facial expression data and capturing facial emotional changes of the user by using a micro-expression recognition algorithm; using a language model to analyze a dialogue and generating language emotional changes; weighting and fusing voice spectrum features, temporal emotional changes, facial emotional changes, and language emotional changes by using an attention mechanism to generate a comprehensive emotional feature vector; storing this time of interactive data, comprising: performing a desensitization operation on interactive data in real time and separating emotional data and business data; isolating and storing the emotional data and the business data in different databases, encrypting the data by using a preset encryption algorithm, and recording all data access behaviors; wherein the desensitization operation is to filter sensitive information based on role permissions dynamically; the multi-modal data comprises interactive voice of the user and facial expression data and body language of the user; judging an emotional intensity by using the emotional state classification model and the comprehensive emotional feature vector, and when the emotional intensity is a medium or low intensity emotion, generating an emotional response strategy in combination with a dialogue history and a hotel scene knowledge graph, comprising: constructing a dialogue context modeling framework based on a graph neural network, dynamically associating emotional states of the user with the dialogue history and the scene background; storing multi-round dialogue information by using a memory enhancement network to track emotional states for a long time; obtaining a hotel scene knowledge graph and using the hotel scene knowledge graph for scene knowledge enhancement to generate an emotional response strategy; wherein when judging the emotional intensity, the emotional states also need to be classified by using an emotional intensity quantification model, and the emotional intensity quantification model has a first preset number of large categories and a second preset number of subcategories.

2. The method for emotional perception and dynamic response of a digital human according to claim 1, wherein, generating a personification behavior by using the emotional response strategy, outputting the personification behavior, collecting interactive feedback of the user, detecting emotional changes to dynamically adjust the response strategy, comprising: Developing an emotion-driven digital human behavior generation engine that dynamically adjusts the digital human's voice tone, facial expression, and body movement, including: Synthesizing speech and adjusting speech synthesis parameters based on emotional intensity, including pitch and speech rate; Generating facial expressions and body movements that match the emotional state through an emotion-action mapping rule base, rendering the robot's facial expressions according to emotional intensity; Dynamically updating model parameters through gradient descent, combining real-time emotional state monitoring algorithms with online learning mechanisms, and adjusting response strategies based on user's immediate interactive feedback data; Wherein the immediate interactive feedback data includes voice interruptions and expression changes.

3. The method for emotional perception and dynamic response of a digital human according to claim 2, wherein, The method further includes: Using the emotional state classification model and the comprehensive emotional feature vector to determine the emotional intensity, and triggering an emergency response mechanism and soothing the user's emotions when the emotional intensity is high.

4. A digital human emotional perception and dynamic response system, characterized in that, The system includes: A collection module for receiving user interaction information and collecting multi-modal data, extracting multiple modal features from the multi-modal data, and generating a comprehensive emotional feature vector by weighting and fusing each modal feature using an attention mechanism; A model construction module for constructing a deep learning model and training the deep learning model using pre-set emotional state data to generate an emotional state classification model; An emotional response strategy generation module for using the emotional state classification model and the comprehensive emotional feature vector to determine the emotional intensity, and generating an emotional response strategy by combining the dialogue history and hotel scene knowledge graph when the emotional intensity is medium or low; An interactive feedback module for generating a personified behavior using the emotional response strategy, outputting the personified behavior, collecting user interaction feedback, detecting emotional changes to dynamically adjust the response strategy, storing this interaction data, and generating a service report; From the multi-modal data, extract multiple modal features and use an attention mechanism to weight and fuse each modal feature to generate a comprehensive emotional feature vector, including: Construct a multi-modal emotional perception network by integrating speech emotion recognition, facial expression analysis, text sentiment analysis, and body language capture; Use a deep convolutional neural network model to extract features from the speech spectrum and combine a long short-term memory network to capture temporal emotional changes; Obtain facial expression data and capture user facial emotional changes through micro-expression recognition algorithms; Use a language model to analyze the dialogue and generate language emotional changes; Weight and fuse the speech spectrum features, temporal emotional changes, facial emotional changes, and language emotional changes through an attention mechanism to generate a comprehensive emotional feature vector; Storing this interaction data includes: Desensitizing the interaction data in real time and separating emotional data from business data; Isolate the emotional data and business data in separate databases, encrypt the data using a pre-set encryption algorithm, and record all data access behaviors; Wherein, the desensitization operation is to filter sensitive information based on role permissions dynamically; The multi-modal data includes user interaction speech and user facial expression data and body language; The emotional intensity is judged by using the emotional state classification model and the comprehensive emotional feature vector. When the emotional intensity is a medium or low intensity emotion, an emotional response strategy is generated by combining the dialogue history and the hotel scene knowledge graph, including: A dialogue context modeling framework based on a graph neural network is constructed to dynamically associate the user's emotional state with the dialogue history and the scene background. A memory enhancement network is used to store multi-round dialogue information for long-term emotional state tracking. A hotel scene knowledge graph is obtained and used for scene knowledge enhancement to generate an emotional response strategy. The emotional intensity is classified by using an emotional intensity quantification model. The emotional state of the emotional intensity quantification model is divided into a first preset number of major categories and a second preset number of subcategories.

5. An emotional perception and dynamic response device for a digital human, characterized by, The device includes a processor and a memory. The memory is used to store one or more program instructions. The processor is used to run one or more program instructions to execute the steps of the digital human emotional perception and dynamic response method according to any one of claims 1 to 3.

6. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, which is executed by the processor to implement the steps of the digital human emotional perception and dynamic response method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Intelligent customer service question-answering system and method

    CN118861223A

  • Intelligent emotional interaction method for service robot

    CN119150099A

  • Method and system for generating digital human video

    CN119180895A