Service robot control method capable of self-adapting to human emotion

Through the adjustment of multimodal emotion recognition model and dynamic behavior strategy, the service robot's adaptability to user emotions in different scenarios is solved, and a more natural human-computer interaction experience is achieved. The robot can adjust the strategy in real time according to user emotions and continuously optimize it through online learning.

CN120422236APending Publication Date: 2025-08-05SENDAO ENERGY (HANGZHOU) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510691442.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

The existing service robots lack adaptability to user emotions in medical, hotel and education scenarios, resulting in a stiff human-computer interaction experience and a lack of natural comfort.

Method used

The multi-modal emotion recognition model is adopted to integrate expression images, voice signals and language content, combined with dynamic behavior strategy mapping mechanism, adjust the robot's voice tone, action speed and service strategy, and optimize the model through the feedback module to achieve real-time adjustment of emotion recognition and behavior.

Benefits of technology

It improves the naturalness of human-computer interaction, enables robots to "understand" and "respond to" human emotions, reduces misjudgment and blunt reactions, and continuously evolves through feedback-driven online learning to adapt to user emotional changes in different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120422236A_ABST
    Figure CN120422236A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of man-machine interaction, in particular to a service robot control method capable of self-adapting to human emotions, which comprises the following steps: acquiring an emotional state of a user in real time through a multi-mode emotion recognition model fusing an expression image, a voice signal and language content; based on the emotional state of the user, a dynamic behavior strategy mapping mechanism is adopted, and the voice intonation, the action speed and the service strategy of the robot are adjusted; and monitoring the response of the user to the adjusted service through a feedback module, and further optimizing an emotion recognition model and a behavior strategy mapping mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of human-computer interaction technology, and in particular to a control method for a service robot that can adapt to human emotions. Background Art

[0002] Existing service robots typically focus solely on command execution, navigation, and basic conversation, paying insufficient attention to user emotions. This results in a stilted, unnatural, and uncomfortably interactive experience. In scenarios like healthcare, hospitality, and education, users experience diverse emotional states and communication styles, creating an urgent need for intelligent control solutions that can adapt to these emotional fluctuations. Summary of the Invention

[0003] In order to overcome the above problems at least to a certain extent, the present application provides a service robot control method that can adapt to human emotions.

[0004] The scheme of this application is as follows:

[0005] A control method for a service robot capable of adapting to human emotions, comprising:

[0006] By integrating facial expressions, voice signals and language content into a multimodal emotion recognition model, the user's emotional state can be acquired in real time.

[0007] Based on the user's emotional state, a dynamic behavior strategy mapping mechanism is used to adjust the robot's voice tone, movement speed and service strategy;

[0008] The feedback module monitors users’ responses to the adjusted services and further optimizes the emotion recognition model and behavior strategy mapping mechanism.

[0009] Preferably, the multimodal emotion recognition model includes:

[0010] Expression recognition submodule, used to extract key features of the user's face;

[0011] Speech emotion analysis submodule, used to extract the acoustic features of user speech;

[0012] The language content understanding submodule is used to analyze the emotional tendency of the user's language text;

[0013] The multimodal emotion recognition model is used to perform weighted fusion on the emotion probabilities output by each submodule to generate a final user emotion score.

[0014] Preferably, the expression recognition submodule adopts a convolutional neural network combined with a key point detection algorithm to improve the recognition accuracy of facial action units.

[0015] Preferably, the speech emotion analysis submodule is based on the integration of time-frequency features and deep learning models, integrating Mel spectrum features and long short-term memory networks.

[0016] Preferably, the language content understanding submodule adopts a pre-trained sentiment analysis model and combines it with transfer learning technology to adapt to the language styles of different fields.

[0017] Preferably, the dynamic behavior strategy mapping mechanism includes:

[0018] Mapping user emotional states to a predefined library of robot behavior parameters;

[0019] Generate robot action sequences and voice parameters based on the mapping results;

[0020] The mapping rules are updated online according to the actual interaction effect.

[0021] Preferably, the service strategy includes but is not limited to customized processes for three different scenarios: medical care, hotel reception, and educational guidance.

[0022] Preferably, it further includes: synchronizing the emotion recognition model and behavior strategy mapping mechanism on a cloud server to achieve multi-robot collaborative learning and strategy sharing.

[0023] The technical solution provided by this application may have the following beneficial effects:

[0024] This technical solution significantly enhances the naturalness of interactions. The robot can "understand" and "respond" to human emotions, making the service experience more like that of a human assistant. The robot can quickly adjust its strategy to address emotional differences within the same user or across different scenarios, reducing misjudgments and blunt responses. Feedback-driven online learning enables the system to continuously evolve in real-world deployments, making it more intelligent with use.

[0025] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0027] Figure 1 This is a flow chart of a method for controlling a service robot that can adapt to human emotions, provided by an embodiment of the present application. DETAILED DESCRIPTION

[0028] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.

[0029] Figure 1 This is a flow chart of a method for controlling a service robot that can adapt to human emotions provided by an embodiment of the present application, with reference to Figure 1 , a service robot control method capable of adapting to human emotions, comprising:

[0030] By integrating facial expressions, voice signals and language content into a multimodal emotion recognition model, the user's emotional state can be acquired in real time.

[0031] Based on the user's emotional state, a dynamic behavior strategy mapping mechanism is used to adjust the robot's voice tone, movement speed, and service strategy;

[0032] The feedback module monitors users’ responses to the adjusted services and further optimizes the emotion recognition model and behavior strategy mapping mechanism.

[0033] This technical solution significantly enhances the naturalness of interactions. The robot can "understand" and "respond" to human emotions, making the service experience more like that of a human assistant. The robot can quickly adjust its strategy to address emotional differences within the same user or across different scenarios, reducing misjudgments and blunt responses. Feedback-driven online learning enables the system to continuously evolve in real-world deployments, making it more intelligent with use.

[0034] It should be noted that the multimodal emotion recognition model includes:

[0035] Expression recognition submodule, used to extract key features of the user's face;

[0036] Speech emotion analysis submodule, used to extract the acoustic features of user speech;

[0037] The language content understanding submodule is used to analyze the emotional tendency of the user's language text;

[0038] The multimodal emotion recognition model is used to weightedly fuse the emotion probabilities output by each sub-module to generate the final user emotion score.

[0039] Expression recognition submodule: Capture facial micro-expressions by detecting key points of the face (such as the corners of the eyes and mouth).

[0040] Speech emotion analysis submodule: extracts acoustic features such as Mel spectrum, pitch, and energy, and then uses a deep model to estimate emotion probability.

[0041] Language content understanding submodule: performs sentiment analysis (positive, neutral, negative) on the user's text or speech-to-text conversion results.

[0042] Probability-weighted fusion: Assign weights to different modalities and fuse their confidence levels in sentiment judgment to obtain a more accurate comprehensive sentiment score.

[0043] In this technical solution, multiple sub-modules are used to complement multi-dimensional information, which makes up for the weaknesses of a single modality (such as only looking at expressions or only listening to voice) and significantly improves recognition accuracy.

[0044] Flexible weight allocation: The weights of each modality can be dynamically adjusted according to the environment (lighting, noise, etc.) to enhance robustness.

[0045] It should be noted that the expression recognition submodule uses a convolutional neural network combined with a key point detection algorithm to improve the recognition accuracy of facial action units.

[0046] The speech sentiment analysis submodule is based on the integration of time-frequency features and deep learning models, integrating Mel spectrum features and long short-term memory networks.

[0047] The language content understanding submodule uses a pre-trained sentiment analysis model and combines it with transfer learning technology to adapt to the language style of different fields.

[0048] Combining classic facial keypoint detection with CNN feature extraction enables the model to capture both geometric structure (keypoint locations) and understand micro-expression textures (CNN output). Even in complex backgrounds or partial occlusions, it can still accurately extract facial action units (APUs), reducing misidentification.

[0049] Mel-spectrograms are used to capture the time-frequency distribution of sound, and time series networks such as LSTM are then used to explore the temporal evolution of emotions. This not only identifies instantaneous acoustic features but also understands how emotions change over the course of speech, improving the consistency and accuracy of speech sentiment analysis.

[0050] We leverage pre-trained models on large-scale corpora (such as Chinese BERT and RoBERTa) for initial sentiment classification, and then fine-tune them using small sample data from specific scenarios (medical, hospitality, and education). This eliminates the need to retrain from scratch and allows for rapid migration to new domains, ensuring the high performance of the sentiment understanding model across a wide range of language styles.

[0051] It should be noted that the dynamic behavior policy mapping mechanism includes:

[0052] Mapping user emotional states to a predefined library of robot behavior parameters;

[0053] Generate robot action sequences and voice parameters based on the mapping results;

[0054] The mapping rules are updated online based on the actual interaction effect.

[0055] Behavioral parameter library: pre-defines parameter ranges such as speaking speed, pitch, movement speed, gesture amplitude, etc. corresponding to different emotions.

[0056] Action and speech generation: The robot synthesizes action sequences and speech synthesis parameters that meet emotional requirements in real time based on mapping parameters.

[0057] Online rule update: Dynamically fine-tune the mapping relationship based on users' immediate feedback (such as satisfaction scores and subsequent emotional changes).

[0058] This technical solution eliminates the need to search for new policy libraries and rapidly generates appropriate feedback within the existing parameter space. It continuously optimizes mapping rules based on this feedback, making the robot's interaction increasingly tailored to specific users or scenarios.

[0059] It should be noted that the service strategy includes but is not limited to customized processes for three different scenarios: medical care, hotel reception, and educational counseling.

[0060] This technical solution provides service process templates tailored to different scenarios. For example, medical care emphasizes slow speech and soothing language; hotel reception emphasizes polite greetings and efficiency; and educational counseling emphasizes interactivity and guided questioning. This allows for a single device to be used across diverse application areas, eliminating the need to rewrite control logic for each industry and reducing development and deployment costs. Emotion recognition and feedback maintain consistency and professionalism across diverse scenarios.

[0061] Furthermore, the method also includes: synchronizing the emotion recognition model and the behavior strategy mapping mechanism on the cloud server to achieve multi-robot collaborative learning and strategy sharing.

[0062] This technical solution synchronizes model updates or mapping improvements on local robots to the cloud, which then distributes them to other robots, enabling experience sharing. Each robot can benefit from the actual operational experience of other robots, shortening the learning cycle for each node. Centralized cloud-based management of metrics and versions ensures that the entire robot swarm maintains the latest and optimal strategies.

[0063] It can be understood that the same or similar parts of the above embodiments can be referenced to each other, and the contents not described in detail in some embodiments can refer to the same or similar contents in other embodiments.

[0064] It should be noted that, in the description of this application, the terms "first", "second", etc. are used for descriptive purposes only and should not be understood as indicating or implying relative importance. In addition, in the description of this application, unless otherwise specified, the meaning of "plurality" refers to at least two.

[0065] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application belong.

[0066] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0067] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.

[0068] In addition, the functional units in the various embodiments of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into a module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0069] The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc.

[0070] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present application. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0071] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.

Claims

1. A service robot control method capable of adapting to human emotions, characterized in that: include: By integrating facial expressions, voice signals and language content into a multimodal emotion recognition model, the user's emotional state can be acquired in real time. Based on the user's emotional state, a dynamic behavior strategy mapping mechanism is used to adjust the robot's voice tone, movement speed and service strategy; The feedback module monitors users’ responses to the adjusted services and further optimizes the emotion recognition model and behavior strategy mapping mechanism.

2. The method according to claim 1, characterized in that The multimodal emotion recognition model includes: Expression recognition submodule, used to extract key features of the user's face; Speech emotion analysis submodule, used to extract the acoustic features of user speech; The language content understanding submodule is used to analyze the emotional tendency of the user's language text; The multimodal emotion recognition model is used to perform weighted fusion on the emotion probabilities output by each submodule to generate a final user emotion score.

3. The method according to claim 2, characterized in that The expression recognition submodule uses a convolutional neural network combined with a key point detection algorithm to improve the recognition accuracy of facial action units.

4. The method according to claim 2, characterized in that The speech emotion analysis submodule is based on the integration of time-frequency features and deep learning models, fusing Mel spectrum features and long short-term memory networks.

5. The method according to claim 2, characterized in that The language content understanding submodule adopts a pre-trained sentiment analysis model and combines it with transfer learning technology to adapt to the language styles of different fields.

6. The method according to claim 1, characterized in that The dynamic behavior strategy mapping mechanism includes: Mapping user emotional states to a predefined library of robot behavior parameters; Generate robot action sequences and voice parameters based on the mapping results; The mapping rules are updated online according to the actual interaction effect.

7. The method according to claim 1, characterized in that The service strategy includes but is not limited to customized processes for three different scenarios: medical care, hotel reception, and educational guidance.

8. The method according to any one of claims 1 to 7, characterized in that Further including: The emotion recognition model and behavior strategy mapping mechanism are synchronized on a cloud server to achieve multi-robot collaborative learning and strategy sharing.

Citation Information

Cited By

  • Optimization method of robot control model and related equipment

    CN120715901A