Emotion expression method and device, equipment and computer readable storage medium
By obtaining user conversation information to generate emoticons, determining emotional labels and intensities, and controlling the robot's output of voice interaction, expressions, and actions, the problem of the robot's limited emotional expression is solved, and more accurate and diverse emotional expressions are achieved.
Patent Information
- Application Number
- CN202510634049.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-09-16
AI Technical Summary
Existing robots' emotional expressions are limited to preset emotional phrases, which are not rich enough and cannot accurately reflect complex emotions.
By obtaining user session information, generating session information containing emoticons, and determining the emotional label and intensity based on the emoticons, the robot is controlled to output corresponding voice interactions, expressions, and actions.
The robot's emotional expression is accurate and rich, and it can respond in a variety of ways according to different emotions and emotional intensities.
Smart Images

Figure CN120645240A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of robotics, and in particular to an emotion expression method and apparatus, device, and computer-readable storage medium. Background Art
[0002] In existing technologies, robots' emotional expression is limited by fixed emotional phrases pre-set on the robot itself, resulting in limited emotional expression and insufficient expression. Therefore, how to accurately realize the emotional expression of robots has become an urgent problem to be improved. Summary of the Invention
[0003] To solve the above technical problems, the present application provides an emotion expression method and apparatus, device, and computer-readable storage medium.
[0004] The present application provides a method for expressing emotions, the method comprising:
[0005] Obtaining first conversation information input by a first user to the robot;
[0006] generating second conversation information output by the robot to the first user according to the first conversation information; the second conversation information includes the first emoticon;
[0007] Determining a first emotion label and / or a first emotion intensity corresponding to the first emoticon;
[0008] Control the robot to output a first emotion-related behavior corresponding to the first emotion label and / or the first emotion intensity, where the first emotion-related behavior includes at least one of outputting corresponding voice interaction information, outputting a corresponding expression, and performing a corresponding action, and the voice interaction information is included in the second conversation information.
[0009] The emotion expression device provided by the present application comprises:
[0010] an acquiring unit, configured to acquire first conversation information input by a first user to the robot;
[0011] A processing unit is configured to generate second conversation information output by the robot to the first user based on the first conversation information; the second conversation information includes a first emoticon; determine a first emotion label and / or a first emotion intensity corresponding to the first emoticon; and control the robot to output a first emotion-related behavior corresponding to the first emotion label and / or the first emotion intensity, where the first emotion-related behavior includes at least one of outputting corresponding voice interaction information, outputting a corresponding expression, and performing a corresponding action, and the voice interaction information is included in the second conversation information.
[0012] The emotion expression device provided in the present application includes: a processor and a memory, the memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the above-mentioned emotion expression method.
[0013] The computer-readable storage medium provided in the present application is used to store a computer program, and the computer program enables a computer to execute the above method.
[0014] In the technical solution of this application, first conversation information input by a first user to a robot is obtained; second conversation information output by the robot to the first user is generated based on the first conversation information; the second conversation information includes a first emoticon; a first emotion label and / or a first emotion intensity corresponding to the first emoticon is determined; and the robot is controlled to output a first emotion-related behavior corresponding to the first emotion label and / or the first emotion intensity, wherein the first emotion-related behavior includes at least one of outputting corresponding voice interaction information, outputting a corresponding expression, and performing a corresponding action, and the voice interaction information is included in the second conversation information. In this way, the first emotion label and / or the first emotion intensity corresponding to the second conversation information can be accurately determined based on the first emoticon, and the corresponding first emotion-related behavior can be accurately obtained, thereby improving the accuracy of the robot's emotional expression. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The drawings described herein are used to provide further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute improper limitations on the present application.
[0016] Figure 1 This is a flow chart of the emotion expression method provided in the embodiment of the present application. Figure 1 ;
[0017] Figure 2 This is a schematic diagram of a sad emoticon provided in an embodiment of the present application;
[0018] Figure 3 This is a flow chart of the emotion expression method provided in the embodiment of the present application. Figure 2 ;
[0019] Figure 4 Schematic diagrams of other emoticons provided in the embodiments of this application;
[0020] Figure 5 This is a schematic diagram of the expression resources provided in the embodiment of the present application;
[0021] Figure 6 This is a schematic diagram of the process of interacting with a robot provided in an embodiment of the present application;
[0022] Figure 7is a schematic diagram of emoticons during robot interaction provided in an embodiment of the present application;
[0023] Figure 8 Schematic diagram of the structure of the emotion expression device provided in an embodiment of the present application;
[0024] Figure 9 This is a schematic structural diagram of an emotion expression device provided in an embodiment of the present application;
[0025] Figure 10 It is a schematic structural diagram of the chip of an embodiment of the present application. DETAILED DESCRIPTION
[0026] The following will describe the technical solutions in the embodiments of this application in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0027] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments, but it is understood that "some embodiments" may be the same subset or different subsets of all possible embodiments, and may be combined with each other without conflict. It should also be noted that the terms "first\second\third" involved in the embodiments of the present application are only used to distinguish similar objects and do not represent a specific ordering of objects. It is understandable that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here. The term "and / or" herein is merely a description of an association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " herein generally indicates that the objects associated before and after are in an "or" relationship. It should also be understood that the "indication" mentioned in the embodiments of the present application may be a direct indication, an indirect indication, or an indication of an association relationship. For example, "A indicates B" can mean that A directly indicates B, for example, B can be obtained through A; it can also mean that A indirectly indicates B, for example, A indicates C, and B can be obtained through C; it can also mean that there is an association relationship between A and B. It should also be understood that the "correspondence" mentioned in the embodiments of this application can mean that there is a direct or indirect correspondence relationship between the two, it can also mean that there is an association relationship between the two, and it can also mean a relationship between indicating and being indicated, configuring and being configured, etc.
[0028] To facilitate understanding of the technical solutions of the embodiments of the present application, the relevant technologies of the embodiments of the present application are described below. The following relevant technologies can be arbitrarily combined with the technical solutions of the embodiments of the present application as optional solutions, and they all fall within the protection scope of the embodiments of the present application.
[0029] Existing emotion recognition technology constructs emotion types and corresponding emotion phrases, matches and extracts emotion phrases from the input information, and then determines the corresponding emotion type through phrase mapping. This method is limited by the number of emotion phrases and requires that the input information contain emotion phrases for recognition. However, humans or robots do not express emotions according to fixed emotion phrases during conversations, and even sentences with exactly the same phrases may express completely different emotions in different contexts, so this method has limitations. Existing emotion intensity calculation technology maps emotion intensity by matching emotion words in the input information, or calculates the intensity value of the corresponding emotion based on the frequency of occurrence of emotion phrases in the input information. Similarly, due to the limitation of emotion phrases, this method also has limitations. Existing robot emotion expression is to calculate the corresponding emotion type and then execute the corresponding emotion type's expression image or action, which is not rich in expression. Therefore, how to achieve the most accurate expression of robot emotions as possible becomes a problem that needs to be considered. To this end, the following technical solutions of the embodiments of the present application are proposed.
[0030] To facilitate understanding of the technical solutions of the embodiments of the present application, the technical solutions of the present application are described in detail below through specific embodiments. The above related technologies can be combined arbitrarily with the technical solutions of the embodiments of the present application as optional solutions, and all of them fall within the scope of protection of the embodiments of the present application. The embodiments of the present application include at least part of the following contents.
[0031] Figure 1 This is a flow chart of the emotion expression method provided in the embodiment of the present application. Figure 1 ,like Figure 1 As shown, the emotion expression method includes the following steps:
[0032] Step 101: Acquire first conversation information input by a first user to a robot.
[0033] In some embodiments, the first conversation information input by the first user to the robot can be text information or voice information. If the input is voice information, it will be converted into text information to facilitate subsequent processing. The specific format of the first conversation information is not specifically limited in this application and can be determined based on actual circumstances.
[0034] In some embodiments, the first session information received by the robot includes the first session information output by the user and the first session information input by the system. The sender of the first session information received by the robot is not specifically limited in this application and can be determined according to actual circumstances.
[0035] Step 102: Generate second conversation information output by the robot to the first user based on the first conversation information; the second conversation information includes the first emoticon.
[0036] In some embodiments, in a text environment, emoticons can easily and accurately convey emotional information compared to text. Therefore, after obtaining the first session information, a second session information corresponding to the first session information is generated based on the first session information. That is, the second session information is a reply information of the first session information, wherein the second session information includes the first emoticon, or the second session information includes the first emoticon and the corresponding text information, and the first emoticon represents the emotion of the text information. Therefore, the emotional information of the second session information can be accurately conveyed based on the first emoticon.
[0037] In some embodiments, the first emoticon can be a picture representing an expression. For example, it can be an emoticon picture or an emoticon package, and also includes self-constructed emoticons that can express emotions. There is no specific limitation on the first emoticon here, and there is no specific limitation on the source of the emoticon package.
[0038] In some embodiments, generating second session information output by a robot to a first user based on first session information includes: identifying keywords in the first session information, retrieving related session information from historical session information between the first user and the robot based on the keywords, generating second session information corresponding to the first session information by performing semantic understanding and language processing on the related session information, and outputting the second session information to the first user.
[0039] In some embodiments, in a conversation environment, the same sentence may express completely different emotions in different contexts. Therefore, when generating the second conversation information, the historical conversation information between the first user and the robot is used as a reference to help understand the user's emotions. That is, the second conversation information output by the robot to the first user is generated based on the first conversation information and the historical conversation information between the first user and the robot.
[0040] In some embodiments, generating second session information output by the robot to the first user based on the first session information and historical session information between the first user and the robot includes: calling a network model to generate the second session information output by the robot to the first user based on the first session information and historical session information between the first user and the robot.
[0041] In some embodiments, a network model can be used to generate second session information based on the first session information and historical session information. The first session information and historical session information are used as input to the network model, and the output of the network model is second session information output by the robot to the first user, where the second session information includes a first emoticon. Specifically, the network model combines historical conversation content and emoticon information to perform sentiment analysis on the user input, and generates a first emoticon that matches the sentiment of the current conversation. Optionally, text information that matches the sentiment of the current conversation is also generated (this text information can serve as the reply content output by the robot to the first user).
[0042] Specifically, keywords in the first conversation information are identified, and related conversation information in the historical conversation information between the first user and the robot is retrieved based on the keywords. Second conversation information corresponding to the first conversation information is generated by performing semantic understanding and language processing on the related conversation information, and the second conversation information is output to the first user.
[0043] In some embodiments, when the network model does not return the second conversation information within a certain period of time, the robot may display a waiting expression to prompt the first user, or the robot may display text to inform the first user to resend the first conversation information and play audio.
[0044] In some embodiments, when the network model is processing text, the robot can display corresponding expressions based on the network model processing process. For example, when the robot is waiting for the network model processing process, it can display the corresponding expression of waiting.
[0045] In some embodiments, the network model is a generative large model. Exemplarily, the network model may be a large language model (LLM). The specific structure of the LLM is not described in detail here.
[0046] In some implementations, the network model needs to be trained so that the network model can generate the second session information including the first emoticon based on the first session information.
[0047] In some implementations, in order to have the network model output the first emoticon, the Prompt project can be used to add output format requirements to the prompt word of the network model, that is, to use emoticons to express emotions in the reply. The position of the emoticon can be at the beginning, end, or center of the reply text. This application does not specifically limit this. For example, the system prompt word format can be shown in Table 1:
[0048]
[0049]
[0050] Table 1: System prompt word styles
[0051] It is understandable that the key points and output formats in the system prompt word style examples include not only the example content, but also the remaining content. The specific precautions and output formats can be determined according to actual conditions, and this application does not make specific limitations on this.
[0052] In some implementations, a supervised fine-tuning (SFT) process can also be used. When the prompt process cannot fully control the network model's emoticon output, the network model can be fine-tuned by simulating training data based on robot conversation scenarios. At least thousands of training data are required. The training data format is shown in Table 2:
[0053]
[0054] Table 2: Training data types
[0055] In some implementations, the network model can be deployed on the robot's computing platform or on a server, and this application does not impose any specific limitations on this.
[0056] In some implementations, the robot may have different personalities, and the network model may generate second conversation information adapted to the robot based on the robot's personality.
[0057] Step 103: Determine a first emotion label and / or a first emotion intensity corresponding to the first emoticon.
[0058] In some implementations, after the second conversation information is acquired, the first emotion tag and / or the first emotion intensity corresponding to the first emoticon is determined according to the first emoticon in the second conversation information.
[0059] For example, the first emotion label (hereinafter referred to as emotion) may be "sad", and the scoring rule for the first emotion intensity (hereinafter referred to as emotion_level) is 1 to 10 points, where 1 to 2 points are mild intensity emotions, 3 to 4 points are moderate intensity emotions, 5 to 7 are severe intensity emotions, and 8 to 10 are extreme intensity emotions. Figure 2 is a schematic diagram of a sad emoticon provided in an embodiment of the present application, and the sad emoticon corresponds to the first emoticon mentioned above, such as Figure 2As shown, there are four sad emoticons. The first sad emoticon specifically represents mild unhappiness, and its corresponding emotion and emotion_level" are {"emotion"="sad","emotion_level"="1"}; the second sad emoticon specifically represents moderate unhappiness, and its corresponding emotion and emotion_level" are {"emotion"="sad","emotion_level"="3"}; the third sad emoticon specifically represents severe unhappiness, and its corresponding emotion and emotion_level" are {"emotion"="sad","emotion_level"="6"}; the fourth sad emoticon specifically represents extreme unhappiness, and its corresponding emotion and emotion_level" are {"emotion"="sad","emotion_level"="9"}.
[0060] It is understandable that emotion labels can also be called emotion type labels.
[0061] In some embodiments, determining the first emotion label and / or first emotion intensity corresponding to the first emoticon includes: determining the first emotion label and / or first emotion intensity corresponding to the first emoticon based on mapping relationship information; wherein the mapping relationship information includes a mapping relationship between at least one group of emoticons and emotion labels and / or emotion intensities.
[0062] In some embodiments, an emotion system is constructed for the robot, which includes mapping relationship information, and the mapping relationship information includes a mapping relationship between at least one set of emoticons and emotion labels and / or emotion intensities. When a first emoticon is obtained, the emotion intensity (i.e., the first emotion intensity) and / or emotion label (i.e., the first emotion label) corresponding to the first emoticon is obtained according to the mapping relationship information.
[0063] For example, the above mapping relationship information can be represented by a table, and an emotional system is constructed for the robot, which includes the robot's basic emotional labels and corresponding emotional intensities, and commonly used emoticons are marked with corresponding emotional labels and emotional intensity labels, and a mapping relationship table between emoticons, emotional labels, and emotional intensities is established.
[0064] Exemplarily, the above-mentioned mapping relationship information can be represented by a table. Referring to Table 3 below, Table 3 below lists the mapping relationships between some basic emoticons and emotional labels and emotional intensities. It should be noted that the mapping relationship information of the embodiment of the present application is not limited to the several mapping relationships listed in Table 3 below, and can include more or fewer mapping relationships based on the several mapping relationships listed in Table 3.
[0065] Emoji Emotion tag (emotion) emotion_level Emoji 1 sad 1 Emoji 2 sad 3 Emoji 3 Happy 1 Emoji 4 Happy 3 Emoji 5 Happy 6
[0066] Table 3: Mapping relationship between emoticons, emotion labels and emotion intensity
[0067] In some implementations, the network model is updated according to the second session information to improve the accuracy of the network model.
[0068] Exemplarily, after the second session information is generated based on the first session information, the second session information can be used as training data for the network model. After obtaining more than a certain amount of second session information, that is, obtaining a certain amount of training data, the network model can be updated using the training data. Exemplarily, the network model can be fine-tuned, or the prompt word project of the network model can be updated. The second session information can also be used as historical session information for the next session and input into the network model to generate reply information for the next session.
[0069] Step 104: Control the robot to output a first emotion-related behavior corresponding to the first emotion label and / or the first emotion intensity, where the first emotion-related behavior includes at least one of outputting corresponding voice interaction information, outputting a corresponding expression, and performing a corresponding action, and the voice interaction information is included in the second conversation information.
[0070] In some embodiments, the robot selects a corresponding first emotion-related behavior based on the first emotion label and / or first emotion intensity corresponding to the first emoticon, and controls the robot to perform the first emotion-related behavior. The first emotion-related information includes the robot broadcasting voice interaction information and performing corresponding expressions and actions. For example, if the first emotion label corresponding to the first emoticon is sad and the first emotion intensity is mildly unhappy, the robot is controlled to broadcast voice interaction information according to the corresponding voice broadcast method and perform the corresponding expressions and actions.
[0071] It is understandable that when controlling the robot to perform the first emotion-related behavior includes voice interaction information, the voice interaction information is included in the second conversation information. That is, the generated second conversation information includes the voice interaction information.
[0072] In some embodiments, if the robot is in silent mode, the operation of broadcasting the interaction information is not performed, but the interaction information is displayed in text form. The interaction information is included in the second conversation information.
[0073] It can be understood that when the second conversation information only includes the first emoticon, the robot performs the corresponding expression and action. When the second conversation information includes the first text information and the corresponding first emoticon, the robot can broadcast the interactive information or display the interactive information according to the corresponding voice broadcast method, and perform the corresponding expression and action, where the interactive information is the first text information.
[0074] In some embodiments, when a robot has different personalities, a corresponding first emotion-related behavior can be selected based on the robot's personality. For example, if the robot's personality is cheerful, an expression resource that leans toward cheerfulness is selected from the expression resource library, and an action resource that leans toward cheerfulness is selected from the action resource library. This application does not specifically limit the specific robot personality.
[0075] In some embodiments, controlling the robot to output voice interaction information corresponding to the first emotion label and / or the first emotion intensity includes: controlling the robot to broadcast the voice interaction information in the voice intonation and / or tone corresponding to the first emotion label and / or the first emotion intensity.
[0076] In some embodiments, the robot utilizes text-to-speech (TTS) technology that can control emotion and emotion intensity, enabling it to synthesize speech with varying emotional tones and / or intonations using the same timbre at varying emotional intensities. Because TTS primarily categorizes emotion labels based on tone and style, which differ from the emotion labels input by the robot, a mapping table has been established between robot emotion labels and TTS emotion labels. Typically, the emotion received by the robot corresponds to the TTS emotion to be invoked. However, when the TTS tone and style vary significantly for the same emotion at varying intensities, matching the TTS emotion is necessary.
[0077] For example, when the emoticon obtained represents mild sadness and moderate sadness, the corresponding TTS emotion label is consistent with the emotion label represented by the emoticon, also expressed as "sad", but when the emoticon obtained represents severe sadness and extreme sadness, the TTS emotion label needs to be mapped to crying, that is, {"emotion"="sad","emotion_level"="6"} is mapped to the corresponding TTS {"TTS_emotion"="cry","emotion_level"="6"}. TTS using the crying emotion style can more realistically express severe and extremely sad emotions.
[0078] In some embodiments, controlling the robot to output an expression corresponding to a first emotion label and / or a first emotion intensity includes: controlling the robot to select a target expression from expression resources corresponding to the first emotion label and / or the first emotion intensity, and calling the selected expression resources to display the target expression, where the target expression includes a static frame or multiple dynamic frames.
[0079] In some embodiments, controlling the robot to perform an action corresponding to a first emotion label and / or a first emotion intensity includes: controlling the robot to select a target action from action resources corresponding to the first emotion label and / or the first emotion intensity, and calling the selected action resources to display the target action.
[0080] In some embodiments, the robot can design matching actions and expression resources based on emotions of varying intensities. To enrich the robot's emotional expression, the same emotion intensity can have multiple expressions and actions, which the robot can randomly select and execute. Therefore, after determining the first emotion label and / or first emotion intensity corresponding to the first emoji, a corresponding expression resource, also known as the target expression, is selected from the expression resources. The selected expression resource is then used to display the target expression. The target expression can include a single static frame or multiple dynamic frames. In other words, the robot can statically display the expression corresponding to a single static frame, or dynamically display the expression corresponding to multiple dynamic frames.
[0081] In some implementations, the robot selects a target action from the action resources corresponding to the first emotion label and / or the first emotion intensity, and calls the selected action resource to display the target action.
[0082] For example, when the robot receives a moderately sad emotion {"emotion"="sad","emotion_level"="3"}, the expression resources include a static frame or multiple dynamic frames of half-drooping eyelids and a static frame or multiple dynamic frames of drooping eyelids and eyes looking to the side, and the action resources include lowering the head and lowering the head and shaking the head. Therefore, according to the first emotion label and / or the first emotion intensity, one of the aforementioned expression resources and action resources is randomly selected for execution.
[0083] It should be noted that even in the absence of emotions, that is, in the "normal" state, there will be a series of resources such as "winking" expressions and "looking at the user" actions that can be randomly selected and executed by the robot.
[0084] In some embodiments, the method further includes: in response to a first interactive operation, controlling the robot to perform a second emotion-related behavior corresponding to the first interactive operation, where the first interactive operation includes one or more of touch, stroking, or key operations perceived by the robot body.
[0085] In some embodiments, the first interaction operation can be a long press or a short press of an AI dialogue key by the first user, where the AI dialogue key is provided on the robot. When the first user presses the AI dialogue key as desired, the robot performs a second emotion-related behavior corresponding to the first interaction operation.
[0086] In some embodiments, when the first user long presses the AI dialogue button for the first time, the robot begins to receive the first user's voice. At this time, the second emotion-related behavior may include vibration, displaying an expression of listening to the voice, and recording. The specific duration of the vibration can be determined according to actual conditions. After the robot receives the first user's voice, it will first verify whether the pass membership has expired. If it has not expired, the robot will play a successful audio message to remind the user that the voice message has been successfully received. If the pass membership has expired, the member needs to be reminded to renew the pass. Relevant text or emoticons and information such as the QR code required for pass renewal are displayed on the interface to remind members to renew the pass.
[0087] In some embodiments, when the first interactive operation is voice input by the first user, after the robot receives the voice input from the first user, the robot sends the voice to the network model, which processes the voice and generates a reply message based on the voice, wherein the reply message includes emoticons. The reply message is sent to the robot, and the robot determines the emotional label and emotional intensity expressed based on the emoticons. However, while the robot is waiting for the network model to process, the second emotion-related behavior can be vibration, displaying a waiting emoji, and looping the waiting audio. At this time, the network model processes the message. If the network model does not return a reply message within the specified time, the robot can vibrate, display relevant text such as "re-enter", display a familiar emoji, and play the missed audio.
[0088] In some embodiments, after the network model returns the reply content, it first determines whether the robot is in silent mode. If it is in silent mode, the robot displays the reply text on the display interface without playing any sound. At this time, the robot can still perform corresponding expressions and actions, waiting for the user to perform the operation again. The robot will perform different second emotion-related behaviors according to the operation performed by the user again.
[0089] For example, when the user presses the intercom button, the robot exits the current state and enters the intercom interface; when the user briefly presses the AI button, the robot returns to the wide-eyed state and plays the corresponding performance of the short AI button press; a long press of the AI button returns to the wide-eyed state and plays the listening performance; pressing the volume up button returns to the wide-eyed state and exits mute, and the next reply will no longer display text; if the user shakes the robot, the robot does not respond. It is understood that different interactive operations have corresponding second emotion-related behaviors, and this application does not specifically limit this.
[0090] In some embodiments, if the robot is not silent, after receiving the second conversation information, the robot is controlled to perform voice broadcasting of interaction information, action display, and expression display.
[0091] The technical solution of the embodiment of the present application is to obtain first conversation information output by a first user to a robot; generate second conversation information output by the robot to the first user based on the first conversation information; the second conversation information includes a first emoticon; determine a first emotion label and / or a first emotion intensity corresponding to the first emoticon; and control the robot to output a first emotion-related behavior corresponding to the first emotion label and / or first emotion intensity, wherein the first emotion-related behavior includes at least one of outputting corresponding voice interaction information, outputting a corresponding expression, and performing a corresponding action, and the voice interaction information is included in the second conversation information. In this way, the first emotion label and / or the first emotion intensity corresponding to the second conversation information can be accurately determined by the first emoticon, and the corresponding first emotion-related behavior can be accurately obtained, thereby improving the accuracy of the robot's emotional expression.
[0092] The technical solutions of the embodiments of the present application are illustrated below with reference to specific application examples.
[0093] As mentioned above, the emotion recognition technology in the existing technology constructs emotion types and corresponding emotion phrases. Due to the differences in the number of emotion phrases and the dialogue scenes and dialogue contexts, this method has limitations. Therefore, the emotion intensity calculation technology and robot emotion expression in the existing technology are also limited. Based on this, the embodiment of the present application proposes an emotion expression method, which uses the generative large model's understanding of emotions and the accuracy of emoticon expression of emotions to achieve more accurate emotion recognition, classification and emotion intensity calculation of the robot. Through different emotion labels and emotion intensities, the robot can select corresponding expressions, actions and corresponding intonations and voice to conduct dialogues, thereby achieving richer emotional expression. Figure 3 This is a flow chart of the emotion expression method provided in the embodiment of the present application. Figure 2 ,like Figure 3 As shown, first, user input / system input is received, and then content replies and emoticons are generated based on the sentiment analysis of the large model. First, the large model generates user emoticons based on historical conversation content, emoticons, and input. The large model generates emoticons for the robot's reply content set, and matches emoticons to corresponding emotions and emotional intensities. Then, the robot expresses according to different emotions and emotional intensities. The robot receives the emotion label, emotional intensity value, and reply content. The robot uses the voice tone and voice corresponding to the emotion and intensity to broadcast the reply content. The robot performs the expression corresponding to the emotion and intensity, and the robot performs the action corresponding to the emotion and intensity. The specific description is as follows:
[0094] (1) Content reply and emoticon generation based on large-scale model sentiment analysis
[0095] The robot receives text information input by the user or the system from the end, first performs sentiment analysis on the user input by combining the historical conversation content and emoticon information through the big model, and then generates reply content and emoticons that match the current conversation mood through the big model.
[0096] In textual contexts, emoticons convey emotions more easily and accurately than text. Similarly, generative models demonstrate greater understanding and accuracy when outputting emotions in the form of emoticons, compared to directly generating emotionally labeled text. To this end, by having the model learn the style of human online conversations, the generated content will include both emoticons and text. Figure 4 It is the schematic diagram of the other emoticons provided by the embodiment of this application. The content generated by the large model is as follows: emoticon 1 + I accidentally fell today. The emoticon 1 is Figure 4 The first emoji in .
[0097] In a conversational environment, the same sentence can express completely different emotions in different contexts. To this end, historical conversation information is introduced as contextual background and fed into the big model to help the big model understand the user's emotions. For example, if the user inputs "You are so smart!", it may be understood as a compliment without the historical conversation context. However, in the context of the historical conversation {Robot: Sorry, I accidentally fell; User: You are so smart!}, the user's words are sarcastic and mocking. Therefore, the big model will generate an emoticon for the user's input to accurately represent the user's current emotion. For example, the user's words in the above context can be converted to {User: Emoji 2 + You are so smart!}, where Emoji 2 is Figure 4 The second emoji in , by introducing historical conversation information and user emojis, can help the large model generate robot reply content and emojis that are more in line with the current context.
[0098] Specifically, if you want the large model to output emoticons, you can use the following method:
[0099] Method 1: Prompt Project: Add an output format to the prompt words of the large model, requiring the use of emoticons to express emotions at the beginning of the reply sentence. For the specific prompt word settings, please refer to the previous example and will not be repeated here.
[0100] Method 2: Large-Model SFT Fine-Tuning Project: If Method 1 cannot fully control the large-model's emoji output, fine-tune the large-model by simulating training data based on robot conversation scenarios. Thousands of training data are required, and the data format can be seen in Table 2.
[0101] The specific method of using the large model to output emoticons can be determined according to actual conditions, and this application does not make any specific restrictions on this.
[0102] (2) Emotion and Emotion Intensity Matching
[0103] The emotional labels and emotional intensities of the emoticons output by the large model are matched and mapped to obtain the emotional labels and intensity values required for the robot's reply.
[0104] The large model has a strong understanding and accuracy of the emotions associated with emoticons. Emoticons themselves contain information about both emotion and intensity. Therefore, we built an emotion system for the robot, which includes basic emotion labels and their corresponding intensity. It also assigns emotion and intensity labels to commonly used emoticons, creating a mapping table between emoticons, emotion labels, and intensity. The scoring rules for emotion intensity are described previously and will not be repeated here.
[0105] (3) Robots’ expression of different emotions and emotional intensities
[0106] Based on the emotion tag and intensity value, the robot selects the corresponding voice tone and voice to announce the reply content, and also performs the corresponding expression and action. It is understood that if the robot is muted, the reply content will not be announced, and the reply content will be displayed in text instead.
[0107] The robot uses text-to-speech (TTS) technology that controls emotion and intensity. This allows it to synthesize speech with varying emotional tones using the same timbre, even at varying emotional intensities. Because TTS primarily uses tone to classify emotions, which differs from the emotion labels input by the robot, a mapping table has been developed between robot and TTS emotion labels. Normally, the emotion label received by the robot is the TTS emotion label to be used. However, when the TTS tone varies significantly for the same emotion at varying intensities, a TTS emotion mapping is required. For example, when an emoji representing mild or moderate sadness is received, the TTS emotion label is also "sad." However, when an emoji representing severe or extreme sadness is received, the TTS emotion label needs to be mapped to "crying." Specifically, {"emotion"="sad","emotion_level"="6"} is mapped to {"TTS_emotion"="cry","emotion_level"="6"}. Using the crying emotion style allows TTS to more realistically convey severe and extreme sadness.
[0108] Furthermore, the robot can design matching action and expression resources based on emotion labels of varying intensities. To enrich the robot's emotional expression, the same emotion intensity can have multiple expressions and actions, which the robot can randomly select and execute. For example, the moderately sad expression resources include e1 (half-drooped eyelids) and e2 (drooping eyelids, eyes looking to the side), and the action resources include a1 (lowering head) and a2 (lowering head and shaking head). When the robot receives the moderately sad emotion {"emotion"="sad","emotion_level"="3"}, it will randomly select one of e1 and e2, and a1 and a2, respectively, based on the emotion and intensity. Even in the "normal" emotion state, the robot can randomly select and execute a series of resources, such as the "wink" expression and the "look at the user" action. Figure 5 This is a schematic diagram of the expression resources provided in the embodiment of this application. Figure 5 Part (a) is the aforementioned expression resource e1. Figure 5 Part (b) is the aforementioned expression resource e2.
[0109] Based on the above content, the embodiment of the present application provides an example flow diagram of robot interaction, such as Figure 6 As shown, the following steps are included:
[0110] Step 601: The robot detects that the user long presses the AI dialogue button
[0111] Step 602: The robot performs the following actions: vibrates for 0.5 seconds (cancels if the vibration affects the reception); plays listening.bin in a loop; and starts recording.
[0112] When the user long presses the AI dialogue button, the robot vibrates to indicate that it can recognize the recording, and then loops the corresponding expression resource "listening" to show that the robot is listening to the voice. In addition, this robot can record for up to 60 seconds. Figure 7 is a schematic diagram of emoticons in the robot interaction process provided by the embodiment of the present application. For example, the listening expression displayed in step 602 is as follows: Figure 7 As shown in part (a).
[0113] Step 603: The robot detects that the user releases the AI dialogue button.
[0114] After recording, the user releases the AI dialogue button.
[0115] Step 604: The robot verifies whether the user's pass membership has expired.
[0116] If the pass membership has expired, execute step 605 , otherwise execute step 606 .
[0117] Step 605: The robot displays the following on its display: "A is no longer able to travel in the human world" and "Emoji 3." Furthermore, the robot displays "APP download QR code" and a prompt: "Scan the code to help A renew his pass."
[0118] When the pass membership has expired, the robot displays the following content in small letters on its display interface: "A can no longer travel in the human world:" and emoticon 3. Furthermore, the robot also displays "APP download QR code" and a prompt message on its display screen - "Quickly scan the code to help A renew his pass", where the APP download QR code can be displayed on the right side of the text. Among them, emoticon 3 is Figure 4 The third emoji in the . A is the name of the robot.
[0119] Step 606: The robot plays the successfully sent audio [send.mp3].
[0120] Step 607: The robot performs the following actions: first vibrate for 0.5 seconds, play waiting.bin in a loop, and play [waiting.mp3] in a loop.
[0121] After receiving the user's voice, the robot needs to process it and get the corresponding reply. At this time, the robot can first vibrate and play the corresponding expression resource of waiting in a loop. At the same time, it can also play the corresponding audio of waiting, such as waiting.mp3. For example, the corresponding expression resource of waiting in step 607 is as follows: Figure 7 (b) The emoticons shown in part. When the robot waits too long, it can switch to other emoticon resources and display the corresponding emoticons. The specific waiting time for switching emoticon resources can be set according to the actual situation.
[0122] Step 608: Determine whether to return reply content.
[0123] Determine whether the backend returns the corresponding reply content. If the return timeout occurs, execute step 609; otherwise, execute step 610.
[0124] Step 609: The robot displays the following content on its display screen: "Oops! I didn't hear it clearly just now. Can you please say it again?", and vibrates for 0.5 seconds, and plays the animation familiar.bin and audio [miss.mp3] once.
[0125] If the return timeout is exceeded, the device may vibrate for 0.5 seconds and display or play the text "Oops! I didn't hear clearly just now, can you please say it again?", play the expression resource corresponding to the familiar expression once, and play the miss audio. Figure 7 The emoticons shown in part (c).
[0126] Step 610: Is the robot in silent mode?
[0127] Determine whether the robot is in silent mode. If it is in silent mode, execute steps 611-612; if it is not in silent mode, execute steps 613-614.
[0128] Step 611: The robot displays the following on its display: Hey, you know what? Sometimes I feel like we're like two little animals in the forest, looking out for each other. For example, last time...
[0129] When the robot is set to silent mode, only text is displayed without playing any sound. For example, the robot displays the following on its screen: Hey, you know what? Sometimes I feel like we're like two little animals in the forest, looking out for each other. For example, last time...
[0130] Step 612: The user performs the next operation.
[0131] Exemplarily, the user operation includes the user performing one or more of the following: 1) pressing the intercom button - exiting and entering the intercom interface; 2) short pressing the AI button - returning to Big Eyes and playing the short press AI performance; 3) long pressing the AI button - returning to Big Eyes and playing the listening performance; 4) increasing the volume - returning to Big Eyes and exiting mute, and no text will be displayed in the next reply; 5) shaking - no response.
[0132] Step 613: The robot receives the reply and drives the large model content to play the expression reply of the large model, and calls the expression file according to the emotions generated by the large model: sorry, happy, sad (sad), angry, afraid (shocked), disgusted, surprised, and coquettish (kiss).
[0133] Step 614: Return to active state, and reset the active state timer to 5 minutes.
[0134] It will be understood that the vibration duration and maximum recording duration mentioned in the embodiments of the present application are examples and can be determined based on actual conditions without any specific limitation.
[0135] According to the technical solution of the embodiment of the present application, in a text environment, emoticons can easily and accurately convey emotional information compared to text. Similarly, compared to directly generating emotion label text, the generative large model has a stronger understanding and accuracy when outputting emotions in the form of emoticons. By utilizing this feature, the robot's emotions can be generated more accurately; by performing emotion classification and emotion intensity calculation on emoticons, the emoticons output by the large model are converted into emotion labels and emotion intensity values that can be used by the robot; the robot selects corresponding expressions and voice intonations and tones for dialogue based on different emotion labels and emotion intensities, achieving richer emotional expression. In this way, the technical solution of the embodiment of the present application is more accurate in emotion recognition and emotion intensity calculation than the previous method of classifying emotions and calculating emotion intensity through emotion phrases; the robot's emotional expression methods in the technical solution of the embodiment of the present application are richer: it supports voice, expression and action expression of different emotions.
[0136] Figure 8 is a schematic diagram of the structure of the emotion expression device provided in the embodiment of the present application, such as Figure 8 As shown, the emotion expression device includes:
[0137] An acquiring unit 801 is configured to acquire first conversation information output by a first user to the robot;
[0138] Processing unit 802 is configured to generate second conversation information output by the robot to the first user based on the first conversation information; the second conversation information includes a first emoticon; determine a first emotion label and / or a first emotion intensity corresponding to the first emoticon; and control the robot to output a first emotion-related behavior corresponding to the first emotion label and / or the first emotion intensity, where the first emotion-related behavior includes at least one of outputting corresponding voice interaction information, outputting a corresponding emoticon, and performing a corresponding action, and the voice interaction information is included in the second conversation information.
[0139] In some embodiments, the processing unit 802 is used to control the robot to announce the voice interaction information according to the voice intonation and / or tone corresponding to the first emotion label and / or the first emotion intensity.
[0140] In some embodiments, the processing unit 802 is used to control the robot to select a target expression from the expression resources corresponding to the first emotion label and / or the first emotion intensity, and call the selected expression resources to display the target expression, where the target expression includes a static frame or multiple dynamic frames.
[0141] In some embodiments, the processing unit 802 is configured to control the robot to select a target action from action resources corresponding to the first emotion label and / or the first emotion intensity, and call the selected action resources to display the target action.
[0142] In some embodiments, the processing unit 802 is configured to identify keywords in the first session information, retrieve associated session information from historical session information between the first user and the robot based on the keywords, generate second session information corresponding to the first session information by performing semantic understanding and language processing on the associated session information, and output the second session information to the first user.
[0143] In some embodiments, the processing unit 802 is used to determine the first emotion label and / or first emotion intensity corresponding to the first emoticon based on mapping relationship information; wherein the mapping relationship information includes a mapping relationship between at least one group of emoticons and emotion labels and / or emotion intensities.
[0144] In some embodiments, the processing unit 802 is used to control the robot to perform a second emotion-related behavior corresponding to a first interaction operation in response to the first interaction operation, where the first interaction operation includes one or more of touch, stroking, or button operations perceived by the robot body.
[0145] Those skilled in the art should understand that Figure 8 The functions implemented by each unit in the emotion expression device shown can be understood by referring to the relevant description of the aforementioned method. Figure 8 The functions of the various units in the illustrated emotion expression device may be implemented by a program running on a processor, or by a specific logic circuit.
[0146] Figure 9 This is a schematic structural diagram of an emotion expression device 900 provided in an embodiment of the present application. Figure 9 The emotion expression device 900 shown includes a processor 910, which can call and run a computer program from a memory to implement the method in the embodiment of the present application.
[0147] Alternatively, as Figure 9 As shown, the emotion expression device 900 may further include a memory 920. The processor 910 may call and run a computer program from the memory 920 to implement the method in the embodiment of the present application.
[0148] The memory 920 may be a separate device independent of the processor 910 , or may be integrated into the processor 910 .
[0149] Alternatively, as Figure 9As shown, the emotion expression device 900 may further include a transceiver 930 , and the processor 910 may control the transceiver 930 to communicate with other devices, specifically, to send information or data to other devices, or to receive information or data sent by other devices.
[0150] The transceiver 930 may include a transmitter and a receiver. The transceiver 930 may further include an antenna, and the number of antennas may be one or more.
[0151] The emotion expression device 900 can implement the corresponding processes implemented by the emotion expression device in each method of the embodiments of the present application. For the sake of brevity, they are not described here.
[0152] Figure 10 It is a schematic structural diagram of the chip of an embodiment of the present application. Figure 10 The chip 1000 shown includes a processor 1010, which can call and run a computer program from a memory to implement the method in the embodiment of the present application.
[0153] Alternatively, as Figure 10 As shown, the chip 1000 may further include a memory 1020. The processor 1010 may call and execute a computer program from the memory 1020 to implement the method in the embodiment of the present application.
[0154] The memory 1020 may be a separate device independent of the processor 1010 , or may be integrated into the processor 1010 .
[0155] Optionally, the chip 1000 may further include an input interface 1030. The processor 1010 may control the input interface 1030 to communicate with other devices or chips, and specifically, may obtain information or data sent by other devices or chips.
[0156] Optionally, the chip 1000 may further include an output interface 1040. The processor 1010 may control the output interface 1040 to communicate with other devices or chips, and specifically, may output information or data to other devices or chips.
[0157] This chip can implement the corresponding processes implemented by the emotion expression device in each method of the embodiments of the present application. For the sake of brevity, they will not be described here.
[0158] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.
[0159] It should be understood that the processor of the embodiment of the present application may be an integrated circuit chip with signal processing capabilities.
[0160] An embodiment of the present application also provides a computer program product, including a computer program.
[0161] When executed by the processor, the computer program implements the corresponding processes implemented by the emotion expression device in each method of the embodiments of the present application. For the sake of brevity, they are not described here in detail.
[0162] An embodiment of the present application also provides a computer-readable storage medium for storing a computer program.
[0163] The computer program enables the computer to execute the corresponding processes implemented by the emotion expression device in each method of the embodiments of the present application. For the sake of brevity, they are not described here in detail.
[0164] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0165] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0166] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0167] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0168] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0169] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0170] The above is only a specific implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the scope of protection of the present application.
Claims
1. A method for expressing emotions, characterized in that: The method comprises: Obtaining first conversation information input by a first user to the robot; generating, according to the first conversation information, second conversation information output by the robot to the first user; the second conversation information including a first emoticon; Determining a first emotion label and / or a first emotion intensity corresponding to the first emoticon; Control the robot to output a first emotion-related behavior corresponding to the first emotion label and / or the first emotion intensity, where the first emotion-related behavior includes at least one of outputting corresponding voice interaction information, outputting a corresponding expression, and performing a corresponding action, and the voice interaction information is included in the second conversation information.
2. The method according to claim 1, characterized in that Controlling the robot to output voice interaction information corresponding to the first emotion label and / or the first emotion intensity includes: The robot is controlled to announce the voice interaction information according to the voice intonation and / or tone corresponding to the first emotion label and / or the first emotion intensity.
3. The method according to claim 1, characterized in that Controlling the robot to output an expression corresponding to the first emotion label and / or the first emotion intensity includes: The robot is controlled to select a target expression from expression resources corresponding to the first emotion label and / or the first emotion intensity, and the selected expression resources are called to display the target expression, where the target expression includes a static frame or multiple dynamic frames.
4. The method according to claim 1, wherein Controlling the robot to perform an action corresponding to the first emotion label and / or the first emotion intensity includes: The robot is controlled to select a target action from action resources corresponding to the first emotion label and / or the first emotion intensity, and the selected action resource is called to display the target action.
5. The method according to any one of claims 1 to 4, characterized in that Generating second conversation information output by the robot to the first user according to the first conversation information includes: Identify keywords in the first conversation information, retrieve related conversation information from historical conversation information between the first user and the robot based on the keywords, generate second conversation information corresponding to the first conversation information by performing semantic understanding and language processing on the related conversation information, and output the second conversation information to the first user.
6. The method according to any one of claims 1 to 5, characterized in that Determining a first emotion label and / or a first emotion intensity corresponding to the first emoticon includes: Determine a first emotion label and / or a first emotion intensity corresponding to the first emoticon according to mapping relationship information; wherein the mapping relationship information includes a mapping relationship between at least one set of emoticons and emotion labels and / or emotion intensities.
7. The method according to any one of claims 1 to 5, characterized in that The method further comprises: In response to a first interactive operation, the robot is controlled to perform a second emotion-related behavior corresponding to the first interactive operation, where the first interactive operation includes one or more of touch, stroking, or key operations perceived by the robot body.
8. An emotion expression device, characterized in that: The device comprises: an acquiring unit, configured to acquire first conversation information input by a first user to the robot; A processing unit is configured to generate, based on the first conversation information, second conversation information output by the robot to the first user; the second conversation information includes a first emoticon; determine a first emotion label and / or a first emotion intensity corresponding to the first emoticon; and control the robot to output a first emotion-related behavior corresponding to the first emotion label and / or the first emotion intensity, the first emotion-related behavior including at least one of outputting corresponding voice interaction information, outputting a corresponding expression, and performing a corresponding action, the voice interaction information being included in the second conversation information.
9. An emotional expression device, characterized in that: include: A processor and a memory, the memory being used to store a computer program, the processor being used to call and run the computer program stored in the memory to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that Used to store a computer program, wherein the computer program causes a computer to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Touch communication doll capable of vividly expressing emotional information
CN101820399A
Multi-mode interaction method and system for intelligent robot
CN106985137A
Emotion reply generation device with emoticons based on deep learning
CN117194621A
Cited By
Chat interaction method and device, equipment and medium
CN121217682A