Presentation device, presentation method, and presentation program
By evaluating nonverbal emotions and controlling nodding motions based on synchronization levels, the presentation device enhances the accuracy and appropriateness of humanoid agent responses, addressing the limitations of relying solely on verbal cues.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- NT T INC
- Filing Date
- 2024-11-05
- Publication Date
- 2026-05-15
AI Technical Summary
Existing presentation devices using humanoid agents fail to accurately interpret user speech due to reliance on verbal cues alone, leading to inappropriate nodding responses that may not align with the user's true understanding or comfort level.
The presentation device incorporates an evaluation unit to assess synchronization with the user based on nonverbal emotions, using a combination of video and audio analysis to control nodding motions of the humanoid agent, considering personality traits and context.
Improves the accuracy of the agent's response actions by aligning them more appropriately with the user's emotional and nonverbal cues, enhancing the understanding and comfort of the interaction.
Smart Images

Figure JP2024039214_15052026_PF_FP_ABST
Abstract
Description
Presentation device, presentation method, and presentation program
[0001] This invention relates to a presentation device, a presentation method, and a presentation program.
[0002] Services have been proposed that use humanoid computer graphics (CG) or robots with bodily expressions (hereinafter referred to as humanoid agents) to provide functions such as assistants, presenters, and counselors. In such situations, when humanoid agents listen to and respond to users, methods are being considered to adjust the type of head movement (nodding or turning the head from side to side), frequency, speed, and timing according to the degree of negativity or positivity of the user's statements, thereby expressing agreement or denial (Non-patent documents 1 and 2).
[0003] Nishida, Ota, Watanabe, Ishii, "A bodily engagement character system that performs reaction actions based on the emotional polarity of words in utterances," Transactions of the Japan Society of Mechanical Engineers, Vol. 83, No. 846, pp. 1-14, 2017. Nishida, Ishii, Watanabe, "A voice-driven bodily engagement character system that performs delayed vocal responses in response to the negative polarity of utterances," Transactions of the Japan Society of Mechanical Engineers, Vol. 87, No. 987, pp. 1-13, 2021.
[0004] Figure 18 shows an example of the configuration of a prior art presentation device. Figure 19 is a flowchart of a prior art presentation method.
[0005] The presentation device 110 described in Non-Patent Documents 1 and 2 has a user speech recognition unit 1151 that converts user utterances input from an input unit 111 or a communication control unit 113 into text. The presentation device 110 also has an emotion polarity analysis unit 1152 that analyzes the emotion polarity of utterances from the text resulting from speech recognition, and a user utterance word emotion dictionary 1141 used for emotion polarity analysis.
[0006] The presentation device 110 includes an agent nodding generation unit 1153 that performs nodding actions for the agent, a body motion database (DB) 1142 that stores the actions performed by the humanoid agent during nodding actions, and an agent speech content DB 1143 that stores the content of the humanoid agent's speech.
[0007] The prompting device 110 acquires the user's spoken voice input (step S201 in FIG. 19), and converts the user's speech into text by voice recognition (step S202 in FIG. 19). Subsequently, the prompting device 110 divides the text of the voice recognition result into words by morphological analysis (step S203 in FIG. 19), refers to the user speech word sentiment dictionary 1141, and calculates the sentiment polarity (negative / positive) of the speech sentence (step S204 in FIG. 19).
[0008] The prompting device 110 selects a body movement corresponding to the sentiment polarity from the body movement DB 1142, and the prompting unit 1154 outputs it as the body movement (for example, nodding) of the humanoid agent (step S205 in FIG. 19). Thereby, the prompting device 110 induces the negative content of the user's speech into positive content.
[0009] In the prior art, since nodding is performed only based on language, there may be cases where the listener returns a nod that does not convey true understanding.
[0010] For example, in the prior art, in a situation where someone says "Well, I think it's okay, but...", tilting their head and speaking in a low voice, the non-verbal cues of the speaker may be missed and it may be judged as a positive statement, and the humanoid agent may respond brightly and in agreement. Also, in the prior art, for example, when an introverted user is asking a question, excessive nodding may cause discomfort, which may be different from the understanding (recognition) of the speaker.
[0011] The present invention has been made in view of the above, and an object thereof is to provide a prompting device, a prompting method, and a prompting program that can improve the accuracy with which an agent interprets the content of a user's speech and enable a more appropriate expression of response actions.
[0012] To solve the above-mentioned problems and achieve the objective, the presenting device of the present invention is characterized by comprising at least an evaluation unit that evaluates the degree of synchronization with the user, which is expressed by the nodding motion of a humanoid agent responding to the user, based on nonverbal emotions, which are the user's emotional information that appears in the user's nonverbal information, and a control unit that controls the nodding motion of the humanoid agent responding to the user, based at least on the synchronization level.
[0013] Furthermore, the presenting method of the present invention is a presenting method performed by a presenting device, and is characterized by including at least the steps of: evaluating the synchronization level, which is the degree of synchronization with the user expressed by the nodding motion of a humanoid agent responding to the user, based on nonverbal emotions, which are the user's emotional information that appears in the user's nonverbal information; and controlling the nodding motion of the humanoid agent responding to the user based at least the synchronization level.
[0014] Furthermore, the program presented by the present invention causes a computer to perform at least the following steps: evaluate the synchronization level, which is the degree of synchronization with the user expressed by the nodding movements of a humanoid agent responding to the user, based on nonverbal emotions, which are the user's emotional information that appears in the user's nonverbal information; and control the nodding movements of the humanoid agent responding to the user based at least the synchronization level.
[0015] According to the present invention, it is possible to improve the accuracy with which the agent interprets the user's statements and enable the expression of more appropriate response actions.
[0016] Figure 1 is a diagram showing an example of the configuration of a presentation device according to an embodiment. Figure 2 is a diagram showing an example of a list of evaluation criteria for the synchronization level. Figure 3 is a diagram showing an example of a list of evaluation criteria for the synchronization level. Figure 4 is a diagram illustrating a list of nodding actions corresponding to the synchronization level. Figure 5 is a diagram illustrating a list of nodding actions corresponding to the synchronization level. Figure 6 is a diagram illustrating an example of updating the list of nodding actions. Figure 7 is a diagram illustrating an example of updating the list of nodding actions. Figure 8 is a diagram showing an example of updating the list of nodding actions. Figure 9 is a diagram showing an example of updating the list of nodding actions. Figure 10 is a diagram showing an example of the processing procedure of the presentation method according to an embodiment. Figure 11 is a diagram showing an example of the processing procedure of the presentation method according to an embodiment. Figure 12 is a diagram showing an example of the processing procedure of the presentation method according to an embodiment. Figure 13 is a diagram showing an example of the processing procedure of the presentation method according to an embodiment. Figure 14 is a diagram showing an example of the processing procedure of the presentation method according to an embodiment. Figure 15 is a diagram showing an example of the processing procedure of the presentation method according to an embodiment. Figure 16 is a diagram showing an example of the processing procedure of the presentation method according to an embodiment. Figure 17 is a diagram showing an example of a computer in which the presentation device is realized by the execution of a program. Figure 18 shows an example of the configuration of a prior art presentation device. Figure 19 is a flowchart of a prior art presentation method.
[0017] Hereinafter, one embodiment of the present invention will be described in detail with reference to the drawings. However, the present invention is not limited to this embodiment. Furthermore, in the drawings, the same parts are denoted by the same reference numerals.
[0018] [Embodiment] [Overview of Presentation Device] The presentation device according to this embodiment analyzes the user's personality traits, intentions, and context from non-verbal information such as video, audio, and pre-registered user information, and uses this in combination with linguistic information to improve the accuracy of the humanoid agent's interpretation of the user's statements and enable the expression of more appropriate response actions. In other words, the presentation device utilizes non-verbal information such as video and audio, and also takes into account the emotions expressed in facial expressions obtained from the video and the intonation of the voice to determine the appropriate response. The presentation device causes the humanoid agent to perform a response that corresponds to the user's personality traits, emotions at the time of speaking, intentions, and the context of the conversation, obtained from non-verbal information such as video and audio. Hereafter, a response will include at least one of a response utterance or a response action. A response utterance refers to verbal responses (for example, saying "That's right," or "I see," in response to what the other person says). A response action refers to head movements such as nodding or shaking the head.
[0019] Here, the presentation device performs an overall evaluation based on emotional information appearing in nonverbal information (hereinafter referred to as nonverbal emotion) and emotional information appearing in the text resulting from speech recognition (hereinafter referred to as verbal emotion), and determines the synchronization level. Based on the synchronization level, the presentation device controls the nodding actions of the humanoid agent responding to the user. The synchronization level is the degree of synchronization with the user expressed by the humanoid agent's nodding actions. The nodding actions are expressed to the user by the humanoid agent's head movements. The presentation device controls the nodding utterances and also controls the type of head movement (nodding or left-right head shake), the number of times, the speed, and the timing of the nodding actions. In this way, the presentation device estimates how much synchronization should be expressed (synchronization level) by taking into account the emotions estimated from facial expressions and vocal intonation, and determines the nodding expression according to the user's emotions. The presentation device may also add hand expressions to the nodding actions, and may control the up-and-down movement of the hands, the number of hand gestures, and the size of the hand gestures, similar to the head movements.
[0020] [Configuration of the Presentation Device] The presentation device according to the embodiment will now be described. Figure 1 is a diagram showing an example of the configuration of the presentation device according to the embodiment.
[0021] As shown in Figure 1, the display device 10 is implemented using a general-purpose computer such as a personal computer, and includes an input unit 11, an output unit 12, a communication control unit 13, a storage unit 14, and a control unit 15.
[0022] The input unit 11 is implemented using input devices such as a keyboard, mouse, camera, and microphone, and in response to input operations by the operator, it inputs various instruction information, such as processing start, to the control unit 15.
[0023] The output unit 12 is implemented by a display device such as a liquid crystal display, a printing device such as a printer, etc. For example, the output unit 12 outputs images and sounds of the humanoid agent's nodding motion.
[0024] The communication control unit 13 is implemented using a NIC (Network Interface Card) or the like, and controls communication between the control unit 15 and external devices via telecommunication lines such as a LAN (Local Area Network) or the Internet.
[0025] The memory unit 14 is implemented using semiconductor memory elements such as RAM (Random Access Memory) or flash memory, or storage devices such as hard disks or optical discs. The memory unit 14 pre-stores processing programs for operating the presentation device 10, as well as data used during the execution of the processing programs, or temporarily stores them each time processing is performed. The memory unit 14 may also be configured to communicate with the control unit 15 via the communication control unit 13.
[0026] The memory unit 14 stores information for interpreting the user's statements and information for controlling the humanoid agent's responses. The memory unit 14 includes a user utterance word emotion dictionary 141, an agent body movement DB 142, an agent utterance content DB 143, a user information DB 144, an agent character DB 145, and a communication history DB 146.
[0027] The user utterance word sentiment dictionary 141 is, for example, a dictionary used for overall evaluation, where text is associated with its corresponding sentiment score. The user utterance word sentiment dictionary 141 may also be a dictionary where text is associated with sentiment labels (positive, negative, neutral).
[0028] The Agent Body Movement DB 142 stores various patterns of humanoid agent body movements, including the type of head movement used for acknowledgment (nodding or shaking the head from side to side), the angle, number of times, speed, timing of nodding and shaking the head, and facial expressions during acknowledgment. For example, the body movements of a humanoid agent are pre-registered in the Agent Body Movement DB 142. Furthermore, if hand movements are added as part of the acknowledgment, the Agent Body Movement DB 142 stores various patterns such as the up-and-down movement of the hand, the number of times the hand is shaken, and the size of the hand shake.
[0029] The Agent Utterance Content DB 143 stores the content of the utterances spoken by the humanoid agent when acknowledging or responding to a conversation. The utterances are pre-registered in the Agent Utterance Content DB 143.
[0030] User Information DB 144 stores information indicating the user's personal characteristics. These characteristics include, for example, the user's attribute information and personality traits. They also include personality traits such as whether the user is introverted and to what extent, whether the user is extroverted and to what extent, and whether the user is neutral. Furthermore, User Information DB 144 may store whether the user's feelings towards their interaction with the humanoid agent are introverted, extroverted, or neutral. The user's personal characteristics are registered in advance. The user's personal information is obtained through methods such as surveys. Alternatively, the user's personal information may be automatically estimated during conversations with the system and updated as needed.
[0031] Personality traits are often represented by the Big Five. The Big Five theory posits that an individual's personality traits consist of four elements: neuroticism (N), extraversion (E), openness to experience (O), agreeableness (A), and conscientiousness (C). Personality traits are analyzed by measuring these Big Five characteristics through questionnaires, etc. (References 1 and 2). Reference 1: “Glossary of Psychological Terms,” [Retrieved October 21, 2024], Internet <URL: https: / / psychologist.x0.com / terms / 154.html> Reference 2: “Big Five,” [Retrieved October 21, 2024], Internet <https: / / psychoterm.jp / basic / personality / bigfive>
[0032] The Agent Character DB 145 stores character information for humanoid agents. This character information may include, for example, character images for various types of humanoid agents, and personal characteristics of various types of humanoid agents (attributes, personality traits (such as being introverted and to what extent, being extroverted and to what extent, being neutral, etc.)). A humanoid agent character may have one or more patterns. If there are multiple patterns, the user can specify a character, or the presentation device 10 can automatically set and switch between characters as needed. The humanoid agent character may also be updated as needed via the communication control unit 13 or the control unit 15.
[0033] The communication history DB 146 stores the history of communication between the user and the humanoid agent. The communication history DB 146 stores audio recordings of the user's interaction with the humanoid agent, as well as video footage of the user interacting with the humanoid agent.
[0034] The control unit 15 is implemented using a CPU (Central Processing Unit) or the like, and executes a processing program stored in memory. As a result, the control unit 15 functions as a user video analysis unit 151, a user voice recognition unit 152, a nonverbal emotion calculation unit 153, a verbal emotion calculation unit 154, an overall evaluation unit 155 (evaluation unit), an agent response generation unit 156 (control unit), and a presentation unit 157, as illustrated in Figure 1, and executes presentation processing. Note that each or part of these functional units may be implemented on different hardware. Furthermore, the control unit 15 may also include other functional units.
[0035] The user video analysis unit 151 acquires video of the user using the humanoid agent from the input unit 11 and the communication control unit 13. The user video analysis unit 151 recognizes the user's facial expressions and body movements from the user's video, creates input information for the nonverbal emotion calculation unit 153, and outputs the created information to the nonverbal emotion calculation unit 153. The user video analysis unit 151 may also detect facial expression labels (such as smiling or neutral) and action labels (such as raising a hand or shaking one's head) using existing recognition technology (for example, Reference 3). Reference 3: MediaGnosis, “Next-Generation Media Processing AI,” [Retrieved October 21, 2024], Internet <URL: https: / / www.rd.ntt / mediagnosis / >
[0036] The user speech recognition unit 152 acquires the user's speech utterances from the input unit 11 when communicating with the humanoid agent. The user speech recognition unit 152 recognizes the user's speech and converts it into text using speech recognition technology. The user speech recognition unit 152 creates the user's speech text, which is the converted speech of the user, as input information for the language sentiment calculation unit 154 and outputs it to the language sentiment calculation unit 154.
[0037] The nonverbal emotion calculation unit 153 calculates a user emotion score as a nonverbal emotion, based on the voice tone information analyzed from the user's voice obtained from the input unit 11, and / or the input from the user video analysis unit 151, which is displayed in the user's nonverbal table. The nonverbal emotion calculation unit 153 outputs the calculated user emotion score to the overall evaluation unit 155. The nonverbal emotion calculation unit 153 may also calculate the result of estimating nonverbal emotions using an existing method for estimating emotions based on voice and video (for example, Reference 4). Reference 4: Akihiko Takashima, Naoki Makishima, Mana Ihori, Tomohiro Tanaka, Shota Orihashi, Ryo Masumura, “Unsupervised Domain Adversarial Training in Angular Space for Facial Expression Recognition,” In Proc. Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC), 1054-1059, 2020.
[0038] The nonverbal emotion calculation unit 153 may calculate nonverbal emotions for each video frame in response to the video input, or it may analyze nonverbal emotions using video and audio of a predetermined time range.
[0039] Furthermore, the nonverbal emotion calculation unit 153 may calculate nonverbal emotions by adjusting the time interval used according to the length of the utterance used for calculating verbal emotions. Also, if meaningful events such as the changing of people or excitement are observed in the video, the nonverbal emotion calculation unit 153 may select that interval and calculate nonverbal emotions.
[0040] Furthermore, the nonverbal emotion calculation unit 153 may analyze the audio and video in combination to analyze the level of excitement in the conversation and the overall atmosphere of the group (silence, movement and conversation, excitement, etc.), and select the interval for calculating nonverbal emotions based on the results.
[0041] Furthermore, the nonverbal emotion calculation unit 153 does not need to use both video and audio when calculating nonverbal emotions; it may use only one of them.
[0042] The nonverbal emotion calculation unit 153 may calculate nonverbal emotions as discrete values in multiple stages such as positive / negative / neutral, or as continuous values between 0.0 and 1.0 or between -1.0 and 1.0.
[0043] Furthermore, the nonverbal emotion calculation unit 153 may, after receiving an emotion estimate as a continuous value, set an arbitrary threshold to convert it into discrete, multi-stage nonverbal emotion values.
[0044] Furthermore, if multiple nonverbal emotions are extracted within the analysis interval, the nonverbal emotion calculation unit 153 determines which nonverbal emotion to adopt using, for example, one of the following methods: the first method, the second method, the third method, or the fourth method.
[0045] The first method is a majority vote for each emotion, that is, when nonverbal emotions are discrete values such as positive / negative, the method adopts the nonverbal emotion with the highest number of detections. The second method adopts the first or last detected emotion. The third method adopts the mean or median of the detected nonverbal emotion values when the nonverbal emotions are between 0.0 and 1.0, with lower values indicating a progression towards negativity. The fourth method adopts the nonverbal emotion at a specific point in time when priority is determined within the analysis interval (for example, when the peak of the above-mentioned surge is determined at a specific time).
[0046] The language emotion calculation unit 154 estimates the emotions contained in the user's utterances from the text input from the user speech recognition unit 152 and calculates them as language emotions. The text input from the user speech recognition unit 152 is the user's spoken text converted from the user's voice by speech recognition. The language emotion calculation unit 154 outputs the calculated language emotions to the comprehensive evaluation unit 155 as input information to the comprehensive evaluation unit 155.
[0047] The language emotion calculation unit 154 calculates, for example, an emotion score for each word in the input user's utterance text from the user utterance word emotion dictionary 141, and calculates the sum or average value of these scores as the language emotion of the utterance.
[0048] Alternatively, the language emotion calculation unit 154 may pre-create a speech emotion estimation model for estimating the emotion of utterances using a learning method such as supervised learning, and adopt the emotion estimated by the speech emotion estimation model for the user's utterance text as the language emotion.
[0049] The language emotion calculation unit 154 may calculate language emotion for each word, or it may calculate it for each certain unit, such as a sentence or an utterance.
[0050] The linguistic emotion calculation unit 154 may calculate linguistic emotion as a discrete value with multiple levels such as positive / negative / neutral, or as a continuous value between 0.0 and 1.0, or between -1.0 and 1.0.
[0051] Furthermore, the language emotion calculation unit 154 may, after receiving the emotion estimate as a continuous value, set an arbitrary threshold to convert it into discrete, multi-stage language emotion values.
[0052] The overall evaluation unit 155 evaluates the synchronization level based at least on the nonverbal emotions calculated by the nonverbal emotion calculation unit 153. The overall evaluation unit 155 comprehensively evaluates the synchronization level based on nonverbal emotions and verbal emotions. Alternatively, the overall evaluation unit 155 may evaluate the synchronization level based only on verbal emotions. The overall evaluation unit 155 calculates the synchronization level based on the calculation results of nonverbal emotions and verbal emotions, and user responses, and outputs it as input information for adjusting the number of responses in the agent response generation unit 156. The overall evaluation unit 155 may correct the synchronization level according to the user's response (level of understanding, response delay time) to the humanoid agent's response actions.
[0053] The agent nodding generation unit 156 controls the humanoid agent's nodding motion based at least on the synchronization level evaluated by the overall evaluation unit 155.
[0054] The agent response generation unit 156 can select a response action corresponding to the synchronization level from the pre-registered agent body movement DB 142 and agent speech content DB 143. Alternatively, a model may be created in advance by supervised learning that has been trained on synchronization levels and appropriate response actions corresponding to those synchronization levels, and the agent response generation unit 156 may use this created model to determine the response that the humanoid agent will express.
[0055] The agent nodding generation unit 156 controls the nodding behavior of the humanoid agent using at least one of the following: the amount of the user's nonverbal emotions, the ratio of the amount of the user's nonverbal emotions to the amount of the user's verbal emotions, the number of times the user nods in response to the humanoid agent's nods, the number of times the user utters nodding phrases, the angle of the user's neck when nodding, and the user's personality traits.
[0056] The agent response generation unit 156 controls the humanoid agent's response behavior based on the humanoid agent's character and / or the type of context the humanoid agent uses (e.g., small talk, explanation).
[0057] The agent nod generation unit 156 controls the humanoid agent's nods to resemble the user's nods. The agent nod generation unit 156 considers the humanoid agent's nods by referring to the mirror effect, which is the phenomenon in which people tend to feel favorably towards others who perform similar gestures.
[0058] The agent response generation unit 156 may correct the control of the humanoid agent's response actions when interacting with the user based on at least one of the following: the user's personality traits, the degree of similarity between the user's and the humanoid agent's personality traits, the user's response to the humanoid agent's response actions (level of understanding, response delay time, etc.), and the user synchronization level. The user synchronization level is the degree to which the user is in sync with the humanoid agent's speech and / or the humanoid agent's behavior.
[0059] The presentation unit 157 outputs the nodding motion generated by the agent nodding generation unit 156 as a physical movement of the humanoid agent.
[0060] [Processing by the Comprehensive Evaluation Unit] [Example of Processing by the Comprehensive Evaluation Unit 1] An example of processing by the Comprehensive Evaluation Unit 155 will be explained. First, the case in which a humanoid agent nods in agreement will be explained.
[0061] In the following examples of operation, we will describe cases where the user expresses positive emotions and the humanoid agent responds with expressions of agreement. However, the processing of the overall evaluation unit 155 and the agent response generation unit 156 is not limited to agreement with positive emotions; it is also possible to respond to negative emotions by offering counterarguments or expressions that help to make them appear positive, or to adjust the degree of agreement according to corrective information.
[0062] Figure 2 shows an example of a list of evaluation criteria for conformity levels. The evaluation criteria list L1 shown in Figure 2 is used, for example, when nonverbal emotions and verbal emotions are discrete binary values.
[0063] The overall evaluation unit 155 evaluates the synchronization level by referring to the evaluation criteria list L1. The overall evaluation unit 155 determines the degree of synchronization with the user through the humanoid agent's nodding actions, according to the user's positive emotions.
[0064] The overall evaluation unit 155 evaluates the conformity level as high if both verbal and nonverbal emotions are positive (True), according to the evaluation criteria list L1. The overall evaluation unit 155 evaluates the conformity level as medium if only one of the verbal or nonverbal emotions is positive. The overall evaluation unit 155 evaluates the conformity level as low if neither the verbal nor nonverbal emotions are positive (False). The evaluation criteria list L1 is, for example, pre-set.
[0065] [Processing Example 2 of the Comprehensive Evaluation Unit] Figure 3 shows an example of a list of evaluation criteria for conformity level. The evaluation criteria list L2 shown in Figure 3 is used when both nonverbal emotion and verbal emotion are continuous values from 0 to 1.
[0066] When using the L2 evaluation criteria list, arbitrary thresholds can be set for nonverbal and verbal emotions, and the level of agreement can be set by dividing nonverbal and verbal emotions into multiple stages according to those thresholds. Note that the number of stages for nonverbal and verbal emotions does not need to be the same; you may subdivide only the one you want to classify more finely.
[0067] For example, evaluation criterion list L2 is an example where nonverbal emotions can be rated on a three-point scale, and verbal emotions can also be rated on a three-point scale. The evaluation criterion list is designed so that the level of agreement increases when each emotion value is high (in some cases, there may be no change). For example, in evaluation criterion list L2, the level of agreement is shown as a value from 1 to 5.
[0068] For example, the overall evaluation unit 155 evaluates the conformity level as 5 if, for example, both verbal and nonverbal emotions are 0.7, according to the evaluation criteria list L2. Also, the overall evaluation unit 155 evaluates the conformity level as 1 if, for example, verbal emotion is 0.7 and nonverbal emotion is 0.3. The evaluation criteria list L2 is, for example, pre-set.
[0069] [Example 3 of the Comprehensive Evaluation Unit's Processing] An example of the Comprehensive Evaluation Unit 155 calculating the synchronization level as a continuous value based on nonverbal emotions and verbal emotions will be described.
[0070] In this example, the synchronization level takes values from 0.0 to 1.0, for example, with a higher value indicating stronger synchronization by the humanoid agent. The comprehensive evaluation unit 155 uses, for example, equation (1) to determine the synchronization level d a Calculate it as a continuous value.
[0071]
[0072] n h This represents the user's nonverbal emotions. h This represents the user's linguistic sentiment. na This represents the degree of influence of nonverbal emotions. va This represents the degree of influence of linguistic emotion. The subscript 'a' in each symbol represents a humanoid agent, and the subscript 'h' represents the user (person).
[0073] wna and w va is a coefficient whose sum is 1 and has a function of adjusting which of the non-verbal emotion and the verbal emotion is strongly reflected in the synchronization level. w na and w va is represented by, for example, Equation (2). The comprehensive evaluation unit 155 automatically adjusts which of the non-verbal emotion and the verbal emotion is strongly reflected in the synchronization level based on the values of the change amount of the input non-verbal emotion (for example, the number of changes in the user's expression) and the change amount of the verbal emotion (for example, the user's speech volume).
[0074]
[0075] When the values of the non-verbal emotion and the verbal emotion are obtained numerically, the comprehensive evaluation unit 155 may use the average, weighted average, or geometric mean of the values of the non-verbal emotion and the verbal emotion as the synchronization level.
[0076] [Processing Example 4 of Comprehensive Evaluation Unit] Next, an example of correcting the synchronization level according to the user reaction will be described. When the comprehensive evaluation unit 155 calculates the synchronization level as a continuous value, it may add a correction term c and correct it as the final synchronization level.
[0077] Also, when a threshold value is set for the value of the correction term c, the comprehensive evaluation unit 155 may perform correction for the discrete synchronization level according to the threshold value of the correction term c.
[0078] Here, the correction term is a value calculated based on, for example, the degree of understanding of the speech content estimated from the video and voice of the user who listened to the speech of the humanoid agent, the delay time of the user's reaction to the speech content, and the like.
[0079] For example, when performing correction based on the degree of understanding, the degree of understanding is given as 0.0 to 1.0, and it is assumed that the larger the value, the better the understanding. In this case, the comprehensive evaluation unit 155 multiplies by the correction in Equation (3) and adjusts the synchronization level so that the degree of synchronization decreases as the degree of understanding decreases.
[0080]
[0081] Note that in this example, c is Equation (4).
[0082]
[0083] Furthermore, the correction term may be calculated for a single user utterance, or it may be calculated based on a cumulative value over a certain period of time.
[0084] For example, if verbal or nonverbal emotions remain positive for an extended period, the overall evaluation unit 155 may adjust the lower limit of the conformity level by one step, or conversely, if negative emotions remain negative for an extended period, it may adjust the upper limit of the conformity level by one step.
[0085] Furthermore, while we have shown an example of incorporating comprehension as an example of calculating the correction term, this is not the only example. The overall evaluation unit 155 may, for example, use the user's satisfaction level with the conversation prior to the point in time when they nod in agreement, the amount of speech they give, and their engagement (empathy) with the conversation. The overall evaluation unit 155 makes adjustments, such as refraining from agreeing, if the estimated satisfaction level is low.
[0086] Furthermore, the comprehensive evaluation unit 155 may calculate the correction term based on a single input, or it may calculate it based on multiple values, such as a combination of comprehension and satisfaction. Existing technologies can be used for acquiring the information that forms the basis of the correction term and for estimation based on video. For example, the method described in Reference 5 can be used for comprehension. Reference 5: Shunichi Kinoshita, Toshiki Onishi, Naoki Azuma, Ryo Ishii, Atsushi Fukayama, Takao Nakamura, and Akihiro Miyata, “A Study of Prediction of Listener's Comprehension Based on Multimodal Information,” In Proceedings of the 23rd ACM International Conference on Intelligent Virtual Agents (IVA '23), Article 30, pp. 1-4, 2023.
[0087] [Processing Example 5 of the Overall Evaluation Unit] The Overall Evaluation Unit 155 processes the user synchronization level d h To find the synchronization level d, aHowever, corrections may be made to make it similar to the user synchronization level.
[0088] In this case, the comprehensive evaluation unit 155 analyzes the nonverbal and verbal emotions of the user during their nodding actions while the humanoid agent is speaking, and determines the degree to which the user is attuned to the humanoid agent's speech as the user attunement level d. h It is calculated as follows.
[0089] The comprehensive evaluation unit 155 analyzes linguistic emotions by focusing on words spoken by the user immediately after the humanoid agent finishes speaking, or during the agent's speech.
[0090] The overall evaluation unit 155 calculates nonverbal emotions, verbal emotions, and user conformity level, which is the normal conformity level d a This is done using the same process as the calculation for [previous calculation].
[0091] However, user synchronization level d h Since this is the user's level of agreement with the agent's speech, the overall evaluation unit 155 determines the normal level of agreement d a Record them separately.
[0092] Furthermore, the overall evaluation unit 155 calculates the user synchronization level d h At this time, the user's nodding or other affirmative actions are recorded in pairs.
[0093] User synchronization level d h Equation (5) is obtained.
[0094]
[0095] The overall evaluation unit 155 determines the user synchronization level d h The coefficient of the tuning level d a By reflecting this, synchronization level d a This can be corrected as shown in equation (6).
[0096]
[0097] [Processing Example 6 of the Comprehensive Evaluation Unit] The Comprehensive Evaluation Unit 155 may also correct the synchronization level according to the user's response (level of understanding, response delay time) to the humanoid agent's nodding actions. For example, the Comprehensive Evaluation Unit 155 corrects the correction function f as shown in equation (7).a Using (c), the tuning level d a This corrects the correction function f. a This is the input to (c). fa This is a weighting coefficient that adjusts the degree of influence of the correction function.
[0098]
[0099] Correction function f a (c) is equation (8).
[0100]
[0101] Correction function f in equation (7) a (c) is a correction that calculates the average level of understanding when c is a vector of user understanding for a fixed time interval, and lowers the conformity level when the level of understanding is low.
[0102] Also, the correction function f a (c) may be formula (9).
[0103]
[0104] The correction function in equation (9) corresponds to the case where c is the level of conformity within a certain time interval, and the average value of the conformity level within that interval is calculated and used as the new bias for the conformity level. The correction function in equation (9) is used, for example, to adjust for a higher level of conformity in conversations that tend to have a high level of conformity (such as small talk), taking the context into account.
[0105] Equation (7) shows the result of applying a correction function to the basic pattern equation (1), but similarly, correction using a correction function is also possible for equation (6).
[0106] The agent response generation unit 156 can use the corrected synchronization level to correct the agent's response to a response similar to the user's response, or to a response that corresponds to the user's reaction (level of understanding, response delay time).
[0107] [Processing of the Agent Nod Generation Unit] [Example of Processing of the Agent Nod Generation Unit 1] Next, an example of processing of the agent nod generation unit 156 will be described. As processing of the agent nod generation unit 156, the process of selecting a nod action by referring to a list of nod actions will be described. The agent nod generation unit 156 determines the number of nods, the number of nod utterances, and the depth of the nod according to a predefined list of actions.
[0108] First, we will explain the process of generating interjections based on discrete synchronization levels. Figure 4 is an example of a list of interjection actions corresponding to synchronization levels.
[0109] The processing example 156 of the agent response generation unit refers to the response action list L11 shown in Figure 4, and based on the synchronization level evaluated by the overall evaluation unit 155, selects a response to be expressed by the humanoid agent from the pre-registered agent body action DB 142 and agent utterance content DB 143, and adjusts how many times that expression is performed in a certain period of time.
[0110] In the agent response generation unit's processing example 156, it is assumed that the greater the number of nods and responses, the stronger the expression of agreement, and the number of nods and responses to be generated is adjusted according to the response action list L11.
[0111] Furthermore, in the processing example 156 of the agent nodding generation unit, the nodding motion and content of the nodding utterances of the humanoid agent may also be selected from each agent's body motion DB 142 and agent utterance content DB 143 based on the synchronization level. The processing example 156 of the agent nodding generation unit may also utilize pre-set fixed nodding motions.
[0112] [Processing Example 2 of Agent Response Generation Unit] Next, we will explain the response generation process based on a continuous synchronization level. Figure 5 is a diagram illustrating a list of response actions corresponding to the synchronization level.
[0113] The agent response generation unit 156 refers to the response action list L21 shown in Figure 5 and, based on the synchronization level value evaluated by the overall evaluation unit 155, selects a response to be expressed by the humanoid agent from the pre-registered agent body action DB 142 and agent utterance content DB 143, and adjusts how many times that expression is performed in a certain period of time.
[0114] According to the list of nodding actions L21, the agent nodding generation unit 156 can set an arbitrary threshold for the synchronization level and, according to that threshold, set the humanoid agent's nodding actions (nodding motions) and speech content in multiple stages.
[0115] The agent nod generation unit 156, in accordance with the nod action list L21, determines that there is no need to perform any particular synchronization expression if the synchronization level is below 0.6, and has the humanoid agent perform a simple nod. The agent nod generation unit 156, in accordance with the nod action list L21, prioritizes user utterances if the synchronization level is below 0.4, and has the humanoid agent express only physical actions without performing any nod utterances. The agent nod generation unit 156, in accordance with the nod action list L21, operates the humanoid agent to indicate attentive listening if the synchronization level is below 0.2.
[0116] In addition, in the list of affirmative action sequences L11 and L21 in Figures 4 and 5, the agent affirmative action generation unit 156 is set to nod as an affirmative action because it is an example of a synchronized expression. However, in the case of a negative expression, it may be changed to a head shake or the like.
[0117] Furthermore, the agent response generation unit 156 may switch the humanoid agent's response actions and responses not only based on a single synchronization level, but also on the amount of change in the synchronization level over time.
[0118] [Processing Example 3 of Agent Response Generation Unit] The agent response generation unit 156 may automatically update the list of response actions based on the characteristics of the user and the agent.
[0119] The agent nod generation unit 156 may update and optimize a pre-set list of nod actions according to the user's personal characteristics, the character of the humanoid agent, and the context. In addition, the agent nod generation unit 156 may adjust not only the number of nod actions and utterances, but also the depth of the nodding action.
[0120] Figures 6 and 7 illustrate examples of updating the list of affirmative responses.
[0121] For example, if the user's personal characteristics are introverted, the agent response generation unit's processing example 156 selects an introverted character for the humanoid agent. The agent response generation unit 156 utilizes a context pre-specified for the humanoid agent. For example, if the humanoid agent is to introduce products or exhibits, the context to be used is set to "explanation" in advance. Also, for example, if the humanoid agent is to provide a communication service acting like the user's friend, the context is set to "small talk." While the context can be set in advance in this way, the system may also analyze the transition of topics during the conversation based on the tendencies of words that appear in the conversation and automatically change the context. For example, the agent response generation unit 156 sets the context to be used to "explanation."
[0122] The agent nod generation unit 156 then modifies the nod action list L21 (Figure 3) to match the introverted agent, as shown in the nod action list L21-1 in Figure 6. In the nod action list L21-1 in Figure 6, the synchronization level threshold is changed so that the stages of the nod action are subdivided within a range where the synchronization level value is greater than 0.6 (see frame W11). For example, the agent nod generation unit 156 updates the highest synchronization level to 0.9 ≤ d < 1.0. In the nod action list L21-1, the upper limit of the number of nods and utterances for each stage of the synchronization level is the same as in the nod action list L21, but the agent nod generation unit 156 may change the upper limit of the number of nods and utterances for each stage of the synchronization level. The processing example 156 of the agent nod generation unit adjusts the depth of the nod to be deeper when the synchronization level is low, corresponding to an introverted user (see frame W13). The depth of the nod is determined by the angle of the user's neck when they nod.
[0123] Furthermore, if the user's personal characteristics are neutral, the agent response generation unit 156 selects an extroverted character for the humanoid agent. For example, the agent response generation unit 156 sets the context to be used to "small talk".
[0124] The agent nod generation unit 156 then modifies the nod action list L21 (Figure 3) to match the extroverted agent, as shown in the nod action list L21-2 in Figure 7. In the nod action list L21-2 in Figure 7, the threshold for the synchronization level is changed (see frame W21). The agent nod generation unit 156 adjusts the upper limit of the number of nods and utterances for the highest synchronization level according to the context (small talk) being used (see, for example, frame W22). The agent nod generation unit 156 also adjusts the depth of the nod to match a neutral user, so that it is of a moderate depth even when the synchronization level is low (see frame W23).
[0125] Furthermore, the user's personal characteristics are recorded in the user information DB144. The user's personal characteristics may be registered in advance when using the system, or they may be automatically estimated during the conversation and updated as needed.
[0126] [Processing Example 4 of the Agent Response Generation Unit] Figures 6 and 7 illustrate an example in which the agent response generation unit 156 updates each column of the response action list in association with the user's personal characteristics, the agent's character, and the context. However, the values of each column may also be calculated by combining these elements.
[0127] For example, the agent nod generation unit 156 adjusts the upper limit of the number of nods and utterances in the nod action list based on a number predetermined according to the context, increasing the number as the user's personal characteristics become more extroverted, and increasing the number as the humanoid agent's character becomes more extroverted.
[0128] The agent nod generation unit 156 then adjusts the upper limit of the number of nods and utterances in the nod action list, based on a number predetermined according to the context, so that the number decreases as the user's personal characteristics become more introverted, and the number decreases as the humanoid agent's character becomes more introverted.
[0129] For example, the agent response generation unit 156 describes a case where there is a discrepancy between the user's personal characteristics and the extroversion of the agent's character.
[0130] In this case, the agent nod generation unit 156 corrects the character of the humanoid agent so that the difference from the user's personal characteristics is within a certain range. The agent nod generation unit 156 adjusts the upper limit of the number of nods and utterances by the humanoid agent so that the number corresponds to the corrected character.
[0131] Furthermore, when the agent response generation unit 156 determines the values of each column in the response action list, it may set certain constraints and adjust the values of each column in the response action list under those conditions.
[0132] When the agent response generation unit 156 updates the list of response actions, it sets an arbitrary upper limit on the number of response actions and adjusts within that range.
[0133] For example, the agent nodding generation unit 156 may set an upper limit on the number of nodding actions based on existing knowledge (Reference 6). Reference 6 sets the upper limit on the number of nodding actions to five, based on human-to-human dialogue. Alternatively, the agent nodding generation unit 156 may automatically set the upper limit on the number of nodding actions based on changes in the user's response during the conversation. Reference 6: Kihara et al., "Examination of Nodding Parameters in Listener Robots," IEICE Technical Report, vol. 116, no. IMQ-494, pp. 63-68, 2017.
[0134] Furthermore, the agent nod generation unit 156 sets constraints on the relationships between each column in the nod action list and adjusts them to satisfy those conditions. For example, the agent nod generation unit 156 sets an inverse relationship between the number of nods and the depth of the nod, so that the more nods there are, the shallower the nod becomes.
[0135] [Processing Example 5 of the Agent Nod Generation Unit] Next, we will explain the process of updating the list of nod actions by applying the user's personal characteristics. The personal characteristics of users using the humanoid agent system are registered in the system in advance. Alternatively, the user's personal characteristics may be obtained in advance by methods such as questionnaire surveys and stored in the user information DB144.
[0136] The agent response generation unit 156 can update the list of response actions based on the individual characteristics of this user.
[0137] In addition, the agent response generation unit 156 may update the list of response actions using both the pre-registered personal characteristics of the user and the personal characteristics of the user estimated during the interaction with the humanoid agent. For example, let's consider a case where there is a discrepancy between the pre-registered personal characteristics of the user and the personal characteristics of the user estimated during the interaction with the humanoid agent. In this case, the agent response generation unit 156 will update the list of response actions to be closer to (stronger in weight of) the personal characteristics of the user assumed during the interaction, assuming that the estimated personal characteristics of the user are closer to the essence of the user's actual personal characteristics. Alternatively, the agent response generation unit 156 may calculate intermediate personal characteristics between the pre-registered personal characteristics of the user and the personal characteristics of the user estimated during the interaction with the humanoid agent, and update the list of response actions based on the calculated intermediate personal characteristics.
[0138] Furthermore, the agent nod generation unit 156 may pre-register the user's preferred nods as personal characteristics using a conventional method (Reference 7), and update the list of nods using the registered nods. Reference 7: Kawana, "The effect of listener nods in dialogue situations on interpersonal attractiveness," Experimental Social Psychology Research, vol. 26, no. 1, pp. 67-76, 1986.
[0139] [Processing Example 6 of the Agent Nod Generation Unit] Next, we will explain the process of updating the list of nod actions according to the user synchronization level.
[0140] The user synchronization level is a value calculated by the comprehensive evaluation unit 155 based on the nonverbal and verbal emotions of the user during their nodding actions while the humanoid agent is speaking. The user synchronization level indicates the degree to which the user is synchronizing with the humanoid agent's speech. The comprehensive evaluation unit 155 records the calculated user synchronization level as a pair with the nodding action the user was performing at that time.
[0141] Based on this pair, the agent nod generation unit 156 updates the list of nod actions so that when the synchronization level is the same as the user synchronization level, it performs the same nod action as the user.
[0142] Figure 8 shows an example of updating the list of nodding actions. The explanation uses the example of a user nodding action where the user synchronization level during humanoid agent speech is 0.9 or higher, the number of noddings is 4, the number of nodding utterances is 4, and the depth of the nodding is "shallow (1)". The depth of the nodding may also be expressed numerically, as shown in the numbers in parentheses in each column for nodding depth. A larger numerical value for nodding depth indicates a deeper nod.
[0143] The agent nod generation unit 156 updates the nod action list L21-1 to the nod action list L21-4 so that it performs the same nod actions as the user. Specifically, as shown in frame W31 of the nod action list L21-4, the agent nod generation unit 156 changes the number of nods corresponding to a synchronization level of 0.9 or higher, the same as the user synchronization level, from 3 to 4, changes the number of nod utterances from 2 to 4, and changes the depth of the nod to shallow.
[0144] In this way, the agent nod generation unit 156 modifies the list of nod actions so that when the humanoid agent is at the same synchronization level as the user, it performs the same nod actions as the user, thereby making the user feel a sense of affinity with the humanoid agent.
[0145] [Example 7 of agent response generation processing] Next, we will explain other processes that update the list of response actions according to the user synchronization level.
[0146] For example, let's consider a case where the user's personality traits (e.g., big5) and the humanoid agent's personality traits are registered in the memory unit 14 beforehand.
[0147] In this case, the agent nod generation unit 156 calculates the similarity α of the personality traits and corrects the nodding expression according to the level of agreement. The similarity α is, for example, the cosine similarity of the vectors when the personality traits are represented as a five-dimensional vector, based on the aforementioned "neuroticism (N)", "extroversion (E)", "openness to experience (O)", "agreeableness (A)", and "conscientiousness (C)" that constitute the personality traits.
[0148] As described above, the comprehensive evaluation unit 155 analyzes the nonverbal and verbal emotions of the user during their nodding actions while the humanoid agent is speaking, calculates the user synchronization level, and records the calculated user synchronization level and the nodding action the user was performing at that time as a pair.
[0149] Based on this pair, the agent nodding generation unit 156 updates the list of nodding actions to perform a nodding action that is a weighted average of the user's nodding action and the humanoid agent's nodding action when the synchronization level is the same as the user's synchronization level.
[0150] When calculating the weighted average of the nodding actions, the weights are determined using a similarity α, so that the lower the similarity in personality (the more different the personalities), the more strongly the user's individual characteristics are reflected. α is the similarity in personality characteristics between the user and the humanoid agent, and takes values between 0.0 and 1.0.
[0151] Figure 9 shows an example of updating the list of nodding actions. The explanation uses the example of a case where the user's nodding actions are 4 times, with 4 nodding utterances, and shallow nodding, when the user synchronization level during the humanoid agent's speech is 0.9 or higher.
[0152] The agent nod generation unit 156 changes the number of nods corresponding to a synchronization level of 0.9 or higher, the same as the user synchronization level, from 3 to (3α+(1-α)4), as shown in frame W41 of the nod action list L21-5. The agent nod generation unit 156 changes the number of nod utterances from 2 to (3α+(1-α)4). The agent nod generation unit 156 changes the depth of the nod from shallow to (α+(1-α)).
[0153] In this way, the agent response generation unit 156 reduces the burden of dialogue caused by differences in the personalities of the user and the humanoid agent by having the humanoid agent perform response actions that strongly reflect the user's personal characteristics, especially when the user and the humanoid agent have different personalities.
[0154] [Processing Example 8 of Agent Antenna Generation Unit] The agent antenna generation unit 156 may also control the antenna operation using a predetermined control formula.
[0155] For example, the agent response generation unit 156 uses formula (1), formula (6), or formula (7) to determine the synchronization level d calculated in the overall evaluation unit 155. a By applying this to the following equations (10) to (12), the number of nods u of the humanoid agent can be calculated. a Number of utterances S a , and the depth of the nod r a We seek.
[0156]
[0157]
[0158]
[0159] w ua This is the amount of change corresponding to the synchronization level. a0 This is the initial value for the number of nods. amax p is the maximum depth of the nod. a This is the ratio of nodding to speaking. ra This is a correction factor for adjusting the digits. Note that the subscript 'a' in each symbol represents a humanoid agent, and the subscript 'h' represents a user (person).
[0160] [Processing Example 9 of Agent Nod Generation Unit] The agent nod generation unit 156 processes the number of nods u of the humanoid agent. a Number of utterances S a , and the depth of the nod r a This can be corrected to make the nodding motion more similar to the user's.
[0161] d h =d a In this process, the agent nod generation unit 156 corrects the humanoid agent's nod motion so that it matches the user's nod motion. Specifically, the agent nod generation unit 156 reflects the coefficients of the equations representing the user's nod motion into the control equation for the humanoid agent's nod motion.
[0162] The agent nod generation unit 156 generates the number of nods u of the user. h Of the equation (13) that shows this, the tuning level d h The coefficient w uh And the initial value u for the number of nods h0 Using equation (14) which reflects the above, the number of nods u of the humanoid agent a To decide.
[0163]
[0164]
[0165] The agent response generation unit 156 generates responses based on the user's number of utterances S. h Of the equations (15) that show this, the ratio p of the number of nods to the number of utterances h Using equation (16), which reflects this, the number of utterances S of the humanoid agent a To decide.
[0166]
[0167]
[0168] The agent nod generation unit 156 generates nods based on the depth r of the user's nod. h Of the equations (17) that show this, the maximum value of the depth of nodding r hmax Using equation (18), which reflects this, the depth of the humanoid agent's nod r aTo decide.
[0169]
[0170]
[0171] The agent nodding generation unit 156 controls the nodding motion using a correction formula, thereby correcting the agent's nodding motion to resemble the user's nodding motion or to a nodding motion that corresponds to the user's response (level of understanding, response delay time).
[0172] [Processing Example 10 of Agent Nod Generation Unit] The agent nod generation unit 156 also takes into account the similarity of the personalities of the user and the humanoid agent, and calculates the number of nods the humanoid agent u a Number of utterances S a , and the depth of the nod r a You may correct this.
[0173] For example, the agent nod generation unit 156 weights the coefficients of each equation representing the user's nodding motion and the coefficients of each equation representing the humanoid agent's nodding motion using the similarity of personality traits α, thereby generating the number of nods u of the humanoid agent. a Number of utterances S a , and the depth of the nod r a This is corrected as shown in equations (19) to (21). The similarity α is, for example, the cosine similarity of the vectors when the personality traits are represented as a 5-dimensional vector based on the aforementioned personality traits: "neuroticism (N)", "extraversion (E)", "openness to experience (O)", "agreeableness (A)", and "conscientiousness (C)". The overall evaluation unit 155 determines the conformity level d a Formula (1), formula (6), or formula (7) shall be used as the calculation formula.
[0174]
[0175]
[0176]
[0177] In this way, the agent response generation unit 156 corrects the response action control formula based on the similarity of personality between the humanoid agent and the user, thereby enabling the humanoid agent to perform a response action that is in line with the user's personality. In this case, the user is more likely to feel a sense of affinity with the humanoid agent and can engage in smooth dialogue with the humanoid agent.
[0178] [Processing Example 11 of Agent Nod Generation Unit] Next, the agent nod generation unit 156 uses a correction function to determine the number of nods u of the humanoid agent. a The tuning level d may be corrected. The overall evaluation unit 155 determines the tuning level d a Formula (1), formula (6), or formula (7) shall be used as the calculation formula.
[0179] For example, the agent response generation unit 156 uses the correction function g as shown in equation (22). a Using (c'), the number of nods u of the humanoid agent. a This corrects the value. C' is the correction function g a This is the input for (c'). ga This is a weighting coefficient that adjusts the degree of influence of the correction function.
[0180]
[0181] Correction function g a (c') is equation (23).
[0182]
[0183] The correction function in equation (23) is such that c' is the score of the user's personality trait, extraversion, and e(c') is a function that converts the extraversion score to a value between -1.0 and 1.0, u am When the average number of nods by an agent is used, a bias corresponding to extraversion is added to the number of nods.
[0184] Correction function g a (c') could also be, for example, equation (24).
[0185] The correction function in equation (24) introduces random variation to the number of nods when c' is the standard deviation of the number of nods in past affirmative actions, and this value is less than or equal to a certain value (β). r(x) is a function that returns a random real number between -1.0 and 1.0.
[0186] The example shown is the case where a correction function is applied to the basic pattern equation (10), but similar corrections using the correction function are also possible for equation (14), which reflects the coefficient of the user's nodding action in the number of nods of the humanoid agent, and for equation (19), which takes into account the user's personality characteristics. Furthermore, the agent nodding generation unit 156 generates the number of utterances S of the humanoid agent. a , and the depth of the nod r a You may also apply a correction function to optimize this as well.
[0187] The agent nod generation unit 156 calculates the number of nods u of the humanoid agent after correction. a Number of utterances S a , and the depth of the nod r a By using this method, the agent's nodding actions can be corrected to resemble the user's nodding actions, or to reflect the user's response (level of understanding, response delay time).
[0188] [Processing Procedure for Presentation Method] Next, the processing procedure for the presentation method according to the embodiment will be described.
[0189] [Processing Procedure Example 1] Figure 10 is a diagram showing an example of the processing procedure for the presentation method according to the embodiment. In Figure 10, the case in which the comprehensive evaluation unit 155 flags nonverbal emotions and verbal emotions and comprehensively evaluates the conformity model based on the flags will be explained as an example.
[0190] The presentation device 10 acquires video input from the user using the humanoid agent (step S11). The presentation device 10 acquires voice input from the user (step S12).
[0191] The presentation device 10 estimates or calculates the user's nonverbal emotions from the video and spoken audio (step S13).
[0192] If the nonverbal emotion is positive (step S14: Yes), the presentation device 10 sets the nonverbal emotion flag to True (step S15).
[0193] If the nonverbal emotion is not positive (Step S14: No), or after processing in Step S15, the presentation device 10 converts the user's utterance into text using speech recognition (Step S16), and divides the text into words using morphological analysis (Step S17). The presentation device queries the user utterance word emotion dictionary 141 and calculates the verbal emotion (Step S18).
[0194] If the verbal emotion is positive (step S19: Yes), the presentation device 10 sets the verbal emotion flag to True (step S20).
[0195] If the verbal emotion is not positive (step S19: No), the presentation device 10 determines whether the user is currently speaking (step S21). If the user is currently speaking (step S21: Yes), the presentation device 10 returns to step S11 and continues setting flags for nonverbal and verbal emotions.
[0196] If the user is not speaking (Step S21: No), or after processing in Step S20, the presentation device 10 comprehensively evaluates the synchronization level, selects a body movement corresponding to the synchronization level from the agent body movement DB 142, and outputs a nodding action from the humanoid agent (Step S22). Based on the flags set for nonverbal and verbal emotions, the presentation device 10 evaluates the synchronization level by referring to, for example, the evaluation criteria list L1 (Figure 2) (for example, using processing example 1 of the comprehensive evaluation unit). The presentation device 10 also refers to the nodding action list L11 (Figure 4) and, based on the evaluated synchronization level, selects a nodding action to be expressed by the humanoid agent from the pre-registered agent body movement DB 142 and agent utterance content DB 143, and adjusts how many times that expression is performed in a certain period of time (for example, using processing example 1 of the comprehensive evaluation unit).
[0197] The flowchart shown in Figure 10 is an example of immediately nodding in agreement when positive linguistic emotion is detected, in order to strongly express agreement. Alternatively, steps S11 to S19 may be repeated multiple times, and the agreement determination may be made based on a majority vote of the results or the total number of positive judgments. In this case, in step S19, immediate agreement expressions may not be made until a certain number of times have been exceeded. Furthermore, as mentioned above, the presentation device 10 may correct the agreement level and the agreement movements of the humanoid agent, and use the correction results to control the agreement movements of the humanoid agent.
[0198] [Processing Procedure Example 2] Figure 11 is a diagram showing an example of the processing procedure of the presentation method according to the embodiment. In Figure 11, the presentation device 10 controls the nodding motion of the humanoid agent by considering only nonverbal emotions.
[0199] Steps S31 to S35 in Figure 11 correspond to steps S11 to S15 in Figure 10.
[0200] If the nonverbal emotion is not positive (step S34: No), or after processing in step S35, the presentation device 10 determines whether the user is speaking or not (step S36). If the user is speaking or not (step S36: Yes), the presentation device 10 returns to step S31 and continues setting the nonverbal emotion flag.
[0201] If the user is not speaking (Step S36: No), the presentation device 10 comprehensively evaluates the synchronization level using nonverbal emotions, selects a body movement corresponding to the synchronization level from the agent body movement DB 142, and outputs a nodding action from the humanoid agent (Step S37). Based on the flags set for nonverbal emotions, the presentation device 10 evaluates the synchronization level by referring to, for example, the evaluation criteria list L1 (Figure 2). The presentation device 10 also refers to, for example, the nodding action list L11 (Figure 4), and based on the evaluated synchronization level, selects a nodding action to be expressed by the humanoid agent from the pre-registered agent body movement DB 142 and agent utterance content DB 143, and adjusts how many times that expression is performed in a certain period of time.
[0202] As mentioned above, the presentation device 10 may correct the synchronization level and the nodding motion of the humanoid agent, and use the correction results to control the nodding motion of the humanoid agent.
[0203] [Processing Procedure Example 3] Figure 12 is a diagram showing an example of the processing procedure of the presentation method according to the embodiment. In Figure 12, the presentation device 10 controls the nodding motion of the humanoid agent by considering only linguistic emotions.
[0204] Step S41 in Figure 12 corresponds to step S12 in Figure 10. Steps S42 to S47 in Figure 12 correspond to steps S16 to S21 in Figure 10.
[0205] If the user is speaking (step S47: Yes), the presentation device 10 returns to step S41 and continues setting the language emotion flag.
[0206] If the user is not speaking (Step S47: No), the presentation device 10 comprehensively evaluates the synchronization level using linguistic emotion, selects a body movement corresponding to the synchronization level from the agent body movement DB 142, and outputs a nodding action from the humanoid agent (Step S48). Based on the flags set for linguistic emotion, the presentation device 10 evaluates the synchronization level by referring to, for example, the evaluation criteria list L1 (Figure 2). The presentation device 10 also refers to, for example, the nodding action list L11 (Figure 4), and based on the evaluated synchronization level, selects a nodding action to be expressed by the humanoid agent from the pre-registered agent body movement DB 142 and agent utterance content DB 143, and adjusts how many times that expression is performed in a certain period of time.
[0207] The flowchart shown in Figure 12 is an example of immediately nodding in agreement when positive linguistic emotion is detected, in order to strongly express agreement. Alternatively, steps S41 to S45 may be repeated multiple times, and the agreement determination may be made based on a majority vote of the results or the total number of positive judgments. In this case, in step S45, immediate agreement expressions may not be made until a certain number of times have been exceeded. Furthermore, as mentioned above, the presentation device 10 may correct the agreement level and the agreement movements of the humanoid agent, and use the correction results to control the agreement movements of the humanoid agent.
[0208] [Processing Procedure Example 4] Figure 13 is a diagram showing an example of the processing procedure for the presentation method according to the embodiment. In Figure 13, the case where nonverbal emotions and verbal emotions are continuous values is shown.
[0209] The presentation device 10 acquires video input from the user using the humanoid agent (step S51). The presentation device 10 acquires voice input from the user (step S52).
[0210] The presentation device 10 estimates or calculates the user's nonverbal emotion value from the video and spoken audio (step S53).
[0211] If the nonverbal emotion is above a threshold (step S54: Yes), the presentation device 10 sets the nonverbal emotion flag to True (step S55).
[0212] If the nonverbal emotion is below the threshold (Step S54: No), or after processing in Step S55, the presentation device 10 converts the user's utterance into text using speech recognition (Step S56), and divides the text into words using morphological analysis (Step S57). The presentation device queries the user utterance word emotion dictionary 141 and calculates the value of the verbal emotion (Step S58).
[0213] If the verbal emotion is above the threshold (step S59: Yes), the presentation device 10 sets the verbal emotion flag to True (step S60).
[0214] If the verbal emotion is below a threshold (step S59: No), the presentation device 10 determines whether or not the user is speaking (step S61). If the user is speaking (step S61: Yes), the presentation device 10 returns to step S51 and continues setting flags for nonverbal and verbal emotions.
[0215] If the user is not speaking (step S61: No), or after processing in step S60, the presentation device 10 evaluates the synchronization level, selects a body movement corresponding to the synchronization level from the agent body movement DB 142, and outputs a nodding motion of the humanoid agent (step S62).
[0216] The flowchart shown in Figure 13 is an example where, in order to strongly express agreement, an immediate nod is made when a verbal emotion above a threshold is detected. Alternatively, steps S51 to S59 may be repeated multiple times, and the nod determination may be made based on a majority vote of the results or the total number of judgments above the threshold. In this case, in step S59, an immediate nod expression may not be made until a certain number of judgments are exceeded. Furthermore, as described above, the presentation device 10 may correct the agreement level and the nod action of the humanoid agent, and use the correction result to control the nod action of the humanoid agent. Also, in processing procedure example 4, as in processing procedure examples 2 and 3, processing can be performed using only verbal emotion or nonverbal emotion. For example, when using only nonverbal emotion, the agreement level can be set by comparing the value of the nonverbal emotion with the threshold, and the nod action can be set using the set agreement level.
[0217] [Processing Procedure Example 5] Figure 14 is a diagram showing an example of the processing procedure for the presentation method according to the embodiment. In Figure 14, an example of calculating the synchronization level as a continuous value based on nonverbal emotions and verbal emotions is explained.
[0218] The presentation device 10 acquires video input from the user using the humanoid agent (step S71). The presentation device 10 acquires voice input from the user (step S72).
[0219] The presentation device 10 estimates or calculates the user's nonverbal emotion value from the video and spoken audio (step S73).
[0220] The presentation device 10 converts the user's utterance into text using speech recognition (step S74), and divides the text into words using morphological analysis (step S75). The presentation device then queries the user utterance word sentiment dictionary 141 and calculates the value of the language sentiment (step S76).
[0221] If the user is speaking (step S77: Yes), the presentation device 10 returns to step S71 and continues to acquire nonverbal emotion and verbal emotion values.
[0222] If the user is not speaking (step S77: No), the presentation device 10 calculates a continuous value of the synchronization level, for example by referring to the evaluation criteria list L2 (Figure 3) or by using formula (1) (for example, using processing example 2 or 3 of the overall evaluation unit). Furthermore, the presentation device 10 may correct the synchronization level using formulas (6), (7) to (9). The presentation device 10 selects a body movement corresponding to the synchronization level from the agent body movement DB 142, and sets and outputs a nodding movement for the humanoid agent, for example by referring to the nodding movement list L21 (step S78) (for example, using processing example 2 of the agent nodding generation unit).
[0223] [Processing Procedure Example 6] Figure 15 is a diagram showing an example of the processing procedure for the presentation method according to the embodiment. In Figure 15, an example is described in which the synchronization level is calculated as a continuous value based on nonverbal emotions and verbal emotions, and further correction is performed.
[0224] Steps S81 to S87 in Figure 15 are the same process as steps S71 to S77 in Figure 14.
[0225] If the user is speaking (step S87: Yes), the presentation device 10 returns to step S81 and continues to acquire nonverbal emotion and verbal emotion values.
[0226] If the user is not speaking (Step S87: No), the presentation device 10 calculates a continuous value of the synchronization level, for example by referring to the evaluation criteria list L2 (Figure 3) or using formula (1). Subsequently, for example, the presentation device 10 calculates a corrected value of the synchronization level based on the user's response, etc., using the method described in processing example 4 of the overall evaluation unit 155 (Step S88). Furthermore, the presentation device 10 may correct the synchronization level using formulas (6), (7) to (9).
[0227] The presentation device 10 selects a body movement corresponding to the corrected synchronization level from the agent body movement DB 142, and for example, refers to the nodding movement list L21 to set and output the nodding movement for the humanoid agent (step S89).
[0228] As mentioned above, the presentation device 10 may correct the nodding motion of the humanoid agent and control the nodding motion of the humanoid agent using the correction result (for example, using processing examples 3, 4, 5, 6, and 7 of the agent nodding generation unit).
[0229] [Processing Procedure Example 7] Figure 16 is a diagram showing an example of the processing procedure of the presentation method according to the embodiment. In Figure 16, an example in which the synchronization level and the acknowledgment action are controlled by continuous values will be explained.
[0230] Steps S91 to S96 in Figure 16 are the same process as steps S71 to S76 in Figure 14.
[0231] The presentation device 10 acquires user information (step S97).
[0232] If the user is speaking (step S98: Yes), the presentation device 10 returns to step S91 and continues to acquire nonverbal emotion and verbal emotion values.
[0233] If the user is not speaking (step S98: No), the presentation device 10 calculates the synchronization level using equation (1). Furthermore, the presentation device 10 may correct the synchronization level using equations (6), (7) to (9).
[0234] The presentation device 10 selects a body movement corresponding to the synchronization level from the agent body movement DB 142, and sets and outputs a nodding motion for the humanoid agent using, for example, formulas (10) to (12) (step S99). The presentation device 10 may also correct the nodding motion of the humanoid agent by appropriately selecting formulas (14), (16), (18), or formulas (19) to (21), or formulas (22) to (24). Furthermore, in processing procedure examples 5, 6, and 7, the presentation device 10 may perform processing using only either verbal emotions or nonverbal emotions, similar to processing procedure examples 2 and 3.
[0235] [Effects of the Embodiment] As described above, the presentation device 10 according to the embodiment performs a comprehensive evaluation based on nonverbal emotions, which are emotional information that appears in nonverbal information, or on verbal and nonverbal emotions that appear in the text of the speech recognition result, and determines the synchronization level. Based on the synchronization level, the presentation device 10 controls the nodding actions of the humanoid agent that responds to the user.
[0236] In this way, the presentation device 10 utilizes non-verbal information such as video and audio to cause the humanoid agent to perform nodding actions that correspond to the user's personality traits, emotions, intentions at the time of speaking, and the context of the conversation, as derived from the non-verbal information. Furthermore, the presentation device 10 analyzes the user's personality traits, intentions, and context from video, audio, and pre-registered user information, and uses this in combination with linguistic information to improve the accuracy of the humanoid agent's interpretation of the user's statements, enabling the humanoid agent to express more appropriate response actions.
[0237] For example, the presentation device 10 can be applied to agent services that provide information while eliciting user information through communication with presenters, counselors, sales representatives, etc. In this case, the presentation device 10 analyzes the user's facial expressions, body movements, and tone of voice from video and audio, and uses this in combination with linguistic information to perform highly accurate negative / positive judgments of what is said, enabling the expression of more appropriate response actions.
[0238] Furthermore, the presentation device 10 takes into account the user's personality traits and can present more favorable nodding actions to the user. It also controls the humanoid agent's nodding actions by taking into account the agent's personality traits, the context of the conversation, and the user's level of understanding. As a result, the presentation device 10 can perform nodding actions that are appropriate for the user, extending the duration of the conversation with the user and enabling communication that elicits more information from the user.
[0239] The presentation device 10 can be used as a method to determine the response that expresses the humanoid agent's own intention when the humanoid agent is operating in real time and autonomously giving a response that is either in agreement or negative.
[0240] The scope of application of the presentation device 10 is not limited to the real-time operation of the humanoid agent, but can also be used to create responses from the humanoid agent to pre-recorded user video and audio.
[0241] Furthermore, in application scenarios other than real-time, the provider of the humanoid agent service can edit the operation of the humanoid agent using the presentation device 10 and modify the service.
[0242] During this editing process, it is possible to simulate the actions of a humanoid agent, and as part of the simulation results, nonverbal emotions, verbal emotions, and overall evaluation results (synchronization level) calculated using the presentation device 10 can be provided to the editor.
[0243] Furthermore, the editor may actively edit the nonverbal emotion, verbal emotion, and overall evaluation values provided by the presentation device 10 and apply them to the humanoid agent to create the humanoid agent's responses.
[0244] Furthermore, the presentation device 10 may pre-collect and record in the user information DB 144 the user's expression tendencies (ratio of nonverbal expression to verbal expression), user information related to nodding (number of nods, number of nodding utterances, neck angle), and user personality traits, which are used to adjust the humanoid agent's nodding behavior. In addition, the presentation device 10 may collect the user's expression tendencies, user information related to nodding behavior, and user personality traits in real time during the dialogue between the humanoid agent and the user. Moreover, the presentation device 10 may record information not only during conversations between the user and the humanoid agent, but also during conversations between people, to obtain the user's expression tendencies, user information related to nodding behavior, and user personality traits.
[0245] [Regarding the System Configuration of the Embodiment] The presentation device 10 is a functional concept and does not necessarily need to be physically configured as shown in the figure. In other words, the specific forms of distribution and integration of the functions of the presentation device 10 are not limited to those shown in the figure, and all or part of it can be configured by functionally or physically distributing or integrating in any unit according to various loads and usage conditions.
[0246] Furthermore, each process performed in the presentation device 10 may be implemented, in whole or in part, by a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and a program that is analyzed and executed by the CPU and GPU. Alternatively, each process performed in the presentation device 10 may be implemented as hardware using wired logic.
[0247] Furthermore, among the processes described in the embodiments, all or part of the processes described as being performed automatically can be performed manually. Alternatively, all or part of the processes described as being performed manually can be performed automatically by known methods. In addition, the processing procedures, control procedures, specific names, and information including various data and parameters described above and illustrated may be changed as appropriate unless otherwise specified.
[0248] [Program] Figure 17 shows an example of a computer in which the presentation device 10 is realized when a program is executed. The computer 1000 has, for example, a memory 1010 and a CPU 1020. The computer 1000 also has a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.
[0249] Memory 1010 includes ROM 1011 and RAM 1012. ROM 1011 stores, for example, a boot program such as a BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to the hard disk drive 1090. The disk drive interface 1040 is connected to the disk drive 1100. For example, a removable storage medium such as a magnetic disk or optical disk is inserted into the disk drive 1100. The serial port interface 1050 is connected to, for example, a mouse 1110 and a keyboard 1120. The video adapter 1060 is connected to, for example, a display 1130.
[0250] The hard disk drive 1090 stores, for example, an OS (Operating System) 1091, an application program 1092, a program module 1093, and program data 1094. That is, the program that defines each process of the presentation device 10 is implemented as a program module 1093 in which code executable by the computer 1000 is written. The program module 1093 is stored, for example, in the hard disk drive 1090. For example, a program module 1093 for performing processes similar to the functional configuration of the presentation device 10 is stored in the hard disk drive 1090. Note that the hard disk drive 1090 may be replaced by an SSD (Solid State Drive).
[0251] Furthermore, the configuration data used in the processing of the above-described embodiment is stored as program data 1094 in, for example, memory 1010 or hard disk drive 1090. The CPU 1020 then reads the program module 1093 and program data 1094 stored in memory 1010 or hard disk drive 1090 into RAM 1012 as needed and executes them.
[0252] Furthermore, the program module 1093 and program data 1094 are not limited to being stored in the hard disk drive 1090; for example, they may be stored in a removable storage medium and read by the CPU 1020 via a disk drive 1100 or the like. Alternatively, the program module 1093 and program data 1094 may be stored in another computer connected via a network (LAN (Local Area Network), WAN (Wide Area Network), etc.). The program module 1093 and program data 1094 may then be read by the CPU 1020 from the other computer via a network interface 1070.
[0253] Although embodiments applying the invention made by the present inventors have been described above, the present invention is not limited by the descriptions and drawings that constitute part of the disclosure of the present invention in these embodiments. That is, all other embodiments, examples, and operational techniques made by those skilled in the art based on these embodiments are included in the scope of the present invention.
[0254] 10, 110 Presentation device 11, 111 Input unit 12, 112 Output unit 13, 113 Communication control unit 14, 114 Storage unit 15, 115 Control unit 141, 1141 User utterance word emotion dictionary 142 Agent body movement DB 143, 1143 Agent utterance content DB 144 User information DB 145 Agent character DB 146 Communication history DB 151 User video analysis unit 152, 1151 User speech recognition unit 153 Nonverbal emotion calculation unit 154 Verbal emotion calculation unit 155 Overall evaluation unit 156, 1153 Agent interjection generation unit 157, 1154 Presentation unit 1142 Body movement DB 1152 Emotion polarity analysis unit
Claims
1. A presentation device comprising: an evaluation unit that evaluates a synchronization level, which is the degree of synchronization with the user expressed by the nodding movements of a humanoid agent responding to the user, based on nonverbal emotions, which are the user's emotional information that appears in the user's nonverbal information; and a control unit that controls the nodding movements of the humanoid agent responding to the user, based on at least the synchronization level.
2. The presentation device according to claim 1, characterized in that the evaluation unit evaluates the synchronization level based on the nonverbal emotion and the verbal emotion, which is emotional information appearing in the text of the speech recognition result for the user's utterance.
3. The presentation device according to claim 2, characterized in that the control unit controls the nodding motion of the humanoid agent using at least one of the ratio of the amount of nonverbal emotion to the amount of verbal emotion of the user, the number of times the user nods in response to the nods of the humanoid agent, the number of times the user speaks in response to the nods, the angle of the user's neck when nodding, and the user's personality traits.
4. The presentation device according to claim 1, characterized in that the control unit controls the nodding motion of the humanoid agent based on the character of the humanoid agent and / or the type of context used by the humanoid agent.
5. The presentation device according to claim 1, characterized in that the evaluation unit corrects the synchronization level in accordance with the user's response to the humanoid agent's nodding motion.
6. The presentation device according to claim 1, characterized in that the control unit corrects the control of the humanoid agent's nodding motion in response to the user based on the user's personality traits, the degree of similarity between the user's personality traits and the humanoid agent's personality traits, and the user's response to the humanoid agent's nodding motion.
7. A presentation method performed by a presentation device, comprising at least the steps of: evaluating a synchronization level, which is the degree of synchronization with the user expressed by the nodding movements of a humanoid agent responding to the user, based on nonverbal emotions, which are emotional information of the user that appears in the user's nonverbal information; and controlling the nodding movements of the humanoid agent responding to the user based at least the synchronization level.
8. A presentation program for causing a computer to perform the following steps: evaluate the degree of synchronization with the user, expressed by the nodding movements of a humanoid agent responding to the user, based on nonverbal emotions, which are the user's emotional information that appears in the user's nonverbal information; and control the nodding movements of the humanoid agent responding to the user, based on at least the synchronization level.