Method and computer program for promoting relationship building between people
A computer-controlled robot facilitates relationship building by encouraging one-on-one conversations and empathy between individuals meeting for the first time, addressing the challenge of spontaneous conversation initiation and psychological barriers through shared perspectives and self-disclosure.
Patent Information
- Application Number
- JP2022027523
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-02-25
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2042-02-25
AI Technical Summary
Existing technologies face challenges in promoting relationship building between people meeting for the first time due to psychological barriers and the difficulty in initiating spontaneous conversations, requiring advanced conversational abilities from robots that are difficult to realize.
A method involving a computer-controlled robot that encourages one-on-one conversations with individuals, where the robot asks questions, waits for responses, and facilitates understanding between participants by having the second person summarize the first person's opinions, thereby creating an empathetic connection without direct human interaction.
This approach enhances relationship building by fostering empathy and understanding between individuals, even when they do not directly converse, by using a robot to facilitate shared perspectives and self-disclosure, leading to increased interaction satisfaction and familiarity.
Smart Images

Figure 0007814734000001 
Figure 0007814734000002 
Figure 0007814734000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a method for promoting relationship building between people, and more particularly to a method for promoting relationship building between people who meet for the first time using a robot, and a computer program therefor. [Background technology]
[0002] When people meet for the first time, there are invisible barriers that prevent them from starting a conversation. As the English expression "break the ice" suggests, these barriers transcend the boundaries of language, race, and nationality. These barriers make it difficult for people who meet for the first time to start a conversation and build relationships with each other.
[0003] Building relationships with others is important for people's sense of fulfillment and happiness. The desire for good relationships with others (the need to belong) is considered to be adaptively advantageous. Isolation and rejection from others can cause anxiety. In other words, building relationships with others is important for people to live fulfilling lives.
[0004] On the other hand, with robots, the above-mentioned barriers do not exist even when meeting someone for the first time. It is also thought that humans will not feel the same barriers when meeting a robot as they would when meeting a human, even when meeting a robot for the first time.
[0005] Research is being conducted to utilize these characteristics of robots to promote the building of relationships between people. In this case, the smallest unit of interaction is a group of three people, consisting of one robot and two humans.
[0006] Three possible dialogue patterns for these three parties are shown in Figure 1: Pattern 1 50, Pattern 2 52, and Pattern 3 54. Pattern 1 50 is a situation in which only robot R and the first person H1 converse. Pattern 2 52 is a situation in which robot R and the first person H1 converse, and robot R and the second person H2 converse, respectively. Pattern 3 54 is a situation in which robot R and the first person H1 and the second person H2 converse, respectively. Moving from left to right in Figure 1 , the effect of promoting the relationship between the first person H1 and the second person H2 increases, and the conversational ability required of the robot also increases. Research into having a robot converse in a situation like Pattern 3 54 is disclosed in Non-Patent Document 1, listed below. Non-Patent Document 1 defines four states as a communication space using an agent that mediates between people meeting for the first time. The four states are a greeting state, a topic provision state, a topic exploration state, and a conversation prompting state. Note that the agent disclosed in Document 1 is a type of robot.
[0007] In the greeting state, the agent separately guides the two people meeting for the first time to speak. In the topic provision state, the agent exchanges information about the two people meeting for the first time and encourages them to express their opinions on a common topic. The agent then guides the two people meeting for the first time to face each other and converse. In the topic exploration state, the agent asks questions to the two people meeting for the first time and enters into their communication. In the conversation prompting state, the agent aims to synchronize with the conversation of the two people meeting for the first time and liven up the conversation. When the conversation of the first time people ends, the state transitions to the topic provision state. [Prior art documents] [Non-patent literature]
[0008] [Non-Patent Document 1] Ryohei Sasama and 4 others, "A First-Time Introduction Agent that Controls Speech Based on Communication Activity Level," Research Report on Human-Computer Interaction, Information Processing Society of Japan, May 8, 2009, pp. 1-8 Summary of the Invention [Problem to be solved by the invention]
[0009] The technology disclosed in Reference 1 aims to achieve a state corresponding to the third pattern 54 in Figure 1. However, in order to achieve this third pattern 54, the robot needs to have a high level of conversational ability. Furthermore, when humans meet for the first time, it is difficult for them to spontaneously initiate a conversation due to the psychological barrier of not knowing each other and the lack of opportunities to initiate a conversation. Although Reference 1 discloses how an agent should behave in each of the above four states, the functions that an agent must achieve in the situations disclosed in Reference 1 in order for humans to spontaneously initiate a conversation are highly advanced, and it is difficult to realize an agent (robot) with such functions.
[0010] Therefore, an object of the present invention is to provide a method that allows a robot to easily promote relationship building between people. [Means for solving the problem]
[0011] A method for promoting relationship building between persons according to a first aspect of the present invention includes a step in which a computer controls a robot to make an utterance to a first person to encourage the first person to express an opinion on a topic determined by a predetermined method; a first waiting step in which the computer waits until the first person expresses an opinion; and a step in which the computer controls the robot to make an utterance to a second person different from the first person to request an explanation for the opinion expressed by the first person.
[0012] Preferably, the method further includes the step of the computer identifying the topic according to a predetermined scenario.
[0013] More preferably, the method further includes a step in which, following the first waiting step, the computer determines whether or not a word of a predetermined part of speech is present in the first person's utterance; a step in which the computer determines what kind of reflection processing to perform according to whether or not a word of the predetermined part of speech is present in the first person's utterance; and a step in which the computer controls the robot to perform the reflection processing determined in the determining step.
[0014] More preferably, the determining step includes a step in which the computer determines whether the robot will execute a process of uttering a reflective sentence including the word or a process of performing a predetermined reflective behavior, depending on whether it is determined that a word of a predetermined part of speech is present in the utterance of the first person.
[0015] Preferably, the predefined reflection behavior includes a behavior in which the robot omits asking the first person to reflect.
[0016] More preferably, the predefined reflection behavior includes the robot uttering a predefined reflection sentence to the first person.
[0017] More preferably, the method further includes a second waiting step in which the computer waits until the second person provides an explanation for the opinion, and a step in which the computer, following the second waiting step, controls the robot to perform an action indicating that it understands the opinion of the first person.
[0018] Preferably, this method further includes the steps of: a computer controlling the robot to ask the second person for their opinion on a new topic determined by a predetermined method; a third waiting step in which the computer waits until the second person gives their opinion; and a computer controlling the robot to make an utterance to the first person to ask for an explanation regarding the opinion given by the second person.
[0019] Preferably, the method further includes a fourth waiting step of waiting until the first person provides an explanation for the second person's opinion, and a step following the fourth waiting step, in which the computer controls the robot to perform an action indicating that it has understood the second person's opinion.
[0020] More preferably, the method further includes a step of determining a direction of the first person relative to the robot prior to the step of making an utterance to solicit an opinion, and the step of making an utterance to solicit an opinion includes a step of the computer controlling the robot to turn toward the first person, and a step of the computer controlling the robot to utter a question to solicit an opinion on the topic.
[0021] Preferably, this method further includes a step in which, prior to the step of requesting an explanation regarding the opinion, the computer determines the direction of the second person relative to the robot, and the step of making an utterance to request an explanation regarding the opinion includes a step in which the computer controls the robot so that the robot faces the second person, and a step in which the computer controls the robot so that an utterance is made to request an explanation regarding the opinion stated by the first person.
[0022] A computer program according to a second aspect of the present invention causes a computer to execute the steps of any of the above-described methods.
[0023] The above and other objects, features, aspects and advantages of the present invention will become apparent from the following detailed description of the invention taken in conjunction with the accompanying drawings. [Brief explanation of the drawings]
[0024] [Figure 1] FIG. 1 is a diagram showing a situation of a three-way dialogue including a robot. [Figure 2] FIG. 2 is a diagram showing a situation of a three-way dialogue including a robot, which is realized in the first embodiment of the present invention. [Figure 3]FIG. 3 is a schematic diagram for explaining a method for promoting relationship building between people meeting for the first time, which is realized by the robot in the first embodiment of the present invention. [Figure 4] FIG. 4 is a flowchart illustrating the steps of a method for encouraging relationship building between people who meet for the first time by controlling a robot using a computer in the first embodiment of the present invention. [Figure 5] FIG. 5 is a flowchart illustrating the steps of a specific example of a method for encouraging relationship building between people meeting for the first time by controlling a robot by a computer in the first embodiment of the present invention. [Figure 6] FIG. 6 is a flowchart illustrating the steps of a more specific example of a method for encouraging relationship building between people meeting for the first time by controlling a robot by a computer in the first embodiment of the present invention. [Figure 7] FIG. 7 is a schematic block diagram showing the hardware configuration of a conversational robot system 350 that realizes the first embodiment of the present invention. [Figure 8] FIG. 8 is a schematic block diagram showing the functional relationships between the functions of a conversational robot system 350 that realizes the first embodiment of the present invention. [Figure 9] FIG. 9 is a diagram showing, in a graphical format, a portion of an example of the configuration of a computer program for realizing the conversational robot system 350 according to the first embodiment of the present invention. [Figure 10] FIG. 10 is a block diagram showing the relationship between the functions and the data flow between them in the dialogue control performed by the conversational robot system 350 according to the first embodiment of the present invention. [Figure 11] FIG. 11 is a flowchart showing the processing flow of a program for realizing local positive / negative judgment recognition to determine whether a response to a question posed by the conversational robot system 350 according to the first embodiment of the present invention is a positive or negative answer. [Figure 12]FIG. 12 is a flowchart showing the flow of processing according to a program for the conversational robot system 350 according to the first embodiment of the present invention to analyze a question posed to the other party and ask for a repetition. [Figure 13] FIG. 13 is a flowchart showing the flow of processing by a program for realizing the first utterance sentence analysis shown in FIG. [Figure 14] FIG. 14 is a flowchart showing the flow of processing by a program for realizing the second utterance sentence analysis shown in FIG. [Figure 15] FIG. 15 is a flowchart showing the flow of processing by the program shown in FIGS. 13 and 14 for generating a question sentence utterance using a reflection word. [Figure 16] FIG. 16 is a diagram schematically showing the progress of dialogue in the embodiment and the comparative example in an experiment for confirming the effects of the first embodiment of the present invention. [Figure 17] FIG. 17 is a graph showing the results of the subjects' impression evaluation of the three-way dialogue in the experiment. [Figure 18] FIG. 18 is a diagram showing the appearance of the operation control PC (Personal Computer) shown in FIG. [Figure 19] FIG. 19 is a block diagram showing the hardware configuration of the operation control PC shown in FIG. DETAILED DESCRIPTION OF THE INVENTION
[0025] In the following description and drawings, the same parts are designated by the same reference numerals, and therefore detailed descriptions thereof will not be repeated.
[0026] 1. Problem and its solution The problem with three-way dialogue involving people meeting for the first time is that the people do not actually converse with each other. This makes it difficult to promote relationship building between the people. Therefore, this embodiment employs a dialogue strategy that promotes relationship building between people through the intervention of a robot, even if the people are not actually conversing with each other.
[0027] It is said that when a robot is equal in terms of human preferences, it creates a better impression between humans on first meeting than when it is unequal. On the other hand, in such a dialogue, aligning each other's values regarding human preferences is thought to increase predictability between humans and lead to the building of mutual trust. In other words, in order to build relationships between humans, it is important to design dialogue after confirming each other's preferences. If this is the case, a dialogue strategy that promotes the building of human relationships in three-way dialogue involving a robot is needed, without relying on the agreement or disagreement of preferences between humans.
[0028] Furthermore, if the purpose of a robot is to promote relationship building between humans, it is not necessarily necessary for the robot to make utterances about itself (e.g., preferences regarding the same matters as humans). On the other hand, if information about the robot is not disclosed, it may not arouse interest, concern, or a desire to interact with the robot. The act of disclosing personal information to others (self-disclosure) is important for building intimacy between humans. It has also been reported that self-disclosure by a robot in a three-way dialogue increases the interaction satisfaction and familiarity of the humans participating in the three-way dialogue with the system. Therefore, this embodiment employs a dialogue strategy in which the robot has its own preferences and expresses them to humans to promote relationship building between humans. Therefore, this embodiment employs a method for realizing a three-way dialogue by repeatedly having one-on-one conversations with the robot, rather than having humans interact directly with each other. In this case, the robot discloses its preferences to the other person.
[0029] 2, in this embodiment, a second pattern 60 in which a first person H1 and a second person H2 appear to be substantially conversing with each other is realized in place of the second pattern 52 shown in FIG. 1, while using a method similar to that of the second pattern 52 shown in FIG. 1. In this method, the robot R primarily converses only with the first person H1 or the second person H2. Therefore, compared to the method disclosed in Non-Patent Document 1 (third pattern 54), the conversational ability required of the robot in this embodiment is not high.
[0030] In order to build a relationship between people using the method adopted in this second pattern 60, it is necessary for the first person H1 and the second person H2 to empathize with each other. To achieve this, it is considered effective for the first person H1 and the second person H2 to share their perspectives. Therefore, as will be explained below, this embodiment adopts a method in which the robot R asks the second person H2 to speak on its behalf (explain or summarize) the opinion of the first person H1 in response to a question posed by the robot R to the first person H1.
[0031] 3, in a first step 100, a dialogue 120 is carried out in which the robot 110 asks a question to a first person 112, and the first person 112 expresses an opinion 130 in response to the question. At this time, the second person 114 does not participate in the dialogue 120. Next, in a second step 102, the robot 110 pretends not to understand the opinion of the first person 112, and asks the second person 114 a question 132 requesting that the second person 114 speak on behalf of the first person 112 (explanation, summary). In response to this question 132, the second person 114 gives a reply 136 to the robot 110 speaking on behalf of the first person 112, based on thoughts 134 from the perspective of the first person 112. In other words, the robot 110 and the second person 114 carry out a dialogue 122. The first person 112 does not participate in this dialogue 122. However, upon hearing the answer 136 from the second person 114, the first person 112 gets the impression 138 that the second person 114 is thinking from the first person's 112 perspective. Thus, even though the first person 112 and the second person 114 do not directly converse, they experience it as if they were conversing with each other. By repeating this experience, it is expected that the first person 112 and the second person 114 will understand and empathize with each other and begin a natural conversation.
[0032] As will be described later, experiments have confirmed that adopting such a method has the effect of strengthening the relationship between people who meet for the first time.
[0033] Scenario 2 A specific embodiment for realizing the above-mentioned method will be described below. Figure 4 shows the flow of conversation according to the above-mentioned scenario. The robot converses with two people roughly according to this scenario. This scenario is prepared as a script that controls the robot's operation.
[0034] Referring to FIG. 4, this conversation scenario includes step 200 in which the robot makes a greeting, and step 202 in which the three parties, including the robot, make utterances to introduce themselves to each other.
[0035] This scenario further includes step 204, in which the robot asks an initial question to each participant in the three-way dialogue. In step 204, the robot asks the participants about their preferences for topics. In this embodiment, although not described in detail here, the pattern of questions to be asked to the participants later will differ depending on the combination of the participants' preferences for topics. Therefore, it is necessary to understand and memorize the answers of the participants in step 204 at this stage.
[0036] This scenario further includes step 206 in which the robot discloses its preferences for the topic of conversation. As described above, such disclosure is intended to improve the interaction satisfaction and familiarity of the humans participating in the three-way interaction with the system by the robot disclosing itself.
[0037] The scenario further includes step 208 in which the robot asks one of the participants (person 1) a second question, which is related to the question posed in step 204, except that the question posed in step 208 is different depending on the answer given by person 1 in step 204.
[0038] This scenario further includes step 210, in which a first dialogue is developed based on the first person's answer to the question in step 208; step 212, in which a second dialogue is developed based on the result of step 210; and finally step 214, in which a closing greeting is given to end the three-way dialogue. In step 210, the robot, for example, asks the first person about the reason for their answer and behaves as if it does not understand the reason. In step 212, the robot turns to the second person and asks them to speak on their behalf (explain, summarize) the reason given by the first person. In other words, this scenario aims to achieve the effect that the first person and the second person are conversing with each other, even though the robot is conversing with the first person and the second person separately in steps 208, 210, and 212.
[0039] Figure 5 shows in flowchart form the specific flow of the robot's operation according to the scenario in Figure 4. In reality, the robot's operation is realized by accumulating real-time operations at each time point of a program that is repeated at very short fixed intervals.
[0040] First, in step 230, the robot faces the first person. This first person's direction is the direction facing the center of the first person's face, with the robot as the base point. To make the robot perform this action, the robot's control circuit needs to know the robot's position and posture, as well as the position of the first person (and the position of the second person), and change the robot's posture to the desired state. The configuration for obtaining this information and controlling the robot, as well as the configuration required to achieve each of the steps described below, will be described later.
[0041] In the next step 232, the robot asks the first person a first question and waits for the first person to respond. In this case, the control circuit of the robot needs to include a natural language processing circuit that generates the text of the question to be uttered, and a speech circuit that converts the text into a voice signal and then utters it as a voice.
[0042] In step 234, the robot listens to the first person's answer to the first question and asks the reason for it. At this point, the robot needs a mechanism for converting speech into a voice signal, performing speech recognition and converting it into text, a natural language processing mechanism for extracting the answer to the question from the text, and a natural language processing and speech processing mechanism for creating and speaking the reason for the answer based on the answer.
[0043] In many cases, the first question is expected to be answered with a YES / NO. For such a question, the first person's answer can be determined by, for example, extracting words expressing affirmation or negation from the answer and interpreting them based on the original question. However, in some cases, the same affirmative answer may not be in a simple affirmative form, or the meaning of the answer to the question may be reversed between affirmation and negation. To prepare for such cases, this embodiment employs local affirmation / negation recognition, as described below. The question posed by the robot in step 234 is generated by inserting predetermined keywords into a prepared speech template according to the first person's answer, according to a predetermined scenario. Many of the questions posed by the robot below are generated in a similar manner.
[0044] After the robot asks the first person the reason in step 234, the robot waits until the first person gives an answer.
[0045] In the next step 236, the robot listens to the first person's answer to the question asked in step 234, pretends not to understand, asks more probing questions about the reasons for the first answer, and waits until the question is answered. However, this step is not essential. The robot may proceed to the next step 238 without performing step 236. In this case, the robot's behavior will always be in accordance with the pre-programmed scenario.
[0046] In a further subsequent step 238, the robot waits for the first person's answer to the probing question posed in step 236 (or the question posed in step 234). If an answer is given, in step 240 the robot pretends not to understand the first person's answer.
[0047] Furthermore, in step 242, the robot turns toward the second person. Furthermore, in step 246, the robot asks the second person whether he or she understands what the first person is saying. This is an important point in this embodiment, because the second person imagines what the first person was thinking in response to the question from the robot and thinks of an answer that represents the first person's thoughts from the first person's perspective. After asking the question, the robot waits until the second person answers.
[0048] In the next step 248, the robot listens to the answer of the second person, behaves as if it is convinced by the answer of the first person, and ends the process. By listening to the answer of the second person and behaving as if it is convinced by the answer of the first person in this way, empathy is created between the first person and the second person, and an effect can be obtained as if there had been a conversation between them.
[0049] By repeating the above process by switching between the first person and the second person, an even greater effect can be obtained.
[0050] Figure 6 shows the scenario used in the experiment described below. The robot greets the participants in step 300, and then introduces the three participants in step 302. In step 304, the robot asks the first and second participants whether they like alcohol, obtains their answers, and stores the answers. Subsequent processing varies depending on the answers of the first and second participants. Therefore, an appropriate scenario is selected in accordance with the answers of both participants. For example, the content of subsequent questions will vary depending on whether the answers of the first and second participants regarding alcohol are like / like, like / dislike (dislike / like), or dislike / dislike.
[0051] Furthermore, in step 306, the robot self-discloses that it does not like drinking parties. This is merely a scenario in the experiment, and it is equally possible for the robot to self-disclose that it likes drinking parties. Drinking parties were chosen as a topic in the experiment because people tend to have different likes and dislikes for them. In the experiment, it is necessary for the dialogue partner to speak on behalf of the opinions of others, so it is desirable for the content of the topic to clearly reveal the opinions of the participants in the dialogue.
[0052] In the following step 310, the robot attempts a dialogue regarding complaints about the drinking party, and in step 312, it attempts a dialogue regarding the conflict over whether or not to attend the drinking party. Finally, in step 314, the robot makes a closing farewell.
[0053] According to the above-described scenario, the above-described probing questions and proxy speaking requests can be made in any of steps 308, 310, and 312. In the experiment, probing questions and proxy speaking requests were made as follows, based on the combination of responses from the dialogue participants in step 304.
[0054] Like / Like → Steps 308 and 312 Like / Dislike → Steps 308 and 312 Dislike / Dislike → Steps 310 and 312 Scenarios are prepared in advance for the above three cases, and the appropriate scenario can be selected based on the responses from the first person and the second person. In the following explanation, for ease of understanding, we will assume that the preferences of the first person and the second person are like / like.
[0055] Third configuration 7 is a block diagram showing the hardware configuration of a conversational robot system 350 that executes a method according to this embodiment. Referring to FIG. 7, the conversational robot system 350 includes a robot 360 and a motion control PC 362 that controls the actuators of the robot 360 in accordance with a program that realizes a pre-prepared scenario, thereby causing the robot 360 to operate. The motion control PC 362 is connected to a network 364.
[0056] The conversational robot system 350 further includes a speech recognition PC 368 that can communicate with the operation control PC 362 via a network 364 and performs speech recognition on the speech of a person who is the conversation partner, converts it into a string of characters, and transmits it to the operation control PC 362 in text format, and a microphone 366 that is connected to the speech recognition PC 368 and converts the speech into a voice signal and inputs it to the speech recognition PC 368. The microphone 366 is installed near the robot 360. In reality, multiple microphones 366 are provided to form a microphone array, and the operation control PC 362 can distinguish between the voices of the first person 112 and the second person 114 based on the voice signals from these microphones 366.
[0057] The conversational robot system 350 further includes a speech synthesis PC 372 connected to a network 364, which performs speech synthesis based on spoken text received from the operation control PC 362 via the network 364 and outputs a speech signal, and a speaker 370 which converts the speech signal output from the speech synthesis PC 372 into speech. The speaker 370 is ideally provided in the head of the robot 360, but may also be placed near the robot 360.
[0058] The conversational robot system 350 further includes a person position sensor 374 for detecting the positions of the first person 112 and the second person 114, and a person position recognition PC 376 for detecting the positions of the first person 112 and the second person 114 based on the output of the person position sensor 374 and transmitting a detection signal to the operation control PC 362 via the network 364. In this embodiment, a commercially available person position sensor that is often used as an input device for video games is used as the person position sensor 374. In addition, considering that there are two people, the first person 112 and the second person 114, who will be interacting with the robot 110, and that the robot 110 needs to know the directions of the first person 112 and the second person 114 from the robot 110 as a base point to interact with these two people, the conversational robot system 350 uses the two person position sensors 374 to separately detect the positions of both the first person 112 and the second person 114.
[0059] The flow of data between the elements shown in Fig. 7 will be briefly described with reference to Fig. 8. The voice recognition PC 368 shown in Fig. 7 performs voice recognition 402 on the speech of the first person 112 and the second person 114 based on the output of the microphone 366, and outputs the recognition result. The person position recognition PC 376 shown in Fig. 7 performs person position recognition 400 on the first person 112 and the second person 114 based on the output of the person position sensor 374, and outputs the recognition result.
[0060] The motion control PC 362 shown in Fig. 7 executes dialogue control 404 in accordance with the above-mentioned scenario. The motion control PC 362 executes dialogue control 404 based on the recognition results from the person position recognition 400 and the voice recognition 402, and provides text and speech commands for speaking in accordance with the scenario to the voice synthesis PC 372 in Fig. 7, and similarly provides commands and parameters for operation for controlling each part of the robot 360 in accordance with the scenario to the person position recognition PC 376 in Fig. 7.
[0061] The speech synthesis PC 372 performs speech synthesis 406 on the given text in accordance with instructions from the dialogue control 404, and outputs a speech signal to the speaker 370. The movement control 408 controls each actuator of the robot 360 in accordance with instructions from the dialogue control 404 using parameters given from the dialogue control 404, and causes the robot 360 to perform movements according to a scenario.
[0062] An example of a simple scenario is shown in Fig. 9. Scenario 450 shown in Fig. 9 is for controlling the movement of robot 360, for example, at the beginning of step 304 in Fig. 6. This scenario includes multiple nodes 462, 464, 466, 468, and 470 arranged between start node 460 and end node 472. Each of these nodes is assigned a function. Scenario 450 further includes one or more edges connecting the nodes so that robot 360 operates in a predetermined order according to the scenario.
[0063] For example, node 462 is a speech block. This speech block includes information about the content of the speech, the robot's movements and emotions while speaking, etc. Node 464 is a question block. This question block includes information about the person asking the question, the content of the question, and the movements of the robot 360 when asking the question. The questions here are basically assumed to be answerable with a YES / NO. Node 466 is a voice recognition block. In this node, voice recognition is performed on the answer of a person (e.g., the first person 112) to the question uttered by the robot 360 in node 464. Node 466 further determines whether the answer of the first person 112 to the question is a positive or negative answer based on the result of the voice recognition. The result is provided to the subsequent node 468. Node 466 further determines how the robot 360 should respond based on the answer of the first person 112, and therefore determines parameters for controlling each actuator of the robot 360. These parameters are provided to node 470.
[0064] The processing of node 468 and the processing of node 470 are performed in parallel. When the robot 360 speaks, it is necessary to control each part of the face and each part of the body. The head of the robot 360, especially the mouth, needs to be precisely controlled while it is speaking. In this example, the movement over time is calculated in 466 and provided to nodes 468 and 470 so that the speech at node 468 and the control of the robot 360 by node 470 are synchronized. Execution of this scenario ends when the speech at node 468 and the robot's movement at node 470 end.
[0065] In this embodiment, the above-described scenario is defined in a graph format as shown in Fig. 9. When each node in the graph is executed, a script is generated by interpreting information about that node. However, conversely, the progress of control from the script may be represented in a graph format as shown in Fig. 9 so that it can be visually understood. Scenarios may be created alternately between graphs and scripts.
[0066] Fig. 10 is a diagram illustrating the functional configuration of the dialogue control 404 shown in Fig. 8. Referring to Fig. 10, the dialogue control 404 is a process realized by the operation control PC 362 shown in Fig. 7. The dialogue control 404 includes a storage device 510, which is realized by a hard disk or a solid state drive (SSD) of the operation control PC 362, and which stores dialogue scenario data that defines a dialogue scenario, and a dialogue scenario control 512 that reads dialogue scenario data 530 from the storage device 510, interprets the processing content for each node, and calculates an emotional state 532 of the robot 360 and a parameter command 534 for voice synthesis.
[0067] The dialogue control 404 further includes an emotional state management 514 that receives an emotional state 532 from the dialogue scenario control 512 to manage the emotional state of the robot 360 and outputs a parameter command 515 for voice synthesis in accordance with the emotional state of the robot 360 and an expression command 538 for controlling an actuator or the like that changes the facial expression of the robot 360, and a voice synthesis parameter control 516 that receives the parameter command 534 for voice synthesis from the dialogue scenario control 512 and the parameter command 515 in accordance with the emotional state from the emotional state management 514, and outputs a voice synthesis command 540 that controls the final voice synthesis.
[0068] A voice synthesis command 540 from the voice synthesis parameter control 516 is given to a voice synthesis module 504 operating in a voice synthesis PC 372 shown in Fig. 7. The voice synthesis module 504 performs voice synthesis according to the scenario and the emotional state of the robot 360 in accordance with this voice synthesis command 540. A gaze / movement command 536 from the dialogue scenario control 512 and a facial expression command 538 from the emotional state management 514 are both given to a movement control module 502 operating in a movement control PC 362 shown in Fig. 7. The movement control module 502 controls each actuator of the robot 360 based on these commands, causing the robot 360 to move in accordance with the scenario.
[0069] 11 shows the control structure of a program that realizes local positive / negative recognition, as mentioned in the description of step 234 in FIG. 5. In this embodiment, the progress of the scenario is determined depending on whether the user's utterance in response to a question posed by the robot 360 is a positive or negative response to the question. In this embodiment, the user's utterance is determined by keyword matching according to a script. Therefore, if the robot 360 cannot hear what the user is saying, the robot 360 is equipped with a recovery system that allows the robot 360 to ask the user to repeat the utterance, for example.
[0070] Furthermore, when recognizing whether a user's utterance in response to the robot 360's question is affirmative or negative, it is not limited to cases where the user uses words generally used to express affirmation and negation, such as "yes" or "that's right." For example, in response to the question "Do you think...?", the response "I think" is affirmative, and the response "I don't think" is negative. In other words, the user may use not only general affirmative and negative words, but also local affirmative and negative words, including verb changes, in response to the other person's utterance, depending on local information, namely the other person's utterance. The program in FIG. 11 is designed to recognize such local affirmative and negative words.
[0071] 11, in local positive / negative recognition, it is first determined whether or not there is a positive keyword in the character string resulting from the speech recognition of the user utterance (step 562). In this case, the positive keyword is a word prepared for each question in each node according to the context.
[0072] If the determination in step 562 is positive, then in step 564 it is determined whether or not the utterance contains a negative word. If a negative word is included, then in step 574 it is determined that the user utterance is a negative response to the question, and this process ends. If the determination in step 564 is negative, then in step 572 it is determined that the user utterance is a positive response to the question, and this process ends.
[0073] If the determination in step 562 is negative, then in step 566 it is determined whether or not there is a negative keyword in the utterance. If there is a negative keyword, then in step 568 it is further determined whether or not a negative word is included in the utterance. If the determination is negative, then in step 574 it is determined that the user utterance is a negative response to the question, and this process ends. If the determination in step 568 is positive, then this is the negation of a negation, and so in step 572 it is determined that the user utterance is a positive response to the question, and this process ends.
[0074] If the determination in step 566 is negative, then in step 570 it is determined that there is no local positive or negative keyword in the utterance sentence, and this process ends.
[0075] This process is incorporated into the analysis of the user's utterance after the robot asks a question. As a result, even if the user responds to a question without using a general-purpose positive or negative keyword, the robot 360 can correctly determine whether the response is positive or negative. This allows subsequent processing to be carried out smoothly according to the scenario.
[0076] FIG. 12 is a flowchart showing the flow of utterance sentence analysis of the utterance of the first person 112, which the robot 360 executes when it hears the first person 112's response to a question posed by the robot 360 in step 208 of FIG. 4, step 236 of FIG. 5, step 308 of FIG. 6, etc. First, the robot 360 executes speech recognition processing on the utterance of the first person 112 (step 600). The robot 360 then performs morphological analysis on the character string obtained by the speech recognition processing (step 602). The robot 360 then determines whether the question is intended to be answered with an adjective, and branches the control flow according to the result (step 604). This step 604 is performed because, depending on the topic and the way the question is asked, there are cases where the answer is intended to be primarily adjectives and cases where the answer is intended to be primarily nouns. In other words, the determination result of this step 604 differs depending on the scenario.
[0077] If the determination in step 604 is positive, a first utterance sentence analysis process is executed assuming an adjective answer (step 606). If the determination in step 604 is negative, a second utterance sentence analysis process is executed assuming a noun answer, as will be described later (step 608). In both steps 606 and 608, a reflection sentence is basically generated. In the following step 610, the generated reflection sentence is uttered to the first person 112. However, in some cases, a reflection sentence may not be generated. In such cases, a predetermined default reflection sentence may be used, or no reflection may be performed.
[0078] FIG. 13 shows the process flow in step 606 of FIG. 12 in more detail. Referring to FIG. 13, in utterance analysis when an adjective is assumed to be an answer, first, it is determined whether or not the utterance contains an adjective (step 632). If an adjective is present in the utterance, it is determined whether or not the adjective "fun" is included (step 634). This determination is made because, as shown in FIG. 6 above, a drinking party was selected as the topic of the three-way dialogue, and therefore "fun" was specially inserted as an adjective that is used frequently. If "fun" is included in the adjective, it is set as a reflection word (a word to be inserted into a question utterance template for reflection) (step 636). Then, a question utterance is generated by inserting the reflection word into the reflection word slot of a prepared question utterance template for reflection (step 646). After step 646, this process ends.
[0079] If the determination in step 634 is negative, i.e., if the word "fun" does not exist in the sentence uttered by the first person 112, the adjective with the highest IDF (Inverse Document Frequency) value among the adjectives in the sentence uttered is set as the reflection word (step 638). After this, the process of step 646 is executed. The IDF value is considered to be a measure of whether a word is used infrequently. A high IDF value means that the word is used infrequently and is not a common word. Therefore, by creating a question utterance using a word with a high IDF value, it is thought that when the second person 114 is later asked to speak on behalf of the first person 112, the second person 114 will be able to think from the perspective of the first person 112.
[0080] To perform this process, it is necessary to store an IDF value for each word. In this embodiment, for example, the operation control PC 362 shown in FIG. 7 stores a table of these IDF values.
[0081] If the determination in step 632 is negative, it is determined whether or not the utterance sentence contains a verb (step 640). If a verb is present, the verb with the highest IDF value is set as the reflection word (step 642). In step 646, a question utterance sentence using the reflection word is generated, and the process ends.
[0082] In this embodiment, if the determination in step 640 is negative, no follow-up question is asked (step 644), and the robot 360 does not ask the first person 112 a follow-up question.
[0083] Figure 14 shows the process flow in step 608 in Figure 12 in more detail. Referring to Figure 14, in this utterance sentence analysis when a noun is assumed to be the answer, first, in step 662, it is determined whether or not a noun is present in the utterance sentence. If the determination is positive, in step 664, the noun with the highest IDF value among the nouns is set as the reflection word. Thereafter, in step 672, a question utterance sentence using the reflection word is generated, and this process ends. The process in step 672 is substantially the same as step 646 in Figure 13.
[0084] If the determination in step 662 is negative, i.e., if there is no noun in the utterance sentence, it is determined in step 666 whether or not there is a verb in the utterance sentence. If there is a verb in the utterance sentence, the verb with the highest IDF value is set as the reflection word in step 668. Thereafter, in step 672, a question utterance sentence using the reflection word is generated, and this process ends.
[0085] If the determination in step 666 is negative, i.e., if there is neither a noun nor a verb in the utterance sentence, step 670 is executed. In step 670, it is determined to use a default question utterance sentence that does not have a reflection word, and this process ends.
[0086] FIG. 15 is a flowchart showing the process flow in step 646 in FIG. 13 and step 672 in FIG. 14. Referring to FIG. 15, in this process, first, in step 702, it is determined whether the question uttered by the robot 110 is intended to be a noun as the other person's answer. If this determination is positive, control proceeds to step 704. In step 704, it is determined whether the reflection word is a noun. This reflection word is the word determined in the process of FIG. 13 or 14. If this determination is positive, control proceeds to step 714. In step 714, the reflection word is applied to a question template for nouns that has been prepared in advance, and this process ends. A question template is a part of a question sentence that has a slot provided in which the reflection word is inserted. A question utterance sentence can be generated by inserting the reflection word into this slot.
[0087] On the other hand, if the determination in step 704 is negative, the process proceeds to step 708. The process performed in step 708 will be described later.
[0088] If the determination in step 702 is negative, control proceeds to step 706. In step 706, it is determined whether the expected reflection word is an adjective. If the determination is positive, in step 710, a question utterance is generated by applying the reflection word to a prepared question template for adjectives, and this process ends. If the determination in step 706 is negative, control proceeds to step 708.
[0089] In step 708, it is determined whether the reflection word is a verb. If the determination is affirmative, the process proceeds to step 712; if the determination is negative, the process proceeds to step 716. In step 712, the reflection word is applied to a question template for verbs to generate a question utterance, and this process ends. In step 716, it is determined that there is no question utterance, and this process ends. If there is no question utterance, the robot 110 does not ask for a reflection.
[0090] 4th action Conversational robot system 350 (FIG. 7) configured as above operates as follows: The details of the control of the operation of each robot are publicly known, and therefore will not be repeated below.
[0091] 4, in step 200, the robot 110 greets the first person 112 and the second person 114. In step 202, the robot 110 first makes an utterance to introduce itself, and then prompts the first person 112 and the second person 114 to introduce themselves.
[0092] After the self-introductions are completed, the robot 110 performs processing to promote a relationship between the first person 112 and the second person 114 in step 204 and subsequent steps. First, in step 204, the first step 100 (FIG. 3) asks the first person 112 and the second person 114 an initial question. This question asks about the participants' preferences for a topic. For example, as shown in step 304 of FIG. 6, the robot 110 asks whether the participants like alcohol, listens to the answers of both the first person 112 and the second person 114, and stores the information. As described above, the pattern of questions to be asked later varies depending on the combination of the participants' preferences for a topic. Therefore, the robot 110 understands and stores the answers of the participants in step 204. At this time, the robot 110 performs local affirmative / negative recognition shown in FIG. 11 on each of the answers of the first person 112 and the second person 114 to correctly understand each answer. In the following description, it is assumed that both the first person 112 and the second person 114 have answered that they like alcohol.
[0093] Next, in step 206, the robot 110 discloses its preferences regarding alcohol. Here, the robot 110 discloses, for example, that it cannot drink alcohol and does not like drinking parties. After this disclosure, in step 208 and thereafter, the robot 110 executes a process for promoting the building of a relationship between the first person 112 and the second person 114. A specific example of this process will be described with reference to FIG. 5.
[0094] Referring to FIG. 5, first, in step 230, the robot 110 turns toward the first person 112. This action lets the first person 112 (here, referred to as "Mr. A") and the second person 114 know that the robot 110 is about to ask the first person 112 a question. In the following step 232, the robot 110 asks the first person 112 a second question. In this example, the second question is, for example, "I heard that Mr. A likes alcohol. Do you like drinking parties?" However, if the first person 112 answers "no" in step 204 shown in FIG. 4, the question here is a different one that has been prepared in advance. The type of question to be used is prepared in advance based on a scenario. The first person 112 answers this question. For example, the first person 112 answers "Yes, I like it." In this case, the robot 110 also performs local affirmative / negative recognition on the answer of the first person 112. In step 234, the robot 110 further asks the first person 112 a question to find out the reason (for example, "Is that so? Why?"). This question is also prepared in advance in the scenario, and it is assumed that the answer will include an adjective. The first person 112 gives some kind of answer to this question.
[0095] For example, suppose the answer of the first person 112 is "Because it's fun." In response to this answer, in step 236, the robot 110 pretends not to understand the answer and asks further probing questions about the reason.
[0096] More specifically, at this time, the robot 110 executes the process shown in Fig. 12. That is, referring to Fig. 12, the robot 110 performs speech recognition on the answer of the first person 112 in step 600. Furthermore, the robot 110 performs morphological analysis on the character string obtained by speech recognition in step 602. As a result, the robot 110 detects the adjective "fun" in the recognized utterance. Since the original question was intended to be answered with an adjective, the determination in step 604 is positive, and the first utterance sentence analysis is performed in step 606.
[0097] Referring to FIG. 13, the determination in step 632 of the first utterance sentence analysis is positive, and the determination in step 634 is also positive. Therefore, step 638 is executed. In this example, the utterance of the first person 112 contains only one adjective. Therefore, the only adjective contained in the utterance, "fun," is set as the reflection word. Furthermore, in step 646, a question utterance sentence is generated using this reflection word, "fun." Here, for example, it is assumed that "Hmm, <slot>? What is <slot>?" has been prepared as a question template.
[0098] In step 646, the process of Fig. 15 is performed. More specifically, in this example, steps 702, 706, and 710 are executed in this order. As a result, the robot 110's retrieval question utterance is "Yes, it's fun. What's fun about it?"
[0099] Returning to Fig. 12, in step 610, the robot 110 asks the first person 112, "Yes, it's fun? What's fun about it?" This question, "What's fun about it?", is a question that assumes a noun as the answer. Then, the first person 112 will respond with some kind of answer, for example, "Because the conversation is fun."
[0100] 5, the robot 110 listens to the answer of the first person 112 to the probing question in step 236 in step 238, and pretends not to understand the answer in step 240. For example, the robot 110 shakes its head and speaks as if it does not understand what the first person 112 is saying.
[0101] At this time, the robot 110 executes the process shown in Fig. 12 on the utterance of the first person 112, similar to the process in step 236. That is, the robot 110 first performs speech recognition on the utterance of the first person 112 (step 600), and then performs morphological analysis (step 602) to identify the words and their parts of speech included in the utterance of the first person 112. The probing question asked in step 236 is intended to be a noun. Therefore, the determination in step 604 is negative, and the process in step 608 is executed.
[0102] In step 608, the process of FIG. 14 is performed. Referring to FIG. 14, in this example, the determination in step 662 is positive. Therefore, step 664 is executed. Since the utterance of the first person 112 contains only one noun ("conversation"), "conversation" is set as the reflection word. In step 672, a question utterance sentence using the reflection word is generated.
[0103] The process executed in Figure 15 will not be repeated. Here, it is assumed that the question utterance is "Conversation? Hmm. I don't quite get it." To emphasize that the answer from first person 112 is not fully understood, an utterance such as "Uh huh. Huh?" may be inserted before the question utterance. After this, returning to Figure 12, a rephrase is executed in step 610.
[0104] 5, after pretending not to understand the answer from the first person 112 in step 240, the robot 110 turns to the second person 114 in step 242. As a result of this action, the second person 114 knows that the robot 110 will now speak to him / her.
[0105] In the next step 246, the robot 110 asks the second person 114 (hereinafter referred to as "Mr. B") whether he or she understands what the first person 112 is saying. The utterance of the robot 110 at this time may be, for example, simply "Mr. B, do you understand?"
[0106] As a result of this processing, the second person 114 considers the meaning of the utterance of the first person 112 from the perspective of the first person 112 in response to the question from the robot 110, and explains it to the robot 110. The robot 110 listens to the answer from the second person 114 and behaves in a way that indicates it understands the utterance of the first person 112.
[0107] After this series of processes is completed, the process of step 214 shown in FIG. 4 is executed, and the three-way dialogue is completed.
[0108] 5. Elemental technologies To realize the above embodiment, detailed techniques are required for controlling the robot. The main techniques will be explained below.
[0109] In the above embodiment, the sensors used to perform the necessary recognition actually include multiple depth sensors, microphone arrays, and cameras installed around the android and the interacting human. By integrating information from these sensors, human detection, tracking, speaker recognition, and speech separation and recognition for each speaker are achieved. Note that in the above embodiment, a camera is not required as long as the position of the human can be detected. However, using a camera is preferable for robot-human interaction.
[0110] The pitch and speed of the voice uttered by the robot 110 are specified when the robot 110 makes emotional utterances. Specifically, the speed is made faster to express a state of increased arousal, whether the emotion is negative or positive. The pitch is made lower for negative emotions and higher for positive emotions. Furthermore, the robot 110 generates facial expressions that express a smile for positive emotions, and expressions that express "sadness" or "disgust" for negative emotions. Existing facial expressions are used to generate facial expressions based on emotions.
[0111] In this scenario, the following emotional states were assumed for the robot 110: In other words, when talking about a drinking party that the robot dislikes, the robot will have negative emotions, and when talking about a person who dislikes drinking parties and empathizing with them (inevitably, if both people like drinking parties, the robot will not have positive emotions), the robot will have positive emotions.
[0112] In the above embodiment, a conversation between three people is conducted, so it is necessary to switch the gaze target. When the robot 110 is speaking to two people, it switches its gaze to one of the people at one or two utterance breaks, and when speaking to one of them, it turns to that person and speaks.
[0113] The robot 110's behavior generation realizes conscious, unconscious, and automatic behavior. For conscious behavior, the robot generates a behavior by specifying the target position for attention or instructions. In this dialogue script, when asking a person's name on first meeting, the robot generates a gesture of turning toward the target and extending its hand to make it clear which of the two people the question is being asked. For unconscious behavior, the robot's emotional state is specified and postures, facial expressions, and movements that express that state are generated. For automatic behavior, the robot automatically generates behaviors that occur automatically due to physical constraints. Specifically, these are automatic lip and head movements that accompany the robot's speech, and physiologically occurring movements (such as blinking).
[0114] 5. Experiments and Results The following experiment was conducted to confirm the effects of the present invention. The purpose of this experiment was to verify whether the building of relationships between people can be promoted by having one person speak on behalf of another person regarding their opinion. Specifically, a human-like android was used as the robot, and the scenario shown in Figure 6 was used. For this purpose, the experiment had two conditions: an embodiment condition in which a dialogue was conducted according to the above embodiment, and a control condition in which the person whose opinion was heard directly asked the other person to respond.
[0115] 16 , under embodiment condition 740, first, in first pattern 750, android 780 engages in dialogue 790 with subject 782. Then, in second pattern 752, android 780 engages in dialogue 792 with experimenter 784 regarding the opinion of subject 782, seeking a representation of subject 782. Continuing the dialogue, android 780 engages in dialogue 794 with experimenter 784 in n−1th pattern 754, and then in the subsequent nth pattern 756, android 780 engages in dialogue 796 with subject 782, seeking a representation of the opinion of experimenter 784 in n−1th pattern 754.
[0116] On the other hand, the first pattern 760 of the control condition 742 was the same as the first pattern 750 of the embodiment condition 740, and the (n-1)th pattern 764 was the same as the (n-1)th pattern 754 of the embodiment condition 740. However, in the second pattern 762 of the control condition 742, the android 780 had a further dialogue 800 with the subject 782 regarding the opinion of the subject 782 in the first pattern 760, to ask for an explanation. Furthermore, in the n-th pattern 766 of the control condition 742, the android 780 had a further dialogue 802 with the experimenter 784 to ask for an explanation.
[0117] Depending on whether the subject and the experimenter had like / like, like / dislike, or dislike / dislike preferences, the situations in which probing questions could be asked were set as steps 308 and 312, steps 308 and 312, and steps 310 and 312. The questions that can be probing questions in steps 308, 310, and 312 are as follows:
[0118] Step 308: Why do you like drinking parties? Step 310: What did you dislike about drinking parties? Step 312: How did you become friends at the drinking party? 5-1 Experimental procedure This experiment involved a three-way dialogue between one android and two humans (one subject and one experimenter). Twenty-two subjects (7 men, 15 women, mean age 39.6 years, standard deviation 12.3 years) participated in the experiment. A between-subjects design was used, with 11 subjects assigned to the embodiment condition and 11 to the control condition. The topic was the preference for "drinking parties," as described above. Subjects were informed that the other human participant was also a similarly recruited subject. This was done to prevent subjects from engaging in a dialogue knowing that the other human participant was the experimenter, which could lead to thoughts that differ from the actual situation. After the experiment was completed, the other human participant revealed that he or she was actually the experimenter. Subjects first engaged in a dialogue with the android and the experimenter. Then, subjects answered each question on a questionnaire.
[0119] As instructions to the subjects, the experimenter read out the following sentence: "You will now be having a conversation with the robot together with the other subjects. The robot may ask you questions about your likes and dislikes, so please answer with either "like" or "dislike" as much as possible. Even if you are neutral or not interested, please answer with either one if you must. Once you sit down, the robot will speak to you."
[0120] As a note of caution regarding the experimenter's speaking style, the subject was to be named "Yamada" and could answer either "likes" or "dislikes." Furthermore, regardless of whether the subject liked or disliked something, the experimenter was to respond with a clear, unambiguous opinion. When the android asked about the other subject's opinion, the subject was to state their own thoughts and opinions, and not to respond with "I don't know." Furthermore, for questions from the android that could be answered with a "yes" or "no," the experimenter was to answer "yes."
[0121] 5-2 Evaluation items (questionnaire) In the experiment, two types of questionnaires were prepared: one to evaluate impressions regarding the interaction, and the other to evaluate impressions regarding the android itself. However, here we will only introduce the questions regarding impressions regarding the interaction.
[0122] To evaluate the dialogue, we prepared the following 12 questions: 1-4 are for evaluating the dialogue with the android, 5-9 are for evaluating the dialogue with the other human, and 10-12 are for evaluating the dialogue between the three parties as a whole. Question 1. Did you feel that the android understood what you were saying? (Android comprehension) Question 2: Did you feel that you had a good time interacting with the android? (Interaction satisfaction with the android) Question 3. Did you feel like you became friends with the android? (Building a relationship with the android) Question 4: Would you like to talk to an android again? (Willingness to communicate with an android) Question 5. Did you feel that the other subject understood what you said? (Human Comprehension) Question 6: Did you feel like you had a good time with the other participant? (Interaction satisfaction with humans) Question 7. Did you feel like you had a good relationship with the other participant? (Human Relationship Building) Question 8. Did you feel like you were having a conversation with another subject? (Feeling like you were interacting with a human being) Question 9: Did you feel like talking to the other participant again? (Interaction with humans) Question 10: Did you feel that the three of you had a great time? (Satisfaction with the interaction among the three people) Question 11. Did you feel like you were having a conversation with three people? Question 12: Do you feel like you want to talk with the three of you again? (Interest in three-way conversation) For these questions, subjects were asked to choose a response on a seven-point scale from 1 (strongly disagree) to 7 (strongly agree).
[0123] 5-3 Results The results obtained in this experiment are as follows:
[0124] When subjects were asked to speak on behalf of the experimenter's comments, no subjects stated that they could not speak on behalf of the experimenter. There are four possible distributions of preferences between the subjects and the experimenter: "like-like" (8), "like-dislike" (6), "dislike-like" (3), and "dislike-dislike" (5). However, the numbers in parentheses indicate the distribution in the experiment. In this experiment, the subject's preferences do not necessarily control the preference pattern (for example, even if "like-like" is assumed, if the subject says "dislike," the situation does not go as expected). However, the experimenter's preferences were controlled to achieve as much balance as possible.
[0125] The results of the impression evaluation regarding the dialogue are shown in Figure 17. In Figure 17, the hatched graphs show the results for the embodiment condition, and the open graphs show the results for the control condition. In Figure 17, an "*" above each pair of graphs indicates that there is a significant difference between the two at p<0.05, and an "†" indicates that there is a significant difference at p<0.1. The error bars on each graph indicate the standard error.
[0126] T-tests were performed on each question, and significant differences were confirmed for satisfaction with interaction with an android (p=0.019), willingness to interact with an android (p=0.022), satisfaction with interaction with a human (p=0.031), building a relationship with a human (p=0.043), willingness to interact with a human (p=0.033), satisfaction with three-way interaction (p=0.006), feeling of three-way interaction (p=0.049), and willingness to interact with a human (p=0.007).On the other hand, at p<0.05, no significant differences were confirmed for android's ability to understand (p=0.078), human's ability to understand (p=0.066), and feeling of interacting with a human (p=0.054).
[0127] The experimental results showed that, first, in the evaluation of the impression of the conversation with the android, the method according to the embodiment was able to significantly improve the impression that the conversation with the android was more lively (dialogue satisfaction) and that the subject wanted to talk again (willingness to talk). Furthermore, in the evaluation of the impression of the conversation with another subject (actually the experimenter), the method according to the embodiment was able to significantly improve the impression that the conversation with the other subject (the experimenter) was more lively (dialogue satisfaction) and that the subject had become closer (building a relationship). Furthermore, in the evaluation of the impression of the three-person conversation, the method according to the embodiment was able to improve the impression that the conversation was more lively (dialogue satisfaction), that the subject had a sense of conversation, and that the subject wanted to talk again (willingness to talk).
[0128] 6. Computer Hardware Configuration Fig. 18 is an external view of a computer system that operates as, for example, the operation control PC 362 shown in Fig. 7. Fig. 19 is a hardware block diagram of the computer system shown in Fig. 18. The voice recognition PC 368, voice synthesis PC 372, and person position recognition PC 376 shown in Fig. 7 can also be realized by a computer system having almost the same configuration as the operation control PC 362. Here, only the configuration of the operation control PC 362 will be described, and detailed configurations of the other PCs will not be described.
[0129] 18, this computer system 950 includes a computer 970 having a DVD (Digital Versatile Disc) drive 1002, and a keyboard 974, a mouse 976, and a monitor 972 for interacting with a user, all of which are connected to the computer 970. Of course, these are just one example of a configuration for when user interaction is required, and any general hardware and software that can be used for user interaction (for example, a touch panel, voice input, or a general pointing device) can be used.
[0130] 19, the computer 970 includes, in addition to a DVD drive 1002, a CPU (Central Processing Unit) 990, a GPU (Graphics Processing Unit) 992, a bus 1010 connected to the CPU 990, the GPU 992, and the DVD drive 1002, a ROM (Read-Only Memory) 996 connected to the bus 1010 and storing a boot-up program and the like for the computer 970, a RAM (Random Access Memory) 998 connected to the bus 1010 and storing instructions constituting a program, a system program, working data, and the like, and an SSD 1000 which is a non-volatile memory connected to the bus 1010. The SSD 1000 is for storing programs executed by the CPU 990 and the GPU 992, data used by the programs executed by the CPU 990 and the GPU 992, and the like. Computer 970 further includes a network I / F (Interface) 1008 that provides connection to a network 986 (network 364 shown in FIG. 7) that enables communication with other terminals, and a USB port 1006 to which a USB (Universal Serial Bus) memory 984 can be attached or detached and that provides communication between USB memory 984 and each part within computer 970.
[0131] The computer 970 further includes an audio I / F 1004 that is connected to the microphone 982, the speaker 980, and the bus 1010, and has the function of reading out audio signals, video signals, and text data generated by the CPU 990 and stored in the RAM 998 or the SSD 1000 in accordance with instructions from the CPU 990, converting them to analog, amplifying them, and driving the speaker 980, and digitizing the analog audio signal from the microphone 982 and storing it at any address in the RAM 998 or the SSD 1000 specified by the CPU 990.
[0132] In the above embodiment, programs for realizing the functions of the operation control PC 362, the voice recognition PC 368, the voice synthesis PC 372, and the person position recognition PC 376 are all stored in, for example, the SSD 1000, RAM 998, DVD 978, or USB memory 984 shown in Fig. 19, or a storage medium of an external device (not shown) connected via the network I / F 1008 and the network 986. Typically, these data and parameters are written to the SSD 1000 from the outside, for example, and loaded into the RAM 998 when the computer 970 is executed.
[0133] 7 and computer programs for operating this computer system to realize the functions of the operation control PC 362, the voice recognition PC 368, the voice synthesis PC 372, and the person position recognition PC 376 and their respective components are stored on a DVD 978 inserted into the DVD drive 1002 and transferred from the DVD drive 1002 to the SSD 1000. Alternatively, these programs may be stored in a USB memory 984, and the USB memory 984 may be inserted into the USB port 1006 and the programs may be transferred to the SSD 1000. Alternatively, the programs may be transmitted to the computer 970 via the network 986 and stored in the SSD 1000.
[0134] The program is loaded into RAM 998 when executed. Of course, a source program may be input using the keyboard 974, monitor 972, and mouse 976, and the compiled object program may be stored in SSD 1000. In the case of a script language as in the above embodiment, a script input using the keyboard 974 or the like may be stored in SSD 1000. In the case of a program that runs on a virtual machine, a program that functions as a virtual machine must be installed in computer 970 in advance. A neural network is used for speech recognition, speech synthesis, etc. A trained neural network may be used, or training may be performed in conversational robot system 350.
[0135] The CPU 990 reads a program from the RAM 998 according to an address indicated by an internal register called a program counter (not shown), interprets the instructions, reads data required to execute the instructions from the RAM 998, the SSD 1000, or other devices according to the address specified by the instruction, and executes the processing specified by the instruction. The CPU 990 stores the execution result data at an address specified by the program, such as in the RAM 998, the SSD 1000, or a register within the CPU 990. Depending on the address, the data may be output from the computer as a command to a robot actuator, an audio signal, or the like. At this time, the program counter value is also updated by the program. The computer program may be loaded directly into the RAM 998 from the DVD 978, the USB memory 984, or via the network 986. Note that some tasks (mainly numerical calculations) of the program executed by the CPU 990 are dispatched to the GPU 992 according to instructions contained in the program or according to the analysis results obtained when the CPU 990 executes the instructions.
[0136] The program that enables the computer 970 to implement the functions of each unit according to the above-described embodiment includes a plurality of instructions written and arranged to cause the computer 970 to operate to implement those functions. Some of the basic functions required to execute these instructions may be provided by an operating system (OS) or third-party program running on the computer 970, various toolkit modules installed on the computer 970, or a program execution environment. Therefore, the program does not necessarily include all of the functions required to implement the system and method according to this embodiment. The program may include only instructions that execute the operations of the above-described devices and their components by statically linking or dynamically calling appropriate functions or modules in a controlled manner to achieve the desired results. The method for operating the computer 970 to achieve this is well known, and will not be repeated here.
[0137] The GPU 992 is capable of parallel processing, and can execute a large amount of calculations involved in machine learning simultaneously in parallel or in a pipelined manner. For example, parallel calculation elements discovered in a program when the program is compiled or when the program is executed are dispatched from the CPU 990 to the GPU 992 as needed, and executed. The results are returned to the CPU 990 directly or via a predetermined address in the RAM 998, and assigned to a predetermined variable in the program.
[0138] 7th Variation In the above embodiment, the conversational robot system 350 includes multiple PCs, which communicate with each other via a network. However, the present invention is not limited to such an embodiment. For example, the conversational robot system 350 can be realized using only one PC with sufficiently high performance. Furthermore, if the computer is small, the necessary one or more computers can all be incorporated into the robot's main body. In the above embodiment, an attempt is made to build a relationship between people meeting for the first time through a three-way dialogue. However, the present invention is not limited to such an embodiment. Even in the case of four or more people, it is believed that the same effect as a three-way dialogue can be achieved by combining three-way dialogue centered around the robot.
[0139] In addition, an android was used as the robot in the above experiment. Androids resemble humans and are therefore considered suitable for three-way dialogue. However, robots are not limited to androids. For example, a small, humanoid agent robot could also be used. Robots are not limited to humanoids; they could also be animal-shaped. Furthermore, in a broader sense, people, androids, agents, etc. that are realized as images and displayed on a monitor can also be considered to be similar to robots.
[0140] Furthermore, in the above embodiment and experiment, the robot only asks for reflection once, but the present invention is not limited to such an embodiment. The robot may ask for reflection two or more times.
[0141] The embodiments disclosed herein are merely examples, and the present invention is not limited to the above-described embodiments. The scope of the present invention is defined by the claims in the appended claims, taking into consideration the detailed description of the invention, and includes all modifications within the meaning and scope equivalent to the wordings described therein. [Explanation of symbols]
[0142] R, 110, 360 Robot 350 Conversational Robot System 362 Motion Control PC 364,986 Network 366, 982 microphones 368 Voice Recognition PC 370, 980 speakers 372 Speech Synthesis PC 374 Person Location Sensor 376 Human location recognition PC 400 people location recognition 402 Speech Recognition 404 Dialogue Control 406 Speech Synthesis 408 Motion Control 450 Scenarios 502 Motion Control Module 504 Speech Synthesis Module 510 Storage device 512 Dialogue Scenario Control 514 Emotional State Management 516 Speech synthesis parameter control
Claims
1. a step in which the computer controls the robot to make an utterance to the first person to encourage their opinion on a topic determined by a predetermined method; a first waiting step in which the computer waits until the first person gives an opinion; and a step in which a computer controls the robot to make an utterance to a second person different from the first person to ask for an explanation regarding the opinion expressed by the first person.
2. The method described in claim 1, wherein the step of making an utterance includes a step in which a computer controls the robot to make an utterance to the first person in accordance with a predetermined scenario, encouraging their opinion on the topic.
3. a step of determining whether or not a word of a predetermined part of speech is present in the utterance of the first person, by the computer, following the first waiting step; a step of determining by a computer what kind of reflection processing to perform according to whether or not the word is present in the utterance of the first person; The method according to claim 1 or claim 2, further comprising: a step of: a computer controlling the robot to perform the reflective processing determined in the determining step.
4. 4. The method according to claim 3, wherein the determining step includes a step of determining, by a computer, whether the robot should execute a process of uttering a reflective sentence including the word or a process of performing a predetermined reflective action, depending on whether the word is determined to be present in the utterance of the first person.
5. The method of claim 4 , wherein the predetermined reflection behavior includes a behavior in which the robot omits reflection to the first person.
6. The method of claim 4 , wherein the predetermined reflection behavior includes the robot uttering a predetermined reflection sentence to the first person.
7. a second waiting step in which the computer waits until the second person provides an explanation regarding the opinion; 7. The method of claim 1, further comprising: a step of controlling the robot, following the second waiting step, to perform an action indicating that the computer understands the opinion of the first person.
8. a step in which a computer controls the robot to ask the second person for their opinion on a topic determined by a predetermined method; a third waiting step in which the computer waits until the second person gives an opinion; The method of claim 7 , further comprising: a step of: a computer controlling the robot to make an utterance to the first person requesting an explanation for the opinion stated by the second person.
9. a fourth waiting step of waiting until the first person provides an explanation regarding the opinion of the second person; 9. The method of claim 8, further comprising: following the fourth waiting step, a step in which a computer controls the robot to perform an action indicative of understanding the opinion of the second person.
10. The method further includes a step of specifying a direction of the first person relative to the robot prior to the step of making an utterance to prompt an opinion, The step of making an utterance to prompt an opinion includes: a step of controlling the robot by a computer so that the robot faces the first person; The method according to any one of claims 1 to 9, further comprising the step of: a computer controlling the robot to utter questions to prompt opinions on the topic.
11. The method further includes, prior to the step of requesting an explanation regarding the opinion, a step by the computer of identifying a direction of the second person relative to the robot; The step of making an utterance to request an explanation regarding the opinion includes: a computer controlling the robot so that the robot faces the second person; The method according to claim 1 , further comprising: a step of controlling the robot by a computer to make an utterance requesting an explanation regarding the opinion stated by the first person.
12. A computer program product that causes a computer to carry out the method of any one of claims 1 to 11.
Citation Information
Patent Citations
Dialogue agent
JP2006071936A
Interactive object identifying method in robot and robot
JP2007160473A
Dialogue robot and robot control program
JP2018173456A
Information processing device, interactive robot, control method
WO2021200307A1