Methods for recovering from failures in dialogue and computer programs
The method allows robots to detect and recover from communication failures with humans by using facial and speech recognition, enabling adaptive emotional expressions and dialogue scenarios to maintain a natural interaction.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-03-31
- Publication Date
- 2026-03-25
AI Technical Summary
Existing technologies struggle to enable robots to recover from communication failures with humans without causing discomfort, particularly when utterance conflicts occur, as they lack the flexibility to adapt responses based on the conversation partner.
A method for detecting communication failures in robots, involving facial and speech recognition, followed by controlled emotional expressions and adaptive dialogue scenarios to recover naturally, including identifying the conversation partner and adjusting utterances to maintain a natural interaction.
Enables robots to recover from communication failures with humans by classifying the partner and adjusting responses, ensuring a natural and comfortable interaction.
Smart Images

Figure 0007835432000001 
Figure 0007835432000002 
Figure 0007835432000003
Abstract
Description
Technical Field
[0001] This invention relates to robot control technology, and particularly to the improvement of communication technology between humans and humanoid robots.
Background Art
[0002] The recent development of robot technology has been remarkable, enabling robots to perform various tasks that were previously difficult. Among them, robots are expected to substitute for human work in daily life or perform tasks that require communication with humans. Therefore, research and development on robots that integrate into society and engage in daily activities involving human interaction have been actively carried out.
[0003] Robots for daily activities are required to have an interactive function. Various robots with an interactive function have been developed so far. Interaction originally refers to the exchange of words between two people facing each other. Therefore, when a human interacts with a robot, it is considered desirable to use a robot that resembles a human as the interaction partner. There is a humanoid robot that has an appearance exactly like a human as an existence that pursues such humanity. Hereinafter, the humanoid robot will be simply referred to as a robot.
[0004] In the communication between humans and robots, there is a published document indicating that for many people, it is difficult to distinguish between a robot and a human when only observing the behavior for a short period of time. Therefore, robots are considered to be a communication medium that can closely interact with humans in society.
[0005] One problem that arises when such robots interact with humans is speech conflict. Speech conflict frequently occurs in conversations between humans as well. In conversations between humans, when a speech conflict occurs, various responses are used to resolve the conflict depending on the situation. However, such flexibility does not exist when the conversation partner is a robot. Therefore, conversations between humans and robots may be quite different from conversations between humans.
[0006] One proposal related to these problems is presented in Patent Document 1, which is listed below. The technology disclosed in the patent document is a technique for estimating who the next speaker is in a conversation involving multiple speakers. This estimation is based on information about changes in the mouth shape of each speaker. Using this method, it may be possible to prevent speech clashes, for example, in remote conferences. [Prior art documents] [Patent Documents]
[0007] [Patent Document 1] Japanese Patent Publication No. 2018-77791 [Overview of the Initiative] [Problems that the invention aims to solve]
[0008] The technology disclosed in Patent Document 1 may be effective in estimating the next speaker in a conversation between humans. However, the technology disclosed in Patent Document 1 cannot solve the problem that when a robot participates in a conversation, its own utterances may conflict with the utterances of the conversation partner. Furthermore, because the conversation partner is a human, there is always a possibility of utterance conflicts occurring between the robot and the human. When such utterance conflicts occur, there is also the problem that the technology disclosed in Patent Document 1 cannot be applied to the problem of how to resolve the utterance conflict without causing discomfort to the conversation partner.
[0009] These problems can essentially be understood as questions of how robots can recover from communication failures with humans without causing discomfort to the other party. For example, if a robot fails to recognize or misrecognizes its conversation partner at the start of a dialogue, it becomes impossible to continue the conversation with the other party. In such cases, too, it is necessary to recover from the communication failure in a natural manner.
[0010] Therefore, this invention aims to provide a method and computer program for recovering from failures in dialogue, which allows robots to recover naturally when they fail to communicate with humans. [Means for solving the problem]
[0011] A method for recovering from a failure in a dialogue according to the first aspect of this invention includes a failure detection step in which a computer detects a failure in communication between a robot and a dialogue partner, and a recovery step in which the computer, in response to the detection of a failure in the failure detection step, controls the robot to engage in a dialogue with the dialogue partner that includes the expression of emotions, in accordance with a predetermined procedure, thereby using the information obtained in the dialogue to recover from the failure.
[0012] Preferably, the failure detection step includes the steps of: selectively controlling the robot to make a first utterance to confirm the result of the identification process with the dialogue partner using a pre-prepared first attitude, depending on whether the confidence level in the identification process of the dialogue partner is higher than a predetermined threshold; controlling the robot to make a second utterance to start the identification procedure with the dialogue partner using a second attitude that appears less confident than the first attitude; and controlling the robot to make a third utterance to start the identification procedure with a second attitude in response to the dialogue partner's response to the first utterance indicating an error in the result of the identification process.
[0013] More preferably, the recovery step includes the computer classifying the conversation partner as an acquaintance to the robot in response that the conversation partner's response to the first utterance indicates that the identification result is correct, and the computer controlling the robot to initiate a conversation according to a pre-prepared scenario for conversation with an acquaintance.
[0014] More preferably, the second and third utterances are the same utterance.
[0015] Preferably, the second utterance is one in which the robot asks whether the person it is talking to is meeting the robot for the first time.
[0016] More preferably, the steps of recovery further include: the step of the computer determining whether the conversation partner's response to the second utterance affirms that the conversation partner is meeting the robot for the first time; the step of the computer controlling the robot so that, in response to a negative response from the conversation partner in the determination step, the computer makes a fourth utterance regarding whether the conversation partner is a person identified by the identification process, with a third attitude pre-prepared to appear even less confident than the second attitude; the step of the computer controlling the robot so that, in response to an affirmative response from the conversation partner to the fourth utterance, the computer classifies the conversation partner as an acquaintance to the robot and begins the conversation with a fourth attitude pre-prepared to appear relieved; and the step of the computer controlling the robot so that, in response to a negative response from the conversation partner to the fourth utterance, the computer displays a fifth attitude pre-prepared to appear disappointed and performs additional identification processing.
[0017] More preferably, the additional identification process includes the steps of: controlling the robot to speak a question asking the conversation partner for their name; generating a determination result by determining whether the name included in the conversation partner's response to the question for their name matches the name of a person registered in a pre-prepared person information database; controlling the robot to classify the conversation partner as an acquaintance to the robot in response to a positive determination result, and to start a conversation according to a scenario for conversation with an acquaintance while displaying a fifth attitude pre-prepared to appear pleased; and controlling the robot to classify the conversation partner as a stranger to the robot in response to a negative determination result, and to start a conversation with the conversation partner according to a scenario pre-prepared for conversation with a stranger.
[0018] Preferably, the additional identification process includes the steps of: controlling the robot to speak a question asking the conversation partner for their name; generating a determination result by determining whether the name included in the conversation partner's response to the question for their name matches the name of a person registered in a pre-prepared person information database; the computer, in response to a positive determination result, performing a process to confirm whether the conversation partner is the same person as a person registered in the person information database, and classifying the conversation partner as an acquaintance or an unknown person to the robot according to the result of the confirmation; the computer, in response to the conversation partner being classified as an acquaintance to the robot, controlling the robot to start a conversation according to a scenario for conversation with an acquaintance while displaying a fifth attitude that has been pre-prepared to appear pleased; and the computer, in response to a negative determination result or the conversation partner being classified as an unknown person to the robot, controlling the robot to start a conversation with the conversation partner according to a scenario that has been pre-prepared for conversation with an unknown person.
[0019] More preferably, the steps to perform the recovery are further, 2 utterance or third utterance The process includes the steps of: controlling the robot with a computer to make a fifth utterance to identify the conversation partner in response to an affirmative response from the conversation partner; the computer generating a determination result regarding whether the information identifying the conversation partner included in the conversation partner's response to the fifth utterance matches the result of the identification process; controlling the robot with a computer to make a sixth utterance to confirm that the conversation partner is an acquaintance of the robot in response to an affirmative determination; and controlling the robot with a computer to classify the conversation partner as an acquaintance of the robot in response to an affirmative response from the conversation partner to the sixth utterance, and to start a conversation according to a scenario for conversation with an acquaintance while displaying a fifth attitude that has been prepared in advance to appear pleased.
[0020] More preferably, the recovery step further includes the step of controlling the robot so that, in response to a negative determination, the computer classifies the conversation partner as an unknown person to the robot and initiates a conversation with the conversation partner according to a pre-prepared scenario for conversation with an unknown person.
[0021] Preferably, the recovery step further includes the step of controlling the robot so that, in response to the interaction partner's response to the sixth utterance being negative, the computer classifies the interaction partner as an unknown person to the robot and initiates a conversation with the interaction partner according to a pre-prepared scenario as a conversation with an unknown person.
[0022] A computer program according to the second aspect of this invention causes a computer to function in order to perform any of the methods described above.
[0023] The above and other objects, features, aspects and advantages of this invention will become apparent from the following detailed description relating to this invention, which will be understood in conjunction with the accompanying drawings.
Brief Description of the Drawings
[0024] [Figure 1] FIG. 1 is a block diagram showing the hardware configuration of a robot system according to a first embodiment of the present invention. [Figure 2] FIG. 2 is a schematic diagram showing an example of a description format of a program executed by the robot system shown in FIG. 1 to control a robot. [Figure 3] FIG. 3 is a flowchart showing a control structure of a program executed when the robot system shown in FIG. 1 starts a conversation with a conversation partner. [Figure 4] FIG. 4 is a schematic diagram showing patterns of speech collisions for which the robot is responsible. [Figure 5] FIG. 5 is a schematic diagram showing patterns of speech collisions for which a person is responsible. [Figure 6] FIG. 6 is a diagram showing the timing for detecting a speech collision for which the robot is responsible. [Figure 7] FIG. 7 is a diagram showing the timing for detecting a speech collision for which a person is responsible. [Figure 8] FIG. 8 is a flowchart showing a control structure of a program executed when the robot system according to the first embodiment detects a speech collision for which the robot is responsible. [Figure 9] FIG. 9 is a flowchart showing a control structure of a program executed when the robot system according to the first embodiment detects a speech collision for which the conversation partner is responsible. [Figure 10] FIG. 10 is a flowchart showing a control structure of a program executed when the robot system according to the first embodiment detects and classifies a speech collision. [Figure 11] FIG. 11 is a diagram for explaining a range to be excluded from a speech collision during speech. [Figure 12] FIG. 12 is a flowchart showing a control structure of a program executed when the robot system according to the second embodiment starts a conversation with a conversation partner. [Figure 13]Figure 13 is a flowchart showing the control structure of the program executed by the robot system according to the third embodiment when it starts a conversation with an interaction partner. [Figure 14] Figure 14 is a flowchart showing the control structure of the program executed by the robot system according to the fourth embodiment when detecting and classifying speech collisions. [Figure 15] Figure 15 is an external view of an example of a computer for realizing a robot system according to each embodiment. [Figure 16] Figure 16 is a block diagram of an example of the computer shown in Figure 15. [Modes for carrying out the invention]
[0025] In the following descriptions and drawings, identical parts are assigned the same reference number. Therefore, detailed descriptions of them will not be repeated.
[0026] First Embodiment 1. Structure Figure 1 shows a block diagram of the hardware configuration of a robot system 100 for communication with humans according to a first embodiment of the present invention. Referring to Figure 1, the robot system 100 includes a camera 60, a microphone 66, a speaker 62, and a robot 110. The robot 110 is a humanoid robot and has actuators in at least the joints of its upper body, and can assume various postures by driving the actuators. In addition, the head of the robot 110 is provided with a plurality of actuators that control the positions of control points defined on the face of the robot 110, and can give the robot various facial expressions by driving these actuators.
[0027] The robot system 100 further includes a Personal Computer (PPC) 116 connected to receive the output of the camera 60 and performing facial recognition on facial images of people in the video output by the camera 60, and a speech recognition PC 118 that receives signals from the microphone 66 and performs speech recognition of the utterances of the person it is interacting with. The Personal Computer (PPC) 116 consists of a neural network and has the function of outputting information indicating which of a predetermined number of people the input facial image belongs to. More specifically, the neural network constituting the Personal Computer (PPC) 116 has as many outputs as there are recognizable people, and outputs the likelihood for each person that the input facial image is the facial image of one of those people. The Personal Computer (PPC) 116 selects the person with the highest likelihood as the facial image recognition result and outputs information corresponding to its identifier. In the following explanation, the likelihood for the person selected as the recognition result is called the confidence level of facial image recognition. The Personal Computer (PPC) 116 outputs not only the identifier of the person who was recognized, but also a predetermined number of other identifiers with high likelihood, along with their likelihood levels.
[0028] The robot system 100 further includes a network 114 to which both a face image recognition PC 116 and a voice recognition PC 118 are connected, and a motion control PC 112 connected to the network 114 for controlling the posture, movements, and facial expressions of the robot 110 by controlling each actuator of the robot 110 according to given control commands. Several movements and facial expressions of the robot 110 are defined in advance, and programs for realizing them are prepared in advance. Examples of facial expressions include confident, unconfident, happy, and disappointed expressions. The motion control PC 112 controls each actuator so that the robot 110 performs the desired movement with the desired facial expression by providing control information such as the duration and magnitude of the movement as arguments to its program.
[0029] To achieve each facial expression, the robot is first made to make various facial expressions using different parameters, and then a questionnaire is administered to multiple subjects regarding what emotions the robot 110 is conveying based on those expressions. Based on the results, the parameters for making the robot 110 make each facial expression can then be determined.
[0030] The robot system 100 further includes a person information database 92 that manages information about a person to be recognized by the face image recognition PC 116 (such as name and affiliation) so that it can be accessed using the person's identifier as a key; an integrated control PC 122 connected to the network 114 and the person information database 92, which controls the overall operation of the robot 110 by calculating the robot's actions and speech content based on information from other PCs connected to the network 114; and a speech synthesis PC 120 connected to the network 114, which performs speech synthesis to produce a specified speech in response to a speech command output from the integrated control PC 122, and provides the speech signal to the speaker 62 to generate sound. The integrated control PC 122 stores a program in a storage device (not shown) that causes the robot 110 to act according to a predetermined scenario. In this embodiment, as will be described later, this program is created as a script represented in graph format.
[0031] In this way, the face image recognition PC 116, voice recognition PC 118, voice synthesis PC 120, integrated control PC 122, and motion control PC 112 work together to control the robot 110 to perform actions that communicate with people.
[0032] Referring to Figure 2, the format of the scenarios stored in the storage device (not shown) of the integrated control PC 122 will be described. In this embodiment, the scenarios are described in a graph format, such as graph 150 shown in Figure 2. Note that Figure 2 is merely for illustrating the scenario description format and is not intended to cause the robot 110 to perform the actions described below.
[0033] Graph 150, shown in Figure 2, consists of blocks representing individual units of robot operation and directed edges connecting each block. Graph 150 includes a start block 160, a speech block 162, a question block 164, and a speech recognition block 166. These blocks are connected in series by directed edges and are executed sequentially along the edges. Similarly, as will be described later for each block, each block contains information for the robot to perform a set of actions. This information includes what the robot should say in that block, the actions it should take, and its emotional state when performing those actions. When the robot reaches each block, it acts based on the information described in that block.
[0034] Graph 150 further includes a facial expression block 168 and an action block 170, which are connected in parallel to each other after the speech recognition block 166, and an end block 172, which is connected after the facial expression block 168 and the action block 170. In the speech recognition block 166, the robot 110 speech-recognizes the dialogue partner's response to the question in block 164, and based on the result, executes block 168 or block 170 with different parameters. In this example, the questions asked in question block 164 are basically questions that can be answered with yes / no, or questions whose category of answer can be predicted (for example, a question asking for the other person's name). In the speech recognition block 166, the dialogue partner's response is classified as yes, no, don't know, neither, no answer, etc. When the dialogue partner's response is yes or no, block 168 operates using that content as a parameter. For example, if the response matches the robot's predicted answer, the robot makes a happy expression, and if it does not match, it makes a surprised expression. If the answer is "I don't know" or "Neither," or if there is no answer, the corresponding parameters are passed to block 170, and the processing of block 170 is executed. In block 170, the robot moves its arms or head according to the parameters.
[0035] In the following embodiments, we assume that the robot 110 is, for example, serving as a receptionist for visitors to a research institute, and we will describe how to recover if communication with a visitor fails during that process. Communication with a visitor as an interaction partner begins with the process of identifying the interaction partner. The embodiments described below relate to recovery processes when the robot fails to identify the interaction partner, and recovery processes when a speech conflict occurs between the interaction partner and the robot during the interaction, as examples of recovery (recovery) from communication failure.
[0036] A. Recovery process from identification failure Figure 3 shows the control structure of a program that enables recovery processing from failure according to a predetermined procedure, which the robot 110 executes when it fails to identify the dialogue partner. As mentioned above, the face image recognition PC 116 shown in Figure 1 performs the identification of the dialogue partner. The face image recognition PC 116 notifies the integrated control PC 122 of the identifier of the identified dialogue partner, its confidence level, and the identifiers of several people who have a low confidence level but are still likely to be the dialogue partner.
[0037] Referring to Figure 3, this program includes a step 360 to receive the result of face image recognition and a step 362 to branch the control flow according to whether the received confidence level is greater than a predetermined threshold. In this case, the threshold cannot be determined in general, as it also depends on the accuracy of face image recognition by the face image recognition PC 116. It is desirable to adjust this threshold based on the actual results of face image recognition.
[0038] This program further includes, in response to the determination in step 362 being affirmative, step 364, which involves reading information about the person corresponding to the identifier identified as the conversation partner from the person information DB92, and confirming the other party's name by speaking the name contained in that information with a confident attitude (facial expression), and step 366, which involves speech recognition of the other party's response to the utterance in step 364, determining whether the other party's response is affirmative or not, and branching the control flow according to the determination result. If the determination in step 366 is affirmative, it means that the identification process of the conversation partner was successful, and that the conversation partner is an acquaintance of the robot. Therefore, in step 368, the robot 110 classifies the conversation partner as an acquaintance and starts a conversation with the other party according to a scenario that has been prepared in advance for conversations with acquaintances. At this time, it is preferable to control the robot 110 to greet with a smile.
[0039] This program further includes step 370, which is executed when the determination in step 362 is negative (i.e., the confidence level is below the threshold) or when the determination in step 366 is negative (i.e., the identification process is incorrect and the other party is not the person identified), in which the program utters a question asking whether the conversation partner is meeting the robot 110 for the first time, in an attitude that appears less confident than when the confidence level is above the threshold; and step 372, which performs speech recognition of the conversation partner's response to the question in step 370 and branches the control flow according to whether the conversation partner's response is affirmative or negative. If the determination in step 372 is affirmative, it means that the conversation partner and the robot 110 are meeting for the first time; if it is negative, it means that the conversation partner and the robot 110 must have interacted before.
[0040] This program further includes step 374, in response to the negation of the determination in step 372, a step in which the robot makes an utterance to the conversation partner to confirm whether the name of the person identified by facial image recognition in step 360 matches the name of the conversation partner, and step 376, which performs speech recognition of the conversation partner's response to the utterance in step 374 and branches the control flow according to whether the response is affirmative or negative. If the determination in step 376 is affirmative, it means that the result of the initial identification process was correct. Therefore, the control proceeds to step 368, classifying the conversation partner as an acquaintance and controlling the robot to start a conversation with the acquaintance. In this case, in step 374, it is preferable to make the robot have an unsure expression. By doing so, from the perspective of the conversation partner, it appears as if the robot is estimating who the conversation partner is. As a result, there is an effect of making the conversation partner feel the intelligence of a humanoid robot.
[0041] This program further includes step 378, which, in response to a negative determination in step 376, controls robot 110 to make an utterance asking the person it is talking to for their name; and step 380, which performs speech recognition on the other party's response to the utterance in step 378, determines whether the name of the other party included in the speech recognition result exists in the person information DB 92, and branches the control flow according to the result. If the determination in step 380 is positive, then this person is an acquaintance of robot 110. Therefore, the control proceeds to step 368. If the determination in step 380 is negative, then although this person should be an acquaintance of robot 110, there is no information about them in the person information DB 92. Therefore, in this program, in step 382 and subsequent processing, robot 110 is controlled to perform processing to gather information about the other party with an apologetic expression.
[0042] The program further includes step 384, which controls the robot 110 to make an utterance asking the other party for their name if the determination in step 372 is affirmative, and step 386, which determines whether the name of the person being spoken to, identified by speech recognition in response to the other party's response, matches the name of one of several people who were listed as candidates in the facial image recognition process in step 360, and branches the control flow according to the determination result.
[0043] If the result of the judgment in step 386 is negative, control proceeds to step 394, where the computer classifies the conversation partner as a stranger (an unknown person). After this, the computer controls the robot 110 to perform a conversation using a script that has been prepared in advance for a conversation with a stranger.
[0044] The program further includes step 388, which controls robot 110 to make a statement to the effect that robot 110 feels like it has met the conversation partner before, with a somewhat uncertain expression, in response to the judgment result of step 386 being affirmative; step 390, which branches the control flow depending on whether the conversation partner's response to this statement is affirmative or not; and step 392, which controls robot 110 to make a happy expression and classifies the conversation partner as an acquaintance of robot 110, if the judgment in step 390 is affirmative. After step 392, robot 110 is controlled to perform a conversation with an acquaintance, similar to step 368. If the judgment in step 390 is negative, the control proceeds to step 394, where the conversation partner is classified as someone being met for the first time.
[0045] In step 372, the other party acknowledges that this is their first encounter with robot 110. Nevertheless, steps 388 and 390 are used to confirm whether the other party has met robot 110 before. This is because some people may find it cumbersome to explain to a robot that they have met before when the other party does not recognize them, and therefore avoid doing so. By including processes like those in steps 388 and 390, it is expected that the conversation partner will feel that robot 110 remembers them and develop a sense of familiarity with robot 110.
[0046] In this way, when the robot fails to identify its conversation partner, it can extract information from the conversation partner to recover from the failure by engaging in a dialogue that includes the expression of emotions. Based on this information, the robot can recover from the failure in a natural way and initiate a dialogue appropriate to the category of the conversation partner.
[0047] B. Recovery process from speech conflicts A common communication failure in dialogue is a clash of utterances. Dialogue between people proceeds with the speaker and listener taking turns. The order in which one person speaks is called the turn of utterance. Normally, the turn of utterances changes naturally, supposedly because some kind of turn-changing rule exists between the speaker and the listener. However, under certain conditions, the turn-changing fails, and the two begin speaking almost simultaneously. This is a clash of utterances.
[0048] It can be assumed that there is an implicit understanding between the two parties in the dialogue that the speaker has the right to speak until it is their turn. This right is also called the right to speak, but in this specification it is called the right to make a statement.
[0049] Speech clashes during turn changes are primarily considered to be the listener's fault. In this embodiment, we consider cases where the robot is responsible for the speech clash and cases where the dialogue partner is responsible. The former is illustrated in Figure 4, and the latter in Figure 5.
[0050] Referring to Figure 4, we will explain a speech collision for which the robot is responsible. In the dialogue partner's speech turn 400, after the dialogue partner makes an utterance 410, while maintaining the speech turn, the robot attempts to make the next utterance 412 after a brief pause. The robot misunderstands this brief pause as the end of the speech turn and attempts to make an utterance 414. As a result, the beginning of utterance 414 and the beginning of utterance 412 overlap in time, causing a speech collision 416.
[0051] Referring to Figure 5, a speech collision for which the dialogue partner is responsible is the opposite situation to that shown in Figure 4. Specifically, after the robot makes utterance 440 within utterance turn 430, it maintains the utterance turn, pauses briefly, and then begins the next utterance 442. The dialogue partner misunderstands this brief pause as the end of the utterance turn and begins the next utterance 444. As a result, the beginning of utterance 442 and the beginning of utterance 444 overlap in time, causing a speech collision 446.
[0052] It should be noted that the mere temporal overlap of utterances between speakers in a dialogue does not automatically constitute a speech conflict. Typically, one speaker may interject with an acknowledgment while the other is speaking. Such interjections should not be considered a speech conflict. Similarly, one speaker may begin speaking before the other has completely finished speaking. In this case, too, if the time between the end of the first speaker's turn and the end of their turn is short, it should not be considered a speech conflict. These issues must be considered when detecting speech conflicts.
[0053] Using the graphs representing the example program shown in Figures 6 and 7, we will explain the situations in which speech collisions are likely to occur and the target intervals (collision detection intervals) for detecting speech collisions. Figures 6 and 7 are the same graph. Figure 6 shows an example of a situation in which a speech collision for which the robot is responsible is likely to occur, and Figure 7 shows an example of a situation in which a speech collision for which the conversation partner is responsible is likely to occur.
[0054] Referring to Figure 6, this graph shows that the starting block on the left is followed by two utterance blocks and two question blocks in that order. Following the question blocks is a speech recognition block that performs speech recognition on the other party's utterance. Following the speech recognition block are three paths. One of these paths is selected based on the response of the conversation partner recognized in the speech recognition block.
[0055] Each of these three paths contains two consecutive utterance blocks. The end of each of these three paths merges into a question block. Following the question block, there is another speech recognition block. Speech recognition of the dialogue partner's utterance in the speech recognition block selects the topic of the dialogue from either topic 1, topic 2, or topic 3, and the execution of this program ends.
[0056] Refer to Figure 6 to see an example of a situation where a speech collision for which the robot is responsible is likely to occur. When the robot operates according to the graph shown in Figure 6, a speech collision is likely to occur immediately after the dialogue partner's speech turn 460. Therefore, a collision detection section 464 is required for speech collisions, enclosing the beginning of the speech block following the dialogue partner's speech turn 460. On the other hand, there is no speech block for the robot after the dialogue partner's speech turn 462. Therefore, a collision detection section for speech collisions is not required for speech turn 462.
[0057] Refer to Figure 7 to see an example of a situation where a speech collision for which the dialogue partner is responsible is likely to occur. When the robot speaks according to the graph shown in Figure 7, three speech blocks follow immediately after the start block on the graph. These three speech blocks constitute the robot's speech turn 480. In the robot's speech turn 480, there is a break in the speech after each speech block. In such areas, the dialogue partner may mistakenly perceive it as the end of the robot's speech turn and speak. As a result, a speech collision for which the dialogue partner is responsible is likely to occur. Therefore, within the speech turn 480 shown in Figure 7, the region including the boundaries of each block is grouped together and designated as the collision detection interval 484.
[0058] Following the robot's speech turn 480, there is a speech recognition block. This section is the dialogue partner's speech turn. Furthermore, there is a robot speech turn 482 that includes the end of this speech recognition block and three paths that are in parallel with each other. Speech collisions for which the dialogue partner is responsible are likely to occur at the beginning of each speech block within this speech turn 482. Therefore, the section including these collisions is collectively designated as the collision detection section 486.
[0059] In this embodiment, overlapping utterances occurring outside of such collision detection intervals are not considered speech collisions. Of course, such overlapping utterances may be treated as speech collisions.
[0060] Figure 8 shows the control structure of a program that enables recovery from a speech collision for which the robot is responsible. Referring to Figure 8, the program includes a step 600 in which the robot interrupts its speech, and a step 602 in which the robot is controlled to make a predefined facial expression (e.g., a surprised expression, a confused expression) that indicates the robot has noticed that a speech collision has occurred. In step 602, the robot may be controlled to make a sound such as "ah".
[0061] The program further includes step 604, following step 602, in which the robot performs a process to relinquish the right to speak to the conversation partner, and step 606, following step 604, in which it determines whether or not another collision has occurred, and if a new collision has occurred, it returns control to step 600, otherwise it terminates this return process.
[0062] Step 604 includes, following step 602, step 610 which determines whether or not there was an interruption in the dialogue partner's utterance and branches the control flow according to the determination result; step 612 which, in response to the determination in step 610 being affirmative, adjusts the utterance turn by yielding the right to speak to the dialogue partner; step 614 which, in response to the determination in step 610 being affirmative and the execution of step 612 being completed, or in response to the failure in step 610 being denied, waits until the end of the dialogue partner's utterance turn is detected; and step 616 which, in response to the end of the dialogue partner's utterance turn being detected, retransmits the information the robot was about to utter and terminates the processing of step 604.
[0063] The reason for determining whether the conversation partner's utterance was interrupted in step 610 is as follows: If the conversation partner does not interrupt their utterance, the robot can recover from the utterance conflict simply by interrupting its own utterance. Therefore, in this case, it is sufficient for the robot to make a gesture indicating that it is relinquishing the right to speak without making any particular utterance. In some cases, even a gesture may not be necessary. Therefore, the conversation can be quickly restored without the robot making any particular utterance.
[0064] On the other hand, if the conversation partner interrupts, the robot needs to more politely relinquish the right to speak to the conversation partner, since it was originally their turn to speak. In this case, the robot should more politely use utterances and gestures to clearly indicate that it is relinquishing the right to speak to the other party. For example, the robot might say "Please go ahead" while extending its hand towards the conversation partner to encourage them to speak. Of course, this is just one example, and many other utterances and gestures can be considered to restore the conversation.
[0065] By performing this process, it is possible to achieve robot behavior that allows for natural dialogue while maintaining the pace of the conversation between the robot and its conversation partner.
[0066] Figure 9 shows the control structure of a program for achieving dialogue recovery (return) from a speech conflict for which the dialogue partner is responsible. Referring to Figure 9, this program includes step 630, which interrupts the robot's speech in response to the detection of a speech conflict; step 632, which controls the robot to make an expression indicating that the robot has noticed the speech conflict; and step 634, which branches the control flow to steps 636 and 638 depending on whether the value of the robot's willingness to speak, which is set as one of the parameters of the robot's emotion in the currently executing speech block, is greater than a predetermined threshold. Step 636 is the process for the robot to relinquish the right to speak to the other party, and step 638 is the process for the robot to maintain its turn to speak. In this embodiment, the value of the robot 110's willingness to speak is set by the system designer when defining the relevant speech block during scenario creation, for example. Of course, the value of the willingness to speak may be set by some means during scenario execution based on other conditions.
[0067] Step 636 includes step 650, which branches the control flow depending on whether or not there was an interruption in the dialogue partner's utterance, and step 652, which adjusts the utterance turn in response to the determination in step 650 being positive. The processing performed in step 652 is the same as the processing performed in step 612 in Figure 8. Step 636 further includes step 654, which waits until the dialogue partner's utterance turn ends in response to the determination in step 650 being positive and step 652 ending, or in response to the determination in step 650 being negative, and step 656, which, in response to the dialogue partner's utterance turn ending, retransmits the information that the robot was trying to convey by utterance when an utterance collision occurred and ends step 636.
[0068] The program further includes step 640, which, in response to the completion of step 636, determines whether a further collision has occurred, returns control to step 630 if the determination is positive, or terminates execution of the program if the determination is negative, indicating that the repair of the dialogue from the speech collision has been completed.
[0069] Step 638 includes step 660, in which a filler is expressed; step 662, following step 660, in which the robot adjusts the turn of speech by informing the other party that it will maintain its turn; and step 664, in which the robot retransmits the information it was trying to convey through speech when a speech collision occurs, and terminates the execution of this program. The filler in step 660 may be a meaningless sound such as "um." The utterance made in step 662 may be anything that makes it clear that the robot will maintain its turn of speech, such as "Let me speak first" or "May I speak first?"
[0070] Figure 10 shows the control structure of a program for determining whether or not a speech collision occurred between the robot and its interaction partner. This program is activated at each time point in the robot's control loop, for example, every 100 milliseconds.
[0071] Referring to Figure 10, this program includes a step 700 that branches the control flow according to a determination of whether or not the robot is recovering from a collision; a step 704 that branches the control flow according to a determination of whether or not the robot's utterance and the dialogue partner's utterance overlap when the determination in step 700 is negative; and a step 706 that branches the control flow according to a determination of whether or not the robot's utterance is an affirmative response when the determination in step 704 is positive. In step 706, the determination of whether or not the robot's utterance is an affirmative response can be made by referring to a dictionary created in advance by collecting utterance texts that are considered affirmative responses. The same applies when the speaker is the dialogue partner.
[0072] If the determination in step 700 is positive, if the determination in step 704 is negative, and if the determination in step 706 is positive, control proceeds to step 702. In step 702, it is determined that no speech collision has occurred, and the processing for when no speech collision has occurred is executed.
[0073] The determination in step 700 indicates that speech collisions will not be detected during recovery from a collision. Furthermore, the determination in step 704 indicates that a speech collision will not occur if the utterances of the robot and the person speaking to each other do not overlap in time. In addition, the determination in step 706 indicates that if the robot's utterance is an acknowledgment, it will not be detected as a speech collision even if it overlaps with the other person's utterance.
[0074] The program further includes step 708, in response to a negative determination in step 706, branching the control flow according to whether the robot retains the right to speak after the robot's last utterance; and step 728, in response to a positive determination in step 708, branching the control flow according to whether the robot's predicted utterance length is greater than or equal to a threshold T1 milliseconds. If the determination in step 728 is negative, it is concluded that no speech collision occurs (step 736).
[0075] The predicted utterance length of the robot, as used here, refers to the maximum length of the robot's current utterance. Even if the robot were to deliver its entire current utterance, if its length is sufficiently short, it will not actually result in an utterance collision. This is why the determination in step 728 is made.
[0076] The program further includes step 730, in response to the determination in step 728 being affirmative, branching the control flow according to whether the robot's current utterance position is a predetermined length portion at the end of the utterance or otherwise, and step 732, in response to the determination in step 730 being negative, i.e., determined to be the beginning of the utterance or the main body of the utterance, branching the control flow according to whether the time during which the robot's utterance and the dialogue partner's utterance overlap is greater than or equal to a threshold T2 milliseconds.
[0077] Even if the robot's predicted utterance length is somewhat long, if the current utterance position is at the end of the utterance, it is not practically necessary to consider it an utterance collision. This is why the determination in step 730 is made. Also, the determination in step 708 indicates that the robot retains the right to speak. Therefore, if the robot's utterance position is not at the end of the utterance, and the overlap time of the utterances exceeds a threshold, it can be concluded that an utterance collision caused by the dialogue partner has occurred. This is why the determination in step 732 is made. For this reason, this program further includes step 734, which determines that if the determination in step 732 is affirmative, an utterance collision caused by the dialogue partner has occurred.
[0078] The program further includes step 736, which determines that no speech conflict occurs if the determination in step 728 is negative, the determination in step 730 is positive, or the determination in step 732 is negative.
[0079] The program further includes step 710, which, in response to a negative determination in step 708, branches the control flow according to whether the dialogue partner's utterance is an affirmative response or not; step 712, which, in response to a negative determination in step 710, branches the control flow according to whether the robot's predicted utterance length is greater than or equal to a threshold T1 milliseconds; step 714, which, in response to a positive determination in step 712, branches the control flow according to whether the robot's current utterance position is at the end, beginning, or main body of the utterance; and step 718, which, in response to a determination in step 714 that the robot's current utterance position is at the main body of the utterance, branches the control flow according to whether the utterance overlap time is greater than or equal to a threshold T2 milliseconds.
[0080] If the conversation partner's utterance is merely an acknowledgment, then even if the robot's utterance and the conversation partner's utterance overlap, it is not necessary to consider that utterance overlap has occurred. The determination in step 710 is for this purpose.
[0081] If, in step 714, the robot's current utterance position is determined to be at the end of the utterance, the control proceeds to step 702 and concludes that no utterance collision has occurred. If, in step 714, the robot's current utterance position is determined to be at the beginning of the utterance, the control proceeds to step 716 and concludes that an utterance collision caused by the robot has occurred. If the determination in step 718 is affirmative, the control proceeds to step 720 and concludes that an utterance collision caused by the dialogue partner has occurred. If the determination in step 718 is negative, the control proceeds to step 722 and concludes that no utterance collision has occurred.
[0082] If the determination in step 708 is negative, the robot does not have the right to speak. Therefore, if in step 712 the robot's predicted utterance length is longer than the threshold and the current utterance position is at the beginning of the utterance, it can be concluded that an utterance collision caused by the robot has occurred. Also, if in step 714 it is determined that the robot's current utterance position is at the end of the utterance, it can be considered that no utterance collision has actually occurred. Furthermore, if the robot's current utterance position is in the body of the utterance and an utterance overlap occurs, it can be determined that this occurred when the dialogue partner started speaking. Therefore, if the overlap time is greater than or equal to the threshold, it can be determined that an utterance caused by the dialogue partner has occurred; otherwise, it can be considered that the dialogue partner ended their utterance immediately, and it can be determined that no utterance collision has actually occurred.
[0083] This program further includes a step 724 that branches the control depending on whether the robot's utterance is a statement indicating the end of a conversation, if the determination in step 712 is negative. If the determination in step 724 is positive, the control proceeds to step 722, and it is concluded that no speech conflict has occurred. If the determination in step 724 is negative, the control proceeds to step 726, and it is concluded that a speech conflict caused by the robot has occurred. An example of an utterance indicating the end of a conversation in step 724 would be a farewell greeting such as "goodbye." Also, threshold values such as 1300 milliseconds for threshold T1 and 1700 milliseconds for threshold T2 may be used.
[0084] Figure 11 shows an example of the beginning 752, body 754, and end 756 of a single robot utterance 750. In the example shown in Figure 11, the beginning 752 is the first 500 milliseconds from the start of the utterance. The end 756 is the last 500 milliseconds from the end of the predicted utterance length. The body 754 is the remaining portion of the utterance 750. Of course, this is just one example, and the lengths of the beginning 752 and end 756 are not limited to those shown in Figure 11. Also, the lengths of the beginning 752 and end 756 do not need to be the same.
[0085] 2. Effects As described above, according to this embodiment, even when the robot fails to identify its conversation partner or when a speech conflict occurs during a conversation, it can repair the conversation and return to a normal conversation by following a set procedure. In this case, the robot uses facial expressions and appropriate gestures to perform the repair dialogue, so from the perspective of the conversation partner, it has the effect of restoring communication in a natural way, similar to when talking to a human. Furthermore, when a speech conflict occurs, unlike in the past, the robot does not always relinquish the right to speak to the conversation partner. By setting the robot's willingness to speak, if the speech conflict is the fault of the conversation partner, depending on the robot's willingness to speak, the robot may either maintain the right to speak and continue speaking, or relinquish the right to speak to the conversation partner. This behavior can be evaluated as being closer to that of a human compared to conventional technology, and has the effect of allowing the conversation partner to return to a normal conversation from a speech conflict in a natural way.
[0086] Second second embodiment When the robot according to the second embodiment fails to identify its interaction partner, it executes the program shown in Figure 12, which represents the control structure, instead of the program executed by the robot according to the first embodiment shown in Figure 3.
[0087] Referring to Figure 12, the program executed by the robot according to the second embodiment differs from that shown in Figure 3 in that, when the determination in step 380 is affirmative, control does not immediately proceed to step 368, but rather the determination in step 770 is made to make the identification of the person more certain. In step 770, it is determined whether the attributes of the person described in the person information retrieved in step 380 from the person information DB 92 match the attributes of the other party. For example, if the record in the person information DB 92 includes the person's gender and date of birth, this information is compared with the gender and age estimated from the face image that was the target of recognition in step 360. In this process, it is not possible to determine whether there is a perfect match or not, but it is possible to calculate the probability (likelihood) that the person in the face image matches the gender recorded in the person information DB 92 and the age calculated from the date of birth by using a trained neural network.
[0088] If this likelihood is above a certain threshold, the decision in step 770 is affirmative; otherwise, it is negative. If the decision in step 770 is affirmative, control proceeds to step 368 and classifies the recognized person as an acquaintance. Otherwise, control proceeds to step 382 and classifies the person as someone who should be an acquaintance but for whom there is no corresponding record in the person information DB92.
[0089] Even if facial image recognition fails to identify the other party, and the name the other party gave in response to the question in step 370 is found in the person information DB92, it does not necessarily mean that the person is the same person as the one recorded in the person information DB92. Inserting the determination in step 770 has the effect of more accurately determining whether or not the person being spoken to is the same person as the one recorded in the person information DB92.
[0090] Third Embodiment The third embodiment, like the second embodiment, differs from the first embodiment in that it executes a program showing the control structure in Figure 13 instead of the program shown in Figure 3 of the first embodiment.
[0091] The program shown in Figure 13 differs from the program shown in Figure 3 in that, when the determination in step 380 is affirmative, control is not immediately transferred to step 368, but rather, as in the second embodiment, a process is provided to make the identification of the dialogue partner more reliable.
[0092] This program includes processing to identify the conversation partner as accurately as possible when multiple records of a person with the same name as the conversation partner are found in the person information DB92.
[0093] More specifically, in addition to the steps shown in Figure 3, this program further includes step 800, which, in response to a positive determination in step 380, branches the control flow depending on whether multiple records of a person with the name given by the conversation partner are found in the person information DB92. If the determination in step 800 is negative, i.e., if only one record is found, the control proceeds to step 368, as in the first embodiment, and the conversation partner is classified as an acquaintance.
[0094] The program further includes step 802 to identify the conversation partner using the record retrieved in the person information DB92, in response to the determination in step 800 being negative, i.e., that only one record was retrieved.
[0095] Step 802 includes step 820, which, in response to the determination in step 800 being positive, i.e., that multiple records were found, executes the following step 822 for multiple records until a predetermined termination condition is met. The termination condition here is that either a record containing information matching the conversation partner is found among the multiple records, or no record containing information matching the conversation partner is found. Step 822 includes step 840, which uses information other than the name from the record being processed to confirm whether the conversation partner is the person recorded in that record, and step 842, which branches the control flow according to whether the determination in step 840 is positive or negative. For example, if the person information DB92 records the affiliation of each person, the robot asks the conversation partner a question such as "Are you Mr. / Ms. B from Department A?" in step 840. If the conversation partner responds positively to this question, the determination in step 842 becomes positive. If the conversation partner responds negatively to this question, the determination in step 842 becomes negative.
[0096] If the determination in step 842 is positive, it is determined that the person being spoken to is the person recorded in that record. Therefore, control proceeds from step 822 to step 368, where the person being spoken to is classified as an acquaintance. If the determination in step 842 is negative, the robot performs the same process using the next record. If the person cannot be identified after performing steps 840 and 842 for all records, control proceeds to step 382. In step 382, the person being spoken to is classified as an acquaintance who is not recorded in the person information DB92, and the conversation begins.
[0097] As described above, according to this third embodiment, even if there are multiple records in the person information DB92 corresponding to the name given by the conversation partner in step 378, if there is a matching person among them, that person can be identified. The possibility of misidentifying a person can be reduced. As a result, the likelihood of the subsequent conversation proceeding smoothly is increased.
[0098] Fourth Embodiment The fourth embodiment is characterized in that, instead of the program shown in Figure 9 of the first embodiment (a program executed by the robot system 100 when a speech collision for which the dialogue partner is responsible is detected), the robot system executes a program whose control structure is shown in Figure 14 to control the robot's movements.
[0099] Referring to Figure 14, the program shown in Figure 14 differs from the program shown in Figure 9 in that it includes an additional step 900 between step 632 and step 634, which branches the control flow depending on whether the utterance collision is the first collision or not. If the determination in step 900 is affirmative, control proceeds to step 636; if the determination in step 900 is negative, control proceeds to step 634. Steps 634 onward are the same as those shown in Figure 1.
[0100] By including step 900, the following effects can be obtained. In the current dialogue between the dialogue partner and the robot, when a speech conflict occurs for the first time due to the dialogue partner's fault, the judgment in step 900 will always be affirmative. Therefore, the robot will always relinquish the right to speak to the dialogue partner. However, in the case of subsequent utterances, the robot will either maintain the right to speak or relinquish it to the other party depending on its willingness to speak. For example, in a dialogue between humans, when a speech conflict occurs, one may relinquish the right to speak to the other party even if it is not their fault. This behavior is thought to be based on the expectation that the other party will do the same when the opposite situation occurs. Such behavior is considered to be very human-like. In this embodiment, by having the robot perform such behavior, the dialogue with the robot becomes closer to a human-like and more natural dialogue from the perspective of the dialogue partner.
[0101] Fifth: Implementation by Computer Figure 15 is an external view of a computer system operating as, for example, the integrated control PC 122 shown in Figure 1. Figure 16 is a hardware block diagram of the computer system shown in Figure 15. The speech recognition PC 118, speech synthesis PC 120, face image recognition PC 116, and motion control PC 112 shown in Figure 1 can also be realized by a computer system with a configuration almost identical to that of the integrated control PC 122. Therefore, only the configuration of the integrated control PC 122 will be described here, and the details of the configurations of the other PCs will not be repeated.
[0102] Referring to Figure 15, this computer system 950 includes a computer 970 having a DVD (Digital Versatile Disc) drive 1002, and a keyboard 974, a mouse 976, and a monitor 972, all connected to the computer 970, for interacting with an intermediary. Of course, these are just one example of a configuration for when intermediary interaction is required, and any general hardware and software (e.g., touch panels, voice input, pointing devices in general) that can be used for interacting with an intermediary to operate the system can be used. These are unnecessary when such intermediary interaction is not anticipated.
[0103] Referring to Figure 16, the computer 970 includes, in addition to the DVD drive 1002, a CPU (Central Processing Unit) 990, a GPU (Graphics Processing Unit) 992, a bus 1010 connected to the CPU 990, GPU 992, and DVD drive 1002, and a ROM (Read-Only Memory) 996 connected to the bus 1010 for storing the computer 970's boot-up program and the like.
[0104] Computer 970 further includes a RAM (Random Access Memory) 998 connected to the bus 1010 for storing program instructions, system programs, and work data, and a non-volatile memory SSD (Solid State Drive) 1000 connected to the bus 1010. The SSD 1000 is for storing programs executed by the CPU 990 and GPU 992, as well as data used by programs executed by the CPU 990 and GPU 992. Computer 970 further includes a network I / F (Interface) 1008 that provides connection to a network 986 (network 114 shown in Figure 1) that enables communication with other terminals, and a USB port 1006 that allows a USB (Universal Serial Bus) memory 984 to be attached and detached, and provides communication between the USB memory 984 and various parts within Computer 970.
[0105] The computer 970 is further connected to the bus 1010 with external devices such as a microphone 982, a speaker 980, and a camera (not shown), and various actuators of a robot, and includes an input / output interface 1004 for input and output between internal components such as the CPU 990 and the external devices.
[0106] In the above embodiment, programs that implement functions such as the operation control PC 112, integrated control PC 122, speech recognition PC 118, speech synthesis PC 120, and face image recognition PC 116 are all stored in storage media of external devices (not shown) connected via network I / F 1008 and network 986, for example, as shown in Figure 16: SSD 1000, RAM 998, DVD 978, USB memory 984, or network I / F 1008 and network 986. Typically, this data and parameters are written to the SSD 1000 from an external source and loaded into the RAM 998 when the computer 970 is running.
[0107] The computer programs for operating this computer system, including the operation control PC 112, integrated control PC 122, speech recognition PC 118, and speech synthesis PC 120 shown in Figure 1, as well as each of these components, are stored on a DVD 978 inserted into the DVD drive 1002 and transferred from the DVD drive 1002 to the SSD 1000. Alternatively, these programs are stored on a USB memory 984, and the USB memory 984 is inserted into the USB port 1006 to transfer the programs to the SSD 1000. Alternatively, these programs may be transmitted to a computer 970 via the network 986 and stored in the SSD 1000.
[0108] The program is loaded into RAM998 when executed. Of course, the source program may be input using the keyboard974, monitor972, and mouse976, and the compiled object program may be stored in SSD1000. In the case of a scripting language as in the above embodiment, the script entered using the keyboard974, etc., may be stored in SSD1000. In the case of a program that runs on a virtual machine, the program that functions as a virtual machine must be installed on the computer970 in advance. Neural networks are used for facial image recognition, speech recognition, and speech synthesis. Neural networks are also used in programs that estimate the gender and age of the person being spoken to from a facial image. For these, a neural network trained by another system may be used, or training may be performed in the robot system 100.
[0109] The CPU990 reads the program from RAM998 according to the address indicated by an internal register called the program counter (not shown), interprets the instructions, reads the data necessary for executing the instructions from RAM998, SSD1000, or other devices according to the address specified by the instructions, and executes the processing specified by the instructions. The CPU990 stores the execution result data at the address specified by the program, such as RAM998, SSD1000, or a register within the CPU990. Depending on the address, the data may be output from the computer as commands to the robot's actuators, audio signals, etc. At this time, the value of the program counter is also updated by the program. The computer program may also be loaded directly into RAM998 from DVD978, USB memory 984, or via network 986. In addition, some tasks (mainly numerical calculations) within the program executed by the CPU990 are dispatched to the GPU992 according to the instructions included in the program or according to the analysis results when the CPU990 executes the instructions.
[0110] The program that enables the computer 970 to implement the functions of each part in each of the embodiments described above includes a plurality of instructions written and arranged to operate the computer 970 to implement those functions. Some of the basic functions necessary to execute these instructions may be provided by the operating system (OS) running on the computer 970, third-party programs, modules of various toolkits installed on the computer 970, or the program execution environment. Therefore, this program does not necessarily have to include all the functions necessary to implement the system and method of this embodiment. This program only needs to include instructions that execute the operations of each of the above-described devices and their components by statically linking or dynamically calling appropriate functions or modules in a controlled manner to obtain the desired results. The method of operating the computer 970 for this purpose is well known and will not be repeated here.
[0111] Furthermore, the GPU992 is capable of parallel processing, allowing it to execute large amounts of calculations associated with machine learning concurrently, in parallel, or in a pipelined manner. For example, parallel computation elements discovered in the program during compilation, or during program execution, are dispatched from the CPU990 to the GPU992 as needed, executed, and the results are returned to the CPU990 either directly or via a predetermined address in RAM998, and assigned to a predetermined variable in the program.
[0112] 6. Other Variations In the above embodiment, the robot 110 operates according to a specific scenario. However, this invention is not limited to such embodiments. Even in embodiments where the robot 110 selects its own actions each time, rather than following a specific scenario, the method according to the above embodiment can be used when engaging in dialogue with another party. Furthermore, in the above embodiment, this invention is not only applicable to cases relating to misidentification of the dialogue partner at the beginning of the dialogue and to speech conflicts during the dialogue. In cases where the dialogue breaks down because one party misunderstands the other's speech, or when the robot's partner does not recognize the robot as a dialogue partner, the robot can be given emotions to respond in the same way as above, allowing for a return to natural dialogue between a person and a robot. Moreover, in the above embodiment, the robot, a physical entity, was one of the parties to the dialogue. However, this invention is not limited to such embodiments. That is, the robot in this invention is not limited to a physical entity. This invention can also be applied to images that mimic the form of a human, such as so-called avatars.
[0113] The embodiments disclosed herein are illustrative and not limited to those embodiments. The scope of the present invention is defined by the claims, with reference to the detailed description of the invention, and includes all modifications within the meaning and scope equivalent to the wording contained herein. [Explanation of symbols]
[0114] 60 Cameras 62,980 speakers 66,982 microphones 92 Person information DB 100 Robot Systems 110 Robots 112 Operation Control PC 114,986 networks 116 Face Image Recognition PC 118 Voice Recognition PC 120 Voice Synthesis PC 122 Integrated Control PC 150 graphs 360, 362, 364, 366, 368, 370, 372, 374, 376, 378, 400, 430, 460, 462 Speech turn 412, 414, 440, 442, 444 utterances 416, 446 Utterance Conflict 480, 482 Robot speech turn 950 Computer Systems 970 Computer
Claims
1. A failure detection step in which the computer detects a failure in communication between the robot and its interaction partner, The computer, in response to the detection of the failure in the failure detection step, controls the robot to engage in an emotionally expressive dialogue with the dialogue partner in accordance with a predetermined procedure, and uses the information obtained in the dialogue to recover from the failure, the computer recovers from the failure. The aforementioned failure detection step is: The computer selectively performs the following steps: controlling the robot to make a first utterance to confirm the result of the identification process with the dialogue partner using a pre-prepared first attitude, depending on whether the confidence level in the identification process of the dialogue partner is higher than a predetermined threshold; and controlling the robot to make a second utterance to start the identification procedure of the dialogue partner using a second attitude that is pre-prepared to appear less confident than the first attitude. A method for recovering from a failure in a dialogue, comprising the steps of: performing a process to control the robot such that, in response to the dialogue partner's response to the first utterance, the robot makes a third utterance to initiate the identification procedure with a second attitude.
2. The step of performing the aforementioned restoration is: The computer, in response to the interaction partner's response to the first utterance indicating that the identification result is correct, classifies the interaction partner as an acquaintance of the robot; The method according to claim 1, further comprising the step of controlling the robot to initiate a conversation according to a pre-prepared scenario for a conversation with an acquaintance.
3. The method according to claim 1 or claim 2, wherein the second utterance and the third utterance are the same utterance.
4. The method according to any one of claims 1 to 3, wherein the second utterance is an utterance by the robot in which it asks whether the person it is talking to is meeting the robot for the first time.
5. The step of performing the aforementioned restoration further includes: The computer determines whether the dialogue partner's response to the second utterance affirms that the dialogue partner is meeting the robot for the first time, In response to the other party's response in the determination step being negative, the computer controls the robot to make a fourth utterance regarding whether or not the other party is the person identified by the identification process, using a third attitude that has been prepared in advance to appear even less confident than the second attitude. The computer controls the robot to classify the conversation partner as an acquaintance of the robot in response to the conversation partner's affirmative response to the fourth utterance, and to begin the conversation with a pre-prepared fourth attitude that appears to show relief. The method according to claim 2, further comprising the step of controlling the robot, in response to the dialogue partner's negative response to the fourth utterance, to perform an additional identification process by displaying a pre-prepared fifth attitude that appears disappointed.
6. The aforementioned additional identification process is The steps include: controlling the robot so that the computer speaks a question to the person it is talking to, asking for their name; The computer generates a determination result by determining whether the name included in the response of the person in the conversation to the question asking for the name matches the name of a person registered in a pre-prepared person information database. The steps include: the computer, in response to the judgment result being positive, classifying the conversation partner as an acquaintance of the robot, and controlling the robot to start a conversation according to a scenario for conversation with an acquaintance while displaying a fifth attitude that has been prepared in advance to appear happy; The method according to claim 5, further comprising the steps of: the computer classifying the conversation partner as an unknown person to the robot in response to the determination result being negative, and controlling the robot to initiate a conversation with the conversation partner according to a pre-prepared scenario for a conversation with an unknown person.
7. The aforementioned additional identification process is The steps include: controlling the robot so that the computer speaks a question to the person it is talking to, asking for their name; The computer generates a determination result by determining whether the name included in the response of the person in the conversation to the question asking for the name matches the name of a person registered in a pre-prepared person information database. The computer, in response to the positive determination result, performs a process to confirm whether the conversation partner is the same person as a person registered in the person information database, and, according to the result of the confirmation, classifies the conversation partner as either an acquaintance or an unknown person to the robot. The steps include controlling the robot so that, in response to the conversation partner being classified as an acquaintance of the robot, it begins a conversation according to a scenario for conversation with an acquaintance while displaying a fifth attitude that has been prepared in advance to appear pleased, The method according to claim 5, further comprising the step of controlling the robot to initiate a conversation with the conversation partner in accordance with a pre-prepared scenario for a conversation with an unknown person, in response to the determination result being negative or the conversation partner being classified as an unknown person to the robot.
8. The step of performing the aforementioned restoration further includes: The steps include: controlling the robot by the computer to make a fifth utterance to identify the conversation partner in response to the second or third utterance being affirmative; The computer generates a determination result regarding whether the information identifying the dialogue partner included in the dialogue partner's response to the fifth utterance matches the result of the identification process. The steps include: controlling the robot so that, in response to the judgment result being positive, the computer makes a sixth utterance to confirm that the conversation partner is an acquaintance of the robot; The method according to claim 5, further comprising the steps of: the computer classifying the conversation partner as an acquaintance to the robot in response to the conversation partner's affirmative response to the sixth utterance; and controlling the robot to initiate a conversation according to a scenario for conversation with an acquaintance while displaying a fifth attitude that has been pre-prepared to appear pleased.
9. The method of claim 8, wherein the recovery step further includes the step of the computer classifying the conversation partner as an unknown person to the robot in response to the determination result being negative, and controlling the robot to initiate a conversation with the conversation partner according to a pre-prepared scenario as a conversation with an unknown person.
10. The method according to claim 8 or 9, wherein the recovery step further includes the step of the computer classifying the conversation partner as an unknown person to the robot in response that the conversation partner's response to the sixth utterance is negative, and controlling the robot to initiate a conversation with the conversation partner according to a pre-prepared scenario as a conversation with an unknown person.
11. A computer program that causes a computer to function to perform the method described in any one of claims 1 to 10.
Citation Information
Patent Citations
Estimation method and estimation system
JP2018077791A
Communication device, communication robot, and communication control program
JP2019000937A
Utterance timing determination device, robot, utterance timing determination method and program
JP2019113696A
Identification device, robot, identification method, and program
JP2020057300A
Identification device, robot, identification method, and storage medium
US20200110968A1