Interactive system and control method

The dialogue system autonomously learns from operator corrections to adapt dialogue in real-time, addressing the limitations of conventional systems by allowing continuous interaction and effective customer engagement.

WO2025143222A1PCT designated stage expired Publication Date: 2025-07-03CYBER AGENT +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/046392
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-28
Filing Date
2024-12-27
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Conventional dialogue systems with autonomous dialogue agents fail to adapt in real-time to changing environments and customer interactions, requiring manual intervention for repetitive corrections, which disrupts the dialogue and may lead to customer dissatisfaction.

Method used

A dialogue system that includes operation intention estimation and action control units to autonomously learn from operator corrections, allowing real-time adjustment of dialogue and behavior through an operator's intended actions, thereby maintaining continuous interaction.

Benefits of technology

Enables simultaneous dialogue expression and correction work without interrupting the conversation, enhancing the system's adaptability and user engagement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024046392_03072025_PF_FP_ABST
    Figure JP2024046392_03072025_PF_FP_ABST
Patent Text Reader

Abstract

This interactive system comprises: one or more interactive agents capable of interacting with a person; an operational intention estimating unit for estimating an operational intention of an operator in accordance with an operation for modifying behavior control information in which behavioral content of the one or more interactive agents is described; and a behavior control unit for controlling interaction by the one or more interactive agents in accordance with the operational intention of the operator estimated by the operational intention estimating unit. 
Need to check novelty before this filing date? Find Prior Art

Description

Dialogue system and control method

[0001] This application claims priority to Japanese Patent Application No. 2023-222213, filed on December 28, 2023, the contents of which are incorporated herein by reference.

[0002] Conventionally, systems have been proposed in which a remote operator operates one or more autonomous dialogue agents to engage in dialogue with a person (see, for example, Non-Patent Documents 1 and 2). In such systems, if a dialogue between one or more autonomous dialogue agents and a person fails, the operator can intervene to continue the dialogue. However, in such systems, the autonomous dialogue agents operate based on preset information, and therefore, if the same interaction as the previous one occurs, the same failure will occur. Therefore, the operator must intervene and address the same failure each time.

[0003] Therefore, systems that autonomously learn from the behavior of an operator have been proposed (see, for example, Non-Patent Documents 3 and 4). Such systems enable an autonomous dialogue agent to improve its processing performance by learning according to the behavior of the operator.

[0004] Dylan F. Glas, et. al, “Teleoperation of Multiple Social Robots”, IEEE Transactions on Systems, Man, and Cybernetics - Part A: Systems and Humans, Vol:42, Issue: 3, May 2012.Kawahara T, et. al, “Semi-autonomous avatar enabling unconstrained parallel conversations - seamless hybrid conversations of WOZ and autonomous dialogue systems”, ADVANCED ROBOTICS2021, VOL. 35, NO. 11, 657-663.M. Doering, et. al, “Modeling Interaction Structure for Robot Imitation Learning of Human Social Behavior”, IEEE Transactions on Human-Machine Systems, Vol:49, Issue:3, June 2019.Malcolm Doering, et. al, “Data-Driven Imitation Learning for a Shopkeeper Robot with Periodically Changing Product Information”, ACM Transactions on Human-Robot Interaction, Vol:10, Issue:4, Article No.:31, pp 1-20.

[0005] In the above-mentioned learning system, learning is performed at predetermined timing, so real-time learning is not possible. In actual customer service, the status of facilities and products changes over time, so real-time intervention and correction of the behavior of the autonomous dialogue agent are necessary. However, with conventional technology, it was sometimes impossible to simultaneously express dialogue to customers and make corrections when intervening.

[0006] In view of the above circumstances, an object of the present invention is to provide a technique that allows both expression of dialogue and correction work.

[0007] A first aspect of the present invention is a dialogue system comprising one or more dialogue agents capable of dialogue with a person, an operational intention estimation unit that estimates an operator's operational intention in response to an operation that modifies behavior control information that describes the behavior of the one or more dialogue agents, and a behavior control unit that controls dialogue by the one or more dialogue agents in response to the operator's operational intention estimated by the operational intention estimation unit.

[0008] In a second aspect of the present invention, in the dialogue system of the first aspect, the one or more dialogue agents carry out dialogue according to the operator's operational intention.

[0009] A third aspect of the present invention is a dialogue system according to the first or second aspect, further comprising a behavior control information modification control unit that, when the behavior control information is modified by the operator, provides the modified behavior control information to a control device that controls the one or more dialogue agents.

[0010] A fourth aspect of the present invention is a dialogue system according to any one of the first to third aspects, wherein when an operation to modify the behavior control information is performed, such as inputting a character string, confirming a character string, operating a mouse on an interface, or registering or deleting character string information, the operation intention estimation unit estimates the operator's operation intention based on the performed operation or a combination of the performed operation and other information.

[0011] In a fifth aspect of the present invention, in the dialogue system of any one of the first to fourth aspects, the behavior control information modification control unit provides the modified behavior control information to a control device that controls the one or more dialogue agents in real time or at a specific timing.

[0012] A sixth aspect of the present invention is a control method for inferring an operator's operational intention in response to an operation to modify behavior control information that describes the behavior of one or more interactive agents that can interact with a person, and for controlling a dialogue between the one or more interactive agents in response to the inferred operator's operational intention.

[0013] The present invention makes it possible to simultaneously express dialogue and perform correction work.

[0014] 1 is a diagram illustrating an example of a configuration of a dialogue system in an embodiment. FIG. 1 is a diagram illustrating an example of a focus word table in an embodiment. FIG. 2 is a diagram illustrating an example of a behavior control information table in an embodiment. FIG. 3 is a diagram illustrating an example of a behavior candidate table in an embodiment. FIG. 4 is a diagram illustrating an example of an utterance information table in an embodiment. FIG. 5 is a diagram illustrating an example of an operation intention utterance information table in an embodiment. FIG. 6 is a diagram illustrating a first example of a display image (image of an operator interface). FIG. 7 is a diagram illustrating a second example of a display image (image of an operator interface). FIG. 8 is a diagram illustrating a second example of a display image (image of an operator interface). FIG. 9 is a sequence diagram illustrating a flow of processing performed by a dialogue system in an embodiment. FIG. 10 is a diagram for explaining an overview of a method for recovering from a dialogue failure and correcting behavior by an interactive agent in an embodiment. FIG. 11 is a diagram illustrating a third example of a display image (image of an operator interface). FIG. 12 is a diagram illustrating a fourth example of a display image (image of an operator interface). FIG. 13 is a sequence diagram illustrating a flow of processing performed by a dialogue system in an embodiment.

[0015] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.

[0016] FIG. 1 is a diagram showing an example of the configuration of a dialogue system 1 according to an embodiment. The dialogue system 1 is a system for grasping the dialogue state of an autonomously controlled dialogue agent, and for recovering from dialogue failures and correcting the behavior of the dialogue agent. Here, the dialogue agent refers to an entity that dialogues on behalf of a user, such as a robot, a two-dimensional agent, or a voice assistant. The two-dimensional agent is, for example, a character displayed on the screen of a display device. The voice assistant is, for example, a microphone speaker. In the following explanation, a case where the dialogue agent is a robot will be described as an example.

[0017] The dialogue system 1 is used to provide product guidance in a commercial facility with multiple stores, such as a department store or a shopping mall. The dialogue system 1 may be used in accommodation facilities such as hotels, or in other facilities that provide guidance to people. In the following explanation, an example will be given in which the dialogue system 1 is applied to a commercial facility. In this case, an interactive agent is installed in the commercial facility where products are sold. The explanation will be given assuming that the operator who operates the interactive agent is located in a location different from the location where the robot is installed (for example, another location in the facility, an office, or a home).

[0018] The dialogue system 1 has two features: a method for grasping the dialogue state of an autonomously controlled dialogue agent, and a method for recovering from dialogue failures and correcting the behavior of the dialogue agent. As the first feature, a method for grasping the dialogue state of an autonomously controlled dialogue agent, the dialogue system 1 expresses the state of autonomous control by imitating the human perceptual process. It is known that humans perceive things in the following sequence: "Select" ⇒ "Organize" ⇒ "Interpret." "Select" refers to selecting an area to focus on, "Organize" refers to organizing components from the selected area, and "Interpret" refers to comparing the organized components with one's own memory to make a judgment. For example, taking a photograph of an animal as an example, a person selects the animal as a region to focus on in the photograph. Next, the person organizes the components, such as the ears, nose, and mouth, of the selected animal. Then, the person determines which animal is in the photograph by comparing them with one's own memory.

[0019] In the dialogue system 1, this concept is used as the perceptual expression of the dialogue agent. More specifically, first, the person or object that the dialogue agent is focusing on and words related to the perceptual expression of the dialogue agent are highlighted on the screen viewed by the operator. Next, executable action candidates (e.g., option display, product information, etc.) based on the highlighted words are displayed on the screen viewed by the operator. After that, the action that the dialogue agent has selected to actually perform is highlighted on the screen viewed by the operator. Then, the action that the dialogue agent actually performed is displayed on the screen viewed by the operator. This allows the operator to quickly grasp the dialogue state by simultaneously viewing the video and audio.

[0020] As a method for recovering from dialogue failures and correcting behavior using a dialogue-type agent, which is the second feature, the dialogue system 1 infers the operator's operational intention based on the operator's operations during the correction work, and causes the operator to make an utterance according to the inferred operational intention. This eliminates the need for the operator to engage in dialogue regarding the correction work, and allows the operator to concentrate on the correction work. A specific configuration for realizing the above two features will be described below.

[0021] (Method for understanding the dialogue state of an autonomously controlled dialogue-type agent) The dialogue system 1 comprises a terminal device 100, a control device 200, a relay server 300, a robot 400, a microphone / speaker 500, a camera 600, a light-emitting unit 700, and a display unit 750.

[0022] The terminal device 100, the control device 200, and the relay server 300 are all connected to each other so as to be able to communicate with each other via a network 800. The network 800 is, for example, a local area network (LAN), a wide area network (WAN), or the Internet. The network 800 may be a network using wireless communication or a network using wired communication. The network 800 may be configured by combining multiple networks. The network 800 may also be a closed communication network such as a virtual private network (VPN).

[0023] Note that network 800 is merely a specific example of a network for realizing communication between devices, and other configurations may be adopted as the network for realizing communication between devices. For example, communication between specific devices may be realized using a network different from the network used for communication between other devices. Specifically, communication between terminal device 100 and relay server 300 may be realized using a network different from the network used for communication between control device 200 and relay server 300.

[0024] 1 , the robot 400, microphone / speaker 500, camera 600, light-emitting unit 700, and display unit 750 are all connected to the control device 200, but the present invention is not limited to this connection configuration. For example, at least one of the robot 400, microphone / speaker 500, camera 600, light-emitting unit 700, and display unit 750 may be communicably connected to the control device 200 via a network 800, or the control device 200, microphone / speaker 500, camera 600, light-emitting unit 700, and display unit 750 may be integrated with the robot 400.

[0025] Some or all of the functional units of the terminal device 100 and the control device 200 are realized as software by a processor such as a CPU (Central Processing Unit) executing a program stored in at least one of the storage units having a non-volatile storage medium (non-transitory storage medium). The program may be recorded on a computer-readable storage medium. Examples of computer-readable storage media include portable media such as a flexible disk, a magneto-optical disk, a ROM (Read Only Memory), and a CD-ROM (Compact Disc Read Only Memory), and non-transitory storage media such as a hard disk built into a computer system.

[0026] Some or all of the functional units of the terminal device 100 and the control device 200 may be realized using hardware including electronic circuits (electronic circuits or circuitry) using, for example, an LSI (Large Scale Integrated circuit), an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), or an FPGA (Field Programmable Gate Array).

[0027] The terminal device 100 is a device operated by an operator. By operating the terminal device 100, the operator can interact with a person via the robot 400. Furthermore, by operating the terminal device 100, the operator can modify behavior control information describing the behavior of the robot 400. Furthermore, the operator can understand the conversation situation of the robot 400 based on image information including a video of the location where the robot 400 is located and sensory information transmitted from the control device 200. Here, the sensory information includes the person detection result, information indicating a person or word to be highlighted, a character string of the person's utterance, a character string indicating candidate actions that the robot 400 can perform, a character string indicating an action actually performed by the robot 400, and the like. The terminal device 100 is configured using an information processing device such as a personal computer, a tablet computer, or a server.

[0028] The control device 200 is a device that controls the operation of one or more robots 400. While Fig. 1 shows a configuration in which one robot 400 is connected to one control device 200, multiple robots 400 may be connected to one control device 200, or one or more robots 400 may be connected to each of multiple control devices 200. The control device 200 is configured using an information processing device such as a personal computer, a tablet computer, or a server.

[0029] The relay server 300 is implemented with a relay function for relaying communication between the terminal device 100 and the control device 200. The relay function may be implemented in the relay server 300 by hardware or by installing software. The relay server 300 executes, for example, WebRTC (Web Real-Time Communication) signaling to realize transmission and reception of image signals and audio signals between the terminal device 100 and the control device 200. The relay server 300 may also function as a WebSocket server for transmitting control information or angle information from the terminal device 100 to the control device 200. The relay server 300 is configured using an information processing device such as a personal computer, an industrial computer, or a server.

[0030] The robot 400 is a conversational agent capable of interacting with humans. The robot 400 performs predetermined actions by controlling various drive mechanisms, light-emitting units, speakers, cameras, and other functions provided in the robot 400 in accordance with control information transmitted by the control device 200. For example, the robot 400 operates by activating drive mechanisms provided at each joint in the neck, shoulders, or arms. The robot 400 may be in the form of an animal that walks by activating drive mechanisms provided at each joint in the shoulders or legs. The robot 400 may be a bipedal robot (humanoid) that walks independently by activating drive mechanisms provided at each joint in the shoulders or legs. The robot 400 may be a mobile robot (agentized robot) that can move on wheels or tracks. The robot 400 may be placed on a flat platform, such as a table or a reception desk.

[0031] The control model of the robot 400 may be a state transition model, a model learned using deep learning, slot filtering, or any other dialogue control model.

[0032] The microphone / speaker 500 is composed of a microphone and a speaker. The microphone / speaker 500 is placed near the robot 400. As a result, the microphone / speaker 500 acquires sounds from the surrounding area where the robot is installed. The microphone / speaker 500 picks up, for example, sounds spoken by a customer facing the robot 400. The microphone / speaker 500 generates an audio signal based on the picked-up sounds. The microphone / speaker 500 outputs the generated audio signal to the control device 200. The microphone / speaker 500 outputs an audio signal transmitted from the terminal device 100. With this configuration, it is possible to make the robot 400 appear to be speaking to the person facing the robot 400. The microphone / speaker 500 may output audio in response to an operation instruction issued by the control unit 203 of the control device 200.

[0033] The camera 600 is a moving image capturing device. The camera 600 is disposed, for example, behind the robot 400. The camera 600 captures moving images of the periphery (vicinity) of the robot 400 so that the back of the robot 400 is captured. That is, the camera 600 captures images of the robot 400 and the user from a third-person perspective. As a result, the camera 600 generates an image (first image) capturing the periphery of the location where the robot 400 is installed. The camera 600 generates an image signal representing the captured moving image and outputs it to the control device 200. The operator of the robot 400 sees the robot 400 and the user from this third-person perspective. The camera 600 may be disposed in any position, such as in the front, back, left, or right, of the robot 400. The camera 600 may be disposed in any position as long as it can capture an image of at least a person (customer) facing the robot 400. For example, the camera 600 may be installed in a position where it can capture first-person perspective images, or it may be installed in a position where it can capture third-person perspective images. As a position where it can capture third-person perspective images, for example, the camera 600 may be placed behind the robot 400. In the case of third-person perspective images, an image including the robot 400 and a person (customer) facing the robot 400 is captured.

[0034] The light-emitting unit 700 is a light-emitting element such as an LED or a light bulb. The light-emitting unit 700 may be, for example, a panel having a predetermined shape and provided with a plurality of light-emitting elements. The predetermined shape may be a circle or a polygon such as a rectangle. The light-emitting unit 700 may emit light in response to an operation instruction issued by the control unit 203 of the control device 200. Note that the light-emitting unit 700 may emit light in different modes depending on the execution content. For example, when the light-emitting unit 700 has a plurality of light-emitting elements, it may be configured to cause some of the light-emitting elements to emit light in response to an operation instruction.

[0035] The display unit 750 is a device that displays guidance information, etc. to a person facing the display unit 750. The display unit 750 is disposed, for example, next to the robot 400.

[0036] Next, we will explain the specific configurations of the terminal device 100 and the control device 200. First, we will explain the configuration of the terminal device 100. The terminal device 100 includes a communication unit 101, an input unit 102, a display unit 103, a microphone 104, a camera 105, a storage unit 106, and a control unit 107.

[0037] The communication unit 101 is a communication device such as a network interface. The communication unit 101 is communicably connected to the network 800 using a predetermined protocol. The communication unit 101 communicates with other devices via the network 800 in accordance with the control of the control unit 107. The other devices in this embodiment are the control device 200 and the relay server 300, but may be devices other than the control device 200 and the relay server 300.

[0038] The input unit 102 is configured using input devices such as a keyboard, a pointing device (a mouse, a tablet, etc.), a button, a touch panel, etc. The input unit 102 is operated by the operator when inputting the operator's instructions to the terminal device 100. The input unit 102 may be an interface for connecting the input device to the terminal device 100. In this case, the input unit 102 inputs an input signal corresponding to the input operation by the operator to the terminal device 100.

[0039] The display unit 103 is an image display device such as a liquid crystal display, an organic EL (Electro Luminescence) display, an electrophoresis display, or a CRT (Cathode Ray Tube) display. The display unit 103 displays information according to the control of the control unit 107. The display unit 103 may be an interface for connecting an image display device to the terminal device 100. In this case, the display unit 103 generates an image signal for displaying an image according to the control of the control unit 107. The display unit 103 outputs the image signal to an image display device connected to the display unit 103. The display unit 103 is one form of a display device.

[0040] The display unit 103 in this embodiment displays on the screen one or more candidate actions to be executed by the robot 400, based on information transmitted from the control device 200 in response to a video input by the control device 200 or a voice input via the microphone / speaker 500. Furthermore, the display unit 103 in this embodiment displays, among the one or more candidate actions, a candidate action executed by the robot 400 in a different format from the other candidate actions. Furthermore, when a voice input is performed via the microphone / speaker 500, the display unit 103 in this embodiment displays a text string as a speech recognition result of the voice input. Furthermore, the display unit 103 in this embodiment highlights specific words included in the text string that is the speech recognition result. The highlighting display may be achieved by, for example, changing the size of the text, changing the color of the text, or surrounding the text, but other expressions may be used as long as the words can be distinguished from other words. Furthermore, the display unit 103 in this embodiment also displays a voice output corresponding to the action executed by the robot 400. Furthermore, the display unit 103 in this embodiment displays on the same screen at least any combination of the video captured by the camera 600, a figure representing a person or object recognized based on the video captured by the camera 600, a character string representing the speech recognition result of the speech input, a character string representing one or more candidate actions, and a character string representing the content of the speech output in response to the action performed by the robot 400. Here, the figure representing the person or object may have any shape, but will be described below as a rectangle.

[0041] The microphone 104 picks up sounds near the terminal device 100. For example, the microphone 104 acquires the voice of the operator of the terminal device 100. The operator of the terminal device 100 speaks to, for example, a person facing the robot 400. The microphone 104 generates a voice signal based on the acquired voice. The microphone 104 outputs the generated voice signal to the terminal device 100.

[0042] The microphone 104 may be an interface for connecting a sound collection device such as an external microphone to the terminal device 100. In this case, the microphone 104 generates an audio signal from the sound input to the sound collection device and outputs the audio signal to the terminal device 100.

[0043] The camera 105 captures video images of the operator of the terminal device 100 and the vicinity of the operator. The camera 105 may be an interface for connecting another camera to the terminal device 100. In this case, the camera 105 generates an image signal in response to the video images captured by the other camera, and inputs image information based on the image signal to the terminal device 100.

[0044] The storage unit 106 is configured using a storage device such as a magnetic hard disk drive or a semiconductor storage device. The storage unit 106 stores information obtained from the control device 200 and operation intention estimation information. The operation intention estimation information is information that can estimate an operation intention according to an operation by an operator. For example, the operation intention estimation information may be in the form of a table in which operation intentions are associated with results when conditions are satisfied, using combinations of the operator's input operations (text input, scrolling, clicking, etc.) and operation areas (search box, result display area, registration button, etc.) on the correction screen, or may be an estimator obtained by learning.

[0045] When an estimator is used, pairs of an information sequence of an input operation and a label or vector representing an operation intention may be used as a learning dataset for learning using machine learning or deep learning. Hereinafter, the operation intention estimation information is also referred to as an operation intention estimation table, which is information in a table format in which operation intentions are associated as a result when a condition is satisfied, based on a combination of an input operation by the operator and an operation area on the editing screen. The storage unit 106 may store audio information acquired by the microphone 104 and images generated by the camera 105.

[0046] The control unit 107 controls the operation of each unit of the terminal device 100. The control unit 107 is configured using a processor such as a CPU and a RAM (Random Access Memory). The processor executes specific programs, causing the control unit 107 to function as a communication control unit 171, a voice recognition unit 172, a video control unit 173, an information provision unit 174, an operational intention estimation unit 175, and a behavior control information modification control unit 176.

[0047] The communication control unit 171 executes a predetermined communication program to communicate with other devices via the communication unit 101. The communication control unit 171 transmits, for example, at least one of voice information, operation information, or an instruction to modify behavior control information to the control device 200 via the communication unit 101. The operation information is information that represents an estimated operation intention. The communication control unit 171 acquires, for example, voice information or image information transmitted from the control device 200 via the communication unit 101.

[0048] The voice recognition unit 172 executes voice recognition processing. The voice recognition processing is processing for generating a character string based on a voice signal. By executing the voice recognition processing, the voice recognition unit 172 generates a character string based on the voice signal output by the microphone 104. The voice recognition unit 172 may generate the character string using a known method.

[0049] The video control unit 173 communicates with other devices by executing a predetermined video control program. For example, the video control unit 173 transmits an audio signal output by the microphone 104 to the control device 200. The video control unit 173 also receives image information from the control device 200. The video control unit 173 causes the display unit 103 to display an image based on the received image information.

[0050] The image information received from the control device 200 includes at least an image of the periphery of the location where the robot 400 is installed. If there is a person around the location where the robot 400 is installed, the person will be included in the image information received from the control device 200. The image information received from the control device 200 may further include perceptual information such as the person detection result, information indicating a person or word to be highlighted, a character string of the person's speech, a character string indicating candidate actions that the robot 400 can perform, and a character string indicating an action actually performed by the robot 400.

[0051] The information providing unit 174 provides the operator with audio information acquired by the control device 200. The audio information acquired by the control device 200 is information related to audio acquired around the location where the robot 400 is installed. The information providing unit 174 may acquire the audio information transmitted from the control device 200 and provide the audio information to the operator at the appropriate timing.

[0052] The operational intention estimation unit 175 estimates the operational intention of the operator in response to an operation to modify behavior control information describing the behavior of the robot 400. Specifically, the operational intention estimation unit 175 estimates the operational intention of the operator based on at least the operational intention estimation information stored in the storage unit 106 and an operation to modify the behavior control information describing the behavior of the robot 400. The operation to modify the behavior control information includes, for example, inputting a character string, a character string confirmation operation, mouse operation on an interface, and registering or deleting character string information. The character string confirmation operation is, for example, an operation to confirm an input character string, such as selecting a predetermined key (e.g., the Return key or the Enter key) or a predetermined button on an interface (e.g., the Confirm button) after inputting a character string. Note that the operation to modify the behavior control information applies to all operational actions on an interface and is not limited to the above-mentioned operations. When the operational intention estimation information is an estimator, the operational intention estimation unit 175 inputs an information sequence indicating an operation to modify the behavior control information into the estimator, and acquires the operational intention output as an estimation result. When the operational intention estimation information is an operational intention estimation table, the operational intention estimation unit 175 refers to the operational intention estimation table and estimates the operator's operational intention based on the operation to modify the behavior control information and information on the area where the operation was performed (operation area). The operational intention estimation unit 175 outputs the estimation result to the communication control unit 171.

[0053] The behavior control information modification control unit 176 modifies the behavior control information held by the control device 200 in response to an operation by the operator.

[0054] The terminal device 100 may control the robot 400 or a device related to the robot 400. When configured in this manner, the terminal device 100 generates control information for controlling the robot 400 or a device related to the robot 400, and transmits the generated control information to the control device 200. The device related to the robot 400 is, for example, an actuator (motor, etc.), a speaker, and a light-emitting unit.

[0055] Next, a description will be given of the control device 200. The control device 200 includes a communication unit 201, a storage unit 202, and a control unit 203.

[0056] The communication unit 201 is a communication device such as a network interface. The communication unit 201 is communicably connected to the network 800 using a predetermined protocol. The communication unit 201 communicates data with other devices via the network 800 under the control of the control unit 203.

[0057] The storage unit 202 is configured using a storage device such as a magnetic hard disk drive or a semiconductor storage device. A focus word table, a behavior control information table, a behavior candidate table, an utterance information table, and an operation intention utterance information table are stored in the storage unit 202. The focus word table, the behavior control information table, the behavior candidate table, the utterance information table, and the operation intention utterance information table will be described below.

[0058] FIG. 2 is a diagram showing an example of a focus word table in the embodiment. The focus word table registers focus words, which are words that should be noted. The focus words represent words that are to be highlighted among the words that make up the content uttered by a person. The focus words are set in advance in the control device 200 by the operator operating the terminal device 100. In FIG. 2, as an example, words such as "photo", "stamp", and "where is it" are registered as focus words.

[0059] FIG. 3 is a diagram showing an example of a behavior control information table in an embodiment. In the behavior control information table, target words, answer contents, and locations are registered in association with each other. The target words represent search words used during searches. Note that multiple similar words may be registered as target words to make searches easier. The answer contents represent words that are output aloud when a word registered as a target word is searched. The location represents the location within the facility of a product identified by the associated target word. The location information may represent a map section within the facility, or may be a combination of a floor number and a position. Note that not all items need to be registered in the behavior control information table. The behavior control information table is one aspect of behavior control information.

[0060] FIG. 4 is a diagram illustrating an example of a behavior candidate table in an embodiment. In the behavior candidate table, triggers and behavior candidates are registered in association with each other. The trigger represents a condition for determining a behavior candidate of the robot 400. In the example illustrated in FIG. 4, triggers such as "a person being located within a predetermined range" and "information (e.g., voice or operation) being input" are shown. The behavior candidate represents a candidate behavior to be performed by the robot 400 when the condition indicated by the trigger is satisfied. In the example illustrated in FIG. 4, the candidate behaviors include "follow" and "display options and guidance information." In the example illustrated in FIG. 4, when "a person being located within a predetermined range," "follow" is selected as the candidate behavior of the robot 400, and when "information (e.g., voice or operation) is input," "display options and guidance information" is selected as the candidate behavior of the robot 400. Note that the behavior candidate table may register a behavior candidate of the robot 400 for each specific word or combination of specific words.

[0061] FIG. 5 is a diagram illustrating an example of an utterance information table in an embodiment. In the utterance information table, actions and utterance contents are registered in association with each other. The action represents an action that the robot 400 has decided to actually perform from among candidate actions. The utterance contents represent the content that the robot 400 is made to utter according to the decided action. In the example illustrated in FIG. 5, when it is decided to perform the “follow” action, the robot 400 is made to utter contents such as “Shall we talk?” or “Are you looking for something?”. Furthermore, in the example illustrated in FIG. 5, when it is decided to perform the “display options” action, the robot 400 is made to utter contents such as “X number of options have been found. Please touch the item you are looking for.” Furthermore, in the example illustrated in FIG. 5, when it is decided to perform the “display guidance information” action, the robot 400 is made to utter contents such as “If it is “...”, it is on the A floor.” Note that the name of the item selected by the person is entered in “...”.

[0062] FIG. 6 is a diagram showing an example of an operational intention utterance information table in the embodiment. In the operational intention utterance information table, conditions (operation intentions) and utterance contents are registered in association with each other. The condition (operation intention) represents the operational intention estimated by the operational intention estimation unit 175. The utterance contents represent the content to be uttered by the robot 400 when the condition (operation intention) is satisfied. In the example shown in FIG. 6, when the condition "A character string "□□" is input into the text search input area" is satisfied, the robot 400 is caused to utter a content such as "Hmm, no candidates are displayed. I'll search using □□." Furthermore, in the example shown in FIG. 6, when the condition "Mouse operation" is satisfied, the robot 400 is caused to utter a content such as "When I searched, it looked like this..." Furthermore, in the example shown in FIG. 6, when the condition "Click (item selection) item "□□"" is satisfied, the robot 400 is caused to utter a content such as "Is this "□□" wrong? " is input to the robot 400. Note that the name of the item selected by the person is input in "□□".

[0063] Returning to Fig. 1 , the explanation will continue. The control unit 203 controls the operation of each unit of the control device 200. The control unit 203 is configured using a processor such as a CPU and RAM. The processor executes specific programs, causing the control unit 203 to function as a voice recognition unit 241, a video recognition unit 242, a sensor recognition unit 243, a communication control unit 244, a sensory information generation unit 245, and a behavior control unit 246.

[0064] The voice recognition unit 241 executes a voice recognition process. By executing the voice recognition process, the voice recognition unit 241 generates a character string based on the voice signal output by the microphone / speaker 500. The voice recognition unit 241 outputs the generated character string to the sensory information generation unit 245. The voice recognition unit 241 may generate the character string using a known method.

[0065] The video recognition unit 242 acquires an image signal output by the camera 600. The video recognition unit 242 outputs the acquired image signal to the sensory information generation unit 245. The video recognition unit 242 also receives an audio signal from the terminal device 100. The video recognition unit 242 outputs the received audio signal to the microphone / speaker 500.

[0066] The sensor recognition unit 243 acquires a sensor signal detected by a sensor (not shown) (for example, a human detection sensor) and outputs the acquired sensor signal to the perception information generation unit 245.

[0067] The communication control unit 244 communicates with other devices by executing a predetermined communication program. For example, the communication control unit 244 receives at least one of audio information and image information from the terminal device 100 via the relay server 300. For example, the communication control unit 244 transmits image information output from the sensory information generation unit 245 to the terminal device 100 via the relay server 300.

[0068] The sensory information generation unit 245 generates sensory information based on at least one of the character string output from the voice recognition unit 241, the image signal output from the video recognition unit 242, the sensor signal output from the sensor recognition unit 243, and the action determined by the action control unit 246. Specifically, the sensory information generation unit 245 generates, as sensory information, a character string of the person's speech and information indicating words to be highlighted, based on the character string output from the voice recognition unit 241. The sensory information generation unit 245 generates, as sensory information, a person detection result, based on the image signal output from the video recognition unit 242. The sensory information generation unit 245 generates, as sensory information, information indicating a person to be highlighted, based on the sensor signal output from the sensor recognition unit 243. The sensory information generation unit 245 generates, as sensory information, a character string indicating candidate actions that the robot 400 can perform and a character string indicating an action actually performed by the robot 400, based on the action determined by the action control unit 246. The sensory information generation unit 245 generates image information including the generated sensory information and the image signal output from the video recognition unit 242. The sensory information generation unit 245 outputs the generated image information to the communication control unit 244.

[0069] The behavior control unit 246 controls the operations of the robot 400, the microphone / speaker 500, the light-emitting unit 700, and the display unit 750 based on the information stored in the memory unit 202. The behavior control unit 246 causes the robot 400 to perform a specific operation by, for example, outputting control information including an operation instruction to the robot 400. The behavior control unit 246 causes the microphone / speaker 500 to output a character string as a voice by, for example, outputting control information including a character string to be output as a voice to the microphone / speaker 500. The behavior control unit 246 causes the light-emitting unit 700 to emit light by, for example, outputting control information including an illumination instruction to the light-emitting unit 700. The behavior control unit 246 causes the display unit 750 to display information by, for example, outputting control information including information to be displayed to the display unit 750.

[0070] Here, a specific example of the operation of the method for grasping the dialogue state of an autonomously controlled dialogue-type agent, which is a first feature of the dialogue system 1, will be described. FIG. 7 is a diagram showing a first example of a display image (an image of the operator interface). In FIG. 7, four display images are shown in FIGS. 7A to 7D, and each of the display images shown in FIGS. 7A to 7D is displayed on the display unit 103 of the terminal device 100 operated by the operator. In FIGS. 7A to 7D, the flow from detecting a person to selecting an action for the robot 400 will be described.

[0071] The display unit 103 displays the display image shown in FIG. 7A when the control device 200 detects a person from an image captured by the camera 600. The display image shown in FIG. 7A displays at least the robot 400 and identification information inf1 that identifies the person. The identification information inf1 is information that indicates a person detected using an existing person detection method. The identification information inf1 is represented, for example, by a rectangle that surrounds the person. Here, the display unit 103 displays the identification information inf1 with a dotted line, as an example. The identification information inf1 is provided by the control device 200 as perceptual information. By displaying the identification information inf1, the operator can understand that a person has been detected.

[0072] When a person is located within a predetermined range from the position of the robot 400 (e.g., within a range of 1.5 m from the position of the robot 400), the display unit 103 displays the display image shown in FIG. 7B. The display image shown in FIG. 7B displays at least the robot 400 and identification information inf1 and inf2 that identify the person. The identification information inf2 is information that indicates a notable person among the detected people. A notable person is, for example, a person located within a range of 1.5 m from the position of the robot 400. Therefore, when a notable person is detected, the display unit 103 displays the identification information inf2 on the screen, surrounding the detected notable person. The identification information inf2 is represented, for example, by a rectangle that surrounds the person. The identification information inf2 is provided from the control device 200 as perceptual information. Here, as an example, the display unit 103 displays the identification information inf2 in a form different from the identification information inf1 (e.g., a solid line).

[0073] When a notable person is detected, the display unit 103 displays the display image shown in FIG. 7C. In the display image shown in FIG. 7C, in addition to the information displayed in FIG. 7B, an action candidate display area R1 is displayed. The action candidate display area R1 is an area in which one or more action candidates to be executed by the robot 400 are displayed. In the example shown in FIG. 7C, the action candidate display area R1 displays candidate information inf3 indicating one action candidate, "face following." The candidate information inf3 to be displayed by the display unit 103 is provided from the control device 200 as perceptual information. The action candidate display area R1 may be displayed in any area on the screen of the display unit 103, but is preferably displayed in an area in which no person is captured (e.g., the upper part of the screen) so as not to interfere with the operator's viewing.

[0074] When candidate information inf3 indicating a candidate action to be performed by the robot 400 is determined from the candidate information inf3 displayed in the action candidate display area R1, the display unit 103 displays the display image shown in Fig. 7(D). In the display image shown in Fig. 7(D), candidate information inf3 indicating the candidate action "face following" determined to be performed by the robot 400 is displayed in an emphasized (e.g., highlighted) manner. Information indicating the action determined to be performed by the robot 400 is provided as perceptual information from the control device 200. This allows the operator to easily understand which action the robot 400 is currently performing by looking at the screen.

[0075] 8 to 10 are diagrams showing a second example of a display image (an image of an operator interface). Five display images are shown in FIGS. 8A, 8B, 9A, 9B, and 10, and the display images shown in FIGS. 8A, 8B, 9A, 9B, and 10 are displayed on the display unit 103 of the terminal device 100 operated by the operator. FIGS. 8A, 8B, 9A, 9B, and 10 will explain the flow from the speech of a person of interest to the display of the results of the action actually taken by the robot 400.

[0076] When the control device 200 acquires the content of a person's speech in real time through voice recognition, the display unit 103 displays the display image shown in FIG. 8A. The display image shown in FIG. 8A displays a person's speech display area R2 in addition to the information displayed in FIG. 7B. The person's speech display area R2 is an area where the content of a person's speech is displayed. The person's speech display area R2 shown in FIG. 8A displays "User: take a photo." This indicates that the person of interest uttered "take a photo." Note that the content of the speech displayed in the person's speech display area R2 is the content acquired by the control device 200 through voice recognition. Therefore, depending on the accuracy of the voice recognition of the control device 200, the content of the person's speech and the content of the speech displayed in the person's speech display area R2 may differ.

[0077] By displaying the speech content in person speech display area R2, the operator can understand what the person has said. Furthermore, if the speech content of the person differs from the content displayed in person speech display area R2, the operator can easily understand that speech recognition has failed. Person speech display area R2 may be displayed in any area on the screen of display unit 103, but it is desirable to display it in an area where no person is imaged (e.g., the top of the screen) so as not to interfere with the operator's viewing.

[0078] When a specific word is included in the speech content of a notable person, the display unit 103 displays the display image shown in FIG. 8B. In the display image shown in FIG. 8B, the specific word "photo" in the speech content displayed in the person speech display area R2 is displayed as emphasized emphasis information inf4. "User: Whoever puts in a photo" is displayed in the person speech display area R2 shown in FIG. 8B. Then, the display unit 103 displays the specific word "photo" in "User: Whoever puts in a photo" displayed in the person speech display area R2 in a different manner from the other words. Information indicating the word to be emphasized is provided from the control device 200 as perceptual information.

[0079] When the person finishes speaking, the display unit 103 displays the display image shown in FIG. 9A. In the display image shown in FIG. 9A, in addition to the person utterance display area R2, an action candidate display area R1 is displayed. The person utterance display area R2 shown in FIG. 9A displays "User: Where is the one that puts photos?". In the display image shown in FIG. 9A, the specific word "where is it" in the utterance content "User: Where is the one that puts photos?" displayed in the person utterance display area R2 is displayed as emphasized information inf5. Then, the display unit 103 displays the specific words "photo" and "where is it" in the utterance content "User: Where is the one that puts photos?" displayed in the person utterance display area R2 in a manner different from the other words. Furthermore, the display unit 103 displays an action candidate display area R1 below the person utterance display area R2, and in the example shown in Fig. 9A, three action candidates, "Display options," "Guide: Photo frame," and "Guide: Album," are displayed in the action candidate display area R1. These three action candidates are determined according to one specific word or a combination of specific words in the utterance content displayed in the person utterance display area R2, and are provided as sensory information by the control device 200.

[0080] When a candidate action to be performed by the robot 400 is determined from the three candidate actions displayed in the candidate action display area R1, the display unit 103 displays the display image shown in Fig. 9(B). In the display image shown in Fig. 9(B), the candidate action "Option Display" determined to be performed by the robot 400 is displayed in a highlighted manner, unlike the other candidate actions "Guidance: Photo Frame" and "Guidance: Album". Information indicating the action to be performed by the robot 400 is provided by the control device 200 as sensory information. This allows the operator to easily understand which action the robot 400 is currently performing by looking at the screen.

[0081] When the determined candidate action is performed by the robot 400, the display unit 103 displays the display image shown in FIG. 10. In the display image shown in FIG. 10, in addition to the candidate action display area R1 and the person utterance display area R2, an action result display area R3 is displayed. The action result display area R3 is an area where the results of the action actually performed by the robot 400 are displayed. In this embodiment, the action result display area R3 displays the utterance of the robot 400 as a result of the action actually performed by the robot 400. The action result display area R3 shown in FIG. 10 indicates that the robot 400 has uttered the following: "X candidates found. Please touch the one you are looking for." The results of the action actually performed by the robot 400 are provided as perceptual information from the control device 200.

[0082] The action result display area R3 may be displayed in any area on the screen of the display unit 103, but is preferably displayed in an area where no person is imaged (e.g., the top of the screen) so as not to interfere with the operator's viewing. In the example shown in this embodiment, the action candidate display area R1, the person utterance display area R2, and the action result display area R3 are displayed in this order from top to bottom of the screen, but the display order of the action candidate display area R1, the person utterance display area R2, and the action result display area R3 may be any order. By displaying the action candidate display area R1, the person utterance display area R2, and the action result display area R3 in this way, the operator can easily understand, by looking at the screen, what action the robot 400 actually performed in response to the content of the person's utterance.

[0083] 11 and 12 are sequence diagrams showing the flow of processing performed by the dialogue system 1 in the embodiment. The robot 400, microphone / speaker 500, camera 600, and light-emitting unit 700 are collectively referred to as the robot, etc.

[0084] The camera 600 captures an image of the robot 400 and its surroundings (step S101). The camera 600 outputs an image signal representing the captured image to the control device 200 in real time. The video recognition unit 242 of the control device 200 acquires the image signal output from the camera 600. The video recognition unit 242 outputs the acquired image signal to the perceptual information generation unit 245. The perceptual information generation unit 245 generates image information based on the image signal output from the video recognition unit 242. In this case, the perceptual information generation unit 245 does not have any perceptual information to add, so it generates the image signal output from the video recognition unit 242 as image information. The perceptual information generation unit 245 outputs the generated image information to the communication control unit 244.

[0085] The communication control unit 244 transmits the image information output from the perception information generation unit 245 to the terminal device 100 (step S102). The video control unit 173 of the terminal device 100 causes the display unit 103 to display an image based on the image information transmitted from the control device 200 (step S103). Thereafter, the processes from step S101 to step S103 are executed until a person is detected.

[0086] The sensory information generator 245 of the control device 200 detects a person based on the image signal output from the camera 600 (step S104). Existing technology can be applied as a method for detecting a person. When the sensory information generator 245 detects a person, it generates image information by adding specific information inf1 that identifies the detected person to the image signal as sensory information. The sensory information generator 245 outputs the generated image information to the communication controller 244. The communication controller 244 transmits the image information output from the sensory information generator 245 to the terminal device 100 (step S105). The video controller 173 of the terminal device 100 causes the display unit 103 to display an image based on the image information transmitted from the control device 200. As a result, for example, the display unit 103 displays the image shown in FIG. 7A (step S106).

[0087] Thereafter, the sensory information generation unit 245 of the control device 200 detects a person of interest based on the sensor signal output from the sensor recognition unit 243 (step S107). When the sensory information generation unit 245 detects a person of interest, it generates image information by adding, as sensory information, specific information inf2 identifying the detected person of interest to the image signal. The sensory information generation unit 245 outputs the generated image information to the communication control unit 244. The communication control unit 244 transmits the image information output from the sensory information generation unit 245 to the terminal device 100 (step S108). The video control unit 173 of the terminal device 100 causes the display unit 103 to display an image based on the image information transmitted from the control device 200. As a result, for example, the display unit 103 displays the image shown in FIG. 7B (step S109).

[0088] Suppose the person of interest then speaks. In this case, the microphone / speaker 500 collects the speech of the person of interest (step S110). The microphone / speaker 500 outputs a voice signal indicating the collected speech of the person of interest to the control device 200 (step S111). The voice recognition unit 241 of the control device 200 performs voice recognition processing on the voice signal output from the microphone / speaker 500 (step S112). As a result, the voice recognition unit 241 generates a character string based on the voice signal output from the microphone / speaker 500. The voice recognition unit 241 outputs the generated character string to the sensory information generation unit 245. The sensory information generation unit 245 generates image information by adding the character string output from the voice recognition unit 241 to an image signal as sensory information. The sensory information generation unit 245 outputs the generated image information to the communication control unit 244. The communication control unit 244 transmits the image information output from the sensory information generation unit 245 to the terminal device 100 (step S113). The video control unit 173 of the terminal device 100 causes the display unit 103 to display an image based on the image information transmitted from the control device 200. As a result, the display unit 103 displays a display image in which the character string output from the voice recognition unit 241 is superimposed on the person speech display area R2 of the image represented by the image signal (step S114). For example, the display unit 103 displays the image shown in FIG. 8A.

[0089] The perceptual information generation unit 245 of the control device 200 refers to the attention word table and selects a specific word from the character string output from the speech recognition unit 241 (step S115). The perceptual information generation unit 245 generates image information by adding emphasis information inf4, which emphasizes the selected specific word, to the image signal as perceptual information. The perceptual information generation unit 245 outputs the generated image information to the communication control unit 244. The communication control unit 244 transmits the image information output from the perceptual information generation unit 245 to the terminal device 100 (step S116). The video control unit 173 of the terminal device 100 causes the display unit 103 to display an image based on the image information transmitted from the control device 200. As a result, the display unit 103 displays a display image in which the specific word identified by the emphasis information inf4 from the character string output from the speech recognition unit 241 is emphasized and superimposed on the person speech display area R2 of the image represented by the image signal (step S117). As a result, the display unit 103 displays, for example, the image shown in FIG. 8B.

[0090] Next, the behavior control unit 246 of the control device 200 refers to the behavior candidate table and selects a behavior candidate for the robot 400 that corresponds to the selected specific word or combination of specific words (step S118). The behavior control unit 246 outputs information about the selected behavior candidate for the robot 400 to the perception information generation unit 245. The perception information generation unit 245 generates image information by adding the information about the behavior candidate for the robot 400 output from the behavior control unit 246 to an image signal as perception information. The perception information generation unit 245 outputs the generated image information to the communication control unit 244. The communication control unit 244 transmits the image information output from the perception information generation unit 245 to the terminal device 100 (step S119). The video control unit 173 of the terminal device 100 causes the display unit 103 to display an image based on the image information transmitted from the control device 200. As a result, the display unit 103 displays a display image in which the information about the behavior candidate for the robot 400 is superimposed on the behavior candidate display area R1 of the image represented by the image signal (step S120). For example, the display unit 103 displays the image shown in FIG.

[0091] Next, the behavior control unit 246 of the control device 200 selects an action to be actually performed by the robot 400 from the candidate actions selected in the process of step S118 (step S121). The selection of an action can be expressed not only by a model for selecting an action, but also by a model for controlling speech and actions, by generating and listing multiple actions as memories. The behavior control unit 246 outputs information about the candidate actions indicating the selected action to the perceptual information generation unit 245. The perceptual information generation unit 245 generates image information by adding the information about the candidate actions indicating the action output from the behavior control unit 246 to an image signal as perceptual information. The perceptual information generation unit 245 outputs the generated image information to the communication control unit 244. The communication control unit 244 transmits the image information output from the perceptual information generation unit 245 to the terminal device 100 (step S122). The video control unit 173 of the terminal device 100 displays an image based on the image information transmitted from the control device 200 on the display unit 103. As a result, the display unit 103 displays a display image in which the behavior candidate indicating the selected behavior is highlighted (step S123). For example, the display unit 103 displays the image shown in FIG.

[0092] The behavior control unit 246 refers to the utterance information table and acquires the utterance content corresponding to the selected behavior. The behavior control unit 246 generates control information for uttering the acquired utterance content. The control device 200 transmits the generated control information to the microphone / speaker 500 (step S124). For example, the control device 200 transmits control information including a character string indicating the utterance content to the microphone / speaker 500. The microphone / speaker 500 outputs the character string included in the control information transmitted from the control device 200 as voice (step S125). This makes it appear as if the robot 400 has spoken.

[0093] Furthermore, the behavior control unit 246 outputs character string information included in the control information transmitted to the microphone / speaker 500 to the sensory information generation unit 245. The sensory information generation unit 245 generates image information by adding the character string information output from the behavior control unit 246 to an image signal as sensory information. The sensory information generation unit 245 outputs the generated image information to the communication control unit 244. The communication control unit 244 transmits the image information output from the sensory information generation unit 245 to the terminal device 100 (step S126). The video control unit 173 of the terminal device 100 causes the display unit 103 to display an image based on the image information transmitted from the control device 200. As a result, the display unit 103 displays a display image in which a character string representing the behavior performed by the robot 400 is superimposed on the behavior result display area R3 of the image represented by the image signal (step S127). For example, the display unit 103 displays the image shown in FIG. 10 .

[0094] The above-described process is the first feature, which is a method for grasping the dialogue state of an autonomously controlled dialogue-type agent. According to the dialogue system 1 configured as described above, at least one of the following is displayed on the screen viewed by the operator: an image captured by the camera 600, a character string representing the speech recognition result of the voice input (e.g., the content of what the person said), a character string representing a candidate action that is a candidate action to be performed by the robot 400, and a character string representing the content of the voice output in response to the action performed by the robot 400. This makes it possible for the operator to easily grasp the dialogue between the robot 400 and a person.

[0095] 7 to 10, it becomes possible to easily grasp the causes of failure in the dialogue of the robot 400. Possible causes of failure in the dialogue of the robot 400 include the following: Cause of failure 1: The person's voice can be heard, but a character string indicating the voice recognition result is not displayed on the screen of the display unit 103. Cause of failure 2: The voice heard differs from the character string indicating the voice recognition result displayed on the screen of the display unit 103. Cause of failure 3: A character string indicating the voice recognition result is displayed on the screen of the display unit 103, but a specific word is not highlighted. Cause of failure 4: The correct action is not displayed as a candidate action. Cause of failure 5: The action that the robot 400 will actually take is determined, but the robot 400 does not take the action.

[0096] Cause of failure 1 is that the operator receives the person's speech, but the character string indicating the speech recognition result is not displayed in the person speech display area R2 of the display unit 103. In this case, the operator can understand that the speech recognition function of the control device 200 is not functioning. As a result, this can be addressed by performing processing such as restarting the application of the control device 200.

[0097] Cause of failure 2 is that the content of the voice that the operator hears is different from the content displayed in the person speech display area R2 of the display unit 103. In this case, the operator can understand that the voice recognition function of the control device 200 is working, but the voice recognition has failed. As a result, this can be addressed by performing processing to support voice recognition.

[0098] Cause of failure 3 is that the character string displayed in the person speech display area R2 of the display unit 103 is correct, but the word is not highlighted. In this case, the operator can understand that the word to be highlighted has not been registered. As a result, the operator can address this issue by registering the word to be highlighted.

[0099] Cause of failure 4 is that the correct action is not displayed as a candidate action in the candidate action display area R1 of the display unit 103. Here, the correct action is an action that is considered appropriate as an action that can be predicted from the character string obtained by speech recognition. An example of a case where the correct action is not displayed as a candidate action is when a request for a certain product is made and information about a completely unrelated product is displayed as a candidate action. In this case, the operator can determine that the list of candidate actions has failed or that there is insufficient action information. As a result, the problem can be addressed by registering new action information.

[0100] Cause of failure 5 occurs when the robot 400 or the like does not take action even though the action to be performed has been decided. In this embodiment, since actions are mainly performed by voice output, in this case the operator can understand that the microphone / speaker 500 is not functioning. As a result, this can be addressed by performing processing such as restarting the application for the microphone / speaker 500.

[0101] As described above, when a dialogue by the robot 400 fails, the operator can easily identify the cause. Since the cause of the failure can be easily identified in this way, it becomes easier to decide how to deal with the problem. This allows the operator to instantly understand what is happening even if they have little knowledge or experience of robot control or systems. Therefore, it becomes easy to understand and take action to deal with the dialogue failure.

[0102] (Method for recovering from a dialogue failure and correcting behavior by an interactive agent) Next, the second feature, a method for recovering from a dialogue failure and correcting behavior by an interactive agent, will be described. Because an autonomously controlled interactive agent operates based on preset information, if the same interaction as a previous failure occurs, the same failure will occur. Therefore, the operator must intervene and deal with the same failure each time. Furthermore, if the robot is controlled to speak after performing correction work, it is likely that the user will leave. Therefore, in this embodiment, the robot recovers from a dialogue failure and corrects behavior while correcting behavior in real time. More specifically, the robot infers the operator's intention in response to the operator's correction work and makes a speech according to the inferred intention. This eliminates the need for the operator to engage in dialogue to make corrections, allowing the operator to concentrate on the correction work.

[0103] FIG. 13 is a diagram illustrating an outline of a method for recovering from a dialogue failure and correcting behavior by an interactive agent according to an embodiment. Assume that an operator operates the terminal device 100 to perform a correction operation. In this case, the operator operates the display unit 103 to display a correction screen. The correction screen is a screen for correcting behavior control information related to the behavior of the robot 400. Assume that the operator enters the character string "album" into the text search input area on the correction screen as a correction operation. In this case, the operation intention estimation unit 175 of the terminal device 100 estimates an operation intention corresponding to the text input into the text search input area in response to the operator's operation for the correction operation. Then, the terminal device 100 transmits information indicating the estimated operation intention to the control device 200. The control device 200 refers to the operation intention utterance information table and acquires the utterance content (e.g., "Hmm, no candidates are displayed. I'll check the album.") associated with the operation intention transmitted from the terminal device 100. The control device 200 outputs the acquired speech content to the microphone / speaker 500. In this way, even though the operator is making corrections, the robot 400 continues the dialogue by outputting a voice such as "Oh, there are no candidates. I'll check the album." This allows the robot 400 to continue the dialogue while the operator is searching to see if information related to "album" has been registered as part of the correction work.

[0104] As another example, suppose the operator "scrolls the mouse" on the edit screen as a correction operation. In this case, the operational intention estimation unit 175 of the terminal device 100, triggered by the operator's operation for the correction operation, estimates the operational intention corresponding to the mouse scrolling. Then, the terminal device 100 transmits information indicating the estimated operational intention to the control device 200. The control device 200 references the operational intention utterance information table and acquires the utterance content (e.g., "When I searched, it looked like this...") associated with the operational intention transmitted from the terminal device 100. The control device 200 outputs the acquired utterance content via the microphone / speaker 500. In this way, even though the operator is performing the correction operation, the robot 400 continues the dialogue by outputting the content such as "When I searched, it looked like this..." as if the operator were performing the correction operation. This allows the robot 400 to continue the dialogue while the operator is performing the search while scrolling the mouse as a correction operation.

[0105] As another example, suppose the operator selects an "item selection" as an operation for a correction task on the correction screen. For example, suppose the operator selects an album (for photos). In this case, the operational intention estimation unit 175 of the terminal device 100 estimates an operational intention corresponding to the selection of the selected item, triggered by the operator's operation for the correction task. The terminal device 100 then transmits information indicating the estimated operational intention to the control device 200. The control device 200 references the operational intention utterance information table and acquires the utterance content (e.g., "Is this 'Album (for photos)' wrong?") associated with the operational intention transmitted from the terminal device 100. The control device 200 then outputs the acquired utterance content via the microphone / speaker 500. In this way, even though the operator is performing a correction task, the robot 400 continues the dialogue by outputting content such as "Is this 'Album (for photos)' wrong?" as if the operator were actually performing the correction task. This allows the robot 400 to continue the dialogue while the operator is performing a search by scrolling with the mouse as part of the correction task.

[0106] FIG. 14 is a diagram showing a third example of a display image (an image of an operator interface). The display image shown in FIG. 14 is an image on the screen of the display unit 103 of the terminal device 100. The display image shown in FIG. 14 displays a display area R11 and a correction button B1. An image of the robot side is displayed in the display area R11. For example, the display area R11 is an area where image information transmitted from the control device 200 such as those shown in FIGS. 7 to 10 is displayed. The correction button B1 is a button for correcting information stored in the memory unit 202 of the control device 200. By correcting the information stored in the memory unit 202 of the control device 200, the behavior of the robot 400 can be corrected. When the operator selects the correction button B1, the display unit 103 displays the display image shown in FIG. 15.

[0107] Fig. 15 is a diagram showing a fourth example of a display image (image of the operator interface). The display image shown in Fig. 15 displays a display area R11, a modify button B1, and a display area R12. The display area R12 is an area where a modification screen for modifying the behavior of the robot 400 is displayed. Modifying the behavior of the robot 400 means modifying the behavior control information. The modification screen displays a search input area R13, a search result display area R14, a behavior control information input area R15, and a register button B2.

[0108] The search input area R13 is an area for inputting a string when performing a text search. For example, the search input area R13 is used to input a string to be searched for (e.g., a stamp). The search result display area R14 is an area for displaying the results of a search for information corresponding to the string input in the search input area R13. For example, the behavior control information correction control unit 176 refers to target words in the behavior control information table in accordance with the string input in the search input area R13 and searches for a record corresponding to the string input in the search input area R13. If a record corresponding to the string input in the search input area R13 is found, the behavior control information correction control unit 176 displays the information registered in the record corresponding to the string input in the search input area R13 in the search result display area R14. Note that the string input in the search input area R13 may match multiple records. In this case, the behavior control information correction control unit 176 displays the information registered in the multiple records obtained as the search results in the search result display area R14.

[0109] On the other hand, if there is no record corresponding to the character string entered in the search input area R13, the behavior control information correction control unit 176 displays nothing in the search result display area R14, or displays a message in the search result display area R14 indicating that there is no relevant information. In the example shown in Fig. 15, the search result display area R14 displays the results of a search corresponding to the character string (e.g., a stamp) entered in the search input area R13.

[0110] The behavior control information input area R15 is an area used to modify information in the behavior control information table. If information corresponding to the character string entered in the search input area R13 is registered in the behavior control information table, the answer content, target word, and location information registered in the behavior control information table are displayed in the answer content, target word, and location fields shown in the behavior control information input area R15. In the example shown in FIG. 15 , all information is registered in the answer content, target word, and location fields shown in the behavior control information input area R15. However, if some information is not registered in the behavior control information table, some information is not displayed in the answer content, target word, and location fields shown in the behavior control information input area R15. For example, in the behavior control information table shown in FIG. 3 , no answer content is registered for the target word "photo." In this case, among the answer content, target word, and location fields shown in the behavior control information input area R15, nothing is displayed in the answer content field, "photo" is displayed in the target word field, and "AA" is displayed in the location field.

[0111] The register button B2 is a button used when registering information entered in the behavior control information input area R15. When the register button B2 is selected after information has been entered in the behavior control information input area R15, the behavior control information correction control unit 176 registers the information entered in the behavior control information input area R15 in the behavior control information table. Specifically, when the register button B2 is selected after information on the target word and information on the response content have been entered in the behavior control information input area R15, the behavior control information correction control unit 176 newly registers a record corresponding to the character string entered in the target word field of the behavior control information input area R15 in the behavior control information table.

[0112] When the register button B2 is selected after answer information has been entered in the action control information input area R15, the action control information correction control unit 176 first refers to the action control information table based on the character string entered in the target word field of the action control information input area R15 and selects a record corresponding to the character string entered in the target word field. The action control information correction control unit 176 then additionally registers the character string entered in the answer field of the action control information input area R15 in the answer field of the selected record.

[0113] When the register button B2 is selected after location information has been entered in the behavior control information input area R15, the behavior control information correction control unit 176 first refers to the behavior control information table based on the character string entered in the target word field of the behavior control information input area R15 and selects a record corresponding to the character string entered in the target word field. The behavior control information correction control unit 176 then additionally registers the character string entered in the location field of the behavior control information input area R15 in the location field of the selected record.

[0114] The above correction process adds new information to the behavior control information table. If information corresponding to the character string entered in the search input area R13 is registered in the behavior control information table as a result, the behavior control information correction control unit 176 displays the information in the search result display area R14 in real time. This allows the operator to select an item displayed in the search result display area R14.

[0115] The correction screen may be provided with a reflection selection button that allows the user to select whether or not to reflect the corrections in real time. When the reflection selection button is ON, the corrections made by the operator are reflected in real time. That is, when the reflection selection button is ON, the behavior control information correction control unit 176 updates the information registered in the behavior control information table stored in the control device 200 based on the corrections. This reflects the corrections in real time. When the reflection selection button is OFF, the corrections made by the operator are not reflected in real time. Therefore, when the reflection selection button is OFF, the behavior control information correction control unit 176 updates the information registered in the behavior control information table stored in the control device 200 with the corrections at a predetermined time. This allows corrections that are considered to need to be reflected in real time to be reflected in real time, and corrections that do not need to be reflected in real time to be reflected later. This reduces the processing load that would otherwise be incurred by reflecting all corrections.

[0116] Furthermore, the correction screen may include a video call area for making inquiries to staff via video call. The staff refers to staff working at the facility where the robot 400 is installed. This allows the operator to consult with staff in real time and correct the behavior control information even in situations that are difficult for the operator to handle (e.g., when it is unclear where a product is located). It is possible that the customer may be kept waiting while the operator is making a video call with the staff. Therefore, when an operation to make a video call is performed, the operation intention estimation unit 175 estimates the operation intention corresponding to the operation of the video call. Based on the operation intention corresponding to the operation of the video call, the control device 200 may notify the customer that a staff inquiry is being made and perform an operation to make effective use of the customer's waiting time. The operation to make effective use of the customer's waiting time may include, for example, a quiz or introduction of recommended products. In this way, by notifying the customer that a staff inquiry is currently being made and allowing the customer to make effective use of the waiting time, the customer can be kept in the location.

[0117] The screen of the display unit 103 of the terminal device 100 in the above-mentioned Figures 14 and 15 is an example, and various input form formats such as text fields, buttons, check boxes, slide bars, tabular input, click coordinate positions on a diagram, etc. can be used as interfaces on the screen of the display unit 103 of the terminal device 100.

[0118] Fig. 16 is a sequence diagram showing the flow of processing performed by the dialogue system 1 in this embodiment. The control device 200, robot 400, microphone / speaker 500, camera 600, and light-emitting unit 700 are collectively referred to as the robot side. Fig. 16 explains an example of dialogue recovery when an error occurs in voice recognition by the control device 200.

[0119] Assume that a customer utters "Where is XX?" (Step S201). The microphone / speaker 500 collects the customer's speech. The microphone / speaker 500 outputs a voice signal indicating the collected customer speech, "Where is XX?", to the control device 200. The voice recognition unit 241 of the control device 200 performs voice recognition processing on the voice signal output from the microphone / speaker 500. Here, assume that the voice recognition result recognized by the voice recognition unit 241 is "Where is △△?" (Step S202). In this case, the control device 200 selects the item candidate that was found for "△△" (Step S203). The control device 200 generates image information including the selected item candidate. The control device 200 provides the generated image information to the terminal device 100 (Step S204).

[0120] The video control unit 173 of the terminal device 100 displays the image information provided by the control device 200 on the display unit 103. As a result, for example, item candidates that were hit by "△△" are presented on the screen of the display unit 103. The operator sees the screen of the display unit 103 and realizes that voice recognition has failed. The operator operates the terminal device 100 to select the correction button B1. In response to the selection of the correction button B1, the behavior control information correction control unit 176 causes the display unit 103 to display a correction screen. The display unit 103 displays the correction screen under the control of the behavior control information correction control unit 176 (step S205). The operator performs a search by entering the character string "〇〇" in the search input area R13 of the displayed correction screen.

[0121] Through this process, the behavior control information correction control unit 176 refers to the target word item in the behavior control information table stored in the storage unit 202 of the control device 200 and searches for a record corresponding to the character string "XX". If there is a record corresponding to the character string "XX", the behavior control information correction control unit 176 displays information about the record corresponding to the character string "XX" in the search result display area R14. As a result, item information is displayed in the search result display area R14. On the other hand, if there is no record corresponding to the character string "XX", the behavior control information correction control unit 176 displays a message in the search result display area R14 indicating that there are no corresponding item candidates. Here, it is assumed that there are corresponding item candidates. The operator selects an item displayed in the search result display area R14 (step S206).

[0122] The operational intention estimation unit 175 of the terminal device 100 estimates the operational intention of the operator based on the operation performed by the operator. In the example described above, the operator first inputs the character string "XX" into the search input area R13, and then selects an item from the search results displayed in the search result display area R14. The operational intention estimation unit 175 then refers to the operational intention estimation table to acquire the operational intention associated with the operation "text input" to modify the behavior control information and the area (operation area) "search window" where the operation was performed. The acquired operational intention is defined as the first operational intention. Next, the operational intention estimation unit 175 refers to the operational intention estimation table to acquire the operational intention associated with the operation "click" to modify the behavior control information and the area (operation area) "result display area" where the operation was performed. The acquired operational intention is defined as the second operational intention. The operational intention estimation unit 175 transmits operation information including the acquired first and second operational intentions and information on the selected item to the control device 200 (step S207).

[0123] The control device 200 receives the operation information transmitted from the terminal device 100. The control device 200 starts uttering in accordance with the first operation intention and the second operation intention contained in the operation information (step S208). Specifically, the control device 200 generates control information for executing the utterance corresponding to the first operation intention. The control device 200 transmits the selected control information to the microphone / speaker 500. For example, the control device 200 transmits control information including a character string to be output to the microphone / speaker 500. The microphone / speaker 500 outputs the character string contained in the control information transmitted from the control device 200 as voice. For example, the microphone / speaker 500 outputs a voice such as, "We haven't found the product you're looking for yet?"

[0124] Thereafter, the control device 200 generates control information for executing an utterance corresponding to the second operation intention. The control device 200 transmits the selected control information to the microphone / speaker 500. For example, the control device 200 transmits control information including a character string to be output to the microphone / speaker 500. The microphone / speaker 500 outputs the character string included in the control information transmitted from the control device 200 as voice. For example, the microphone / speaker 500 outputs voice such as "Is this XX?"

[0125] Furthermore, the control device 200 causes the display unit 750 to display information about the item included in the operation information. The display unit 750 displays the information about the item under the control of the control device 200 (step S209). Assume that the customer selects an item displayed on the display unit 750 (step S210). The control device 200 causes the display unit 750 to display guidance information for guiding the customer to the location where the selected item is located. The display unit 750 displays the guidance information under the control of the control device 200 (step S211). The guidance information may be, for example, a map showing the location where the selected item is located, or a character string indicating the location where the selected item is located (e.g., second floor).

[0126] According to the dialogue system 1 configured as described above, the operator's operational intention is estimated in response to an operation to modify the behavior control information, and the dialogue by the robot 400 is controlled in accordance with the estimated operator's operational intention. In this way, while the operator is modifying the behavior control information, the control device 200 performs control in accordance with the operator's operational intention. This eliminates the need for the operator to simultaneously express a dialogue and perform a modification. Furthermore, since the dialogue continues even while a modification is being performed, it is possible to reduce the chance of the person with whom the dialogue is taking place leaving. This makes it possible to simultaneously express a dialogue and perform a modification.

[0127] <Modification 1> In the above-described embodiment, a configuration has been described in which it is assumed that the robot 400 is remotely controlled (a configuration in which the sensory information of the robot 400 is displayed on the terminal device 100 located at a location distant from where the robot 400 is installed). In contrast, the present invention is also applicable to a case in which the robot 400 is not remotely controlled and only autonomously interacts. In such a configuration, the sensory information of the robot 400 may be displayed on a display unit provided on the robot 400. With such a configuration, by showing a display according to the sensory information to the user, it is possible to clarify how the robot 400 is moving, and this can be used to gain trust or for maintenance.

[0128] <Variation 2> In the dialogue system 1, if the second feature, which is the method for recovering from dialogue failures and correcting behavior using an interactive agent, is not implemented, the terminal device 100 does not need to be equipped with the operation intention estimation unit 175 and the behavior control information correction control unit 176.

[0129] <Variation 3> In the above-described embodiment, the configuration in which the control device 200 uses each table (attention word table, behavior control information table, action candidate table, utterance information table, and operation intention utterance information table) is merely an example. For example, the configuration in which the control device 200 uses each table is merely an implementation example when a state transition model is used. When a deep learning model, generative AI, or the like is used, each table may not be used. In other words, some or all of each table does not need to be stored in the storage unit 202. For example, the control device 200 may not have a attention word table, but may use a deep learning model that understands utterances and display the weight of the attention mechanism of the deep learning model using color. For example, the control device 200 may not have a behavior candidate table, but may generate one or more action candidates using a generative model and arrange the generated action candidates as candidates in the action candidate display area R1. For example, the control device 200 may not have an utterance information table, but may automatically generate utterance content using a deep learning model. In this way, each table is merely information necessary to drive the dialogue using a state transition model, and if a deep learning model or generative model is used as the dialogue control model, the intermediate products and multiple outputs that result from them may also be displayed.

[0130] When the weights of the attention mechanism are displayed as described above, failure cause 3 in the above-described embodiment can be addressed by noticing that a particular word is not emphasized, and by changing the model, correcting the training data, changing the examples or prompts passed to the model, etc. When a deep learning model or generative AI is used as described above, failure cause 4 in the above-described embodiment can be addressed by changing the model, correcting the training data, changing the examples or prompts passed to the model, etc.

[0131] Although an embodiment of the present invention has been described above in detail with reference to the drawings, the specific configuration is not limited to this embodiment, and includes designs within the scope of the gist of the present invention.

[0132] The present invention can be applied to a technology that uses one or more conversational agents that can converse with a person.

[0133] 1...Dialogue system, 100...Terminal device, 101...Communication unit, 102...Input unit, 103...Display unit, 104...Microphone, 105...Camera, 106...Memory unit, 107...Control unit, 171...Communication control unit, 172...Speech recognition unit, 173...Video control unit, 174...Information provision unit, 175...Operation intention estimation unit, 176...Action control information modification control unit, 200...Control device, 201...Communication unit, 202...Memory unit, 203...Control unit, 241...Speech recognition unit, 242...Video recognition unit, 243...Sensor recognition unit, 244...Communication control unit, 245...Perception information generation unit, 246...Action control unit, 300...Relay server, 400...Robot, 500...Microphone / speaker, 600...Camera, 700...Light emitting unit, 800...Network

Claims

1. One or more dialogue agents capable of interacting with a person, an operation intention estimation unit that estimates an operator's operation intention according to an operation for modifying operation control information describing the action content of the one or more dialogue agents, and an action control unit that controls the dialogue by the one or more dialogue agents according to the operator's operation intention estimated by the operation intention estimation unit. A dialogue system comprising:

2. The one or more dialogue agents perform a dialogue with content according to the operator's operation intention. The dialogue system according to claim 1.

3. When the operation control information is modified by the operator, an operation control information modification control unit that provides the modified operation control information to a control device that controls the one or more dialogue agents. The dialogue system according to claim 1 or 2, further comprising:

4. When any one of character string input, character string confirmation operation, mouse operation on an interface, registration or deletion of character string information is performed as an operation for modifying the operation control information, the operation intention estimation unit estimates the operator's operation intention by the executed operation or a combination of the executed operation and other information. The dialogue system according to claim 1 or 2.

5. The operation control information modification control unit provides the modified operation control information to a control device that controls the one or more dialogue agents in real time or at a specific timing. The dialogue system according to claim 3.

6. A control method for estimating an operator's operation intention according to an operation for modifying operation control information describing the action content of one or more dialogue agents capable of interacting with a person, and controlling the dialogue by the one or more dialogue agents according to the estimated operator's operation intention.

Citation Information

Patent Citations

  • Control device, control method, annotator presentation device, presentation method to annotator, program and communication system

    JP2022158193A