Interaction system and control method

The dialogue system addresses real-time learning and correction challenges by estimating operator intentions to facilitate simultaneous dialogue expression and correction, enhancing interaction quality and reducing operator intervention.

JP2025104424APending Publication Date: 2025-07-10CYBER AGENT +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023222213
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-28
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

Conventional systems with autonomous interactive agents fail to perform real-time learning and dialogue modification, leading to repetitive intervention by operators for dialogue failures, which disrupts customer interactions.

Method used

A dialogue system with dialogue agents that estimate operator intentions through input operations, allowing real-time correction and control of dialogue actions, enabling simultaneous dialogue expression and correction work without operator intervention.

Benefits of technology

Enables seamless dialogue continuation and efficient correction of autonomous agent behaviors in real-time, improving customer interaction quality and reducing operator workload.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025104424000001_ABST
    Figure 2025104424000001_ABST
Patent Text Reader

Abstract

To achieve both representation of an interaction and correction operations.SOLUTION: An interaction system comprises: one or more interaction agents that can interact with a person; an operation intention estimation unit that estimates an operator's operation intention according to an operation to correct behavior control information in which the details of behaviors of the one or more interaction agents are described; and a behavior control unit that controls the interaction performed by the one or more interaction agents according to the operator's operation intention estimated by the operation intention estimation unit.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an interactive system and a control method.

Background Art

[0002] Conventionally, a system has been proposed in which an operator located at a remote location operates one or more autonomous interactive agents to interact with a person (see, for example, Non-Patent Documents 1 and 2). In such a system, when the interaction between one or more autonomous interactive agents and a person fails, the operator can intervene to continue the interaction. However, in such a system, since the autonomous interactive agent operates based on preset information, if the same interaction as the one that has failed once occurs, the same failure will occur. Therefore, the operator needs to intervene and respond to the same failure each time.

[0003] Therefore, a system that autonomously learns from the behavior of an operator has been proposed (see, for example, Non-Patent Documents 3 and 4). With such a system, the autonomous interactive agent can improve its processing performance by learning according to the behavior of the operator.

Prior Art Documents

Non-Patent Documents

[0004]

Non-Patent Document 1

Non-Patent Document 2

[0005] In the system that performs the above learning, since the learning is performed at a predetermined timing, real-time learning is not performed. In actual customer service, since the situation of the facility and products changes over time, real-time intervention and modification of the behavior of the autonomous dialogue agent are required. However, in the conventional technology, there are cases where it is impossible to balance the expression of dialogue with the customer at the time of intervention and the modification work.

[0006] In view of the above circumstances, an object of the present invention is to provide a technology capable of achieving both dialogue expression and correction work.

Means for Solving the Problems

[0007] Aspect 1 of the present invention is a dialogue system including one or more dialogue agents capable of interacting with a person, an operation intention estimation unit that estimates an operator's operation intention according to an operation for correcting action control information in which the action contents of the one or more dialogue agents are described, and an action control unit that controls the dialogue by the one or more dialogue agents according to the operator's operation intention estimated by the operation intention estimation unit.

[0008] Aspect 2 of the present invention is the dialogue system of Aspect 1, in which the one or more dialogue agents conduct a dialogue with contents according to the operator's operation intention.

[0009] Aspect 3 of the present invention is the dialogue system of Aspect 1 or 2, further including an action control information correction control unit that, when the action control information is corrected by the operator, provides the corrected action control information to a control device that controls the one or more dialogue agents.

[0010] Aspect 4 of the present invention is the dialogue system according to any one of Aspects 1 to 3, in which the operation intention estimation unit estimates the operator's operation intention by an operation executed when any one of character string input, character string confirmation operation, mouse operation on an interface, registration or deletion of character string information is performed as an operation for correcting the action control information, or by a combination of the executed operation and other information.

[0011] Aspect 5 of the present invention is the dialogue system according to any one of Aspects 1 to 4, in which the action control information correction control unit provides the corrected action control information to a control device that controls the one or more dialogue agents in real time or at a specific timing.

[0012] Aspect 6 of the present invention is a control method that estimates the operator's operation intention according to an operation for modifying the action control information describing the action content of one or more dialogue agents capable of interacting with a person, and controls the dialogue by the one or more dialogue agents according to the estimated operator's operation intention.

Advantages of the Invention

[0013] According to the present invention, it becomes possible to achieve both dialogue expression and modification work.

Brief Description of the Drawings

[0014]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Embodiment for Carrying Out the Invention

[0015] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.

[0016] FIG. 1 is a diagram showing a configuration example of a dialogue system 1 in an embodiment. The dialogue system 1 is a system for grasping the dialogue state of an autonomously controlled interactive agent and performing recovery of dialogue failure and behavior correction by the interactive agent. Here, the interactive agent means an object that conducts dialogue on behalf of a user such as a robot, an agent represented in two dimensions, and a voice assistant. The agent represented in two dimensions is, for example, a character displayed on the screen of a display device. The voice assistant is, for example, a microphone speaker. In the following description, the case where the interactive agent is a robot will be described as an example.

[0017] The dialogue system 1 is used, for example, in a commercial facility where a plurality of stores such as a department store or a shopping mall are provided, for guiding customers to products. Note that the dialogue system 1 may be used in an accommodation facility such as a hotel or other facilities as long as it is a facility that guides people. In the following description, the case where the dialogue system 1 is applied to a commercial facility will be described as an example. In this case, the interactive agent is installed inside a commercial facility that sells products. The operator who operates the interactive agent is described as being located at a location different from the location where the robot is installed (for example, another location inside the facility, an office, or a home).

[0018] The dialogue system 1 has two features: a method for grasping the dialogue state of a self-controlled dialogue agent, and a method for recovering from dialogue failures and correcting the behavior by the dialogue agent. As the first feature, the method for grasping the dialogue state of a self-controlled dialogue agent, the dialogue system 1 represents the state of self-control by imitating the human perception process. As a human perception process, it is known that humans perceive things in the flow of "Select" ⇒ "Organize" ⇒ "Interpret". "Select" is to select the area to be focused on, "Organize" is to organize the components from the selected area, and "Interpret" is to judge things by comparing the organized components with one's own memory. For example, taking a photo of an animal as an example, a person selects the animal as the area to be focused on from the photo. Next, the person organizes the components such as the ears, nose, and mouth of the selected animal. Then, by comparing with one's own memory, the person judges which animal is shown in the photo.

[0019] The dialogue system 1 uses such a concept as the perceptual representation of the dialogue agent. More specifically, first, the person or object that the dialogue agent is focusing on and the words related to the perceptual representation of the dialogue agent are highlighted on the screen that the operator is looking at. Next, candidates for executable actions (for example, option display, product guidance, etc.) based on the highlighted words are displayed on the screen that the operator is looking at. After that, the action selected as the one that the dialogue agent will actually perform is highlighted on the screen that the operator is looking at. And the action actually performed by the dialogue agent is displayed on the screen that the operator is looking at. Thereby, the operator can quickly grasp the dialogue state by seeing the video and the audio at the same time.

[0020] As a method for recovering from a dialogue failure and correcting behavior by an interactive agent, which is the second feature, in the dialogue system 1, the operation intention of the operator is estimated according to the operation during the operator's correction work, and a speech corresponding to the estimated operation intention is made to be spoken. As a result, the operator does not need to conduct a dialogue in the correction response and can concentrate on the correction work. Hereinafter, a specific configuration for realizing the above two features will be described.

[0021] (Method for grasping the dialogue state of an autonomously controlled interactive agent) The dialogue system 1 includes a terminal device 100, a control device 200, a relay server 300, a robot 400, a microphone - speaker 500, a camera 600, a light - emitting unit 700, and a display unit 750.

[0022] The terminal device 100, the control device 200, and the relay server 300 are all communicably connected via a network 800. The network 800 is, for example, a network such as a LAN (Local Area Network), a WAN (Wide Area Network), or the Internet. The network 800 may be a network using wireless communication or a network using wired communication. The network 800 may be configured by combining a plurality of networks. The network 800 may be a closed - area communication network such as a VPN (Virtual Private Network).

[0023] Note that the network 800 is only a specific example of the network for realizing the communication of each device, and other configurations may be adopted as the network for realizing the communication of each device. For example, the communication between specific devices may be realized using a network different from the network used for the communication between other devices. Specifically, the communication between the terminal device 100 and the relay server 300 may be realized using a network different from the communication between each of the control device 200 and the relay server 300.

[0024] In addition, in FIG. 1, the robot 400, the microphone speaker 500, the camera 600, the light emitting unit 700, and the display unit 750 are all connected to the control device 200, but such a connection form is not limited thereto. For example, at least one of the robot 400, the microphone speaker 500, the camera 600, the light emitting unit 700, and the display unit 750 may be communicably connected to the control device 200 via the network 800, or the control device 200, the microphone speaker 500, the camera 600, the light emitting unit 700, and the display unit 750 may be integrated with the robot 400.

[0025] Some or all of the functional units of the terminal device 100 and the control device 200 are realized as software by a processor such as a CPU (Central Processing Unit) executing a program stored in at least one of the storage units having a non-volatile recording medium (non-temporary recording medium). The program may be recorded on a computer-readable recording medium. A computer-readable recording medium is a non-temporary recording medium such as a portable medium such as a flexible disk, a magneto-optical disk, a ROM (Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or a storage device such as a hard disk incorporated in a computer system.

[0026] Some or all of the functional units of the terminal device 100 and the control device 200 may be realized using hardware including an electronic circuit (electronic circuit or circuitry) using, for example, an LSI (Large Scale Integrated circuit), an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), or an FPGA (Field Programmable Gate Array).

[0027] The terminal device 100 is a device operated by an operator. The operator can interact with a person via the robot 400 by operating the terminal device 100. Further, the operator can modify the action control information describing the action content of the robot 400 by operating the terminal device 100. Further, the operator can grasp the interaction situation of the robot 400 based on the image information including the video and perception information of the location where the robot 400 is located, which is transmitted from the control device 200. Here, the perception information includes the detection result of a person, the information indicating the person or word to be highlighted, the character string of the speech content of the person, the character string indicating the candidates for the actions that the robot 400 can execute, the character string indicating the actions actually taken by the robot 400, and the like. The terminal device 100 is configured using an information processing device such as a personal computer, a tablet computer, or a server.

[0028] The control device 200 is a device that controls the operation of one or more robots 400. In FIG. 1, a configuration is shown in which one robot 400 is connected to one control device 200, but a plurality of robots 400 may be connected to one control device 200, or one or more robots 400 may be connected to each of a plurality of control devices 200. The control device 200 is configured using an information processing device such as a personal computer, a tablet computer, or a server.

[0029] The relay server 300 is implemented with a relay function for relaying communication between the terminal device 100 and the control device 200. The relay function may be implemented in the relay server 300 by hardware or by installing software. The relay server 300 realizes the transmission and reception of image signals and audio signals between the terminal device 100 and the control device 200, for example, by executing WebRTC (Web Real-Time Communication) signaling. Also, the relay server 300 may function as a WebSocket server for transmitting control information or angle information from the terminal device 100 to the control device 200. The relay server 300 is configured using an information processing device such as a personal computer, an industrial computer, or a server.

[0030] The robot 400 is an interaction agent capable of interacting with people. The robot 400 executes a predetermined operation by controlling functions provided in the robot 400, such as each drive mechanism, a light-emitting unit, a speaker, or a camera, according to the control information transmitted by the control device 200. For example, the robot 400 operates by operating drive mechanisms provided at respective joint portions of the neck, shoulders, or arms. The robot 400 may be in the shape of an animal that walks by operating drive mechanisms provided at respective joint portions such as the shoulders or legs. The robot 400 may be a robot (humanoid) that walks autonomously by operating drive mechanisms provided at respective joint portions such as the shoulders or legs. The robot 400 may be a mobile robot (agentified robot) that can move on wheels or an endless track. The robot 400 may be installed, for example, on a plate-shaped base such as a table or a reception desk.

[0031] The control model of the robot 400 may be a state transition model, a model learned using deep learning, a slot filtering, or other dialogue control models may be used.

[0032] The microphone speaker 500 is composed of a microphone and a speaker. The microphone speaker 500 is arranged near the robot 400. Thereby, the microphone speaker 500 acquires the voices around the place where the robot is installed. The microphone speaker 500 picks up, for example, the voice spoken by a customer facing the robot 400. The microphone speaker 500 generates an audio signal based on the picked-up voice. The microphone speaker 500 outputs the generated audio signal to the control device 200. The microphone speaker 500 outputs the audio signal transmitted from the terminal device 100. With such a configuration, the robot 400 can appear as if it is talking to the person facing the robot 400. The microphone speaker 500 may output voice according to an operation instruction instructed by the control unit 203 of the control device 200.

[0033] The camera 600 is an imaging device for moving images. The camera 600 is arranged, for example, behind the robot 400. The camera 600 captures a moving image around (near) the robot 400 so that the rear view of the robot 400 is reflected. That is, the camera 600 captures the robot 400 and the user from the perspective of a third party. Thereby, the camera 600 generates an image (first image) of the surroundings of the place where the robot 400 is installed. The camera 600 generates an image signal indicating the captured moving image and outputs it to the control device 200. From this perspective of a third party, the operator of the robot 400 will see the robot 400 and the user. Note that the camera 600 may be provided at any position in front of, behind, to the left, or to the right of the robot 400. The camera 600 may be at any position as long as it can at least capture the person (customer) facing the robot 400. For example, the camera 600 may be installed at a position where it can capture a first-person perspective video, or at a position where it can capture a third-person perspective video. As a position where a third-person perspective video can be captured, for example, the camera 600 may be arranged behind the robot 400. In the case of a third-person perspective video, a video including the person (customer) facing the robot 400 and the robot 400 is captured.

[0034] The light-emitting unit 700 is a light-emitting member such as an LED or a light bulb. The light-emitting unit 700 may include a plurality of light-emitting members on a panel having a predetermined shape, for example. The predetermined shape may be circular or a polygon such as a quadrilateral. The light-emitting unit 700 may emit light according to an operation instruction instructed by the control unit 203 of the control device 200. Note that the light-emitting unit 700 may emit light in different modes according to the execution content. For example, when the light-emitting unit 700 includes a plurality of light-emitting members, it may be configured to cause some of the light-emitting members to emit light according to an operation instruction.

[0035] The display unit 750 is a device that displays guidance information or the like to a person facing it. The display unit 750 is disposed, for example, beside the robot 400.

[0036] Next, the specific configurations of the terminal device 100 and the control device 200 will be described. First, the configuration of the terminal device 100 will be described. The terminal device 100 includes a communication unit 101, an input unit 102, a display unit 103, a microphone 104, a camera 105, a storage unit 106, and a control unit 107.

[0037] The communication unit 101 is a communication device such as a network interface. The communication unit 101 is communicably connected to the network 800 using a predetermined protocol. The communication unit 101 communicates with other devices via the network 800 according to the control of the control unit 107. The other devices in the present embodiment are the control device 200 and the relay server 300, but may be devices other than the control device 200 and the relay server 300.

[0038] The input unit 102 is configured using an input device such as a keyboard, a pointing device (mouse, tablet, etc.), a button, or a touch panel. The input unit 102 is operated by an operator when inputting an instruction of the operator to the terminal device 100. The input unit 102 may be an interface for connecting an input device to the terminal device 100. In this case, the input unit 102 inputs an input signal corresponding to an input operation by the operator to the terminal device 100.

[0039] The display unit 103 is an image display device such as a liquid crystal display, an organic EL (Electro Luminescence) display, an electrophoretic display, or a CRT (Cathode Ray Tube) display. The display unit 103 displays information according to the control of the control unit 107. The display unit 103 may be an interface for connecting an image display device to the terminal device 100. In this case, the display unit 103 generates an image signal for displaying an image according to the control of the control unit 107. The display unit 103 outputs the image signal to the image display device connected to the display unit 103. The display unit 103 is an aspect of the display device.

[0040] The display unit 103 in the present embodiment displays, on the screen, one or more action candidates that are candidates for the robot 400 to execute, based on the information transmitted from the control device 200 triggered by video input from the control device 200 or voice input via the microphone speaker 500. Further, the display unit 103 in the present embodiment displays the action candidate executed by the robot 400 in a manner different from other action candidates among the one or more action candidates. Further, the display unit 103 in the present embodiment, when voice input is performed via the microphone speaker 500, displays the voice recognition result of the voice input as a character string. Further, the display unit 103 in the present embodiment performs a display that emphasizes a specific word included in the character string that is the voice recognition result. The emphasized display is, for example, changing the size of the character, changing the color of the character, surrounding the character, etc., but other expressions may be used as long as they can be distinguished from other words. Further, the display unit 103 in the present embodiment further displays the content output by voice according to the action executed by the robot 400. Further, the display unit 103 in the present embodiment displays, on the same screen, at least any combination of the video captured by the camera 600, the figure indicating the person or object recognized based on the video captured by the camera 600, the character string of the voice recognition result of the voice input, the character string indicating one or more action candidates, and the character string indicating the content output by voice according to the action executed by the robot 400. Here, the figure indicating the person or object may have any shape, but will be described as a rectangle in the following description.

[0041] Microphone 104 picks up the sound near the terminal device 100. For example, microphone 104 acquires the voice of the operator of terminal device 100. The operator of terminal device 100, for example, speaks to a person facing robot 400. Microphone 104 generates an audio signal based on the acquired sound. Microphone 104 outputs the generated audio signal to terminal device 100.

[0042] Note that microphone 104 may be an interface for connecting a sound collection device such as an external microphone to terminal device 100. In this case, microphone 104 generates an audio signal from the sound input in the sound collection device and outputs it to terminal device 100.

[0043] Camera 105 captures a moving image of the operator of terminal device 100 and the vicinity of the operator. Camera 105 may be an interface for connecting another camera to terminal device 100. In this case, camera 105 generates an image signal according to the moving image captured by another camera and inputs image information based on the image signal to terminal device 100.

[0044] The storage unit 106 is configured using a storage device such as a magnetic hard disk device and a semiconductor storage device. The storage unit 106 stores information obtained from the control device 200 and operation intention estimation information. The operation intention estimation information is information capable of estimating an operation intention corresponding to the operation of the operator. For example, the operation intention estimation information may be in the form of a table in which an operation intention is associated as a result when a condition is satisfied, with the condition being a combination of an input operation (text input, scroll, click, etc.) of the operator and an operation area (search window, result display area, registration button, etc.) on the correction screen, or may be an estimator obtained by learning.

[0045] When using an estimator, machine learning or deep learning may be used to train a pair of an information sequence of an input operation and a label or vector representing an operation intention as a learning data set. Hereinafter, the operation intention estimation information is also referred to as an operation intention estimation table in a table format in which an operation intention is associated as a result when a condition is satisfied on the condition of a combination of an input operation of an operator and an operation area on a correction screen. The storage unit 106 may store voice information acquired by the microphone 104 and images generated by the camera 105.

[0046] The control unit 107 controls the operations of each part of the terminal device 100. The control unit 107 is configured by using a processor such as a CPU and a RAM (Random Access Memory). The control unit 107 functions as a communication control unit 171, a voice recognition unit 172, a video control unit 173, an information providing unit 174, an operation intention estimation unit 175, and an action control information correction control unit 176 when the processor executes a specific program.

[0047] The communication control unit 171 communicates with other devices via the communication unit 101 by executing a predetermined communication program. The communication control unit 171 transmits, for example, at least any one of voice information, operation information, or a correction instruction of action control information to the control device 200 via the communication unit 101. The operation information is information representing the estimated operation intention. The communication control unit 171 acquires, for example, voice information or image information transmitted from the control device 200 via the communication unit 101.

[0048] The voice recognition unit 172 executes voice recognition processing. The voice recognition processing is processing for generating a character string based on a voice signal. The voice recognition unit 172 generates a character string based on the voice signal output by the microphone 104 by executing the voice recognition processing. The voice recognition unit 172 may generate a character string by using a known method.

[0049] The video control unit 173 communicates with other devices by executing a predetermined video control program. For example, the video control unit 173 transmits the audio signal output by the microphone 104 to the control device 200. Also, the video control unit 173 receives image information from the control device 200. The video control unit 173 causes the display unit 103 to display an image based on the received image information.

[0050] The image information received from the control device 200 includes at least an image of the surroundings of the location where the robot 400 is installed. If there are people around the location where the robot 400 is installed, the image information received from the control device 200 will include images of the people. The image information received from the control device 200 may further include perceptual information such as the detection result of a person, information indicating the person or word to be highlighted, the character string of the speech content of the person, the character string indicating the candidate actions that the robot 400 can perform, and the character string indicating the actions actually taken by the robot 400.

[0051] The information providing unit 174 provides the operator with the audio information acquired by the control device 200. The audio information acquired by the control device 200 is information regarding the audio acquired around the location where the robot 400 is installed. The information providing unit 174 may acquire the audio information transmitted from the control device 200 and provide the audio information to the operator at the appropriate timing.

[0052] The operation intention estimation unit 175 estimates the operation intention of the operator according to an operation for modifying the action control information in which the action content of the robot 400 is described. Specifically, the operation intention estimation unit 175 estimates the operation intention of the operator based on at least the operation intention estimation information stored in the storage unit 106 and an operation for modifying the action control information in which the action content of the robot 400 is described. The operation for modifying the action control information includes character string input, string determination operation, mouse operation on the interface, registration or deletion of string information, etc. The string determination operation is, for example, an operation for determining the input character string. For example, after the character string is input, a predetermined key (e.g., Return key or Enter key) is selected or a predetermined button on the interface (e.g., OK button) is selected. Note that the operation for modifying the action control information targets all operation actions on the interface and is not limited to the above-described operations. When the operation intention estimation information is an estimator, the operation intention estimation unit 175 obtains, as an estimation result, the operation intention output by inputting an information sequence indicating the operation for modifying the action control information to the estimator. When the operation intention estimation information is an operation intention estimation table, the operation intention estimation unit 175 refers to the operation intention estimation table and estimates the operation intention of the operator based on the operation for modifying the action control information and the information on the area (operation area) where the operation is performed. The operation intention estimation unit 175 outputs the estimation result to the communication control unit 171.

[0053] The action control information modification control unit 176 modifies the action control information held by the control device 200 according to the operation of the operator.

[0054] Note that the terminal device 100 may control the robot 400 or devices related to the robot 400. When configured in this way, the terminal device 100 generates control information for controlling the robot 400 or devices related to the robot 400, and transmits the generated control information to the control device 200. Devices related to the robot 400 are, for example, actuators (motors, etc.), speakers, and light emitting units.

[0055] Next, the control device 200 will be described. The control device 200 includes a communication unit 201, a storage unit 202, and a control unit 203.

[0056] The communication unit 201 is a communication device such as a network interface. The communication unit 201 is communicably connected to the network 800 according to a predetermined protocol. The communication unit 201 performs data communication with other devices via the network 800 under the control of the control unit 203.

[0057] The storage unit 202 is configured using a storage device such as a magnetic hard disk device or a semiconductor storage device. The storage unit 202 stores a highlighted word table, action control information table, action candidate table, utterance information table, and operation intention utterance information table. Hereinafter, each of the highlighted word table, action control information table, action candidate table, utterance information table, and operation intention utterance information table will be described.

[0058] FIG. 2 is a diagram showing an example of the highlighted word table in the embodiment. In the highlighted word table, highlighted words to be noted are registered. The highlighted word represents a word to be highlighted among the words constituting the content spoken by a person. The highlighted word is preset in the control device 200 by the operator operating the terminal device 100. In FIG. 2, as an example, words such as "photo", "seal", and "where is it" are registered as highlighted words.

[0059] FIG. 3 is a diagram showing an example of an action control information table in the embodiment. In the action control information table, a target word, a response content, and a location are registered in association with each other. The target word represents a search word used at the time of search. Note that a plurality of similar words or the like may be registered as the target word so as to facilitate searching. The response content represents a word to be output as voice when the word registered in the target word is searched. The location represents the placement location in the facility of the product specified by the associated target word. The location information may represent a section of the map within the facility, or may be a combination of the floor number and the position. Note that not all items need to be registered in the action control information table. The action control information table is one aspect of the action control information.

[0060] FIG. 4 is a diagram showing an example of an action candidate table in the embodiment. In the action candidate table, a trigger and action candidates are registered in association with each other. The trigger represents a condition for determining the action candidates of the robot 400. In the example shown in FIG. 4, “a person is located within a predetermined range” and “input of information (for example, voice or operation)” are shown as triggers. The action candidates represent candidates for actions that the robot 400 executes when the conditions indicated by the triggers are satisfied. In the example shown in FIG. 4, “follow” and “display of options, display of guidance information” are shown as action candidates. In the example shown in FIG. 4, when “a person is located within a predetermined range”, “follow” is selected as the action candidate of the robot 400, and when “input of information (for example, voice or operation)” is made, “display of options, display of guidance information” is selected as the action candidate of the robot 400. Note that the action candidate table may register the action candidates of the robot 400 for each one specific word or combination of specific words.

[0061] FIG. 5 is a diagram showing an example of a speech information table in the embodiment. In the speech information table, an action and a speech content are registered in association with each other. The action represents an action determined by the robot 400 to actually perform among the action candidates. The speech content represents the content to be spoken by the robot 400 according to the determined action. In the example shown in FIG. 5, when it is determined to perform the action of "following", it is shown that the robot 400 is made to speak contents such as "Shall we not talk?" or "Are you looking for something?". Further, in the example shown in FIG. 5, when it is determined to perform the action of "displaying options", it is shown that the robot 400 is made to speak contents such as "〇 candidates have been found. Please touch the item you are looking for". Further, in the example shown in FIG. 5, when it is determined to perform the action of "displaying guidance information", it is shown that the robot 400 is made to speak contents such as "If it is "...", it is on the A floor". Note that the name of the item selected by the person is input to "…".

[0062] FIG. 6 is a diagram showing an example of an operation intention speech information table in the embodiment. In the operation intention speech information table, a condition (operation intention) and a speech content are registered in association with each other. The condition (operation intention) represents an operation intention estimated by the operation intention estimation unit 175. The speech content represents the content to be spoken by the robot 400 when the condition (operation intention) is satisfied. In the example shown in FIG. 6, when the condition "a character string is input to the text search input area "□□"" is satisfied, it is shown that the robot 400 is made to speak contents such as "Oh, no candidates have come out yet. I will search for □□". Further, in the example shown in FIG. 6, when the condition "mouse operation" is satisfied, it is shown that the robot 400 is made to speak contents such as "If you search like this...". Further, in the example shown in FIG. 6, when the condition "click (item selection) item "□□"" is satisfied, it is shown that the robot 400 is made to speak contents such as "Is this "□□" different?". Note that the name of the item selected by the person is input to "□□".

[0063] Returning to FIG. 1, the description will continue. The control unit 203 controls the operations of each part of the control device 200. The control unit 203 is composed of a processor such as a CPU and a RAM. By executing a specific program by the processor, the control unit 203 functions as an audio recognition unit 241, a video recognition unit 242, a sensor recognition unit 243, a communication control unit 244, a perception information generation unit 245, and an action control unit 246.

[0064] The audio recognition unit 241 executes audio recognition processing. By executing the audio recognition processing, the audio recognition unit 241 generates a character string based on the audio signal output by the microphone-speaker 500. The audio recognition unit 241 outputs the generated character string to the perception information generation unit 245. The audio recognition unit 241 may generate a character string using a known method.

[0065] The video recognition unit 242 acquires the image signal output by the camera 600. The video recognition unit 242 outputs the acquired image signal to the perception information generation unit 245. Also, the video recognition unit 242 receives an audio signal from the terminal device 100. The video recognition unit 242 outputs the received audio signal to the microphone-speaker 500.

[0066] The sensor recognition unit 243 acquires the sensor signal detected by a sensor (for example, a human sensor, etc.) not shown. The sensor recognition unit 243 outputs the acquired sensor signal to the perception information generation unit 245.

[0067] The communication control unit 244 communicates with other devices by executing a predetermined communication program. The communication control unit 244 receives, for example, at least one of audio information or image information from the terminal device 100 via the relay server 300. The communication control unit 244 transmits, for example, the image information output from the perception information generation unit 245 to the terminal device 100 via the relay server 300.

[0068] The perception information generation unit 245 generates perception information based on at least any one of the character string output from the speech recognition unit 241, the image signal output from the video recognition unit 242, the sensor signal output from the sensor recognition unit 243, and the action determined by the action control unit 246. Specifically, the perception information generation unit 245 generates, as perception information, the character string of the speech content of a person and the information indicating the word to be highlighted based on the character string output from the speech recognition unit 241. The perception information generation unit 245 generates, as perception information, the detection result of a person based on the image signal output from the video recognition unit 242. The perception information generation unit 245 generates, as perception information, the information indicating the person to be highlighted based on the sensor signal output from the sensor recognition unit 243. The perception information generation unit 245 generates, as perception information, the character string indicating the candidates for the actions that the robot 400 can execute and the character string indicating the actions actually performed by the robot 400 based on the action determined by the action control unit 246. The perception information generation unit 245 generates image information including the generated perception information and the image signal output from the video recognition unit 242. The perception information generation unit 245 outputs the generated image information to the communication control unit 244.

[0069] The action control unit 246 controls the operations of the robot 400, the microphone - speaker 500, the light - emitting unit 700, and the display unit 750 based on the information stored in the storage unit 202. The action control unit 246 causes the robot 400 to execute a specific action by, for example, outputting control information including an operation instruction to the robot 400. The action control unit 246 causes the microphone - speaker 500 to output a character string as voice by, for example, outputting control information including the character string to be voice - output to the microphone - speaker 500. The action control unit 246 causes the light - emitting unit 700 to emit light by, for example, outputting control information including an emission instruction to the light - emitting unit 700. The action control unit 246 causes the display unit 750 to display information by, for example, outputting control information including the information to be displayed to the display unit 750.

[0070] Here, a specific operation example of a method for grasping the dialogue state of an autonomously controlled interactive agent, which is the first feature of the dialogue system 1, will be described. FIG. 7 is a diagram showing a first example of a display image (image of the operator interface). In FIG. 7, four display images are shown in FIGS. 7(A) to 7(D), and each display image shown in FIGS. 7(A) to 7(D) is displayed on the display unit 103 of the terminal device 100 operated by the operator. In FIGS. 7(A) to 7(D), the flow from the detection of a person to the action selection of the robot 400 will be described.

[0071] When the control device 200 detects a person from the image captured by the camera 600, the display unit 103 displays the display image shown in FIG. 7(A). In the display image shown in FIG. 7(A), at least the robot 400 and specific information inf1 for identifying the person are displayed. The specific information inf1 is information indicating a person detected by an existing person detection method. The specific information inf1 is represented by, for example, a rectangle surrounding the person. Here, as an example, the display unit 103 displays the specific information inf1 with a dotted line. The specific information inf1 is provided from the control device 200 as perceptual information. By displaying the specific information inf1, the operator can grasp that a person has been detected.

[0072] When a person is located within a predetermined range (for example, a range of 1.5 m from the position of the robot 400) from the position of the robot 400, the display unit 103 displays the display image shown in FIG. 7(B). In the display image shown in FIG. 7(B), at least the robot 400 and specific information inf1 and inf2 for identifying the person are displayed. The specific information inf2 is information indicating a person to be noted among the detected persons. The person to be noted is, for example, a person located within a range of 1.5 m from the position of the robot 400. Therefore, when a person to be noted is detected, the display unit 103 displays the specific information inf2 surrounding the detected person to be noted on the screen. The specific information inf2 is represented by, for example, a rectangle surrounding the person. The specific information inf2 is provided from the control device 200 as perceptual information. Here, as an example, the display unit 103 displays the specific information inf2 in a manner different from the specific information inf1 (for example, a solid line).

[0073] When a person to be noted is detected, the display unit 103 displays the display image shown in FIG. 7(C). In the display image shown in FIG. 7(C), in addition to the information shown in FIG. 7(B), an action candidate display area R1 is displayed. The action candidate display area R1 is an area in which one or more action candidates that are candidates for the robot 400 to execute are displayed. In the example shown in FIG. 7(C), candidate information inf3 indicating one action candidate "face tracking" is displayed in the action candidate display area R1. The candidate information inf3 to be displayed by the display unit 103 is provided from the control device 200 as perception information. The action candidate display area R1 may be displayed in any area on the screen of the display unit 103, but it is desirable to be displayed in an area where a person is not imaged (for example, the upper part of the screen) so as not to interfere with the operator's viewing.

[0074] When candidate information inf3 indicating the action candidate to be executed by the robot 400 is determined from among the candidate information inf3 displayed in the action candidate display area R1, the display unit 103 displays the display image shown in FIG. 7(D). In the display image shown in FIG. 7(D), the candidate information inf3 indicating the action candidate "face tracking" determined to be executed by the robot 400 is emphasized (for example, highlighted) and displayed. Information indicating the action determined to be executed by the robot 400 is provided from the control device 200 as perception information. Thereby, the operator can easily grasp which action the robot 400 is currently performing by looking at the screen.

[0075] FIGS. 8 to 10 are diagrams showing a second example of the display image (image of the operator interface). In FIGS. 8 to 10, five display images are shown in FIGS. 8(A), 8(B), 9(A), 9(B), and 10, and each display image shown in FIGS. 8(A), 8(B), 9(A), 9(B), and 10 is displayed on the display unit 103 of the terminal device 100 operated by the operator. In FIGS. 8(A), 8(B), 9(A), 9(B), and 10, the flow from the speech of the person to be noted to the display of the result of the action actually performed by the robot 400 will be described.

[0076] When the control device 200 obtains in real time the content of what a person has spoken through voice recognition, the display unit 103 displays the display image shown in Fig. 8(A). In the display image shown in Fig. 8(A), in addition to the information shown in Fig. 7(B), a person speech display area R2 is displayed. The person speech display area R2 is an area where the content of what a person has spoken is displayed. In the person speech display area R2 shown in Fig. 8(A), "User: Photo" is displayed. This indicates that the person to be noted has spoken "Photo". Note that the speech content displayed in the person speech display area R2 is the content obtained by the control device 200 through voice recognition. Therefore, depending on the accuracy of the voice recognition of the control device 200, there may be a case where the speech content of a person is different from the speech content displayed in the person speech display area R2.

[0077] By displaying the speech content in the person speech display area R2, the operator can grasp the content of what the person has spoken. Furthermore, when the content of what the person has spoken is different from the content displayed in the person speech display area R2, the operator can easily grasp that the voice recognition has failed. The person speech display area R2 may be displayed in any area on the screen of the display unit 103, but it is desirable to be displayed in an area where the person is not imaged (for example, the upper part of the screen) so as not to interfere with the operator's viewing.

[0078] When a specific word is included in the speech content of the person to be noted, the display unit 103 displays the display image shown in Fig. 8(B). In the display image shown in Fig. 8(B), the specific word "Photo" in the speech content displayed in the person speech display area R2 is displayed as emphasized information inf4. In the person speech display area R2 shown in Fig. 8(B), "User: The thing that puts in photos" is displayed. Then, the display unit 103 displays the specific word "Photo" in "User: The thing that puts in photos" displayed in the person speech display area R2 in a manner different from other words. The information indicating the word to be emphasized is provided from the control device 200 as perceptual information.

[0079] When the speech of the person ends, the display unit 103 displays the display image shown in Fig. 9(A). In the display image shown in Fig. 9(A), in addition to the person speech display area R2, an action candidate display area R1 is displayed. In the person speech display area R2 shown in Fig. 9(A), "User: Where is the thing for inserting photos?" is displayed. In the display image shown in Fig. 9(A), among the speech content "User: Where is the thing for inserting photos?" displayed in the person speech display area R2, a specific word "Where is it?" is displayed as emphasized information inf5. Then, the display unit 103 displays a specific word "photo" and "Where is it?" in the speech content "User: Where is the thing for inserting photos?" displayed in the person speech display area R2 in a manner different from other words. Furthermore, the display unit 103 displays the action candidate display area R1 below the person speech display area R2. In the example shown in Fig. 9(A), three action candidates, "Option display", "Guide: Photo stand", and "Guide: Album", are displayed in the action candidate display area R1. These three action candidates are determined according to one specific word or a combination of specific words in the speech content displayed in the person speech display area R2 and are provided from the control device 200 as perceptual information.

[0080] When the action candidate that the robot 400 should execute is determined from the three action candidates displayed in the action candidate display area R1, the display unit 103 displays the display image shown in Fig. 9(B). In the display image shown in Fig. 9(B), the action candidate "Option display" determined to be executed by the robot 400 is displayed emphasized, different from the other action candidates "Guide: Photo stand" and "Guide: Album". The information indicating the action that the robot 400 should execute is provided from the control device 200 as perceptual information. Thereby, the operator can easily grasp which action the robot 400 is currently performing by looking at the screen.

[0081] When the determined action candidate is executed by the robot 400, the display unit 103 displays the display image shown in FIG. 10. In the display image shown in FIG. 10, in addition to the action candidate display area R1 and the human speech display area R2, an action result display area R3 is displayed. The action result display area R3 is an area where the result of the action actually performed by the robot 400 is displayed. In the present embodiment, in the action result display area R3, the speech content of the robot 400 is displayed as the result of the action actually performed by the robot 400. In the action result display area R3 shown in FIG. 10, it is shown that the robot 400 has spoken "0 candidates were found. Please touch what you are looking for." The result of the action actually performed by the robot 400 is provided from the control device 200 as perceptual information.

[0082] The action result display area R3 may be displayed in any area on the screen of the display unit 103, but it is preferably displayed in an area where no person is imaged (for example, the upper part of the screen) so as not to interfere with the operator's viewing. In the example shown in the present embodiment, the action candidate display area R1, the human speech display area R2, and the action result display area R3 are displayed from the top of the screen in this order, but the display order of the action candidate display area R1, the human speech display area R2, and the action result display area R3 may be in any order. In this way, by displaying the action candidate display area R1, the human speech display area R2, and the action result display area R3, the operator can easily grasp what actions the robot 400 actually took in response to the human speech content by looking at the screen.

[0083] FIGS. 11 and 12 are sequence diagrams showing the flow of processing performed by the dialogue system 1 in the embodiment. The robot 400, the microphone / speaker 500, the camera 600, and the light emitting unit 700 are collectively referred to as robots and the like.

[0084] The camera 600 images the surroundings of the robot 400 including the robot 400 (step S101). The camera 600 outputs an image signal indicating the captured image to the control device 200 in real time. The video recognition unit 242 of the control device 200 acquires the image signal output from the camera 600. The video recognition unit 242 outputs the acquired image signal to the perception information generation unit 245. The perception information generation unit 245 generates image information based on the image signal output from the video recognition unit 242. Here, since there is no perception information to be given, the perception information generation unit 245 generates the image signal output from the video recognition unit 242 as image information. The perception information generation unit 245 outputs the generated image information to the communication control unit 244.

[0085] The communication control unit 244 transmits the image information output from the perception information generation unit 245 to the terminal device 100 (step S102). The video control unit 173 of the terminal device 100 causes the display unit 103 to display an image based on the image information transmitted from the control device 200 (step S103). Thereafter, the processing from step S101 to step S103 is executed until a person is detected.

[0086] The perception information generation unit 245 of the control device 200 detects a person based on the image signal output from the camera 600 (step S104). Existing techniques can be applied to the method of detecting a person. When the perception information generation unit 245 detects a person, it generates image information in which specific information inf1 for specifying the detected person is attached as perception information to the image signal. The perception information generation unit 245 outputs the generated image information to the communication control unit 244. The communication control unit 244 transmits the image information output from the perception information generation unit 245 to the terminal device 100 (step S105). The video control unit 173 of the terminal device 100 causes the display unit 103 to display an image based on the image information transmitted from the control device 200. Thereby, for example, the display unit 103 displays the image shown in FIG. 7(A) (step S106).

[0087] After that, the perception information generation unit 245 of the control device 200 detects a person to be noted based on the sensor signals output from the sensor recognition unit 243 (step S107). When the perception information generation unit 245 detects a person to be noted, it generates image information by attaching specific information inf2 for specifying the detected person to be noted as perception information to the image signal. The perception information generation unit 245 outputs the generated image information to the communication control unit 244. The communication control unit 244 transmits the image information output from the perception information generation unit 245 to the terminal device 100 (step S108). The video control unit 173 of the terminal device 100 causes the display unit 103 to display an image based on the image information transmitted from the control device 200. Thereby, for example, the display unit 103 displays the image shown in FIG. 7(B) (step S109).

[0088] After that, assume that the person of interest speaks. In this case, the microphone speaker 500 collects the speech of the person of interest (step S110). The microphone speaker 500 outputs an audio signal indicating the content of the speech of the collected person of interest to the control device 200 (step S111). The speech recognition unit 241 of the control device 200 performs speech recognition processing on the audio signal output from the microphone speaker 500 (step S112). Thereby, the speech recognition unit 241 generates a character string based on the audio signal output from the microphone speaker 500. The speech recognition unit 241 outputs the generated character string to the perception information generation unit 245. The perception information generation unit 245 generates image information by attaching the character string output from the speech recognition unit 241 as perception information to the image signal. The perception information generation unit 245 outputs the generated image information to the communication control unit 244. The communication control unit 244 transmits the image information output from the perception information generation unit 245 to the terminal device 100 (step S113). The video control unit 173 of the terminal device 100 causes the display unit 103 to display an image based on the image information transmitted from the control device 200. Thereby, the display unit 103 displays a display image in which the character string output from the speech recognition unit 241 is superimposed on the person speech display area R2 of the image represented by the image signal (step S114). For example, the display unit 103 displays the image shown in FIG. 8(A).

[0089] The perception information generation unit 245 of the control device 200 refers to the attention word table and selects a specific word from the character string output from the speech recognition unit 241 (step S115). The perception information generation unit 245 generates image information in which emphasis information inf4 that emphasizes the selected specific word is added to the image signal as perception information. The perception information generation unit 245 outputs the generated image information to the communication control unit 244. The communication control unit 244 transmits the image information output from the perception information generation unit 245 to the terminal device 100 (step S116). The video control unit 173 of the terminal device 100 causes the display unit 103 to display an image based on the image information transmitted from the control device 200. As a result, the display unit 103 displays a display image in which a specific word specified by the emphasis information inf4 among the character strings output from the speech recognition unit 241 is emphasized and superimposed on the person speech display area R2 of the image represented by the image signal (step S117). As a result, for example, the display unit 103 displays the image shown in FIG. 8(B).

[0090] Next, the action control unit 246 of the control device 200 refers to the action candidate table and selects an action candidate for the robot 400 according to the selected one specific word or combination of specific words (step S118). The action control unit 246 outputs information on the selected action candidate for the robot 400 to the perception information generation unit 245. The perception information generation unit 245 generates image information in which information on the action candidate for the robot 400 output from the action control unit 246 is added to the image signal as perception information. The perception information generation unit 245 outputs the generated image information to the communication control unit 244. The communication control unit 244 transmits the image information output from the perception information generation unit 245 to the terminal device 100 (step S119). The video control unit 173 of the terminal device 100 causes the display unit 103 to display an image based on the image information transmitted from the control device 200. As a result, the display unit 103 displays a display image in which information on the action candidate for the robot 400 is superimposed on the action candidate display area R1 of the image represented by the image signal (step S120). For example, the display unit 103 displays the image shown in FIG. 9(A).

[0091] Next, the action control unit 246 of the control device 200 selects an action to be actually executed by the robot 400 from among the action candidates selected in the process of step S118 (step S121). The selection of an action can be expressed not only by a model that selects an action but also by a model that controls speech and actions by generating and enumerating a plurality of actions as memories. The action control unit 246 outputs information on the action candidate indicating the selected action to the perception information generation unit 245. The perception information generation unit 245 generates image information in which the information on the action candidate indicating the action output from the action control unit 246 is added to the image signal as perception information. The perception information generation unit 245 outputs the generated image information to the communication control unit 244. The communication control unit 244 transmits the image information output from the perception information generation unit 245 to the terminal device 100 (step S122). The video control unit 173 of the terminal device 100 causes the display unit 103 to display an image based on the image information transmitted from the control device 200. As a result, the display unit 103 displays a display image in which the action candidate indicating the selected action is emphasized (step S123). For example, the display unit 103 displays the image shown in FIG. 9(B).

[0092] The action control unit 246 refers to the speech information table and acquires the speech content corresponding to the selected action. The action control unit 246 generates control information for uttering the acquired speech content. The control device 200 transmits the generated control information to the microphone / speaker 500 (step S124). For example, the control device 200 transmits control information including a character string indicating the speech content to the microphone / speaker 500. The microphone / speaker 500 outputs the character string included in the control information transmitted from the control device 200 as speech (step S125). As a result, it is possible to make it appear as if the robot 400 has spoken.

[0093] Furthermore, the action control unit 246 outputs the information of the character string included in the control information transmitted to the microphone / speaker 500 to the perception information generation unit 245. The perception information generation unit 245 generates image information in which the information of the character string output from the action control unit 246 is attached to the image signal as perception information. The perception information generation unit 245 outputs the generated image information to the communication control unit 244. The communication control unit 244 transmits the image information output from the perception information generation unit 245 to the terminal device 100 (step S126). The video control unit 173 of the terminal device 100 causes the display unit 103 to display an image based on the image information transmitted from the control device 200. As a result, the display unit 103 displays a display image in which the character string representing the action executed by the robot 400 is superimposed on the action result display area R3 of the image represented by the image signal (step S127). For example, the display unit 103 displays the image shown in FIG. 10.

[0094] The above-described process is a method for grasping the dialogue state of the autonomously controlled interactive agent, which is the first feature. According to the dialogue system 1 configured as described above, the operator can see on the screen the video captured by the camera 600, the speech recognition result of the voice input as a character string (for example, the content spoken by a person), the character string indicating the action candidate that is a candidate for the action performed by the robot 400, and at least one of the character strings indicating the content output as voice according to the action executed by the robot 400. Therefore, it becomes possible to easily grasp the dialogue between the robot 400 and the person.

[0095] Furthermore, by displaying the display images shown in FIGS. 7 to 10, it becomes possible to easily grasp the cause of the failure of the dialogue of the robot 400. The following are considered as the causes of the failure of the dialogue of the robot 400. · Failure cause 1: The voice spoken by the person can be heard, but the character string indicating the speech recognition result is not displayed on the screen of the display unit 103. · Failure cause 2: The audible voice is different from the character string indicating the speech recognition result displayed on the screen of the display unit 103. · Cause of failure 3: A string indicating the voice recognition result is displayed on the screen of the display unit 103, but a specific word is not highlighted. · Cause of failure 4: The correct action is not displayed as an action candidate. · Cause of failure 5: The action that the robot 400 actually performs is determined, but the robot 400 does not act.

[0096] The cause of failure 1 is that the operator can hear the voice of a person's speech, but a string indicating the voice recognition result is not displayed in the person's speech display area R2 of the display unit 103. In this case, the operator can grasp that the voice recognition function of the control device 200 is not functioning. As a result, it is possible to respond by performing a process such as restarting the application of the control device 200.

[0097] The cause of failure 2 is that the content displayed in the person's speech display area R2 of the display unit 103 is different from the content of the voice that the operator can hear. In this case, the operator can grasp that the voice recognition function of the control device 200 is operating, but the voice recognition has failed. As a result, it is possible to respond by performing a process to support the voice recognition.

[0098] The cause of failure 3 is that the string displayed in the person's speech display area R2 of the display unit 103 is correct, but the word is not emphasized. In this case, the operator can grasp that the word to be highlighted is not registered. As a result, it is possible to respond by registering the word to be highlighted.

[0099] The failure cause 4 is that the correct action is not displayed as an action candidate in the action candidate display area R1 of the display unit 103. Here, the correct action is an action that is considered appropriate as an action assumed from the voice-recognized character string. As a case where the correct action is not displayed as an action candidate, for example, when a person says they want a certain product, a guide for a completely unrelated product is displayed as an action candidate. In this case, the operator can recognize that the enumeration of action candidates has failed or that the action information is insufficient. As a result, it can be dealt with by newly registering the action information.

[0100] The failure cause 5 is that, despite the action to be executed being determined, the robot 400 or the like does not take an action. In this embodiment, since the action by voice output is the main one, in this case, the operator can recognize that the microphone speaker 500 is not functioning. As a result, it can be dealt with by performing processing such as restarting the application of the microphone speaker 500.

[0101] As described above, when the dialogue by the robot 400 fails, the operator can easily identify the cause. Since the cause of such failure can be easily identified, it becomes easier to determine a countermeasure. As a result, even when there is little knowledge or experience regarding robot control or the system, it is possible to instantly grasp what is happening. Therefore, it is possible to easily grasp the action to be taken as a response to the dialogue failure and to execute it.

[0102] (Method for recovering from dialogue failure and correcting behavior by a dialogue agent) Next, a method for recovering from dialogue failures and modifying behavior by an interactive agent, which is the second feature, will be described. Since an autonomously controlled interactive agent operates based on preset information, it will perform the same failure as when the same conversation as the one that has failed once occurs. Therefore, the operator needs to intervene and respond to the same failure each time. Furthermore, if the robot is controlled to speak after the correction work is performed, it is conceivable that the user will leave. Therefore, in the present embodiment, the recovery from dialogue failures by the robot, the modification of behavior, and the correction work of actions are made compatible in real time. More specifically, on the robot side, the operation intention is estimated according to the operation of the operator's correction work, and a speech corresponding to the estimated operation intention is performed. As a result, the operator does not need to conduct a dialogue in the correction response and can concentrate on the correction work.

[0103] FIG. 13 is a diagram for explaining an outline of a method for recovering from a dialogue failure and correcting behavior by an interactive agent in an embodiment. Assume that an operator who is an operator operates the terminal device 100 to perform a correction operation. In this case, the operator operates the screen of the display unit 103 to display a correction screen. The correction screen is a screen for correcting action control information related to the action of the robot 400. Here, assume that the operator inputs the character string "album" as a correction operation in the text search input area on the correction screen. In this case, the operation intention estimation unit 175 of the terminal device 100 estimates an operation intention corresponding to the input of text to the text search input area, triggered by the operator's correction operation. Then, the terminal device 100 transmits information indicating the estimated operation intention to the control device 200. The control device 200 refers to the operation intention utterance information table and acquires the utterance content (for example, "Oh, no candidates have come out yet. I will search by album.") associated with the operation intention transmitted from the terminal device 100. The control device 200 outputs the acquired utterance content to the microphone-speaker 500 as voice. In this way, even though the operator is performing a correction operation, the dialogue continues as if the robot 400 outputs voice such as "Oh, no candidates have come out yet. I will search by album." As a result, during the operation of searching whether information related to "album" is registered as a correction operation by the operator, the robot 400 can continue the dialogue.

[0104] As another example, assume that the operator performs "mouse scrolling" as an operation for the correction work on the correction screen. In this case, upon the operation of the operator's correction work, the operation intention estimation unit 175 of the terminal device 100 estimates an operation intention corresponding to the mouse scrolling. Then, the terminal device 100 transmits information indicating the estimated operation intention to the control device 200. The control device 200 refers to the operation intention speech information table and acquires the speech content (for example, "Let's search and see...") associated with the operation intention transmitted from the terminal device 100. The control device 200 causes the microphone-speaker 500 to output the acquired speech content as voice. In this way, even though the operator is performing the correction work, the dialogue continues as if the robot 400 outputs voice content such as "Let's search and see...". Thereby, during the work where the operator is searching while scrolling the mouse as a correction work, the robot 400 can continue the dialogue.

[0105] As another example, assume that the operator selects "item selection" as an operation for the correction work on the correction screen. For example, assume that an album (for photos) is selected as the item. In this case, upon the operation of the operator's correction work, the operation intention estimation unit 175 of the terminal device 100 estimates an operation intention corresponding to the selection of the selected item. Then, the terminal device 100 transmits information indicating the estimated operation intention to the control device 200. The control device 200 refers to the operation intention speech information table and acquires the speech content (for example, "Is this 'album (for photos)' different?") associated with the operation intention transmitted from the terminal device 100. The control device 200 causes the microphone-speaker 500 to output the acquired speech content as voice. In this way, even though the operator is performing the correction work, the dialogue continues as if the robot 400 outputs voice content such as "Is this 'album (for photos)' different?". Thereby, during the work where the operator is searching while scrolling the mouse as a correction work, the robot 400 can continue the dialogue.

[0106] FIG. 14 is a diagram showing a third example of a display image (image of an operator interface). The display image shown in FIG. 14 is an image of the screen of the display unit 103 of the terminal device 100. In the display image shown in FIG. 14, a display area R11 and a correction button B1 are displayed. In the display area R11, a video on the robot side is displayed. For example, the display area R11 is an area where image information transmitted from the control device 200 as shown in FIGS. 7 to 10 is displayed. The correction button B1 is a button for correcting information stored in the storage unit 202 of the control device 200. By correcting the information stored in the storage unit 202 of the control device 200, the behavior of the robot 400 can be corrected. When the operator selects the correction button B1, the display unit 103 displays the display image shown in FIG. 15.

[0107] FIG. 15 is a diagram showing a fourth example of a display image (image of an operator interface). In the display image shown in FIG. 15, a display area R11, a correction button B1, and a display area R12 are displayed. The display area R12 is an area where a correction screen for correcting the behavior of the robot 400 is displayed. Correcting the behavior of the robot 400 means correcting the action control information. The correction screen displays a search input area R13, a search result display area R14, an action control information input area R15, and a registration button B2.

[0108] The search input area R13 is an area where a character string is input when performing a text search. For example, a character string to be searched (e.g., a seal, etc.) is input into the search input area R13. The search result display area R14 is an area where the result of searching for information corresponding to the character string input into the search input area R13 is displayed. For example, the action control information modification control unit 176 refers to the target word in the action control information table according to the character string input into the search input area R13, and searches for whether there is a record corresponding to the character string input into the search input area R13. If there is a record corresponding to the character string input into the search input area R13, the action control information modification control unit 176 causes the information registered in the record corresponding to the character string input into the search input area R13 to be displayed in the search result display area R14. Note that depending on the character string input into the search input area R13, there may be a plurality of corresponding records. In this case, the action control information modification control unit 176 causes the information registered in the plurality of records obtained as search results to be displayed in the search result display area R14.

[0109] On the other hand, when there is no record corresponding to the character string input into the search input area R13, the action control information modification control unit 176 either displays nothing in the search result display area R14 or causes a message indicating that there is no corresponding information to be displayed in the search result display area R14. In the example shown in FIG. 15, the result of searching corresponding to the character string (e.g., a seal) input into the search input area R13 is displayed in the search result display area R14.

[0110] The action control information input area R15 is an area used to modify the information in the action control information table. When the information corresponding to the character string input in the search input area R13 is registered in the action control information table, the response content, target word, and location items shown in the action control information input area R15 display the response content, target word, and location information registered in the action control information table. In the example shown in FIG. 15, all the information is registered in the response content, target word, and location items shown in the action control information input area R15. However, if some information is not registered in the action control information table, some of the information will not be displayed in the response content, target word, and location items shown in the action control information input area R15. For example, in the action control information table shown in FIG. 3, no response content is registered for the target word "photo". In this case, among the response content, target word, and location items shown in the action control information input area R15, nothing is displayed in the response content item, "photo" is displayed in the target word item, and "AA" is displayed in the location item.

[0111] The registration button B2 is a button used when registering the information input in the action control information input area R15. When the registration button B2 is selected after information is input in the action control information input area R15, the action control information modification control unit 176 registers the information input in the action control information input area R15 in the action control information table. Specifically, when the registration button B2 is selected after the target word information and response content information, etc. are input in the action control information input area R15, the action control information modification control unit 176 newly registers in the action control information table the record corresponding to the character string input in the target word area of the action control information input area R15.

[0112] When the registration button B2 is selected after the information on the answer content is input into the action control information input area R15, first, the action control information modification control unit 176 refers to the action control information table based on the character string input in the target word area of the action control information input area R15 and selects a record corresponding to the character string input in the target word area. Then, the action control information modification control unit 176 additionally registers the character string input in the answer content area of the action control information input area R15 in the answer content item of the selected record.

[0113] When the registration button B2 is selected after the location information is input into the action control information input area R15, first, the action control information modification control unit 176 refers to the action control information table based on the character string input in the target word area of the action control information input area R15 and selects a record corresponding to the character string input in the target word area. Then, the action control information modification control unit 176 additionally registers the character string input in the location area of the action control information input area R15 in the location item of the selected record.

[0114] New information is additionally registered in the action control information table by the above modification process. As a result, when the information corresponding to the character string input in the search input area R13 is registered in the action control information table, the action control information modification control unit 176 displays the information in the search result display area R14 in real time. Thereby, the operator can select the item displayed in the search result display area R14.

[0115] In addition, the correction screen may be provided with a reflection selection button that allows the user to select whether to reflect the correction content in real time. When the reflection selection button is ON, the content after correction by the operator is reflected in real time. That is, when the reflection selection button is ON, the action control information correction control unit 176 updates the information registered in the action control information table stored in the control device 200 based on the corrected content. Thereby, the corrected content is reflected in real time. When the reflection selection button is OFF, the content after correction by the operator is not reflected in real time. Therefore, when the reflection selection button is OFF, the action control information correction control unit 176 updates the information registered in the action control information table stored in the control device 200 at a timing when a predetermined time has elapsed based on the corrected content. Thereby, it is possible to reflect in real time the correction content that is considered necessary to be reflected in real time, and to reflect later the correction content that is not necessary to be reflected in real time. As a result, the processing load caused by reflecting all the correction content can be reduced.

[0116] Furthermore, the correction screen may be provided with a video call area for the staff to make inquiries via video call. The staff are those who work at the facility where the robot 400 is installed. This enables real-time consultation with the staff and modification of the action control information even in situations where it is difficult for the operator to handle (for example, when it is unclear where the products are placed). Note that there is a possibility of keeping the customer waiting while the operator is on a video call with the staff. Therefore, when an operation to make a video call is performed, the operation intention estimation unit 175 estimates an operation intention corresponding to the video call operation. The control device 200 may perform an operation for notifying that a staff inquiry is being made and effectively utilizing the customer's waiting time based on the operation intention corresponding to the video call operation. Operations for effectively utilizing the customer's waiting time include, for example, quizzes or introductions of recommended products. In this way, by notifying the customer that a staff inquiry is currently being made and effectively utilizing the waiting time, the customer can be kept on the spot.

[0117] The screens of the display unit 103 of the terminal device 100 in FIGS. 14 and 15 described above are examples. As interfaces on the screen of the display unit 103 of the terminal device 100, various input form formats such as text fields, buttons, check boxes, slide bars, table-form inputs, click coordinate positions on figures, etc. can be used.

[0118] FIG. 16 is a sequence diagram showing the flow of processing performed by the dialogue system 1 in the embodiment. The control device 200, the robot 400, the microphone-speaker 500, the camera 600, and the light emitting unit 700 are collectively referred to as the robot side. Note that FIG. 16 will explain, as an example, dialogue recovery when an error occurs in voice recognition by the control device 200.

[0119] Suppose a customer says, "Where is XX?" (Step S201). The microphone speaker 500 collects the customer's utterance. The microphone speaker 500 outputs an audio signal indicating the content of the customer's utterance "Where is XX?" to the control device 200. The speech recognition unit 241 of the control device 200 performs speech recognition processing on the audio signal output from the microphone speaker 500. Here, suppose that as a result of speech recognition by the speech recognition unit 241, it is recognized as "Where is YY?" (Step S202). In this case, the control device 200 selects an item candidate that hits with "YY" (Step S203). The control device 200 generates image information including the selected item candidate. The control device 200 provides the generated image information to the terminal device 100 (Step S204).

[0120] The video control unit 173 of the terminal device 100 displays the image information provided from the control device 200 on the display unit 103. As a result, for example, on the screen of the display unit 103, item candidates that hit with "YY" are presented. The operator grasps that speech recognition has failed by looking at the screen of the display unit 103. The operator operates the terminal device 100 to select the correction button B1. The action control information correction control unit 176 causes the display unit 103 to display a correction screen in response to the selection of the correction button B1. The display unit 103 displays the correction screen according to the control of the action control information correction control unit 176 (Step S205). The operator inputs the character string "XX" into the search input area R13 of the displayed correction screen to perform a search.

[0121] Through this process, the action control information modification control unit 176 refers to the item of the target word in the action control information table stored in the storage unit 202 of the control device 200, and searches for a record corresponding to the character string "〇〇". If there is a record corresponding to the character string "〇〇", the action control information modification control unit 176 displays the information of the record corresponding to the character string "〇〇" in the search result display area R14. As a result, the information of the item is displayed in the search result display area R14. On the other hand, if there is no record corresponding to the character string "〇〇", the action control information modification control unit 176 displays in the search result display area R14 that there is no corresponding item candidate. Here, it is assumed that there was a corresponding item candidate. The operator selects the item displayed in the search result display area R14 (step S206).

[0122] The operation intention estimation unit 175 of the terminal device 100 estimates the operation intention of the operator according to the operation performed by the operator. In the above-described example, the operator first performs an operation of inputting the character string "〇〇" in the search input area R13, and then performs an operation of selecting an item from the search results displayed in the search result display area R14. Therefore, the operation intention estimation unit 175 refers to the operation intention estimation table and acquires the operation intention associated with the operation "text input" for modifying the action control information and the operation area (operation area) "search window" where the operation was performed. The operation intention obtained here is regarded as the first operation intention. Next, the operation intention estimation unit 175 refers to the operation intention estimation table and acquires the operation intention associated with the operation "click" for modifying the action control information and the operation area (operation area) "result display area" where the operation was performed. The operation intention obtained here is regarded as the second operation intention. The operation intention estimation unit 175 transmits operation information including the acquired first operation intention and second operation intention and the information of the selected item to the control device 200 (step S207).

[0123] The control device 200 receives the operation information transmitted from the terminal device 100. The control device 200 starts speaking according to each of the first operation intention and the second operation intention included in the operation information (step S208). Specifically, the control device 200 generates control information for causing the speaking corresponding to the first operation intention to be executed. The control device 200 transmits the selected control information to the microphone-speaker 500. For example, the control device 200 transmits control information including the character string to be output to the microphone-speaker 500. The microphone-speaker 500 outputs the character string included in the control information transmitted from the control device 200 as voice. For example, the microphone-speaker 500 outputs voice with the content such as "Don't you have the product you are looking for?".

[0124] After that, the control device 200 generates control information for causing the speaking corresponding to the second operation intention to be executed. The control device 200 transmits the selected control information to the microphone-speaker 500. For example, the control device 200 transmits control information including the character string to be output to the microphone-speaker 500. The microphone-speaker 500 outputs the character string included in the control information transmitted from the control device 200 as voice. For example, the microphone-speaker 500 outputs voice with the content such as "Is this XX?".

[0125] Furthermore, the control device 200 causes the display unit 750 to display the information of the item included in the operation information. The display unit 750 displays the information of the item according to the control of the control device 200 (step S209). Suppose the customer selects the item displayed on the display unit 750 (step S210). The control device 200 causes the display unit 750 to display guidance information for guiding the location where the selected item is arranged. The display unit 750 displays the guidance information according to the control of the control device 200 (step S211). The guidance information may be, for example, a map showing the location where the selected item is arranged, or a character string (for example, the second floor, etc.) showing the location where the selected item is arranged.

[0126] According to the dialogue system 1 configured as described above, in response to an operation for modifying the action control information, the operation intention of the operator is estimated, and the dialogue by the robot 400 is controlled according to the estimated operation intention of the operator. In this way, while the operator is performing the modification work of the action control information, the control device 200 performs control according to the operation intention of the operator. Thereby, the operator does not need to perform the dialogue expression and the modification work at the same time. Furthermore, since the dialogue continues even while the modification work is being performed, it is possible to reduce the situation where the person engaged in the dialogue leaves. Therefore, it becomes possible to achieve both the dialogue expression and the modification work.

[0127] <Modification Example 1> In the above-described embodiment, the configuration has been described on the premise of remotely operating the robot 400 (a configuration in which the perceptual information of the robot 400 is displayed on the terminal device 100 located at a location away from the location where the robot 400 is installed). On the other hand, the present invention is applicable also when the robot 400 performs only autonomous dialogue without remotely operating the robot 400. When configured in this way, the perceptual information of the robot 400 may be displayed on a display unit provided in the robot 400. By being configured in this way, by showing a display according to the perceptual information to the user side, it is possible to clarify how it is moving and to be used for obtaining trust or for maintenance.

[0128] <Modification Example 2> In the dialogue system 1, when the method for recovering from dialogue failure and correcting behavior by the dialogue-type agent, which is the second feature, is not performed, the terminal device 100 does not necessarily need to include the operation intention estimation unit 175 and the action control information modification control unit 176.

[0129] <Modification Example 3> In the above-described embodiment, the configuration in which the control device 200 uses each table (the target word table, the action control information table, the action candidate table, the utterance information table, and the operation intention utterance information table) is an example. For example, the configuration in which the control device 200 uses each table is merely an implementation example when using a state transition model. When using a deep learning model, generative AI, etc., it may not be necessary to use each table. That is, it is not necessary for some or all of each table to be stored in the storage unit 202. For example, the control device 200 may not have a target word table and may display the weights of the attention mechanism of the deep learning model in color using a deep learning model that understands utterances. For example, the control device 200 may not have an action candidate table and may generate one or more action candidates using a generation model and arrange the generated action candidates as candidates in the action candidate display area R1. For example, the control device 200 may not have an utterance information table and may automatically generate the utterance content using a deep learning model. Thus, each table is merely information necessary for driving the dialogue using a state transition model, and when using a deep learning model or a generation model as the dialogue control model, intermediate products or a plurality of outputs derived therefrom may be displayed.

[0130] Note that when displaying the weights of the attention mechanism as described above, in failure cause 3 in the above-described embodiment, by observing that a specific word is not emphasized, measures such as changing the model, modifying the learning data, changing the examples or prompts passed to the model can be taken. When using a deep learning model or generative AI as described above, in failure cause 4 in the above-described embodiment, measures such as changing the model, modifying the learning data, changing the examples or prompts passed to the model can be taken.

[0131] As described above, the embodiments of the present invention have been described in detail with reference to the drawings. However, the specific configuration is not limited to this embodiment, and designs and the like within the scope not departing from the gist of the present invention are also included.

Explanation of Reference Numerals

[0132] 1... Control system, 100... Terminal device, 101... Communication unit, 102... Input unit, 103... Display unit, 104... Microphone, 105... Camera, 106... Memory unit, 107... Control unit, 171... Communication control unit, 172... Speech recognition unit, 173... Video control unit, 174... Information providing unit, 175... Operation intention estimation unit, 176... Action control information correction control unit, 200... Control device, 201... Communication unit, 202... Memory unit, 203... Control unit, 241... Speech recognition unit, 242... Video recognition unit, 243... Sensor recognition unit, 244... Communication control unit, 245... Perception information generation unit, 246... Action control unit, 300... Relay server, 400... Robot, 500... Microphone / Speaker, 600... Camera, 700... Light emitting unit, 800... Network

Claims

1. One or more dialogue agents capable of interacting with a person, An operation intention estimation unit that estimates the operator's operation intention according to an operation for modifying operation control information describing the action content of the one or more dialogue agents, An action control unit that controls the dialogue by the one or more dialogue agents according to the operation intention of the operator estimated by the operation intention estimation unit, A dialogue system comprising:

2. The one or more dialogue agents, Perform a dialogue with content according to the operation intention of the operator, The dialogue system according to Claim 1.

3. When the operation control information is modified by the operator, an operation control information modification control unit that provides the modified operation control information to a control device that controls the one or more dialogue agents, Further comprising, The dialogue system according to Claim 1 or 2.

4. The operation intention estimation unit, As an operation for modifying the operation control information, when any one of character string input, character string confirmation operation, mouse operation on the interface, registration or deletion of character string information is performed, the operation intention of the operator is estimated by the executed operation or a combination of the executed operation and other information, The dialogue system according to Claim 1 or 2.

5. The operation control information modification control unit, Provides the modified operation control information to a control device that controls the one or more dialogue agents in real time or at a specific timing, The dialogue system according to Claim 3.

6. Estimate the operator's operation intention according to an operation for modifying operation control information describing the action content of one or more dialogue agents capable of interacting with a person, A control method for controlling the dialogue by the one or more dialogue agents according to the estimated operation intention of the operator.