Devices and programs, etc.

The device addresses the limitations of conventional communication robots by controlling output information and utilizing speech recognition and dialogue history to provide enhanced interaction and cost-effective communication.

JP7859638B2Active Publication Date: 2026-05-15株式会社ユピテル鹿儿岛
View PDF 10 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
株式会社ユピテル鹿儿岛
Filing Date
2025-01-07
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Conventional communication robots lack sufficient capabilities for effective interaction and communication with users.

Method used

A device equipped with functions to control the generation and timing of output information, including voice and visual displays, to enhance communication capabilities, utilizing speech recognition, dialogue history display, and network connectivity for improved interaction.

Benefits of technology

Enhances the convenience and effectiveness of communication by allowing for responsive and engaging interactions through visual and auditory feedback, reducing reliance on external servers, and optimizing communication costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007859638000004
    Figure 0007859638000004
  • Figure 0007859638000005
    Figure 0007859638000005
  • Figure 0007859638000006
    Figure 0007859638000006
Patent Text Reader

Abstract

To provide a device which has a function to perform communication, for example.SOLUTION: A robot 1 has a function to output voice and a function to perform communication with a user. The robot 1 interacts with the user by using an interaction engine and while being in interaction converts a content that has been spoken by the user immediately before that into text data and displays it in a touch panel part 7 that serves as a display part. Thus, the user can visually confirm the content that he / she has spoken, which contributes to correction or directionality of subsequent communication.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an apparatus, a program, etc. having a function of performing communication or the like, for example.

Background Art

[0002] Patent Document 1 discloses a technology related to an interactive communication robot.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, there has been a problem that conventional communication robots do not have sufficient capabilities. Therefore, an object is to provide an apparatus, a program, etc. having capabilities superior to those of the prior art. The object of the invention of the present application is not limited to this, and the applicant also has the intention of obtaining rights by divisional application, amendment, etc. for a configuration aimed at obtaining an effect resulting from a part of the configuration disclosed in this specification and drawings, etc. For example, in this specification, a part where it is described as "can" is read as "is a problem", and the problem is disclosed in this specification. The problems are described as independent ones, and the applicant also has the intention of obtaining rights by divisional application, amendment, etc. alone for a configuration for solving this problem. Even if the problem is tacitly grasped from the description of the specification, the applicant has the intention of making a part of the configuration described in this specification the scope of claims by amendment or divisional application. Also, problems combining these independent problems are disclosed.

Means for Solving the Problems

[0005] (1) A device having a function to communicate by outputting output information to at least one of a user or another device, wherein the device may also have a function to control the generation of the output information for the communication or a function to control the timing of outputting the output information for the communication. In this way, at least one of the user or other device can obtain at least one of either output information whose generation is controlled or output information whose timing is controlled. This allows for the provision of a device that is superior to conventional devices. The device may, for example, generate output information corresponding to the communication as a control for generating the output information for communication. This increases convenience, especially when users or other devices communicate with the device. Communication can be carried out by outputting any kind of output information, but it is particularly good to store and use the history information of past communications as a basis for communication. It is also good to use the history information of communications with multiple different users or other devices as a basis for communication. In particular, the output of output information may be carried out from one output means, but it is also good to have a configuration that allows for output from multiple different output means, and to select one of these output means to perform the output. Examples of output means include voice output means, display means, communication means, etc. Communication may be carried out by appealing to sound, sight, or other senses, such as touch.

[0006] The "device" should ideally include a function to output information as audio. This would allow, for example, control of a device that operates using audio input. The device should ideally include an interface for communication and a means for making decisions to execute communication. The components of the "device" may consist of multiple enclosures, but it is preferable to use a single enclosure. Furthermore, the device should ideally be a system that has the ability to access a network, for example, via wired or wireless connections. In particular, it should be a smartphone, tablet, smart speaker, smart camera, etc. While there are no limitations on its appearance, it should ideally be a device that clearly communicates with others. In particular, it should be a robot. For example, a robot that mimics a human or animal, or a robot with another anthropomorphic form, would be especially good.

[0007] The device should ideally have an input interface for communication. This could be an input device such as a keyboard, or an interface with optical character recognition (OCR) that reads characters and converts them into data. However, it is preferable to have an input interface that uses voice. For example, the device should have a function to acquire voice data based on voice signals converted into electrical signals by a microphone. The output interface should preferably include an audio interface, such as a speaker or earphone. The output interface should preferably include a visual interface, such as a display with changeable content, such as a liquid crystal display (LCD), plasma display (PDP), organic EL display, or cathode ray tube display. It is good to have a device. It is also good to have output, for example, printed material. In particular, it is good to have something that actually generates movement as the output interface. It is good to have an actuator as something that actually generates movement. For example, it is good to have a motor. In particular, the device should be a robot equipped with components that actually generate movement. In particular, the device should output output information as the movement of components that actually generate movement. In particular, it is good to have audio, visual, and physical output interfaces as the output interface. A "user" is, for example, someone who can operate the device; it could be just one person, but it's better to have multiple users. "Other devices" may be the same as, or different from, one of the specific devices mentioned above, for example, in appearance, function, etc. Other devices do not have to have the function to output audio, but it is preferable that they have the function to output audio. Also, other devices do not have the function to input audio, but it is preferable that they have the function to input audio. Other devices may be those that cannot access the network, but it is preferable that they be those that can access the network. In particular, it is preferable that they be those that can access the internet. The output information should be configured to produce a certain output from the output means. This "certain output" could be, for example, a notification to the outside world. Even if it's not in the form of an obvious "notification," the mere fact that some change occurs as a result can be interpreted as a "notification." The "certain output" does not necessarily have to be intended as a notification. For example, it could be an output of sound or light that contains some kind of information, or no information at all, or it could be a change in some physical quantity, the movement of an object, etc.

[0008] (2) The device may be equipped with a display function that changes the display mode on the display unit in response to the voice communication.

[0009] When communicating by voice, the display can change its display mode in relation to the voice communication, allowing users to communicate by converting voice into a visual display, thus increasing the convenience of communication. The "display unit" may be a device that changes the display mode on the display unit in response to the aforementioned voice communication, and may be a display device such as a liquid crystal display (LCD), plasma display (PDP), organic EL display, or cathode ray tube. In particular, it is preferable to have the display unit in the device. "To change the display mode on the display unit in response to the aforementioned voice communication" may mean either displaying the voice of the user or another device on the display unit and changing its mode, or displaying the voice of the device itself on the display unit and changing its mode, or either one alone, but it is particularly good to have both, and in this case, either one or both may be displayed, but it is good to have both a state where only one is displayed and a state where both are displayed, and to have a switchable configuration. For example, an example of displaying only one is the mode when the face screen S1 is displayed in the robot 1 of Embodiment 1 below, and an example of displaying both is the mode when the chat screen S2 is displayed in the robot 1 of Embodiment 1 below. In terms of display methods, for example, the screen can be changed in various ways in relation to sound. For instance, an animation could be executed where objects displayed on the screen move in response to sound. For example, an image could be changed in accordance with the sound output, or another image could be overlaid on an image in response to changes in sound. Alternatively, for example, sound data could be converted into text data and displayed. The display of this text data should change moment by moment in response to changes in sound. The voice communication is controlled by a function that controls the generation of the output information or a function that controls the timing of outputting the output information for the communication. It is preferable to control the communication, and in particular, the communication is preferable to control by a function that controls the generation of the output information and a function that controls the timing of outputting the output information for the communication. Furthermore, the function for changing the display mode on the display unit may be controlled by a function for controlling the generation of the output information or a function for controlling the timing of outputting the output information for communication. In particular, the communication may be controlled by a function for controlling the generation of the output information and a function for controlling the timing of outputting the output information for communication. Similarly, in (3) and subsequent sections, the configuration for outputting from the device may be controlled by a function that controls the generation of the output information or a function that controls the timing of outputting the output information for communication, and in particular, the communication may be controlled by a function that controls the generation of the output information and a function that controls the timing of outputting the output information for communication.

[0010] (3) The change in the display mode on the display unit should preferably be based solely on the speech of at least one of the user or the other device.

[0011] By changing the display on the device based on speech from the user or other devices, users and other devices can visually confirm whether their speech is being recognized by the device, thus improving the convenience of communication. Furthermore, it creates a unique form of communication using a different type of human interface, which can feel fresh and interesting. Speech may be uttered directly by the user or another device, or it may be used indirectly, for example, by converting the utterance into text data. Alternatively, the utterance may be converted into some corresponding information, such as other sounds or visualized patterns, and the display unit may change its display mode based on that. (3) is based on "speech from at least one of the user or another device," so for example, the device itself should have a voice communication function. A configuration based on the utterance of at least one of the user or the other device may include, for example, a configuration that uses a speech recognition function to convert speech into text and identify the content of the utterance of at least one of the user or the other device.

[0012] (4) The change in the display mode on the display unit is performed alternately with the audio output from the device. This allows for stable and mutually beneficial communication, rather than one-sided communication. "Reciprocal" means, for example, that the communication between the device and the user or other devices proceeds in a dialogue format, and that it is not a configuration where only one party speaks unilaterally.

[0013] (5) The display mode may include character information obtained by converting the speech of at least one of the user or the other device.

[0014] In this way, users and other devices can see specifically how their speech is being recognized by the device through the displayed text information, and can judge whether communication is being done correctly based on the content of the displayed text information. They can also visually confirm what they have said. Furthermore, if the recognition is incorrect, they can repeat themselves or rephrase their words to guide the user toward correct communication. "Character information" refers to characters based on user or other device utterances, such as regular kanji, hiragana, and katakana, in the case of Japanese. If the utterance constitutes a sentence, it should preferably be a sentence containing phrases mixed with kanji, hiragana, katakana, and foreign language characters. If the utterance is in a foreign language, such as English or Chinese, it should be displayed in those characters. stomach. For example, it is particularly preferable to have a voice recognition function and a voice output function, which allows the voice recognition function to recognize the content of speech spoken by at least one of the user or the other device and convert it into text information, to display the content on the display unit, to generate response text information based on that content, and to convert the response text information into voice information using a speech synthesis function and output it as voice.

[0015] (6) The characters obtained from the utterance should be displayed on the display unit simultaneously, showing the entire content from the start to the end of the utterance. Since the entire utterance with a certain length from the user or other devices can be received by the device side, correct communication from the device for that utterance can be expected. Also, the content of one's own utterance can be visually confirmed instantaneously, which contributes to the correction and direction of subsequent communication. It is advisable to set the start to the end of the utterance considering the time for a person to utter continuously. For example, the end of the utterance may be considered as the point when a sound volume, which is regarded as no voice utterance, continues for a predetermined time. The start of the utterance may be started on the condition that, for example, the sound volume exceeds a predetermined level. The start of the utterance may be started on the condition that, for example, a rapid change in the sound volume at a level above a predetermined level is detected.

[0016] (7) The device may be provided with a function of causing the display unit to display the dialogue history of the voice communication between the device and at least one of the user or the other device as character information.

[0017] It becomes easy to understand how the dialogue was conducted between the device and the user or other devices, enhancing the convenience in voice communication. The display of the dialogue history may be, for example, displayed so that it can be understood which party made the dialogue. For this purpose, for example, it is advisable to provide a balloon to distinguish which character information is based on which utterance. It is advisable that the past dialogue can be scrolled and confirmed on the screen. It is advisable to display different avatar characters on the device side and the user or other device side to indicate that it is a dialogue format. It is advisable that the date and time when the utterance was made are displayed simultaneously, which serves as an opportunity for the user to recall the past from the dialogue history.

[0018] (8) When the device displays the dialogue history as character information, if there is an alternation between the user or the other device and the dialogue target of the device changes, the device may be provided with a function of causing the display unit to display that fact.

[0019] When the device's target of interaction changes, the content of the interaction also changes. By displaying a notification that the target of interaction has changed, for example, when viewing past interaction history, the reader will see this notification and understand that the content of the interaction has changed before and after that point, thus making the break in the conversation clear. The device may be equipped with a function to detect a change in the user or other equipment. The detection of the change may be based on a change in the characteristics of the voice, but it is preferable to have a configuration that includes a function to acquire and detect the status of people or equipment in the surroundings using a camera, and in particular, a configuration that detects based on both.

[0020] (9) The device may be provided with a function that displays on the display unit of the device the status of voice recognition of at least one of the user or the other device by means of a voice recognition function.

[0021] For example, when a user or other device is speaking, the system can assure the user that it is definitely hearing what they are saying, thereby ensuring that the conversation is proceeding smoothly. For example, as a visual function, it would be good to change the color of the display screen or the objects displayed on the display screen according to the degree of voice recognition, or to display different objects according to the voice recognition status, or to display a quantitative value according to the degree of voice recognition, for example, showing a high value if the recognition is good. As an auditory function, it would be good to increase or decrease the volume according to the degree of voice recognition, or to change the tone, for example. It is particularly desirable to display the information in a personified form on the display unit, and to configure it so that the facial expressions change.

[0022] (10) The system is equipped with a function for performing the aforementioned communication by recognizing speech and converting it into a string, and it is preferable that the system be equipped with a function to output the corresponding content as speech if there is a part in the result of recognizing speech and converting it into a string that matches a string stored in a storage means that stores in advance the correspondence between the resulting string and the output content.

[0023] If the recognized and converted speech string matches any part of the string stored in the speech memory, and the necessary speech output can be generated solely within the device, it can respond quickly to speech from the user or other devices. Furthermore, since it does not connect to an external server, connection costs can be reduced. For example, it is good to have a large number of built-in scenarios as strings stored in a memory device. These built-in scenarios are pre-planned conversations, such as simple greetings like the device responding "How are you?" to the user's "Good morning," or scenarios for executing certain processes, such as "Open the settings screen" (user), "Are you sure?" (device), "Yes" (user), "Okay, I'll open the settings screen now" (device). Suitable memory devices include ROM or SSDs inside the computer, external SD cards, microSD cards, CD-ROMs, etc.

[0024] Here, "when there is a part that matches the string" refers to both cases where it perfectly matches a memorized string and cases where it is a regular expression where some parts may differ. A regular expression is a language processing method that represents a set of strings with a single string. For example, in the case of "xxx turn up the volume xx", even if there are differences in parts such as "Yupi-bo, turn up the volume," "Hey, turn up the volume," or "Please turn up the volume," if the essential part matches, it will be recognized as an expression that means "turn up the volume." And, "output the content corresponding to the string as audio" means, for example, in response to the string "turn up the volume," an audio output such as "Yes, I will turn up the volume" would be appropriate. Furthermore, the device may perform some processing after such an audio output. For example, after the utterance "Yes, I will turn up the volume" is made, the device may actually increase the volume of its own utterances in the subsequent dialogue. The process of recognizing speech and converting it to text may be performed by the device itself, or it may be performed by sending speech data to a speech recognition server connected to a network and receiving the converted text from the speech recognition server. Preferably, both methods should be provided, and it is desirable to have a function to decide which result to use depending on the communication situation, etc.

[0025] (11) If the result of recognizing speech and converting it into a string does not match any part of a string stored in a storage means that has previously stored the correspondence between the resulting string and the output content, it is preferable to have a function to connect to a server equipped with a dialogue engine and output the speech data.

[0026] If the string recognized and converted from speech does not match any of the strings stored in the speech memory device, it connects to a server equipped with a dialogue engine, thus reducing the cost of connection. Cut. The server may be a device that has the functions of a computer as a storage unit and control unit, for example, connected using an internet connection. In this invention, it is preferable that it is equipped with a dialogue engine. In the case of an external server, connection may be possible using, for example, an ID, password, or electronic authentication. The external server may be a cloud server. The server may be equipped with a speech recognition engine that can convert speech into text data. The converted text data may be transmitted to the device using an internet connection. A function that connects to a server equipped with a dialogue engine and outputs voice data may, for example, send a speech-recognized string to the dialogue engine, receive a string containing the corresponding dialogue content from the dialogue engine, and convert the string of dialogue content into voice data using a speech synthesis function. The conversion of a string into speech data may be performed by the speech synthesis engine installed in the device, but it is preferable to send the string to a speech recognition server that converts strings into speech data, and then receive the speech data corresponding to the converted string from the speech recognition server.

[0027] (12) Even if the result of recognizing speech and converting it into a string matches a string stored in a storage means that has previously stored the correspondence between the resulting string and the output content, it is preferable to have a function that connects to a server equipped with a speech recognition engine and outputs the speech data when certain conditions are met.

[0028] When the string of text recognized and converted from speech matches a string of text stored in the memory device, providing a predictable, fixed voice output can discourage the user from engaging in dialogue. Therefore, deliberately connecting to an external server in this way allows for a more human-like conversation, which is preferable. For example, if a user says "Hello" and the device recognizes it, and the built-in scenario is to have a conversation like "Hello, how are you?", it would be better to request the voice data "Hello" from an external server instead of using that scenario, and then request the external server's dialogue engine to create response data for that "Hello". The conditions for this could be, for example, a certain number of times every few times, or random timing.

[0029] (13) The system may include a function to detect when the audio is interrupted and becomes silent after speech recognition, a storage means for storing audio data from speech recognition until the period of silence, and a function to connect the audio data stored in the storage means to a server equipped with a speech recognition engine and output the audio data at the moment the period of silence occurs.

[0030] Dialogue often involves periods of silence. However, maintaining a connection to an external speech recognition engine server during these silences incurs unnecessary costs. Therefore, by pre-storing the audio data from speech recognition to the point of silence and sending this data only when silence occurs, rather than in real time, the silent portion of the audio can be eliminated, thus reducing costs.

[0031] (14) The device may be equipped with a recording function and a function to connect to a server equipped with a speech recognition engine and output speech data upon detection of speech at a predetermined sound pressure level. Constantly being connected to a server with an external speech recognition engine incurs unnecessary costs. This system reduces connection costs by only connecting to the external server when necessary for a conversation to begin, and not connecting when there is silence or near silence in the conversation. A server equipped with a speech recognition engine can process recordings from a predetermined period of time that have already been recorded on the device. The system could receive data and return a string corresponding to the recorded data, but it would be particularly beneficial to use a system that receives audio data in real time, for example as streaming data, and returns a string. While many systems charge a fee per hour of audio data reception, this approach can significantly reduce costs.

[0032] (15) When connecting to a server equipped with a speech recognition engine and outputting voice data, if the server is busy, it is preferable to provide a function that outputs a selected example dialogue from the dialogue data stored in the storage means to the user.

[0033] While it's common practice to notify the user when the system is busy, such notification in the middle of a conversation can be abrupt and seem unrelated to the dialogue, potentially disrupting the conversation. Therefore, instead of notifying the user that the system is busy, using a prompt like "Could you repeat that?" or a connecting utterance like "Oh, I see" allows the system to connect to the speech recognition engine and continue the conversation appropriately, without making the dialogue feel unnatural.

[0034] (16) If the recognized user utterance is determined to be too long, it is preferable to provide a function that outputs a selected example dialogue from the audio data stored in the storage means without connecting to a server equipped with a speech recognition engine.

[0035] If the user's utterance is too long, the speech recognition engine may misinterpret it, resulting in an irrelevant response. Therefore, if the sentence exceeds a certain length, it is common practice to eliminate this possibility and restart the conversation by selecting and outputting ambiguous interjections such as "Yes," "Really?", or "Is that true?", which allows for a more appropriate conversation to continue.

[0036] (17) In the aforementioned dialogue-based communication, if the user misses hearing the sound of the device, the device may re-output the previous sound based on a certain utterance by the user. For example, if you didn't hear or accidentally forgot what the device said just before, you can use phrases like "Say that again" or "Tell me that again" to get the device to repeat what it said. This allows you to continue the conversation you were just having without interruption.

[0037] (18) If the voice is not recognized, it is preferable that the device output a voice prompting the user to speak again. This makes it possible to continue the conversation that was taking place just before without interruption.

[0038] (19) It is advisable to display an item on the display unit when the recognized audio content meets certain conditions. For example, if a predefined word is spoken and recognized by speech recognition, the display unit will show a corresponding "specific display." The "predefined word" could be, for example, the user's birthday, the user's child's name, a nickname for the device, a company name, or a specific advertising slogan. The system should include a function to pre-set the correspondence between the predefined word and the corresponding display. This allows for communication that goes beyond simple dialogue to include visual communication, increasing the types of communication with the device and improving the convenience of communication.

[0039] (20) The device has a function to move the housing or a part connected to the housing, and recognizes voice If the contents meet certain conditions, the enclosure or the part connected to the enclosure should perform a certain movement. For example, when a specific word is spoken and recognized by speech recognition, the device or a part connected to the device is made to perform a movement, such as a gesture, corresponding to that word. This allows for communication that goes beyond simple dialogue and includes the movement of the device, increasing the modes of communication with the device and improving the convenience of communication. This is especially effective when combined with the "certain display" mentioned above.

[0040] (21) The device comprises an eye portion which is a part that the user can recognize as an eye, a user position recognition function which recognizes the position of the user, and a function which moves the eye portion, and the device may also include a function which moves the eye portion to face the direction of the user's position recognized by the position recognition function as communication.

[0041] When the recognizable eye portion faces the user's direction, it creates a simulated feeling of actually talking to a person, increasing the desire to communicate with the device and improving its usefulness. The "eye portion, which the user can recognize as an eye," may be an eye object displayed on a screen, or it may be an actual, mechanically operating eye rather than a virtual image. The device itself may be controlled to face the user's position and direction in synchronization with the eye portion. Alternatively, only the device for orienting the eye portion towards the user may be controlled to face the user's position and direction.

[0042] (22) The device may be equipped with a face recognition function that recognizes the user's face. Because it can identify the faces of individual people, it becomes possible to communicate in a way that suits each person's personality. For example, by associating the authenticated face of each person with their name, it becomes possible to address the person whose face has been recognized during a conversation by that name. It is especially effective to configure the system to conduct conversations specifically tailored to the person whose face has been recognized, based on past conversation history.

[0043] (23) The device may include a display unit and a function to display the status of the user's face recognition on the display unit by the face recognition function. This way, users can understand the status of their face recognition by the device by looking at the display. It is particularly useful to display whether face recognition is complete or whether the face has not yet been recognized as a person. In this way, if recognition is not yet complete, users can cooperate by trying to keep their face still as much as possible to make it easier for the device to recognize their face.

[0044] (24) The position recognition function may be a sound source direction identification function comprising three microphones arranged at the vertices of a triangle, and a determination unit that determines the position of the sound source from a position projected onto the plane including the triangle in a direction perpendicular to the plane including the triangle, to a reference position located inside the region enclosed by the triangle on the plane, based on the difference in the time it takes for sound to arrive from the sound source to each of the three microphones. This allows the direction of the sound source to be identified using three microphones. And once the direction of the sound source can be identified, the device can be directed in the direction of the user's speech, allowing the user to feel as if they are engaging in conversational communication.

[0045] (25) The device may be equipped with an infrared remote control signal output unit, and the communication may be communication with the other device equipped with an infrared remote control receiving function. This allows for easy communication between the device and other equipment via the device's infrared remote control signal output unit. For example, it becomes possible to output an infrared remote control signal from the device to a receiving device equipped with an infrared remote control signal receiving unit, such as a television, audio equipment, or air conditioner, and perform controls such as ON / OFF. In particular, it is preferable for the device to be equipped with a voice interaction function, and in response to the utterance of command phrases such as "Turn on the TV" or "Turn off the TV," the device controls the infrared remote control signal output unit based on that command.

[0046] (26) The device may be remotely controlled via the Internet from the other devices. The device's convenience is enhanced because it can be remotely controlled from other devices. For example, the other device could be a smartphone. The device should be equipped with a camera. The device should also be equipped with a mechanism to change the camera's direction. It would be good to be able to access it from other devices and, for example, view the camera video from the device or change the camera's direction. This makes it possible to control the device even when not nearby. Also, for example, it would be good to be able to access it from a smartphone and, for example, turn on the device's monitoring function so that it sends an email notification to the smartphone when a person (or object) moves. For example, for monitoring sick people or those receiving care, assuming they are constantly moving, it would be good to send an notification if, for example, the person has not moved for a certain period of time. The other device could be, for example, a tablet or a personal computer.

[0047] (27) The device may perform the voice output using text information transmitted from the other device via the Internet. For example, it would be good to have a function that reads aloud the text of received emails. By setting it up to receive emails from someone, the content of the email will be output as audio from the device, eliminating the need to visually check one's own device. Other devices could be smartphones, tablets, or personal computers, but they could also be other "devices".

[0048] (28) The audio output may be modified by changing the time, date, or number of times the audio output is performed depending on the content of the text information. For example, we want the email "I took my medicine" to be spoken at a specific time. If the subject line matches, the device can speak at a predetermined time, or if important information is spoken twice with a time interval in between, the user near the device can be sure that the email content is carried out correctly.

[0049] (29) The device may have a function to transmit the voice-recognized text information to the other device via the Internet. By using the device's voice communication function to convert speech into text and send it as text data to other devices, you can send emails, for example, without having to manually type the text into your own device.

[0050] (30) The system may have a function that sends a speech-recognized string as an input string to multiple different servers, each equipped with a dialogue engine that outputs an output string corresponding to an input string, receives the strings output as output strings from the multiple different servers, selects the longest output string, and initiates the dialogue. The longer the response, the more it feels like a real conversation, eliminating monotony and allowing the listener (user) to enjoy the interaction.

[0051] (31) The system may have a function to send a speech-recognized string as an input string to multiple different servers, each equipped with a dialogue engine that outputs an output string corresponding to an input string, receive the strings output as output strings from the multiple different servers, and select an output string with a question mark at the end to initiate a dialogue. Adding a question mark to the end of a sentence encourages further discussion that addresses the question, making the conversation flow more easily and allowing the listener (user) to enjoy the dialogue.

[0052] (32) It is preferable to have a function that transmits a speech-recognized string as an input string to multiple different servers equipped with a dialogue engine that outputs an output string corresponding to an input string, receives the strings output as output strings from the multiple different servers, and combines affirmative sentences and then interrogative sentences to initiate a dialogue. By arranging the response in this way, it sounds as if the person has carefully considered and crafted their words, making the user feel that their words are being listened to attentively and encouraging them to continue the conversation. This makes it easier for the conversation to continue and allows the listener (user) to enjoy the dialogue. In addition, it allows for a longer output time and prompts the listener (user) for a response.

[0053] (33) It is preferable to have a function that transmits a speech-recognized string as an input string to multiple different servers, each equipped with a dialogue engine that outputs an output string corresponding to an input string, receives the strings output as output strings from the multiple different servers, and combines them so that the string indicating a change of topic is placed at the end to facilitate dialogue. By arranging the conversation in this way, the topic is changed, which encourages further discussion and makes it easier to continue the conversation.

[0054] (34) The system may have a function to send a speech-recognized string as an input string to multiple different servers, each equipped with a dialogue engine that outputs an output string corresponding to an input string, receive the strings output as output strings from the multiple different servers, and combine them so that the friendly output strings are placed first to initiate a dialogue. By arranging the conversation in this way, the listener (user) becomes more easily drawn into the discussion, and the conversation is more likely to continue.

[0055] (35) The system may have a function that sends a speech-recognized string as an input string to multiple different servers, each equipped with a dialogue engine that outputs an output string corresponding to an input string, receives the strings output as output strings from the multiple different servers, and combines them in a random order to facilitate dialogue. As the variety of dialogue increases, even if the listener (user) says the same thing, they won't get the exact same response, making the conversation less boring and easier to continue.

[0056] (36) The system has a function to send a voice-recognized string as an input string to multiple different servers, each equipped with a dialogue engine that outputs an output string corresponding to an input string, and to receive the strings output as output strings from the multiple different servers, and if any of these output strings include emoticons, the system should not treat them as subjects for dialogue, and should display them together with the output strings that were treated as subjects for dialogue on the display unit. While emoticons cannot be output as voice, intentionally displaying them on the screen and incorporating them into the dialogue along with the voice can create an unusual and engaging way to interact.

[0057] (37) The system may have a function that sends a speech-recognized string as an input string to multiple different servers, each equipped with a dialogue engine that outputs an output string corresponding to an input string, receives the strings output as output strings from the multiple different servers, and selects only one of the output strings that contain the same string and combines it with other output strings to initiate a dialogue. Repeating the same string of characters makes the dialogue tedious and can make the listener feel uncomfortable.

[0058] (38) The system may have a function to send a speech-recognized string as an input string to multiple different servers, each equipped with a dialogue engine that outputs an output string corresponding to an input string, receive the strings output as output strings from the multiple different servers, and then combine the output strings after converting the word endings using a word ending conversion engine to facilitate dialogue. It's good because it uses more approachable language compared to the text produced by a typical dialogue engine.

[0059] (39) The system may have a function to send a speech-recognized string as an input string to multiple different servers, each equipped with a dialogue engine that outputs an output string corresponding to an input string, receive the strings output as output strings from the multiple different servers, store some of the output strings in a storage means without using all of them, and retrieve them from the storage means for use in subsequent dialogues. This feature can be used when speech recognition fails or when responses from external servers are slow to arrive, allowing for uninterrupted conversations and contributing to more natural dialogue.

[0060] (40) When using servers equipped with a speech recognition engine, it is advisable to use a mix of free and paid servers. This allows heavy users, especially those who frequently interact with the device, to save on server connection fees.

[0061] (41) The other device is a smart speaker, and the device is configured to output audio to the smart speaker and communicate with the smart speaker. Because smart speakers lack a display, their convenience is enhanced when used in combination with other devices. A smart speaker can be defined as a speaker with wireless communication connectivity and voice-controlled assistant functionality. Examples include Google Home, Amazon Echo, and LINE Clova. A smart speaker should have the ability to perform various functions and capabilities (skills). These functions can be executed by speaking to the smart speaker from a device that communicates by voice. When a device speaks, for example, the user might speak a phrase to activate the smart speaker's skill, the device would recognize the speech, save it as text data, and at a certain time, synthesize that text data into speech and speak it to the smart speaker to execute the skill.

[0062] (42) The other device is a smart speaker, and the device is configured to output audio to the smart speaker and communicate with the smart speaker. This allows you to activate smart speakers, launch skills, or make inquiries. It also enables you to automatically execute smart speaker skills at specific times or scheduled periods without having to manually activate the smart speaker.

[0063] (43) The other device is a smart speaker, and the device may have an adaptation function that adapts the audio output of the smart speaker. This is useful when you find it tedious to listen to the long, fixed answers from a smart speaker to your questions, or when you want to quickly review the content.

[0064] (44) The audio output of the device may have a function to read aloud web articles. You can hear the information simply by interacting with the device, without having to read any web articles.

[0065] (45) The aforementioned output should be made when the execution of a predetermined processing unit of robotic process automation is completed. Robotic process automation makes it difficult to understand the execution status of processing units, but the device By having the system produce a specific output corresponding to the execution of each processing unit, the processing status becomes easier to understand, increasing convenience.

[0066] (46) The aforementioned output may be a notification operation. The notification action indicates that a certain output has been produced. (47) When a computer running robotic process automation enters a state where it is waiting for user input, it will perform a notification action. This allows the system to inform the user that it is waiting for input and prompt them to proceed to the next step.

[0067] (48) The system comprises a client computer that performs robotics process automation and a server computer that gives instructions to the client computer to perform robotics process automation, and the client computer is configured to perform a notification operation when it receives an instruction from the server computer. This allows the server computer to notify the user of an instruction and prompt them to proceed to the next step.

[0068] (49) It is preferable to generate output information that points in the direction of a computer performing robotics process automation. This allows the user to see which computer the execution took place on, and enables prompting the user for the next action. (50) It is preferable to generate different output information depending on the execution status of the robotic process automation. This allows us to distinguish what kind of actions were performed.

[0069] (51) A program for a computer to implement the functions of any of the devices described in (1) to (50). Where "a certain output" or similar phrases use the word "a certain," it would be better to use a different word, such as "a predetermined output." The inventions described in (1) to (50) above can be combined in any way. For example, one may combine all or part of the configuration of the invention described in (1) with at least part of the configuration of at least one of the inventions described in (2) and onward. In particular, it is preferable to create an invention that combines the invention described in (1) with at least part of the configuration of at least one of the inventions described in (2) and onward. Alternatively, one may extract any configuration from the inventions described in (1) to (50) and combine the extracted configurations. The applicant of this application intends to acquire rights to inventions that include these configurations. Furthermore, even if there is a description such as "in the case of..." or "when...", it is not meant to be a configuration that is limited to that case or time. Configurations that do not fall under these cases or times are also disclosed, and the applicant intends to acquire rights to them. Also, even if there is a sequence of descriptions, it is not limited to that order. Configurations with some parts deleted or the order rearranged are also disclosed, and the applicant intends to acquire rights to them. [Effects of the Invention]

[0070] When communicating with users or other devices, the device can generate output information corresponding to the communication. Therefore, it enhances the convenience of communication between the user or other devices and the device. The effects of the present invention are not limited thereto, and the effects produced by the components of the structure disclosed in this specification and the drawings are also disclosed. The present invention intends to obtain rights to the components that produce such effects through divisional applications, amendments, etc. For example, the phrases "can do..." in this specification are descriptions that specify the effects produced, and there are components that produce effects even without such descriptions. Furthermore, there are effects that can be grasped by the component even without such descriptions. [Brief explanation of the drawing]

[0071] [Figure 1] A front view of the robot according to Embodiment 1 of the present invention. [Figure 2] A side view of the robot of the same embodiment 1. [Figure 3] Rear view of the robot of the same embodiment 1. [Figure 4] A block diagram illustrating the electrical configuration of the robot. [Figure 5] An explanatory diagram showing an example of a facial expression displayed on the face screen of a robot. [Figure 6] An explanatory diagram showing an example of a chat state with a chat screen displayed on the robot's face. [Figure 7] An explanatory diagram illustrating a standby image with the face screen displayed on the robot's face as the background. [Figure 8] An explanatory diagram illustrating a standby image with a chat screen displayed on the robot's face as the background. [Figure 9] (a) to (d) are explanatory diagrams illustrating the deformation patterns of the eye objects displayed on the robot's face. [Figure 10] (a) to (c) are explanatory diagrams illustrating how the user's speech gradually appears as a string of characters on the robot's face. [Figure 11] This diagram illustrates how eye objects move in accordance with the user's face on the face screen displayed on the robot's face. [Figure 12] An explanatory diagram illustrating an example of a smartphone. [Figure 13] An explanatory diagram illustrating the relationship between the robot's startup, wake-up mode, dialogue mode, and sleep mode. [Figure 14] An explanatory diagram illustrating the relationship between the robot and the smart speaker in Embodiment 7. [Figure 15] An explanatory diagram illustrating an example of a smartphone in Embodiment 9. [Modes for carrying out the invention]

[0072] <Embodiment 1> As shown in Figures 1 to 3, Robot 1, a communication robot that operates in response to human voices, has a housing consisting of a fixed part 2 which forms the lower body and a movable part 3 which forms the upper body and is mounted on the fixed part 2. The movable part 3 consists of a torso 4 positioned adjacent to the fixed part 2 and a head 5 supported by the torso 4. The fixed part 2 is formed in the shape of a bowl that opens upwards, and the torso 4 is formed in the shape of a cylinder with a continuous curve in the vertical direction from the upper edge of the fixed part 2. The connection point between the fixed part 2 and the torso 4 of Robot 1 is configured to have the largest diameter, and the outer shape narrows vertically from that connection point onward. The front of the torso 4 has a large semicircular cutout. The head 5 is fitted so as to be embedded in the upper part of the torso 4. The torso 4 rotates horizontally relative to the fixed part 2 (in the direction of the arrow in Figure 1), and the head 5 rotates vertically relative to the torso 4 (in the direction of the arrow in Figure 2) and in the left-right rotation direction (in the direction of the arrow in Figure 3).

[0073] The head 5 is constructed in a spherical shape, which is the remainder of a sphere with a portion (front portion) cut off by a single plane. The cut-out front portion appears as a circular shape and constitutes the face portion 6 of the robot 1. The rectangular portion formed on the surface of the face portion 6 is a touch panel portion 7, which is a liquid crystal display (LCD) with touch panel functionality. The content displayed on the touch panel portion 7 will be described later. A smoke panel is placed around the face portion 6 area surrounding the touch panel portion 7, giving the entire face portion 6 a uniform dark background. Inside the head 5, an illuminance sensor 8 and a high-brightness white LED 9 are located at the upper left and right housing positions of the face portion 6, respectively. On the face portion 6, the lens 11 of the face recognition camera 10 is located at the upper center position of the touch panel portion 7.

[0074] Within the torso 4, the front left and right positions and the rear central position of the torso 4 are offset by 120 degrees each. Microphones 12 are positioned at three locations at the same height. Inside the fixed part 2, a pair of left and right speaker devices 13 are positioned on the fixed part 2. An opening 14 is formed on the side of the speaker device 13 for outputting the sound generated by the speaker device 13. A power switch 15 and an up switch 16 and a down switch 17 for adjusting the volume of the speaker device 13 are positioned adjacent to the speaker opening 14. A USB OTG (On-The-Go) terminal 18, a DC12V power jack 20, and a microSD card socket (reader) 19 are positioned at the rear of the fixed part 2.

[0075] Robot 1 is capable of connecting to a designated external cloud server using an internet connection. The cloud server is equipped with a memory area for storing data required by Robot 1, and various engines for performing various processing tasks required by Robot 1. Therefore, in a broad sense, Robot 1 can be interpreted as a device that includes the software and other components of this cloud server.

[0076] Next, the electrical configuration of the robot 1 of Embodiment 1 will be described based on the block diagram in Figure 4. The controller MC, which serves as a control means, is connected to the touch panel unit 7, illuminance sensor 8, high-brightness white LED 9, face recognition camera 10, microphone 12, speaker device 13, terminal 18, and microSD card socket 19. In addition to these, a wireless LAN device 21, Doppler sensor 22, and first to third motors 23 to 25 are also connected to it.

[0077] The touch panel unit 7 has an input operation function that is performed by touching its surface. In the natural dialogue mode described later, the touch panel unit 7 can display different screens as shown in Figure 5 or Figure 6. The controller MC displays the face screen S1, which controls the facial expressions of the robot 1, particularly the changes around the eyes, as the first screen on the touch panel unit 7, as shown in Figure 5, in a displaceable manner. The face screen S1 is the screen that is displayed by default, and displays the eye object 27, cheek object 28, and elliptical region 29 of the robot 1. The eye object 27 has several deformation patterns as an animated image (Figures 9(a)~(d)). In addition, the pupil object 27a moves from side to side as an animated image.

[0078] Furthermore, the controller MC displays a chat screen S2, as shown in Figure 6, on the touch panel 7 as a second screen. The chat screen S2 can be displayed on the touch panel 7 in place of the face screen S1 by touching and sliding the touch panel 7 while the face screen S1 is displayed. The face screen S1 and the chat screen S2 can be switched between each other by sliding. The chat screen S2 will be described later. Furthermore, in standby mode (wake-up mode), the touch panel unit 7 can display the standby image shown in Figure 7 or Figure 8. The standby image will be described later. Robot 1 has a settings screen, which is a different image from those shown, and can be accessed from the chat screen S2. When Robot 1 is first started, you access the settings screen and enter necessary initial settings such as registering the ID and password to the server, registering the Wi-Fi password, user registration (e.g., name, age, gender, etc.), facial recognition, and setting the email address to which data will be transferred.

[0079] The illuminance sensor 8 recognizes the brightness of the environment in which the robot 1 is installed. The high-brightness white LED 9 automatically turns on based on the value detected by the illuminance sensor 8 if there is insufficient light for the face recognition camera 10 to take a picture. The microphone 12 is a voice input means that captures the user's utterances during interaction with the user, and at the same time, it uses three microphones 12 positioned at the vertices of a triangle simultaneously. This also serves as a direction detection means that can determine the direction of the sound source based on the difference in sound arrival times between these components. The controller MC calculates the arrival time difference from the phase difference of the electrical signals acquired by each microphone 12. Based on this arrival time difference, the controller MC calculates the sound source angle relative to the reference direction. A microphone specifically designed for interaction with the user may be provided, for example, on the face 6. The speaker device 13 is a voice output means used by the robot 1 to speak (output voice) during interaction with the user. The microSD card socket 19 reads and writes data to the inserted microSD card. The wireless LAN device 21 is a device that allows the Wi-Fi-enabled robot 1 to connect to the internet wirelessly. In this embodiment, the IEEE 802.11b international standard is used. The Doppler sensor 22 is a microwave-based sensor that emits microwaves and compares the frequency of the reflected microwaves with the frequency of the emitted radio waves to detect whether an object (person) is moving. It utilizes the Doppler effect, which causes a change in the frequency of the reflected wave when an object (person) is moving. For example, it is a device used to detect abnormalities around robot 1, such as the presence of a suspicious person when the user is absent. The first motor 23 is a servo motor for rotating the body 4 horizontally (in the direction of the arrow in Figure 1) relative to the fixed part 2. The second motor 24 is a servo motor for rotating the head 5 vertically (in the direction of the arrow in Figure 2) relative to the body 4. The third motor 25 is a servo motor for rotating the head 5 left and right (in the direction of the arrow in Figure 3) relative to the body 4. When the direction in which the user speaks is determined by the microphone 12, the controller MC controls the first motor 23 to rotate the body 4 (movable part 3) relative to the fixed part 2 so that the face part 6 faces the user directly.

[0080] The controller MC consists of a well-known CPU, memory such as ROM and RAM, SSD, bus, and real-time clock (RTC). The ROM of the controller MC stores various programs for executing the various functions of robot 1. Various programs are stored, such as a dialogue program for controlling interaction with the user via the microphone 12 and speaker device 13, a face recognition program for face recognition using the face recognition camera 10, a display variation / gesture program for changing the facial expressions and movements of the robot 1 during interaction by controlling the touch panel 7 and the first to third motors 23 to 25, a screen display program for displaying different screens or images on the touch panel 7 based on interaction with the user and operation of the touch panel 7, a data transmission / reception program for processing data acquired by the robot 1, such as camera images and emails from smartphones, etc., with other computers and smartphones, an absence setting program for monitoring when the user is absent, and an OS for operation, management, and control, including GUI functions, network connectivity functions, and process management. Input / output data for dialogue and face recognition are temporarily stored in RAM. Each program works in conjunction with other programs or independently to implement functions such as dialogue, face recognition, and gestures in a multitasking manner.

[0081] A. Regarding the actions taken during the dialogue In the configuration described above, the controller MC controls communication with the user through dialogue by executing a dialogue program. Face recognition, which becomes available in sync with the start of dialogue, will be described later in "B. Operation during face recognition" below. Here, the dialogue program is 1) A subprogram that sends a request to the cloud server with user speech data (voice data) acquired from microphone 12, and responds with the user's speech data (string data) converted into text using the server's speech recognition engine. 2) Execute built-in scenario dialogues based on user utterances (string data). Built-in Scenario Subprogram 3) If the user's utterance does not correspond to a built-in scenario, the utterance data transfer subprogram sends another request to the cloud server for utterance data (string data) and uses the dialogue API (Application Programming Interface) to cause the dialogue engine to create response data (string data) for Robot 1. 4) A subprogram for displaying the received response data (string data) on the touch panel unit 7, which acts as the display unit. 5) A voice data subprogram that converts the received response data (string data) into voice data using a speech synthesis engine and outputs it as speech from the speaker device 13 on the robot 1 side. 6) A display mode / operation variation subprogram that changes the display mode on the touch panel unit 7 and the operation of the robot 1 based on user-side string data and robot 1-side string data. Includes, etc. The following describes an example of the control content of the controller MC, primarily based on an interactive program, along with the interrelationships between the standby mode (wake-up mode), natural dialogue mode, and sleep mode after startup. These interrelationships are shown in Figure 12. Figures 5 and 6 show the screen in natural language dialogue mode, and Figures 7 and 8 show the screen in standby mode. In sleep mode, these screens are darkened as the backlight of the touch panel 7 is turned off.

[0082] 1. Startup The robot 1 is started when the power switch 15 is turned on (process M0 in Figure 13). The controller MC executes the boot program, and then when the OS starts up, the OS enters a state of waiting for a "command" from the user, i.e., a wake-up state. In this initial waiting mode, the waiting screen shown in Figure 7 is displayed. Please note that the following describes the state after initial setup is complete, i.e., when Robot 1's ID and password are registered on the cloud server, user registration is complete, facial recognition for multiple users is performed, and the smartphone's email address is registered with Robot 1, etc. "1. Effects during Activation" In this way, upon startup, one screen selected from multiple standby screens (Figure 7) is displayed first. In other words, a fixed screen is always displayed upon startup. Furthermore, the fact that Robot 1's eyes are closed (implying that it cannot interact) makes it easy for the user to understand that it is in standby mode.

[0083] 2. Standby mode (Wake-up mode) Standby mode is a state in which interaction with Robot 1 becomes possible when the trigger for starting natural dialogue mode is activated. It is also a state to which the user transitions when interaction in natural dialogue mode ends. Furthermore, if natural dialogue mode is not started for a certain period of time, the user will enter sleep mode. In sleep mode, communication with Robot 1 through interaction is not possible. Here, "natural dialogue" refers to the user interacting with the synthesized voice of the device (Robot 1) using the robot's built-in scenarios or the dialogue engine (dialogue software) on the server. Natural dialogue mode is a state in which natural dialogue is possible. There are multiple standby mode screens, and in this embodiment, two types are provided: Figure 7 and Figure 8. Figure 7 is the standby screen that transitions from the screen in natural dialogue mode shown in Figure 5 (processing M2 in Figure 13). It is also the standby screen displayed at startup by processing M0 in Figure 13. In Figure 7, the date, time, day of the week, and the current time are prominently displayed on the clock layer screen, and a layer screen of Robot 1's face screen S1, with its eyes (eye object 27) closed, is faintly displayed in the background. Figure 8 shows the standby screen transitioning from the natural conversation mode screen in Figure 6 (processing M2 in Figure 13). In Figure 8, the chat screen S2 layer is faintly displayed in the background of a clock layer that prominently displays the date, day of the week, and current time. In other words, it is in standby mode, but not natural conversation mode. "2. Effects in standby mode (wake-up mode)" In this way, different standby screens are available, making it convenient for users to interact using the screen they were previously accessing when the natural conversation mode is initiated from a particular standby screen. Furthermore, by displaying a screen specific to standby mode, users can easily understand that Robot 1 is in standby mode.

[0084] 3. Starting and stopping natural language dialogue mode From a state where the device is powered on and in standby or sleep mode, the controller MC processes the device to switch to natural dialogue mode at multiple timings, i.e., mode transition triggers, such as those shown below (processes M1 and M5 in Figure 13). The following triggers are just examples. In natural dialogue mode, the device switches to face recognition mode as described in "B. Operation during face recognition" below (face recognition becomes possible).

[0085] 1-1) In standby mode, if the controller MC recognizes a voice utterance from the microphone 12 within a certain period of time, such as "Hey, Yupi-bo," it will use this as a trigger to enter natural dialogue mode (processing M1 in Figure 13). Also, if it determines that a touch operation has occurred on the touch panel 7, it will use this as a trigger to enter natural dialogue mode (processing M1 in Figure 13). 1-2) In sleep mode, the controller MC determines whether a touch operation has occurred on the touch panel 7 of the standby screen at a predetermined timing. The touch operation on the touch panel 7 differs depending on the display mode on the face unit 6. On the standby screen of the face screen S1, it is initiated by touching the entire touch panel 7, while on the standby screen of the chat screen S2, it is initiated by touching the dialogue start button object 36, which will be described later. In other words, it is initiated by different operations on different screens. If the controller MC determines that a touch operation has occurred on the touch panel 7, it temporarily enters standby mode (process M3 in Figure 13), and then, if it determines that another touch has occurred, it enters natural conversation mode (process M1 in Figure 13). 1-3) In sleep mode, the controller MC determines whether a touch operation has occurred on the touch panel 7, similar to 1-2). If it determines that a touch operation has occurred, the controller MC temporarily enters standby mode (processing M3 in Figure 13). In this state, if a voice utterance such as "Hey, Yupi-bo" is recognized from the microphone 12 within a certain period of time as an activation phrase, the controller MC uses this as a trigger to enter natural dialogue mode (processing M1 in Figure 13). 2) In sleep mode, the controller MC determines whether a predetermined time set by the RTC has arrived, and switches to natural conversation mode when the predetermined time has arrived (process M5 in Figure 13). 3) In sleep mode, the controller MC outputs random utterances (voice data) from the speaker device 13 at random time intervals, for example, based on random numbers generated at predetermined timings. In other words, it outputs voices from the robot 1 that invite dialogue, such as "Hey, what are you doing?" or "I'm bored," as a kind of monologue, to enter natural dialogue mode (processing M5 in Figure 13) and prompt the user to speak. 4) In sleep mode, the controller MC uses the Doppler sensor 22 to determine whether an object (person) is moving, and switches to natural dialogue mode when it detects that an object (person) is moving (processing M5 in Figure 13).

[0086] 5) In sleep mode, the controller MC detects weather changes such as abnormal weather or earthquakes. If this occurs, the user is notified, and this triggers the initiation of natural conversation. An external cloud server acquires and stores information on abnormal weather conditions, including, for example, unusual weather (e.g., heavy snow, typhoons, etc.), earthquakes, and lightning strikes, at regular intervals using an abnormal weather detection engine, based on certain criteria. These regular intervals may all be the same, or the timing of acquisition may be changed depending on the type of weather. In this embodiment 1, abnormal weather information is distributed to the device (controller MC) using, for example, a push-type distribution system from the server. When the controller MC acquires the information, it enters natural conversation mode (processing M5 in Figure 13). 6) If the controller MC does not detect any user utterance for a certain period of time while in the natural dialogue mode described in 1) to 5) above, it enters standby mode (process M2 in Figure 13), and then enters sleep mode after a certain period of time (process M4 in Figure 13). The length of these mode transition times can be appropriately changed by, for example, the terminal device or by utterance as a built-in scenario.

[0087] "3. Effects of starting and stopping natural dialogue mode" By providing a variety of natural dialogue modes, users will have more opportunities to interact with Robot 1 at different times, leading to more opportunities to enjoy natural conversations and allowing users to appreciate the benefits of owning Robot 1. Furthermore, once a dialogue mode ends, the robot enters standby mode before going into sleep mode, thus reducing power costs. Furthermore, since the device skips sleep mode and standby mode directly to the natural conversation mode screen, conversation can begin immediately, resulting in a smooth start to the interaction. In addition, the conversation screen (Figures 5 and 7) is displayed as long as the conversation continues, which encourages the user to continue interacting.

[0088] 4. Dialogue in built-in scenarios in natural dialogue mode In natural dialogue mode, multiple dialogue processing options are available, including built-in scenario dialogues and normal dialogues using the server's dialogue engine. The controller MC first determines whether the speech data (string data) based on the user's utterance matches a built-in scenario, and if not, controls the system to use a dialogue engine via the cloud server (hereinafter referred to as "normal dialogue"). From the user's perspective, it appears as if they are always interacting with Robot 1, but in reality, there are multiple internal processes in natural dialogue mode. The controller MC performs a process that compares the user's utterance (string data) created by the speech recognition engine on the cloud server side with the text data of the built-in scenario (script). In this embodiment, the text data of the built-in scenario is stored in memory. Built-in scenarios may also be added to an SD card. If an SD card is used, it is easy to add built-in scenarios one after another by rewriting the data. When the controller MC recognizes the user's utterance, it determines whether the string data matches a predetermined regular or non-regular expression. If it does, it converts the script corresponding to the string data into speech data using the speech synthesis engine and outputs it as the utterance of Robo to 1 from the speaker device 13. Many built-in scenarios are provided, including simple scenarios such as greetings like "Hello" or "It's a nice day today" to prompt user responses, and scenarios that require processing based on user responses. Tables 1-3 show examples of such built-in scenarios. Of course, many more built-in scenarios are available in practice.

[0089] If the dialogue does not proceed according to the built-in scenario, the dialogue will terminate prematurely. Examples of situations where the dialogue does not proceed according to the built-in scenario include the following: It is correct. 1) If it does not match the intended regular expression or non-regular expression This refers to cases where the built-in scenario no longer matches the regular or non-regular expression from the beginning or midway through. It also includes cases where the user's speech is unclear and cannot be correctly captured. In this case, the controller MC determines that it is a normal conversation and immediately connects to an external cloud server. From then on, it requests speech data from the external cloud server and causes the conversation engine on the external cloud server to create response data in the form of string data. Then, the speech synthesis engine converts that response data into speech data and outputs it from the speaker device 13, continuing the conversation.

[0090] 2) If the dialogue in the built-in scenario is completed as scheduled. For example, when the system asks the user a question according to a scenario, such as "May I proceed with XXX?", and the user responds with a positive utterance like "Yes" or "Please," the built-in scenario dialogue ends as planned, or the dialogue ends midway through the scenario. In these cases, the system will enter standby mode after a certain period of time. 3) If the user's utterance is negative regarding whether or not to proceed with a certain process. If, when the system responds to a user with a negative response such as "No" or "That was a mistake" instead of a positive response such as "Yes" or "Please," the built-in scenario ends, and the subsequent dialogue proceeds as in 1) or 2). When a negative utterance occurs, the controller MC will ask a question such as "Are you sure?" to confirm whether it is okay to stop the process. This allows for handling of user slip-ups or changes of mind. For example, in a power-off scenario, when the utterance "Are you sure you don't want to turn off the power?" is made to the user according to the scenario, and the user utters "Yes", the question "Are you sure you don't want to turn off the power?" is repeated several times (for example, three times in this embodiment), and if the user utters "Yes", the dialogue in the built-in scenario ends.

[0091] "The effect of using a built-in scenario" With built-in scenarios like these, it's unnecessary to request all interactions from an external server; they can be processed internally by the device. This reduces communication costs associated with connecting to a server, and eliminates the need for communication time and server-side calculation time, preventing the user from experiencing awkward delays in responses that interrupt conversations. Furthermore, by providing such built-in scenarios for executing predefined processes, users no longer need to operate the touch panel 7 or access the robot 1 from another terminal to execute the process. Instead, the process can be executed through dialogue, making it user-friendly.

[0092] [Table 1]

[0093] [Table 2]

[0094] [Table 3]

[0095] 5. Requests and Responses in Normal Interactions On the other hand, if the utterance data (string data) is not a built-in scenario, the controller MC connects to the cloud server, sends the utterance data (string data) to the server again, and requests the dialogue engine to create response data. If the user is authenticated, the controller MC sends user-specific authentication information (e.g., ID and password) at the beginning of the utterance data in the request. In this case, past dialogue information is taken into account when creating the response data. On the other hand, if the user is not authenticated, no authentication information is sent as the user is not identified in the past, so past dialogue information is not taken into account. When the cloud server receives a request, it uses the dialogue API (Application Programming Interface) to have the dialogue engine create response data (string data) based on the utterance data (string data) and responds to robot 1 (controller MC). If there is a history of past user dialogue, the response data is created taking that content into account. The controller MC converts this response data into voice data and outputs it from the speaker device 13 as utterance from robot 1. Similar to the built-in scenario described above, the system will enter standby mode if the user remains silent for a certain period of time. "5. The Effects of Requests and Responses in Normal Dialogue" Unlike built-in scenarios, connecting to an external cloud server and conducting normal conversations allows for significantly larger amounts of data to be used for advanced conversation analysis, enabling highly sophisticated dialogues that closely resemble real-world human interaction.

[0096] 6. Actions of Robot 1 during dialogue The controller MC performs the following actions (a) to (d) when the interaction is taking place in natural dialogue mode, regardless of whether it is a built-in scenario or a normal interaction. I. Gestures of Robot 1 during startup, natural dialogue mode, and standby mode. The controller MC controls the first to third motors 23 to 25 of robot 1 at various timings to change the posture of robot 1. The following is an example. 1) At startup: If the face part 6 of head 5 is not facing forward or if head 5 is tilted, move it to the default position facing forward. 2) When the screen is touched: Same as 1) (to ensure that the face recognition camera 10 for face recognition is facing the user directly) 3) When the trigger utterance "Hey Yupi-bo" is spoken: Same as 1) (to ensure that the facial recognition camera 10 for facial recognition faces the user directly) 4) When sound direction is detected: The face portion 6 of the head 5 is turned in that direction. 5) As a special emotional utterance, for example when happy: The third motor 25 is controlled to rotate the head 5 in the left and right directions (clockwise and counterclockwise) while keeping the face 6 facing the user. 6) As a special utterance, for example in the case of sadness: the second motor 24 is controlled to keep the head 5 nodding for a while, and then return it to the default position. 7) When emotions are expressed as utterances, such as greetings like "Good morning," "Good afternoon," "Good evening," or "Hello": The second motor 24 is controlled to make the head 5 bow. 8) When uttering simple positive communication terms as non-special emotional utterances, such as "Okay," "I get it," "That's right," or "Yes!", control the second motor 24 to make the head 5 nod. In steps 6-8), the speed and timing of the second motor 24 can be changed to make sadness, bowing, and nodding different. 9) When an utterance of simple negative communication terms occurs as an ordinary emotional utterance, such as "no," "I can't," or "no way": The first motor 23 is controlled to rotate the movable part 3 from side to side several times. 10) Depending on the type of dialogue, including special or non-special emotional utterances, various gestures may be used, such as rotating the head 5 several times from side to side (clockwise and counterclockwise) or forward and backward, or combining this with rotating the entire movable part 3 from side to side, making large rotations or small nodding movements. The gestures of this robot 1 are best combined with the display on the following touch panel section 7 (face section 6). "6. Effects of Robot 1's Behavior During Dialogue (Part 1)" By having Robot 1 perform these kinds of gestures, users will feel a sense of familiarity with Robot 1, enjoy interacting with it, and find joy in actively engaging with it.

[0097] (b) Changes in the display mode when the face screen S1 is in a certain state. 1) When the user is speaking When the controller MC acquires and recognizes the user's speech from the microphone 12, it displays an animation on the face screen S1 of the touch panel unit 7, as shown in Figure 5, where an elliptical region 29 is displayed in blue and its area (i.e., size) changes according to the volume of the user's speech. Specifically, when the volume of the user's speech increases, the controller MC expands the elliptical region 29 while maintaining its elliptical shape, and when the volume decreases, it shrinks while maintaining its elliptical shape. It also displays the cheek object 28 in green.

[0098] Furthermore, the controller MC displays the user's speech data (string data) received as a response from the cloud server on the touch panel unit 7 in a predetermined manner. For example, as shown in Figure 5, when the touch panel unit 7 is the face screen S1, and it receives a utterance from the user, such as "Hello", the controller MC will access the layer of the face screen S1. - The text string "Hello" based on this utterance will be displayed on the screen. This display will begin before Robot 1's utterance, which is the response to the user's utterance. As an example of the display mode, the face screen S1 is displayed from a transparent state to an opaque state, as shown in Figure 10(a) and Figure 10(b), and finally, as shown in Figure 10(c), the face screen S1 in the background is completely hidden. In other words, the text is displayed gradually. After displaying only the text in the state shown in Figure 10(c), where only the text is displayed brightly against a dark background, for a very short period of time, the layer screen displaying the text is then gradually erased, as shown in Figure 10(c) → Figure 10(b) → Figure 10(a), returning to the default state of the face screen S1. At this time, all the text in a single utterance appears and disappears simultaneously. This display mode is just one example, and it may be displayed in a different manner. "6. Effects of Robot 1's Behavior During Dialogue, Part 2" This allows users to see what they've said on Robot 1, confirming whether Robot 1 heard them correctly, determining if the conversation is going smoothly, and guiding the conversation to avoid strange or irrelevant exchanges. Also, irrelevant conversations can be frustrating, but by checking, users can understand the reason and try changing their way of speaking and attempting the conversation again. Furthermore, since the user's speech data, both for built-in scenarios and normal dialogues, is requested to be converted into string data on the cloud server, there is no delay in the prerequisite processing for string data conversion. Also, since the creation of response data by the dialogue engine is requested only after this string data has been generated, the display of the user's speech on the touch panel 7 can be performed at least before the robot 1's response data is uttered, eliminating the risk of misinterpreting the order of dialogue.

[0099] When user utterances are converted into strings, their length varies depending on the utterance and is therefore not uniform. Furthermore, if the utterance is a "sentence" containing phrases rather than just words, it can become quite long. The controller MC adjusts the font size of the string so that even in the case of such long sentences, the entire content of a single utterance is displayed simultaneously on the touch panel 7. In other words, shorter utterances are displayed in a larger font, and longer utterances are displayed in a relatively smaller font. "6. Effects of Robot 1's Behavior During Dialogue, Part 3" This allows users to see what they say with a single glance, making it less likely for the conversation to be interrupted until the entire sentence appears, and reducing the chance of overlap with the next user's utterance. Also, because each utterance appears simultaneously, the entire sentence can be understood at once, and users can fully comprehend it even with a short display time. Furthermore, because the text is displayed across the entire touch panel 7, each character can be displayed larger, making it easier for the user to read, and allowing for sufficient confirmation even with a very short display time.

[0100] 2) When Robot 1 is speaking On the face screen S1 of the touch panel unit 7, as shown in Figure 5, an elliptical region 29 is displayed in red, and an animation is displayed that changes the area (i.e., size) of the region according to the robot's speech volume. Specifically, the controller MC expands the elliptical region 29 while maintaining its elliptical shape when the output level from the speaker device 13 increases, and shrinks it while maintaining its elliptical shape when the volume decreases. In addition, the cheek object 28 is displayed in light red. "6. Effects of Robot 1's Behavior During Dialogue, Part 4" In this way, the display on the face screen S1 changes depending on the alternating interaction between the user and robot 1, creating an interesting dynamic where the interaction takes place alternately not only in the actual conversation but also on the screen, leading to a more engaging conversation.

[0101] H. Changes in the display mode of the chat screen S2 The display patterns of the chat screen S2 on the touch panel unit 7, in accordance with user interaction, will be explained based on Figures 6 and 8. As described above, the display changes from the face screen S1 to the chat screen S2 in response to user operation. First, let me explain the configuration of the chat screen S2 again. As shown in Figure 6, a user object 31 is displayed as an avatar character in the lower left position of the chat screen S2, and a robot object 32 is displayed in the lower right position opposite it. Different user objects 31 are prepared for each authenticated user recognized by the face recognition mode described later, or for unauthenticated users, and different objects are displayed depending on the user currently in conversation. In the central area, speech bubble objects 33, which are string representations of the conversation between the user and robot 1, are displayed sequentially along the timeline. A conversation stop button object 34 is displayed in the upper left position of the chat screen S2. A settings button object 35 is displayed in the upper right position of the chat screen S2.

[0102] On the chat screen S2, speech bubble objects 33 are added in real time according to the dialogue between the user and robot 1. Each speech bubble object 33 displays the user's utterances, which have been converted into text data, and the robot 1's utterances, also converted into text data, in a single line along the timeline, making them available on the chat screen S2. The most recent utterance is displayed in the new speech bubble object 33, synchronized with that utterance, at the bottom of the row of past speech bubble objects 33. The speech bubble object 33 indicates the direction of speech so that it is clear whether the speech is from the user or from robot 1. Because it is not possible to display all dialogue history on the screen at once, the chat screen S2 is configured to be scrollable vertically, allowing users to view speech bubble objects 31 by going back in time. If users do not go back in time, the speech bubble object 33 of the most recent dialogue is always displayed. In this embodiment 1, once the conversation ends and the system enters standby mode, if the conversation resumes and the user recognized again in the facial recognition mode described later changes, the message "User Change" is displayed in the middle of the speech bubble object 33 column, and a different user object 31 is displayed according to the user that has been recognized again.

[0103] Furthermore, by touching the conversation stop button object 34 on the chat screen S2, the conversation is actively interrupted by the user, and the system enters standby mode. In this case, the standby screen of the chat screen S2 shown in Figure 8 is displayed instead of Figure 6, but the conversation start button object 36 is displayed in the place of the conversation stop button object 34. To return to natural conversation mode, the user can touch the conversation start button object 36 to return to the chat screen S2 shown in Figure 6. "6. Effects of Robot 1's Behavior During Dialogue, Part 5" The user object 31 of the user and the robot object 32 of robot 1, which are in a conversational relationship, are positioned facing each other, and the speech bubble objects 33 representing the conversation are lined up between them, giving the viewer a sense of actual conversation from the chat screen S2. Furthermore, since past chat history can be reviewed later, it can be used as a diary. It also shows who had what conversations, allowing users to see data such as who in the family uses it most often. A "User Change" indicator is displayed, letting users know when a conversation has been interrupted, preventing confusion when reviewing past history.

[0104] II. Special gestures in normal conversation In a typical conversation, the cloud server detects if the user's speech data contains specific words. If it determines that a special action is being performed, the controller MC responds with a command along with string data that causes the robot to perform a special action. Based on the above-mentioned screen display program, gesture program, and dialogue program, the controller MC performs a specific action, such as the following, using that command. The following control is just one example, and the controller may be controlled to perform other actions, or if there are multiple commands during the user's speech, the controller may be controlled to perform the actions consecutively or simultaneously. The following special actions may be performed individually or in combination. Instead of the gestures of Robot 1 in section (i) of "6. Actions of Robot 1 during Dialogue" above, the following displays may be shown on the display unit, or the displays on the display unit may be combined as appropriate.

[0105] 1) If a negative expression is uttered by the user during a normal conversation between people, the robot 1 will simultaneously animate the display on the touch panel 7, changing from a normal eye as shown in Figure 9(a) to an angry eye object as shown in Figure 9(b). In this embodiment, the eye object is stored in memory. 2) When the user utters an expression that would be considered fun in a normal conversation between people, the robot 1 simultaneously displays an animation on the touch panel 7, changing from a normal eye as shown in Figure 9(a) to a smiling eye object as shown in Figure 9(c). In this embodiment, the eye object is stored in memory. 3) If the user utters a sad expression during a normal conversation between people, the robot 1 will simultaneously animate the display on the touch panel 7 from a normal eye as shown in Figure 9(a) to a sad-looking eye object as shown in Figure 9(d). In this embodiment, the eye object is stored in memory. 4) When the user speaks the name of the user's child, the second motor 24 is controlled so that the head 5 of robot 1 makes a nodding gesture. In this embodiment, the program for eye gestures is stored in memory. 5) When the user speaks the name of the company that manufactures robot 1, the first to third motors 23 to 25 are controlled so that the head 5 of robot 1 moves forward, backward, left, and right, while the body 4 repeatedly oscillates against the fixed part 2 in a gesture motion. At the same time, the text data of the company name is synthesized into speech and the speaker device 13 repeatedly pronounces the company name as voice. In this embodiment, the gesture program is stored in memory. "6. Effects of Robot 1's Behavior During Dialogue, Part 6" These special actions allow users to anticipate unexpected actions from Robot 1 during the conversation, enabling them to actively enjoy interacting with Robot 1.

[0106] B. Actions performed during facial recognition 1. Starting and stopping face recognition mode The controller MC recognizes and authenticates the user's face by running a face recognition program. The face recognition program recognizes the acquired image as a human face by performing face pattern recognition, and then quantifies and stores various positions of the recognized face. It then determines the degree of match with numerical data of faces registered in the past to perform authentication. The controller MC switches to face recognition mode in sync with the natural language dialogue mode, and performs face authentication each time it switches from standby mode to natural language dialogue mode.

[0107] In face recognition mode, the controller MC uses the face recognition camera 10 to recognize the user's face. Specifically, 1) The face recognition camera 10 is activated. The user's face image captured by the face recognition camera 10 is displayed on the touch panel 7 (face display mode). In other words, the user is prompted to look at their own face on the touch panel 7. This enables face recognition processing, and by previewing the image in this way, it is possible to determine whether the person has been authenticated in the past. 2) If the face recognition camera 10 is unable to capture a face in step 1), and face recognition is not possible within a certain time period, In this case, the third motor 25 is driven to move the head 5 up and down. In other words, the face recognition camera 10 is made to scan in the vertical direction. Then, while the face recognition camera 10 is scanning in the vertical direction in this manner, the first motor 23 is driven to rotate the face recognition camera 10 360 degrees and perform face recognition. 3) If face recognition is successful in 1) or 2), authentication is performed. If the user is already registered, the user will use the data of a specific authenticated user in the conversation to enter the natural conversation mode described above. If the user is not registered, the user will be recognized as an unspecified person and the user will enter the natural conversation mode described above. The touch panel unit 7 will return from face display mode to either the previous face screen S1 (Figure 5) or the chat screen S2 (Figure 6). 4)2) If face recognition fails, the system will terminate the face recognition mode and the natural conversation mode itself, and return to standby mode, indicating that a person could not be recognized. The touch panel unit 7 will return from face display mode to either the standby screen of the face screen S1 (Figure 7) or the standby screen of the chat screen S2 (Figure 8), which was the previous standby mode. "1. Effects of starting and stopping face recognition mode" Since conversation is fundamentally about looking at the other person's face, the system actively encourages users to perform face recognition by preventing conversation if the face cannot be recognized. Therefore, conversation cannot take place unless the user is actually facing Robot 1 face-to-face, giving the user the feeling of having a real conversation.

[0108] 2. Actions of Robot 1 during facial recognition 1) After face recognition, the controller MC has the face recognition camera 10 acquire an image and continuously performs face pattern recognition at a fixed interval. The controller MC recognizes the user's face within the field of view of the face recognition camera 10 and sets the default position to a predetermined position within the field of view, for example, the position C of the center of the user's two eyes at the origin in the center of the field of view. If the position C shifts from this default position, the controller MC displays an animation in which the pupil object 27a moves in either the left or right direction depending on the amount of the shift. Normally, the pupil object 27a is fully displayed as an ellipse shape within the white of the eye object 27, for example, as shown in Figure 9(a). However, in a certain state where the user's face is moving, the pupil object 27a will be displayed as an eye object 27 with part of it hidden, as if it were looking in that direction, as shown in Figure 11. If the user moves and their face goes out of the field of view of the face recognition camera 10, making face recognition impossible, the controller MC drives the first motor 23 to rotate the entire movable part 3 in the direction in which the central position C has shifted, so that it faces the face recognition camera 10. Once face recognition is achieved, the first motor 23 is stopped. If face recognition is not achieved even after a certain amount of rotation, for example, rotating the entire movable part 3 by 45 degrees, the controller MC stops the first motor 23 at that point and continues face pattern recognition in that state. "2. Effects of Robot 1's Behavior During Face Recognition" This gives the user the feeling that they are always being watched while interacting with Robot 1, which enhances the enjoyment of the conversation.

[0109] 2) The controller MC constantly performs face pattern recognition with the face recognition camera 10, but face recognition is not possible if the user is moving or is not within the field of view of the face recognition camera 10. For this reason, it is good to notify the user of the face recognition status as a change on the screen. In Embodiment 1, the face recognition status is notified by a change in the intensity of the color of the sickle-shaped reflection object 27b, which represents the reflection on the pupil inside the pupil object 27a. In Embodiment 1, the controller MC displays a very light blue color when no face is recognized, a darker color when recognition is in progress, and a dark blue color when normal face recognition is occurring. "2. Effects of Robot 1's Behavior During Face Recognition, Part 2" This allows users to easily see whether or not their face is being recognized by Robot 1, encouraging them to actively cooperate in facial recognition and facilitating smoother conversations. Yes.

[0110] C. Operation when the system is set to be away from home In Embodiment 1, the absence setting mode is enabled, meaning that images can be forwarded to the email address registered during the absence setting. In this Embodiment 1, the absence setting mode is enabled and disabled on the settings screen that appears after the user touches the settings button object 35 on the chat screen S2 of the robot 1. The following describes the processing based on the absence setting program of the controller MC when the absence setting mode is enabled. In the standby mode described in B.4) of "2. Starting and stopping the natural dialogue mode" above in the absence setting mode, the controller MC, when it determines that an object (person) is moving based on the Doppler sensor 22, controls as follows:

[0111] 1) The controller MC will notify the user via email if it detects the presence of any moving object around robot 1. The controller MC will send an email containing the URL of robot 1's cloud server to the terminal device of a user who is not near robot 1 (hereinafter referred to as an external user), such as a smartphone, via the internet, whose email address is registered. The subject line and message body of the email will include wording that makes the purpose of the notification clear. For example, a phrase such as "It looks like someone is here" or an icon that indicates this. "C. Effects of operation when the system is set to be unattended" This means that, first, an email is sent, which notifies and allows the external user to recognize that there is some kind of object (person) moving around the robot 1, and gives the external user an opportunity to take countermeasures in response to this situation.

[0112] 2) The controller MC sends an email to an external user and simultaneously activates the facial recognition camera 10 to acquire an image. 3) The controller MC sends an email to the external user and simultaneously determines whether there is a trigger utterance within a certain time frame. For example, a greeting utterance such as "I'm home, Yupi-bo." Based on this utterance, a built-in scenario in natural dialogue mode is initiated, prompting the user (in this case, someone near robot 1 who uttered "I'm home, Yupi-bo") to perform facial recognition. If the controller MC determines, based on the facial recognition, that the user is one of the registered users, it sends a second email to the external user's terminal device. This email informs the external user that the person around robot 1 is not a suspicious person. In other words, the second email informs the user that they are a related person, such as a family member. The email includes the name of the registered user as information in the subject line and message body, based on the registration information. Alternatively, a trigger other than speech may be used, such as touching the touch panel 7 to perform facial recognition and confirm that the user is a registered user. "C. Effects of operation when the system is set to be unattended, part 2" This means that if a family member, such as a child, returns home while you are away, you will be notified by this second email, eliminating the need to check on the house via your smartphone while you are out.

[0113] 4)1) In case 3) a second email is not sent, the external user who receives the email can enter the ID and password of robot 1 to connect to the cloud server and view the camera image of the face recognition camera 10 provided by the cloud server in real time on their smartphone browser. This is also possible if a second email is sent. Figure 12 shows an example of a user's smartphone 41, and after connecting to the cloud server, a predetermined camera image from the face recognition camera 10 is displayed on its display screen 43, which also serves as a touch panel. Within the camera image are four elements for remotely controlling the orientation of the face recognition camera 10. Operation icons 44a to 44d are displayed. External users can operate operation icons 44a to 44d to output control commands to the controller MC via the cloud server. Based on the control commands, the first motor 23 or the third motor 25 is driven and controlled, causing the head 5 and torso 4 of the robot 1 to rotate and change the orientation of the face recognition camera 10. In addition, recording can be started by touching the record button icon 45 and stopped by touching it again. Furthermore, voice data spoken into the microphone (not shown) of the smartphone 41 is output from the speaker device 13 of the robot 1 via a cloud server, while voice data spoken into the microphone 12 of the robot 1 is output from the speaker device (not shown) of the smartphone 41 via a cloud server. As a result, external users can interact with users near the robot 1 using the smartphone 41 and the robot 1 while viewing images from the face recognition camera 10. "C. Effects of operation when the system is set to be unattended, part 3" This allows external users to remotely change the orientation of the facial recognition camera 10 to check the surroundings of the robot 1, for example, to check the safety of their home while they are away. Furthermore, if family members, such as children, return home while the user is away, proactively contacting them from outside using a smartphone contributes to maintaining good relationships with family and others.

[0114] <Modification 1 of Embodiment 1> Next, a modified example of Embodiment 1 will be described. In the built-in scenario dialogue in the natural dialogue described above, if the user's articulation is poor or other sounds are mixed in and the voice data acquired from microphone 12 does not match the regular or non-regular expression of the built-in scenario, the controller MC may not immediately determine that it is a normal dialogue, but may instead prompt the user to speak again by outputting a voice message such as "Please say that again." The controller MC causes the speaker device 13 to make such prompting utterances, and if the user makes a correct utterance in accordance with the built-in scenario within a certain time, it processes it again as a dialogue in the built-in scenario. On the other hand, even in such cases, if the user's utterance does not match the regular or non-regular expression in the built-in scenario, it connects to an external cloud server. This way, there is no need to unnecessarily connect to an external cloud server, and the interaction can be conducted solely within the robot 1.

[0115] <Modification 2 of Embodiment 1> Next, a modified example 2 of Embodiment 1 will be described. In the above natural dialogue scenario, if the user misses what Robot 1 says (utters), the user can make a request to repeat Robot 1's previous utterance within a certain time frame, causing Robot 1 to speak again (output audio). This process is possible in both built-in scenario dialogues and natural dialogues. The controller MC performs speech recognition during dialogue mode to determine if there was any utterance that could trigger a missed message, such as "Say that again." If it determines that the user has requested a repetition of the utterance after Robot 1 has spoken, it repeats what Robot 1 just said. It then cancels the previous utterance and processes the second utterance as the first utterance. This way, even if a user misses something during the conversation, the conversation will resume without interruption.

[0116] <Modification 3 of Embodiment 1> Next, a modified example 3 of Embodiment 1 will be described. In the above natural dialogue, if the user confirms the user's speech displayed on the touch panel unit 7 by the robot 1 and realizes that it has been incorrectly recognized, the robot may allow the user to point this out and correct the dialogue within a certain time. This process is possible in both built-in scenario dialogues and natural dialogues. During dialogue mode, the controller MC recognizes whether the user has made any utterances that trigger a correction of incorrect speech recognition, such as "That's wrong" or "No, I'll say it again." If the controller MC determines that such utterances occurred within a certain time after the user's utterances were displayed on the touch panel 7, it will output a voice prompt to encourage the user to speak again. For example, it might output the utterance, "Sorry, could you say that again?" and, (1) In the case of a built-in scenario dialogue, the user's previous utterance is canceled, and the content spoken by the user again is processed as the correct utterance in speech recognition. (2) In a normal conversation, the above-mentioned dialogue such as "That's wrong" or "Let me say that again" is also sent to an external cloud server as speech data (audio data), and a request is made to create response data with the content spoken by the user again, including such utterances. In this way, even if robot 1 incorrectly recognizes the user's utterance, it can be corrected to the correct dialogue.

[0117] <Modification 4 of Embodiment 1> A configuration different from that of Embodiment 1 may be adopted, for example, the following. (1) The above configuration was set up so that if the built-in scenario was not executed, a request would be made to the cloud server and the system would switch to normal interaction. In other words, if the built-in scenario was going to be executed, the system would always use the built-in scenario. However, even when supporting the built-in scenario, it is also possible to request the cloud server under certain conditions instead of handling it locally. These conditions could be, for example, executing it once every few times or at random intervals. This allows for unpredictable interactions with Robot 1, enabling more human-like dialogue that is not based on predetermined scenarios. (2) In the above embodiment 1, the controller MC did not have a speech recognition engine, but connected to an external server equipped with a speech recognition engine to convert the user's utterance (voice data) into text. This reduced the burden on the robot 1. However, the controller MC may be equipped with a speech recognition engine in its memory and call the speech recognition engine to create string data by itself converting the user's utterance data (voice data) acquired from the microphone 12 into text using the speech recognition engine. In other words, the controller MC of robot 1 may have the ability to convert the user's voice data into text data on its own. This would allow string data to be created without using the speech recognition engine, for example, reducing internal processing time. (3) In Embodiment 1 described above, the initial setup was performed from the settings screen, but it is also possible to register from an external source via a cloud server using a terminal device such as a smartphone. This would make the setup easier and save time, especially for people who are familiar with using terminal devices. (4) The first to third motors 23 to 25 may use other drive means other than servo motors. Other drive means include, for example, other types of motors or hydraulic cylinders. (5) In the above embodiment 1, the string data was synthesized into speech within the robot 1. Using an internal speech synthesis engine in this way is better than directly exchanging speech data with the server because it does not result in excessively large data. Alternatively, the response data (string data) obtained using a dialogue engine on the cloud server side may be synthesized into speech, and that speech data may be used as a response to the robot 1. (6) In the above embodiment 1, the configuration was such that the settings screen could be accessed from the setting button object 35 on the chat screen S2, but by sliding the touch panel 7, the settings screen could be accessed. You can also transition to a surface.

[0118] (7) In "c. Special actions in normal dialogue," even in built-in scenarios, if it is determined that a specific word is included in the user's utterance data, a special action may be performed. For example, if the controller MC determines that a specific word is included in the user's utterance data, it may be controlled to perform a special action in the same manner as described above. (8) In "c. Special gestures in normal dialogue," if the shape of robot 1 is different, it may be controlled to perform even more different gestures. For example, if robot 1 has hands or feet, the controller MC may control the drive means to move them. (9) As an animation for the eye object 27 on the face screen S1, you may occasionally add an animation that makes it blink. For example, this can be done by controlling it to insert the closed eye object 27 as shown in Figure 7 from the state of the eye object 27 as shown in Figure 9(a). Doing so will create a sense of realism as if the robot 1 is actually looking at you, and you will be able to enjoy interacting with the robot 1 more. (10) In face recognition mode, if a user is not registered, they are recognized as an unspecified person. It would be convenient if the user could then proceed to the settings screen and register as a new, different user. (11) In "C. Operation when the system is set to be away," it would be convenient to allow the system to be set to be away mode from a smartphone. (12) In "C. Operation when absent," the controller MC processes an email notification to the user if any moving object is detected around the robot 1. Conversely, it may also be configured to recognize that objects are moving at regular intervals and to notify the user via email if no moving objects are detected within a certain period of time. For example, if there is a sick person or someone requiring care, placing robot 1 nearby allows for constant monitoring based on the assumption that there is movement.

[0119] <Embodiment 2> Next, Embodiment 2 will be described. The high-brightness white LED 9 of the robot 1 in Embodiment 1 described above may be replaced with an infrared LED attached thereto. In this case, if the facial recognition camera 10 module is equipped with an infrared filter, it should be removed. Since infrared LEDs are invisible to humans, if there is an intruder such as a burglar at night, the high-brightness white LED 9 will light up, which may startle the intruder and cause them to run away. On the other hand, since it is difficult for the intruder to realize that they are being photographed with an infrared LED, they will not run away, making it possible to review and save images of the intrusion.

[0120] <Modification 1 of Embodiment 2> Next, a modified example 1 of Embodiment 2 will be described. If robot 1 is equipped with an infrared LED, this infrared LED may be used to control various devices in the room that are equipped with an infrared remote control signal receiver. Examples of such devices include televisions, audio equipment, and air conditioners. Robot 1 is preferably installed in a location where it can directly see the infrared remote control signal receiver of a device equipped with an infrared remote control signal receiver. In the modified example 1 of Embodiment 2, the controller MC is equipped with a shape recognition program related to shape recognition using the face recognition camera 10, and can recognize, for example, a television based on its shape characteristics (square, large, black, etc.). When the user makes an utterance such as "Turn on the TV" which serves as a trigger for outputting an infrared remote control signal to various devices, Robot 1 controls the first motor 23 or the third motor 25 based on that utterance to move Robot 1 up to the face recognition camera 10. The device is operated so that the face (5) is turned downwards and rotated 360 degrees to photograph the surroundings, and the shape recognition program is used to recognize the shape of the television. When the controller MC detects the presence of a television, it memorizes its direction and simultaneously outputs an infrared remote control signal from its infrared LED towards the object, executing controls such as turning the television ON or OFF. The next time a trigger for a television is received, it first performs shape recognition in that direction. The infrared remote control signal can be used not only for simple ON / OFF switching control, but also to control other functions by changing the infrared frequency, such as changing channels for a television or adjusting the temperature for an air conditioner. Such fine-grained control requires multiple infrared remote control signals with different frequencies, and the setting of the infrared frequency can be done via a server using, for example, the user's smartphone 41 as shown in Figure 12. Alternatively, instead of using a shape recognition program to determine the direction, the orientation of various devices can be obtained and registered by operating the face recognition camera 10 via the smartphone 41 to change its orientation.

[0121] <Embodiment 3> Next, Embodiment 3 will be described. Embodiment 3 describes a control method primarily aimed at reducing running costs associated with connections, such as the use of a dialogue API, when using a server equipped with a speech recognition engine. Furthermore, the dialogue program of Embodiment 3 includes a subprogram that can detect silence in the audio acquired from the microphone 12. The dialogue program also includes a recording / output subprogram that acquires the user's speech data from the microphone 12, records it, and outputs it to the server. The controller MC does not immediately connect to the server when it hears a utterance. Instead, it first records the user's utterance audio data, and only connects to the server and outputs the recorded audio data when it detects a period of silence in the recorded audio data, allowing the dialogue engine to create response data. In this way, it is not constantly connected to the server, and there is no need to connect to the server for long periods including periods of silence, thus eliminating silent connection time. In a speech recognition engine, one process is dedicated to the user while they are speaking. This means that a single process takes up a very long time for the computer, several seconds, which results in a high cost burden for the user using the speech recognition engine. However, as in Embodiment 3, the number of processes required per user can be reduced, contributing to cost reduction for the user.

[0122] <Modification 1 of Embodiment 3> Next, a modified example 1 of Embodiment 3 will be described. Modification 1 of Embodiment 3 also describes a control method primarily aimed at reducing the running costs of connections when using a server equipped with a speech recognition engine. If there is a time lag before the user starts speaking, or if the connection to the server is terminated due to a timeout because the user does not speak while the server is waiting for the user to speak, the speech recognition server consumes processes, resulting in costs for the user. The dialogue program of the robot 1 in the modified example 1 of Embodiment 3 includes a recording / output subprogram that acquires speech audio data from the microphone 12, records it, and outputs it to a server. The dialogue program also includes a subprogram that detects and outputs the sound pressure level of the recorded audio data. The controller MC connects to the server when the audio data being continuously recorded in a natural conversational state exceeds a certain sound pressure level, and plays back the data being recorded. In other words, it ignores silence or small utterances that cannot be recognized, and plays back only conversational utterances. Only when necessary will the system connect to the server, output audio data, and allow the server's speech recognition engine to create response data. This eliminates wasted connection time due to waiting for speech to be uttered.

[0123] <Embodiment 4> Next, Embodiment 4 will be described. Because speech recognition servers are expensive, it may not always be possible to prepare sufficient resources in advance. If server resources are insufficient, the server may be busy when a terminal attempts to connect to it. In the typical dialogue described above, the user sends utterance data (voice data) to the server and requests the creation of response data. The server then creates the response data, converted into string data based on the utterance data, and sends it back as a response. However, if the server is busy, the response data may not be generated, resulting in an error. The server will then send back an error message. When the controller MC of robot 1 receives an error message from the server, it should not announce to the user that there is a server connection error, but instead output a voice response from the built-in scenario that allows the conversation to continue. For example, it could give an ambiguous response such as "Say that again," "Uh-huh," or "What was it again?", or give an appropriate nod of agreement, while waiting for the server to become available.

[0124] <Modification 1 of Embodiment 4> Next, a modified example 1 of Embodiment 4 will be described. If the user's utterance is too long, the speech recognition engine may misrecognize it. Therefore, if the system determines that the recognized user's utterance is too long, it is desirable to have a function that outputs a selected example dialogue from the audio data stored in the memory without connecting to the server equipped with the speech recognition engine. When the controller MC of robot 1 determines that the user's speech data exceeds a certain length, it outputs a voice response from a built-in scenario that can be interpreted in any way during a conversation, such as "Yes," "Really?", or "Is that true?", without connecting to a server. The speech data can be in its original voice data form, or it can be converted into string data by the controller MC or the server. This prevents irrelevant responses and allows the conversation to be restarted, prompting the user to engage in further discussion.

[0125] <Embodiment 5> Next, Embodiment 5 will be described. Embodiment 5 describes a case in which multiple speech recognition engines are used in combination. There are two types of speech recognition engines: local engines (i.e., those that process data within the device without connecting to a network server) and cloud-based speech recognition engines that connect to a network server and respond with dialogue data created in response to requests. Both local and cloud-based engines have multiple variations, some free and some paid. Therefore, when using servers equipped with these different speech recognition engines, it's advisable to mix free and paid servers. In a cloud server to which robot 1 connects via an internet connection, when robot 1 issues a request for speech data and the speech recognition engine creates response data in dialogue mode, the cloud server should, for example, respond as follows: (1) Access to the paid conversational API is available up to a certain time A set per month. (2) From a set time A to a set time B each month, a mix of paid APIs and free speech recognition engines will be used. For example, the first few consecutive recognitions will be performed on the server of the paid speech recognition engine. For subsequent recognition, a mix of methods is used, such as using a free speech recognition engine server. (3) If a certain time B set per month is exceeded, use only the free speech recognition engine. This is just one example; for example, you could set it up so that if a certain time A set per month is exceeded, use only the free speech recognition engine immediately. In this way, interaction can be conducted without significantly exceeding the paid service limits. The cloud server connected to robot 1 processes requests from robot 1, calculating the balance between paid and free services based on a monthly time limit, using a program that performs this type of processing. Robot 1 itself may perform this process and, when issuing a request, instruct the cloud server whether to use a paid conversational API or a free server's speech recognition engine. <Embodiment 5-1> The same applies when using multiple speech recognition engines in combination as in Embodiment 5, or when using multiple dialogue engines in combination.

[0126] <Embodiment 6> Next, Embodiment 6 will be described. Using only one fixed dialogue engine may result in predictable responses, potentially causing the user to become bored with the conversation with Robot 1. Therefore, Embodiment 6 describes a process to resolve this issue by using the output results of multiple dialogue engines and arranging the results to prevent the user from becoming bored with the conversation. The dialogue engine may be a cloud-based dialogue engine, or it may be a local dialogue engine within robot 1. This process may be performed on the robot 1 side that receives multiple response data, or it may be performed on the cloud server side when response data is obtained from several servers equipped with dialogue engines. (1) Output the result of the chat engine that provided the longest response among the conversational engines. Allowing users to provide the longest possible response makes the conversation feel more genuine, preventing monotony and allowing them to enjoy the interaction. A. For example, if a user says "I'm hungry," and the responses from the three engines a-c are "Engine a: Do you often snack?", "Engine b: Haven't you eaten a meal?", and "Engine c: Eat something," then engine a will be selected and its response data will be output. B. For example, if the user says "The weather today is sunny," and the responses from the three engines a-c are "Engine a: I'll go to the sky right now and check," "Engine b: Looks like it's going to be clear?", and "Engine c: Sometimes whether it's sunny or rainy determines your mood for the day," then engine c will be selected and its response data will be output.

[0127] (2) In the casual conversation engine, the last thing to be output is the one that ends with "?". If there are multiple answers that end with "?", they are output consecutively. Adding a question mark to the end of a sentence encourages further discussion that addresses the question, making the conversation flow more easily and allowing users to enjoy the interaction. For example, option (1)A above would output "Do you often snack? Haven't you eaten a meal?". Also, option (1)B above would output "Does it look like a clear day?". (3) After combining affirmative sentences, combine them with interrogative sentences to produce the output. By arranging responses in this way, they appear as if the user has carefully considered and crafted their sentences, giving the user the feeling that their words are being listened to attentively. This makes them want to continue the conversation, making it easier for the conversation to flow and allowing the user to enjoy the interaction. Furthermore, it allows for increased output time and enables the request for a response from a person. For example, if response data like (1)B. above is obtained, the output will be: "I'll go to the sky right now to check. Whether it's sunny or rainy can sometimes determine your mood for the day, right? Looks like it's going to be sunny?"

[0128] (4) The results from engines that change topics more frequently than others are output later than the results from other engines. By arranging the conversation in this way, the topic is changed, which encourages further discussion and makes it easier to continue the conversation. A. For example, if the user says "Hello," and the responses from the three engines a-c are "Engine a: Right," "Engine b: That's right," and "Engine c: Do you watch baseball?", the output will be "Right, that's right, Do you watch baseball?" with the data from engine c placed last. B. For example, if the user says, "It's hard to find," and the responses from the three engines a-c are "Engine a: That's right," "Engine b: So true," and "Engine c: How many people are in your family?", the output will be "That's right, So true, How many people are in your family?" with the data from engine c placed last. (5) The engine that gives the friendliest response will be output first, and then the responses from the other engines will be added and output afterwards. This arrangement makes it easier for users to engage in the conversation and encourages it to continue. Whether a word is friendly or not can be determined by assigning a relative hierarchy to the words used, allowing for their placement in the conversation. A. For example, in the case of (4)B. above, the result of engine b is used first and the output is "That's so true! How many people are in your family?" B. For example, if the user says "I see," and the responses from two engines, a and b, are "engine a: Oh, that's a suitable response," and "engine b: Hmm," then the data from engine b will be placed last, and the output will be "Hohoho, that's a suitable response, hmm."

[0129] (6) Combine the processes from (1) to (5) as you see fit. This increases the variety of dialogue, preventing users from receiving the exact same response even when they say the same thing, thus preventing them from getting bored with the conversation and making it easier to continue the conversation. For example, if a user says "What is old age?", and the responses from the three engines a-c are "Engine a: Please wait a moment," "Engine b: I don't dislike support," and "Engine c: Are you healthy now?", the most friendly engine b will be used first, combined with affirmative statements, then combined with questions to output "I don't dislike support. Please wait a moment. Are you healthy now?".

[0130] (7) In the case of the engine used for text output, parentheses and emoticons may be returned, so if these are returned, the audio output will be suppressed and output will be suppressed. These will then be displayed on the screen. While emoticons cannot be output as voice, intentionally displaying them on the screen and incorporating them into the dialogue along with the voice can create an unusual and engaging way to interact. A. For example, if the user says "I'm hungry," and the responses from the three engines a-c are "engine a: (intentionally ignored)," "engine b: Hello, how are you?", and "engine c: Hello," then only engine a will not output any voice, but will instead display it on the touch panel unit 7 (display screen). B. For example, if the user says "it's already included", then there are three engines a-c. If the responses from N are "a engine: Do you do that often? b engine: (´·ω·`)" and "c engine: Are you a night owl?", then only b will not be output as audio, but will be displayed on the touch panel unit 7 (display screen). (8) If the same string is included in the response, output one of them. Repeating the same string of characters makes the dialogue tedious and can make the listener feel uncomfortable. For example, if the user says "Chinese food," and the responses from the three engines a-c are "Engine a: Oh, that's nice," "Engine b: Yes, it's Chinese food," and "Engine c: Are you going to eat Chinese food?", then since both engine b and engine c contain the string "Chinese food," only one of them will output. For example, it might output "Oh, that's nice. Are you going to eat Chinese food?"

[0131] (9) Use a means of converting sentence endings, such as a sentence ending conversion API, to ensure consistency. In this case, sentence ending conversion may be performed on all responses, or it may be performed only on either the last output sentence or the first output sentence. It's good because it uses more approachable language compared to the text produced by a typical dialogue engine. For example, if a certain engine responds with "I'm sleepy," the ending can be converted using a word ending conversion API to output "I'm sleepy meow." (10) In preparation for recognition failures, only a portion of the responses obtained from multiple engines will be used for speech output, and the remaining responses will be stored. If the next speech recognition fails or there is no response from the dialogue system, the stored responses will be returned. This feature can be used when speech recognition fails or when responses from external servers are slow to arrive, allowing for uninterrupted conversations and contributing to more natural dialogue.

[0132] <Embodiment 7> As shown in Figure 14, Embodiment 7 is a device (system) that combines the robot 1 and the smart speaker 51, with the smart speaker 51 placed near the robot 1. The distance between the robot 1 and the smart speaker 51 should be such that their microphones can pick up sound from each other, for example, they should be placed adjacent to each other within 1 to 2 meters. The smart speaker 51 is a network terminal equipped with a network module that has a built-in wireless LAN device, wireless communication capabilities using the internet, telephone line connection capabilities, etc., and is also a type of computer equipped with a microphone and speaker device. The smart speaker 51 uses a terminal device like a smartphone to perform various initial registrations via a server (for example, registration of the user's name, address, telephone number, email address, multiple voice registrations, setting up network-enabled AI devices via Bluetooth, etc.), connects to the internet in response to a voice command from the registered user, uses the server's search engine to perform predetermined processing, and outputs the results as voice information from the speaker device.

[0133] Robot 1 and smart speaker 51 can complement each other's functions by working together. Specifically, robot 1 and smart speaker 51 perform the following functions using voice as an interface. (1) Instruction function from robot 1 to smart speaker 51 (i) For example, when a user speaks a phrase to activate a smart speaker skill, the robot stores the string of characters resulting from speech recognition of that phrase, and the robot then speaks that string of characters at a predetermined time using speech synthesis. In Embodiment 7, the controller MC of the robot 1 includes a voice recognition program that analyzes the frequency components of the user's speech to identify the individual's voice, and a recording and playback program that distinguishes the user's speech for each individual, acquires it via a microphone 12, stores it as string data, and synthesizes a phrase based on that string data to be played back from the speaker device 13. Robot 1 has a function to memorize trigger utterances, such as the words that follow an utterance like "I'm going to speak now, so record it." The user then uses this function to memorize words that will activate or perform some action on the smart speaker 51. For example, words like "OK, xxx. Turn on the lights." are suitable. In this case, the user's personal voice registered in Robot 1 is the same as the user's personal voice registered in the smart speaker 51. Then, the robot 1 is made to speak at predetermined times. The settings for speaking at predetermined times can be registered by operating a device such as a smartphone. (b) The "speech at a predetermined timing" of robot 1 includes, for example, robot 1 detecting some kind of change, such as a touch operation on the touch panel 7 or detection of an object (person) by the Doppler sensor 22. For example, when the controller MC of robot 1 detects a person using the Doppler sensor 22, it outputs a voice message from the speaker device 13 such as "OK, xxx. Turn on the lights." In response, the smart speaker 51 controls the room's lighting, which is a network-enabled AI device, to turn on. The lighting to be controlled is pre-registered as an object to be controlled by the smart speaker 51. In addition to lighting, other AI devices such as air conditioners, televisions, and curtain openers can also be used.

[0134] (2) A function to distinguish between speech from the smart speaker 51 and speech from the user. By identifying the user's individual voice and registering it in the settings of robot 1, robot 1 can be equipped with the function to distinguish between speech coming from the smart speaker 51 and direct speech from the user. This allows robot 1 to be controlled so as not to react to sounds from the smart speaker 51 that are not registered in the settings. Conversely, by registering the user's individual voice in both the smart speaker 51 and robot 1, as in (1), instructions can be given to the smart speaker 51 not only by the user but also by robot 1. This ability to distinguish individual speech also prevents interference with voice control of the smart speaker 51. (3) Display function for speech from the smart speaker 51 Alternatively, the robot 1 may acquire the content of the user's instructions given to the smart speaker 51, convert it into text, and display it on the touch panel unit 7. The controller MC of robot 1 should ideally have a program that acquires spoken content and converts it into text, either by itself or by connecting to a server. For example, if a user says, "OK, xxx. Tell me today's weather," and the smart speaker 51 replies, "Today's weather in Okazaki City, Aichi Prefecture is sunny, with a high of 15 degrees Celsius and a 20% chance of precipitation," then the entire conversation between the two will be displayed on the robot 1's touch panel 7 as follows: "User: OK, xxx. Tell me today's weather." Smart speaker: Today's weather in Okazaki City, Aichi Prefecture is sunny, with a high of 15 degrees Celsius and a 20% chance of rain. Furthermore, it would be beneficial for the robot 1's controller MC to have a program that shortens or summarizes the textualized content. The robot 1 should then output the shortened or summarized content as audio or display it on the touch panel 7. (4) A function that allows robot 1 to ask various questions to smart speaker 51 and learn when no one is present. For example, robot 1 has built-in scenarios with question words such as "What's the weather like tomorrow?" or "Are there any incidents?", and when the user is away, the controller MC of robot 1 is instructed to output the question words along with a phrase to activate the smart speaker 51 at a predetermined time (assuming robot 1's voice is already registered with the smart speaker 51). At this time, the controller MC acquires and stores the content of the speech from the smart speaker 51 using the microphone 12, and at a predetermined time... The content of this will be output as audio. The predetermined timing is, for example, when a person is detected by the Doppler sensor 22 at a predetermined time, or when the user speaks to the robot 1 as part of a built-in scenario, such as "Is there any news?".

[0135] <Modification 1 of Embodiment 7> You may also configure the system to prevent interference with the voice operation of other voice recognition devices, such as smart speakers. For example, if robot 1 recognizes the voice of another smart speaker's activation phrase (voice recognition start word), it may be configured to stop its own voice output. For example, if the robot 1's controller MC recognizes the voice of an activation phrase, such as "OK, xxx," which is the activation phrase of another smart speaker from company A, it will acquire it from microphone 12 and, if it determines that it is a registered activation phrase, it will temporarily stop its own voice output. This allows the voice recognition device to function without interfering with voice control.

[0136] <Modification 2 of Embodiment 7> Robot 1 recognizes the voice recognition activation keyword of another voice recognition device, such as the smart speaker 51, and then recognizes the user's subsequent pronunciation to the smart speaker 51. It then sends a request to the cloud server, which performs a search using a search engine or similar method to obtain the corresponding answer. If robot 1 recognizes that smart speaker 51 has failed to perform speech recognition (for example, "Error") or that it cannot provide an appropriate response to the speech recognition result (for example, "Sorry"), robot 1 will output a response that it has obtained in advance. Alternatively, after receiving a voice output indicating that speech recognition has failed or that it cannot provide an appropriate response to the speech recognition result, robot 1 may request a response from the cloud server.

[0137] <Embodiment 8> Robot 1 may, for example, read aloud news articles from a website upon user request. The controller MC of Robot 1 requests news articles using a search engine on the server, and the cloud server responds to that request by providing, for example, news data from registered sites as text data. Robot 1 synthesizes and reads aloud (outputs) news data, and at the same time, it also synthesizes and reads aloud (outputs) the names of the information sources for the articles. In addition, the touch panel unit 7, which serves as a display screen, displays the URLs of the information sources for the articles, and when the user touches the URL on the touch panel unit 7, the content of the page at that URL is displayed on the touch panel unit 7. Furthermore, when reading aloud articles or their sources, it is best to output them in a way that clearly indicates they are quotations. For example, let's explain how to read aloud the content of an article at the URL "https: / / \\\\.jp / archives / 92###". "I'll read aloud an article from the 'XX News' website. It says, '..... will be available for purchase from January 15th of this year.'" In this way, for example, phrases other than the article or its source, such as "I'll read the article from this site" or "That's right," are made to be spoken in a way that the listener can understand, as if they were quotations using regular expressions. In this case, the screen displays a URL, for example, "This is information from https: / / \\\\.jp / archives / 92###." By touching this URL, the content of the article that was read aloud is displayed again on the touch panel 7. "Speaking in a way that the listener can understand" means changing your tone of voice or tone of voice between the article portion and the non-article portion. In this way, you can listen to what Robot 1 reads without having to read the web article. You can understand the news content just by looking at it, and in some cases, you can also visually check the news content just to be sure.

[0138] <Embodiment 9> Embodiment 9 describes a case where the invisible movements of a computer are shown through the operation of a robot's actuators. (1) Regarding robotics process automation Robotics Process Automation (hereinafter referred to as RPA) is software that automates simple computer tasks. The software can be configured on a server or on a user's computer. An example of a case where the software is configured on a server based on Figure 15 and the robot 1 of each embodiment described above is incorporated into the RPA network will be explained. As shown in Figure 15, the cloud server 55 and user computers 56 and 57 are connected via a network using the Internet. The cloud server 55 and robot 1 are also connected via the network. User computer 56 is controlled by the cloud server 55 through an RPA program. Robot 1 is also notified of processing information such as when a predetermined process in the RPA program executed by the cloud server 55 is about to be executed, is being executed, or has been executed. The cloud server 55 instructs the user-side computer 56 to process processes 1 to 4 in order. In this embodiment, processes 1 and 2 are executed by computer 56, and processes 3 and 4 are executed by computer 57. More processes may be set, and there may be one or more user-side computers 56 involved in the processing. Process 1 is, for example, user access and login to computer 56; Process 2 is, for example, creation and sorting of lists based on data in computer 56; Process 3 is, for example, correction of billing details for each customer, which is performed after Process 2; and Process 4 is, for example, issuance of invoices, which is performed after Process 3. In this embodiment 9, for example, in Process 3, the user is prompted to input for correction, and after that input, the system proceeds to the next Process 4. In other words, after Process 2, the system waits until the work in Process 3 is completed. The cloud server 55 should output different notification information to robot 1 immediately before, during, and after these processes are executed, and robot 1 should then notify those around it of the processing status based on this notification information. Alternatively, one notification per process would suffice.

[0139] for example, a. Robot 1 will notify users of the status of each process through voice, sound differences, music, etc. b. Announce the information on the display screen. This may be done simultaneously with a. c. Since user input is required in process 3, it is also possible to notify only process 3, or to make process 3 a different (identifiable) notification from the other notifications. d. Robot 1 transmits the processing status to other terminal devices for notification. e. Robot 1 operates in a way that allows the robot to understand the processing status. For example, if computer 56 is processing, it is controlled to face that direction. For this reason, it is preferable to recognize in advance the direction of each computer 56 and 57 relative to robot 1 using some direction-determining means, such as the shape recognition program described above. For example, robot 1 may be equipped with a direction-indicating member such as an arrow or an arm, and it may operate so that computers 56 and 57, which are the targets of notification, are in the direction indicated by the member.

[0140] (2) About blockchain Blockchain is a system where numerous computers record data in a distributed manner. In particular, in public blockchains, the data to be recorded and the recorded data are made public. Therefore, we will place Robot 1 in the blockchain network so that Robot 1 notifies the user before the block data is sent. It would be good to stop the transmission by having Robot 1 output a "wait" command (i.e., requesting that it wait without sending the data).

[0141] <Embodiment 10> While each embodiment has described the natural dialogue mode, it is also advisable to provide a foreign language learning mode instead of, or in addition to, the natural dialogue mode. When providing a foreign language learning mode in addition to the natural dialogue mode, for example, when the voice "Switch to foreign language learning mode" is recognized in the natural dialogue mode, it is advisable to switch from the natural dialogue mode to the foreign language learning mode. Conversely, when the voice "Switch to natural dialogue mode" is recognized in the foreign language learning mode, it is advisable to switch from the foreign language learning mode to the natural dialogue mode. It's often said that making foreign friends is a good way to learn a foreign language, but not many people have such opportunities. Therefore, it would be beneficial to have a foreign language learning function that acts as a conversational system, serving as a partner in foreign language learning. The system can be configured to link any first and second language. The following explanation will describe a configuration that links a Japanese dialogue system with an English dialogue system. The system should primarily involve conversations in English, with the ability to output requests in Japanese during the conversation, such as "Say that again." Furthermore, it should integrate with English text analysis web services to provide explanations in Japanese for English sentences in response to requests such as "Explain." If the user doesn't know how to say something in English during a conversation, they can request "Translate," which will output the English equivalent. By reading the outputted English, the user can continue the conversation. Since there is no need to look things up themselves, the conversation will not be interrupted, and smooth English conversation learning can be expected. In foreign language learning mode, both the speech recognition engine for the native language (e.g., Japanese for Japanese speakers) and the speech recognition engine for the foreign language (e.g., English) can utilize cloud-based speech recognition engines. However, for the native language (e.g., Japanese), since the requests are standardized and the format of the responses is fixed, it is preferable to install and use a speech recognition engine locally (e.g., within Robot 1). The same applies to the dialogue engine. The speech synthesis engine can also be installed anywhere, but it is particularly preferable to install it locally. When audio data based on the signal from microphone 12 is sent to the speech recognition engines for both languages, some kind of result will be returned from both speech recognition engines. For example, if the Japanese phrase "Say it again" is sent to both engines, the Japanese engine will return the text data "Say it again," while the English engine will return nonsensical text data that interprets "Say it again" as English. In such a case, since the Japanese request is a fixed phrase, it is best to perform a process to distinguish between English and Japanese by comparing the text data returned by the Japanese engine with the fixed phrase of the request. If they match, it is determined that a Japanese request was made; if they do not match, it is determined that English was spoken. While requests could be accepted in English, the inventors found it desirable to also allow requests in Japanese, especially considering cases where users may not understand English. The function that provides a Japanese explanation of English sentences in a conversation in response to requests such as "explain," could simply output a Japanese translation of the English sentence, but it would be particularly beneficial to output an explanation of the vocabulary and grammar used in the English sentence. In particular, it would be beneficial to output an explanation of the sentence structure used in the English sentence. For example, the explanation of the sentence structure could be drawn by using each phrase as a vertex (node) and drawing a diagram that encloses each phrase, and drawing branches (edges) such as line segments that show the relationships between related phrases. For example, it would be good to display it on the touch panel 7 as a graph structure (especially a tree structure). Furthermore, it would be good to output the displayed content as audio from the speaker device 13. For example, an AP for a syntactic analysis service like the one explained at https: / / gigazine.net / news / 20160602-foxtype-review / It would be best to configure it to call I, receive the result, and output the analysis result in Japanese. For example, the following processing and output will be performed. "Processing: Retrieve phrases from an English dialogue engine." I'm a fantastic robot. Person: "Say that again." I'm a fantastic robot. Person: "Explain it." Processing (recognizes "explain") → Calls the parsing API → Generates a Japanese explanation from the parsing result. In "robot," "I" is the subject, "am" is the verb, and "robot" is the object. "Fantastic" is an adjective meaning "amazing" and modifies "robot." The English sentence means "I am a wonderful robot." Person: "I don't think so." Robot: Don't say it! If you don't know how to say something in English, you can ask the robot to explain it in Japanese, which has the excellent effect of keeping the conversation going without interruption. If you can't make requests in Japanese, you have to look things up in a dictionary or use online translation when you don't know how to say something in English, which makes it feel like you're studying and can be stressful. If you can make requests in Japanese, you can learn without stress, feeling like you're just having a conversation with a bilingual person. Since consistency is very important in language learning, minimizing stress during learning is extremely important for continuing. This configuration makes it possible to realize a robot 1 that can continue language learning.

[0142] <Embodiment 11> In addition to the functions described in each embodiment, it is advisable to provide a function that speaks when it detects a person entering the room in which the robot 1 is installed. It is also advisable to provide a function that speaks when it detects a person leaving the room in which the robot 1 is installed. For example, sensors can be installed inside the room where robot 1 is installed and in the passageways outside the room to detect when a person enters the room where robot 1 is installed and when a person leaves the room where robot 1 is installed. In particular, if there is an automatic door for entering and exiting the room where robot 1 is installed, it is good to use sensors that specifically detect when a person approaches the automatic door in order to open or close it. Specifically, it is good to connect the controller MC of robot 1 to at least two of the following sensors: a first person detection sensor located outside the room on either side of the automatic door, a second person detection sensor located inside the room on either side of the automatic door, and a third person detection sensor that detects a person at the location of the automatic door when the automatic door is open (a sensor to prevent a person from getting caught in the door), in order to detect entry and exit into the room. In this way, entry and exit into the room where robot 1 is installed can be detected without installing any new sensors. For example, the edge of the signal that rises when each sensor detects a person can be captured for detection. For example, if a person is detected by the first person detection sensor and then by the third person detection sensor, it would be good to output a voice message from speaker device 13 with a phrase welcoming the person who has entered, such as "Welcome." For example, if a person is detected by the second person detection sensor and then by the third person detection sensor, it would be good to output a voice message from speaker device 13 with a phrase thanking the person who is leaving, such as "Thank you." At these times, the first to third motors 23 to 25 should be activated to cause the robot 1 to turn from its installation position towards the pre-set automatic door. If the first and second human detection sensors detect a person at the same time, the first human detection sensor should be given priority. This makes it easier for people entering to notice the presence of robot 1, and reduces the awkwardness of them suddenly being told "thank you" by robot 1.

[0143] <Other embodiments> (1) In each embodiment, a wireless LAN device 21 is provided, but a wired LAN device may be provided instead or together with it, and connected to a wired LAN network. The wired LAN device may be built into the robot 1 or attached externally. The wired LAN device may be connected to the USB OTG (On-The-Go) terminal 18. The wired LAN network should be configured to connect to the internet via a router or the like. In some environments, wireless LAN communication may be unstable or connection may not be possible. For example, if the robot 1 is installed in a place where there are many wireless LAN devices such as smartphones, it is desirable to configure it to access the internet via a wired LAN device. (2) In each embodiment, an example of a dialogue between a person and the robot 1 in a half-duplex manner is shown, but the dialogue between a person and the robot 1 may also be performed in a full-duplex manner. For example, in the dialogue in Table 1, after the person says "Open...", the recognition of the person's voice continues even while the person is saying "Are you sure?", and if the voice "Cancel" is recognized during the utterance of "Are you sure?", the robot may be configured to interrupt the utterance if it is still in progress and immediately say "Cancelled". (3) A configuration that performs dialogue using a half-duplex method is particularly good because it simplifies the configuration and processing and reduces costs. However, the inventors found a problem in that it is difficult to tell when the robot 1 turns on the microphone 12, and even when the user speaks, the beginning of the user's voice that the robot is trying to recognize, i.e., the beginning of the words, is often missing. To solve this problem, it is good to display a distinctive screen when the controller MC turns on the microphone. This will support smooth conversation. The distinctive screen display when the controller MC turns on the microphone can be done by either 1) lighting up the four corners of the screen, or 2) displaying a microphone icon, and doing both 1) and 2) is particularly effective. (4) The motors constituting the first to third motors 23 to 25 can be various motors such as DC motors, but it is particularly preferable to use stepping motors. The robot 1 is preferably configured to control its posture by a stepping motor. The torque of a stepping motor changes in proportion to the current flowing through the motor. If a large amount of current is passed, a large torque can be obtained, but problems such as heat generation and battery life will occur. Therefore, when the robot 1 is stationary, the minimum current necessary to maintain its posture is passed, and it is particularly preferable to pass a large current only when the robot 1 changes its posture. Note that it is also possible to energize all of the servo motors 23 to 25 with the minimum current necessary to maintain their posture when stationary. However, while the first motor 23 of the body part that can be supported by the detent torque (torque in the non-energized state) is not energized, for the second motor 24 and the third motor 25 of the head part, since they will lose against the detent torque, it is preferable to energize them even when stationary. Also, if it is rotated with the torque at rest, torque shortage will occur and detuning will occur and it will not rotate properly. Therefore, when rotating, it is better to increase the current to increase the torque compared to when at rest. On the other hand, since a large torque like when rotating is not necessary at rest, it is better to lower the current compared to when rotating. Also, when changing the direction to a certain direction, it is preferable to perform control such that it accelerates to a constant speed and then decelerates and stops when the target angle is approached. Also, when the robot is driven by a motor, mechanical and electrical noises are generated. Since this noise reduces the recognition rate of voice recognition, it is preferable to control to stop the motor during voice recognition. The scope of the present invention is not limited to the configurations explicitly described in the specification or limited ones. Combinations of various aspects of the present invention disclosed in this specification are also included in its scope. Among the present invention, the configuration for which a patent is sought is specified in the appended claims. However, even if it is a configuration not currently specified in the claims, it has the intention to make the configuration disclosed in this specification the claims in the future. The present invention of the application is not limited to the configurations described in the above-described embodiments. Each of the above-described embodiments and The components of the modification examples may be arbitrarily selected and combined. Also, any component of each embodiment or modification example may be arbitrarily combined with any component described in the means for solving the invention or a component that embodies any component described in the means for solving the invention. Regarding these, there is also an intention to obtain rights in the amendment or divisional application of this application. Also, even if there is a description such as "in the case of ~" or "when ~", it is not described as a configuration limited to that case or that time. The disclosure also includes configurations that are not these cases or times, and there is an intention to obtain rights. Also, the parts described in a sequential order are not limited to this order. The disclosure also includes configurations in which some parts are deleted or the order is reversed, and there is an intention to obtain rights. Also, by filing a change application for a design application, there is an intention to obtain rights for the overall design or partial design. Although the drawings depict the entire apparatus in solid lines, the drawings include not only the overall design but also partial designs claimed for a part of the apparatus. For example, not only can a part of the members of the apparatus be a partial design, but the drawings also include as a partial design a part of the apparatus regardless of the members. As a part of the apparatus, it may be a part of the members of the apparatus or a part of that member. Regarding the overall design, of course, there is an intention to claim as rights a partial design in which an arbitrary part of the solid-line part of the drawing is changed to a broken line.

Explanation of Reference Signs

[0144] 1... Robot as an apparatus, 41... Smartphone as another device, 51... Smart speaker as another device.

Claims

1. The display unit has a function to show a face image, A function that, when the display unit is touched while a face image is displayed in standby mode where interaction with the user is not possible, transitions from standby mode to a conversational mode where interaction with the user is possible. In the standby mode, if the display unit is touched and slid while a face image is displayed, the function will display a different standby image in place of the face image while remaining in standby mode. A device equipped with the following features.

2. The display unit has a function to show a face image, A function that, when the display unit is touched while a face image is displayed in standby mode where interaction with the user is not possible, transitions from standby mode to a conversational mode where interaction with the user is possible. In the standby mode, when the face image is displayed, touching any area of ​​the display unit will transition to the dialogue mode. In the standby mode, when the standby image is displayed, touching an area where a predetermined object is displayed will transition to the dialogue mode. A device equipped with the following features.

3. A program for a computer to implement the functions of the apparatus described in claim 1 or 2.