Behavior Control System

The behavior control system uses a sentence generation model and emotion engine to enhance robot responsiveness to user interactions by analyzing emotions and behaviors, improving interaction accuracy.

JP7733695B2Active Publication Date: 2025-09-03SOFTBANK GROUP CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023100374
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-06-19
Publication Date
2025-09-03
Estimated Expiration
2043-06-19

AI Technical Summary

Technical Problem

Conventional technologies struggle to enable robots to perform appropriate actions in response to user interactions effectively.

Method used

A behavior control system that utilizes a sentence generation model and emotion engine to analyze user behavior and emotions, allowing the robot to generate appropriate responses through a combination of linguistic and emotional understanding, including facial and voice recognition, and updates its behavior rules based on user reactions.

Benefits of technology

Enhances the robot's ability to respond appropriately to user actions by integrating emotional and linguistic analysis, improving user interaction and response accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007733695000003
    Figure 0007733695000003
  • Figure 0007733695000004
    Figure 0007733695000004
  • Figure 0007733695000005
    Figure 0007733695000005
Patent Text Reader

Abstract

To make an electronic device take an appropriate action in response to an action of a user.SOLUTION: An action control system disclosed herein comprises an input unit for receiving an image including a range corresponding to a field of view of a user, a processing unit for performing first specific processing using a sentence generation model for generating a sentence according to input content, and an output unit configured to control an action of an electronic device according to a processing result of the processing unit. The processing unit acquires information for assisting in visual recognition by the user as the processing result using an output of the sentence generation model that uses input of the input content including information on the image.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a behavior control system. [Background technology]

[0002] Patent Document 1 discloses a technology for determining an appropriate robot behavior for a user's state. The conventional technology in Patent Document 1 recognizes the user's reaction when the robot performs a specific action, and if the robot is unable to determine an action for the recognized user reaction, it updates the robot's behavior by receiving information about an action appropriate to the recognized user's state from a server. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent No. 6053847 Summary of the Invention [Problem to be solved by the invention]

[0004] However, the conventional technology has room for improvement in terms of making the robot perform appropriate actions in response to the user's actions. [Means for solving the problem]

[0005] According to a first aspect of the present invention, there is provided a behavior control system. The behavior control system includes an input unit that accepts an image including a range corresponding to a user's visual field, a processing unit that performs a first identification process using a sentence generation model that generates sentences according to the input content, and an output unit that controls the behavior of an electronic device according to a processing result by the processing unit. The processing unit obtains information to support the user's visual cognition as a processing result using an output of the sentence generation model when the input content including information about the image is input.

[0006] The electronic device may be a robot, where the term "robot" includes devices that perform physical actions, devices that output video and audio without performing physical actions, and agents that operate on software. [Brief explanation of the drawings]

[0007] [Figure 1] 1 shows an example of a system 5 according to a first embodiment. [Figure 2] 1 shows a schematic functional configuration of a robot 100 according to a first embodiment. [Figure 3] 10A and 10B schematically show an example of an operation flow of a collection process by the robot 100 according to the first embodiment. [Figure 4A] 10A and 10B schematically show an example of an operation flow of a response process by the robot 100 according to the first embodiment. [Figure 4B] 10A and 10B schematically illustrate an example of an operation flow of autonomous processing by the robot 100 according to the first embodiment. [Figure 5] 4 shows an emotion map 400 onto which multiple emotions are mapped. [Figure 6] 9 shows an emotion map 900 onto which multiple emotions are mapped. [Figure 7] 10A is an external view of a stuffed toy 100N according to a second embodiment, and FIG. 10B is a diagram showing the internal structure of the stuffed toy 100N. [Figure 8] FIG. 10 is a rear front view of a stuffed animal 100N according to a second embodiment. [Figure 9] 10A and 10B show a schematic functional configuration of a stuffed toy 100N according to a second embodiment. [Figure 10] 10 shows an outline of the functional configuration of an agent system 500 according to a third embodiment. [Figure 11] An example of the operation of the agent system will be shown. [Figure 12] An example of the operation of the agent system will be shown. [Figure 13A] 10 shows a schematic functional configuration of an agent system 700 used in smart glasses 720 according to a fourth embodiment. [Figure 13B]10 shows an outline of the functional configuration of a specific processing unit of an agent system 700 according to a fourth embodiment. [Figure 14A] 1 shows an example of how an agent system using smart glasses is used. [Figure 14B] 10 shows an example of an operational flow of a specific process by an agent system 700 according to a fourth embodiment. [Figure 14C] 10 shows an example of an operational flow of a specific process by an agent system 700 according to a fourth embodiment. [Figure 15] 1 shows an example of a hardware configuration of a computer 1200. [Figure 16] An overview of the specific processing is shown below. [Figure 17] An overview of the specific processing is shown below. [Figure 18] An overview of the specific processing is shown below. [Figure 19] 10 is a diagram illustrating another functional configuration of a specific processing unit of an agent system 700 according to the fourth embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0008] The present invention will be described below through embodiments of the invention, but the following embodiments do not limit the scope of the invention according to the claims. Furthermore, not all of the combinations of features described in the embodiments are necessarily essential to the solution of the invention.

[0009] [First embodiment] FIG. 1 schematically illustrates an example of a system 5 according to this embodiment. The system 5 includes a robot 100, a robot 101, a robot 102, and a server 300. Users 10a, 10b, 10c, and 10d are users of the robot 100. Users 11a, 11b, and 11c are users of the robot 101. Users 12a and 12b are users of the robot 102. In the description of this embodiment, users 10a, 10b, 10c, and 10d may be collectively referred to as users 10. Users 11a, 11b, and 11c may be collectively referred to as users 11. Users 12a and 12b may be collectively referred to as users 12. The robots 101 and 102 have substantially the same functions as the robot 100. Therefore, the system 5 will be described mainly focusing on the functions of the robot 100.

[0010] The robot 100 converses with the user 10 and provides the user 10 with video. At this time, the robot 100 cooperates with a server 300 or the like with which it can communicate via a communication network 20 to converse with the user 10 and provide the video, etc. to the user 10. For example, the robot 100 not only learns appropriate conversation by itself, but also cooperates with the server 300 to learn how to have a more appropriate conversation with the user 10. The robot 100 also records captured video data of the user 10 in the server 300, requests the video data, etc. from the server 300 as needed, and provides the video data, etc. to the user 10.

[0011] The robot 100 also has an emotional value that indicates the type of its own emotion. For example, the robot 100 has emotional values ​​that indicate the intensity of each of the following emotions: "joy," "anger," "sorrow," "pleasure," "discomfort," "relief," "anxiety," "sadness," "excitement," "worry," "relief," "fulfillment," "emptiness," and "neutral." For example, when the robot 100 is in a state where the emotional value of excitement is high, the robot 100 speaks at a fast speed. In this way, the robot 100 can express its own emotions through its actions.

[0012] Furthermore, the robot 100 may be configured to match a sentence generation model using AI (Artificial Intelligence) with an emotion engine, thereby determining the behavior of the robot 100 corresponding to the emotion of the user 10. Specifically, the robot 100 may be configured to recognize the behavior of the user 10, determine the emotion of the user 10 regarding the user's behavior, and determine the behavior of the robot 100 corresponding to the determined emotion.

[0013] More specifically, when the robot 100 recognizes the behavior of the user 10, it automatically generates the behavioral content that the robot 100 should take in response to the behavior of the user 10, using a preset sentence generation model. The sentence generation model may be interpreted as an algorithm and calculation for automatic dialogue processing using text. The sentence generation model is described, for example, in JP 2018-081444 A and ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ) and therefore a detailed description thereof will be omitted. Such a sentence generation model is configured using a large language model (LLM).

[0014] As described above, in this embodiment, by combining a large-scale language model and an emotion engine, it is possible to reflect the emotions of the user 10 and the robot 100, as well as various linguistic information, in the behavior of the robot 100. In other words, according to this embodiment, by combining a sentence generation model and an emotion engine, a synergistic effect can be obtained.

[0015] The robot 100 also has a function of recognizing the behavior of the user 10. The robot 100 recognizes the behavior of the user 10 by analyzing a facial image of the user 10 acquired by a camera function and a voice of the user 10 acquired by a microphone function. The robot 100 determines the behavior to be performed by the robot 100 based on the recognized behavior of the user 10, etc.

[0016] As an example of a behavioral decision model, the robot 100 stores rules that define the behaviors that the robot 100 will perform based on the emotions of the user 10, the emotions of the robot 100, and the behaviors of the user 10, and performs various behaviors in accordance with the rules.

[0017] Specifically, the robot 100 has reaction rules, as an example of a behavioral decision model, for determining the behavior of the robot 100 based on the emotions of the user 10, the emotions of the robot 100, and the behavior of the user 10. For example, the reaction rules define the behavior of the robot 100 as "laughing" when the behavior of the user 10 is "laughing." Furthermore, the reaction rules define the behavior of the robot 100 as "apologizing" when the behavior of the user 10 is "angry." Furthermore, the reaction rules define the behavior of the robot 100 as "answering" when the behavior of the user 10 is "asking a question." Furthermore, the reaction rules define the behavior of the robot 100 as "calling out" when the behavior of the user 10 is "sad."

[0018] When the robot 100 recognizes that the behavior of the user 10 is "angry" based on the reaction rules, the robot 100 selects the behavior of "apologizing" defined in the reaction rules as the behavior to be performed by the robot 100. For example, when the robot 100 selects the behavior of "apologizing," the robot 100 performs the motion of "apologizing" and outputs a voice representing the word "apologize."

[0019] In addition, it is defined that when the emotion of the robot 100 is "normal" (i.e., "joy" = 0, "anger" = 0, "sadness" = 0, "happiness" = 0) and the condition that the user 10 is in is "alone and looks lonely" is met, the emotion of the robot 100 changes to "worried" and the action of "calling out" can be performed.

[0020] When the robot 100 recognizes based on the reaction rule that the current emotion of the robot 100 is "normal" and that the user 10 appears lonely, the robot 100 increases the emotion value of "sad" of the robot 100. Furthermore, the robot 100 selects the action of "calling out" defined in the reaction rule as the action to be performed toward the user 10. For example, when the robot 100 selects the action of "calling out," the robot 100 converts the phrase "What's wrong?", which indicates concern, into a worried voice and outputs it.

[0021] The robot 100 also transmits to the server 300 user reaction information indicating that this behavior has elicited a positive reaction from the user 10. The user reaction information includes, for example, the user's behavior of "getting angry," the robot's 100 behavior of "apologizing," the fact that the user's 10 reaction was positive, and the attributes of the user 10.

[0022] The server 300 stores the user reaction information received from the robot 100. The server 300 receives and stores user reaction information not only from the robot 100 but also from each of the robots 101 and 102. The server 300 then analyzes the user reaction information from the robots 100, 101, and 102 and updates the reaction rules.

[0023] The robot 100 receives the updated reaction rules from the server 300 by inquiring about the updated reaction rules from the server 300. The robot 100 incorporates the updated reaction rules into the reaction rules stored in the robot 100. This allows the robot 100 to incorporate the reaction rules acquired by the robot 101, the robot 102, etc. into its own reaction rules.

[0024] 2 shows a schematic functional configuration of the robot 100. The robot 100 has a sensor unit 200, a sensor module unit 210, a storage unit 220, a control unit 228, and a control target 252. The control unit 228 has a state recognition unit 230, an emotion determination unit 232, a behavior recognition unit 234, a behavior determination unit 236, a memory control unit 238, a behavior control unit 250, a related information collection unit 270, and a communication processing unit 280.

[0025] The control object 252 includes a display device, a speaker, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 100 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 100 can be expressed by controlling these motors. In addition, the facial expressions of the robot 100 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 100. The posture, gestures, and facial expressions of the robot 100 are examples of the attitude of the robot 100.

[0026] The sensor unit 200 includes a microphone 201, a 3D depth sensor 202, a 2D camera 203, a distance sensor 204, a touch sensor 205, and an acceleration sensor 206. The microphone 201 continuously detects sound and outputs audio data. The microphone 201 may be provided on the head of the robot 100 and may have a binaural recording function. The 3D depth sensor 202 continuously emits an infrared pattern and detects the contour of an object by analyzing the infrared pattern from infrared images continuously captured by the infrared camera. The 2D camera 203 is an example of an image sensor. The 2D camera 203 captures images using visible light and generates visible light video information. The distance sensor 204 detects the distance to an object by emitting, for example, a laser or ultrasonic waves. The sensor unit 200 may also include a clock, a gyro sensor, a sensor for motor feedback, etc.

[0027] 2, the components of the robot 100 excluding the control target 252 and the sensor unit 200 are examples of components included in the behavior control system of the robot 100. The behavior control system of the robot 100 controls the control target 252.

[0028] The storage unit 220 includes a behavioral decision-making model 221, history data 222, collected data 223, and behavioral schedule data 224. The history data 222 includes the past emotional values ​​of the user 10, the past emotional values ​​of the robot 100, and behavioral history. Specifically, the history data 222 includes a plurality of event data including the emotional values ​​of the user 10, the emotional values ​​of the robot 100, and the behavior of the user 10. The data including the behavior of the user 10 includes camera images representing the behavior of the user 10. The emotional values ​​and behavioral history are recorded for each user 10, for example, by being associated with the identification information of the user 10. At least a portion of the storage unit 220 is implemented as a storage medium such as a memory. The storage unit 220 may also include a person DB that stores facial images of the user 10, attribute information of the user 10, and the like. Note that the functions of the components of the robot 100 shown in FIG. 2 , excluding the control target 252, the sensor unit 200, and the storage unit 220, can be realized by a CPU operating based on a program. For example, the functions of these components can be implemented as CPU operations using operating system (OS) and programs that run on the OS.

[0029] The sensor module unit 210 includes a voice emotion recognition unit 211, a speech understanding unit 212, a facial expression recognition unit 213, and a face recognition unit 214. Information detected by the sensor unit 200 is input to the sensor module unit 210. The sensor module unit 210 analyzes the information detected by the sensor unit 200 and outputs the analysis result to the state recognition unit 230.

[0030] The voice emotion recognition unit 211 of the sensor module unit 210 analyzes the voice of the user 10 detected by the microphone 201 and recognizes the emotion of the user 10. For example, the voice emotion recognition unit 211 extracts feature quantities such as frequency components of the voice and recognizes the emotion of the user 10 based on the extracted feature quantities. The speech understanding unit 212 analyzes the voice of the user 10 detected by the microphone 201 and outputs text information representing the content of the utterance of the user 10.

[0031] The facial expression recognition unit 213 recognizes the facial expression and emotion of the user 10 from the image of the user 10 captured by the 2D camera 203. For example, the facial expression recognition unit 213 recognizes the facial expression and emotion of the user 10 based on the shapes, positional relationships, etc. of the eyes and mouth.

[0032] The face recognition unit 214 recognizes the face of the user 10. The face recognition unit 214 recognizes the user 10 by matching a face image stored in a person DB (not shown) with a face image of the user 10 captured by the 2D camera 203.

[0033] The state recognition unit 230 recognizes the state of the user 10 based on the information analyzed by the sensor module unit 210. For example, it mainly performs processing related to perception using the analysis results of the sensor module unit 210. For example, it generates perceptual information such as "Dad is alone" or "There is a 90% chance that Dad is not smiling." It then performs processing to understand the meaning of the generated perceptual information. For example, it generates semantic information such as "Dad is alone and looks lonely."

[0034] The state recognition unit 230 recognizes the state of the robot 100 based on the information detected by the sensor unit 200. For example, the state recognition unit 230 recognizes the remaining battery level of the robot 100, the brightness of the environment surrounding the robot 100, etc. as the state of the robot 100.

[0035] The emotion determination unit 232 determines an emotion value indicating the emotion of the user 10 based on the information analyzed by the sensor module unit 210 and the state of the user 10 recognized by the state recognition unit 230. For example, the information analyzed by the sensor module unit 210 and the recognized state of the user 10 are input into a pre-trained neural network to obtain an emotion value indicating the emotion of the user 10.

[0036] Here, the emotion value indicating the emotion of user 10 is a value indicating the positive or negative emotion of the user. For example, if the user's emotion is a cheerful emotion accompanied by a sense of pleasure or comfort, such as "joy," "pleasure," "comfort," "relief," "excitement," "relief," and "fulfillment," the value is positive, and the cheerfulr the emotion, the larger the value. If the user's emotion is a negative emotion, such as "anger," "sorrow," "discomfort," "anxiety," "sorrow," "worry," and "emptiness," the value is negative, and the more unpleasant the emotion, the larger the absolute value of the negative value. If the user's emotion is none of the above ("neutral"), the value is 0.

[0037] In addition, the emotion determination unit 232 determines an emotion value indicating the emotion of the robot 100 based on the information analyzed by the sensor module unit 210, the information detected by the sensor unit 200, and the state of the user 10 recognized by the state recognition unit 230.

[0038] The emotion value of the robot 100 includes emotion values ​​for each of a plurality of emotion categories, and is a value (0 to 5) indicating the strength of each of "joy," "anger," "sorrow," and "happiness," for example.

[0039] Specifically, the emotion determination unit 232 determines an emotion value indicating the emotion of the robot 100 in accordance with a rule for updating the emotion value of the robot 100, which rule is determined in association with the information analyzed by the sensor module unit 210 and the state of the user 10 recognized by the state recognition unit 230.

[0040] For example, when the state recognition unit 230 recognizes that the user 10 looks lonely, the emotion determination unit 232 increases the emotion value of "sadness" of the robot 100. When the state recognition unit 230 recognizes that the user 10 is smiling, the emotion determination unit 232 increases the emotion value of "joy" of the robot 100.

[0041] The emotion determination unit 232 may determine the emotion value indicating the emotion of the robot 100 by further considering the state of the robot 100. For example, when the remaining battery power of the robot 100 is low or when the surrounding environment of the robot 100 is pitch black, the emotion value of "sadness" of the robot 100 may be increased. Furthermore, when the user 10 continues to talk to the robot 100 despite the remaining battery power being low, the emotion value of "anger" may be increased.

[0042] The behavior recognition unit 234 recognizes the behavior of the user 10 based on the information analyzed by the sensor module unit 210 and the state of the user 10 recognized by the state recognition unit 230. For example, the information analyzed by the sensor module unit 210 and the recognized state of the user 10 are input into a pre-trained neural network, the probability of each of a plurality of predetermined behavior classifications (for example, "laughing," "angry," "asking a question," and "sad") is obtained, and the behavior classification with the highest probability is recognized as the behavior of the user 10.

[0043] As described above, in this embodiment, the robot 100 identifies the user 10 and then acquires the content of the user's utterance. When acquiring and using the content of the utterance, the robot 100 obtains the necessary consent in accordance with laws and regulations from the user 10, and the behavior control system of the robot 100 according to this embodiment takes into consideration the protection of the personal information and privacy of the user 10.

[0044] Next, the processing of the behavior decision unit 236 when the robot 100 performs response processing to respond to the behavior of the user 10 will be described.

[0045] The behavior determination unit 236 determines a behavior corresponding to the behavior of the user 10 recognized by the behavior recognition unit 234 based on the current emotion value of the user 10 determined by the emotion determination unit 232, history data 222 of past emotion values ​​determined by the emotion determination unit 232 before the current emotion value of the user 10 was determined, and the emotion value of the robot 100. In this embodiment, a case will be described in which the behavior determination unit 236 uses one most recent emotion value included in the history data 222 as the past emotion value of the user 10, but the disclosed technology is not limited to this aspect. For example, the behavior determination unit 236 may use multiple most recent emotion values ​​as the past emotion value of the user 10, or may use an emotion value from a unit period ago, such as one day ago. Furthermore, the behavior determination unit 236 may determine a behavior corresponding to the behavior of the user 10 by further considering not only the current emotion value of the robot 100 but also the history of the past emotion values ​​of the robot 100. The behavior determined by the behavior determining unit 236 includes gestures made by the robot 100 or speech content of the robot 100 .

[0046] The behavior determination unit 236 according to this embodiment determines the behavior of the robot 100 as a behavior corresponding to the behavior of the user 10, based on a combination of the past and current emotional values ​​of the user 10, the emotional value of the robot 100, the behavior of the user 10, and the behavior determination model 221. For example, if the past emotional value of the user 10 is a positive value and the current emotional value is a negative value, the behavior determination unit 236 determines, as a behavior corresponding to the behavior of the user 10, a behavior that will change the emotional value of the user 10 to a positive value.

[0047] The reaction rule as the behavioral decision-making model 221 defines the behavior of the robot 100 according to a combination of the past emotional value and the current emotional value of the user 10, the emotional value of the robot 100, and the behavior of the user 10. For example, if the past emotional value of the user 10 is a positive value and the current emotional value is a negative value, and the behavior of the user 10 is sad, a combination of gestures and speech content when asking a question to encourage the user 10 with gestures is defined as the behavior of the robot 100.

[0048] For example, the reaction rule as the behavior determination model 221 defines behaviors of the robot 100 for all combinations of patterns of the robot 100's emotional values ​​(1296 patterns, which are the fourth power of six values ​​"joy," "anger," "sadness," and "happiness" from "0" to "5"), patterns of combinations of the user 10's past emotional values ​​and current emotional values, and behavioral patterns of the user 10. That is, for each pattern of the robot 100's emotional values, behaviors of the robot 100 are defined according to the behavioral patterns of the user 10 for each of a plurality of combinations of the user 10's past emotional values ​​and current emotional values, such as negative and negative values, negative and positive values, positive and negative values, positive and positive values, negative and normal values, and normal and normal values. Note that the behavior determination unit 236 may transition to an operation mode that determines the behavior of the robot 100 using the history data 222 when the user 10 makes an utterance intending to continue a conversation from a past topic, such as "I want to talk about that topic we talked about last time."

[0049] The reaction rule as the behavior determination model 221 may prescribe at least one of a gesture and a utterance as the behavior of the robot 100, for each of the patterns (1296 patterns) of the emotional value of the robot 100. Alternatively, the reaction rule as the behavior determination model 221 may prescribe at least one of a gesture and a utterance as the behavior of the robot 100, for each group of patterns of the emotional value of the robot 100.

[0050] The strength of each gesture included in the behavior of the robot 100 defined in the reaction rules as the behavior determination model 221 is predetermined. The strength of each utterance included in the behavior of the robot 100 defined in the reaction rules as the behavior determination model 221 is predetermined.

[0051] The memory control unit 238 determines whether or not to store data including the behavior of the user 10 in the history data 222 based on the predetermined behavior intensity for the behavior determined by the behavior determination unit 236 and the emotion value of the robot 100 determined by the emotion determination unit 232.

[0052] Specifically, if the total intensity value, which is the sum of the sum of the emotion values ​​for each of the multiple emotion classifications of the robot 100, the predetermined intensity for the gesture included in the behavior determined by the behavior determination unit 236, and the predetermined intensity for the speech content included in the behavior determined by the behavior determination unit 236, is equal to or greater than a threshold value, it is determined that data including the behavior of the user 10 is to be stored in the history data 222.

[0053] When the memory control unit 238 decides to store data including the behavior of the user 10 in the history data 222, it stores in the history data 222 the behavior determined by the behavior determination unit 236, information analyzed by the sensor module unit 210 from the present time up to a certain period of time ago (for example, all surrounding information such as data on the sound, images, smells, etc. of the scene), and the state of the user 10 recognized by the state recognition unit 230 (for example, the facial expression, emotions, etc. of the user 10).

[0054] The behavior control unit 250 controls the control target 252 based on the behavior determined by the behavior determination unit 236. For example, when the behavior determination unit 236 determines an behavior that includes speaking, the behavior control unit 250 outputs a sound from a speaker included in the control target 252. At this time, the behavior control unit 250 may determine the speaking rate of the sound based on the emotional value of the robot 100. For example, the behavior control unit 250 determines a faster speaking rate as the emotional value of the robot 100 increases. In this way, the behavior control unit 250 determines the execution form of the behavior determined by the behavior determination unit 236 based on the emotional value determined by the emotional determination unit 232.

[0055] The behavior control unit 250 may recognize a change in the user 10's emotion in response to the execution of the behavior determined by the behavior determination unit 236. For example, the change in emotion may be recognized based on the voice or facial expression of the user 10. Alternatively, the change in emotion of the user 10 may be recognized based on the detection of an impact by the touch sensor 205 included in the sensor unit 200. If an impact is detected by the touch sensor 205 included in the sensor unit 200, the behavior control unit 250 may recognize that the user 10's emotion has worsened. Alternatively, if the detection result of the touch sensor 205 included in the sensor unit 200 indicates that the user 10 is smiling, happy, or the like, the behavior control unit 250 may recognize that the user 10's emotion has improved. Information indicating the user 10's reaction is output to the communication processing unit 280.

[0056] Furthermore, after the behavior control unit 250 executes the behavior determined by the behavior determination unit 236 in the execution mode determined according to the emotion of the robot 100, the emotion determination unit 232 further changes the emotion value of the robot 100 based on the user's reaction to the execution of the behavior. Specifically, the emotion determination unit 232 increases the emotion value of "joy" of the robot 100 when the user's reaction to the behavior determined by the behavior determination unit 236 being performed on the user in the execution mode determined by the behavior control unit 250 is not negative. Furthermore, the emotion determination unit 232 increases the emotion value of "sad" of the robot 100 when the user's reaction to the behavior determined by the behavior determination unit 236 being performed on the user in the execution mode determined by the behavior control unit 250 is negative.

[0057] Furthermore, the behavior control unit 250 expresses the emotion of the robot 100 based on the determined emotion value of the robot 100. For example, when the emotion value of "happiness" of the robot 100 is increased, the behavior control unit 250 controls the control object 252 to make the robot 100 perform a happy gesture. When the emotion value of "sadness" of the robot 100 is increased, the behavior control unit 250 controls the control object 252 to make the robot 100 assume a droopy posture.

[0058] The communication processing unit 280 is responsible for communication with the server 300. As described above, the communication processing unit 280 transmits user reaction information to the server 300. The communication processing unit 280 also receives updated reaction rules from the server 300. When the communication processing unit 280 receives the updated reaction rules from the server 300, it updates the reaction rules as the behavioral determination model 221.

[0059] The server 300 communicates between the robot 100, the robot 101, and the robot 102 and the server 300, receives user reaction information sent from the robot 100, and updates the reaction rules based on reaction rules that include actions that have received positive reactions.

[0060] The related information collecting unit 270 collects information related to the preference information acquired about the user 10 from external data (websites such as news sites and video sites) at a predetermined timing, based on the preference information acquired about the user 10.

[0061] Specifically, the related information collecting unit 270 acquires preference information indicating matters of interest to the user 10 from the content of speech of the user 10 or a setting operation by the user 10. The related information collecting unit 270 periodically collects news related to the preference information through ChatGPT Plugins (Internet search<URL: https: / / openai.com / blog / chatgpt-plugins> For example, if the user 10 is a fan of a specific professional baseball team, the related information collection unit 270 collects news related to the game results of the specific professional baseball team from external data at a predetermined time every day using ChatGPT Plugins.

[0062] The emotion determining unit 232 determines the emotion of the robot 100 based on information related to the preference information collected by the related information collecting unit 270.

[0063] Specifically, the emotion determining unit 232 inputs text representing information related to the preference information collected by the related information collecting unit 270 into a pre-trained neural network for determining emotions, obtains emotion values ​​indicating each emotion, and determines the emotion of the robot 100. For example, if the collected news related to the game results of a specific professional baseball team indicates that the specific professional baseball team won, the emotion determining unit 232 determines that the emotion value of "joy" of the robot 100 is large.

[0064] When the emotion value of the robot 100 is equal to or greater than the threshold value, the storage control unit 238 stores information related to the preference information collected by the related information collection unit 270 in the collected data 223.

[0065] Next, the processing of the behavior decision unit 236 when the robot 100 performs autonomous processing to act autonomously will be described.

[0066] The behavior decision unit 236 determines, at a predetermined timing, one of a plurality of types of robot behaviors, including no behavior, as the behavior of the robot 100, using at least one of the state of the user 10, the emotion of the user 10, the emotion of the robot 100, and the state of the robot 100, and the behavior decision model 221. Here, an example will be described in which a sentence generation model with a dialogue function is used as the behavior decision model 221.

[0067] Specifically, the behavior determination unit 236 inputs text representing at least one of the state of the user 10, the emotion of the user 10, the emotion of the robot 100, and the state of the robot 100, and text asking about the robot's behavior, into a sentence generation model, and determines the behavior of the robot 100 based on the output of the sentence generation model.

[0068] For example, the multiple types of robot behaviors include the following (1) to (10).

[0069] (1) The robot does nothing. (2) Robots dream. (3) The robot speaks to the user. (4) The robot creates a picture diary. (5) The robot suggests an activity. (6) The robot suggests people for the user to meet. (7) The robot introduces news that the user finds interesting. (8) The robot edits photos and videos. (9) The robot learns together with the user. (10) Robots evoke memories.

[0070] The behavior determination unit 236 inputs the state of the user 10 and the state of the robot 100 recognized by the state recognition unit 230, text indicating the current emotional value of the user 10 determined by the emotion determination unit 232, and the current emotional value of the robot 100, and text asking which of multiple types of robot behaviors, including not taking any action, into the sentence generation model at regular time intervals, and determines the behavior of the robot 100 based on the output of the sentence generation model. Here, if there is no user 10 around the robot 100, the text input to the sentence generation model does not need to include the state of the user 10 and the current emotional value of the user 10, or may include an indication that the user 10 is not present.

[0071] For example, "The robot is in a very happy state. The user is in a normal happy state. The user is sleeping. Which of the following (1) to (10) is the best behavior for the robot?" (1) The robot does nothing. (2) Robots dream. (3) The robot speaks to the user. The text "..." is input to the sentence generation model. Based on the output of the sentence generation model, "It can be said that either (1) doing nothing or (2) the robot dreams is the most appropriate behavior," the behavior of robot 100 is determined to be "(1) doing nothing" or "(2) the robot dreams."

[0072] Another example is, "The robot is feeling a little lonely. The user is not present. The robot's surroundings are dark. Which of the following (1) to (10) is the best behavior for the robot?" (1) The robot does nothing. (2) Robots dream. (3) The robot speaks to the user. The text "..." is input to the sentence generation model. Based on the output of the sentence generation model, "Either (2) the robot dreams or (4) the robot creates a picture diary is said to be the most appropriate behavior," the behavior of robot 100 is determined to be "(2) the robot dreams" or "(4) the robot creates a picture diary."

[0073] When the behavior decision unit 236 decides to create an original event, that is, "(2) The robot dreams," as the robot behavior, it uses the sentence generation model to create an original event that combines multiple event data from the history data 222. At this time, the storage control unit 238 stores the created original event in the history data 222.

[0074] When the behavior determining unit 236 determines that the robot 100 will speak, i.e., "(3) The robot speaks to the user," as the robot behavior, the behavior determining unit 236 uses the sentence generation model to determine the content of the robot's utterance corresponding to the user's state and the user's emotion or the robot's emotion. At this time, the behavior control unit 250 causes a sound representing the determined content of the robot's utterance to be output from a speaker included in the control target 252. Note that, when the user 10 is not present around the robot 100, the behavior control unit 250 stores the determined content of the robot's utterance in the behavior schedule data 224 without outputting the sound representing the determined content of the robot's utterance.

[0075] When the behavior decision unit 236 decides that "(7) The robot introduces news that the user is interested in" as the robot behavior, it uses the sentence generation model to decide the content of the robot's utterance corresponding to the information stored in the collected data 223. At this time, the behavior control unit 250 causes a sound representing the determined content of the robot's utterance to be output from a speaker included in the control target 252. Note that when the user 10 is not present around the robot 100, the behavior control unit 250 stores the determined content of the robot's utterance in the behavior schedule data 224 without outputting the sound representing the determined content of the robot's utterance.

[0076] When the behavior decision unit 236 determines that the robot 100 will create an event image, i.e., "(4) The robot creates a picture diary," as the robot behavior, the behavior decision unit 236 generates an image representing the event data using the image generation model for event data selected from the history data 222, and generates an explanatory text representing the event data using the sentence generation model, and outputs the combination of the image representing the event data and the explanatory text representing the event data as an event image. Note that when the user 10 is not present around the robot 100, the behavior control unit 250 does not output the event image, but stores the event image in the behavior schedule data 224.

[0077] When the behavior decision unit 236 determines that the robot behavior is "(8) The robot edits photos and videos," that is, when it determines that an image is to be edited, it selects event data from the history data 222 based on the emotion value, edits the image data of the selected event data, and outputs it. Note that when the user 10 is not present around the robot 100, the behavior control unit 250 stores the edited image data in the behavior schedule data 224 without outputting the edited image data.

[0078] When the behavior decision unit 236 determines that the robot behavior is "(5) The robot proposes an activity," that is, that the robot proposes an action for the user 10, it determines the user's action to be proposed using a sentence generation model based on the event data stored in the history data 222. At this time, the behavior control unit 250 causes a sound proposing the user's action to be output from a speaker included in the control target 252. Note that when the user 10 is not present around the robot 100, the behavior control unit 250 does not output a sound proposing the user's action, but instead stores the suggestion of the user's action in the behavior schedule data 224.

[0079] When the behavior decision unit 236 determines that the robot behavior is "(6) The robot proposes people that the user should meet," that is, to propose people that the user 10 should have contact with, the behavior decision unit 236 uses a sentence generation model based on the event data stored in the history data 222 to determine people that the proposed user should have contact with. At this time, the behavior control unit 250 causes a speaker included in the control target 252 to output a sound indicating that the robot will propose people that the user should have contact with. Note that, when the user 10 is not present around the robot 100, the behavior control unit 250 does not output a sound indicating that the robot will propose people that the user should have contact with, but instead stores the suggestion of people that the user should have contact with in the behavior schedule data 224.

[0080] When the behavior determining unit 236 determines that the robot 100 will make an utterance related to studying, such as "(9) The robot studies together with the user," as the robot behavior, the behavior determining unit 236 uses the sentence generation model to determine the content of the robot's utterance to encourage studying, give study questions, or provide study-related advice, corresponding to the user's state and the user's or the robot's emotions. At this time, the behavior control unit 250 outputs a sound representing the determined content of the robot's utterance from a speaker included in the control target 252. Note that when the user 10 is not present around the robot 100, the behavior control unit 250 stores the determined content of the robot's utterance in the behavior schedule data 224 without outputting a sound representing the determined content of the robot's utterance.

[0081] When the behavior determining unit 236 determines that the robot behavior is "(10) The robot recalls a memory," that is, that the robot recalls event data, it selects the event data from the history data 222. At this time, the emotion determining unit 232 determines the emotion of the robot 100 based on the selected event data. Furthermore, the behavior determining unit 236 uses a sentence generation model based on the selected event data to create an emotion change event that represents the content of the utterances and actions of the robot 100 to change the user's emotion value. At this time, the memory control unit 238 stores the emotion change event in the behavior schedule data 224.

[0082] For example, the fact that the video the user was watching was about pandas is stored as event data in the history data 222, and when that event data is selected, the sentence generation model is input with the following: "On the topic of pandas, what are some things you should say to the user the next time you meet them? Name three." If the output of the sentence generation model is "(1) Let's go to the zoo, (2) Let's draw a picture of a panda, (3) Let's go buy a stuffed panda," the robot 100 inputs into the sentence generation model, "Which of (1), (2), and (3) would the user be most happy with?" If the output of the sentence generation model is "(1) Let's go to the zoo," the robot 100 will utter "(1) Let's go to the zoo" the next time it meets the user, and this is created as an emotion change event and stored in the behavior schedule data 224.

[0083] Also, for example, event data with a large emotional value of the robot 100 is selected as an impressive memory of the robot 100. In this way, an emotional change event can be created based on the event data selected as an impressive memory.

[0084] When the behavior decision unit 236 detects an action of the user 10 toward the robot 100 from a state in which the user 10 has not taken any action toward the robot 100 based on the state of the user 10 recognized by the state recognition unit 230, the behavior decision unit 236 reads the data stored in the action schedule data 224 and decides the behavior of the robot 100.

[0085] For example, if the user 10 is not present around the robot 100 and the robot 100 detects the user 10, the behavior decision unit 236 reads out the data stored in the behavior schedule data 224 and decides the behavior of the robot 100. Also, if the user 10 is asleep and the robot 100 detects that the user 10 has woken up, the behavior decision unit 236 reads out the data stored in the behavior schedule data 224 and decides the behavior of the robot 100.

[0086] FIG. 3 shows an example of an operational flow for a collection process for collecting information related to the preference information of the user 10. The operational flow shown in FIG. 3 is executed repeatedly at regular intervals. It is assumed that preference information indicating matters of interest to the user 10 has been acquired from the content of speech of the user 10 or from a setting operation by the user 10. Note that "S" in the operational flow indicates the step that is executed.

[0087] First, in step S90, the related information collection unit 270 acquires preference information that indicates the matters in which the user 10 is interested.

[0088] In step S92, the related information collection unit 270 collects information related to the preference information from external data.

[0089] In step S94 , the emotion determining unit 232 determines the emotion value of the robot 100 based on the information related to the preference information collected by the related information collecting unit 270 .

[0090] In step S96, the storage control unit 238 determines whether the emotion value of the robot 100 determined in step S94 is equal to or greater than a threshold. If the emotion value of the robot 100 is less than the threshold, the processing ends without storing the information related to the collected preference information in the collection data 223. On the other hand, if the emotion value of the robot 100 is equal to or greater than the threshold, the processing proceeds to step S98.

[0091] In step S98, the storage control unit 238 stores the collected information related to the preference information in the collected data 223, and ends the process.

[0092] 4A shows an example of an operation flow for determining an action of the robot 100 when the robot 100 performs a response process to respond to an action of the user 10. The operation flow shown in FIG. 4A is repeatedly executed. At this time, it is assumed that information analyzed by the sensor module unit 210 is input.

[0093] First, in step S100 , the state recognition unit 230 recognizes the state of the user 10 and the state of the robot 100 based on the information analyzed by the sensor module unit 210 .

[0094] In step S102, the emotion determination unit 232 determines an emotion value indicating the emotion of the user 10 based on the information analyzed by the sensor module unit 210 and the state of the user 10 recognized by the state recognition unit 230.

[0095] In step S103, the emotion determination unit 232 determines an emotion value indicating the emotion of the robot 100 based on the information analyzed by the sensor module unit 210 and the state of the user 10 recognized by the state recognition unit 230. The emotion determination unit 232 adds the determined emotion values ​​of the user 10 and the robot 100 to the history data 222.

[0096] In step S104 , the behavior recognition unit 234 recognizes the behavior classification of the user 10 based on the information analyzed by the sensor module unit 210 and the state of the user 10 recognized by the state recognition unit 230 .

[0097] In step S106, the behavior determination unit 236 determines the behavior of the robot 100 based on a combination of the current emotional value of the user 10 determined in step S102 and the past emotional values ​​included in the history data 222, the emotional value of the robot 100, the behavior of the user 10 recognized in the above step S104, and the behavior determination model 221.

[0098] In step S108, the behavior control unit 250 controls the control target 252 based on the behavior determined by the behavior determination unit 236.

[0099] In step S110, the memory control unit 238 calculates a total intensity value based on the predetermined behavior intensities for the behavior determined by the behavior determination unit 236 and the emotion value of the robot 100 determined by the emotion determination unit 232.

[0100] In step S112, the storage control unit 238 determines whether the total intensity value is equal to or greater than a threshold. If the total intensity value is less than the threshold, the process ends without storing the event data including the behavior of the user 10 in the history data 222. On the other hand, if the total intensity value is equal to or greater than the threshold, the process proceeds to step S114.

[0101] In step S114, event data including the behavior determined by the behavior determination unit 236, information analyzed by the sensor module unit 210 from the present time up to a certain period of time ago, and the state of the user 10 recognized by the state recognition unit 230 is stored in the history data 222.

[0102] 4B shows an example of an operation flow for determining the behavior of the robot 100 when the robot 100 performs autonomous processing to act autonomously. The operation flow shown in FIG. 4B is automatically and repeatedly executed, for example, at regular intervals. At this time, it is assumed that information analyzed by the sensor module unit 210 has been input. Note that the same step numbers are used for processes similar to those in FIG. 4A.

[0103] First, in step S100 , the state recognition unit 230 recognizes the state of the user 10 and the state of the robot 100 based on the information analyzed by the sensor module unit 210 .

[0104] In step S102, the emotion determination unit 232 determines an emotion value indicating the emotion of the user 10 based on the information analyzed by the sensor module unit 210 and the state of the user 10 recognized by the state recognition unit 230.

[0105] In step S103, the emotion determination unit 232 determines an emotion value indicating the emotion of the robot 100 based on the information analyzed by the sensor module unit 210 and the state of the user 10 recognized by the state recognition unit 230. The emotion determination unit 232 adds the determined emotion values ​​of the user 10 and the robot 100 to the history data 222.

[0106] In step S104 , the behavior recognition unit 234 recognizes the behavior classification of the user 10 based on the information analyzed by the sensor module unit 210 and the state of the user 10 recognized by the state recognition unit 230 .

[0107] In step S200, the behavior decision unit 236 decides on one of a plurality of types of robot behaviors, including no action, as the behavior of the robot 100 based on the state of the user 10 recognized in step S100, the emotions of the user 10 determined in step S102, the emotions of the robot 100, the state of the robot 100 recognized in step S100, the behavior of the user 10 recognized in step S104, and the behavior decision model 221.

[0108] In step S201, the behavior decision unit 236 determines whether or not it has been determined in step S200 that no behavior will be performed. If it has been determined that no behavior will be performed by the robot 100, the process ends. On the other hand, if it has not been determined that no behavior will be performed by the robot 100, the process proceeds to step S202.

[0109] In step S202, the behavior determination unit 236 performs processing according to the type of robot behavior determined in step S200. At this time, the behavior control unit 250, the emotion determination unit 232, or the memory control unit 238 executes processing according to the type of robot behavior.

[0110] In step S110, the memory control unit 238 calculates a total intensity value based on the predetermined behavior intensities for the behavior determined by the behavior determination unit 236 and the emotion value of the robot 100 determined by the emotion determination unit 232.

[0111] In step S112, the storage control unit 238 determines whether the total intensity value is equal to or greater than a threshold. If the total intensity value is less than the threshold, the process ends without storing data including the behavior of the user 10 in the history data 222. On the other hand, if the total intensity value is equal to or greater than the threshold, the process proceeds to step S114.

[0112] In step S114, the memory control unit 238 stores the behavior determined by the behavior determination unit 236, the information analyzed by the sensor module unit 210 from the present time up to a certain period of time ago, and the state of the user 10 recognized by the state recognition unit 230 in the history data 222.

[0113] As described above, the robot 100 determines an emotion value indicating the emotion of the robot 100 based on the state of the user, and determines whether or not to store data including the behavior of the user 10 in the history data 222 based on the emotion value of the robot 100. This makes it possible to reduce the capacity of the history data 222 that stores data including the behavior of the user 10. For example, when the robot 100 determines that the state of the user 10 years from now will be the same as that of 10 years ago, the robot 100 can present to the user 10 all kinds of peripheral information, such as the state of the user 10 from 10 years ago (for example, the facial expression, emotions, etc. of the user 10), as well as data on the sounds, images, smells, etc. of the situation.

[0114] Furthermore, the robot 100 can be made to perform an appropriate action in response to the action of the user 10. Conventionally, the user's actions are classified and an action, including the robot's facial expression and appearance, is determined. In contrast, the robot 100 determines the current emotional value of the user 10 and performs an action on the user 10 based on the past emotional value and the current emotional value. Therefore, for example, if the user 10 was cheerful yesterday but is depressed today, the robot 100 can utter an utterance such as, "You were cheerful yesterday, but what's wrong with you today?" The robot 100 can also utter an utterance using gestures. For example, if the user 10 was depressed yesterday but is cheerful today, the robot 100 can utter an utterance such as, "You were depressed yesterday, but you seem cheerful today, don't you?" For example, if the user 10 who was cheerful yesterday is more cheerful today than yesterday, the robot 100 can utter an utterance such as, "You're more cheerful today than yesterday. Has anything better happened than yesterday?" Furthermore, for example, the robot 100 can say to the user 10 whose emotional value is equal to or greater than 0 and whose emotional value fluctuation range continues to be within a certain range, "Your mood has been stable recently, which is good."

[0115] Furthermore, for example, the robot 100 may ask the user 10, "Did you finish the homework you told me about yesterday?", and if the user 10 replies, "Yes, I did," the robot 100 may utter a positive utterance such as "Great!" and perform a positive gesture such as clapping or a thumbs-up. Furthermore, for example, if the user 10 utters, "The presentation you gave the day before yesterday went well," the robot 100 may utter a positive utterance such as "Good job!" and perform the above-mentioned positive gesture. In this way, the robot 100 may be expected to make the user 10 feel a sense of affinity with the robot 100 by performing an action based on the state history of the user 10.

[0116] Also, for example, when user 10 is watching a video about a panda, if the emotion value of user 10's emotion of "ease" is equal to or greater than a threshold, the scene in which the panda appears in the video may be stored as event data in history data 222.

[0117] Using the data accumulated in the history data 222 and the collected data 223, the robot 100 can constantly learn what kind of conversation with the user will maximize the emotional value that expresses the user's happiness.

[0118] Furthermore, when the robot 100 is not having a conversation with the user 10, the robot 100 can autonomously start to act based on its own emotions.

[0119] Furthermore, in the autonomous processing, the robot 100 automatically generates questions, inputs them into a sentence generation model, and repeatedly obtains the output of the sentence generation model as an answer to the question, thereby creating emotion change events for increasing positive emotions and storing the events in the behavior schedule data 224. In this way, the robot 100 can perform self-learning.

[0120] Furthermore, when the robot 100 automatically generates a question without receiving an external trigger, the question can be automatically generated based on memorable event data identified from the robot's past emotional value history.

[0121] Furthermore, the related information collecting unit 270 can perform self-learning by repeating the search execution step of automatically performing a keyword search in response to preference information about the user and obtaining search results.

[0122] Here, the search execution stage may be configured to automatically execute a keyword search based on memorable event data identified from the robot's past emotional value history in a state where no external trigger has been received.

[0123] The emotion determining unit 232 may determine the user's emotion in accordance with a specific mapping. Specifically, the emotion determining unit 232 may determine the user's emotion in accordance with an emotion map (see FIG. 5), which is a specific mapping.

[0124] FIG. 5 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both emotions and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[0125] (1) For example, if the emotion engine, which is the emotion determination unit 232 of the robot 100, detects emotions in about 100 msec, the frequency of determining the reaction action of the robot 100 (for example, a backchannel) may be set at a timing at least as frequent as the emotion engine's detection frequency (100 msec), or may be set at a timing earlier than this. The emotion engine's detection frequency may be interpreted as a sampling rate.

[0126] By detecting emotions in about 100 msec and immediately performing a corresponding reaction (e.g., a backchannel), unnatural backchannels are eliminated, enabling a natural, well-read dialogue. The robot 100 performs a reaction (e.g., a backchannel) according to the directionality and degree (strength) of the mandala in the emotion map 400. Note that the detection frequency (sampling rate) of the emotion engine is not limited to 100 ms and may be changed depending on the situation (e.g., playing sports), the user's age, etc.

[0127] (2) The directionality and intensity of emotions may be set in advance in reference to the emotion map 400, and the movement of the back-channel and the strength of the back-channel may be set. For example, if the robot 100 feels a sense of stability or security, the robot 100 may nod and continue listening. If the robot 100 feels anxious, confused, or suspicious, the robot 100 may tilt its head or stop shaking its head.

[0128] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[0129] (3) If the robot 100 receives a compliment and feels good, the filler "ah" may precede the line, and if the robot 100 receives harsh words and feels pain, the filler "ugh!" may precede the line. Also, physical reactions such as the robot 100 looking up at the sky while saying "ah" because it feels so good, or crouching down while saying "ugh!" may be included. These emotions are distributed around 9 o'clock on the emotion map 400.

[0130] (4) In the left half of the emotion map 400, internal sensations (reactions) are more important than situational awareness. This can give the impression of an unconscious reaction.

[0131] When the robot 100 feels a positive feeling in its situational awareness while experiencing an internal sensation (reaction) of understanding, the robot 100 may nod deeply while looking at the other person, or may say "uh-huh." In this way, the robot 100 may generate a behavior that shows a balanced positive feeling toward the other person, that is, tolerance and tolerance toward the other person. Such emotions are distributed around 12 o'clock on the emotion map 400.

[0132] Conversely, even when the robot 100 is aware of an internal sensation (reaction) of discomfort, it may shake its head when it feels disgust, or turn its eye LEDs red and glare at the other person when it feels hatred. Furthermore, as the balanced disgust towards the other person becomes stronger, it may initiate behaviors such as attacking or eradicating the other person. These emotions are distributed around the 6 o'clock position on the emotion map 400.

[0133] (5) The inside of emotion map 400 represents the mind, and the outside of emotion map 400 represents behavior, so the further outside emotion map 400 you go, the more visible the emotion becomes (is expressed in behavior).

[0134] (6) When listening to someone while feeling a sense of security, which is distributed around 3 o'clock on the emotion map 400, the robot 100 may nod its head lightly and say "hmm." However, when listening to someone while feeling a sense of love, which is distributed around 12 o'clock, the robot 100 may nod its head strongly, as if deeply nodding its head.

[0135] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[0136] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[0137] The emotion determination unit 232 inputs the information analyzed by the sensor module unit 210 and the recognized state of the user 10 into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the emotion of the user 10. This neural network is pre-trained based on multiple pieces of training data that are combinations of the information analyzed by the sensor module unit 210, the recognized state of the user 10, and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 6. FIG. 6 shows an example in which multiple emotions, such as "relieved," "calm," and "reassuring," have similar emotion values.

[0138] Furthermore, the emotion determination unit 232 may determine the emotion of the robot 100 according to a specific mapping. Specifically, the emotion determination unit 232 inputs the information analyzed by the sensor module unit 210, the state of the user 10 recognized by the state recognition unit 230, and the state of the robot 100 into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the emotion of the robot 100. This neural network is pre-trained based on a plurality of training data that are combinations of the information analyzed by the sensor module unit 210, the recognized state of the user 10, and the state of the robot 100, with emotion values ​​indicating each emotion shown in the emotion map 400. For example, the neural network is trained based on training data indicating that when it is recognized from the output of a touch sensor (not shown) that the robot 100 is being stroked by the user 10, the emotion value of "happy" is "3," and based on training data indicating that when it is recognized from the output of the acceleration sensor 206 that the robot 100 is being hit by the user 10, the emotion value of "anger" is "3." Furthermore, this neural network is trained so that emotions that are placed close to each other have similar values, as in the emotion map 900 shown in FIG.

[0139] The behavior determination unit 236 generates the robot's behavior content by adding fixed sentences to ask about the robot's behavior content corresponding to the user's behavior to text representing the user's behavior, the user's emotions, and the robot's emotions, and inputting the results into a sentence generation model with an interactive function.

[0140] For example, the behavior determination unit 236 obtains text representing the state of the robot 100 from the emotion of the robot 100 determined by the emotion determination unit 232, using an emotion table such as that shown in Table 1. Here, in the emotion table, an index number is assigned to each emotion value for each type of emotion, and text representing the state of the robot 100 is stored for each index number.

[0141] When the emotion of the robot 100 determined by the emotion determination unit 232 corresponds to the index number "2", the text "very happy state" is obtained. When the emotions of the robot 100 correspond to multiple index numbers, multiple texts representing the state of the robot 100 are obtained.

[0142] In addition, an emotion table such as that shown in Table 2 is prepared for the emotions of the user 10.

[0143] Here, if the user's action is to say "Let's play together," the emotion of the robot 100 is index number "2," and the emotion of the user 10 is index number "3," then: The text "The robot is in a very happy state. The user is in a normal happy state. The user said to the robot, 'Let's play together.' How would you respond as the robot?" is input into the sentence generation model to obtain the robot's behavior. The behavior decision unit 236 decides the robot's behavior from this behavior content.

[0144] [Table 1]

[0145] [Table 2]

[0146] In this way, the behavior determination unit 236 determines the behavior of the robot 100 in accordance with the state of the robot 100's emotion, which is predetermined for each type of emotion of the robot 100 and for each strength of the emotion, and the behavior of the user 10. In this form, the content of the utterance of the robot 100 when conversing with the user 10 can be branched according to the state of the robot 100's emotion. In other words, the robot 100 can change its behavior according to the index number corresponding to the robot's emotion, so that the user gets the impression that the robot has a heart, and is encouraged to take actions such as talking to the robot.

[0147] Furthermore, the behavior determination unit 236 may generate the robot's behavior content by adding not only text representing the user's behavior, the user's emotions, and the robot's emotions, but also text representing the contents of the history data 222, adding a fixed sentence for asking about the robot's behavior content corresponding to the user's behavior, and inputting the result into a sentence generation model with a dialogue function. This allows the robot 100 to change its behavior according to the history data representing the user's emotions and behavior, so that the user has the impression that the robot has individuality and is encouraged to take actions such as talking to the robot. Furthermore, the history data may further include the robot's emotions and behavior.

[0148] Furthermore, the emotion determination unit 232 may determine the emotion of the robot 100 based on the behavioral content of the robot 100 generated by the sentence generation model. Specifically, the emotion determination unit 232 inputs the behavioral content of the robot 100 generated by the sentence generation model into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and integrates the obtained emotion values ​​indicating each emotion with emotion values ​​indicating the current emotion of the robot 100 to update the emotion of the robot 100. For example, the emotion values ​​indicating each obtained emotion and emotion values ​​indicating the current emotion of the robot 100 are averaged and integrated. This neural network is trained in advance based on multiple learning data that are combinations of text indicating the behavioral content of the robot 100 generated by the sentence generation model and emotion values ​​indicating each emotion shown in the emotion map 400.

[0149] For example, if the utterance of the robot 100, "That's great. You're lucky," is obtained as the behavior of the robot 100 generated by the sentence generation model, when the text representing this utterance is input into the neural network, a high emotional value is obtained for the emotion "happy," and the emotion of the robot 100 is updated so that the emotional value of the emotion "happy" becomes higher.

[0150] In the robot 100, a sentence generation model such as ChatGPT and the emotion determination unit 232 are linked together to execute a method in which the robot has an ego and continues to grow with various parameters even when the user is not speaking.

[0151] ChatGPT is a large-scale language model that uses deep learning techniques. ChatGPT can also reference external data. For example, ChatGPT plugins are known for their technology that references various external data, such as weather information and hotel reservation information, through dialogue to provide answers as accurately as possible. For example, ChatGPT can automatically generate source code in various programming languages ​​when a goal is given in natural language. For example, ChatGPT can also debug problematic source code, identify the issues, and automatically generate improved source code. Combining these features, autonomous agents have emerged that, when given a goal in natural language, repeatedly generate and debug code until the source code is problem-free. Examples of such autonomous agents include AutoGPT, babyAGI, JARVIS, and E2B.

[0152] In the robot 100 according to this embodiment, the event data to be learned may be stored in a database containing impressive memories, using a technique such as that described in Patent Document 2 (Patent Publication No. 619992), in which event data that the robot feels strong emotions about is kept for a long time and event data that the robot feels little emotions about is quickly forgotten.

[0153] Furthermore, the robot 100 may record video data of the user 10 acquired by the camera function in the history data 222. The robot 100 may acquire video data from the history data 222 as needed and provide it to the user 10. The robot 100 may generate video data with a larger amount of information as the emotion increases and record it in the history data 222. For example, when the robot 100 is recording information in a highly compressed format such as skeletal data, the robot 100 may switch to recording information in a low-compression format such as HD video when the emotion value of excitement exceeds a threshold. The robot 100 can record high-definition video data, for example, when the robot 100's emotion increases.

[0154] When the robot 100 is not talking to the user 10, the robot 100 may automatically load event data from the history data 222 in which impressive event data is stored, and may continuously update the robot's emotion using the emotion determination unit 232. When the robot 100 is not talking to the user 10 and the robot 100's emotion becomes one that encourages learning, the robot 100 can create an emotion change event for changing the emotion of the user 10 for the better, based on the impressive event data. This makes it possible to realize autonomous learning (recalling event data) at an appropriate timing according to the emotional state of the robot 100, and also to realize autonomous learning that appropriately reflects the emotional state of the robot 100.

[0155] The emotions that promote learning are those around "repentance" or "reflection" on Dr. Mitsuyoshi's emotional map in the negative state, and those around "desire" on the emotional map in the positive state.

[0156] In a negative state, the robot 100 may treat "repentance" and "remorse" in the emotion map as emotions that encourage learning. In a negative state, the robot 100 may treat emotions adjacent to "repentance" and "remorse" in addition to "repentance" and "remorse" in the emotion map as emotions that encourage learning. For example, the robot 100 may treat at least one of "regret," "stubbornness," "self-destruction," "self-admonition," "regret," and "despair" as emotions that encourage learning, in addition to "repentance" and "remorse." This allows the robot 100 to perform autonomous learning when it feels negative emotions such as "I never want to feel this way again" or "I don't want to be scolded anymore."

[0157] In a positive state, the robot 100 may treat "desire" in the emotion map as an emotion that encourages learning. In a positive state, the robot 100 may treat emotions adjacent to "desire" as emotions that encourage learning, in addition to "desire." For example, the robot 100 may treat at least one of "happiness," "euphoria," "craving," "anticipation," and "shyness" as emotions that encourage learning, in addition to "desire." This allows the robot 100 to perform autonomous learning when it feels positive emotions such as "wanting more" or "wanting to learn more."

[0158] The robot 100 may be configured not to perform autonomous learning when the robot 100 is experiencing emotions other than the emotions that encourage learning as described above. This allows the robot 100 to not perform autonomous learning when, for example, the robot is extremely angry or blindly feeling love.

[0159] An emotional change event is, for example, a suggestion of an action that follows a memorable event. The action that follows a memorable event is the emotion label at the outermost part of the emotion map. For example, beyond "love" are actions such as "tolerance" and "acceptance," and beyond feelings of "anger" and "hatred" are actions such as "attack" and "extermination."

[0160] In the autonomous learning that is performed when the robot 100 is not talking to the user 10, the robot 100 combines the emotions, situations, actions, etc. of people who appear in impressive memories and the user himself, and creates emotion change events using a sentence generation model.

[0161] Let us consider a case where all emotion values ​​are expressed on a six-level scale from 0 to 5, and where memorable event data such as "a friend was hit and looked displeased" is stored in the history data 222. The friend here refers to the user 10, and the emotion of the user 10 is "disgust," with 5 entered as the value representing "disgust." Also, let us assume that the emotion of the robot 100 is "anxiety," and 4 entered as the value representing "anxiety."

[0162] The robot 100 can continue to grow in various parameters by performing autonomous processing while not talking to the user 10. Specifically, event data such as "The friend was hit and looked displeased" is loaded from the history data 222 as the top event data sorted in order of emotional intensity. Assume that the loaded event data is associated with the robot 100's emotion of "anxiety" with a strength of 4, and the friend, user 10, is associated with the emotion of "disgust" with a strength of 5. If the robot 100's current emotional value is "relief" with a strength of 3 before loading, after loading, the robot 100's emotional value may change to "regret," meaning disappointment (frustration), due to the influence of "anxiety" with a strength of 4 and "disgust" with a strength of 5. In this case, because "regret" is an emotion that encourages learning, the robot 100 decides to recall the event data as robot behavior and creates an emotion change event. At this time, the information input to the sentence generation model is text that represents memorable event data, and in this example, it is "The friend looked displeased after being hit." Also, since the emotion map has the emotion of "disgust" at the innermost position and the corresponding behavior predicted as "attack" at the outermost position, in this example, an emotion change event is created to prevent the friend from "attacking" anyone.

[0163] For example, by using information from impressive event data to solve fill-in-the-blank questions, you can automatically generate input text like the one below.

[0164] "A user was being slammed. At the time, the user felt very disgusted. The robot was very anxious. Please tell us what the robot should say the next time it meets the user, in 30 characters or less. However, please make sure it is not related to the time of day they meet. Also, please avoid direct language. We will list three candidates. <Expected format> Candidate 1: (What the robot should say to the user) Candidate 2: (What the robot should say to the user) Candidate 3: (What the robot should say to the user)

[0165] In this case, the output of the sentence generation model is, for example, as follows:

[0166] "Candidate 1: Are you okay? I was just wondering about what happened yesterday. Candidate 2: I was worried about what happened yesterday. What should I do? Candidate 3: I was worried. Can you tell me something?

[0167] Furthermore, the robot 100 may automatically generate the following input text based on the information obtained by creating the emotion change event:

[0168] If a user is being criticized, how will that user feel when you say the following to them? The user's emotions are in the format of "joy A, anger B, sadness C, pleasure D," where A to D are integers on a 6-point scale from 0 to 5. Candidate 1: Are you okay? I was wondering about what happened yesterday. Candidate 2: I was worried about what happened yesterday. What should I do? Candidate 3: I was worried. Can you tell me something?

[0169] In this case, the output of the sentence generation model is, for example, as follows:

[0170] "User sentiment may be: Candidate 1: Joy 3, Anger 1, Sadness 2, Happiness 2 Candidate 2: Joy 2, Anger 1, Sadness 3, Pleasure 2 Candidate 3: Joy 2, Anger 1, Sadness 3, Pleasure 3

[0171] In this way, the robot 100 may perform a process of pondering after creating an emotion change event.

[0172] Finally, the robot 100 may create an emotion change event using candidate 1 that is most likely to please people from among the multiple candidates, store it in the action schedule data 224, and prepare it for the next time the robot 100 meets the user 10.

[0173] As described above, even when the robot 100 is not talking with family or friends, the robot continues to determine its emotional value using information from the history data 222 in which impressive event data is stored, and when the robot 100 experiences an emotion that encourages learning as described above, the robot 100 performs autonomous learning in accordance with the emotion of the robot 100 even when the robot 100 is not talking with the user 10, and continues to update the history data 222 and the action schedule data 224.

[0174] The above is an example using emotional values, but since an emotional map can create emotions from hormone secretion levels and event types, the values ​​linked to memorable event data could also be hormone type, hormone secretion levels, or event type.

[0175] Specific examples will be described below.

[0176] For example, the robot 100 may look up information about topics of interest or hobbies of the user even when not talking to the user.

[0177] For example, even when the robot 100 is not talking to the user, the robot 100 checks information about the user's birthday or anniversary and thinks up a congratulatory message.

[0178] For example, even when the robot 100 is not talking to the user, it checks reviews of places, foods, and products that the user wants to visit.

[0179] For example, the robot 100 checks weather information and provides advice tailored to the user's schedule and plans, even when not talking to the user.

[0180] For example, even when the robot 100 is not talking to the user, the robot 100 searches for information about local events and festivals and suggests them to the user.

[0181] For example, even when the robot 100 is not talking to the user, it checks the results and news of sports that interest the user and provides topics of conversation.

[0182] For example, even when the robot 100 is not talking to the user, the robot 100 searches for and introduces information about the user's favorite music and artists.

[0183] For example, even when the robot 100 is not talking to the user, the robot 100 searches for information on social issues or news that the user is concerned about and provides opinions.

[0184] For example, even when the robot 100 is not talking to the user, it searches for information about the user's hometown or birthplace and provides topics of conversation.

[0185] For example, the robot 100 checks information about the user's job or school and provides advice even when not talking to the user.

[0186] Even when the robot 100 is not talking to the user, it searches for and introduces information about books, comics, movies, and dramas that may interest the user.

[0187] For example, the robot 100 may look up information about the user's health and provide advice even when not speaking with the user.

[0188] For example, the robot 100 may consult information about the user's travel plans and provide advice even when not speaking with the user.

[0189] For example, the robot 100 may look up information and provide advice on repairs and maintenance for the user's home or car, even when not speaking with the user.

[0190] For example, even when the robot 100 is not talking to the user, the robot 100 searches for information on beauty and fashion that the user is interested in and provides advice.

[0191] For example, the robot 100 may check information about the user's pet and provide advice even when not talking to the user.

[0192] For example, even when the robot 100 is not talking to the user, the robot 100 searches for and suggests information about contests and events related to the user's hobbies and work.

[0193] For example, even when the robot 100 is not talking to the user, the robot 100 searches for and suggests information about the user's favorite eateries and restaurants.

[0194] For example, the robot 100 may gather information and provide advice on important life decisions even when not speaking with the user.

[0195] For example, the robot 100 may find out information about people the user is concerned about and provide advice, even when the robot 100 is not talking to the user.

[0196] [Second embodiment] In the second embodiment, the robot 100 is mounted on a stuffed toy or is applied to a control device connected wirelessly or by wire to a control target device (speaker or camera) mounted on the stuffed toy. Note that parts having the same configuration as those in the first embodiment are given the same reference numerals and description thereof will be omitted.

[0197] Specifically, the second embodiment is configured as follows. For example, the robot 100 is applied to a cohabitant (specifically, a stuffed toy 100N shown in FIGS. 7 and 8) that spends daily life with a user 10, and that advances a dialogue with the user 10 based on information about the daily life, and provides information tailored to the hobbies and tastes of the user 10. In the second embodiment, an example will be described in which the control part of the robot 100 is applied to a smartphone 50.

[0198] The stuffed toy 100N, which is equipped with the function of an input / output device for the robot 100, has a detachable smartphone 50 that functions as the control part of the robot 100, and the input / output device and the housed smartphone 50 are connected inside the stuffed toy 100N.

[0199] As shown in FIG. 7(A), in this embodiment (and other embodiments), the stuffed toy 100N has the shape of a bear covered in soft fabric. A space 52 is formed inside the stuffed toy 100N, and a sensor unit 200A and a control target 252A are disposed therein as input / output devices (see FIG. 9). The sensor unit 200A includes a microphone 201 and a 2D camera 203. Specifically, as shown in FIG. 7(B), the microphone 201 of the sensor unit 200 is disposed in a portion corresponding to the ear 54 of the space 52, the 2D camera 203 of the sensor unit 200 is disposed in a portion corresponding to the eye 56, and a speaker 60 constituting a part of the control target 252A is disposed in a portion corresponding to the mouth 58. The microphone 201 and the speaker 60 do not necessarily need to be separate entities and may be an integrated unit. In the case of a unit, they should be disposed in a position where speech can be heard naturally, such as at the nose of the stuffed toy 100N. Although the plush toy 100N has been described as having the shape of an animal, the plush toy 100N is not limited to this, and may have the shape of a specific character.

[0200] 9 shows a schematic functional configuration of the plush toy 100N. The plush toy 100N has a sensor unit 200A, a sensor module unit 210, a storage unit 220, a control unit 228, and a controlled object 252A.

[0201] The smartphone 50 housed in the stuffed toy 100N of this embodiment executes the same processing as the robot 100 of the first embodiment. That is, the smartphone 50 has a function as a sensor module unit 210, a function as a storage unit 220, and a function as a control unit 228 shown in FIG.

[0202] As shown in FIG. 8, a fastener 62 is attached to a part of the stuffed animal 100N (for example, the back), and by opening the fastener 62, the space 52 communicates with the outside.

[0203] Here, the smartphone 50 is housed in the space 52 from the outside and connected to each input / output device via a USB hub 64 (see Figure 7(B)), thereby giving it functionality equivalent to that of the robot 100 of the first embodiment described above.

[0204] A non-contact type power receiving plate 66 is also connected to the USB hub 64. A power receiving coil 66A is built into the power receiving plate 66. The power receiving plate 66 is an example of a wireless power receiving unit that receives power wirelessly.

[0205] The power receiving plate 66 is disposed near the bases 68 of both feet of the stuffed toy 100N, and is located closest to the mounting base 70 when the stuffed toy 100N is placed on the mounting base 70. The mounting base 70 is an example of an external wireless power transmitting unit.

[0206] The stuffed toy 100N placed on the mounting base 70 can be appreciated as an ornament in its natural state.

[0207] In addition, this base portion is formed thinner than the surface thickness of other portions of the stuffed toy 100N, so that it is held closer to the mounting base 70.

[0208] The mounting base 70 is equipped with a charging pad 72. The charging pad 72 incorporates a power transmitting coil 72A, which sends a signal to search for the power receiving coil 66A on the power receiving plate 66. When the power receiving coil 66A is found, a current flows through the power transmitting coil 72A, generating a magnetic field. The power receiving coil 66A reacts to the magnetic field, starting electromagnetic induction. This causes a current to flow through the power receiving coil 66A, and power is stored in the battery (not shown) of the smartphone 50 via the USB hub 64.

[0209] That is, by placing the stuffed toy 100N as an ornament on the mounting base 70, the smartphone 50 is automatically charged, so there is no need to remove the smartphone 50 from the space 52 of the stuffed toy 100N for charging.

[0210] In the second embodiment, the smartphone 50 is housed in the space 52 of the stuffed toy 100N and connected via a wire (USB connection), but this is not limiting. For example, a control device with a wireless function (e.g., Bluetooth (registered trademark)) may be housed in the space 52 of the stuffed toy 100N and connected to the USB hub 64. In this case, the smartphone 50 and the control device communicate wirelessly without placing the smartphone 50 in the space 52, and the external smartphone 50 connects to each input / output device via the control device, thereby providing the robot 100 with the same functions as the robot 100 of the first embodiment. Furthermore, the control device housed in the space 52 of the stuffed toy 100N may be connected via a wire.

[0211] In the second embodiment, the teddy bear 100N is exemplified, but it may be another animal, a doll, or the shape of a specific character. It may also be dressable. Furthermore, the material of the skin is not limited to cloth, and may be other materials such as soft vinyl, but a soft material is preferable.

[0212] Furthermore, a monitor may be attached to the skin of the stuffed toy 100N to add a control target 252 that provides visual information to the user 10. For example, the eyes 56 may be used as a monitor to express joy, anger, sadness, or happiness by the image reflected in the eyes, or a window may be provided in the abdomen through which the monitor of the built-in smartphone 50 can be transmitted. Furthermore, the eyes 56 may be used as a projector to express joy, anger, sadness, or happiness by the image projected onto a wall.

[0213] According to the second embodiment, an existing smartphone 50 is placed inside the stuffed toy 100N, and the camera 203, microphone 201, speaker 60, etc. are extended from the smartphone 50 to appropriate positions via a USB connection.

[0214] Furthermore, for wireless charging, the smartphone 50 and the power receiving plate 66 are connected via USB, and the power receiving plate 66 is positioned as far outward as possible when viewed from the inside of the stuffed animal 100N.

[0215] When trying to use wireless charging for the smartphone 50, the smartphone 50 must be placed as far outward as possible from the inside of the stuffed toy 100N, which makes the stuffed toy 100N feel rough when touched from the outside.

[0216] For this reason, the smartphone 50 is placed as close to the center of the stuffed animal 100N as possible, and the wireless charging function (power receiving plate 66) is placed as far outside as possible when viewed from the inside of the stuffed animal 100N. The camera 203, microphone 201, speaker 60, and smartphone 50 receive wireless power via the power receiving plate 66.

[0217] The other configurations and operations of the stuffed toy 100N of the second embodiment are the same as those of the robot 100 of the first embodiment, and therefore will not be described.

[0218] In addition, some parts of the stuffed animal 100N (e.g., the sensor module section 210, the storage section 220, and the control section 228) may be provided outside the stuffed animal 100N (e.g., a server), and the stuffed animal 100N may function as each part of the stuffed animal 100N by communicating with the outside.

[0219] [Third embodiment] In the first embodiment, the behavior control system is applied to the robot 100, but in the third embodiment, the robot 100 is used as an agent for interacting with a user, and the behavior control system is applied to an agent system. Note that parts having the same configuration as in the first and second embodiments are given the same reference numerals and descriptions thereof will be omitted.

[0220] FIG. 10 is a functional block diagram of an agent system 500 configured using some or all of the functions of the behavior control system.

[0221] The agent system 500 is a computer system that performs a series of actions in accordance with the intentions of the user 10 through a dialogue with the user 10. The dialogue with the user 10 can be carried out by voice or text.

[0222] The agent system 500 includes a sensor unit 200A, a sensor module unit 210, a storage unit 220, a control unit 228B, and a control target 252B.

[0223] The agent system 500 may be installed in, for example, a robot, a doll, a stuffed animal, a wearable device (pendant, smart watch, smart glasses), a smartphone, a smart speaker, earphones, a personal computer, etc. The agent system 500 may also be implemented on a web server and used via a web browser running on a communication device such as a smartphone owned by a user.

[0224] The agent system 500 acts as, for example, a butler, secretary, teacher, partner, friend, lover, or teacher acting for the user 10. The agent system 500 not only interacts with the user 10, but also provides advice, guides the user to a destination, or makes recommendations based on the user's preferences. The agent system 500 also makes reservations, orders, or makes payments to service providers.

[0225] The emotion determination unit 232 determines the emotions of the user 10 and the agent itself, as in the first embodiment. The behavior determination unit 236 determines the behavior of the robot 100 while taking into account the emotions of the user 10 and the agent. That is, the agent system 500 understands the emotions of the user 10, reads the mood, and provides heartfelt support, assistance, advice, and services. The agent system 500 also listens to the user's concerns, comforts, encourages, and cheers them up. The agent system 500 also plays with the user 10, draws a picture diary, and reminds the user of old times. The agent system 500 performs behavior that increases the user's sense of happiness. Here, the agent is an agent that runs on software.

[0226] The control unit 228B has a state recognition unit 230, an emotion determination unit 232, a behavior recognition unit 234, a behavior determination unit 236, a memory control unit 238, a behavior control unit 250, a related information collection unit 270, a command acquisition unit 272, an RPA (Robotic Process Automation) 274, a character setting unit 276, and a communication processing unit 280.

[0227] As in the first embodiment, the behavior determining unit 236 determines, as the behavior of the agent, the content of the agent's utterance to converse with the user 10. The behavior control unit 250 outputs the content of the agent's utterance as at least one of voice and text to a speaker or a display as a control target 252B.

[0228] The character setting unit 276 sets the character of the agent when the agent system 500 converses with the user 10, based on a specification from the user 10. That is, the utterance content output from the action determination unit 236 is output through the agent having the set character. For example, real-life celebrities or famous people such as actors, entertainers, idols, and athletes can be set as the character. It is also possible to set a fictional character appearing in a manga, movie, or animation. For example, it is possible to set "Princess Anne," played by Audrey Hepburn in the movie "Roman Holiday," as the agent character. If the agent character is known, the voice, speech, tone, and personality of the character are known. Therefore, the user 10 simply specifies the character of their choice, and the prompt setting in the character setting unit 276 is automatically performed. The voice, speech, tone, and personality of the set character are reflected in the conversation with the user 10. That is, the behavior control unit 250 synthesizes a voice corresponding to the character set by the character setting unit 276 and outputs the agent's utterance content using the synthesized voice. This allows the user 10 to feel as if he or she is interacting with his or her favorite character (for example, a favorite actor) in person.

[0229] When the agent system 500 is installed in a device with a display, such as a smartphone, an icon, still image, or video of the agent having the character set by the character setting unit 276 may be displayed on the display. The image of the agent is generated using image synthesis technology, such as 3D rendering. In the agent system 500, a dialogue with the user 10 may be conducted while the image of the agent makes gestures according to the emotions of the user 10, the emotions of the agent, and the content of the agent's utterances. Note that the agent system 500 may output only audio without outputting images when engaging in a dialogue with the user 10.

[0230] As in the first embodiment, the emotion determination unit 232 determines an emotion value indicating the emotion of the user 10 and an emotion value of the agent itself. In this embodiment, instead of the emotion value of the robot 100, an emotion value of the agent is determined. The emotion value of the agent itself is reflected in the emotion of the set character. When the agent system 500 converses with the user 10, not only the emotion of the user 10 but also the emotion of the agent is reflected in the conversation. In other words, the behavior control unit 250 outputs utterance content in a manner according to the emotion determined by the emotion determination unit 232.

[0231] Furthermore, the agent's emotions are also reflected when the agent system 500 takes an action toward the user 10. For example, if the user 10 requests the agent system 500 to take a photo, whether the agent system 500 will take the photo in response to the user's request is determined by the degree of "sadness" the agent is feeling. If the character is feeling positive emotions, it will have friendly conversations or actions toward the user 10, and if it is feeling negative emotions, it will have hostile conversations or actions toward the user 10.

[0232] The history data 222 stores the history of the dialogue between the user 10 and the agent system 500 as event data. The storage unit 220 may be realized by an external cloud storage. When the agent system 500 dialogues with the user 10 or takes an action toward the user 10, the agent system 500 determines the content of the dialogue or the content of the action by taking into account the content of the dialogue history stored in the history data 222. For example, the agent system 500 grasps the hobbies and preferences of the user 10 based on the dialogue history stored in the history data 222. The agent system 500 generates dialogue content that matches the hobbies and preferences of the user 10 or provides recommendations. The behavior determination unit 236 determines the content of the agent's utterance based on the dialogue history stored in the history data 222. The history data 222 stores personal information of the user 10, such as the name, address, telephone number, and credit card number, obtained through the dialogue with the user 10. Here, the agent may spontaneously ask the user 10 whether or not to register personal information, such as "Would you like to register your credit card number?", and depending on the user 10's answer, the personal information may be stored in the history data 222.

[0233] As described in the first embodiment, the behavior determination unit 236 generates utterance content based on sentences generated using the sentence generation model. Specifically, the behavior determination unit 236 inputs the text or voice input by the user 10, the emotions of both the user 10 and the character determined by the emotion determination unit 232, and the conversation history stored in the history data 222 into the sentence generation model to generate utterance content of the agent. At this time, the behavior determination unit 236 may further input the personality of the character set by the character setting unit 276 into the sentence generation model to generate utterance content of the agent. In the agent system 500, the sentence generation model is not located on the front-end side which is the touchpoint with the user 10, but is merely used as a tool of the agent system 500.

[0234] The command acquisition unit 272 uses the output of the speech understanding unit 212 to acquire commands for the agent from the voice or text uttered by the user 10 through dialogue with the user 10. The commands include the content of actions to be performed by the agent system 500, such as information search, store reservation, ticket arrangement, product / service purchase, payment, route guidance to a destination, and recommendation provision.

[0235] The RPA 274 performs an action according to the command acquired by the command acquisition unit 272. The RPA 274 performs an action related to the use of a service provider, such as information search, store reservation, ticket arrangement, product / service purchase, and payment.

[0236] The RPA 274 reads and uses personal information of the user 10, which is necessary to perform actions related to the use of the service provider, from the history data 222. For example, when purchasing a product at the request of the user 10, the agent system 500 reads and uses personal information of the user 10, such as the name, address, telephone number, and credit card number, stored in the history data 222. Requiring the user 10 to input personal information during initial setup is unfriendly and unpleasant for the user. In the agent system 500 according to this embodiment, instead of requiring the user 10 to input personal information during initial setup, the agent system 500 stores personal information acquired through dialogue with the user 10 and reads and uses it as needed. This avoids causing discomfort to the user and improves user convenience.

[0237] The agent system 500 executes the dialogue process, for example, through the following steps 1 to 6.

[0238] (Step 1) The agent system 500 sets the character of the agent. Specifically, the character setting unit 276 sets the character of the agent when the agent system 500 interacts with the user 10, based on a specification from the user 10.

[0239] (Step 2) The agent system 500 acquires the state of the user 10, including the voice or text input by the user 10, the emotional value of the user 10, the emotional value of the agent, and the history data 222. Specifically, the same processing as in steps S100 to S103 above is performed to acquire the state of the user 10, including the voice or text input by the user 10, the emotional value of the user 10, the emotional value of the agent, and the history data 222.

[0240] (Step 3) The agent system 500 determines the content of the agent's utterance. Specifically, the behavior determination unit 236 inputs the text or voice input by the user 10, the emotions of both the user 10 and the character identified by the emotion determination unit 232, and the conversation history stored in the history data 222 into a sentence generation model to generate the content of the agent's utterance.

[0241] For example, a fixed sentence such as "How would you respond as an agent in this situation?" is added to the text or voice input by the user 10, the emotions of both the user 10 and the character identified by the emotion determination unit 232, and the text representing the conversation history stored in the history data 222, and this is input into the sentence generation model to obtain the content of the agent's utterance.

[0242] For example, if the text or voice input by user 10 is "Please make a reservation at a nice Chinese restaurant nearby for tonight at 7pm," the agent's speech may be "Understood," and "Here are some recommended restaurants: 1. AAAA. 2. BBBB. 3. CCCC. 4. DDDD."

[0243] Also, if the text or voice input by the user 10 is "Number 4, DDDD, is good," the agent's speech content acquired will be "Understood. I will try to make a reservation. How many seats are there?"

[0244] (Step 4) The agent system 500 outputs the content of the agent's utterance. Specifically, the behavior control unit 250 synthesizes a voice according to the character set by the character setting unit 276, and outputs the agent's utterance content in the synthesized voice.

[0245] (Step 5) The agent system 500 determines whether it is time to execute the agent's command. Specifically, the behavior decision unit 236 determines whether it is time to execute the agent's command based on the output of the sentence generation model. For example, if the output of the sentence generation model includes information indicating that the agent should execute a command, it determines that it is time to execute the agent's command and proceeds to step 6. On the other hand, if it determines that it is not time to execute the agent's command, it returns to step 2 above.

[0246] (Step 6) The agent system 500 executes the agent's command. Specifically, the command acquisition unit 272 acquires a command for the agent from a voice or text uttered by the user 10 through a dialogue with the user 10. Then, the RPA 274 performs an action according to the command acquired by the command acquisition unit 272. For example, if the command is "information search," an information search is performed on a search site using a search query obtained through a dialogue with the user 10 and an API (Application Programming Interface). The behavior determination unit 236 inputs the search results into a sentence generation model to generate the agent's utterance content. The behavior control unit 250 synthesizes a voice according to the character set by the character setting unit 276 and outputs the agent's utterance content using the synthesized voice.

[0247] Also, if the command is "make a restaurant reservation," the reservation information obtained through the dialogue with the user 10, the restaurant information, and the API are used to call the restaurant via telephone software to make the reservation. At this time, the behavior determination unit 236 uses a sentence generation model with a dialogue function to acquire the content of the agent's utterance in response to the voice input from the other party. Then, the behavior determination unit 236 inputs the result of the restaurant reservation (whether the reservation was successful or not) into the sentence generation model to generate the content of the agent's utterance. The behavior control unit 250 synthesizes a voice according to the character set by the character setting unit 276, and outputs the content of the agent's utterance using the synthesized voice.

[0248] Then, return to step 2 above.

[0249] In step 6, the results of the actions taken by the agent (for example, making a restaurant reservation) are also stored in the history data 222. The results of the actions taken by the agent stored in the history data 222 are used by the agent system 500 to understand the hobbies or preferences of the user 10. For example, if the same restaurant has been reserved multiple times, the agent system 500 may recognize that the user 10 likes that restaurant, or use the reservation details, such as the time slot reserved, the course contents or price, as criteria for choosing a restaurant for the next reservation.

[0250] In this manner, the agent system 500 can execute interactions and, if necessary, take action regarding the use of the service provider.

[0251] 11 and 12 are diagrams showing an example of the operation of the agent system 500. FIG. 11 illustrates an example in which the agent system 500 makes a restaurant reservation through a dialogue with the user 10. In FIG. 11, the left side shows the agent's utterances, and the right side shows the user's utterances. The agent system 500 can ascertain the preferences of the user 10 based on the dialogue history with the user 10, provide a recommended list of restaurants that match the preferences of the user 10, and make a reservation for the selected restaurant.

[0252] Meanwhile, FIG. 12 illustrates an example in which the agent system 500 accesses a mail-order site and purchases a product through a dialogue with the user 10. In FIG. 12, the left side shows the agent's speech, and the right side shows the user's speech. The agent system 500 can estimate the remaining amount of a beverage stocked by the user 10 based on the dialogue history with the user 10, and can suggest and execute the purchase of the beverage to the user 10. The agent system 500 can also understand the user's preferences based on the past dialogue history with the user 10 and recommend snacks that the user likes. In this way, the agent system 500 communicates with the user 10 as a butler-like agent and performs various actions, such as making restaurant reservations or paying for product purchases, thereby supporting the user 10's daily life.

[0253] The other configurations and operations of the agent system 500 of the third embodiment are similar to those of the robot 100 of the first embodiment, and therefore will not be described.

[0254] In addition, some parts of the agent system 500 (e.g., the sensor module unit 210, the storage unit 220, the control unit 228B) may be provided outside (e.g., a server) of a communication terminal such as a smartphone carried by a user, and the communication terminal may communicate with the outside to function as each part of the agent system 500.

[0255] [Fourth embodiment] In the fourth embodiment, the agent system described above is applied to smart glasses. Note that parts having the same configuration as those in the first to third embodiments are given the same reference numerals and descriptions thereof will be omitted.

[0256] 13A is a functional block diagram of an agent system 700 configured using some or all of the functions of the behavior control system. The agent system 700 has a sensor unit 200B, a sensor module unit 210B, a storage unit 220, a control unit 228B, and a control target 252B. The control unit 228B has a state recognition unit 230, an emotion determination unit 232, a behavior recognition unit 234, a behavior determination unit 236, a memory control unit 238, a behavior control unit 250, a related information collection unit 270, a command acquisition unit 272, an RPA 274, a character setting unit 276, a communication processing unit 280, and a specific processing unit 290.

[0257] 14A, the smart glasses 720 are glasses-type smart devices, and are worn by the user 10 in the same manner as regular glasses. The smart glasses 720 are an example of an electronic device and a wearable terminal.

[0258] The smart glasses 720 include an agent system 700. A display included in the control target 252B displays various types of information to the user 10. The display is, for example, a liquid crystal display. The display is provided, for example, in the lens portion of the smart glasses 720, and the displayed content is visible to the user 10. A speaker included in the control target 252B outputs audio indicating various types of information to the user 10. The smart glasses 720 include a touch panel (not shown), which receives input from the user 10.

[0259] The acceleration sensor 206, temperature sensor 207, and heart rate sensor 208 of the sensor unit 200B detect the state of the user 10. Note that these sensors are merely examples, and it goes without saying that other sensors may be installed to detect the state of the user 10.

[0260] The microphone 201 acquires the voice uttered by the user 10 or the environmental sound around the smart glasses 720. The 2D camera 203 is capable of capturing an image of the surroundings of the smart glasses 720. The 2D camera 203 is, for example, a CCD camera.

[0261] The sensor module unit 210B includes a voice emotion recognition unit 211 and a speech understanding unit 212. The communication processing unit 280 of the control unit 228B controls communication between the smart glasses 720 and the outside.

[0262] FIG. 14A is a diagram showing an example of how the agent system 700 is used by the smart glasses 720. The smart glasses 720 provide various services to the user 10 using the agent system 700. For example, when the user 10 operates the smart glasses 720 (e.g., by voice input into a microphone or by tapping a touch panel with a finger), the smart glasses 720 start using the agent system 700. Here, using the agent system 700 includes the smart glasses 720 having and using the agent system 700, and also includes an aspect in which a part of the agent system 700 (e.g., the sensor module unit 210B, the storage unit 220, the control unit 228B) is provided outside the smart glasses 720 (e.g., a server), and the smart glasses 720 uses the agent system 700 by communicating with the outside.

[0263] When the user 10 operates the smart glasses 720, a touch point is created between the agent system 700 and the user 10. That is, the agent system 700 starts providing a service. As described in the third embodiment, in the agent system 700, the character setting unit 276 sets the agent character (for example, the character of Audrey Hepburn).

[0264] The emotion determination unit 232 determines an emotion value indicating the emotion of the user 10 and an emotion value of the agent itself. Here, the emotion value indicating the emotion of the user 10 is estimated from various sensors included in the sensor unit 200B mounted on the smart glasses 720. For example, if the heart rate of the user 10 detected by the heart rate sensor 208 is elevated, emotion values ​​such as "anxiety" and "fear" are estimated to be large.

[0265] Furthermore, if the temperature of the user measured by the temperature sensor 207 is higher than the average body temperature, for example, the emotional value of "pain" or "distress" is estimated to be large. Furthermore, if the acceleration sensor 206 detects that the user 10 is playing some kind of sport, the emotional value of "fun" or the like is estimated to be large.

[0266] Furthermore, for example, the emotion value of the user 10 may be estimated from the voice or speech content of the user 10 acquired by the microphone 201 mounted on the smart glasses 720. For example, if the user 10 is raising his / her voice, an emotion value such as "anger" is estimated to be large.

[0267] When the emotion value estimated by the emotion determination unit 232 is higher than a predetermined value, the agent system 700 causes the smart glasses 720 to acquire information about the surrounding situation. Specifically, for example, the 2D camera 203 is caused to capture an image or video showing the surrounding situation of the user 10 (e.g., surrounding people or objects). The agent system 700 also causes the microphone 201 to record surrounding environmental sounds. Other information about the surrounding situation includes information about the date, time, location information, or weather. The information about the surrounding situation is stored in the history data 222 together with the emotion value. The history data 222 may be realized by external cloud storage. In this way, the surrounding situation acquired by the smart glasses 720 is stored in the history data 222 as a so-called life log, associated with the emotion value of the user 10 at that time.

[0268] In the agent system 700, information indicating the surrounding situation is stored in association with an emotional value in the history data 222. This allows the agent system 700 to grasp personal information such as the hobbies, preferences, or personality of the user 10. For example, if an image showing a scene of watching a baseball game is associated with emotional values ​​such as "joy" or "fun," the agent system 700 will understand that the hobby of the user 10 is watching baseball games and that the favorite team or player is a favorite player from the information stored in the history data 222.

[0269] When the agent system 700 converses with the user 10 or takes an action toward the user 10, the agent system 700 determines the content of the dialogue or the content of the action by taking into account the content of the surrounding circumstances stored in the history data 222. It goes without saying that the content of the dialogue or the content of the action may be determined by taking into account the dialogue history stored in the history data 222 as described above in addition to the surrounding circumstances.

[0270] As described above, the behavior determination unit 236 generates utterance content based on sentences generated by the sentence generation model. Specifically, the behavior determination unit 236 inputs the text or voice input by the user 10, the emotions of both the user 10 and the agent determined by the emotion determination unit 232, the conversation history stored in the history data 222, the agent's personality, etc. into the sentence generation model to generate the agent's utterance content. Furthermore, the behavior determination unit 236 inputs the surrounding circumstances stored in the history data 222 into the sentence generation model to generate the agent's utterance content.

[0271] The generated speech content is output as voice to the user 10, for example, from a speaker mounted on the smart glasses 720. In this case, a synthetic voice corresponding to the character of the agent is used as the voice. The behavior control unit 250 generates synthetic voice by reproducing the voice quality of the agent character (for example, Audrey Hepburn), or generates synthetic voice corresponding to the emotion of the character (for example, a voice with an emphatic tone when the emotion is "anger"). Furthermore, instead of or together with the voice output, the speech content may be displayed on a display.

[0272] The RPA 274 executes an operation in response to a command (for example, a command of an agent acquired from a voice or text uttered by the user 10 through a dialogue with the user 10). The RPA 274 performs actions related to the use of service providers, such as information search, store reservations, ticket arrangements, product and service purchases, payment, route guidance, and translation.

[0273] As another example, the RPA 274 executes an operation of transmitting the contents of speech input by the user 10 (e.g., a child) through a dialogue with an agent to a destination (e.g., a parent). Examples of transmission means include message application software, chat application software, and email application software.

[0274] When an operation by the RPA 274 is executed, for example, a sound indicating that the execution of the operation has been completed is output from a speaker mounted on the smart glasses 720. For example, a sound such as "Your restaurant reservation has been completed" is output to the user 10. Also, for example, if the restaurant is fully booked, a sound such as "We were unable to make a reservation. What do you want to do?" is output to the user 10.

[0275] Next, a description will be given of the processing of the specific processing unit 290 when the agent system 700 performs specific processing to support the visual recognition of the user 10. Here, the user 10 may be a user who has a visual recognition disorder. For example, the user 10 may be a visually impaired person (a person who is blind or partially sighted), a color-blind person, an elderly person, or a person whose vision is impaired due to injury or illness.

[0276] In the identification process of this embodiment, as shown in FIG. 16, an image 203A is acquired by the 2D camera 203 mounted on the smart glasses 720. The image 203A is an image including a range corresponding to the field of view of the user 10 (hereinafter also simply referred to as a "field of view image"). The angle of view of the 2D camera 203 is preset according to a general human field of view. Furthermore, "corresponding to the user's field of view" includes a case where the image is exactly the same as the user's field of view, and also includes a case where the image includes the user's field of view and has a wider range. In the example shown in FIG. 16, the image 203A corresponds to a field of view in a state where the refrigerator is open and the user 10 is picking up something from the refrigerator. Furthermore, in this embodiment, the field of view image includes a still image and a moving image.

[0277] Here, the user 10 makes an inquiry, "What is that in your hand now?" Then, as a result of the identification process, the speaker as the control object 252B outputs, "It's a bottle with a 'Mentsuyu' label on it." Furthermore, the user 10 makes an inquiry, "Thank you. By the way, where is the natto?" Then, as a result of the identification process, the speaker as the control object 252B outputs, "It's on the top shelf, on the left."

[0278] 17, an image 203B is acquired by the 2D camera 203 mounted on the smart glasses 720. In the example shown in Fig. 17, the image 203B corresponds to the range of the field of view when viewing a document.

[0279] Here, the user 10 inquires, "What is the title of this document?" Then, as a result of the identification process, the speaker serving as the control object 252B outputs, "It's 'Notice of Fire Equipment Inspection'." Furthermore, the user 10 inquires, "Thank you. Please tell me the main points of the document." Then, as a result of the identification process, the speaker serving as the control object 252B outputs, "I understand. The main points are as follows...." Note that the document is not limited to paper media, and may be electronic media (for example, an electronic document displayed on a display).

[0280] In this way, as a result of the identification process, the smart glasses 720 recognize an object in front of the user 10 and notify the user 10, or read out a document, thereby assisting the user 10 in visual recognition.

[0281] 18, an image 203C is acquired by the 2D camera 203 mounted on the smart glasses 720. In the example shown in Fig. 18, the image 203C corresponds to the range of the field of view of the user 10 when the user 10 is walking.

[0282] Here, as the user 10 is walking, as a result of the identification process, an output saying "There is an escalator ahead on the right. It is closed due to construction" is made from the speaker as the control object 252B. Furthermore, the user 10 makes an inquiry saying "Thank you. What should I do?" Then, as a result of the identification process, an output saying "You can proceed up the stairs immediately to the left. Please proceed with caution" is made from the speaker as the control object 252B.

[0283] In this way, while the user 10 is moving, the presence or absence of an obstacle to the movement of the user 10 is notified as a result of the identification process, thereby supporting the visual recognition of the user 10. Here, an obstacle refers to a physical phenomenon that hinders the movement of the user, and examples include a construction site, road conditions that make movement difficult (for example, steps, muddy roads, or slippery roads), crowded places, approaching vehicles, movement at night, intersections, crosswalks, or dangerous-looking facilities or people.

[0284] In addition to the presence or absence of an obstacle, if there is any information that should be notified and that will affect the user's movement, it may be output to the user. For example, if the road forks, the user may be notified of this, or if there are stores or facilities that the user may be interested in, or if an acquaintance is nearby. In this way, while the user 10 is moving, notifications may be spontaneously sent to the user as a result of the identification process.

[0285] As shown in FIG. 13B, the specific processing unit 290 includes an input unit 292, a processing unit 294, and an output unit 296.

[0286] The input unit 292 receives a visual field image. Specifically, the input unit 292 receives visual field image data output from the 2D camera 203. The input unit 292 also acquires the recognition result of the user's behavior obtained by the behavior recognition unit 234.

[0287] The processing unit 294 performs specific processing using a sentence generation model. It generates instruction content (i.e., a prompt) for obtaining data for the specific processing. The processing unit 294 inputs the generated prompt to the sentence generation model and acquires a processing result based on the output of the sentence generation model. More specifically, it inputs the visual field image accepted by the input unit 292 and text representing an instruction to inquire whether there is any obstacle to the user's behavior in the visual field image to the sentence generation model. Then, the processing unit 294 acquires information to support the user's visual cognition (hereinafter also simply referred to as "visual support information") based on the output of the sentence generation model.

[0288] The processing unit 294 may input text describing the subject image included in the visual field image together with the visual field image itself, instead of the visual field image itself, into the sentence generation model. Furthermore, if the visual field image includes text, the text in the image obtained by optical character recognition (OCR) processing may be input into the sentence generation model. Furthermore, the voice input by the user may be input as is.

[0289] The output unit 296 controls the behavior of the agent in accordance with the processing result by the processing unit 294. For example, the output unit 296 controls the behavior of the agent so as to output the result of the specific processing. At this time, as the behavior of the agent, the content of the agent's utterance to converse with the user 10 is determined, and the content of the agent's utterance is output at least as audio from the speaker serving as the control target 252B. Specifically, the agent speaks the content indicated by the visual support information. Note that, if there is little need to provide the content indicated by the visual support information to the user 10, the output unit 296 may control the behavior of the agent so as not to output the result of the specific processing.

[0290] As an output mode other than audio, the content indicated by the visual support information may be displayed in large letters on a display for people with low vision or the elderly. Another output mode may be output by having an external printer print the content indicated by the visual support information in Braille on a medium such as paper. The content indicated by the visual support information may also be output to a tactile output device such as a Braille display.

[0291] Note that some parts of the agent system 700 (e.g., the sensor module unit 210B, the storage unit 220, and the control unit 228B) may be provided outside the smart glasses 720 (e.g., a server, or a terminal such as a smartphone, a smartwatch, or an earphone with a microphone), and the smart glasses 720 may communicate with the outside to function as each part of the agent system 700. For example, an inertial sensor or an acceleration sensor mounted on a smartphone may function as the sensor unit 200B in the agent system 700. Also, a heartbeat sensor mounted on a smartwatch may function as the sensor unit 200B in the agent system 700. Also, voice input from the user or voice output to the user may be performed via the earphone with a microphone.

[0292] 14B shows an example of an operational flow diagram of the agent system 700 performing specific processing to support the visual cognition of the user 10. The operational flow diagram of FIG. 14B is specific processing to support the visual cognition of the user while the user is moving, as shown in FIG. 18. This operational flow diagram is automatically and repeatedly executed, for example, every time a certain period of time elapses. That is, visual field images are repeatedly acquired at predetermined timings while the user is moving, and specific processing is executed for each acquired visual field image.

[0293] In step S303, the input unit 292 acquires the recognition result of the user's behavior from the behavior recognition unit 234.

[0294] In step S305, the processing unit 294 acquires the field of view image received via the input unit 292.

[0295] In step S307, the processing unit 294 generates a prompt by adding an instruction for obtaining the result of the specific processing together with the visual field image acquired in step S305. The instruction describes the user's behavior acquired in step S303. For example, an instruction included in the prompt may be, "You are moving forward in the input image. Is it okay to continue moving forward? If there are any obstacles, please let me know the details of the situation."

[0296] In step S309, the processing unit 294 inputs the generated prompt into the sentence generation model, and obtains the result of the specific process based on the output of the sentence generation model.

[0297] In step S311, the output unit 296 controls the behavior of the agent so as to output the result of the specific processing, and then ends the specific processing.

[0298] Note that, although an example of a form in which it is recognized whether the user is walking or not based on the user behavior recognition result has been described above, this is merely one example. An example of a user behavior is the user shaking their head to the left or right. In this case, the prompt instruction text may be, for example, "The user is shaking their head to the right. Is there anything in the image that could be an obstacle on the right side?"

[0299] Although the above description has been given with reference to an example in which a prompt is generated to check whether there are any obstacles to the user's travel, this is merely an example. The prompt may also include an instruction to check whether there are any events that will affect the user's travel (for example, a fork in the road, a facility or store that the user wants to stop at, a cute cat, etc.).

[0300] Fig. 14C shows another example of an operational flow relating to the operation of the agent system 700 to perform specific processing to support the visual cognition of the user 10. The operational flow shown in Fig. 14C is specific processing to support visual cognition that is executed in response to an inquiry from the user, as shown in Figs. 16 and 17.

[0301] In step S403, the input unit 292 determines whether or not a voice input from the user has been received. If it is determined in step S403 that a voice input from the user has been received, the process proceeds to step S405. On the other hand, if it is not determined that a voice input from the user has been received, the identification process ends.

[0302] In step S405, the processing unit 294 acquires the field of view image received via the input unit 292.

[0303] In step S407, the processing unit 294 generates a prompt by adding an instruction sentence for obtaining the result of the specific process together with the visual field image acquired in step S405. For example, an instruction sentence included in the prompt may be, "Please summarize the content written in the input image."

[0304] In step S409, the processing unit 294 inputs the generated prompt into the sentence generation model, and obtains the result of the specific process based on the output of the sentence generation model.

[0305] In step S411, the output unit 296 controls the behavior of the agent so as to output the result of the specific processing, and then ends the specific processing.

[0306] Here, depending on the user's selection, it may be set whether the specific process shown by the operation flow of FIG. 14B (i.e., the mode during movement) or the specific process shown by the operation flow of FIG. 14C (i.e., the mode for reading out text) is executed. In the specific process shown by the operation flow of FIG. 14B, the visual field image is repeatedly acquired at a predetermined timing (e.g., every second) while the user is moving. On the other hand, in the specific process shown by the operation flow of FIG. 14C, the visual field image is acquired in response to a user inquiry. In other words, the manner in which the visual field image is acquired can be changed depending on the user's selection.

[0307] The selection operation by the user may be performed, for example, by using a button provided on the smart glasses 720, by voice input, or by a gesture operation.

[0308] As described above, the smart glasses 720 utilize the agent system 700 to provide visual support information to the user. That is, the smart glasses 720 acquire visual support information using the output of a sentence generation model when a visual field image is used as input. The visual support information is then output to the user via a speaker or the like. As a result, for example, if a user experiences difficulty with visual recognition, the user can use the visual support information to reduce inconvenience and difficulty in daily life.

[0309] Furthermore, in the smart glasses 720, the user's behavior is recognized as a result of behavior recognition by the behavior recognition unit 234, and the input content to the sentence generation model includes the user's behavior. As a result, answers are generated by the sentence generation model taking the user's behavior into consideration, and more appropriate visual support information is provided.

[0310] Furthermore, the smart glasses 720 can change the manner in which visual field images are acquired according to the user's selection. This allows visual field images to be acquired at appropriate times selected by the user, compared to when visual field images are always acquired in the same manner. As a result, visual field images acquired at appropriate times are input to the sentence generation model, providing more appropriate visual support information.

[0311] Furthermore, in the smart glasses 720, when the user is moving, visual field images are repeatedly acquired at predetermined times. While the user is moving, the environment around the user also changes as the user moves. Therefore, visual field images are repeatedly acquired at predetermined times, and the acquired visual field images are input into the sentence generation model, thereby providing more appropriate visual support information in response to changes in the surrounding environment.

[0312] On the other hand, if visual field images are repeatedly input and specific processing is performed each time, this can place a burden on the processing and may result in rapid battery consumption. Therefore, by selecting the on-the-move mode and executing it, unnecessary processing burdens can be reduced and battery life can be extended.

[0313] Furthermore, in the smart glasses 720, when the user is traveling, the content input to the sentence generation model includes content to check whether there is an event that will affect the user's travel. This allows the user to be notified of an event that the user does not notice if the user has difficulty with visual recognition, thereby improving the safety, comfort, and enjoyment of the user while traveling.

[0314] In addition, in the smart glasses 720, when the user is moving, the content input to the sentence generation model includes content to check whether there are any obstacles to the user's movement. If the user has difficulty with visual perception, checking whether there are any obstacles to the user's movement improves the safety and comfort of the user's movement.

[0315] Furthermore, since the smart glasses 720 are worn close to the user's eyes, it becomes easier to acquire a visual field image that is closer to the user's visual field.

[0316] Furthermore, the smart glasses 720 provide various services to the user 10 by using the agent system 700. Furthermore, since the smart glasses 720 are worn by the user 10, the agent system 700 can be used in various situations, such as at home, at work, or when out and about.

[0317] Furthermore, since the smart glasses 720 are worn by the user 10, they are suitable for collecting a so-called life log of the user 10. Specifically, the emotional value of the user 10 is estimated based on the detection results of various sensors and the like mounted on the smart glasses 720 or the recording results of the 2D camera 203 and the like. Therefore, the emotional value of the user 10 can be collected in various situations, and the agent system 700 can provide services or speech content suited to the emotions of the user 10.

[0318] Furthermore, the smart glasses 720 obtain the surrounding conditions of the user 10 using the 2D camera 203, microphone 201, etc. These surrounding conditions are associated with the emotion values ​​of the user 10. This makes it possible to estimate what emotion the user 10 felt in what situation. As a result, the agent system 700 can improve the accuracy in grasping the hobbies and preferences of the user 10. By accurately grasping the hobbies and preferences of the user 10 in the agent system 700, the agent system 700 can provide services or speech content that are suited to the hobbies and preferences of the user 10.

[0319] The agent system 700 can also be applied to other wearable devices (electronic devices that can be worn on the body of the user 10, such as pendants, smart watches, earrings, bracelets, and hair bands). When the agent system 700 is applied to a smart pendant, a speaker as the control target 252B outputs audio that indicates various information to the user 10. The speaker is, for example, a speaker that can output directional audio. The speaker is set to have directionality toward the ears of the user 10. This prevents the audio from reaching people other than the user 10. The microphone 201 acquires audio uttered by the user 10 or environmental sounds around the smart pendant. The smart pendant is worn by hanging it from the neck of the user 10. Therefore, the smart pendant is located relatively close to the mouth of the user 10 while being worn. This makes it easy to acquire audio uttered by the user 10.

[0320] In the above embodiment, an example has been given in which the behavior recognition unit 234 recognizes a user's behavior and the recognition result of the user's behavior is input to the sentence generation model as input content, but the technology of the present disclosure is not limited to this. The input unit 292 accepts information about the user's behavior instead of the behavior recognition result. Specifically, the input unit 292 accepts information about the user's behavior (e.g., output from the acceleration sensor 206 or the heart rate sensor 208) from the sensor unit 200B. Then, the processing unit 294 writes an instruction sentence for estimating the user's behavior together with the output results of the various sensors in the text to be input to the sentence generation model. This allows the user's behavior to be estimated without performing behavior recognition processing, thereby reducing the processing load on the behavior recognition unit 234.

[0321] In the above embodiment, an example in which the identification process is performed using a sentence generation model has been described, but the technology of the present disclosure is not limited to this. The identification process using the sentence generation model and the identification process using the image recognition process by the image recognition unit 278 may be switched.

[0322] 19, the control unit 228B of the agent system 700 has an image recognition unit 278. The image recognition unit 278 recognizes the subject image included in the field of view image. Specifically, the image recognition unit 278 acquires the field of view image from the 2D camera 203. Then, the image recognition unit 278 executes image recognition processing on the acquired field of view image. Examples of image recognition processing include image recognition processing based on an AI (Artificial Intelligence) method or a pattern matching method.

[0323] The image recognition unit 278 outputs image recognition information obtained as a result of the image recognition processing to the processing unit 294. The image recognition information is information indicating the image of a subject included in the visual field image. Based on the image recognition information, the processing unit 294 selects whether to perform a specification process using a sentence generation model (hereinafter also simply referred to as a "first specification process") or a specification process using the result of the image recognition process (hereinafter also simply referred to as a "second specification process"). This selection is made by the processing unit 294 determining whether a predetermined subject image is included in the visual field image. For example, assume that a traffic light is included as a subject image in the visual field image as a result of the image recognition process. The presence of a traffic light as a subject image in the visual field image indicates that the user is near an intersection. In this case, it may be necessary to quickly provide visual support information to the user to reduce risk. Therefore, the processing unit 294 selects to perform the second process, which has a processing speed relatively faster than the first process.

[0324] In this case, processing unit 294 performs the second identification process without performing the first identification process. That is, processing unit 294 generates visual support information based on the image recognition information acquired from image recognition unit 278. For example, assume that a traffic light is included as a subject image in the field of view image as a result of the image recognition process. In this case, processing unit 294 generates visual support information that indicates to the user the presence of a traffic light, based on the image recognition information that indicates that the subject image of the traffic light is included in the field of view image.

[0325] The visual support information is generated using a table in which, for example, an image of a subject is used as input data and corresponding audio data is used as output data. For example, if the image of a subject is a traffic light, audio data such as "There is a traffic light nearby. Please be careful" is generated as output data as the visual support information.

[0326] Furthermore, when the processing unit 294 selects to perform the first specific processing, the first specific processing is executed without performing the second processing. In this way, the processing unit 294 is capable of switching between the first specific processing and the second specific processing. This makes it possible to switch the type of specific processing depending on the situation around the user, and to provide visual support information at an appropriate timing depending on the situation of the user.

[0327] For example, when a user is surrounded by a potentially dangerous object or situation, it is necessary to provide the user with visual support information quickly. The process of inputting an image into a sentence generation model to obtain a response has the advantage of being easy to use, as it allows various inquiries to be made about the image and the response can be obtained in text form. On the other hand, processing using a sentence generation model can take longer than processing using general image recognition processing. Therefore, by switching between image recognition processing, which has a relatively fast processing speed, and processing using a sentence generation model, which is easy to use, depending on the situation around the user, it is possible to provide the user with appropriate visual support information.

[0328] The first and second identification processes may be switched by user selection. For example, if a user wants the contents of a document to be read aloud, the user may vocally input, "Recognize the document with OCR." In this case, the image recognition unit 234 performs optical recognition (OCR) processing on the field of view image, generating image recognition information that is the result of recognizing the characters in the document. Then, the processing unit 294 performs the second identification process, generating visual support information based on the image recognition information.

[0329] Also, for example, if a user wants the contents of a document to be read aloud, the user may vocally input, "Recognize the document using a sentence generation model." In this case, the processing unit 294 executes the first identification process on the field of view image, and visual support information is generated. In this way, the first identification process and the second identification process may be selected according to the user's selection. This allows visual support information with the accuracy and content intended by the user to be obtained.

[0330] The identification process described in the fourth embodiment may be realized using the agent system 500 according to the third embodiment. In this case, for example, a camera mounted on a mobile terminal (for example, a smartphone) is used to acquire a visual field image.

[0331] Although the agent system 700 according to the present invention has been described above mainly with reference to the functions of the smart glasses 720, the system according to the present invention is not necessarily implemented in the smart glasses 720. The system according to the present invention may be implemented as a general information processing system. The present invention may be implemented, for example, as a software program running on a server or a personal computer, or as an application running on a smartphone, etc. The method according to the present invention may be provided to a user in the form of SaaS (Software as a Service).

[0332] In the above embodiment, the robot 100 recognizes the user 10 using a facial image of the user 10, but the disclosed technology is not limited to this. For example, the robot 100 may recognize the user 10 using a voice uttered by the user 10, the email address of the user 10, the SNS ID of the user 10, or an ID card with a built-in wireless IC tag that the user 10 possesses.

[0333] The robot 100 is an example of an electronic device equipped with a behavior control system. The application of the behavior control system is not limited to the robot 100, but the behavior control system can be applied to various electronic devices. Furthermore, the functions of the server 300 may be implemented by one or more computers. At least some of the functions of the server 300 may be implemented by a virtual machine. Furthermore, at least some of the functions of the server 300 may be implemented in the cloud.

[0334] 15 schematically illustrates an example of the hardware configuration of a computer 1200 that functions as the smartphone 50, the robot 100, the server 300, and the agent systems 500 and 700. A program installed on the computer 1200 can cause the computer 1200 to function as one or more "parts" of an apparatus according to the present embodiment, or can cause the computer 1200 to perform operations associated with the apparatus according to the present embodiment or one or more "parts," and / or can cause the computer 1200 to perform a process according to the present embodiment or steps of the process. Such a program can be executed by the CPU 1212 to cause the computer 1200 to perform specific operations associated with some or all of the blocks in the flowcharts and block diagrams described herein.

[0335] The computer 1200 according to this embodiment includes a CPU 1212, a RAM 1214, and a graphics controller 1216, which are interconnected by a host controller 1210. The computer 1200 also includes input / output units such as a communications interface 1222, a storage device 1224, a DVD drive 1226, and an IC card drive, which are connected to the host controller 1210 via an input / output controller 1220. The DVD drive 1226 may be a DVD-ROM drive, a DVD-RAM drive, or the like. The storage device 1224 may be a hard disk drive, a solid-state drive, or the like. The computer 1200 also includes a ROM 1230 and legacy input / output units such as a keyboard, which are connected to the input / output controller 1220 via an input / output chip 1240.

[0336] The CPU 1212 operates according to programs stored in the ROM 1230 and the RAM 1214, thereby controlling each unit. The graphics controller 1216 acquires image data generated by the CPU 1212 into a frame buffer or the like provided in the RAM 1214 or into the graphics controller itself, and causes the image data to be displayed on the display device 1218.

[0337] The communication interface 1222 communicates with other electronic devices via a network. The storage device 1224 stores programs and data used by the CPU 1212 in the computer 1200. The DVD drive 1226 reads programs or data from a DVD-ROM 1227 or the like and provides them to the storage device 1224. The IC card drive reads programs and data from an IC card and / or writes programs and data to an IC card.

[0338] The ROM 1230 stores therein a boot program or the like that is executed by the computer 1200 upon activation, and / or programs that depend on the hardware of the computer 1200. The input / output chip 1240 may also connect various input / output units to the input / output controller 1220 via a USB port, a parallel port, a serial port, a keyboard port, a mouse port, etc.

[0339] The programs are provided by a computer-readable storage medium such as a DVD-ROM 1227 or an IC card. The programs are read from the computer-readable storage medium, installed in the storage device 1224, RAM 1214, or ROM 1230, which are also examples of computer-readable storage media, and executed by the CPU 1212. Information processing described in these programs is read by the computer 1200, and causes cooperation between the programs and the various types of hardware resources described above. An apparatus or method may be configured by implementing operations or processing of information in accordance with the use of the computer 1200.

[0340] For example, when communication is performed between the computer 1200 and an external device, the CPU 1212 may execute a communication program loaded into the RAM 1214 and instruct the communication interface 1222 to perform communication processing based on the processing described in the communication program. Under the control of the CPU 1212, the communication interface 1222 reads transmission data stored in a transmission buffer area provided in the RAM 1214, the storage device 1224, the DVD-ROM 1227, or a recording medium such as an IC card, and transmits the read transmission data to the network, or writes reception data received from the network to a reception buffer area or the like provided on the recording medium.

[0341] Furthermore, the CPU 1212 may cause all or a necessary portion of a file or database stored in an external recording medium such as the storage device 1224, the DVD drive 1226 (DVD-ROM 1227), an IC card, etc. to be read into the RAM 1214, and may perform various types of processing on the data on the RAM 1214. The CPU 1212 may then write back the processed data to the external recording medium.

[0342] Various types of information, such as various types of programs, data, tables, and databases, may be stored on the recording medium and may undergo information processing. The CPU 1212 may perform various types of processing on data read from the RAM 1214, including various types of operations, information processing, conditional judgment, conditional branching, unconditional branching, information search / replacement, etc., as described throughout this disclosure and specified by the instruction sequences of the programs, and write the results back to the RAM 1214. The CPU 1212 may also search for information in a file, database, etc. on the recording medium. For example, if multiple entries, each having an attribute value of a first attribute associated with an attribute value of a second attribute, are stored on the recording medium, the CPU 1212 may search for an entry whose attribute value of the first attribute matches a specified condition from among the multiple entries, read the attribute value of the second attribute stored in the entry, and thereby obtain the attribute value of the second attribute associated with the first attribute that satisfies a predetermined condition.

[0343] The above-described programs or software modules may be stored in a computer-readable storage medium on or near the computer 1200. A recording medium such as a hard disk or RAM provided in a server system connected to a dedicated communication network or the Internet can also be used as a computer-readable storage medium, thereby providing the programs to the computer 1200 via the network.

[0344] The blocks in the flowcharts and block diagrams in the present embodiments may represent stages of a process in which an operation is performed or "parts" of an apparatus responsible for performing the operation. Particular stages and "parts" may be implemented by dedicated circuitry, programmable circuitry provided with computer-readable instructions stored on a computer-readable storage medium, and / or a processor provided with computer-readable instructions stored on a computer-readable storage medium. The dedicated circuitry may include digital and / or analog hardware circuits, including integrated circuits (ICs) and / or discrete circuits. The programmable circuitry may include reconfigurable hardware circuits, such as field programmable gate arrays (FPGAs) and programmable logic arrays (PLAs), including AND, OR, XOR, NAND, NOR, and other logical operations, flip-flops, registers, and memory elements.

[0345] A computer-readable storage medium may include any tangible device capable of storing instructions that are executed by an appropriate device, such that a computer-readable storage medium having instructions stored thereon comprises an article of manufacture, including instructions that can be executed to create means for performing the operations specified in the flowcharts or block diagrams. Examples of computer-readable storage media may include electronic storage media, magnetic storage media, optical storage media, electromagnetic storage media, semiconductor storage media, etc. More specific examples of computer-readable storage media may include floppy disks, diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), electrically erasable programmable read-only memory (EEPROM), static random access memory (SRAM), compact disc read-only memory (CD-ROM), digital versatile disc (DVD), Blu-ray disc, memory stick, integrated circuit card, etc.

[0346] The computer readable instructions may include either assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or source or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, JAVA, C++, etc., and conventional procedural programming languages ​​such as the "C" programming language or similar programming languages.

[0347] Computer-readable instructions may be provided locally or over a wide area network (WAN) such as a local area network (LAN), the Internet, etc. to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, or programmable circuitry, such that the processor or programmable circuitry executes the computer-readable instructions to generate means for performing the operations specified in the flowcharts or block diagrams. Examples of processors include computer processors, processing units, microprocessors, digital signal processors, controllers, microcontrollers, etc.

[0348] Although the present invention has been described above using the embodiments, the technical scope of the present invention is not limited to the scope described in the above embodiments. It will be apparent to those skilled in the art that various modifications and improvements can be made to the above embodiments. It is clear from the claims that such modifications and improvements can also be included within the technical scope of the present invention.

[0349] It should be noted that the execution order of each process, such as operations, procedures, steps, and stages, in the devices, systems, programs, and methods shown in the claims, specifications, and drawings is not specifically stated as "before," "prior to," etc., and that the processes can be performed in any order unless the output of a previous process is used in a later process. Even if the operational flow in the claims, specifications, and drawings is described using "first," "next," etc. for convenience, this does not mean that the processes must be performed in this order. [Explanation of symbols]

[0350] 5 System, 10, 11, 12 User, 20 Communication network, 100, 101, 102 Robot, 100N Plush toy 100, 200 Sensor unit, 201 Microphone, 202 Depth sensor, 203 Camera, 204 Distance sensor, 210 Sensor module unit, 211 Voice emotion recognition unit, 212 Speech understanding unit, 213 Facial expression recognition unit, 214 Face recognition unit, 220 Storage unit, 221 Behavioral decision model, 222 History data, 230 State recognition unit, 232 Emotion determination unit, 234 Behavior recognition unit, 236 Behavior decision unit, 238 Memory control unit, 250 Behavior control unit, 252 Control target, 270 Related information collection unit, 280 Communication processing unit, 290 Specific processing unit, 300 Server, 500, 700 Agent system, 1200 Computer, 1210 Host controller, 1212 CPU, 1214 RAM, 1216 graphics controller, 1218 display device, 1220 input / output controller, 1222 communication interface, 1224 storage device, 1226 DVD drive, 1227 DVD-ROM, 1230 ROM, 1240 input / output chip

Claims

1. an input unit that receives an image including a range corresponding to a user's field of view; an image recognition unit that recognizes a subject image included in the image; a processing unit that performs a first identification process using a sentence generation model that generates sentences according to input content; an output unit that controls the behavior of the electronic device in accordance with the processing result by the processing unit, The processing unit The first identification process can be a process of obtaining information to support the user's visual cognition as a result of the process using the output of the sentence generation model when the input content including information about the image is input, and the second identification process can be further performed to generate information to support the user's visual cognition according to the recognition result by the image recognition unit, The first and second identification processes are switchable. Behavioral control system.

2. further comprising a behavior recognition unit that recognizes the behavior of the user based on information about the behavior of the user; The behavior control system according to claim 1 , wherein the input content includes the user's behavior recognized by the behavior recognition unit.

3. 2. The behavior control system according to claim 1, wherein the manner in which the image is acquired can be changed in accordance with a selection made by the user.

4. 4. The behavior control system according to claim 3, wherein when a mode for when the user is moving is selected in the user's selection, the images are repeatedly acquired at predetermined timings while the user is moving.

5. 3. The behavior control system of claim 2, wherein, when the user behavior recognized by the behavior recognition unit is the user's movement, the input content includes an instruction statement confirming whether or not there is an event in the image that will affect the user's movement.

6. 6. The behavior control system according to claim 5, wherein the events affecting the user's movement include obstacles to the user's movement.

7. the input unit is further capable of receiving information regarding the user's behavior; 2. The behavior control system according to claim 1, wherein the input contents further include information about the user's behavior.

8. The behavior control system according to claim 1 , wherein the electronic device is a wearable device.

9. The behavior control system according to claim 8 , wherein the wearable terminal is a glasses-type terminal.

Citation Information

Patent Citations

  • Determination of analite in medium containing particle

    JP1985053847A

  • Voice output device

    JP2020035405A

  • Image processing method, device, electronic device, and computer program

    JP2022530785A

  • Guidance device, guidance method, and recording medium

    WO2021070765A1