Center device, dialogue system, and dialogue method

The dialogue system addresses labor shortages and interaction quality issues by using sensors to estimate user states and adjust dialogue frequency and nature based on user profiles, ensuring personalized and efficient interactions.

JP2025144468APending Publication Date: 2025-10-02SECOM CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024044267
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-19
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Existing communication systems face challenges in providing personalized and efficient dialogue responses due to labor shortages and varying quality of interactions, whether through human operators or automated systems, especially in casual conversations.

Method used

A dialogue system that utilizes sensors to estimate user states, generates tailored speech messages based on user profiles and responses, and adjusts the frequency and nature of interactions to match individual user preferences and needs.

Benefits of technology

Enables personalized and efficient dialogue tailored to individual users, reducing operational burden and improving interaction quality by adapting to user reactions and preferences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025144468000001_ABST
    Figure 2025144468000001_ABST
Patent Text Reader

Abstract

To perform a dialogue depending on an individual user.SOLUTION: A dialogue system comprises a dialogue device and one or more sensors existing in a user's living space. The dialogue device includes: sensor data reception means for receiving sensor data output from the sensors; state estimation means for estimating a state on the user in the living space on the basis of the sensor data; message generation means for generating an utterance message depending on the state estimated by the state estimation means; output means for performing voice output of the generated utterance message; input means for performing voice input of the user's response message to the utterance message of which the voice output is performed; and determination means for determining necessity or frequency of the utterance message to be output by voice for the state estimated by the state estimation means, on the basis of the response message.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a dialogue system and dialogue method using a terminal device having a voice input / output function, and a center device. [Background technology]

[0002] Loneliness is a factor that leads to serious health risks such as dementia, and the increase in single-person households (especially elderly people living alone) coupled with the recent increased risk of infectious diseases makes loneliness more likely to occur, making alleviating loneliness an important social issue. Casual conversation and other everyday interactions are thought to be effective in alleviating loneliness.

[0003] In recent years, communication services have been proposed that aim to alleviate feelings of loneliness by installing devices with voice input and output functions (e.g., interactive robots) in the homes of elderly people, especially those living alone, and engaging in everyday conversations such as casual voice chat through these devices.

[0004] In this system, when a user speaks to a device at home, the content is converted into a string of characters using voice recognition technology and sent as a message to the management center of the service provider. At the management center, an operator who notices the receipt of the message types a reply into the center device, which is then sent back to the user. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Japanese Patent Publication No. 2021-157419 [Patent Document 2] Japanese Patent Application Laid-Open No. 2007-286376 Summary of the Invention [Problem to be solved by the invention]

[0006] There are currently both human-assisted interactive services and automated response services using communication robots. However, labor shortages due to the declining labor force are also a social issue, and there are limits to how much human operators can provide. When an operator types a reply to a user, they must understand the message and assess the user's condition based on their memory and experience of previous messages exchanged with the user before creating a reply. However, the more users a single operator handles, the more difficult this becomes. Furthermore, the operational burden increases, as does the difficulty of handing over tasks when an operator is replaced. On the other hand, communication robots that do not require human intervention may be disliked by users due to the low quality of their interactions (e.g., inappropriate responses, formulaic responses, etc.).

[0007] Patent Document 1 proposes a UI that supports the operator's responses by displaying candidate responses for each category (affirmative, counterargument, topic change) along with an evaluation value. However, when an operator responds to a message from a user, it is a heavy burden for the operator to think of a response message for each message, and there is a risk that the quality of the response messages will vary depending on the operator. Although Patent Document 1 displays candidate responses to the operator, the operator must check each candidate and decide on a response message, which still places a burden on the operator.

[0008] Patent Document 2 proposes a technology that switches between automatic responses and responses via an operator by requesting a remote support device for response assistance when a robot cannot decide on a response message. However, in conversations where it is not easy for a robot to determine the appropriateness of a response message, such as everyday conversations such as casual conversations, switching to an operator response each time may place an excessive burden on the operator.

[0009] In view of the above circumstances, an object of the present invention is to provide a dialogue system that is capable of conducting dialogue tailored to individual users. [Means for solving the problem]

[0010] A dialogue system according to one aspect of the present invention comprises: one or more sensors in the user's living space; an interactive device; A dialogue system comprising: The dialogue device a sensor data receiving means for receiving sensor data output from the sensor; a state estimation means for estimating a state related to a user in the living space based on the sensor data; a message generating means for generating a speech message according to the state estimated by the state estimating means; an output means for audibly outputting the generated speech message; an input means for inputting a user's response message by voice in response to the spoken message outputted by voice; a determination means for determining whether or not the speech message to be output by voice in response to the state estimated by the state estimation means is necessary or how often the speech message is to be output by voice, based on the response message; It has.

[0011] When a user's state is estimated, some users feel happy to be spoken to about that state, while other users feel annoyed that they do not want to be spoken to much when they are in that state. Therefore, according to this embodiment, a center device estimates the user's state (user state, user behavior, home state) based on the detection results from sensors installed in the user's home, and a terminal device outputs a speech message according to the estimation result. The center device determines whether or not (or how often) the user needs to speak about the state based on the user's response message to the speech message. This makes it possible to easily understand through natural communication between the user and the terminal device whether or not the detected state is one in which the user wants to be spoken to.

[0012] The determination means may determine a polarity indicating a negative / positive reaction of the user to the speech message based on the response message, and determine the necessity or the frequency based on the polarity.

[0013] This makes it possible to change the frequency of utterances to be reduced (or to be unnecessary) in the case of a negative response (including no response), for example.

[0014] The determination means may determine a polarity value representing the magnitude of the user's negative / positive reaction to the speech message based on the response message, and determine the necessity or frequency based on the polarity value.

[0015] This makes it possible to change the frequency of utterances so that, for example, the stronger the "strength" of a negative response, the less frequently the response is spoken.

[0016] the dialogue device further comprises a storage means for storing profile information representing characteristics of the user with respect to each state; The determination means may determine the user's negative / positive reaction to the speech message based on the response message and change the profile information, and may determine the necessity or frequency of the speech message based on the profile information.

[0017] This allows the user profile of the user (for example, the average polarity value for each state over a predetermined period in the past, the number of positive or negative responses (the number of times each polarity) etc.) to be changed depending on the determination result.

[0018] As the profile information, information relating to the reaction for each time period of a day for each state of the user is stored; The determining means may determine the necessity or frequency of the spoken message based on the profile information of the time period corresponding to the estimated time of the state estimated by the state estimating means.

[0019] This allows, for example, a user who tends to give positive (favorable) response messages regarding a predetermined state in the morning time zone to have the frequency of utterance messages regarding the predetermined state increased in the morning time zone.

[0020] As the profile information, an average polarity value of all states of the user and a polarity value relating to each state of the user are stored; The determining means may determine whether or not the spoken message is necessary or how often the spoken message is required when the state estimating means estimates the predetermined state, based on the difference between the average polarity value and the polarity value in the predetermined state.

[0021] This allows us to determine the necessity or frequency of a spoken message for each state based on whether there is a significant deviation from the user's general tendency to respond negatively / positively (i.e., the average polarity value for all states).

[0022] an importance level indicating the degree of importance of each of the states is stored in association with each of the states; The determining means may determine whether or not the spoken message is necessary or how often the spoken message is to be sent based on the importance of the state estimated by the state estimating means.

[0023] This makes it possible to change the speech frequency depending on the importance of the estimated state, for example, not to reduce the speech frequency too much if the state is important.

[0024] The determining means may determine whether or not the spoken message is necessary or how often the spoken message is necessary based on the security status of the user.

[0025] This means that the user is emotionally on guard, so by sending spoken messages as frequently as possible, communication that is sensitive to the user's anxieties can be realized.

[0026] The message generating means may determine the strength of the speech message based on the profile information.

[0027] For example, when the average polarity value is large, the message is likely to be a strong message because it is a message that is needed by the user. Conversely, when the average polarity value is small, the message is likely to be a weak message because it is a message that is likely to be avoided by the user.

[0028] A center device according to one aspect of the present invention comprises: a sensor data receiving means for receiving sensor data output from one or more sensors in the user's living space; a state estimation means for estimating a state related to a user in the living space based on the sensor data; a message generating means for generating a speech message according to the state estimated by the state estimating means; a response sending means for sending the speech message to a terminal device which exchanges messages with a user in a two-way manner; a message receiving means for receiving a response message from the user in response to the spoken message from the terminal device; a determining means for determining whether or not the spoken message is necessary or how often the spoken message is to be sent when the state is estimated by the state estimating means based on the response message; It is equipped with:

[0029] A dialogue method according to one aspect of the present invention includes: The computer of the center device receiving sensor data output from one or more sensors in the user's living space; estimating a state of a user in the living space based on the sensor data; generating a speech message according to the estimated state; transmitting the spoken message to a terminal device that exchanges messages with a user; receiving a response message from the user to the spoken message from the terminal device; determining whether or not the speech message is necessary or how often the speech message is necessary when the state is estimated based on the response message; Execute. [Effects of the Invention]

[0030] According to the present invention, it is possible to provide a dialogue system that is capable of conducting dialogue according to individual users.

[0031] The effects described here are not necessarily limited to those described above, and may be any of the effects described in the present invention. [Brief explanation of the drawings]

[0032] [Figure 1] 1 is a diagram showing a configuration of a dialogue system according to an embodiment of the present invention. [Figure 2] FIG. 2 is a diagram illustrating a hardware configuration of a center device. [Figure 3] An overview of the dialogue system is shown below. [Figure 4] The operation flow of the dialogue system is shown below. [Figure 5] 10 shows an example of the configuration of sensor information. [Figure 6] 10 shows an example of the configuration of estimation condition information. [Figure 7] 10 shows an example of the configuration of estimation information. [Figure 8] 1 illustrates an embodiment of a dialogue system. [Figure 9] 10 shows an example of the configuration of user profile information. DETAILED DESCRIPTION OF THE INVENTION

[0033] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.

[0034] 1. Dialogue system configuration

[0035] FIG. 1 is a diagram showing the configuration of a dialogue system according to this embodiment.

[0036] The dialogue system 10 realizes a communication service in which everyday conversations such as casual chatting are mainly conducted. The dialogue system 10 is difficult to standardize responses, anticipates a wide variety of responses, and realizes everyday conversations such as continuous casual chatting with multiple users. Unlike automatic response systems for product guidance, contract procedures, etc., which can be implemented without the human intervention of an operator, the dialogue system 10 requires the human intervention of an operator to ensure the quality of the dialogue, as emotions such as empathy and intentions are important.

[0037] The dialogue system 10 is composed of a center device 100 as a server device managed by a business operator that provides services through the dialogue system 10, an operator device 160 connected to the center device 100, multiple terminal devices 200 connected to a network so as to be able to communicate with the center device 100, and one or more sensors 300 in the user's living space.

[0038] The terminal device 200 is installed in the home of a user (e.g., an elderly person) of this system. The terminal device 200 has an input means 201, a message sending means 202, and a response output means 203. The input means 201 inputs a response user message from the user via a microphone M. The message sending means 202 sends the response message input via the input means 201 to the center device 100 via the network. The response output means 203 receives a spoken message from the center device 100 via the network, and outputs the message as voice to the user from a speaker S.

[0039] The terminal device 200 may have at least the above components, but may also be an interactive robot with a small doll-like appearance that makes the user feel like they are having a conversation with a person (especially everyday conversation such as casual chatting).

[0040] The one or more sensors 300 in the user's living space may be, for example, a door sensor, a PIR (Passive Infrared Ray) motion sensor, a smart tap, a CO2 sensor, an environmental sensor, or a vital sensor mounted on a wearable device (such as a smart watch or healthcare device). A smart tap (smart plug) is a plug-type device equipped with a wireless network communication function, which is connected between an outlet and the cable of a home appliance, and connects the connected home appliance to the IoT. The living space refers to a space inside a house, a building, a premises, or a room.

[0041] The center device 100 is managed by a business operator that provides services using the dialogue system 10, and includes at least a sensor data receiving means 101, a state estimating means 102, a message generating means 103, a response transmitting means 104, a message receiving means 105, a determining means 106, and a speech necessity determining means 107. It also includes a control means for controlling these means. These are realized by well-known hardware (so-called server computers or personal computers) or software stored in a storage means 140, as appropriate.

[0042] The storage means 140 stores a dictionary 141 , one or more pieces of profile information 142 , sensor information 310 , estimation condition information 320 , and estimation information 330 .

[0043] The one or more pieces of profile information 142 are different for each user. For example, the profile information is information that represents the characteristics of the user, and may be data on the user's lifestyle-related behavior (such as sleeping, eating, bathing, etc.), mood, activity level, and / or physical symptoms. For example, the profile information may include a behavior probability (the probability of performing a specific behavior at a specific time).

[0044] The operator device 160 is managed by a business that provides the service through the dialogue system 10, and displays candidates for messages to be spoken to the operator. The center device 100 and the operator device 160 are in a server-client relationship.

[0045] The center device 100 and the operator device 160 may be integrated into one hardware unit.

[0046] 2. Hardware configuration of the center device

[0047] FIG. 2 is a diagram showing the hardware configuration of the center device.

[0048] As shown in the figure, the center device 100 includes a CPU (Central Processing Unit) 11, a ROM (Read Only Memory) 12, a RAM (Random Access Memory) 13, an input / output interface 15, and a bus 14 connecting these together.

[0049] The CPU 11 accesses the RAM 13 and other memory as needed, and performs various arithmetic processing while controlling all of the blocks in the center device 100. The ROM 12 is a non-volatile memory that permanently stores firmware such as the OS, programs, and various evaluation indexes to be executed by the CPU 11. The RAM 13 is used as a working area for the CPU 11, and temporarily stores the OS, various applications currently being executed, and various data currently being processed.

[0050] The input / output interface 15 is connected to a display unit 16, an operation reception unit 17, a storage unit 18, a communication unit 19, and the like.

[0051] The display unit 16 is a display device that uses, for example, an LCD (Liquid Crystal Display), an OLED (Organic Electro Luminescence Display), a CRT (Cathode Ray Tube), or the like.

[0052] The operation reception unit 17 is, for example, a pointing device such as a mouse, a keyboard, a touch panel, or other input device.

[0053] The storage unit 18 is a non-volatile memory such as a hard disk drive (HDD), a flash memory (solid state drive (SSD)), or other solid-state memory. In addition to the OS, the storage unit 18 stores schedule information, various flags, evaluation indicators, and software for implementing each means of the center device 100. The storage unit 18 also stores applications and other programs and databases for exchanging messages with the terminal device 200, including generating a response message.

[0054] The communication unit 19 is, for example, a NIC (Network Interface Card) for Ethernet (registered trademark) or one of various modules for wireless communication such as a wireless LAN, and is responsible for communication processing with the terminal device 200.

[0055] The hardware configuration of the operator device 160 and the terminal device 200 is basically the same as that of the center device 100, but the terminal device 200 has a microphone M and a speaker S on the front surface as described above.

[0056] 3. Operational flow of the dialogue system

[0057] Figure 3 shows an overview of the dialogue system. Figure 4 shows the operation flow of the dialogue system. Figure 5 shows an example of the configuration of sensor information.

[0058] The sensor data receiving means 101 of the center device 100 receives the sensor data 313 and the sensor ID 312 output from the multiple sensors 300, associates the sensor data 313 and the sensor ID 312 with the property ID 311 of the property (an ID that identifies the living space in which the sensor 300 is installed), and updates the sensor information 310 in the storage means 140 (step S1). The sensor data receiving means 101 may receive the sensor data 313 and the sensor ID 312 directly from the multiple sensors 300, or may receive the sensor data 313 and the sensor ID 312 via the terminal device 200 or another device that serves as a hub (such as a smartphone, not shown). When the sensor data 313 and the sensor ID 312 are received at the hub, the sensor data 313 and the sensor ID 312 may be associated with the property ID 311 at the hub and transmitted to the center device 100.

[0059] Fig. 6 shows an example of the configuration of estimation condition information, and Fig. 7 shows an example of the configuration of estimation information.

[0060] The state estimation means 102 estimates the state of the user in the living space based on the sensor information 310 and updates the estimated information 330 (step S2). The state of the user may include the user's health state, the user's behavioral state, the state of the user's living space, etc. As shown in Fig. 7, the estimated information 330 describes a user ID 331, an estimated date and time 332, and an estimated state 333 in association with each other.

[0061] As a first example, the state estimation means 102 may estimate a state related to a user based on estimation condition information 320 as a rule. The estimation condition information 320 is information in which a state 321, an estimation condition 322, and an importance 323 are associated with each other. The importance 323 has a higher value as the state relates to an accident or health. The importance 323 is set in advance for each type of state.

[0062] If the state (or a combination thereof) estimated from the sensor data 313 or the sensor output data 313 in the sensor information 310 matches the estimation condition 322 in the estimation condition information 320, the state estimation means 102 estimates that the user is in that state 321.

[0063] As a second example, the state estimation means 102 may estimate the state of the user using AI. That is, the state estimation means 102 may estimate the state by inputting sensor data 313 to a state estimation AI that has been trained in advance. The state estimation AI is a machine learning model that receives sensor data 313 as input and outputs an estimated result of the user's state observed by a sensor. By inputting time-series data from a sensor that can observe behavior, the state estimation AI outputs a score for each user state class. The state corresponding to the class with the highest score (and a score exceeding a threshold) is estimated as the user's state. When it is considered that a user may have multiple states (for example, eating and watching TV may occur at the same time), multiple classes with the highest scores exceeding a threshold are estimated as the current state. Alternatively, a composite state may be defined as a new class.

[0064] The speech necessity determination means 107 determines whether speech is necessary based on the state 321 (state type) and importance 323 estimated by the state estimation means 102 (step S3). For example, the speech necessity determination means 107 may determine that speech is unnecessary for a state 321 with a low importance 323. The speech necessity determination means 107 may also score the importance 323 and determine the necessity based on a threshold. In this case, the speech necessity determination means 107 may determine whether to use automatic speech (automatically sending a message) or manual speech based on the estimated state 333 (type). Manual speech means that a message to be spoken is confirmed and corrected by an operator, and then sent in response to a message transmission instruction from the operator.

[0065] When it is determined that an utterance is required (step S3, Yes), the message generation means 103 generates a utterance message corresponding to the state estimated by the state estimation means 102 (step S4). The utterance message may be a standard message stored in advance corresponding to the state. Alternatively, it may be an individual message generated for each user based on the user's profile information. Furthermore, the message generation means 103 may generate the utterance message by having an operator input the utterance message using an input device such as a keyboard. In this case, the operator may create or modify the utterance message based on an automatically generated standard message or individual message.

[0066] The response sending means 104 sends a speech message to the terminal device 200 (step S5). When sending a speech message at a specified time, the response sending means 104 may store the speech message in a message queue together with the scheduled transmission time, and send the message based on the message queue. For example, the response sending means 104 may send a speech message about a condition detected during the nighttime hours the next morning. After sending the speech message, the response sending means 104 updates the message log (records the speech message). The response sending means 104 may determine whether to send the speech message automatically or manually by an operator based on the user's profile information.

[0067] The response transmitting means 104 may convert the spoken message of text data into a spoken message of voice data using a voice synthesis technique and transmit the converted message to the terminal device 200, or may transmit the spoken message of text data to the terminal device 200.

[0068] The response output means 203 of the terminal device 200 receives the speech message from the center device 100 and outputs the speech message from the speaker S at a predetermined timing (for example, immediately after receiving the message, when a human presence sensor detects a person around the terminal device 200). The response output means 203 may receive and output a speech message of voice data from the center device 100, or may receive a speech message of text data from the center device 100 and convert it into a speech message of voice data using voice synthesis technology and output it.

[0069] The input means 201 of the terminal device 200 inputs a response message from the user via the microphone M. The response message is a message in response to a spoken message by the user, and is voice data. The message sending means 202 of the terminal device 200 sends the received response message to the center device 100 via the network. The message sending means 202 may perform voice recognition on the voice data response message, convert it into a text data response message, and send it to the center device 100, or may send the voice data response message to the center device 100.

[0070] The message receiving means 105 of the center device 100 receives a user's response message to the spoken message from the terminal device 200 via the communication unit 19 (step S6, Yes). The message receiving means 105 may receive the response message in voice data and convert it into a response message in text data, or may receive a response message in text data converted from voice data by the terminal device 200. The message receiving means 105 updates the message log (records the response message) based on the received response message. If a predetermined threshold time has passed without receiving a response message after sending the spoken message (timeout), the message receiving means 105 may determine that there is no response and proceed to the next process.

[0071] Based on the response message from the terminal device 200, the determination means 106 determines whether or not a speech message is necessary or how often it is necessary when the state estimation means 102 estimates the state (step S7). Specifically, based on the response message, the determination means 106 determines the polarity indicating a favorable or negative (hostile) response to the speech message, and determines whether or not or how often a speech message is necessary based on the polarity. In the case of a negative response (including no response), the speech frequency may be changed to be lower (or no speech is necessary).

[0072] The determination means 106 determines the polarity of the response message, for example, by using a dictionary. First, the determination means 106 performs morphological analysis on the response message to extract one or more characteristic words and phrases. At this time, the characteristic words and phrases are extracted based on a dictionary 141 stored in advance in the storage means 140.

[0073] The dictionary 141 stores phrases (words), one or more states, and polarity values ​​in association with each other. Phrases frequently used in each state of each user may be registered in the dictionary 141 as individual characteristic phrases for each user. Words related to the state of the user may be registered in the dictionary 141 in advance as characteristic phrases, or may be registered as individual characteristic phrases based on the dialogue history with the user.

[0074] Next, the determination means 106 calculates a polarity value representing the magnitude (strength) of the user's negative (hostile) / friendly reaction to the speech message based on the response message, and determines the necessity or frequency based on the polarity value. The polarity value may be calculated based on the strength of the expression, the volume of the voice, etc. Specifically, the determination means 106 calculates the polarity of the response message based on the polarity value of each extracted feature phrase, and determines whether the reaction to the speech message is positive or negative (hostile). If the polarity of the response message is positive, it is determined to be a positive reaction, and if the polarity is negative, it is determined to be a negative (hostile) reaction. If the message is ignored, it is determined to be a negative reaction. The determination means 106 updates the average polarity value related to the state in the user's profile information 142 based on the calculated polarity value of the response message. Note that instead of the polarity value, only the polarity may be stored in association with the phrase (word). In this case, the determining means 106 skips updating the average polarity value, and determines whether or not a spoken message is necessary or how often it is necessary when the state is estimated by the state estimating means 102 based on the polarity.

[0075] FIG. 9 shows an example of the configuration of user profile information.

[0076] The user profile information 142 includes a user ID 341, a state 342, and an average polarity value 343, which are recorded in association with one another.

[0077] For example, the utterance message for the sleep-deprived state 342 is "It seems like you didn't sleep well last night. Are you sleep-deprived?" The user's response message to this is "Don't mind, it's none of my business!" In this case, the determination means 106 extracts "don't mind," "none of my business," and "none of my business" as characteristic phrases. The determination means 106 calculates the polarity value of the response message itself based on the sum of the polarity values ​​of the characteristic phrases, and determines whether the utterance message is favorable or unfavorable. The determination means 106 updates the average polarity value 343 for this state in the user's profile information 142 based on the calculated polarity value of the response message.

[0078] Alternatively, the determination means 106 may determine the user's state estimated by the state estimation means 102 using other determination methods, such as emotion estimation, without using a dictionary. The determination means 106 may use an AI that recognizes emotions from text or voice (audio). For example, the determination means 106 may perform an emotion analysis of the response message based on the text data or audio data of the response message, determine the polarity of the response message based on the emotion analysis results, and determine whether the response to the spoken message is favorable or unfavorable (hostile). In this case, the determination means 106 calculates the polarity value of the response message based on the score value output by the AI ​​and determines whether the response to the spoken message is favorable or unfavorable. The determination means 106 updates the average polarity value related to the state in the user's profile information 142 based on the determined polarity value of the response message.

[0079] In this way, the determination means 106 determines whether the reaction to the spoken message is favorable or negative (hostile) based on at least one of the polarity of the content (whether or not the spoken message contains a predetermined positive / negative term) and the polarity of the emotion (based on the emotion estimation result).

[0080] The degree of determination will now be explained. In a dictionary-based determination method, the determination means 106 determines the polarity value of the response message by summing the polarity values ​​of each feature phrase, while in an emotion estimation determination method, the determination means 106 determines the polarity value based on the score output by AI. This is not a limitation; the determination means 106 may also determine the absolute value of the polarity value based on the time interval between the output of the spoken message and the reception of the response message. For example, if the time interval is long, the determination means 106 may determine the response message as an "ambiguous response message" and calculate a small absolute value for the polarity value. Furthermore, the determination means 106 may also determine the response message as an "ambiguous response message" when the voice volume of the response message is low and calculate a small absolute value for the polarity value. Conversely, if the voice volume of the response message is high, the determination means 106 may determine the response message as a "response message with a clear intention" and calculate a large absolute value for the polarity value.

[0081] The determination means 106 changes the profile information 142 as necessary (step S8) based on the determination result by the determination means 106 (whether the reaction to the spoken message was favorable or negative (hostile)) (step S7). The profile information 142 is information that stores information based on the determination result of the determination means 106 for each state of the user.

[0082] Furthermore, the determination means 106 may feed back the updated profile information 142 to the state estimation means 102, and thereafter the state estimation means 102 may estimate the state based on the updated profile information 142 (step S2). That is, a determination history of the determination means 106 for each state estimated by the state estimation means 102 may be stored as the profile information 142, and the state estimation means 102 may estimate the state based on the determination history. The determination means 106 may personalize its determination to improve the accuracy of state estimation.

[0083] Furthermore, the determination means 106 may feed back the updated profile information 142 to the message generation means 103, and thereafter the message generation means 103 may generate an individual message as an utterance message based on the updated profile information 142 (step S4). That is, a determination history of the determination means 106 for each state estimated by the state estimation means 102 may be stored as the profile information 142, and the message generation means 103 may generate an utterance message based on the determination history. The message generation means 103 may personalize the utterance message to improve user convenience.

[0084] The determination means 106 may change the profile information based on the determination result and determine whether or not speech is required based on the profile information. The profile information may include information related to the determination result of the determination means 106, specifically, the average polarity value for each state over a predetermined period in the past, the number of positive or negative responses (the number of times each polarity was used), etc. The determination means 106 may determine whether or not speech is required based on, for example, the average polarity value of the estimated state (normalized to a value between -1 and 1). If the average polarity value of the estimated state is large, it may determine that a speech message should be generated and transmitted each time that state is estimated. Conversely, if a state with a small average polarity value (a value close to -1) is estimated, it may determine that generation and transmission of a speech message is unnecessary.

[0085] The determination means 106 may determine the speech frequency for each state of the user based on the profile information and determine whether or not a speech is necessary based on the speech frequency. For example, assume that the speech frequency for the "insufficient sleep" state is set to once every three days as the default. In this case, a speech message related to the "insufficient sleep" state is generated and transmitted once every three days. Here, if the average polarity value for the user's "insufficient sleep" state is large, the setting is changed so that a speech message related to the "insufficient sleep" state is generated and transmitted once per day. Similarly, if the average polarity value for the "forgot to hang out laundry" state is small, even if the default frequency is set to once per day, the setting is changed to once per week. Note that the necessity or frequency of a speech message when a predetermined state is estimated may be determined based on the polarity (polarity value) of the most recent (immediately preceding) response message related to the predetermined state, without using profile information.

[0086] As the profile information, information related to the determination result of the determination means 106 for each state of the user for each time period of a day may be stored. The determination means 106 may determine the necessity or frequency of a spoken message based on the estimated time of the state by the state estimation means 102 and the profile information. For example, the profile information may record an average polarity value for each time period of a day (e.g., every hour). In this case, when the state estimation means 102 estimates a state related to the user, the determination means 106 may determine which time period the estimated time falls in and determine the necessity or frequency of a spoken message based on the average polarity value for that time period. This makes it possible to increase the frequency of spoken messages related to a predetermined state for a user who tends to send positive (favorable) response messages in the morning time period. Note that the frequency may be changed based on the difference between the average polarity value for all time periods and the average polarity value for each time period.

[0087] The profile information may include an average polarity value for all states of the user and a polarity value for each state of the user. The determination means 106 may determine the necessity or frequency of a spoken message when the state estimation means 102 estimates the predetermined state based on the difference between the average polarity value for all states of the user and the polarity value for the predetermined state. The "average polarity value for all states" may be the average polarity value for all states of the user (average polarity value for all states). In this case, when the state estimation means 102 estimates the state related to the user, the determination means 106 may determine the necessity or frequency of a spoken message based on the difference between the average polarity value for the state and the average polarity value for all states. This makes it possible to determine the necessity or frequency of a spoken message related to each state based on whether the state significantly deviates from the user's general tendency to respond negatively / favorably (i.e., the average polarity value for all states).

[0088] The storage means 140 may store an importance level corresponding to each state and indicating the degree of importance of the state. The determination means 106 may determine the necessity or frequency of a speech message based on the importance level of the state estimated by the state estimation means 102. For example, the importance level of each state may be stored in advance in association with each other, and the determination may be made as to the necessity of generating and transmitting a speech message by scoring based on the average polarity value and importance level of the estimated state and determining the threshold value. Alternatively, the frequency of generating and transmitting a speech message may be calculated based on the average polarity value and importance level of the estimated state, and the speech message may be generated and transmitted based on the frequency.

[0089] The determination means 106 may determine whether or not a spoken message is necessary or how often it is necessary based on the security status of the user. That is, in addition to the sensor status, the determination means 106 may receive the security status of a security device installed in the user's home, and determine whether or not a spoken message is necessary or how often it is necessary based on the security status estimated by the status estimation means 102. For example, if the user sets security for a specific area partially while at home (for example, sets security for only the first floor) or sets security for sensors at entrances and exits, it can be said that the user is emotionally on guard, and therefore, by sending spoken messages as frequently as possible, communication that is sensitive to the user's anxieties can be realized.

[0090] The message generating unit 103 may determine the strength of the speech message based on the profile information. Specifically, the "strength of the speech message" may be the strength of the expression (e.g., whether or not the intent of the question is conveyed, its specificity, and the strength of conveying that it is a question), the volume of the voice, etc. When the average polarity value is large, the speech message is needed by the user, so it can be a strong speech message. Conversely, when the average polarity value is small, the speech message is likely to be avoided by the user, so it can be a "weak" speech message. For example, for a "lack of sleep" state, a strong speech message could be "You seemed to wake up several times yesterday. Did something happen?", i.e., a question format based on the sensor detection results. An intermediate level speech message could be "You didn't seem to sleep well last night. Are you sleepy?", i.e., a question that is not specific but conveys the recognition that the user did not sleep well. An example of a weak speech message could be "Good morning. Did you sleep well last night?", i.e., a question that could be interpreted as a greeting.

[0091] 4. Working Example

[0092] FIG. 8 shows an embodiment of the dialogue system.

[0093] The state estimation means 102 of the center device 100 estimates the state (lack of sleep) of the user based on the state estimated from the sensor data (detecting two wake-ups after going to sleep) and updates the estimated information 330 (step S2). When the speech necessity determination means 107 determines that speech is necessary based on the state (lack of sleep) (step S3, Yes), the message generation means 103 generates a speech message according to the state (lack of sleep), such as "It seems you didn't sleep well last night. Are you feeling unwell?" (step S4). The response transmission means 104 transmits the speech message to the terminal device 200 (step S5).

[0094] The message receiving means 105 receives the user's response message "That's right, thanks for worrying about me!" from the terminal device 200 (step S6, Yes). The determining means 106 determines the utterance frequency to be a relatively high frequency of once a day based on the response message "That's right, thanks for worrying about me!" (step S7), and changes the profile information 142 (step S8).

[0095] On the other hand, if the message receiving means 105 determines a timeout (ignores) without receiving a user response message from the terminal device 200, the determining means 106 determines the speech frequency to be a medium frequency of once every two days (step S7) and changes the profile information 142 (step S8).

[0096] Meanwhile, the message receiving means 105 receives the user's response message "None of your business!" from the terminal device 200 (step S6, Yes). The determining means 106 determines the utterance frequency to be the lowest, "No utterance for this state," based on the response message "None of your business!" (step S7), and changes the profile information 142 (step S8).

[0097] 5. Conclusion

[0098] When a user's state is estimated, some users feel happy to be spoken to about that state, while other users feel annoyed that they do not want to be spoken to much in that state. Therefore, according to this embodiment, the center device 100 estimates the user's state (user state, user behavior, home state) based on the detection results from the sensors 300 installed in the user's home, and the terminal device 200 outputs a speech message according to the estimation result. The center device 100 determines whether or not (or how often) the user needs to speak about the state based on the user's response message to the speech message. This makes it possible to easily understand through natural communication between the user and the terminal device 200 whether or not the detected state is one in which the user wants to be spoken to.

[0099] Although the embodiments and modified examples of the present technology have been described above, the present technology is not limited to the above-described embodiments, and it goes without saying that various modifications can be made within the scope of the gist of the present technology.

[0100] In the above-described embodiment and each modified example, the frequency of speech or the necessity of speech for a user is determined based on the user's profile information 142. However, without being limited to this, the frequency of speech or the necessity of speech for a predetermined user may be determined using the profile information 142 of other users. For example, if the polarity values ​​of response messages to speech messages issued when a predetermined state is determined in the profile information 142 of multiple users are generally small, it may be determined that speech messages for that state are generally not popular, and the default speech frequency of speech messages for that state may be changed to be lower, or the expression or strength of the speech messages may be changed to be weaker (gentlerer).

[0101] In the above-described embodiment and each modification, the center device 100 and the terminal device 200 are different devices. However, this is not limiting. The center device 100 and the terminal device 200 may be configured as the same device (e.g., an interactive device) and installed in the user's living space. That is, the various means of the center device 100 and the storage means 140 may all be provided within the terminal device 200. In this case, the operator device 160 is assumed to be communicably connected to the interactive device via a network. Furthermore, the interactive device may not be an integrated device, but may be configured as separate devices, such as a main device and a robot. For example, the main device may have the same functions as the center device 100, and the robot may have the same functions as the terminal device 200, and each may be installed in the user's living space. In addition, in the above-described embodiment and each modification, a speech message transmitted from the center device 100 is received by the terminal device 200, and the response output means 203 of the terminal device 200 outputs the received speech message to the user via a speaker S. However, without being limited to this, the terminal device 200 may store the audio data (text data) of the spoken message in association with an identifier (spoken message ID) for identifying the spoken message, and by receiving the spoken message ID of the spoken message to be spoken from the center device 100, output the spoken message associated with the spoken message ID from the speaker S.

[0102] A dialogue system according to one embodiment of the present invention can contribute to solving social issues such as loneliness and isolation among the elderly, extending healthy lifespan, improving the quality of life (QoL) of the elderly, and a declining labor force. Furthermore, the dialogue system according to one embodiment of the present invention can contribute to achieving Goal 3 of the Sustainable Development Goals (SDGs) adopted by the United Nations, "Ensure good health and well-being for all." [Explanation of symbols]

[0103] 10 Dialogue Systems 100 Center Device 101 Sensor data receiving means 102 State estimation means 103 Message Generation Method 104 Response sending means 105 Message Receiving Method 106 Judgment means 107 Means for determining whether or not to speak 140 Memory means 141 Dictionaries 142 Profile Information 160 Operator Device 200 Terminal Device 201 Input Method 202 Message sending method 203 Response output means 300 sensors 310 Sensor Information 320 Estimated Condition Information 330 Estimated Information

Claims

1. one or more sensors in the user's living space; an interactive device; A dialogue system comprising: The dialogue device a sensor data receiving means for receiving sensor data output from the sensor; a state estimation means for estimating a state related to a user in the living space based on the sensor data; a message generating means for generating a speech message according to the state estimated by the state estimating means; an output means for audibly outputting the generated speech message; an input means for inputting a user's response message by voice in response to the spoken message outputted by voice; a determination means for determining whether or not the speech message to be output by voice in response to the state estimated by the state estimation means is necessary or how often the speech message is to be output by voice, based on the response message; have Dialogue system.

2. 2. The dialogue system according to claim 1, The determination means determines the polarity of the user's reaction to the speech message, whether negative or positive, based on the response message, and determines the necessity or frequency based on the polarity. Dialogue system.

3. Dialogue system according to claim 2, The determining means determines a polarity value representing the magnitude of the user's negative / positive reaction to the speech message based on the response message, and determines the necessity or the frequency based on the polarity value. Dialogue system.

4. Dialogue system according to any one of claims 1 to 3, the dialogue device further comprises a storage means for storing profile information representing characteristics of the user with respect to each state; The determining means determines, based on the response message, whether the user has a negative or positive reaction to the speech message, and changes the profile information to determine whether or not the speech message is necessary or how often it is necessary to send the speech message based on the profile information. Dialogue system.

5. Dialogue system according to claim 4, As the profile information, information relating to the reaction for each time period of a day for each state of the user is stored; The determining means determines whether or not the message is necessary or how often the message is to be uttered based on the profile information of the time period corresponding to the estimated time of the state estimated by the state estimating means. Dialogue system.

6. Dialogue system according to claim 4, As the profile information, an average polarity value of all states of the user and a polarity value relating to each state of the user are stored; The determining means determines whether or not the spoken message is necessary or how often the spoken message is necessary when the state estimating means estimates the predetermined state based on the difference between the average polarity value and the polarity value of the predetermined state. Dialogue system.

7. Dialogue system according to any one of claims 1 to 3, an importance level indicating the degree of importance of each of the states is stored in association with each of the states; The determining means determines whether or not the message needs to be uttered or the frequency of the message based on the importance of the state estimated by the state estimating means. Dialogue system.

8. Dialogue system according to any one of claims 1 to 3, The determining means determines whether or not the spoken message is necessary or how often the spoken message is to be sent based on the security status of the user. Dialogue system.

9. Dialogue system according to claim 4, The message generating means determines the strength of the speech message based on the profile information. Dialogue system.

10. a sensor data receiving means for receiving sensor data output from one or more sensors in the user's living space; a state estimation means for estimating a state related to a user in the living space based on the sensor data; a message generating means for generating a speech message according to the state estimated by the state estimating means; a response sending means for sending the speech message to a terminal device which exchanges messages with a user in a two-way manner; a message receiving means for receiving a response message from the user in response to the spoken message from the terminal device; a determining means for determining whether or not the spoken message is necessary or how often the spoken message is to be sent when the state is estimated by the state estimating means based on the response message; Equipped with Center device.

11. The computer of the center device receiving sensor data output from one or more sensors in a user's living space; estimating a state of a user in the living space based on the sensor data; generating a speech message according to the estimated state; transmitting the spoken message to a terminal device that exchanges messages with a user; receiving a response message from the user to the spoken message from the terminal device; determining whether or not the speech message is necessary or how often the speech message is necessary when the state is estimated based on the response message; Run How to interact.

Citation Information

Patent Citations

  • Voice guide system

    JP2007286376A

  • Interactive business support system and interactive business support method

    JP2021157419A