Dialogue system, dialogue method and center device

The dialogue system addresses labor shortages and quality issues by using sensors and user feedback to personalize and improve interaction quality in communication systems for alleviating loneliness.

JP2025144467APending Publication Date: 2025-10-02SECOM CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024044266
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-19
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Existing communication systems for alleviating loneliness, such as those involving human operators and automated robots, face challenges with labor shortages and varying response quality, making it difficult to provide tailored and high-quality interactions.

Method used

A dialogue system utilizing sensors in the user's living space, a terminal device, and a center device that estimates user states, generates and evaluates responses, and adjusts based on user feedback to improve interaction quality.

Benefits of technology

The system provides personalized and accurate dialogue tailored to individual users, enhancing user interaction quality and reducing operational burdens by refining response accuracy through user feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025144467000001_ABST
    Figure 2025144467000001_ABST
Patent Text Reader

Abstract

To provide a dialog system that can perform dialogues according to individual users.SOLUTION: A dialogue system is the dialogue system provided with: one or a plurality of sensors present in a living space of a user; and a dialogue device. The dialogue device has: sensor data reception means for receiving sensor data to be output from the sensor; state estimation means for estimating a state on the user inside the living space on the basis of the sensor data; message generation means for generating an utterance message in accordance with the state the state estimation means estimates; output means for vocally outputting the generated utterance message; input means for vocally inputting a reply message of the user to the vocally outputted utterance message; and determination means for evaluating the state the state estimation means estimates based on the reply message to output an evaluation result.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a dialogue system and dialogue method using a terminal device having a voice input / output function, and a center device. [Background technology]

[0002] Loneliness is a factor that leads to serious health risks such as dementia, and the increase in single-person households (especially elderly people living alone) coupled with the recent increased risk of infectious diseases makes loneliness more likely to occur, making alleviating loneliness an important social issue. Casual conversation and other everyday interactions are thought to be effective in alleviating loneliness.

[0003] In recent years, communication services have been proposed that aim to alleviate feelings of loneliness by installing devices with voice input and output functions (e.g., interactive robots) in the homes of elderly people, especially those living alone, and engaging in everyday conversations such as casual voice chat through these devices.

[0004] In this system, when a user speaks to a device at home, the content is converted into a string of characters using voice recognition technology and sent as a message to the management center of the service provider. At the management center, an operator who notices the receipt of the message types a reply into the center device, which is then sent back to the user. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Japanese Patent Publication No. 2021-157419 [Patent Document 2] Japanese Patent Application Laid-Open No. 2007-286376 Summary of the Invention [Problem to be solved by the invention]

[0006] There are currently both human-assisted interactive services and automated response services using communication robots. However, labor shortages due to the declining labor force are also a social issue, and there are limits to how much human operators can provide. When an operator types a reply to a user, they must understand the message and assess the user's condition based on their memory and experience of previous messages exchanged with the user before creating a reply. However, the more users a single operator handles, the more difficult this becomes. Furthermore, the operational burden increases, as does the difficulty of handing over tasks when an operator is replaced. On the other hand, communication robots that do not require human intervention may be disliked by users due to the low quality of their interactions (e.g., inappropriate responses, formulaic responses, etc.).

[0007] Patent Document 1 proposes a UI that supports the operator's responses by displaying candidate responses for each category (affirmative, counterargument, topic change) along with an evaluation value. However, when an operator responds to a message from a user, it is a heavy burden for the operator to think of a response message for each message, and there is a risk that the quality of the response messages will vary depending on the operator. Although Patent Document 1 displays candidate responses to the operator, the operator must check each candidate and decide on a response message, which still places a burden on the operator.

[0008] Patent Document 2 proposes a technology that switches between automatic responses and responses via an operator by requesting a remote support device for response assistance when a robot cannot decide on a response message. However, in conversations where it is not easy for a robot to determine the appropriateness of a response message, such as everyday conversations such as casual conversations, switching to an operator response each time may place an excessive burden on the operator.

[0009] In view of the above circumstances, an object of the present invention is to provide a dialogue system that is capable of conducting dialogue tailored to individual users. [Means for solving the problem]

[0010] A dialogue system according to one aspect of the present invention comprises: one or more sensors in the user's living space; An interactive device; A dialogue system comprising: The dialogue device a sensor data receiving means for receiving sensor data output from the sensor; a state estimation means for estimating a state related to a user in the living space based on the sensor data; a message generating means for generating a speech message according to the state estimated by the state estimating means; an output means for audibly outputting the generated speech message; an input means for inputting a user's response message by voice in response to the spoken message outputted by voice; a determination means for evaluating the state estimated by the state estimation means based on the response message and outputting an evaluation result; It has.

[0011] When estimating a user's state using sensors, it is operationally difficult to confirm whether the estimation result is correct. Therefore, according to this embodiment, a center device estimates the user's state (user state, user behavior, and home state) based on detection results from sensors installed in the user's home, and a terminal device outputs a speech message corresponding to the estimation result. The center device determines whether to change the estimation result based on the user's response message to the speech message, i.e., determines whether the estimation result is correct and changes the reliability of the estimation result. This allows the user to easily understand the accuracy of the estimation result of the user's state through natural communication between the user and the terminal device. Furthermore, by being able to ask questions in a timely manner while the information is still fresh in the user's mind, higher-quality information, such as factors related to the user's state, can be obtained compared to periodic questions, such as those asked during weekly visits or medical examinations.

[0012] The determining means may determine a polarity indicating a positive / negative reaction of the user to the speech message based on the response message, and evaluate the state based on the polarity.

[0013] As a result, if the polarity of the response message is positive, it is determined that the response is a positive reaction and the estimation result is correct, and if the polarity is negative, it is determined that the response is a negative reaction and the estimation result is incorrect.

[0014] The message generating means may determine whether to generate a speech message following the voice utterance according to the evaluation result.

[0015] This allows the dialogue with the user to continue with appropriate content.

[0016] the dialogue device further comprises a storage means for storing profile information representing characteristics of the user with respect to each state; the state estimation means estimates the state based on the profile information; The determining means may change the profile information based on the evaluation result.

[0017] This allows the dialogue with the user to continue with appropriate content.

[0018] the dialogue device further comprises a storage means for storing profile information representing characteristics of the user with respect to each state; The message generating means may determine whether to automatically transmit the speech message based on the profile information or to manually transmit the speech message based on an instruction from an operator. The storage means may store, as the profile information, a determination history of the determination means for each of the states estimated by the state estimation means.

[0019] This allows updated profile information to be fed back, and thereafter, the state can be estimated based on the updated profile information, thereby personalizing the judgment and improving the accuracy of state estimation.

[0020] the dialogue device further comprises a storage means for storing profile information representing characteristics of the user with respect to each state; The message generating means may generate the speech message based on the profile information.

[0021] This allows updated profile information to be fed back, and thereafter, individual messages can be generated as spoken messages based on the updated profile information, thereby personalizing spoken messages and improving user convenience.

[0022] the determining means identifies a cause of the state estimated by the state estimating means based on the response message; the storage means stores a factor history of the factor identified by the determination means as the profile information; The message generating means may generate the speech message based on the factor history.

[0023] This allows the updated profile information to be fed back, and thereafter, the dialogue with the user can be continued with appropriate content based on the updated profile information.

[0024] A center device according to one aspect of the present invention comprises: a sensor data receiving means for receiving sensor data output from one or more sensors in the user's living space; a state estimation means for estimating a state related to a user in the living space based on the sensor data; a message generating means for generating a speech message according to the state estimated by the state estimating means; a response sending means for sending the speech message to a terminal device which exchanges messages with a user in a two-way manner; a message receiving means for receiving a response message from the user in response to the spoken message from the terminal device; a determination means for evaluating the state estimated by the state estimation means based on the response message and outputting an evaluation result; a sensor data receiving means for receiving sensor data output from one or more sensors in the user's living space; a state estimation means for estimating a state related to a user in the living space based on the sensor data; a message generating means for generating a speech message according to the state estimated by the state estimating means; a response sending means for sending the speech message to a terminal device which exchanges messages with a user in a two-way manner; a message receiving means for receiving a response message from the user in response to the spoken message from the terminal device; a determination means for evaluating the state estimated by the state estimation means based on the response message and outputting an evaluation result; It is equipped with:

[0025] A dialogue method according to one aspect of the present invention includes: The computer of the center device receiving sensor data output from one or more sensors in the user's living space; estimating a state of a user in the living space based on the sensor data; generating a speech message according to the estimated state; transmitting the spoken message to a terminal device that exchanges messages with a user; receiving a response message from the user to the spoken message from the terminal device; evaluating the estimated state based on the response message and outputting an evaluation result; Execute. [Effects of the Invention]

[0026] According to the present invention, it is possible to provide a dialogue system that is capable of conducting dialogue tailored to individual users.

[0027] The effects described here are not necessarily limited to those described above, and may be any of the effects described in the present invention. [Brief explanation of the drawings]

[0028] [Figure 1] 1 is a diagram showing a configuration of a dialogue system according to an embodiment of the present invention. [Figure 2] FIG. 2 is a diagram illustrating a hardware configuration of a center device. [Figure 3] An overview of the dialogue system is shown below. [Figure 4] The operation flow of the dialogue system is shown below. [Figure 5] 10 shows an example of the configuration of sensor information. [Figure 6] 10 shows an example of the configuration of estimation condition information. [Figure 7] 10 shows an example of the configuration of estimation information. [Figure 8] 1 shows an embodiment of a dialogue system. DETAILED DESCRIPTION OF THE INVENTION

[0029] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.

[0030] 1. Dialogue system configuration

[0031] FIG. 1 is a diagram showing the configuration of a dialogue system according to this embodiment.

[0032] The dialogue system 10 realizes a communication service in which everyday conversations such as casual chatting are mainly conducted. The dialogue system 10 is difficult to standardize responses, anticipates a wide variety of responses, and realizes everyday conversations such as continuous casual chatting with multiple users. Unlike automatic response systems for product guidance, contract procedures, etc., which can be implemented without the human intervention of an operator, the dialogue system 10 requires the human intervention of an operator to ensure the quality of the dialogue, as emotions such as empathy and intentions are important.

[0033] The dialogue system 10 is composed of a center device 100 as a server device managed by a business operator that provides services through the dialogue system 10, an operator device 160 connected to the center device 100, multiple terminal devices 200 connected to a network so as to be able to communicate with the center device 100, and one or more sensors 300 in the user's living space.

[0034] The terminal device 200 is installed in the home of a user (e.g., an elderly person) of this system. The terminal device 200 has an input means 201, a message sending means 202, and a response output means 203. The input means 201 inputs a response message from the user via a microphone M. The message sending means 202 sends the response message input via the input means 201 to the center device 100 via the network. The response output means 203 receives a spoken message from the center device 100 via the network, and outputs the message as voice from a speaker S to the user.

[0035] The terminal device 200 may have at least the above components, but may also be an interactive robot with a small doll-like appearance that makes the user feel like they are having a conversation with a person (especially everyday conversation such as casual chatting).

[0036] The one or more sensors 300 in the user's living space may be, for example, a door sensor, a PIR (Passive Infrared Ray) motion sensor, a smart tap, a CO2 sensor, an environmental sensor, or a vital sensor mounted on a wearable device (such as a smart watch or healthcare device). A smart tap (smart plug) is a plug-type device equipped with a wireless network communication function, which is connected between an outlet and the cable of a home appliance, and connects the connected home appliance to the IoT. The living space refers to a space inside a house, a building, a premises, or a room.

[0037] The center device 100 is managed by a business operator that provides services using the dialogue system 10, and includes at least a sensor data receiving means 101, a state estimating means 102, a message generating means 103, a response transmitting means 104, a message receiving means 105, a determining means 106, and a speech necessity determining means 107. It also includes a control means for controlling these means. These are realized by well-known hardware (so-called server computers or personal computers) or software stored in a storage means 140, as appropriate.

[0038] The storage means 140 stores a dictionary 141 , one or more pieces of profile information 142 , sensor information 310 , estimation condition information 320 , and estimation information 330 .

[0039] The one or more pieces of profile information 142 are different for each user. For example, the profile information is information that represents the characteristics of the user, and may be data on lifestyle-related behavior (such as sleeping, eating, bathing, etc.), mood, activity level, and / or physical symptoms. For example, the profile information may include a behavior probability (the probability of performing a specific behavior at a specific time).

[0040] The operator device 160 is managed by a business that provides the service through the dialogue system 10, and displays candidates for messages to be spoken to the operator. The center device 100 and the operator device 160 are in a server-client relationship.

[0041] The center device 100 and the operator device 160 may be integrated into one hardware unit.

[0042] 2. Hardware configuration of the center device

[0043] FIG. 2 is a diagram showing the hardware configuration of the center device.

[0044] As shown in the figure, the center device 100 includes a CPU (Central Processing Unit) 11, a ROM (Read Only Memory) 12, a RAM (Random Access Memory) 13, an input / output interface 15, and a bus 14 connecting these together.

[0045] The CPU 11 accesses the RAM 13 and other memory as needed, and performs various arithmetic processing while controlling all of the blocks in the center device 100. The ROM 12 is a non-volatile memory that permanently stores firmware such as the OS, programs, and various evaluation indexes to be executed by the CPU 11. The RAM 13 is used as a working area for the CPU 11, and temporarily stores the OS, various applications currently being executed, and various data currently being processed.

[0046] The input / output interface 15 is connected to a display unit 16, an operation reception unit 17, a storage unit 18, a communication unit 19, and the like.

[0047] The display unit 16 is a display device that uses, for example, an LCD (Liquid Crystal Display), an OLED (Organic Electro Luminescence Display), a CRT (Cathode Ray Tube), or the like.

[0048] The operation reception unit 17 is, for example, a pointing device such as a mouse, a keyboard, a touch panel, or other input device.

[0049] The storage unit 18 is a non-volatile memory such as a hard disk drive (HDD), a flash memory (solid state drive (SSD)), or other solid-state memory. In addition to the OS, the storage unit 18 stores schedule information, various flags, evaluation indicators, and software for implementing each means of the center device 100. The storage unit 18 also stores applications and other programs and databases for exchanging messages with the terminal device 200, including generating a response message.

[0050] The communication unit 19 is, for example, a NIC (Network Interface Card) for Ethernet (registered trademark) or various modules for wireless communication such as wireless LAN, and is responsible for communication processing with the terminal device 200.

[0051] The hardware configuration of the operator device 160 and the terminal device 200 is basically the same as that of the center device 100, but the terminal device 200 has a microphone M and a speaker S on the front surface as described above.

[0052] 3. Operational flow of the dialogue system

[0053] Figure 3 shows an overview of the dialogue system. Figure 4 shows the operation flow of the dialogue system. Figure 5 shows an example of the configuration of sensor information.

[0054] The sensor data receiving means 101 of the center device 100 receives the sensor data 313 and the sensor ID 312 output from the multiple sensors 300, associates the sensor data 313 and the sensor ID 312 with the property ID 311 of the property (an ID that identifies the living space in which the sensor 300 is installed), and updates the sensor information 310 in the storage means 140 (step S1). The sensor data receiving means 101 may receive the sensor data 313 and the sensor ID 312 directly from the multiple sensors 300, or may receive the sensor data 313 and the sensor ID 312 via the terminal device 200 or another device that serves as a hub (such as a smartphone, not shown). When the sensor data 313 and the sensor ID 312 are received at the hub, the sensor data 313 and the sensor ID 312 may be associated with the property ID 311 at the hub and transmitted to the center device 100.

[0055] Fig. 6 shows an example of the configuration of estimation condition information, and Fig. 7 shows an example of the configuration of estimation information.

[0056] The state estimation means 102 estimates the state of the user in the living space based on the sensor information 310 and updates the estimated information 330 (step S2). The state of the user may include the user's health state, the user's behavioral state, the state of the user's living space, etc. As shown in FIG. 7A, the estimated information 330 describes a user ID 331, an estimated date and time 332, and an estimated state 333 (including a state 334 and a reliability 335) in association with each other.

[0057] As a first example, the state estimation means 102 may estimate a state related to a user based on estimation condition information 320 as a rule. The estimation condition information 320 is information in which a state 321, an estimation condition 322, and an importance 323 are associated with each other. The importance 323 has a higher value as the state relates to an accident or health. The importance 323 is set in advance for each type of state.

[0058] If a state (or a combination thereof) estimated from the sensor data 313 or the sensor output data 313 of the sensor information 310 matches an estimation condition 322 of the estimation condition information 320, the state estimation means 102 estimates that the user is in that state 321. The state estimation means 102 may calculate the reliability of the state.

[0059] For example, for the "microwave left off" state 321A, the condition 322A is "when it is detected that the kitchen is unoccupied for a predetermined period of time after the microwave ignition is detected (a sudden rise in the range hood temperature) until the microwave is turned off." The longer the "predetermined period" is, the higher the reliability (0.7 for 5 minutes, 0.9 for 15 minutes, etc.).

[0060] Also, in the case of the "insufficient sleep" state 321B, if the number of times room movement is detected between the time sleep is detected and the time of wake-up is "two", the reliability is low, but if it is "three or more times" (condition 322B), the reliability is high.

[0061] Similarly, in the case of the "laundry forgotten to hang out" state 321C, the laundry will be left in the "forgot to hang out" state (reliability: low) for about 30 minutes after the washing machine has stopped, but if it is left unattended for an hour (condition 322C), the laundry will be left in the "forgot to hang out" state (reliability: high).

[0062] As a second example, the state estimation means 102 may estimate the state of the user using AI. That is, the state estimation means 102 may estimate the state by inputting sensor data 313 to a state estimation AI that has been trained in advance. The state estimation AI is a machine learning model that receives sensor data 313 as input and outputs an estimated result of the user's state observed by a sensor. By inputting time-series data from a sensor that can observe behavior, the state estimation AI outputs a score for each user state class. The state corresponding to the class with the highest score (and a score exceeding a threshold) is estimated as the user's state. When it is considered that a user may have multiple states (for example, eating and watching TV may occur at the same time), multiple classes with the highest scores exceeding a threshold are estimated as the current state. Alternatively, a composite state may be defined as a new class.

[0063] The speech necessity determination means 107 determines whether speech is necessary based on the states 321 and 334 (state type), importance 323, and reliability 335 estimated by the state estimation means 102 (step S3). For example, the speech necessity determination means 107 may determine that speech is unnecessary for the state 321 with low importance 323 or the state 334 with low reliability 335. The speech necessity determination means 107 may also score both the importance 323 and the reliability 335 to determine whether speech is necessary based on a threshold. In this case, the speech necessity determination means 107 may determine whether automatic speech (automatically sending a message) or manual speech is to be performed based on at least one of the estimated state 334 (type), importance 323, and its reliability 335. Manual speech means that a message to be spoken is confirmed and corrected by an operator, and then sent in response to a message transmission instruction from the operator.

[0064] When it is determined that an utterance is required (step S3, Yes), the message generation means 103 generates a utterance message corresponding to the state estimated by the state estimation means 102 (step S4). The utterance message may be a standard message stored in advance corresponding to the state. Alternatively, it may be an individual message generated for each user based on the user's profile information. Furthermore, the message generation means 103 may generate the utterance message by having an operator input the utterance message using an input device such as a keyboard. In this case, the operator may create or modify the utterance message based on an automatically generated standard message or individual message.

[0065] The response sending means 104 sends a speech message to the terminal device 200 (step S5). When sending a speech message at a specified time, the response sending means 104 may store the speech message in a message queue together with the scheduled transmission time, and send the message based on the message queue. For example, the response sending means 104 may send a speech message about a condition detected during the nighttime hours the next morning. After sending the speech message, the response sending means 104 updates the message log (records the speech message). The response sending means 104 may determine whether to send the speech message automatically or manually by an operator based on the user's profile information.

[0066] The response transmitting means 104 may convert the spoken message of text data into a spoken message of voice data using a voice synthesis technique and transmit the converted message to the terminal device 200, or may transmit the spoken message of text data to the terminal device 200.

[0067] The response output means 203 of the terminal device 200 receives the speech message from the center device 100 and outputs the speech message from the speaker S at a predetermined timing (for example, immediately after receiving the message, when a human presence sensor detects a person around the terminal device 200). The response output means 203 may receive and output a speech message of voice data from the center device 100, or may receive a speech message of text data from the center device 100 and convert it into a speech message of voice data using voice synthesis technology and output it.

[0068] The input means 201 of the terminal device 200 inputs a response message from the user via the microphone M. The response message is a message in response to a spoken message by the user, and is voice data. The message sending means 202 of the terminal device 200 sends the received response message to the center device 100 via the network. The message sending means 202 may perform voice recognition on the voice data response message, convert it into a text data response message, and send it to the center device 100, or may send the voice data response message to the center device 100.

[0069] The message receiving means 105 of the center device 100 receives a user's response message to the spoken message from the terminal device 200 via the communication unit 19 (step S6, Yes). The message receiving means 105 may receive the response message in voice data and convert it into a response message in text data, or may receive a response message in text data converted from voice data by the terminal device 200. The message receiving means 105 updates the message log (records the response message) based on the received response message. If a predetermined threshold time has passed without receiving a response message after sending the spoken message (timeout), the message receiving means 105 may determine that there is no response and proceed to the next process.

[0070] The determination means 106 evaluates the user's state estimated by the state estimation means 102 based on the response message from the terminal device 200 and outputs the evaluation result. For example, the determination means 106 may evaluate the accuracy of the estimation result or determine whether the reliability of the estimation result needs to be changed. If the determination means 106 determines that a state change is necessary, it changes the state estimated by the state estimation means 102 (step S7). The determination means 106 determines the accuracy of the estimation result (step S2) by the state estimation means 102 based on the response message. Specifically, the determination means 106 determines the polarity indicating the user's positive / negative reaction to the spoken message based on the response message, and evaluates the state based on the polarity. The determination means 106 updates the estimation information 330 based on the evaluation result, as shown in (B) of FIG. 7. If there is no response, the determination means 106 does not determine the estimation result and does not need to change the estimated state and reliability.

[0071] The determination means 106, for example, uses a dictionary to determine the user's state estimated by the state estimation means 102. First, the determination means 106 performs morphological analysis on the response message to extract one or more characteristic words and phrases. At this time, the characteristic words and phrases are extracted based on a dictionary 141 stored in advance in the storage means 140.

[0072] The dictionary 141 stores phrases (words), one or more states, and polarity values ​​in association with each other. Phrases frequently used in each state of each user may be registered in the dictionary 141 as individual characteristic phrases for each user. Words related to the state of the user may be registered in the dictionary 141 in advance as characteristic phrases, or may be registered as individual characteristic phrases based on the dialogue history with the user.

[0073] Next, the determination means 106 calculates the polarity of the response message based on the polarity value of each extracted feature phrase, and determines whether the estimation result is correct. If the polarity of the response message is positive, the determination means 106 determines that the estimation result is correct as a positive reaction, and if the polarity is negative, the determination means 106 determines that the estimation result is incorrect as a negative reaction.

[0074] For example, the spoken message is "You didn't seem to sleep well last night. Are you sleep-deprived?" The user's response message to this is "I see, I've been having a hard time falling asleep lately..." In this case, the determination means 106 extracts "I see" and "I can't fall asleep" as characteristic phrases. Since each characteristic phrase has a positive polarity (affirmative response) to the "sleep-deprived" state, the determination means 106 determines that the estimation result is correct. As shown in (B) of FIG. 7, the determination means 106 changes the reliability from 0.8 to 1.

[0075] When the polarity values ​​of the characteristic phrases are a mixture of positive and negative, the determination means 106 may determine the polarity of the response message based on the polarity value calculated by summing up the polarity values. For example, even if the response message has a positive polarity, if the sum polarity value is small, the response message may be considered an "ambiguous response message (positive)" and the increase in reliability may be reduced. In other words, the absolute value of the polarity value indicates the certainty of whether the response message is positive or negative.

[0076] Alternatively, the determination means 106 may determine the user's state estimated by the state estimation means 102 using other determination methods, such as emotion estimation, without using a dictionary. The determination means 106 may use an AI that recognizes emotions from text or voice (audio). For example, the determination means 106 may perform emotion analysis on the response message based on the text data or audio data of the response message, determine the polarity of the response message based on the emotion analysis result, and determine whether the estimation result is correct. In this case, the determination means 106 may calculate a polarity value of the response message based on the score value output by the AI, and determine an increase in reliability based on the polarity value.

[0077] In this way, the determination means 106 may determine the estimation result based on at least one of the polarity of the content (whether or not a predetermined positive / negative term corresponding to the spoken message is included) and the polarity of the emotion (based on the emotion estimation result).

[0078] The degree of determination will now be explained. In a dictionary-based determination method, the determination means 106 determines the polarity value of the response message by summing the polarity values ​​of each feature phrase. In a emotion estimation determination method, the determination means 106 determines the polarity value based on the score output by AI. This is not a limitation. The determination means 106 may also determine the absolute value of the polarity value based on the time interval between the output of the spoken message and the reception of the response message. For example, if the time interval is long, the determination means 106 may determine the response message as an "ambiguous response message" and calculate a small absolute value for the polarity value. In this case, the increase in reliability will be small. Furthermore, the determination means 106 may also determine the response message as an "ambiguous response message" when the voice volume is low and calculate a small absolute value for the polarity value. Conversely, if the voice volume of the response message is high, the determination means 106 may determine the response message as a "response message with a clear intention" and calculate a large absolute value for the polarity value.

[0079] The determination means 106 changes the profile information 142 as necessary (step S8) based on the evaluation result of the user's state by the determination means 106 (step S7). The profile information 142 is information that stores information based on the evaluation result of the determination means 106 for each state of the user.

[0080] Furthermore, the determination means 106 may feed back the updated profile information 142 to the state estimation means 102, and thereafter the state estimation means 102 may estimate the state based on the updated profile information 142 (step S2). That is, a determination history of the determination means 106 for each state estimated by the state estimation means 102 may be stored as the profile information 142, and the state estimation means 102 may estimate the state based on the determination history. The determination means 106 may personalize its determination to improve the accuracy of state estimation.

[0081] As a first example, in the case of rule-based state estimation, the profile information 142 may store a determination history by the determination means 106 for each state estimated by the state estimation means. The error estimation rate for each state may be calculated and stored as the determination history. For example, a 78% error estimation rate for the “insufficient sleep” state, a 40% error estimation rate for the “forgot to turn off the microwave” state, etc. may be recorded as the determination history. Then, for a state with a large error estimation rate (“insufficient sleep”), the estimation conditions may be changed to make it less likely that the state will be estimated (to make the estimation conditions stricter). For example, while the basic estimation condition is that the “insufficient sleep” state is estimated when “movements from one room to another are detected two or more times between the time of sleep detection and the time of wake-up,” the estimation condition for user A may be changed to “three or more times” because the error estimation rate for the “insufficient sleep” state of user A is high. Note that the error estimation rate for each time period may be stored as the determination history. In this case, stricter estimation conditions may be applied only to time periods with a high error estimation rate.

[0082] As a second example, in the case of state estimation by AI (machine learning), the profile information 142 may be stored as a determination history in which the correct state determined by the determination means 106 is associated (annotated) with the sensor data from which the state was estimated. The accuracy of state estimation may be improved by re-learning (additional learning) the AI ​​using the determination history as learning data. Alternatively, as in the first example, the error estimation rate for each state may be stored, and the threshold for state estimation (score threshold) may be set higher for states with a higher error estimation rate.

[0083] The following describes the case where an "ambiguous response message" is determined based on the time interval between the output of the spoken message and the reception of the response message (step S7). In the first example (state estimation by rules), the degree to which the estimation conditions are made stricter may be reduced for ambiguous response messages. In the second example (relearning / additional learning by AI), ambiguous response messages whose absolute polarity values ​​are below a threshold value may not be included in the learning data for additional learning.

[0084] Furthermore, the determination means 106 may feed back the updated profile information 142 to the message generation means 103, and thereafter the message generation means 103 may generate an individual message as an utterance message based on the updated profile information 142 (step S4). That is, a cause history for each state estimated by the state estimation means 102 may be stored as the profile information 142, and the message generation means 103 may generate an utterance message based on the cause history. The message generation means 103 may personalize the utterance message to improve user convenience.

[0085] As the profile information 142, a factor history based on user responses may be stored for each state estimated by the state estimation means 102. For example, factors extracted and generated from past response messages such as "I woke up because it was hot and humid" may be stored in association with the "lack of sleep" state. Alternatively, the past response messages themselves may be stored without generating factors. Then, when a "lack of sleep" state is newly estimated, the message generation means 103 may generate a speech message based on the factor history. For example, a speech message such as "Were you unable to sleep because it was hot and humid last night?" may be generated. In the case of manual speech by the operator, the factor information may be displayed on the operator's screen to prompt the operator to create a response message related to the factor, or an example sentence (automatically generated) presented to the operator may be generated based on the factor history.

[0086] Furthermore, the message generation means 103 may determine whether to generate a speech message following the voice utterance, depending on the evaluation result by the determination means 106. A determination history of the determination means 106 for each state estimated by the state estimation means 102 may be stored as the profile information 142. The message generation means 103 may determine, based on the profile information 142, whether to automatically transmit the speech message or manually transmit it based on an instruction from an operator.

[0087] For example, if it is determined that the estimation result is correct or the error estimation rate is small, the message generation means 103 may generate a new speech message in response to the response message and automatically transmit it. At this time, the speech message may be generated based on the content of the response message. This allows the dialogue with the user to continue with appropriate content. On the other hand, if the estimation result is incorrect or the error estimation rate is large, the message generation means 103 may prompt the operator to create an speech message and have the operator manually transmit it.

[0088] 4. Working Example

[0089] FIG. 8 shows an embodiment of the dialogue system.

[0090] The state estimation means 102 of the center device 100 estimates the state of the user (lack of sleep, reliability 0.8) based on the state estimated from the sensor data (detecting two wake-ups after going to sleep) and updates the estimation information 330 (step S2). When the utterance necessity determination means 107 determines that an utterance is necessary based on the state (lack of sleep, reliability 0.8) (step S3, Yes), the message generation means 103 generates a utterance message according to the state (lack of sleep), such as "It seems you didn't sleep well last night. Are you feeling unwell?" (step S4). The response transmission means 104 transmits the utterance message to the terminal device 200 (step S5).

[0091] The message receiving means 105 receives the user's response message "It's fine, I just woke up because it was hot and humid" from the terminal device 200 (step S6, Yes). The determining means 106 determines that the user's state estimated by the state estimating means 102 has a "reliability of 0 (false)" based on the response message "It's fine, I just woke up because it was hot and humid" (step S7), and changes the profile information 142 (step S8).

[0092] Meanwhile, the message receiving means 105 receives the user's response message "That's right, I've been having a hard time falling asleep lately" from the terminal device 200 (step S6, Yes). The determining means 106 determines that the user's state estimated by the state estimating means 102 has "reliability 1 (positive)" based on the response message "That's right, I've been having a hard time falling asleep lately" (step S7), and changes the profile information 142 (step S8).

[0093] 5. Conclusion

[0094] When estimating a user's state using a sensor, it is operationally difficult to confirm whether the estimation result is correct. Therefore, according to this embodiment, the center device 100 estimates the user's state (user state, user behavior, and home state) based on the detection results from the sensor 300 installed in the user's home, and the terminal device 200 outputs a speech message corresponding to the estimation result. The center device 100 determines whether to change the estimation result based on the user's response message to the speech message, i.e., determines whether the estimation result is correct and changes the reliability of the estimation result. This allows the user to easily understand the accuracy of the estimation result of the user's state through natural communication between the user and the terminal device 200. Furthermore, by being able to ask questions in a timely manner while the information is still fresh in the user's mind, higher-quality information, such as factors related to the user's state, can be obtained compared to periodic questions, such as those asked during a weekly visit or medical examination.

[0095] Although the embodiments and modified examples of the present technology have been described above, the present technology is not limited to the above-described embodiments, and it goes without saying that various modifications can be made within the scope of the gist of the present technology.

[0096] In the above embodiment and each modification, a speech message for a user is generated and a state of the user is estimated based on the user's profile information 142. However, this is not limiting, and a speech message for a predetermined user may be generated and a state may be estimated using profile information 142 of other users. For example, if the erroneous estimation rate for a predetermined state is high in the profile information 142 of multiple users, the basic estimation condition may be considered to be incorrect, and the estimation condition for each user may be changed.

[0097] In the above-described embodiment and each modification, the center device 100 and the terminal device 200 are different devices. However, this is not limiting. The center device 100 and the terminal device 200 may be configured as the same device (e.g., an interactive device) and installed in the user's living space. That is, the various means of the center device 100 and the storage means 140 may all be provided within the terminal device 200. In this case, the operator device 160 is assumed to be communicably connected to the interactive device via a network. Furthermore, the interactive device may not be an integrated device, but may be configured as separate devices, such as a main device and a robot. For example, the main device may have the same functions as the center device 100, and the robot may have the same functions as the terminal device 200, and each may be installed in the user's living space. In addition, in the above-described embodiment and each modification, a speech message transmitted from the center device 100 is received by the terminal device 200, and the response output means 203 of the terminal device 200 outputs the received speech message to the user via a speaker S. However, without being limited to this, the terminal device 200 may store the audio data (text data) of the spoken message in association with an identifier (spoken message ID) for identifying the spoken message, and by receiving the spoken message ID of the spoken message to be spoken from the center device 100, output the spoken message associated with the spoken message ID from the speaker S.

[0098] A dialogue system according to one embodiment of the present invention can contribute to solving social issues such as loneliness and isolation among the elderly, extending healthy lifespan, improving the quality of life (QoL) of the elderly, and a declining labor force. Furthermore, the dialogue system according to one embodiment of the present invention can contribute to achieving Goal 3 of the Sustainable Development Goals (SDGs) adopted by the United Nations, "Ensure good health and well-being for all." [Explanation of symbols]

[0099] 10 Dialogue Systems 100 Center Device 101 Sensor data receiving means 102 State estimation means 103 Message Generation Method 104 Response sending means 105 Message Receiving Method 106 Judgment means 107 Means for determining whether or not to speak 140 Memory means 141 Dictionaries 142 Profile Information 160 Operator Device 200 Terminal Device 201 Input Method 202 Message Transmission Method 203 Response output means 300 sensors 310 Sensor Information 320 Estimated Condition Information 330 Estimated Information

Claims

1. one or more sensors in the user's living space; an interactive device; A dialogue system comprising: The dialogue device a sensor data receiving means for receiving sensor data output from the sensor; a state estimation means for estimating a state related to a user in the living space based on the sensor data; a message generating means for generating a speech message according to the state estimated by the state estimating means; an output means for audibly outputting the generated speech message; an input means for inputting a user's response message by voice in response to the spoken message outputted by voice; a determination means for evaluating the state estimated by the state estimation means based on the response message and outputting an evaluation result; have Dialogue system.

2. 2. A dialogue system according to claim 1, The determining means determines a polarity indicating a positive / negative reaction of the user to the speech message based on the response message, and evaluates the state based on the polarity. Dialogue system.

3. 3. A dialogue system according to claim 1 or 2, The message generating means determines whether to generate a speech message following the voice utterance according to the evaluation result. Dialogue system.

4. 2. A dialogue system according to claim 1, the dialogue device further comprises a storage means for storing profile information representing characteristics of the user with respect to each state; the state estimation means estimates the state based on the profile information; The determining means changes the profile information based on the evaluation result. Dialogue system.

5. 2. The dialogue system according to claim 1, the dialogue device further comprises a storage means for storing profile information representing characteristics of the user with respect to each state; The message generating means determines whether to automatically transmit the speech message based on the profile information or to manually transmit the speech message based on an instruction from an operator. Dialogue system.

6. 6. A dialogue system according to claim 4 or 5, The storage means stores, as the profile information, a determination history of the determination means for each of the states estimated by the state estimation means. Dialogue system.

7. 3. A dialogue system according to claim 1 or 2, the dialogue device further comprises a storage means for storing profile information representing characteristics of the user with respect to each state; The message generating means generates the speech message based on the profile information. Dialogue system.

8. 8. A dialogue system according to claim 7, the determining means identifies a cause of the state estimated by the state estimating means based on the response message; the storage means stores a factor history of the factor identified by the determination means as the profile information; The message generating means generates the speech message based on the factor history. Dialogue system.

9. a sensor data receiving means for receiving sensor data output from one or more sensors in the user's living space; a state estimation means for estimating a state related to a user in the living space based on the sensor data; a message generating means for generating a speech message according to the state estimated by the state estimating means; a response sending means for sending the speech message to a terminal device which exchanges messages with a user in a two-way manner; a message receiving means for receiving a response message from the user in response to the spoken message from the terminal device; a determination means for evaluating the state estimated by the state estimation means based on the response message and outputting an evaluation result; Equipped with Center device.

10. The computer of the center device receiving sensor data output from one or more sensors in a user's living space; estimating a state of a user in the living space based on the sensor data; generating a speech message according to the estimated state; transmitting the spoken message to a terminal device that exchanges messages with a user; receiving a response message from the user to the spoken message from the terminal device; evaluating the estimated state based on the response message and outputting an evaluation result; Run How to interact.

Citation Information

Patent Citations

  • Voice guide system

    JP2007286376A

  • Interactive business support system and interactive business support method

    JP2021157419A