system
A system using speech recognition and generative AI models addresses loneliness and health monitoring for elderly individuals, providing timely health updates and emergency notifications to family members.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-02
- Publication Date
- 2026-04-14
AI Technical Summary
Elderly individuals living alone experience loneliness and difficulty in monitoring their health status, making it challenging for family members to provide timely support and reassurance.
A system that utilizes speech recognition, natural language processing, and generative AI models to convert user voice into text, analyze health and mood, generate responses, accumulate health data, and notify family members or emergency contacts.
Alleviates feelings of loneliness among the elderly, effectively monitors their health, and provides timely information to family members, ensuring peace of mind and prompt responses to emergencies.
Smart Images

Figure 2026064726000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In recent years, in a society with a progressing aging population, the problems of loneliness and health management among the elderly have become more serious. In particular, elderly people living alone have few daily conversation partners, which deepens their sense of loneliness, and it is difficult for them to grasp changes in their health status by themselves. Also, family members living far away find it difficult to frequently check on the daily life and health status of the elderly and often feel anxious. In such a situation, there is a need for a system that can alleviate the loneliness of the elderly, grasp their health status, and provide a sense of security to the family.
Means for Solving the Problems
[0005] The present invention solves the above problems by providing a system that includes means for recognizing a user's voice and converting it into text data, means for analyzing the converted text data and generating an appropriate response, means for providing the generated response to the user, means for analyzing the conversation content and inferring the user's health condition and mood, means for accumulating the user's health data and generating reports periodically, means for notifying family members of the generated reports, and means for detecting an emergency situation involving the user and notifying designated contacts. As a result, elderly people can enjoy everyday conversations, their health condition can be easily monitored, and their families can live their daily lives with peace of mind.
[0006] "User" refers to individuals who use this system, especially elderly people.
[0007] "Speech recognition" refers to a technology that converts a user's voice into a digital signal and then converts its content into text data.
[0008] "Text data" refers to data in string format generated by speech recognition.
[0009] "Analysis" refers to the process of scrutinizing information based on text data and extracting patterns and important elements.
[0010] "Response generation" refers to the technology of creating appropriate responses based on analyzed data.
[0011] "Reports" refer to reports that are generated periodically regarding the user's health status and conversation content.
[0012] "Family" refers to relatives or guardians registered as the user's emergency contact.
[0013] "Health status" refers to the user's physical and mental condition.
[0014] An "emergency situation" refers to a situation where a user faces a sudden deterioration in their health or other circumstances requiring urgent attention.
[0015] "Notification" refers to the process of quickly transmitting important information to specified contact points.
Brief Explanation of Drawings
[0016] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which multiple emotions are mapped. [Figure 10] It shows an emotion map to which multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when the emotion engine is combined. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.
Embodiment for Carrying out the Invention
[0017] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0020] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0021] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0024] [First Embodiment]
[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0037] This invention relates to a dialogue system using speech recognition technology, and more particularly to specific embodiments for alleviating feelings of loneliness among the elderly, monitoring their health status, and providing information to their families.
[0038] User registration and initial setup
[0039] The elderly person (hereinafter referred to as the user) first activates the device (hereinafter referred to as the terminal) obtained from the gacha machine. The terminal displays an initial setup screen, and the user enters their name, address, emergency contact information, and health information. This information is sent from the terminal to the server, and the server stores the information in a database. In this way, user-specific information is accumulated on the server.
[0040] Everyday conversation and data collection
[0041] Users can enjoy natural conversations with the device. For example, if a user asks, "What's the weather like today?", the device recognizes this voice and converts it into text data. Next, the device analyzes the converted text data and generates an appropriate response, such as "It's sunny today." At the same time, the conversation is sent to a server, which analyzes the content to infer the user's health and mood, and stores this information in a database.
[0042] Counseling and health management
[0043] When a user asks a question about their daily health, for example, "I haven't had much of an appetite lately, what should I do?", the device recognizes this voice and converts it into text data. Next, the device analyzes the converted text data, and the AI generates an appropriate answer, such as, "It might be because of the heat. Try drinking more water." This question and answer are also sent to the server and stored as the user's health data.
[0044] Providing information to family members
[0045] The server automatically generates periodic reports on the user's health and mood, and notifies family members. Notifications are sent via email or a dedicated app. This allows family members to understand the user's condition in real time, providing them with peace of mind.
[0046] Response when a problem occurs
[0047] If a user experiences a sudden change in their health or an emergency, for example, by telling the device "My chest hurts," the device will use voice recognition to convert this information into text. If the analysis determines that it is an emergency, the device will send this information to a server. The server will immediately notify family members and registered emergency contacts (e.g., medical institutions). Family members will then be able to receive the notification and respond quickly.
[0048] As a concrete example, consider the case of Mr. Tanaka, an elderly person, using the service. When Mr. Tanaka tells the device, "I didn't sleep well last night," that information is sent to the server, and the analysis results are notified to his family. Mr. Tanaka's son can then review the report and take the necessary actions.
[0049] In this way, the present invention realizes a system that can alleviate the user's feelings of loneliness, effectively manage their health, and provide a sense of security to their family.
[0050] The following describes the processing flow.
[0051] Step 1:
[0052] The user turns on the device (terminal) they obtained from the gacha machine.
[0053] Step 2:
[0054] The device displays the initial setup screen. The user enters their name, address, emergency contact information, and health information.
[0055] Step 3:
[0056] The terminal sends the entered information to the server.
[0057] Step 4:
[0058] The server saves the received information to the database.
[0059] Step 5:
[0060] The device notifies the user that the initial setup is complete.
[0061] Step 6:
[0062] The user begins a casual conversation with the device. For example, they might ask, "What's the weather like today?"
[0063] Step 7:
[0064] The device recognizes the user's voice and converts it into text data.
[0065] Step 8:
[0066] The terminal analyzes the converted text data and generates an appropriate response. It generates the response, "It's sunny today."
[0067] Step 9:
[0068] The terminal provides the user with the generated response.
[0069] Step 10:
[0070] The device sends the conversation content to the server.
[0071] Step 11:
[0072] The server analyzes conversation data to infer the user's health status and mood.
[0073] Step 12:
[0074] The server stores the inferred data in the database.
[0075] Step 13:
[0076] Users ask health-related questions using the device. For example, "I haven't had much of an appetite lately, what should I do?"
[0077] Step 14:
[0078] The device recognizes the user's voice and converts it into text data.
[0079] Step 15:
[0080] The device analyzes the converted text data, and the AI generates an appropriate response. For example, it might generate a response like, "It might be due to the heat. Try drinking more water."
[0081] Step 16:
[0082] The terminal provides the user with the generated response.
[0083] Step 17:
[0084] The device sends the question and answer data to the server.
[0085] Step 18:
[0086] The server collects health data and incorporates it into reports.
[0087] Step 19:
[0088] The server periodically generates reports on the user's health status and mood.
[0089] Step 20:
[0090] The server notifies family members of the generated report. Notifications are sent via email or a dedicated app.
[0091] Step 21:
[0092] The user can use the device to report a sudden change in their health or an emergency. For example, they might say, "My chest hurts."
[0093] Step 22:
[0094] The device recognizes the user's voice and converts it into text data. It determines that an emergency situation has occurred.
[0095] Step 23:
[0096] The device sends emergency information to the server.
[0097] Step 24:
[0098] The server receives the emergency call and notifies the family and designated emergency contacts.
[0099] Step 25:
[0100] Family members receive the notification and begin taking action.
[0101] (Example 1)
[0102] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0103] Challenges include elderly people experiencing loneliness in their daily lives and difficulty in responding early to changes in their health. Furthermore, it is difficult for families to monitor the health and mood changes of elderly people in real time, and a system is needed that can respond appropriately in emergencies.
[0104] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0105] In this invention, the server includes means for recognizing the user's voice and converting it into text data, means for analyzing the converted text data and generating an appropriate response, means for providing the generated response to the user, means for analyzing the conversation content and inferring the user's health status and mood, means for accumulating the user's health data and generating reports periodically, means for notifying family members of the generated reports, and means for detecting user emergencies and notifying designated contacts. This makes it possible to alleviate the user's feelings of loneliness, effectively manage their health, and provide reassurance to their family.
[0106] "Means of recognizing user speech and converting it into text data" refers to technology that takes user-generated speech data as input and converts that speech into text format through digital processing.
[0107] "Means of analyzing converted text data and generating appropriate responses" refers to the process of analyzing text data converted from speech using machine learning and natural language processing technologies to generate responses that are suitable for the user's questions and requests.
[0108] "Means of providing generated answers to users" refers to technologies that provide users with generated text-based answers in audio or text display formats.
[0109] "Methods for analyzing conversation content to infer a user's health status and mood" refers to technologies that collect and analyze conversation content with users as data, and then use that information to infer the user's health status and mood.
[0110] "Means for accumulating user health data and generating periodic reports" refers to technology that accumulates user health data obtained through conversations and other data collection methods, and generates periodic reports based on this data.
[0111] "Means of notifying family members of generated reports" refers to technology that notifies family members of generated health status and mood reports via email or a dedicated app.
[0112] "Means for detecting user emergencies and notifying designated contacts" refers to technologies that use voice recognition technology, sensors, etc., to detect user health emergencies and quickly notify designated contacts (family or medical institutions) of that information.
[0113] Modes for carrying out the invention
[0114] This invention is a dialogue system that recognizes the user's voice and generates appropriate responses, with the aim of alleviating feelings of loneliness among the elderly, monitoring their health status, and providing information to their families. This system is composed of speech recognition technology, natural language processing technology, and generative AI models.
[0115] User registration and initial setup
[0116] The user first starts the device (hereinafter referred to as the terminal) and enters their name, address, emergency contact information, and health information on the initial setup screen. This information is sent from the terminal to the server, which stores the information in a database. The server accumulates user-specific information and uses this information to perform subsequent processing.
[0117] Everyday conversation and data collection
[0118] Users can engage in natural voice interactions with the device. For example, if a user asks, "What's the weather like today?", the device recognizes the speech and converts it into text data. This process utilizes Google® Speech-to-Text API and IBM Watson® Speech to Text. The device analyzes the converted text data and generates an appropriate response using generative AI models such as OpenAI® GPT-3® and Microsoft® Azure® Cognitive Services. The user is then provided with a response such as, "It's sunny today." The conversation is also sent to a server, which analyzes the content to infer the user's health and mood, and stores this information in a database.
[0119] Counseling and health management
[0120] When a user asks a question about their daily health, for example, "I haven't had much of an appetite lately, what should I do?", the device recognizes the speech, converts it into text data, and uses a generative AI model to generate an appropriate answer. The user might receive an answer such as, "It might be because of the heat. Try drinking more water." These questions and answers are also sent to the server and stored in a database.
[0121] Providing information to family members
[0122] The server automatically generates periodic reports on the user's health and mood, and notifies family members. Notifications are sent via email or a dedicated app, using APIs from SendGrid and Twilio. This allows family members to understand the user's status in real time and gain peace of mind.
[0123] Response when a problem occurs
[0124] If a user experiences a sudden change in their health or an emergency, for example, by telling the device "My chest hurts," the device will use voice recognition to convert this information into text. If the analysis determines that it is an emergency, the device will send this information to a server. The server will immediately notify family members and registered emergency contacts, allowing family members to respond quickly.
[0125] As a concrete example, if an elderly user says, "I didn't sleep well last night," that information is sent to the server, and the analysis results are notified to their family. The family can then review the report and take the necessary actions. An example of a prompt message would be, "I haven't been sleeping well lately. What should I do?"
[0126] In this way, the present invention realizes a system that alleviates the user's feelings of loneliness, effectively manages their health, and provides a sense of security to their family.
[0127] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0128] Step 1:
[0129] The user starts up the device. The device automatically displays the initial setup screen. The initial setup screen displays input forms for entering name, address, emergency contact information, and health information.
[0130] Input: User input on the terminal (name, address, emergency contact information, health information)
[0131] Output: Confirmation message displayed on the device upon completion of initial setup.
[0132] Specific operation: The user enters the required information into the input form displayed on the screen and confirms it.
[0133] Step 2:
[0134] The device collects the information entered by the user and sends it to the server using the HTTPS protocol. The server receives the information and stores it in a database.
[0135] Input: Personal information entered by the user (name, address, emergency contact information, health information)
[0136] Output: User information stored in the database
[0137] Specific operation: The terminal converts the input data into JSON format and sends it to the server. The server parses the received data, generates an SQL statement to save it to the database, and executes it.
[0138] Step 3:
[0139] The user asks the device, "What's the weather like today?" The device captures the user's voice using its microphone.
[0140] Input: Voice input from the user (question content)
[0141] Output: Audio file
[0142] Specific operation: The device's microphone captures the user's voice as digital data and temporarily stores it.
[0143] Step 4:
[0144] The device uses the Google Speech-to-Text API to convert the captured audio into text data. The converted text data is stored on the device.
[0145] Input: Audio file
[0146] Output: Text data
[0147] Specific operation: The device sends the audio file to the cloud API, receives the returned text data, and saves it.
[0148] Step 5:
[0149] The terminal analyzes the converted text data and generates an answer using a generative AI model such as OpenAI GPT-3. After an appropriate answer is generated, it is saved in text format.
[0150] Input: Text data (converted question content)
[0151] Output: Text data (generated response)
[0152] Specific operation: The device sends text data to the NLP engine and receives and saves the returned response text.
[0153] Step 6:
[0154] The generated text response is converted to speech using the Google Text-to-Speech API and then sent back to the user.
[0155] Input: Text data (generated response)
[0156] Output: Audio data
[0157] Specific operation: The device sends text data to a cloud API, and the returned audio data is sent to the playback device.
[0158] Step 7:
[0159] The terminal sends the conversation content with the user to the server via the HTTPS protocol. The server analyzes the received data and stores it in a database.
[0160] Input: User and device conversation history
[0161] Output: Conversation history stored in the database
[0162] Specific operation: The terminal converts the conversation data into JSON format and sends it to the server. The server parses the received data, generates SQL statements, and saves them to the database.
[0163] Step 8:
[0164] The user asks a health-related question to the device: "I haven't had much of an appetite lately, what should I do?" The device captures the audio and converts it into text data.
[0165] Input: Voice input from the user (health-related questions)
[0166] Output: Text data
[0167] Specific operation: The user's voice is captured using the device's microphone and converted into text data using the Google Speech-to-Text API.
[0168] Step 9:
[0169] The device analyzes the converted text data and uses a generative AI model to generate an appropriate response. For example, it might generate a response like, "It might be due to the heat. Try drinking more water."
[0170] Input: Text data (converted question content)
[0171] Output: Text data (generated response)
[0172] Specific operation: The device sends text data to the NLP engine and saves the returned response text.
[0173] Step 10:
[0174] The generated response is converted into speech and provided to the user. The user receives a response such as, "It might be because of the heat. Try drinking more water."
[0175] Input: Text data (generated response)
[0176] Output: Audio data
[0177] Specific operation: The device sends text data to the Google Text-to-Speech API and plays the returned audio data.
[0178] Step 11:
[0179] The terminal sends the user's question and answer exchange to the server, which then stores this information in a database.
[0180] Input: User and device question and answer history
[0181] Output: History of questions and answers stored in the database
[0182] Specific operation: The terminal converts the question and answer data into JSON format and sends it to the server. The server parses the data, generates SQL statements, and saves them to the database.
[0183] Step 12:
[0184] The server periodically analyzes information in the database and generates reports on the user's health and mood. These reports utilize libraries such as Python's Pandas library.
[0185] Input: Conversation history and health data stored in the database
[0186] Output: Report on health status and mood
[0187] Specific operation: The server runs a Python script to analyze the stored data and generate periodic reports.
[0188] Step 13:
[0189] The server notifies family members of the generated report. Notifications are sent via email or SMS using APIs such as SendGrid or Twilio.
[0190] Input: Generated report
[0191] Output: Notification to family members (email or SMS)
[0192] Specific operation: The server calls the SendGrid or Twilio APIs to send the generated report to the family.
[0193] Step 14:
[0194] The user reports an emergency, such as "I have chest pain," to the device. The device captures the audio and converts it into text data.
[0195] Input: Voice input from the user (reporting an emergency)
[0196] Output: Text data
[0197] Specific operation: The device captures the audio as digital data and converts it into text data using the Google Speech-to-Text API.
[0198] Step 15:
[0199] The terminal sends the converted text data to the server, which then performs analysis to determine if an emergency has occurred.
[0200] Input: Text data (report of an emergency)
[0201] Output: Emergency Determination Result
[0202] Specific operation: The terminal converts the text data into JSON format and sends it to the server. The server runs an analysis algorithm to determine if it is an emergency.
[0203] Step 16:
[0204] If the server determines an emergency is occurring, it will notify registered emergency contacts. It will also use the Twilio API to notify family members and medical institutions via SMS.
[0205] Input: Emergency situation determination result
[0206] Output: Notification to emergency contacts (SMS)
[0207] Specific operation: The server calls the Twilio API to send a notification containing details of the emergency to the family or healthcare provider.
[0208] The above outlines the specific processing steps of the system. This allows users to alleviate feelings of loneliness, manage their health effectively, and provide reassurance to their families.
[0209] (Application Example 1)
[0210] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0211] There is a need for an effective system to alleviate feelings of loneliness among the elderly, monitor their health, and provide information to their families. In particular, there is a lack of means for the elderly to consult about their health on a daily basis or to receive prompt assistance in emergencies.
[0212] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0213] In this invention, the server includes means for recognizing the user's voice and converting it into text data; means for analyzing the converted text data and generating appropriate responses; means for providing the generated responses to the user; means for analyzing the conversation content and inferring the user's health condition and mood; means for accumulating the user's health data and generating reports periodically; means for notifying family members of the generated reports; means for detecting user emergencies and notifying designated contacts; means for answering health-related questions based on voice input; and means for analyzing the results of the question-answering and using a generative AI model to support the user's health management. This enables support for the daily lives of the elderly, as well as rapid management of their health condition and provision of information to their families.
[0214] "A means of recognizing speech and converting it into text data" refers to a technology that captures the speech spoken by a user as a digital signal, performs analysis processing, and converts that speech into a corresponding string of characters.
[0215] "Means for analyzing converted text data and generating appropriate responses" refers to technology that receives string data converted from speech, understands and analyzes its content, and generates appropriate responses to user questions and statements.
[0216] "Means of providing generated answers to users" refers to means of providing answers obtained through analysis and generation to users in the form of audio or text.
[0217] "Methods for analyzing conversation content to infer a user's health status and mood" refers to methods for evaluating and inferring a user's current health status and emotions by analyzing the user's statements, facial expressions, voice tone, etc., based on the user's dialogue history.
[0218] "Methods for accumulating user health data and generating reports periodically" refers to methods for storing health-related information obtained from users in a database and periodically generating reports summarizing their health status and its trends.
[0219] "Means for notifying family members of generated reports" refers to means of sending and notifying the user's family of the regularly generated user health status reports via email or a dedicated application.
[0220] "Means for detecting user emergencies and notifying designated contacts" refers to a means of quickly notifying pre-registered emergency contacts (e.g., family members or medical institutions) when it is determined that a user is in an emergency situation.
[0221] "A means of providing health-related question-and-answer services based on voice input" refers to a technology that generates responses to specific health-related questions based on voice input from the user.
[0222] "Methods using generative AI models to analyze the results of question-answering and provide support for user health management" refers to methods using generative AI models to analyze user health information collected through voice dialogue and use that data to support user health management.
[0223] This invention relates to a dialogue system for supporting the daily lives of elderly people and effectively managing their health. A specific embodiment includes the following configuration.
[0224] Basic system configuration
[0225] The system of the present invention includes a terminal for user voice input, a server for processing the input voice, and a generative AI model. This configuration enables monitoring of the user's health status and provision of appropriate information.
[0226] Hardware and software usage
[0227] Hardware: Mobile devices such as smartphones that allow the user to input voice commands. This includes a microphone and internet connectivity.
[0228] Software: Python, speech recognition libraries (e.g., SpeechRecognition), natural language processing libraries (e.g., the transformers library).
[0229] Speech recognition and text conversion
[0230] The device has a microphone to receive voice input from the user. When the user speaks, the voice is collected by the device's microphone and converted into text data using a speech recognition library. This conversion can be performed using, for example, the Google Speech Recognition API.
[0231] Text data analysis and response generation
[0232] The converted text data is sent to the server. On the server, a natural language processing library is used to analyze the text data and generate appropriate answers to the user's questions and comments. The generated answers are then provided to the user again in either audio or text format.
[0233] Prediction of health status and data accumulation
[0234] The server has an algorithm built in that analyzes conversation content to infer the user's health status and mood. For example, if a user says, "I haven't had much of an appetite lately, what should I do?", the system will collect information about their health and generate periodic reports.
[0235] Providing information to families and emergency response
[0236] The generated reports are regularly notified to the user's family, for example, via email or a dedicated application. Furthermore, if the user faces an emergency, such as saying "my chest hurts," that information is immediately sent to the server and notified to pre-registered emergency contacts. This notification allows for a swift response.
[0237] Applications of Generative AI Tables
[0238] By using generative AI models, health management support can be further enhanced. When the user's voice input is a specific health question, the generative AI model can provide a more accurate answer. For example, if the prompt is "I haven't had much appetite lately, what should I do?", the generative AI model will generate specific advice such as "It might be because of the heat. Try drinking more water."
[0239] This system will alleviate feelings of loneliness that elderly people experience in their daily lives and allow for effective management of their health. Furthermore, it will enable the rapid provision of information to family members and medical institutions, realizing comprehensive support.
[0240] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0241] Step 1:
[0242] The user speaks into the device. At this time, the device's microphone captures the user's voice. The input is the user's voice data, and the output is prepared voice data for speech recognition. Specifically, the microphone collects the voice data and converts the voice signal into digital data in real time.
[0243] Step 2:
[0244] The device converts the collected audio data into text data using a speech recognition library (e.g., SpeechRecognition). The input is pre-prepared audio data, and the output is string data. Specifically, this data is sent to a cloud-based API, where the audio is analyzed and converted back into text.
[0245] Step 3:
[0246] The converted text data is sent to the server. The server receives the text data and performs analysis using a natural language processing library (e.g., transformers). The input is text data, and the output is the analysis result. Specifically, this analysis prepares the system to generate appropriate answers to the user's questions.
[0247] Step 4:
[0248] The server generates appropriate answers using a generative AI model based on the analysis results. The input is text data containing the user's question, and the output is the answer text. Specifically, it uses a generative AI model (e.g., BERT model) to understand the intent of the question and generate an appropriate answer.
[0249] Step 5:
[0250] The server sends the generated response to the terminal. The input is the generated response text, and the output is the text data sent back to the terminal. Specifically, this response text is sent to the terminal and is ready to be sent back to the user.
[0251] Step 6:
[0252] The terminal provides the user with the received response text. The input is text data received from the server, and the output is the response audio or text provided to the user. Specifically, the terminal speaks the text to the user using a speech output device (e.g., a speaker) or provides it as text information on a display device.
[0253] Step 7:
[0254] The server infers the user's health status and mood based on the conversation content. The input is analyzed text data, and the output is health status or mood evaluation data. Specifically, it uses machine learning algorithms to analyze the text content and infer the user's health status and emotions.
[0255] Step 8:
[0256] The server stores user health data and generates reports periodically. The input is health information obtained from conversations, and the output is a health report generated periodically. Specifically, it stores information in a database and automatically generates reports at regular intervals.
[0257] Step 9:
[0258] The generated report is notified to the family. The input is the generated health report, and the output is notification data sent via email or a dedicated application. Specifically, this involves sending an email with the report attached or notifying the family through a dedicated application.
[0259] Step 10:
[0260] If a user experiences an emergency, the server will notify the designated contacts. The input is text data indicating the emergency, and the output is emergency notification data. Specifically, the server activates a system that immediately sends an alert to the emergency contacts.
[0261] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0262] This invention is a system that recognizes a user's voice, converts it into text data, and then analyzes that data to generate appropriate responses. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it can more accurately predict the user's health condition and mood, and adjust the advice and information provided according to those emotions.
[0263] User registration and initial setup
[0264] The elderly person (hereinafter referred to as the user) turns on the device (hereinafter referred to as the terminal) obtained from the gacha machine and displays the initial setup screen. The user enters their name, address, emergency contact information, and health information, and the terminal sends this information to the server. The server saves the received information in its database and notifies the user via the terminal that the initial setup is complete.
[0265] Everyday conversation and data collection
[0266] Users enjoy natural conversations with the device. For example, they might ask, "What's the weather like today?" The device recognizes the user's voice and converts it into text data. Next, it analyzes the converted text data and generates an appropriate response. During the response generation process, an emotion engine analyzes the user's emotions from their voice and text and generates a response that matches those emotions. For example, if the user asks the question in a cheerful voice, it will provide a positive response such as, "It's sunny today. Have a great day!" At the same time, the conversation content is sent to a server and used to estimate the user's health and mood.
[0267] Counseling and health management
[0268] When a user asks a health-related question, for example, "I haven't had much of an appetite lately, what should I do?", the device recognizes this voice, converts it into text data, and analyzes it. The AI then generates an appropriate response, such as, "It might be because of the heat. Try drinking more water." The emotion engine analyzes the user's emotions and, if the user is feeling down, provides a gentle response such as, "Don't worry, try eating a little at a time." This consultation and response are sent to the server and stored in the health data.
[0269] Providing information to family members
[0270] The server automatically generates periodic reports on the user's health and mood, and notifies family members. Notifications are sent via email or a dedicated app. This allows family members to understand the user's condition in real time and gain peace of mind.
[0271] Response when a problem occurs
[0272] If a user experiences a sudden change in their health or an emergency, they can tell the device, for example, "My chest hurts." The device recognizes this voice and converts it into text data. The emotion engine analyzes the user's urgency, and if it determines it is an emergency, it sends the information to the server. The server immediately notifies family members and designated emergency contacts (e.g., medical institutions). Family members receive the notification and can respond quickly.
[0273] As a concrete example, consider the case of Mr. Tanaka, an elderly person, using this system. When Mr. Tanaka speaks to the device and says, "I didn't sleep well last night," that information is sent to the server, and the analysis results are notified to his family. Furthermore, the emotion engine reads feelings of anxiety and fatigue from the tone of Mr. Tanaka's voice and provides a gentle message such as, "That's worrying. Please rest a little today." Mr. Tanaka's son can then review the report and take any necessary action.
[0274] In this way, the present invention realizes a system that can alleviate the user's feelings of loneliness, effectively manage their health, and provide a sense of security to their family. By combining emotional engines, even more sophisticated responses become possible, enabling the provision of support tailored to the user's emotions.
[0275] The following describes the processing flow.
[0276] Step 1:
[0277] The user turns on the device (terminal) they obtained from the gacha machine.
[0278] Step 2:
[0279] The device displays the initial setup screen. The user enters their name, address, emergency contact information, and health information.
[0280] Step 3:
[0281] The terminal sends the input information to the server.
[0282] Step 4:
[0283] The server saves the received information in the database.
[0284] Step 5:
[0285] The terminal notifies the user of the completion of the initial settings.
[0286] Step 6:
[0287] The user starts a daily conversation with the terminal. For example, ask "What's the weather today?".
[0288] Step 7:
[0289] The terminal recognizes the user's voice and converts it into text data.
[0290] Step 8:
[0291] The terminal analyzes the converted text data and generates an appropriate response. Generate a response such as "It's sunny today."
[0292] Step 9:
[0293] The terminal provides the generated response to the user.
[0294] Step 10:
[0295] The terminal sends the conversation content to the server.
[0296] Step 11:
[0297] The server analyzes the conversation data and estimates the user's health condition and mood.
[0298] Step 12:
[0299] The server stores the inferred data in the database.
[0300] Step 13:
[0301] The user asks a health-related question on the terminal. For example, ask "I don't have an appetite recently. What should I do?"
[0302] Step 14:
[0303] The terminal recognizes the user's voice and converts it into text data.
[0304] Step 15:
[0305] The terminal analyzes the converted text data, and the AI generates an appropriate answer. Generate a response such as "It might be due to the heat. Try drinking more fluids."
[0306] Step 16:
[0307] The terminal provides the generated answer to the user.
[0308] Step 17:
[0309] The terminal sends the question and its answer data to the server.
[0310] Step 18:
[0311] The server stores the health data and reflects it in the report.
[0312] Step 19:
[0313] The server periodically generates reports on the user's health status and mood.
[0314] Step 20:
[0315] The server notifies family members of the generated report. Notifications are sent via email or a dedicated app.
[0316] Step 21:
[0317] The user can use the device to report a sudden change in their health or an emergency. For example, they might say, "My chest hurts."
[0318] Step 22:
[0319] The device recognizes the user's voice and converts it into text data. It determines that an emergency situation has occurred.
[0320] Step 23:
[0321] The device sends emergency information to the server.
[0322] Step 24:
[0323] The server receives the emergency call and notifies the family and designated emergency contacts.
[0324] Step 25:
[0325] Family members receive the notification and begin taking action.
[0326] Step 26:
[0327] The emotion engine analyzes the user's voice tone and text data to identify their emotions.
[0328] Step 27:
[0329] The device adjusts the tone and content of its responses based on the analysis results of the emotion engine.
[0330] Step 28:
[0331] The device provides users with responses that are tailored to their emotions. For example, if a user is feeling down, it might respond in a gentle tone, "Don't worry, try eating a little at a time."
[0332] Step 29:
[0333] The device sends emotional data to the server.
[0334] Step 30:
[0335] The server analyzes sentiment data and reflects particularly noteworthy changes and trends in a report.
[0336] (Example 2)
[0337] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0338] In modern society, managing the health and daily communication of the elderly are challenges, and for elderly people living alone, loneliness and dealing with changes in their health are particularly difficult. Furthermore, it is difficult for family members and caregivers to understand the elderly person's condition in real time, creating a need for systems that can respond quickly to emergencies. Additionally, a function that analyzes the elderly person's emotions and provides appropriate support based on those emotions is also crucial.
[0339] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0340] In this invention, the server includes means for recognizing the user's voice and converting it into text data; means for analyzing the converted text data and generating an appropriate response; means for providing the generated response to the user; means for analyzing the conversation content and inferring the user's health status and mood; means for accumulating the user's health data and generating reports periodically; means for notifying family members of the generated reports; means for detecting user emergencies and notifying designated contacts; means for analyzing the user's emotions and generating responses appropriate to those emotions; means for analyzing the user's voice and text data to infer emotions and adjusting responses based on those emotions; and means for providing responses in a gentle tone when the user is feeling down. This not only alleviates the user's feelings of loneliness and effectively manages their health, but also provides reassurance to family members and enables support tailored to the user's emotions.
[0341] "Means for recognizing speech and converting it into text data" refers to devices or software that analyze the speech spoken by a user and convert it into a digital text format.
[0342] "Means for analyzing text data and generating appropriate responses" refers to devices or software that understand the content of text data obtained through speech recognition and generate the most appropriate response to a user's question or request.
[0343] "Means of providing generated answers to users" refers to devices or software that have the function of displaying or reading aloud the system-generated answers to the user.
[0344] "Means for analyzing conversation content to infer a user's health status and mood" refers to devices or software that analyze the content and tone of conversations with a user to infer the user's current health status and mental mood.
[0345] "Means for accumulating user health data and generating periodic reports" refers to devices or software that have the function of storing collected user health data and creating periodic reports based on that data.
[0346] "Means of notifying family members of generated reports" refers to devices or software that have the function of sending the created user health report to family members via email or a dedicated app.
[0347] "Means for detecting user emergencies and notifying designated contacts" refers to devices or software that have the function of detecting when a user is in an emergency and automatically sending notifications to pre-designated emergency contacts.
[0348] "Means of analyzing user emotions and generating responses that correspond to those emotions" refers to devices or software that analyze a user's emotions from their tone of voice and word choice, and generate responses that are appropriate to those emotions.
[0349] "Means of analyzing user voice and text data to infer emotions and adjust responses based on those emotions" refers to devices or software that have the function of reading emotions based on the voice and text data spoken by the user and adjusting the system's response content accordingly.
[0350] "Means of providing a gentle tone of response when a user is feeling down" refers to devices or software that have a function to respond in a gentle and friendly tone when the system determines that a user is feeling down.
[0351] This invention relates to a speech recognition system for the elderly that converts the user's voice into text data, analyzes that data, and then uses an emotion engine to recognize emotions and provide appropriate responses and advice. The system uses the following hardware and software.
[0352] User registration and initial setup
[0353] The user first powers on the device and enters their name, address, emergency contact information, and health information using the initial setup screen that appears. The device sends this information to the server. The server stores the received information in a database and notifies the user via the device when the initial setup is complete. A tablet or smart speaker can be used as the device during this process. The transmitted information uses a secure communication protocol such as HTTPS.
[0354] Everyday conversation and data collection
[0355] When a user asks, "What's the weather like today?", the device uses a speech recognition engine (e.g., Google Cloud Speech-to-Text) to convert the speech into text data. This text data is then sent to a server, where the server's analysis engine analyzes its content. A generative AI model (e.g., OpenAI GPT-3) is used for the analysis. During this process, an emotion engine (e.g., IBM Watson Tone Analyzer) analyzes the user's emotions, and for positive voices, it generates a positive response such as, "It's sunny today. Have a nice day!" The response is then sent from the server to the device and communicated to the user via voice or screen display.
[0356] Counseling and health management
[0357] When a user asks, "I haven't had much of an appetite lately, what should I do?", the device uses speech recognition to convert it into text data and send it to the server. The server analyzes the text data, and an emotion engine analyzes the user's emotions. If the user is feeling down, the generative AI model generates a gentle response such as, "Don't worry, try eating a little at a time." The response is provided to the user through the device, and the content is stored on the server as health data.
[0358] Providing information to family members
[0359] The server automatically generates periodic reports on the user's health and mood, and notifies family members. These notifications are sent via email or a dedicated app. Family members receive this information in real time, allowing them to understand the user's condition. This feature helps alleviate feelings of isolation for the user and provides reassurance to family members.
[0360] Response when a problem occurs
[0361] When a user tells the device they are experiencing chest pain, the device recognizes the voice and converts it into text. The server's emotion engine determines the urgency, and if it is determined to be a high-priority emergency, the server immediately notifies family members and designated emergency contacts. The notification is sent via SMS, email, or a dedicated app, allowing family members to respond quickly.
[0362] Examples and prompts for generative AI models
[0363] Specific example
[0364] When an elderly user says, "I didn't sleep well last night," this information is converted into text data by the device and sent to the server. The server's analysis engine and emotion engine analyze this information and notify the family. In addition, the user is provided with a kind message such as, "That's worrying. Please get some rest today," and the family is informed of the user's condition.
[0365] Examples of prompts for generative AI models
[0366] Please generate responses for when a user says, "I've lost my appetite lately, what should I do?" Include a gentle, supportive response for when the user is feeling down.
[0367] answer:
[0368] 1. "It might be due to the heat. Try drinking more water."
[0369] 2. "Don't worry, try eating a little at a time."
[0370] Thus, the present invention is a system that supports the user's health management and provides a sense of security to their family through user voice recognition and emotion analysis.
[0371] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0372] Detailed step-by-step explanation of the program processing
[0373] User registration and initial setup
[0374] Step 1:
[0375] The user powers on the device. The input is pressing the device's power button. The output is the display of the initial setup screen. The device starts up and displays the initial setup screen.
[0376] Step 2:
[0377] The user enters their name, address, emergency contact information, and health information on the initial setup screen. The input is the user's information (name, address, emergency contact information, health information). The output is the display of this information in the device's input fields.
[0378] Step 3:
[0379] The terminal sends the input information to the server. The input is the information entered by the user. The output is the data stream of the information sent to the server. Communication is secure using the HTTPS protocol.
[0380] Step 4:
[0381] The server saves the received information to the database. The input is user information sent from the device. The output is user information saved in the database. Upon successful saving, an initial setup completion message is generated.
[0382] Step 5:
[0383] The server notifies the terminal that the initial setup is complete. The input is the initial setup completion message. The output is the notification to the terminal. The terminal receives the message and displays completion to the user.
[0384] Everyday conversation and data collection
[0385] Step 1:
[0386] The user speaks to the device, saying, "What's the weather like today?" The input is the user's voice. The output is the device receiving the voice input.
[0387] Step 2:
[0388] The device uses a speech recognition engine (e.g., Google Cloud Speech-to-Text) to convert speech into text data. The input is the user's speech data. The output is text data.
[0389] Step 3:
[0390] The terminal sends the converted text data to the server. The input is the text data on the terminal. The output is the text data sent to the server.
[0391] Step 4:
[0392] The server's analysis engine analyzes text data and generates appropriate responses. The input is the text data sent to the server. The output is the generated response data. A generative AI model (e.g., OpenAI GPT-3) is used for analysis.
[0393] Step 5:
[0394] The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions. The input is text data. The output is the emotion analysis result. If the user asks a question in a cheerful voice, it is judged to be a positive emotion.
[0395] Step 6:
[0396] The server adjusts the response based on the analysis results and generates a response. The input is the analysis results and sentiment information. The output is the adjusted response.
[0397] Step 7:
[0398] The server sends the generated response to the terminal. The input is the adjusted response data. The output is the response sent to the terminal.
[0399] Step 8:
[0400] The device displays the received response on the screen and outputs it as audio using a text-to-speech engine (e.g., Amazon Polly). The input is the response data received from the server. The output is displayed to the user and read aloud.
[0401] Counseling and health management
[0402] Step 1:
[0403] The user speaks into the device saying, "I haven't had much of an appetite lately, what should I do?" The input is the user's voice. The output is the device receiving the voice input.
[0404] Step 2:
[0405] The device recognizes the user's voice and converts it into text data. The input is the user's voice data. The output is text data.
[0406] Step 3:
[0407] The terminal sends text data to the server. The input is the text data on the terminal. The output is the text data sent to the server.
[0408] Step 4:
[0409] The server processes text data to analyze health information. The input is the transmitted text data. The output is the health information analysis result.
[0410] Step 5:
[0411] The server's emotion engine analyzes the user's emotions. The input is text data. The output is the emotion analysis result.
[0412] Step 6:
[0413] The server generates appropriate responses based on health information analysis results and emotional information. Using a generative AI model, it generates responses in a gentle tone, such as, "Don't worry, try eating a little at a time." The input is the analysis results and emotional information. The output is the generated response data.
[0414] Step 7:
[0415] The server sends the generated response to the terminal. The input is the generated response data. The output is the response sent to the terminal.
[0416] Step 8:
[0417] The device communicates the received responses to the user and stores the content as health data on the server. Input consists of the response data received from the server and the resulting audio or display. Output is the stored health data.
[0418] Providing information to family members
[0419] Step 1:
[0420] The server automatically generates periodic reports on the user's health status and mood. The input is accumulated health data. The output is automatically generated report data.
[0421] Step 2:
[0422] The server notifies the family of the generated report. Notifications are sent via email or a dedicated app. The input is the generated report data. The output is the notification to the family.
[0423] Response when a problem occurs
[0424] Step 1:
[0425] The user says "My chest hurts" to the device. The input is the user's voice. The output is the device receiving the voice input.
[0426] Step 2:
[0427] The device recognizes speech and converts it into text data. The input is the user's voice data. The output is text data.
[0428] Step 3:
[0429] The terminal sends text data to the server. The input is the text data on the terminal. The output is the transmission of text data to the server.
[0430] Step 4:
[0431] The server's emotion engine determines the urgency. The input is text data. The output is the urgency assessment result.
[0432] Step 5:
[0433] If the server determines the situation is urgent, it will notify family members and designated emergency contacts. Notifications will be sent via SMS or a dedicated app. Input is an emergency notification message. Output is a notification to family members and emergency contacts.
[0434] In this way, the system of the present invention recognizes the user's voice, converts it into text data, generates an appropriate response based on the analysis results, and provides an emotion-appropriate response through an emotion engine. This enables health management for the elderly, information provision to family members, and rapid response to emergencies.
[0435] (Application Example 2)
[0436] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0437] There is a problem in that elderly users and those requiring health management have difficulty choosing appropriate meals that suit their emotions and health condition. Furthermore, there is a lack of means for family members and caregivers to monitor health conditions in real time and respond quickly. Therefore, there is a need for a system that provides accurate meal suggestions based on the user's emotions and health condition, and that can instantly share necessary information with family members and caregivers.
[0438] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0439] In this invention, the server includes means for recognizing the user's voice and converting it into text data; means for analyzing the converted text data and generating an appropriate response; means for providing the generated response to the user; means for analyzing the conversation content and inferring the user's health condition and mood; means for accumulating the user's health data and generating reports periodically; means for notifying family members of the generated reports; means for detecting the user's emergency and notifying designated contacts; means for suggesting meal menus based on the user's emotional data; and means for automatically placing a food delivery order if the user approves the suggested meal menu. This enables accurate meal suggestions tailored to the user's emotions and health condition, as well as rapid information sharing.
[0440] "Speech recognition" is a technology that converts speech into text data.
[0441] "Text data" refers to string data generated by speech recognition.
[0442] "Emotional data" refers to data obtained as a result of analyzing and evaluating users' emotions.
[0443] "Health status" refers to information indicating the current state of the user's physical and mental health.
[0444] "Mood" refers to information that represents the user's emotional state at that particular time.
[0445] A "meal menu" refers to a selection of appropriate meals suggested based on the user's health condition and mood.
[0446] "Food delivery" is a service that takes orders and delivers meals to customers' homes.
[0447] A "report" is a document that summarizes information about a user's health status and mood.
[0448] An "emergency situation" refers to a serious situation that affects the user's health.
[0449] "Family" refers to the user's relatives and close family members.
[0450] "Notification" is the act of informing other devices or people of specific information.
[0451] "Approval" refers to the act of a user accepting a proposed meal menu.
[0452] A system for realizing this invention includes the following means:
[0453] Speech recognition and analysis
[0454] The user uses a smartphone and speaks into the device, expressing feelings such as "I'm tired today." The user's voice is captured via the microphone and converted into text data using speech recognition technology. For example, the "speech_recognition" library is used for this purpose.
[0455] Emotion analysis
[0456] The converted text data is analyzed using an emotion analysis engine. This engine extracts emotional data from the user's words and evaluates the user's current mood and health status. The "emotion_recognition" library can be used for emotion analysis.
[0457] Predicting health status and mood
[0458] Based on emotional data, the user's health status and mood are inferred. This result is sent to a server and stored in a user-specific database. This allows for the accumulation of user health data, and reports are generated periodically.
[0459] Suggestions for meal menus
[0460] Based on user emotional data and assessments of their health status, meal menus are identified. For example, if a user rates themselves as "tired," a nutritious, energy-boosting meal is suggested. This is selected from a pre-configured database of recommended recipes.
[0461] Automatic ordering and notifications
[0462] Once the user approves the suggested meal menu, the order is automatically sent to the food delivery service. The order details and the user's health information are stored on the server and notified to family members or caregivers via a dedicated app. In the event of an emergency, an immediate notification is sent to the designated contact person.
[0463] Hardware and software used
[0464] Hardware: Smartphones (e.g., iPhone®, ANDROID®)
[0465] Software: "speech_recognition" library, "emotion_recognition" library
[0466] Examples
[0467] When Ms. Tanaka says "I'm tired today" into her smartphone, the voice is converted into text data and sent to the emotion engine. The engine generates emotion data for "tired" and displays a suggestion: "Here's a recommended energy boost menu: Banana smoothie and chicken sandwich." When Ms. Tanaka approves this, the order is automatically sent and her family is notified.
[0468] Example of a prompt
[0469] "I want to create a food delivery service that analyzes emotions from voice and suggests meals. First, imagine a scenario where a user says, 'I'm tired today.' Based on that, please write a program that suggests the most suitable meal menu."
[0470] This invention aims to realize a system that enables accurate meal suggestions tailored to the user's emotions and health condition, as well as rapid information sharing.
[0471] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0472] Step 1:
[0473] The user speaks into the device.
[0474] Input: User's voice
[0475] Output: Audio data (acoustic signal) from the microphone
[0476] Specific action: The user speaks "I'm tired today" into the smartphone's microphone. The microphone captures this audio.
[0477] Step 2:
[0478] The device performs speech recognition and converts the speech data into text data.
[0479] Input: Audio data (acoustic signal)
[0480] Output: Recognized text data
[0481] Specific operation: The captured audio data is converted into text data by speech recognition software (e.g., the "speech_recognition" library).
[0482] Step 3:
[0483] The device analyzes text data and generates sentiment data.
[0484] Input: Text data
[0485] Output: Emotional data (tired, happy, etc.)
[0486] Specific operation: The generated text data is sent to an emotion analysis engine (e.g., the "emotion_recognition" library), which analyzes the text to determine if the emotion is "tired".
[0487] Step 4:
[0488] The server infers the user's health status and mood from emotional data.
[0489] Input: Sentiment data
[0490] Output: User's health status and mood
[0491] Specific operation: By analyzing emotional data, the user's health status and mood are inferred, and this information is stored in the server's database. For example, emotional data such as "tired" is used to infer a "fatigued state."
[0492] Step 5:
[0493] The server suggests an appropriate meal plan based on the server's estimated health condition and mood.
[0494] Input: User's health status and mood
[0495] Output: Suggested meal menus
[0496] Specific operation: The server selects and suggests a suitable meal menu for the user (e.g., banana smoothie and chicken sandwich) from the recipe information stored in the database.
[0497] Step 6:
[0498] The user approves the suggested meal menu.
[0499] Input: Suggested meal menu
[0500] Output: User approval (confirmation information)
[0501] Specific action: The user reviews the meal menu displayed on their smartphone screen and presses the approve button.
[0502] Step 7:
[0503] The server automatically sends the order to the food delivery service.
[0504] Input: User authorization information
[0505] Output: Food delivery order information
[0506] Specific operation: Based on the user's authorization information, the server automatically sends an order to the food delivery service. The order includes the selected meal menu and the user's delivery address information.
[0507] Step 8:
[0508] The server notifies the family of the generated report.
[0509] Input: User's health status, mood, and dietary information
[0510] Output: Notification report to family
[0511] Specific operation: The server generates reports containing the user's health status, mood, and suggested and approved meal information, and notifies family members via a dedicated app or email.
[0512] The above outlines the processing steps of this system. In each step, data processing or calculations are performed based on the appropriate input data to obtain the final output.
[0513] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0514] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0515] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0516] [Second Embodiment]
[0517] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0518] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0519] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0520] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0521] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0522] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0523] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0524] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0525] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0526] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0527] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0528] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0529] This invention relates to a dialogue system using speech recognition technology, and more particularly to specific embodiments for alleviating feelings of loneliness among the elderly, monitoring their health status, and providing information to their families.
[0530] User registration and initial setup
[0531] The elderly person (hereinafter referred to as the user) first activates the device (hereinafter referred to as the terminal) obtained from the gacha machine. The terminal displays an initial setup screen, and the user enters their name, address, emergency contact information, and health information. This information is sent from the terminal to the server, and the server stores the information in a database. In this way, user-specific information is accumulated on the server.
[0532] Everyday conversation and data collection
[0533] Users can enjoy natural conversations with the device. For example, if a user asks, "What's the weather like today?", the device recognizes this voice and converts it into text data. Next, the device analyzes the converted text data and generates an appropriate response, such as "It's sunny today." At the same time, the conversation is sent to a server, which analyzes the content to infer the user's health and mood, and stores this information in a database.
[0534] Counseling and health management
[0535] When a user asks a question about their daily health, for example, "I haven't had much of an appetite lately, what should I do?", the device recognizes this voice and converts it into text data. Next, the device analyzes the converted text data, and the AI generates an appropriate answer, such as, "It might be because of the heat. Try drinking more water." This question and answer are also sent to the server and stored as the user's health data.
[0536] Providing information to family members
[0537] The server automatically generates periodic reports on the user's health and mood, and notifies family members. Notifications are sent via email or a dedicated app. This allows family members to understand the user's condition in real time, providing them with peace of mind.
[0538] Response when a problem occurs
[0539] If a user experiences a sudden change in their health or an emergency, for example, by telling the device "My chest hurts," the device will use voice recognition to convert this information into text. If the analysis determines that it is an emergency, the device will send this information to a server. The server will immediately notify family members and registered emergency contacts (e.g., medical institutions). Family members will then be able to receive the notification and respond quickly.
[0540] As a concrete example, consider the case of Mr. Tanaka, an elderly person, using the service. When Mr. Tanaka tells the device, "I didn't sleep well last night," that information is sent to the server, and the analysis results are notified to his family. Mr. Tanaka's son can then review the report and take the necessary actions.
[0541] In this way, the present invention realizes a system that can alleviate the user's feelings of loneliness, effectively manage their health, and provide a sense of security to their family.
[0542] The following describes the processing flow.
[0543] Step 1:
[0544] The user turns on the device (terminal) they obtained from the gacha machine.
[0545] Step 2:
[0546] The device displays the initial setup screen. The user enters their name, address, emergency contact information, and health information.
[0547] Step 3:
[0548] The terminal sends the entered information to the server.
[0549] Step 4:
[0550] The server saves the received information to the database.
[0551] Step 5:
[0552] The device notifies the user that the initial setup is complete.
[0553] Step 6:
[0554] The user begins a casual conversation with the device. For example, they might ask, "What's the weather like today?"
[0555] Step 7:
[0556] The device recognizes the user's voice and converts it into text data.
[0557] Step 8:
[0558] The terminal analyzes the converted text data and generates an appropriate response. It generates the response, "It's sunny today."
[0559] Step 9:
[0560] The terminal provides the user with the generated response.
[0561] Step 10:
[0562] The device sends the conversation content to the server.
[0563] Step 11:
[0564] The server analyzes conversation data to infer the user's health status and mood.
[0565] Step 12:
[0566] The server stores the inferred data in the database.
[0567] Step 13:
[0568] Users ask health-related questions using the device. For example, "I haven't had much of an appetite lately, what should I do?"
[0569] Step 14:
[0570] The device recognizes the user's voice and converts it into text data.
[0571] Step 15:
[0572] The device analyzes the converted text data, and the AI generates an appropriate response. For example, it might generate a response like, "It might be due to the heat. Try drinking more water."
[0573] Step 16:
[0574] The terminal provides the user with the generated response.
[0575] Step 17:
[0576] The device sends the question and answer data to the server.
[0577] Step 18:
[0578] The server collects health data and incorporates it into reports.
[0579] Step 19:
[0580] The server periodically generates reports on the user's health status and mood.
[0581] Step 20:
[0582] The server notifies family members of the generated report. Notifications are sent via email or a dedicated app.
[0583] Step 21:
[0584] The user can use the device to report a sudden change in their health or an emergency. For example, they might say, "My chest hurts."
[0585] Step 22:
[0586] The device recognizes the user's voice and converts it into text data. It determines that an emergency situation has occurred.
[0587] Step 23:
[0588] The device sends emergency information to the server.
[0589] Step 24:
[0590] The server receives the emergency call and notifies the family and designated emergency contacts.
[0591] Step 25:
[0592] Family members receive the notification and begin taking action.
[0593] (Example 1)
[0594] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0595] Challenges include elderly people experiencing loneliness in their daily lives and difficulty in responding early to changes in their health. Furthermore, it is difficult for families to monitor the health and mood changes of elderly people in real time, and a system is needed that can respond appropriately in emergencies.
[0596] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0597] In this invention, the server includes means for recognizing the user's voice and converting it into text data, means for analyzing the converted text data and generating an appropriate response, means for providing the generated response to the user, means for analyzing the conversation content and inferring the user's health status and mood, means for accumulating the user's health data and generating reports periodically, means for notifying family members of the generated reports, and means for detecting user emergencies and notifying designated contacts. This makes it possible to alleviate the user's feelings of loneliness, effectively manage their health, and provide reassurance to their family.
[0598] "Means of recognizing user speech and converting it into text data" refers to technology that takes user-generated speech data as input and converts that speech into text format through digital processing.
[0599] "Means of analyzing converted text data and generating appropriate responses" refers to the process of analyzing text data converted from speech using machine learning and natural language processing technologies to generate responses that are suitable for the user's questions and requests.
[0600] "Means of providing generated answers to users" refers to technologies that provide users with generated text-based answers in audio or text display formats.
[0601] "Methods for analyzing conversation content to infer a user's health status and mood" refers to technologies that collect and analyze conversation content with users as data, and then use that information to infer the user's health status and mood.
[0602] "Means for accumulating user health data and generating periodic reports" refers to technology that accumulates user health data obtained through conversations and other data collection methods, and generates periodic reports based on this data.
[0603] "Means of notifying family members of generated reports" refers to technology that notifies family members of generated health status and mood reports via email or a dedicated app.
[0604] "Means for detecting user emergencies and notifying designated contacts" refers to technologies that use voice recognition technology, sensors, etc., to detect user health emergencies and quickly notify designated contacts (family or medical institutions) of that information.
[0605] Modes for carrying out the invention
[0606] This invention is a dialogue system that recognizes the user's voice and generates appropriate responses, with the aim of alleviating feelings of loneliness among the elderly, monitoring their health status, and providing information to their families. This system is composed of speech recognition technology, natural language processing technology, and generative AI models.
[0607] User registration and initial setup
[0608] The user first starts the device (hereinafter referred to as the terminal) and enters their name, address, emergency contact information, and health information on the initial setup screen. This information is sent from the terminal to the server, which stores the information in a database. The server accumulates user-specific information and uses this information to perform subsequent processing.
[0609] Everyday conversation and data collection
[0610] Users can engage in natural voice interactions with the device. For example, if a user asks, "What's the weather like today?", the device recognizes the speech and converts it into text data. This process utilizes the Google Speech-to-Text API or IBM Watson Speech to Text. The device then analyzes the converted text data and generates an appropriate response using generative AI models such as OpenAI GPT-3 or Microsoft Azure Cognitive Services. The user is then provided with a response such as, "It's sunny today." The conversation is also sent to a server, which analyzes the content to infer the user's health and mood, and stores this information in a database.
[0611] Counseling and health management
[0612] When a user asks a question about their daily health, for example, "I haven't had much of an appetite lately, what should I do?", the device recognizes the speech, converts it into text data, and uses a generative AI model to generate an appropriate answer. The user might receive an answer such as, "It might be because of the heat. Try drinking more water." These questions and answers are also sent to the server and stored in a database.
[0613] Providing information to family members
[0614] The server automatically generates periodic reports on the user's health and mood, and notifies family members. Notifications are sent via email or a dedicated app, using APIs from SendGrid and Twilio. This allows family members to understand the user's status in real time and gain peace of mind.
[0615] Response when a problem occurs
[0616] If a user experiences a sudden change in their health or an emergency, for example, by telling the device "My chest hurts," the device will use voice recognition to convert this information into text. If the analysis determines that it is an emergency, the device will send this information to a server. The server will immediately notify family members and registered emergency contacts, allowing family members to respond quickly.
[0617] As a concrete example, if an elderly user says, "I didn't sleep well last night," that information is sent to the server, and the analysis results are notified to their family. The family can then review the report and take the necessary actions. An example of a prompt message would be, "I haven't been sleeping well lately. What should I do?"
[0618] In this way, the present invention realizes a system that alleviates the user's feelings of loneliness, effectively manages their health, and provides a sense of security to their family.
[0619] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0620] Step 1:
[0621] The user starts up the device. The device automatically displays the initial setup screen. The initial setup screen displays input forms for entering name, address, emergency contact information, and health information.
[0622] Input: User input on the terminal (name, address, emergency contact information, health information)
[0623] Output: Confirmation message displayed on the device upon completion of initial setup.
[0624] Specific operation: The user enters the required information into the input form displayed on the screen and confirms it.
[0625] Step 2:
[0626] The device collects the information entered by the user and sends it to the server using the HTTPS protocol. The server receives the information and stores it in a database.
[0627] Input: Personal information entered by the user (name, address, emergency contact information, health information)
[0628] Output: User information stored in the database
[0629] Specific operation: The terminal converts the input data into JSON format and sends it to the server. The server parses the received data, generates an SQL statement to save it to the database, and executes it.
[0630] Step 3:
[0631] The user asks the device, "What's the weather like today?" The device captures the user's voice using its microphone.
[0632] Input: Voice input from the user (question content)
[0633] Output: Audio file
[0634] Specific operation: The device's microphone captures the user's voice as digital data and temporarily stores it.
[0635] Step 4:
[0636] The device uses the Google Speech-to-Text API to convert the captured audio into text data. The converted text data is stored on the device.
[0637] Input: Audio file
[0638] Output: Text data
[0639] Specific operation: The device sends the audio file to the cloud API, receives the returned text data, and saves it.
[0640] Step 5:
[0641] The terminal analyzes the converted text data and generates an answer using a generative AI model such as OpenAI GPT-3. After an appropriate answer is generated, it is saved in text format.
[0642] Input: Text data (converted question content)
[0643] Output: Text data (generated response)
[0644] Specific operation: The device sends text data to the NLP engine and receives and saves the returned response text.
[0645] Step 6:
[0646] The generated text response is converted to speech using the Google Text-to-Speech API and then sent back to the user.
[0647] Input: Text data (generated response)
[0648] Output: Audio data
[0649] Specific operation: The device sends text data to a cloud API, and the returned audio data is sent to the playback device.
[0650] Step 7:
[0651] The terminal sends the conversation content with the user to the server via the HTTPS protocol. The server analyzes the received data and stores it in a database.
[0652] Input: User and device conversation history
[0653] Output: Conversation history stored in the database
[0654] Specific operation: The terminal converts the conversation data into JSON format and sends it to the server. The server parses the received data, generates SQL statements, and saves them to the database.
[0655] Step 8:
[0656] The user asks a health-related question to the device: "I haven't had much of an appetite lately, what should I do?" The device captures the audio and converts it into text data.
[0657] Input: Voice input from the user (health-related questions)
[0658] Output: Text data
[0659] Specific operation: The user's voice is captured using the device's microphone and converted into text data using the Google Speech-to-Text API.
[0660] Step 9:
[0661] The device analyzes the converted text data and uses a generative AI model to generate an appropriate response. For example, it might generate a response like, "It might be due to the heat. Try drinking more water."
[0662] Input: Text data (converted question content)
[0663] Output: Text data (generated response)
[0664] Specific operation: The device sends text data to the NLP engine and saves the returned response text.
[0665] Step 10:
[0666] The generated response is converted into speech and provided to the user. The user receives a response such as, "It might be because of the heat. Try drinking more water."
[0667] Input: Text data (generated response)
[0668] Output: Audio data
[0669] Specific operation: The device sends text data to the Google Text-to-Speech API and plays the returned audio data.
[0670] Step 11:
[0671] The terminal sends the user's question and answer exchange to the server, which then stores this information in a database.
[0672] Input: User and device question and answer history
[0673] Output: History of questions and answers stored in the database
[0674] Specific operation: The terminal converts the question and answer data into JSON format and sends it to the server. The server parses the data, generates SQL statements, and saves them to the database.
[0675] Step 12:
[0676] The server periodically analyzes information in the database and generates reports on the user's health and mood. These reports utilize libraries such as Python's Pandas library.
[0677] Input: Conversation history and health data stored in the database
[0678] Output: Report on health status and mood
[0679] Specific operation: The server runs a Python script to analyze the stored data and generate periodic reports.
[0680] Step 13:
[0681] The server notifies family members of the generated report. Notifications are sent via email or SMS using APIs such as SendGrid or Twilio.
[0682] Input: Generated report
[0683] Output: Notification to family members (email or SMS)
[0684] Specific operation: The server calls the SendGrid or Twilio APIs to send the generated report to the family.
[0685] Step 14:
[0686] The user reports an emergency, such as "I have chest pain," to the device. The device captures the audio and converts it into text data.
[0687] Input: Voice input from the user (reporting an emergency)
[0688] Output: Text data
[0689] Specific operation: The device captures the audio as digital data and converts it into text data using the Google Speech-to-Text API.
[0690] Step 15:
[0691] The terminal sends the converted text data to the server, which then performs analysis to determine if an emergency has occurred.
[0692] Input: Text data (report of an emergency)
[0693] Output: Emergency Determination Result
[0694] Specific operation: The terminal converts the text data into JSON format and sends it to the server. The server runs an analysis algorithm to determine if it is an emergency.
[0695] Step 16:
[0696] If the server determines an emergency is occurring, it will notify registered emergency contacts. It will also use the Twilio API to notify family members and medical institutions via SMS.
[0697] Input: Emergency situation determination result
[0698] Output: Notification to emergency contacts (SMS)
[0699] Specific operation: The server calls the Twilio API to send a notification containing details of the emergency to the family or healthcare provider.
[0700] The above outlines the specific processing steps of the system. This allows users to alleviate feelings of loneliness, manage their health effectively, and provide reassurance to their families.
[0701] (Application Example 1)
[0702] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0703] There is a need for an effective system to alleviate feelings of loneliness among the elderly, monitor their health, and provide information to their families. In particular, there is a lack of means for the elderly to consult about their health on a daily basis or to receive prompt assistance in emergencies.
[0704] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0705] In this invention, the server includes means for recognizing the user's voice and converting it into text data; means for analyzing the converted text data and generating appropriate responses; means for providing the generated responses to the user; means for analyzing the conversation content and inferring the user's health condition and mood; means for accumulating the user's health data and generating reports periodically; means for notifying family members of the generated reports; means for detecting user emergencies and notifying designated contacts; means for answering health-related questions based on voice input; and means for analyzing the results of the question-answering and using a generative AI model to support the user's health management. This enables support for the daily lives of the elderly, as well as rapid management of their health condition and provision of information to their families.
[0706] "A means of recognizing speech and converting it into text data" refers to a technology that captures the speech spoken by a user as a digital signal, performs analysis processing, and converts that speech into a corresponding string of characters.
[0707] "Means for analyzing converted text data and generating appropriate responses" refers to technology that receives string data converted from speech, understands and analyzes its content, and generates appropriate responses to user questions and statements.
[0708] "Means of providing generated answers to users" refers to means of providing answers obtained through analysis and generation to users in the form of audio or text.
[0709] "Methods for analyzing conversation content to infer a user's health status and mood" refers to methods for evaluating and inferring a user's current health status and emotions by analyzing the user's statements, facial expressions, voice tone, etc., based on the user's dialogue history.
[0710] "Methods for accumulating user health data and generating reports periodically" refers to methods for storing health-related information obtained from users in a database and periodically generating reports summarizing their health status and its trends.
[0711] "Means for notifying family members of generated reports" refers to means of sending and notifying the user's family of the regularly generated user health status reports via email or a dedicated application.
[0712] "Means for detecting user emergencies and notifying designated contacts" refers to a means of quickly notifying pre-registered emergency contacts (e.g., family members or medical institutions) when it is determined that a user is in an emergency situation.
[0713] "A means of providing health-related question-and-answer services based on voice input" refers to a technology that generates responses to specific health-related questions based on voice input from the user.
[0714] "Methods using generative AI models to analyze the results of question-answering and provide support for user health management" refers to methods using generative AI models to analyze user health information collected through voice dialogue and use that data to support user health management.
[0715] This invention relates to a dialogue system for supporting the daily lives of elderly people and effectively managing their health. A specific embodiment includes the following configuration.
[0716] Basic system configuration
[0717] The system of the present invention includes a terminal for user voice input, a server for processing the input voice, and a generative AI model. This configuration enables monitoring of the user's health status and provision of appropriate information.
[0718] Hardware and software usage
[0719] Hardware: Mobile devices such as smartphones that allow the user to input voice commands. This includes a microphone and internet connectivity.
[0720] Software: Python, speech recognition libraries (e.g., SpeechRecognition), natural language processing libraries (e.g., the transformers library).
[0721] Speech recognition and text conversion
[0722] The device has a microphone to receive voice input from the user. When the user speaks, the voice is collected by the device's microphone and converted into text data using a speech recognition library. This conversion can be performed using, for example, the Google Speech Recognition API.
[0723] Text data analysis and response generation
[0724] The converted text data is sent to the server. On the server, a natural language processing library is used to analyze the text data and generate appropriate answers to the user's questions and comments. The generated answers are then provided to the user again in either audio or text format.
[0725] Prediction of health status and data accumulation
[0726] The server has an algorithm built in that analyzes conversation content to infer the user's health status and mood. For example, if a user says, "I haven't had much of an appetite lately, what should I do?", the system will collect information about their health and generate periodic reports.
[0727] Providing information to families and emergency response
[0728] The generated reports are regularly notified to the user's family, for example, via email or a dedicated application. Furthermore, if the user faces an emergency, such as saying "my chest hurts," that information is immediately sent to the server and notified to pre-registered emergency contacts. This notification allows for a swift response.
[0729] Applications of Generative AI Tables
[0730] By using generative AI models, health management support can be further enhanced. When the user's voice input is a specific health question, the generative AI model can provide a more accurate answer. For example, if the prompt is "I haven't had much appetite lately, what should I do?", the generative AI model will generate specific advice such as "It might be because of the heat. Try drinking more water."
[0731] This system will alleviate feelings of loneliness that elderly people experience in their daily lives and allow for effective management of their health. Furthermore, it will enable the rapid provision of information to family members and medical institutions, realizing comprehensive support.
[0732] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0733] Step 1:
[0734] The user speaks into the device. At this time, the device's microphone captures the user's voice. The input is the user's voice data, and the output is prepared voice data for speech recognition. Specifically, the microphone collects the voice data and converts the voice signal into digital data in real time.
[0735] Step 2:
[0736] The device converts the collected audio data into text data using a speech recognition library (e.g., SpeechRecognition). The input is pre-prepared audio data, and the output is string data. Specifically, this data is sent to a cloud-based API, where the audio is analyzed and converted back into text.
[0737] Step 3:
[0738] The converted text data is sent to the server. The server receives the text data and performs analysis using a natural language processing library (e.g., transformers). The input is text data, and the output is the analysis result. Specifically, this analysis prepares the system to generate appropriate answers to the user's questions.
[0739] Step 4:
[0740] The server generates appropriate answers using a generative AI model based on the analysis results. The input is text data containing the user's question, and the output is the answer text. Specifically, it uses a generative AI model (e.g., BERT model) to understand the intent of the question and generate an appropriate answer.
[0741] Step 5:
[0742] The server sends the generated response to the terminal. The input is the generated response text, and the output is the text data sent back to the terminal. Specifically, this response text is sent to the terminal and is ready to be sent back to the user.
[0743] Step 6:
[0744] The terminal provides the user with the received response text. The input is text data received from the server, and the output is the response audio or text provided to the user. Specifically, the terminal speaks the text to the user using a speech output device (e.g., a speaker) or provides it as text information on a display device.
[0745] Step 7:
[0746] The server infers the user's health status and mood based on the conversation content. The input is analyzed text data, and the output is health status or mood evaluation data. Specifically, it uses machine learning algorithms to analyze the text content and infer the user's health status and emotions.
[0747] Step 8:
[0748] The server stores user health data and generates reports periodically. The input is health information obtained from conversations, and the output is a health report generated periodically. Specifically, it stores information in a database and automatically generates reports at regular intervals.
[0749] Step 9:
[0750] The generated report is notified to the family. The input is the generated health report, and the output is notification data sent via email or a dedicated application. Specifically, this involves sending an email with the report attached or notifying the family through a dedicated application.
[0751] Step 10:
[0752] If a user experiences an emergency, the server will notify the designated contacts. The input is text data indicating the emergency, and the output is emergency notification data. Specifically, the server activates a system that immediately sends an alert to the emergency contacts.
[0753] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0754] This invention is a system that recognizes a user's voice, converts it into text data, and then analyzes that data to generate appropriate responses. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it can more accurately predict the user's health condition and mood, and adjust the advice and information provided according to those emotions.
[0755] User registration and initial setup
[0756] The elderly person (hereinafter referred to as the user) turns on the device (hereinafter referred to as the terminal) obtained from the gacha machine and displays the initial setup screen. The user enters their name, address, emergency contact information, and health information, and the terminal sends this information to the server. The server saves the received information in its database and notifies the user via the terminal that the initial setup is complete.
[0757] Everyday conversation and data collection
[0758] Users enjoy natural conversations with the device. For example, they might ask, "What's the weather like today?" The device recognizes the user's voice and converts it into text data. Next, it analyzes the converted text data and generates an appropriate response. During the response generation process, an emotion engine analyzes the user's emotions from their voice and text and generates a response that matches those emotions. For example, if the user asks the question in a cheerful voice, it will provide a positive response such as, "It's sunny today. Have a great day!" At the same time, the conversation content is sent to a server and used to estimate the user's health and mood.
[0759] Counseling and health management
[0760] When a user asks a health-related question, for example, "I haven't had much of an appetite lately, what should I do?", the device recognizes this voice, converts it into text data, and analyzes it. The AI then generates an appropriate response, such as, "It might be because of the heat. Try drinking more water." The emotion engine analyzes the user's emotions and, if the user is feeling down, provides a gentle response such as, "Don't worry, try eating a little at a time." This consultation and response are sent to the server and stored in the health data.
[0761] Providing information to family members
[0762] The server automatically generates periodic reports on the user's health and mood, and notifies family members. Notifications are sent via email or a dedicated app. This allows family members to understand the user's condition in real time and gain peace of mind.
[0763] Response when a problem occurs
[0764] If a user experiences a sudden change in their health or an emergency, they can tell the device, for example, "My chest hurts." The device recognizes this voice and converts it into text data. The emotion engine analyzes the user's urgency, and if it determines it is an emergency, it sends the information to the server. The server immediately notifies family members and designated emergency contacts (e.g., medical institutions). Family members receive the notification and can respond quickly.
[0765] As a concrete example, consider the case of Mr. Tanaka, an elderly person, using this system. When Mr. Tanaka speaks to the device and says, "I didn't sleep well last night," that information is sent to the server, and the analysis results are notified to his family. Furthermore, the emotion engine reads feelings of anxiety and fatigue from the tone of Mr. Tanaka's voice and provides a gentle message such as, "That's worrying. Please rest a little today." Mr. Tanaka's son can then review the report and take any necessary action.
[0766] In this way, the present invention realizes a system that can alleviate the user's feelings of loneliness, effectively manage their health, and provide a sense of security to their family. By combining emotional engines, even more sophisticated responses become possible, enabling the provision of support tailored to the user's emotions.
[0767] The following describes the processing flow.
[0768] Step 1:
[0769] The user turns on the device (terminal) they obtained from the gacha machine.
[0770] Step 2:
[0771] The device displays the initial setup screen. The user enters their name, address, emergency contact information, and health information.
[0772] Step 3:
[0773] The terminal sends the entered information to the server.
[0774] Step 4:
[0775] The server saves the received information to the database.
[0776] Step 5:
[0777] The device notifies the user that the initial setup is complete.
[0778] Step 6:
[0779] The user begins a casual conversation with the device. For example, they might ask, "What's the weather like today?"
[0780] Step 7:
[0781] The device recognizes the user's voice and converts it into text data.
[0782] Step 8:
[0783] The terminal analyzes the converted text data and generates an appropriate response. It generates the response, "It's sunny today."
[0784] Step 9:
[0785] The terminal provides the user with the generated response.
[0786] Step 10:
[0787] The device sends the conversation content to the server.
[0788] Step 11:
[0789] The server analyzes conversation data to infer the user's health status and mood.
[0790] Step 12:
[0791] The server stores the inferred data in the database.
[0792] Step 13:
[0793] The user asks health-related questions to the device. For example, they might ask, "I haven't had much of an appetite lately, what should I do?"
[0794] Step 14:
[0795] The device recognizes the user's voice and converts it into text data.
[0796] Step 15:
[0797] The device analyzes the converted text data, and the AI generates an appropriate response. For example, it might generate a response like, "It might be due to the heat. Try drinking more water."
[0798] Step 16:
[0799] The terminal provides the user with the generated response.
[0800] Step 17:
[0801] The device sends the question and answer data to the server.
[0802] Step 18:
[0803] The server collects health data and incorporates it into reports.
[0804] Step 19:
[0805] The server periodically generates reports on the user's health status and mood.
[0806] Step 20:
[0807] The server notifies family members of the generated report. Notifications are sent via email or a dedicated app.
[0808] Step 21:
[0809] The user can use the device to report a sudden change in their health or an emergency. For example, they might say, "My chest hurts."
[0810] Step 22:
[0811] The device recognizes the user's voice and converts it into text data. It determines that an emergency situation has occurred.
[0812] Step 23:
[0813] The device sends emergency information to the server.
[0814] Step 24:
[0815] The server receives the emergency call and notifies the family and designated emergency contacts.
[0816] Step 25:
[0817] Family members receive the notification and begin taking action.
[0818] Step 26:
[0819] The emotion engine analyzes the user's voice tone and text data to identify their emotions.
[0820] Step 27:
[0821] The device adjusts the tone and content of its responses based on the analysis results of the emotion engine.
[0822] Step 28:
[0823] The device provides users with responses that are tailored to their emotions. For example, if a user is feeling down, it might respond in a gentle tone, "Don't worry, try eating a little at a time."
[0824] Step 29:
[0825] The device sends emotional data to the server.
[0826] Step 30:
[0827] The server analyzes sentiment data and reflects particularly noteworthy changes and trends in a report.
[0828] (Example 2)
[0829] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0830] In modern society, managing the health and daily communication of the elderly are challenges, and for elderly people living alone, loneliness and dealing with changes in their health are particularly difficult. Furthermore, it is difficult for family members and caregivers to understand the elderly person's condition in real time, creating a need for systems that can respond quickly to emergencies. Additionally, a function that analyzes the elderly person's emotions and provides appropriate support based on those emotions is also crucial.
[0831] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0832] In this invention, the server includes means for recognizing the user's voice and converting it into text data; means for analyzing the converted text data and generating an appropriate response; means for providing the generated response to the user; means for analyzing the conversation content and inferring the user's health status and mood; means for accumulating the user's health data and generating reports periodically; means for notifying family members of the generated reports; means for detecting user emergencies and notifying designated contacts; means for analyzing the user's emotions and generating responses appropriate to those emotions; means for analyzing the user's voice and text data to infer emotions and adjusting responses based on those emotions; and means for providing responses in a gentle tone when the user is feeling down. This not only alleviates the user's feelings of loneliness and effectively manages their health, but also provides reassurance to family members and enables support tailored to the user's emotions.
[0833] "Means for recognizing speech and converting it into text data" refers to devices or software that analyze the speech spoken by a user and convert it into a digital text format.
[0834] "Means for analyzing text data and generating appropriate responses" refers to devices or software that understand the content of text data obtained through speech recognition and generate the most appropriate response to a user's question or request.
[0835] "Means of providing generated answers to users" refers to devices or software that have the function of displaying or reading aloud the system-generated answers to the user.
[0836] "Means for analyzing conversation content to infer a user's health status and mood" refers to devices or software that analyze the content and tone of conversations with a user to infer the user's current health status and mental mood.
[0837] "Means for accumulating user health data and generating periodic reports" refers to devices or software that have the function of storing collected user health data and creating periodic reports based on that data.
[0838] "Means of notifying family members of generated reports" refers to devices or software that have the function of sending the created user health report to family members via email or a dedicated app.
[0839] "Means for detecting user emergencies and notifying designated contacts" refers to devices or software that have the function of detecting when a user is in an emergency and automatically sending notifications to pre-designated emergency contacts.
[0840] "Means of analyzing user emotions and generating responses that correspond to those emotions" refers to devices or software that analyze a user's emotions from their tone of voice and word choice, and generate responses that are appropriate to those emotions.
[0841] "Means of analyzing user voice and text data to infer emotions and adjust responses based on those emotions" refers to devices or software that have the function of reading emotions based on the voice and text data spoken by the user and adjusting the system's response content accordingly.
[0842] "Means of providing a gentle tone of response when a user is feeling down" refers to devices or software that have a function to respond in a gentle and friendly tone when the system determines that a user is feeling down.
[0843] This invention relates to a speech recognition system for the elderly that converts the user's voice into text data, analyzes that data, and then uses an emotion engine to recognize emotions and provide appropriate responses and advice. The system uses the following hardware and software.
[0844] User registration and initial setup
[0845] The user first powers on the device and enters their name, address, emergency contact information, and health information using the initial setup screen that appears. The device sends this information to the server. The server stores the received information in a database and notifies the user via the device when the initial setup is complete. A tablet or smart speaker can be used as the device during this process. The transmitted information uses a secure communication protocol such as HTTPS.
[0846] Everyday conversation and data collection
[0847] When a user asks, "What's the weather like today?", the device uses a speech recognition engine (e.g., Google Cloud Speech-to-Text) to convert the speech into text data. This text data is then sent to a server, where the server's analysis engine analyzes its content. A generative AI model (e.g., OpenAI GPT-3) is used for the analysis. During this process, an emotion engine (e.g., IBM Watson Tone Analyzer) analyzes the user's emotions, and for positive voices, it generates a positive response such as, "It's sunny today. Have a nice day!" The response is then sent from the server to the device and communicated to the user via voice or screen display.
[0848] Counseling and health management
[0849] When a user asks, "I haven't had much of an appetite lately, what should I do?", the device uses speech recognition to convert it into text data and send it to the server. The server analyzes the text data, and an emotion engine analyzes the user's emotions. If the user is feeling down, the generative AI model generates a gentle response such as, "Don't worry, try eating a little at a time." The response is provided to the user through the device, and the content is stored on the server as health data.
[0850] Providing information to family members
[0851] The server automatically generates periodic reports on the user's health and mood, and notifies family members. These notifications are sent via email or a dedicated app. Family members receive this information in real time, allowing them to understand the user's condition. This feature helps alleviate feelings of isolation for the user and provides reassurance to family members.
[0852] Response when a problem occurs
[0853] When a user tells the device they are experiencing chest pain, the device recognizes the voice and converts it into text. The server's emotion engine determines the urgency, and if it is determined to be a high-priority emergency, the server immediately notifies family members and designated emergency contacts. The notification is sent via SMS, email, or a dedicated app, allowing family members to respond quickly.
[0854] Examples and prompts for generative AI models
[0855] Specific example
[0856] When an elderly user says, "I didn't sleep well last night," this information is converted into text data by the device and sent to the server. The server's analysis engine and emotion engine analyze this information and notify the family. In addition, the user is provided with a kind message such as, "That's worrying. Please get some rest today," and the family is informed of the user's condition.
[0857] Examples of prompts for generative AI models
[0858] Please generate responses for when a user says, "I've lost my appetite lately, what should I do?" Include a gentle, supportive response for when the user is feeling down.
[0859] answer:
[0860] 1. "It might be due to the heat. Try drinking more water."
[0861] 2. "Don't worry, try eating a little at a time."
[0862] Thus, the present invention is a system that supports the user's health management and provides a sense of security to their family through user voice recognition and emotion analysis.
[0863] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0864] Detailed step-by-step explanation of the program processing
[0865] User registration and initial setup
[0866] Step 1:
[0867] The user powers on the device. The input is pressing the device's power button. The output is the display of the initial setup screen. The device starts up and displays the initial setup screen.
[0868] Step 2:
[0869] The user enters their name, address, emergency contact information, and health information on the initial setup screen. The input is the user's information (name, address, emergency contact information, health information). The output is the display of this information in the device's input fields.
[0870] Step 3:
[0871] The terminal sends the input information to the server. The input is the information entered by the user. The output is the data stream of the information sent to the server. Communication is secure using the HTTPS protocol.
[0872] Step 4:
[0873] The server saves the received information to the database. The input is user information sent from the device. The output is user information saved in the database. Upon successful saving, an initial setup completion message is generated.
[0874] Step 5:
[0875] The server notifies the terminal that the initial setup is complete. The input is the initial setup completion message. The output is the notification to the terminal. The terminal receives the message and displays completion to the user.
[0876] Everyday conversation and data collection
[0877] Step 1:
[0878] The user speaks to the device, saying, "What's the weather like today?" The input is the user's voice. The output is the device receiving the voice input.
[0879] Step 2:
[0880] The device uses a speech recognition engine (e.g., Google Cloud Speech-to-Text) to convert speech into text data. The input is the user's speech data. The output is text data.
[0881] Step 3:
[0882] The terminal sends the converted text data to the server. The input is the text data on the terminal. The output is the text data sent to the server.
[0883] Step 4:
[0884] The server's analysis engine analyzes text data and generates appropriate responses. The input is the text data sent to the server. The output is the generated response data. A generative AI model (e.g., OpenAI GPT-3) is used for analysis.
[0885] Step 5:
[0886] The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions. The input is text data. The output is the emotion analysis result. If the user asks a question in a cheerful voice, it is judged to be a positive emotion.
[0887] Step 6:
[0888] The server adjusts the response based on the analysis results and generates a response. The input is the analysis results and sentiment information. The output is the adjusted response.
[0889] Step 7:
[0890] The server sends the generated response to the terminal. The input is the adjusted response data. The output is the response sent to the terminal.
[0891] Step 8:
[0892] The device displays the received response on the screen and outputs it as audio using a text-to-speech engine (e.g., Amazon Polly). The input is the response data received from the server. The output is displayed to the user and read aloud.
[0893] Counseling and health management
[0894] Step 1:
[0895] The user speaks into the device saying, "I haven't had much of an appetite lately, what should I do?" The input is the user's voice. The output is the device receiving the voice input.
[0896] Step 2:
[0897] The device recognizes the user's voice and converts it into text data. The input is the user's voice data. The output is text data.
[0898] Step 3:
[0899] The terminal sends text data to the server. The input is the text data on the terminal. The output is the text data sent to the server.
[0900] Step 4:
[0901] The server processes text data to analyze health information. The input is the transmitted text data. The output is the health information analysis result.
[0902] Step 5:
[0903] The server's emotion engine analyzes the user's emotions. The input is text data. The output is the emotion analysis result.
[0904] Step 6:
[0905] The server generates appropriate responses based on health information analysis results and emotional information. Using a generative AI model, it generates responses in a gentle tone, such as, "Don't worry, try eating a little at a time." The input is the analysis results and emotional information. The output is the generated response data.
[0906] Step 7:
[0907] The server sends the generated response to the terminal. The input is the generated response data. The output is the response sent to the terminal.
[0908] Step 8:
[0909] The device communicates the received responses to the user and stores the content as health data on the server. Input consists of the response data received from the server and the resulting audio or display. Output is the stored health data.
[0910] Providing information to family members
[0911] Step 1:
[0912] The server automatically generates periodic reports on the user's health status and mood. The input is accumulated health data. The output is automatically generated report data.
[0913] Step 2:
[0914] The server notifies the family of the generated report. Notifications are sent via email or a dedicated app. The input is the generated report data. The output is the notification to the family.
[0915] Response when a problem occurs
[0916] Step 1:
[0917] The user says "My chest hurts" to the device. The input is the user's voice. The output is the device receiving the voice input.
[0918] Step 2:
[0919] The device recognizes speech and converts it into text data. The input is the user's voice data. The output is text data.
[0920] Step 3:
[0921] The terminal sends text data to the server. The input is the text data on the terminal. The output is the transmission of text data to the server.
[0922] Step 4:
[0923] The server's emotion engine determines the urgency. The input is text data. The output is the urgency assessment result.
[0924] Step 5:
[0925] If the server determines the situation is urgent, it will notify family members and designated emergency contacts. Notifications will be sent via SMS or a dedicated app. Input is an emergency notification message. Output is a notification to family members and emergency contacts.
[0926] In this way, the system of the present invention recognizes the user's voice, converts it into text data, generates an appropriate response based on the analysis results, and provides an emotion-appropriate response through an emotion engine. This enables health management for the elderly, information provision to family members, and rapid response to emergencies.
[0927] (Application Example 2)
[0928] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0929] There is a problem in that elderly users and those requiring health management have difficulty choosing appropriate meals that suit their emotions and health condition. Furthermore, there is a lack of means for family members and caregivers to monitor health conditions in real time and respond quickly. Therefore, there is a need for a system that provides accurate meal suggestions based on the user's emotions and health condition, and that can instantly share necessary information with family members and caregivers.
[0930] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0931] In this invention, the server includes means for recognizing the user's voice and converting it into text data; means for analyzing the converted text data and generating an appropriate response; means for providing the generated response to the user; means for analyzing the conversation content and inferring the user's health condition and mood; means for accumulating the user's health data and generating reports periodically; means for notifying family members of the generated reports; means for detecting the user's emergency and notifying designated contacts; means for suggesting meal menus based on the user's emotional data; and means for automatically placing a food delivery order if the user approves the suggested meal menu. This enables accurate meal suggestions tailored to the user's emotions and health condition, as well as rapid information sharing.
[0932] "Speech recognition" is a technology that converts speech into text data.
[0933] "Text data" refers to string data generated by speech recognition.
[0934] "Emotional data" refers to data obtained as a result of analyzing and evaluating users' emotions.
[0935] "Health status" refers to information indicating the current state of the user's physical and mental health.
[0936] "Mood" refers to information that represents the user's emotional state at that particular time.
[0937] A "meal menu" refers to a selection of appropriate meals suggested based on the user's health condition and mood.
[0938] "Food delivery" is a service that takes orders and delivers meals to customers' homes.
[0939] A "report" is a document that summarizes information about a user's health status and mood.
[0940] An "emergency situation" refers to a serious situation that affects the user's health.
[0941] "Family" refers to the user's relatives and close family members.
[0942] "Notification" is the act of informing other devices or people of specific information.
[0943] "Approval" refers to the act of a user accepting a proposed meal menu.
[0944] A system for realizing this invention includes the following means:
[0945] Speech recognition and analysis
[0946] The user uses a smartphone and speaks into the device, expressing feelings such as "I'm tired today." The user's voice is captured via the microphone and converted into text data using speech recognition technology. For example, the "speech_recognition" library is used for this purpose.
[0947] Emotion analysis
[0948] The converted text data is analyzed using an emotion analysis engine. This engine extracts emotional data from the user's words and evaluates the user's current mood and health status. The "emotion_recognition" library can be used for emotion analysis.
[0949] Predicting health status and mood
[0950] Based on emotional data, the user's health status and mood are inferred. This result is sent to a server and stored in a user-specific database. This allows for the accumulation of user health data, and reports are generated periodically.
[0951] Suggestions for meal menus
[0952] Based on user emotional data and assessments of their health status, meal menus are identified. For example, if a user rates themselves as "tired," a nutritious, energy-boosting meal is suggested. This is selected from a pre-configured database of recommended recipes.
[0953] Automatic ordering and notifications
[0954] Once the user approves the suggested meal menu, the order is automatically sent to the food delivery service. The order details and the user's health information are stored on the server and notified to family members or caregivers via a dedicated app. In the event of an emergency, an immediate notification is sent to the designated contact person.
[0955] Hardware and software used
[0956] Hardware: Smartphones (e.g., iPhone, Android)
[0957] Software: "speech_recognition" library, "emotion_recognition" library
[0958] Examples
[0959] When Ms. Tanaka says "I'm tired today" into her smartphone, the voice is converted into text data and sent to the emotion engine. The engine generates emotion data for "tired" and displays a suggestion: "Here's a recommended energy boost menu: Banana smoothie and chicken sandwich." When Ms. Tanaka approves this, the order is automatically sent and her family is notified.
[0960] Example of a prompt
[0961] "I want to create a food delivery service that analyzes emotions from voice and suggests meals. First, imagine a scenario where a user says, 'I'm tired today.' Based on that, please write a program that suggests the most suitable meal menu."
[0962] This invention aims to realize a system that enables accurate meal suggestions tailored to the user's emotions and health condition, as well as rapid information sharing.
[0963] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0964] Step 1:
[0965] The user speaks into the device.
[0966] Input: User's voice
[0967] Output: Audio data (acoustic signal) from the microphone
[0968] Specific action: The user speaks "I'm tired today" into the smartphone's microphone. The microphone captures this audio.
[0969] Step 2:
[0970] The device performs speech recognition and converts the speech data into text data.
[0971] Input: Audio data (acoustic signal)
[0972] Output: Recognized text data
[0973] Specific operation: The captured audio data is converted into text data by speech recognition software (e.g., the "speech_recognition" library).
[0974] Step 3:
[0975] The device analyzes text data and generates sentiment data.
[0976] Input: Text data
[0977] Output: Emotional data (tired, happy, etc.)
[0978] Specific operation: The generated text data is sent to an emotion analysis engine (e.g., the "emotion_recognition" library), which analyzes the text to determine if the emotion is "tired".
[0979] Step 4:
[0980] The server infers the user's health status and mood from emotional data.
[0981] Input: Sentiment data
[0982] Output: User's health status and mood
[0983] Specific operation: By analyzing emotional data, the user's health status and mood are inferred, and this information is stored in the server's database. For example, emotional data such as "tired" is used to infer a "fatigued state."
[0984] Step 5:
[0985] The server suggests an appropriate meal plan based on the server's estimated health condition and mood.
[0986] Input: User's health status and mood
[0987] Output: Suggested meal menus
[0988] Specific operation: The server selects and suggests a suitable meal menu for the user (e.g., banana smoothie and chicken sandwich) from the recipe information stored in the database.
[0989] Step 6:
[0990] The user approves the suggested meal menu.
[0991] Input: Suggested meal menu
[0992] Output: User approval (confirmation information)
[0993] Specific action: The user reviews the meal menu displayed on their smartphone screen and presses the approve button.
[0994] Step 7:
[0995] The server automatically sends the order to the food delivery service.
[0996] Input: User authorization information
[0997] Output: Food delivery order information
[0998] Specific operation: Based on the user's authorization information, the server automatically sends an order to the food delivery service. The order includes the selected meal menu and the user's delivery address information.
[0999] Step 8:
[1000] The server notifies the family of the generated report.
[1001] Input: User's health status, mood, and dietary information
[1002] Output: Notification report to family
[1003] Specific operation: The server generates reports containing the user's health status, mood, and suggested and approved meal information, and notifies family members via a dedicated app or email.
[1004] The above outlines the processing steps of this system. In each step, data processing or calculations are performed based on the appropriate input data to obtain the final output.
[1005] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1006] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1007] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[1008] [Third Embodiment]
[1009] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[1010] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[1011] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1012] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[1013] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1014] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1015] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1016] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1017] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1018] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1019] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1020] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[1021] This invention relates to a dialogue system using speech recognition technology, and more particularly to specific embodiments for alleviating feelings of loneliness among the elderly, monitoring their health status, and providing information to their families.
[1022] User registration and initial setup
[1023] The elderly person (hereinafter referred to as the user) first activates the device (hereinafter referred to as the terminal) obtained from the gacha machine. The terminal displays an initial setup screen, and the user enters their name, address, emergency contact information, and health information. This information is sent from the terminal to the server, and the server stores the information in a database. In this way, user-specific information is accumulated on the server.
[1024] Everyday conversation and data collection
[1025] Users can enjoy natural conversations with the device. For example, if a user asks, "What's the weather like today?", the device recognizes this voice and converts it into text data. Next, the device analyzes the converted text data and generates an appropriate response, such as "It's sunny today." At the same time, the conversation is sent to a server, which analyzes the content to infer the user's health and mood, and stores this information in a database.
[1026] Counseling and health management
[1027] When a user asks a question about their daily health, for example, "I haven't had much of an appetite lately, what should I do?", the device recognizes this voice and converts it into text data. Next, the device analyzes the converted text data, and the AI generates an appropriate answer, such as, "It might be because of the heat. Try drinking more water." This question and answer are also sent to the server and stored as the user's health data.
[1028] Providing information to family members
[1029] The server automatically generates periodic reports on the user's health and mood, and notifies family members. Notifications are sent via email or a dedicated app. This allows family members to understand the user's condition in real time, providing them with peace of mind.
[1030] Response when a problem occurs
[1031] If a user experiences a sudden change in their health or an emergency, for example, by telling the device "My chest hurts," the device will use voice recognition to convert this information into text. If the analysis determines that it is an emergency, the device will send this information to a server. The server will immediately notify family members and registered emergency contacts (e.g., medical institutions). Family members will then be able to receive the notification and respond quickly.
[1032] As a concrete example, consider the case of Mr. Tanaka, an elderly person, using the service. When Mr. Tanaka tells the device, "I didn't sleep well last night," that information is sent to the server, and the analysis results are notified to his family. Mr. Tanaka's son can then review the report and take the necessary actions.
[1033] In this way, the present invention realizes a system that can alleviate the user's feelings of loneliness, effectively manage their health, and provide a sense of security to their family.
[1034] The following describes the processing flow.
[1035] Step 1:
[1036] The user turns on the device (terminal) they obtained from the gacha machine.
[1037] Step 2:
[1038] The device displays the initial setup screen. The user enters their name, address, emergency contact information, and health information.
[1039] Step 3:
[1040] The terminal sends the entered information to the server.
[1041] Step 4:
[1042] The server saves the received information to the database.
[1043] Step 5:
[1044] The device notifies the user that the initial setup is complete.
[1045] Step 6:
[1046] The user begins a casual conversation with the device. For example, they might ask, "What's the weather like today?"
[1047] Step 7:
[1048] The device recognizes the user's voice and converts it into text data.
[1049] Step 8:
[1050] The terminal analyzes the converted text data and generates an appropriate response. It generates the response, "It's sunny today."
[1051] Step 9:
[1052] The terminal provides the user with the generated response.
[1053] Step 10:
[1054] The device sends the conversation content to the server.
[1055] Step 11:
[1056] The server analyzes conversation data to infer the user's health status and mood.
[1057] Step 12:
[1058] The server stores the inferred data in the database.
[1059] Step 13:
[1060] Users ask health-related questions using the device. For example, "I haven't had much of an appetite lately, what should I do?"
[1061] Step 14:
[1062] The device recognizes the user's voice and converts it into text data.
[1063] Step 15:
[1064] The device analyzes the converted text data, and the AI generates an appropriate response. For example, it might generate a response like, "It might be due to the heat. Try drinking more water."
[1065] Step 16:
[1066] The terminal provides the user with the generated response.
[1067] Step 17:
[1068] The device sends the question and answer data to the server.
[1069] Step 18:
[1070] The server collects health data and incorporates it into reports.
[1071] Step 19:
[1072] The server periodically generates reports on the user's health status and mood.
[1073] Step 20:
[1074] The server notifies family members of the generated report. Notifications are sent via email or a dedicated app.
[1075] Step 21:
[1076] The user can use the device to report a sudden change in their health or an emergency. For example, they might say, "My chest hurts."
[1077] Step 22:
[1078] The device recognizes the user's voice and converts it into text data. It determines that an emergency situation has occurred.
[1079] Step 23:
[1080] The device sends emergency information to the server.
[1081] Step 24:
[1082] The server receives the emergency call and notifies the family and designated emergency contacts.
[1083] Step 25:
[1084] Family members receive the notification and begin taking action.
[1085] (Example 1)
[1086] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1087] Challenges include elderly people experiencing loneliness in their daily lives and difficulty in responding early to changes in their health. Furthermore, it is difficult for families to monitor the health and mood changes of elderly people in real time, and a system is needed that can respond appropriately in emergencies.
[1088] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1089] In this invention, the server includes means for recognizing the user's voice and converting it into text data, means for analyzing the converted text data and generating an appropriate response, means for providing the generated response to the user, means for analyzing the conversation content and inferring the user's health status and mood, means for accumulating the user's health data and generating reports periodically, means for notifying family members of the generated reports, and means for detecting user emergencies and notifying designated contacts. This makes it possible to alleviate the user's feelings of loneliness, effectively manage their health, and provide reassurance to their family.
[1090] "Means of recognizing user speech and converting it into text data" refers to technology that takes user-generated speech data as input and converts that speech into text format through digital processing.
[1091] "Means of analyzing converted text data and generating appropriate responses" refers to the process of analyzing text data converted from speech using machine learning and natural language processing technologies to generate responses that are suitable for the user's questions and requests.
[1092] "Means of providing generated answers to users" refers to technologies that provide users with generated text-based answers in audio or text display formats.
[1093] "Methods for analyzing conversation content to infer a user's health status and mood" refers to technologies that collect and analyze conversation content with users as data, and then use that information to infer the user's health status and mood.
[1094] "Means for accumulating user health data and generating periodic reports" refers to technology that accumulates user health data obtained through conversations and other data collection methods, and generates periodic reports based on this data.
[1095] "Means of notifying family members of generated reports" refers to technology that notifies family members of generated health status and mood reports via email or a dedicated app.
[1096] "Means for detecting user emergencies and notifying designated contacts" refers to technologies that use voice recognition technology, sensors, etc., to detect user health emergencies and quickly notify designated contacts (family or medical institutions) of that information.
[1097] Modes for carrying out the invention
[1098] This invention is a dialogue system that recognizes the user's voice and generates appropriate responses, with the aim of alleviating feelings of loneliness among the elderly, monitoring their health status, and providing information to their families. This system is composed of speech recognition technology, natural language processing technology, and generative AI models.
[1099] User registration and initial setup
[1100] The user first starts the device (hereinafter referred to as the terminal) and enters their name, address, emergency contact information, and health information on the initial setup screen. This information is sent from the terminal to the server, which stores the information in a database. The server accumulates user-specific information and uses this information to perform subsequent processing.
[1101] Everyday conversation and data collection
[1102] Users can engage in natural voice interactions with the device. For example, if a user asks, "What's the weather like today?", the device recognizes the speech and converts it into text data. This process utilizes the Google Speech-to-Text API or IBM Watson Speech to Text. The device then analyzes the converted text data and generates an appropriate response using generative AI models such as OpenAI GPT-3 or Microsoft Azure Cognitive Services. The user is then provided with a response such as, "It's sunny today." The conversation is also sent to a server, which analyzes the content to infer the user's health and mood, and stores this information in a database.
[1103] Counseling and health management
[1104] When a user asks a question about their daily health, for example, "I haven't had much of an appetite lately, what should I do?", the device recognizes the speech, converts it into text data, and uses a generative AI model to generate an appropriate answer. The user might receive an answer such as, "It might be because of the heat. Try drinking more water." These questions and answers are also sent to the server and stored in a database.
[1105] Providing information to family members
[1106] The server automatically generates periodic reports on the user's health and mood, and notifies family members. Notifications are sent via email or a dedicated app, using APIs from SendGrid and Twilio. This allows family members to understand the user's status in real time and gain peace of mind.
[1107] Response when a problem occurs
[1108] If a user experiences a sudden change in their health or an emergency, for example, by telling the device "My chest hurts," the device will use voice recognition to convert this information into text. If the analysis determines that it is an emergency, the device will send this information to a server. The server will immediately notify family members and registered emergency contacts, allowing family members to respond quickly.
[1109] As a concrete example, if an elderly user says, "I didn't sleep well last night," that information is sent to the server, and the analysis results are notified to their family. The family can then review the report and take the necessary actions. An example of a prompt message would be, "I haven't been sleeping well lately. What should I do?"
[1110] In this way, the present invention realizes a system that alleviates the user's feelings of loneliness, effectively manages their health, and provides a sense of security to their family.
[1111] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1112] Step 1:
[1113] The user starts up the device. The device automatically displays the initial setup screen. The initial setup screen displays input forms for entering name, address, emergency contact information, and health information.
[1114] Input: User input on the terminal (name, address, emergency contact information, health information)
[1115] Output: Confirmation message displayed on the device upon completion of initial setup.
[1116] Specific operation: The user enters the required information into the input form displayed on the screen and confirms it.
[1117] Step 2:
[1118] The device collects the information entered by the user and sends it to the server using the HTTPS protocol. The server receives the information and stores it in a database.
[1119] Input: Personal information entered by the user (name, address, emergency contact information, health information)
[1120] Output: User information stored in the database
[1121] Specific operation: The terminal converts the input data into JSON format and sends it to the server. The server parses the received data, generates an SQL statement to save it to the database, and executes it.
[1122] Step 3:
[1123] The user asks the device, "What's the weather like today?" The device captures the user's voice using its microphone.
[1124] Input: Voice input from the user (question content)
[1125] Output: Audio file
[1126] Specific operation: The device's microphone captures the user's voice as digital data and temporarily stores it.
[1127] Step 4:
[1128] The device uses the Google Speech-to-Text API to convert the captured audio into text data. The converted text data is stored on the device.
[1129] Input: Audio file
[1130] Output: Text data
[1131] Specific operation: The device sends the audio file to the cloud API, receives the returned text data, and saves it.
[1132] Step 5:
[1133] The terminal analyzes the converted text data and generates an answer using a generative AI model such as OpenAI GPT-3. After an appropriate answer is generated, it is saved in text format.
[1134] Input: Text data (converted question content)
[1135] Output: Text data (generated response)
[1136] Specific operation: The device sends text data to the NLP engine and receives and saves the returned response text.
[1137] Step 6:
[1138] The generated text response is converted to speech using the Google Text-to-Speech API and then sent back to the user.
[1139] Input: Text data (generated response)
[1140] Output: Audio data
[1141] Specific operation: The device sends text data to a cloud API, and the returned audio data is sent to the playback device.
[1142] Step 7:
[1143] The terminal sends the conversation content with the user to the server via the HTTPS protocol. The server analyzes the received data and stores it in a database.
[1144] Input: User and device conversation history
[1145] Output: Conversation history stored in the database
[1146] Specific operation: The terminal converts the conversation data into JSON format and sends it to the server. The server parses the received data, generates SQL statements, and saves them to the database.
[1147] Step 8:
[1148] The user asks a health-related question to the device: "I haven't had much of an appetite lately, what should I do?" The device captures the audio and converts it into text data.
[1149] Input: Voice input from the user (health-related questions)
[1150] Output: Text data
[1151] Specific operation: The user's voice is captured using the device's microphone and converted into text data using the Google Speech-to-Text API.
[1152] Step 9:
[1153] The device analyzes the converted text data and uses a generative AI model to generate an appropriate response. For example, it might generate a response like, "It might be due to the heat. Try drinking more water."
[1154] Input: Text data (converted question content)
[1155] Output: Text data (generated response)
[1156] Specific operation: The device sends text data to the NLP engine and saves the returned response text.
[1157] Step 10:
[1158] The generated response is converted into speech and provided to the user. The user receives a response such as, "It might be because of the heat. Try drinking more water."
[1159] Input: Text data (generated response)
[1160] Output: Audio data
[1161] Specific operation: The device sends text data to the Google Text-to-Speech API and plays the returned audio data.
[1162] Step 11:
[1163] The terminal sends the user's question and answer exchange to the server, which then stores this information in a database.
[1164] Input: User and device question and answer history
[1165] Output: History of questions and answers stored in the database
[1166] Specific operation: The terminal converts the question and answer data into JSON format and sends it to the server. The server parses the data, generates SQL statements, and saves them to the database.
[1167] Step 12:
[1168] The server periodically analyzes information in the database and generates reports on the user's health and mood. These reports utilize libraries such as Python's Pandas library.
[1169] Input: Conversation history and health data stored in the database
[1170] Output: Report on health status and mood
[1171] Specific operation: The server runs a Python script to analyze the stored data and generate periodic reports.
[1172] Step 13:
[1173] The server notifies family members of the generated report. Notifications are sent via email or SMS using APIs such as SendGrid or Twilio.
[1174] Input: Generated report
[1175] Output: Notification to family members (email or SMS)
[1176] Specific operation: The server calls the SendGrid or Twilio APIs to send the generated report to the family.
[1177] Step 14:
[1178] The user reports an emergency, such as "I have chest pain," to the device. The device captures the audio and converts it into text data.
[1179] Input: Voice input from the user (reporting an emergency)
[1180] Output: Text data
[1181] Specific operation: The device captures the audio as digital data and converts it into text data using the Google Speech-to-Text API.
[1182] Step 15:
[1183] The terminal sends the converted text data to the server, which then performs analysis to determine if an emergency has occurred.
[1184] Input: Text data (report of an emergency)
[1185] Output: Emergency Determination Result
[1186] Specific operation: The terminal converts the text data into JSON format and sends it to the server. The server runs an analysis algorithm to determine if it is an emergency.
[1187] Step 16:
[1188] If the server determines an emergency is occurring, it will notify registered emergency contacts. It will also use the Twilio API to notify family members and medical institutions via SMS.
[1189] Input: Emergency situation determination result
[1190] Output: Notification to emergency contacts (SMS)
[1191] Specific operation: The server calls the Twilio API to send a notification containing details of the emergency to the family or healthcare provider.
[1192] The above outlines the specific processing steps of the system. This allows users to alleviate feelings of loneliness, manage their health effectively, and provide reassurance to their families.
[1193] (Application Example 1)
[1194] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1195] There is a need for an effective system to alleviate feelings of loneliness among the elderly, monitor their health, and provide information to their families. In particular, there is a lack of means for the elderly to consult about their health on a daily basis or to receive prompt assistance in emergencies.
[1196] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1197] In this invention, the server includes means for recognizing the user's voice and converting it into text data; means for analyzing the converted text data and generating appropriate responses; means for providing the generated responses to the user; means for analyzing the conversation content and inferring the user's health condition and mood; means for accumulating the user's health data and generating reports periodically; means for notifying family members of the generated reports; means for detecting user emergencies and notifying designated contacts; means for answering health-related questions based on voice input; and means for analyzing the results of the question-answering and using a generative AI model to support the user's health management. This enables support for the daily lives of the elderly, as well as rapid management of their health condition and provision of information to their families.
[1198] "A means of recognizing speech and converting it into text data" refers to a technology that captures the speech spoken by a user as a digital signal, performs analysis processing, and converts that speech into a corresponding string of characters.
[1199] "Means for analyzing converted text data and generating appropriate responses" refers to technology that receives string data converted from speech, understands and analyzes its content, and generates appropriate responses to user questions and statements.
[1200] "Means of providing generated answers to users" refers to means of providing answers obtained through analysis and generation to users in the form of audio or text.
[1201] "Methods for analyzing conversation content to infer a user's health status and mood" refers to methods for evaluating and inferring a user's current health status and emotions by analyzing the user's statements, facial expressions, voice tone, etc., based on the user's dialogue history.
[1202] "Methods for accumulating user health data and generating reports periodically" refers to methods for storing health-related information obtained from users in a database and periodically generating reports summarizing their health status and its trends.
[1203] "Means for notifying family members of generated reports" refers to means of sending and notifying the user's family of the regularly generated user health status reports via email or a dedicated application.
[1204] "Means for detecting user emergencies and notifying designated contacts" refers to a means of quickly notifying pre-registered emergency contacts (e.g., family members or medical institutions) when it is determined that a user is in an emergency situation.
[1205] "A means of providing health-related question-and-answer services based on voice input" refers to a technology that generates responses to specific health-related questions based on voice input from the user.
[1206] "Methods using generative AI models to analyze the results of question-answering and provide support for user health management" refers to methods using generative AI models to analyze user health information collected through voice dialogue and use that data to support user health management.
[1207] This invention relates to a dialogue system for supporting the daily lives of elderly people and effectively managing their health. A specific embodiment includes the following configuration.
[1208] Basic system configuration
[1209] The system of the present invention includes a terminal for user voice input, a server for processing the input voice, and a generative AI model. This configuration enables monitoring of the user's health status and provision of appropriate information.
[1210] Hardware and software usage
[1211] Hardware: Mobile devices such as smartphones that allow the user to input voice commands. This includes a microphone and internet connectivity.
[1212] Software: Python, speech recognition libraries (e.g., SpeechRecognition), natural language processing libraries (e.g., the transformers library).
[1213] Speech recognition and text conversion
[1214] The device has a microphone to receive voice input from the user. When the user speaks, the voice is collected by the device's microphone and converted into text data using a speech recognition library. This conversion can be performed using, for example, the Google Speech Recognition API.
[1215] Text data analysis and response generation
[1216] The converted text data is sent to the server. On the server, a natural language processing library is used to analyze the text data and generate appropriate answers to the user's questions and comments. The generated answers are then provided to the user again in either audio or text format.
[1217] Prediction of health status and data accumulation
[1218] The server has an algorithm built in that analyzes conversation content to infer the user's health status and mood. For example, if a user says, "I haven't had much of an appetite lately, what should I do?", the system will collect information about their health and generate periodic reports.
[1219] Providing information to families and emergency response
[1220] The generated reports are regularly notified to the user's family, for example, via email or a dedicated application. Furthermore, if the user faces an emergency, such as saying "my chest hurts," that information is immediately sent to the server and notified to pre-registered emergency contacts. This notification allows for a swift response.
[1221] Applications of Generative AI Tables
[1222] By using generative AI models, health management support can be further enhanced. When the user's voice input is a specific health question, the generative AI model can provide a more accurate answer. For example, if the prompt is "I haven't had much appetite lately, what should I do?", the generative AI model will generate specific advice such as "It might be because of the heat. Try drinking more water."
[1223] This system will alleviate feelings of loneliness that elderly people experience in their daily lives and allow for effective management of their health. Furthermore, it will enable the rapid provision of information to family members and medical institutions, realizing comprehensive support.
[1224] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1225] Step 1:
[1226] The user speaks into the device. At this time, the device's microphone captures the user's voice. The input is the user's voice data, and the output is prepared voice data for speech recognition. Specifically, the microphone collects the voice data and converts the voice signal into digital data in real time.
[1227] Step 2:
[1228] The device converts the collected audio data into text data using a speech recognition library (e.g., SpeechRecognition). The input is pre-prepared audio data, and the output is string data. Specifically, this data is sent to a cloud-based API, where the audio is analyzed and converted back into text.
[1229] Step 3:
[1230] The converted text data is sent to the server. The server receives the text data and performs analysis using a natural language processing library (e.g., transformers). The input is text data, and the output is the analysis result. Specifically, this analysis prepares the system to generate appropriate answers to the user's questions.
[1231] Step 4:
[1232] The server generates appropriate answers using a generative AI model based on the analysis results. The input is text data containing the user's question, and the output is the answer text. Specifically, it uses a generative AI model (e.g., BERT model) to understand the intent of the question and generate an appropriate answer.
[1233] Step 5:
[1234] The server sends the generated response to the terminal. The input is the generated response text, and the output is the text data sent back to the terminal. Specifically, this response text is sent to the terminal and is ready to be sent back to the user.
[1235] Step 6:
[1236] The terminal provides the user with the received response text. The input is text data received from the server, and the output is the response audio or text provided to the user. Specifically, the terminal speaks the text to the user using a speech output device (e.g., a speaker) or provides it as text information on a display device.
[1237] Step 7:
[1238] The server infers the user's health status and mood based on the conversation content. The input is analyzed text data, and the output is health status or mood evaluation data. Specifically, it uses machine learning algorithms to analyze the text content and infer the user's health status and emotions.
[1239] Step 8:
[1240] The server stores user health data and generates reports periodically. The input is health information obtained from conversations, and the output is a health report generated periodically. Specifically, it stores information in a database and automatically generates reports at regular intervals.
[1241] Step 9:
[1242] The generated report is notified to the family. The input is the generated health report, and the output is notification data sent via email or a dedicated application. Specifically, this involves sending an email with the report attached or notifying the family through a dedicated application.
[1243] Step 10:
[1244] If a user experiences an emergency, the server will notify the designated contacts. The input is text data indicating the emergency, and the output is emergency notification data. Specifically, the server activates a system that immediately sends an alert to the emergency contacts.
[1245] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1246] This invention is a system that recognizes a user's voice, converts it into text data, and then analyzes that data to generate appropriate responses. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it can more accurately predict the user's health condition and mood, and adjust the advice and information provided according to those emotions.
[1247] User registration and initial setup
[1248] The elderly person (hereinafter referred to as the user) turns on the device (hereinafter referred to as the terminal) obtained from the gacha machine and displays the initial setup screen. The user enters their name, address, emergency contact information, and health information, and the terminal sends this information to the server. The server saves the received information in its database and notifies the user via the terminal that the initial setup is complete.
[1249] Everyday conversation and data collection
[1250] Users enjoy natural conversations with the device. For example, they might ask, "What's the weather like today?" The device recognizes the user's voice and converts it into text data. Next, it analyzes the converted text data and generates an appropriate response. During the response generation process, an emotion engine analyzes the user's emotions from their voice and text and generates a response that matches those emotions. For example, if the user asks the question in a cheerful voice, it will provide a positive response such as, "It's sunny today. Have a great day!" At the same time, the conversation content is sent to a server and used to estimate the user's health and mood.
[1251] Counseling and health management
[1252] When a user asks a health-related question, for example, "I haven't had much of an appetite lately, what should I do?", the device recognizes this voice, converts it into text data, and analyzes it. The AI then generates an appropriate response, such as, "It might be because of the heat. Try drinking more water." The emotion engine analyzes the user's emotions and, if the user is feeling down, provides a gentle response such as, "Don't worry, try eating a little at a time." This consultation and response are sent to the server and stored in the health data.
[1253] Providing information to family members
[1254] The server automatically generates periodic reports on the user's health and mood, and notifies family members. Notifications are sent via email or a dedicated app. This allows family members to understand the user's condition in real time and gain peace of mind.
[1255] Response when a problem occurs
[1256] If a user experiences a sudden change in their health or an emergency, they can tell the device, for example, "My chest hurts." The device recognizes this voice and converts it into text data. The emotion engine analyzes the user's urgency, and if it determines it is an emergency, it sends the information to the server. The server immediately notifies family members and designated emergency contacts (e.g., medical institutions). Family members receive the notification and can respond quickly.
[1257] As a concrete example, consider the case of Mr. Tanaka, an elderly person, using this system. When Mr. Tanaka speaks to the device and says, "I didn't sleep well last night," that information is sent to the server, and the analysis results are notified to his family. Furthermore, the emotion engine reads feelings of anxiety and fatigue from the tone of Mr. Tanaka's voice and provides a gentle message such as, "That's worrying. Please rest a little today." Mr. Tanaka's son can then review the report and take any necessary action.
[1258] In this way, the present invention realizes a system that can alleviate the user's feelings of loneliness, effectively manage their health, and provide a sense of security to their family. By combining emotional engines, even more sophisticated responses become possible, enabling the provision of support tailored to the user's emotions.
[1259] The following describes the processing flow.
[1260] Step 1:
[1261] The user turns on the device (terminal) they obtained from the gacha machine.
[1262] Step 2:
[1263] The device displays the initial setup screen. The user enters their name, address, emergency contact information, and health information.
[1264] Step 3:
[1265] The terminal sends the entered information to the server.
[1266] Step 4:
[1267] The server saves the received information to the database.
[1268] Step 5:
[1269] The device notifies the user that the initial setup is complete.
[1270] Step 6:
[1271] The user begins a casual conversation with the device. For example, they might ask, "What's the weather like today?"
[1272] Step 7:
[1273] The device recognizes the user's voice and converts it into text data.
[1274] Step 8:
[1275] The terminal analyzes the converted text data and generates an appropriate response. It generates the response, "It's sunny today."
[1276] Step 9:
[1277] The terminal provides the user with the generated response.
[1278] Step 10:
[1279] The device sends the conversation content to the server.
[1280] Step 11:
[1281] The server analyzes conversation data to infer the user's health status and mood.
[1282] Step 12:
[1283] The server stores the inferred data in the database.
[1284] Step 13:
[1285] The user asks health-related questions to the device. For example, they might ask, "I haven't had much of an appetite lately, what should I do?"
[1286] Step 14:
[1287] The device recognizes the user's voice and converts it into text data.
[1288] Step 15:
[1289] The device analyzes the converted text data, and the AI generates an appropriate response. For example, it might generate a response like, "It might be due to the heat. Try drinking more water."
[1290] Step 16:
[1291] The terminal provides the user with the generated response.
[1292] Step 17:
[1293] The device sends the question and answer data to the server.
[1294] Step 18:
[1295] The server collects health data and incorporates it into reports.
[1296] Step 19:
[1297] The server periodically generates reports on the user's health status and mood.
[1298] Step 20:
[1299] The server notifies family members of the generated report. Notifications are sent via email or a dedicated app.
[1300] Step 21:
[1301] The user can use the device to report a sudden change in their health or an emergency. For example, they might say, "My chest hurts."
[1302] Step 22:
[1303] The device recognizes the user's voice and converts it into text data. It determines that an emergency situation has occurred.
[1304] Step 23:
[1305] The device sends emergency information to the server.
[1306] Step 24:
[1307] The server receives the emergency call and notifies the family and designated emergency contacts.
[1308] Step 25:
[1309] Family members receive the notification and begin taking action.
[1310] Step 26:
[1311] The emotion engine analyzes the user's voice tone and text data to identify their emotions.
[1312] Step 27:
[1313] The device adjusts the tone and content of its responses based on the analysis results of the emotion engine.
[1314] Step 28:
[1315] The device provides users with responses that are tailored to their emotions. For example, if a user is feeling down, it might respond in a gentle tone, "Don't worry, try eating a little at a time."
[1316] Step 29:
[1317] The device sends emotional data to the server.
[1318] Step 30:
[1319] The server analyzes sentiment data and reflects particularly noteworthy changes and trends in a report.
[1320] (Example 2)
[1321] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1322] In modern society, managing the health and daily communication of the elderly are challenges, and for elderly people living alone, loneliness and dealing with changes in their health are particularly difficult. Furthermore, it is difficult for family members and caregivers to understand the elderly person's condition in real time, creating a need for systems that can respond quickly to emergencies. Additionally, a function that analyzes the elderly person's emotions and provides appropriate support based on those emotions is also crucial.
[1323] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1324] In this invention, the server includes means for recognizing the user's voice and converting it into text data; means for analyzing the converted text data and generating an appropriate response; means for providing the generated response to the user; means for analyzing the conversation content and inferring the user's health status and mood; means for accumulating the user's health data and generating reports periodically; means for notifying family members of the generated reports; means for detecting user emergencies and notifying designated contacts; means for analyzing the user's emotions and generating responses appropriate to those emotions; means for analyzing the user's voice and text data to infer emotions and adjusting responses based on those emotions; and means for providing responses in a gentle tone when the user is feeling down. This not only alleviates the user's feelings of loneliness and effectively manages their health, but also provides reassurance to family members and enables support tailored to the user's emotions.
[1325] "Means for recognizing speech and converting it into text data" refers to devices or software that analyze the speech spoken by a user and convert it into a digital text format.
[1326] "Means for analyzing text data and generating appropriate responses" refers to devices or software that understand the content of text data obtained through speech recognition and generate the most appropriate response to a user's question or request.
[1327] "Means of providing generated answers to users" refers to devices or software that have the function of displaying or reading aloud the system-generated answers to the user.
[1328] "Means for analyzing conversation content to infer a user's health status and mood" refers to devices or software that analyze the content and tone of conversations with a user to infer the user's current health status and mental mood.
[1329] "Means for accumulating user health data and generating periodic reports" refers to devices or software that have the function of storing collected user health data and creating periodic reports based on that data.
[1330] "Means of notifying family members of generated reports" refers to devices or software that have the function of sending the created user health report to family members via email or a dedicated app.
[1331] "Means for detecting user emergencies and notifying designated contacts" refers to devices or software that have the function of detecting when a user is in an emergency and automatically sending notifications to pre-designated emergency contacts.
[1332] "Means of analyzing user emotions and generating responses that correspond to those emotions" refers to devices or software that analyze a user's emotions from their tone of voice and word choice, and generate responses that are appropriate to those emotions.
[1333] "Means of analyzing user voice and text data to infer emotions and adjust responses based on those emotions" refers to devices or software that have the function of reading emotions based on the voice and text data spoken by the user and adjusting the system's response content accordingly.
[1334] "Means of providing a gentle tone of response when a user is feeling down" refers to devices or software that have a function to respond in a gentle and friendly tone when the system determines that a user is feeling down.
[1335] This invention relates to a speech recognition system for the elderly that converts the user's voice into text data, analyzes that data, and then uses an emotion engine to recognize emotions and provide appropriate responses and advice. The system uses the following hardware and software.
[1336] User registration and initial setup
[1337] The user first powers on the device and enters their name, address, emergency contact information, and health information using the initial setup screen that appears. The device sends this information to the server. The server stores the received information in a database and notifies the user via the device when the initial setup is complete. A tablet or smart speaker can be used as the device during this process. The transmitted information uses a secure communication protocol such as HTTPS.
[1338] Everyday conversation and data collection
[1339] When a user asks, "What's the weather like today?", the device uses a speech recognition engine (e.g., Google Cloud Speech-to-Text) to convert the speech into text data. This text data is then sent to a server, where the server's analysis engine analyzes its content. A generative AI model (e.g., OpenAI GPT-3) is used for the analysis. During this process, an emotion engine (e.g., IBM Watson Tone Analyzer) analyzes the user's emotions, and for positive voices, it generates a positive response such as, "It's sunny today. Have a nice day!" The response is then sent from the server to the device and communicated to the user via voice or screen display.
[1340] Counseling and health management
[1341] When a user asks, "I haven't had much of an appetite lately, what should I do?", the device uses speech recognition to convert it into text data and send it to the server. The server analyzes the text data, and an emotion engine analyzes the user's emotions. If the user is feeling down, the generative AI model generates a gentle response such as, "Don't worry, try eating a little at a time." The response is provided to the user through the device, and the content is stored on the server as health data.
[1342] Providing information to family members
[1343] The server automatically generates periodic reports on the user's health and mood, and notifies family members. These notifications are sent via email or a dedicated app. Family members receive this information in real time, allowing them to understand the user's condition. This feature helps alleviate feelings of isolation for the user and provides reassurance to family members.
[1344] Response when a problem occurs
[1345] When a user tells the device they are experiencing chest pain, the device recognizes the voice and converts it into text. The server's emotion engine determines the urgency, and if it is determined to be a high-priority emergency, the server immediately notifies family members and designated emergency contacts. The notification is sent via SMS, email, or a dedicated app, allowing family members to respond quickly.
[1346] Examples and prompts for generative AI models
[1347] Specific example
[1348] When an elderly user says, "I didn't sleep well last night," this information is converted into text data by the device and sent to the server. The server's analysis engine and emotion engine analyze this information and notify the family. In addition, the user is provided with a kind message such as, "That's worrying. Please get some rest today," and the family is informed of the user's condition.
[1349] Examples of prompts for generative AI models
[1350] Please generate responses for when a user says, "I've lost my appetite lately, what should I do?" Include a gentle, supportive response for when the user is feeling down.
[1351] answer:
[1352] 1. "It might be due to the heat. Try drinking more water."
[1353] 2. "Don't worry, try eating a little at a time."
[1354] Thus, the present invention is a system that supports the user's health management and provides a sense of security to their family through user voice recognition and emotion analysis.
[1355] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1356] Detailed step-by-step explanation of the program processing
[1357] User registration and initial setup
[1358] Step 1:
[1359] The user powers on the device. The input is pressing the device's power button. The output is the display of the initial setup screen. The device starts up and displays the initial setup screen.
[1360] Step 2:
[1361] The user enters their name, address, emergency contact information, and health information on the initial setup screen. The input is the user's information (name, address, emergency contact information, health information). The output is the display of this information in the device's input fields.
[1362] Step 3:
[1363] The terminal sends the input information to the server. The input is the information entered by the user. The output is the data stream of the information sent to the server. Communication is secure using the HTTPS protocol.
[1364] Step 4:
[1365] The server saves the received information to the database. The input is user information sent from the device. The output is user information saved in the database. Upon successful saving, an initial setup completion message is generated.
[1366] Step 5:
[1367] The server notifies the terminal that the initial setup is complete. The input is the initial setup completion message. The output is the notification to the terminal. The terminal receives the message and displays completion to the user.
[1368] Everyday conversation and data collection
[1369] Step 1:
[1370] The user speaks to the device, saying, "What's the weather like today?" The input is the user's voice. The output is the device receiving the voice input.
[1371] Step 2:
[1372] The device uses a speech recognition engine (e.g., Google Cloud Speech-to-Text) to convert speech into text data. The input is the user's speech data. The output is text data.
[1373] Step 3:
[1374] The terminal sends the converted text data to the server. The input is the text data on the terminal. The output is the text data sent to the server.
[1375] Step 4:
[1376] The server's analysis engine analyzes text data and generates appropriate responses. The input is the text data sent to the server. The output is the generated response data. A generative AI model (e.g., OpenAI GPT-3) is used for analysis.
[1377] Step 5:
[1378] The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions. The input is text data. The output is the emotion analysis result. If the user asks a question in a cheerful voice, it is judged to be a positive emotion.
[1379] Step 6:
[1380] The server adjusts the response based on the analysis results and generates a response. The input is the analysis results and sentiment information. The output is the adjusted response.
[1381] Step 7:
[1382] The server sends the generated response to the terminal. The input is the adjusted response data. The output is the response sent to the terminal.
[1383] Step 8:
[1384] The device displays the received response on the screen and outputs it as audio using a text-to-speech engine (e.g., Amazon Polly). The input is the response data received from the server. The output is displayed to the user and read aloud.
[1385] Counseling and health management
[1386] Step 1:
[1387] The user speaks into the device saying, "I haven't had much of an appetite lately, what should I do?" The input is the user's voice. The output is the device receiving the voice input.
[1388] Step 2:
[1389] The device recognizes the user's voice and converts it into text data. The input is the user's voice data. The output is text data.
[1390] Step 3:
[1391] The terminal sends text data to the server. The input is the text data on the terminal. The output is the text data sent to the server.
[1392] Step 4:
[1393] The server processes text data to analyze health information. The input is the transmitted text data. The output is the health information analysis result.
[1394] Step 5:
[1395] The server's emotion engine analyzes the user's emotions. The input is text data. The output is the emotion analysis result.
[1396] Step 6:
[1397] The server generates appropriate responses based on health information analysis results and emotional information. Using a generative AI model, it generates responses in a gentle tone, such as, "Don't worry, try eating a little at a time." The input is the analysis results and emotional information. The output is the generated response data.
[1398] Step 7:
[1399] The server sends the generated response to the terminal. The input is the generated response data. The output is the response sent to the terminal.
[1400] Step 8:
[1401] The device communicates the received responses to the user and stores the content as health data on the server. Input consists of the response data received from the server and the resulting audio or display. Output is the stored health data.
[1402] Providing information to family members
[1403] Step 1:
[1404] The server automatically generates periodic reports on the user's health status and mood. The input is accumulated health data. The output is automatically generated report data.
[1405] Step 2:
[1406] The server notifies the family of the generated report. Notifications are sent via email or a dedicated app. The input is the generated report data. The output is the notification to the family.
[1407] Response when a problem occurs
[1408] Step 1:
[1409] The user says "My chest hurts" to the device. The input is the user's voice. The output is the device receiving the voice input.
[1410] Step 2:
[1411] The device recognizes speech and converts it into text data. The input is the user's voice data. The output is text data.
[1412] Step 3:
[1413] The terminal sends text data to the server. The input is the text data on the terminal. The output is the transmission of text data to the server.
[1414] Step 4:
[1415] The server's emotion engine determines the urgency. The input is text data. The output is the urgency assessment result.
[1416] Step 5:
[1417] If the server determines the situation is urgent, it will notify family members and designated emergency contacts. Notifications will be sent via SMS or a dedicated app. Input is an emergency notification message. Output is a notification to family members and emergency contacts.
[1418] In this way, the system of the present invention recognizes the user's voice, converts it into text data, generates an appropriate response based on the analysis results, and provides an emotion-appropriate response through an emotion engine. This enables health management for the elderly, information provision to family members, and rapid response to emergencies.
[1419] (Application Example 2)
[1420] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1421] There is a problem in that elderly users and those requiring health management have difficulty choosing appropriate meals that suit their emotions and health condition. Furthermore, there is a lack of means for family members and caregivers to monitor health conditions in real time and respond quickly. Therefore, there is a need for a system that provides accurate meal suggestions based on the user's emotions and health condition, and that can instantly share necessary information with family members and caregivers.
[1422] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[1423] In this invention, the server includes means for recognizing the user's voice and converting it into text data; means for analyzing the converted text data and generating an appropriate response; means for providing the generated response to the user; means for analyzing the conversation content and inferring the user's health condition and mood; means for accumulating the user's health data and generating reports periodically; means for notifying family members of the generated reports; means for detecting the user's emergency and notifying designated contacts; means for suggesting meal menus based on the user's emotional data; and means for automatically placing a food delivery order if the user approves the suggested meal menu. This enables accurate meal suggestions tailored to the user's emotions and health condition, as well as rapid information sharing.
[1424] "Speech recognition" is a technology that converts speech into text data.
[1425] "Text data" refers to string data generated by speech recognition.
[1426] "Emotional data" refers to data obtained as a result of analyzing and evaluating users' emotions.
[1427] "Health status" refers to information indicating the current state of the user's physical and mental health.
[1428] "Mood" refers to information that represents the user's emotional state at that particular time.
[1429] A "meal menu" refers to a selection of appropriate meals suggested based on the user's health condition and mood.
[1430] "Food delivery" is a service that takes orders and delivers meals to customers' homes.
[1431] A "report" is a document that summarizes information about a user's health status and mood.
[1432] An "emergency situation" refers to a serious situation that affects the user's health.
[1433] "Family" refers to the user's relatives and close family members.
[1434] "Notification" is the act of informing other devices or people of specific information.
[1435] "Approval" refers to the act of a user accepting a proposed meal menu.
[1436] A system for realizing this invention includes the following means:
[1437] Speech recognition and analysis
[1438] The user uses a smartphone and speaks into the device, expressing feelings such as "I'm tired today." The user's voice is captured via the microphone and converted into text data using speech recognition technology. For example, the "speech_recognition" library is used for this purpose.
[1439] Emotion analysis
[1440] The converted text data is analyzed using an emotion analysis engine. This engine extracts emotional data from the user's words and evaluates the user's current mood and health status. The "emotion_recognition" library can be used for emotion analysis.
[1441] Predicting health status and mood
[1442] Based on emotional data, the user's health status and mood are inferred. This result is sent to a server and stored in a user-specific database. This allows for the accumulation of user health data, and reports are generated periodically.
[1443] Suggestions for meal menus
[1444] Based on user emotional data and assessments of their health status, meal menus are identified. For example, if a user rates themselves as "tired," a nutritious, energy-boosting meal is suggested. This is selected from a pre-configured database of recommended recipes.
[1445] Automatic ordering and notifications
[1446] Once the user approves the suggested meal menu, the order is automatically sent to the food delivery service. The order details and the user's health information are stored on the server and notified to family members or caregivers via a dedicated app. In the event of an emergency, an immediate notification is sent to the designated contact person.
[1447] Hardware and software used
[1448] Hardware: Smartphones (e.g., iPhone, Android)
[1449] Software: "speech_recognition" library, "emotion_recognition" library
[1450] Examples
[1451] When Ms. Tanaka says "I'm tired today" into her smartphone, the voice is converted into text data and sent to the emotion engine. The engine generates emotion data for "tired" and displays a suggestion: "Here's a recommended energy boost menu: Banana smoothie and chicken sandwich." When Ms. Tanaka approves this, the order is automatically sent and her family is notified.
[1452] Example of a prompt
[1453] "I want to create a food delivery service that analyzes emotions from voice and suggests meals. First, imagine a scenario where a user says, 'I'm tired today.' Based on that, please write a program that suggests the most suitable meal menu."
[1454] This invention aims to realize a system that enables accurate meal suggestions tailored to the user's emotions and health condition, as well as rapid information sharing.
[1455] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1456] Step 1:
[1457] The user speaks into the device.
[1458] Input: User's voice
[1459] Output: Audio data (acoustic signal) from the microphone
[1460] Specific action: The user speaks "I'm tired today" into the smartphone's microphone. The microphone captures this audio.
[1461] Step 2:
[1462] The device performs speech recognition and converts the speech data into text data.
[1463] Input: Audio data (acoustic signal)
[1464] Output: Recognized text data
[1465] Specific operation: The captured audio data is converted into text data by speech recognition software (e.g., the "speech_recognition" library).
[1466] Step 3:
[1467] The device analyzes text data and generates sentiment data.
[1468] Input: Text data
[1469] Output: Emotional data (tired, happy, etc.)
[1470] Specific operation: The generated text data is sent to an emotion analysis engine (e.g., the "emotion_recognition" library), which analyzes the text to determine if the emotion is "tired".
[1471] Step 4:
[1472] The server infers the user's health status and mood from emotional data.
[1473] Input: Sentiment data
[1474] Output: User's health status and mood
[1475] Specific operation: By analyzing emotional data, the user's health status and mood are inferred, and this information is stored in the server's database. For example, emotional data such as "tired" is used to infer a "fatigued state."
[1476] Step 5:
[1477] The server suggests an appropriate meal plan based on the server's estimated health condition and mood.
[1478] Input: User's health status and mood
[1479] Output: Suggested meal menus
[1480] Specific operation: The server selects and suggests a suitable meal menu for the user (e.g., banana smoothie and chicken sandwich) from the recipe information stored in the database.
[1481] Step 6:
[1482] The user approves the suggested meal menu.
[1483] Input: Suggested meal menu
[1484] Output: User approval (confirmation information)
[1485] Specific action: The user reviews the meal menu displayed on their smartphone screen and presses the approve button.
[1486] Step 7:
[1487] The server automatically sends the order to the food delivery service.
[1488] Input: User authorization information
[1489] Output: Food delivery order information
[1490] Specific operation: Based on the user's authorization information, the server automatically sends an order to the food delivery service. The order includes the selected meal menu and the user's delivery address information.
[1491] Step 8:
[1492] The server notifies the family of the generated report.
[1493] Input: User's health status, mood, and dietary information
[1494] Output: Notification report to family
[1495] Specific operation: The server generates reports containing the user's health status, mood, and suggested and approved meal information, and notifies family members via a dedicated app or email.
[1496] The above outlines the processing steps of this system. In each step, data processing or calculations are performed based on the appropriate input data to obtain the final output.
[1497] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1498] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1499] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[1500] [Fourth Embodiment]
[1501] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[1502] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1503] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1504] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[1505] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1506] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1507] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1508] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[1509] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1510] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1511] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1512] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1513] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1514] This invention relates to a dialogue system using speech recognition technology, and more particularly to specific embodiments for alleviating feelings of loneliness among the elderly, monitoring their health status, and providing information to their families.
[1515] User registration and initial setup
[1516] The elderly person (hereinafter referred to as the user) first activates the device (hereinafter referred to as the terminal) obtained from the gacha machine. The terminal displays an initial setup screen, and the user enters their name, address, emergency contact information, and health information. This information is sent from the terminal to the server, and the server stores the information in a database. In this way, user-specific information is accumulated on the server.
[1517] Everyday conversation and data collection
[1518] Users can enjoy natural conversations with the device. For example, if a user asks, "What's the weather like today?", the device recognizes this voice and converts it into text data. Next, the device analyzes the converted text data and generates an appropriate response, such as "It's sunny today." At the same time, the conversation is sent to a server, which analyzes the content to infer the user's health and mood, and stores this information in a database.
[1519] Counseling and health management
[1520] When a user asks a question about their daily health, for example, "I haven't had much of an appetite lately, what should I do?", the device recognizes this voice and converts it into text data. Next, the device analyzes the converted text data, and the AI generates an appropriate answer, such as, "It might be because of the heat. Try drinking more water." This question and answer are also sent to the server and stored as the user's health data.
[1521] Providing information to family members
[1522] The server automatically generates periodic reports on the user's health and mood, and notifies family members. Notifications are sent via email or a dedicated app. This allows family members to understand the user's condition in real time, providing them with peace of mind.
[1523] Response when a problem occurs
[1524] If a user experiences a sudden change in their health or an emergency, for example, by telling the device "My chest hurts," the device will use voice recognition to convert this information into text. If the analysis determines that it is an emergency, the device will send this information to a server. The server will immediately notify family members and registered emergency contacts (e.g., medical institutions). Family members will then be able to receive the notification and respond quickly.
[1525] As a concrete example, consider the case of Mr. Tanaka, an elderly person, using the service. When Mr. Tanaka tells the device, "I didn't sleep well last night," that information is sent to the server, and the analysis results are notified to his family. Mr. Tanaka's son can then review the report and take the necessary actions.
[1526] In this way, the present invention realizes a system that can alleviate the user's feelings of loneliness, effectively manage their health, and provide a sense of security to their family.
[1527] The following describes the processing flow.
[1528] Step 1:
[1529] The user turns on the device (terminal) they obtained from the gacha machine.
[1530] Step 2:
[1531] The device displays the initial setup screen. The user enters their name, address, emergency contact information, and health information.
[1532] Step 3:
[1533] The terminal sends the entered information to the server.
[1534] Step 4:
[1535] The server saves the received information to the database.
[1536] Step 5:
[1537] The device notifies the user that the initial setup is complete.
[1538] Step 6:
[1539] The user begins a casual conversation with the device. For example, they might ask, "What's the weather like today?"
[1540] Step 7:
[1541] The device recognizes the user's voice and converts it into text data.
[1542] Step 8:
[1543] The terminal analyzes the converted text data and generates an appropriate response. It generates the response, "It's sunny today."
[1544] Step 9:
[1545] The terminal provides the user with the generated response.
[1546] Step 10:
[1547] The device sends the conversation content to the server.
[1548] Step 11:
[1549] The server analyzes conversation data to infer the user's health status and mood.
[1550] Step 12:
[1551] The server stores the inferred data in the database.
[1552] Step 13:
[1553] Users ask health-related questions using the device. For example, "I haven't had much of an appetite lately, what should I do?"
[1554] Step 14:
[1555] The device recognizes the user's voice and converts it into text data.
[1556] Step 15:
[1557] The device analyzes the converted text data, and the AI generates an appropriate response. For example, it might generate a response like, "It might be due to the heat. Try drinking more water."
[1558] Step 16:
[1559] The terminal provides the user with the generated response.
[1560] Step 17:
[1561] The device sends the question and answer data to the server.
[1562] Step 18:
[1563] The server collects health data and incorporates it into reports.
[1564] Step 19:
[1565] The server periodically generates reports on the user's health status and mood.
[1566] Step 20:
[1567] The server notifies family members of the generated report. Notifications are sent via email or a dedicated app.
[1568] Step 21:
[1569] The user can use the device to report a sudden change in their health or an emergency. For example, they might say, "My chest hurts."
[1570] Step 22:
[1571] The device recognizes the user's voice and converts it into text data. It determines that an emergency situation has occurred.
[1572] Step 23:
[1573] The device sends emergency information to the server.
[1574] Step 24:
[1575] The server receives the emergency call and notifies the family and designated emergency contacts.
[1576] Step 25:
[1577] Family members receive the notification and begin taking action.
[1578] (Example 1)
[1579] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1580] Challenges include elderly people experiencing loneliness in their daily lives and difficulty in responding early to changes in their health. Furthermore, it is difficult for families to monitor the health and mood changes of elderly people in real time, and a system is needed that can respond appropriately in emergencies.
[1581] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1582] In this invention, the server includes means for recognizing the user's voice and converting it into text data, means for analyzing the converted text data and generating an appropriate response, means for providing the generated response to the user, means for analyzing the conversation content and inferring the user's health status and mood, means for accumulating the user's health data and generating reports periodically, means for notifying family members of the generated reports, and means for detecting user emergencies and notifying designated contacts. This makes it possible to alleviate the user's feelings of loneliness, effectively manage their health, and provide reassurance to their family.
[1583] "Means of recognizing user speech and converting it into text data" refers to technology that takes user-generated speech data as input and converts that speech into text format through digital processing.
[1584] "Means of analyzing converted text data and generating appropriate responses" refers to the process of analyzing text data converted from speech using machine learning and natural language processing technologies to generate responses that are suitable for the user's questions and requests.
[1585] "Means of providing generated answers to users" refers to technologies that provide users with generated text-based answers in audio or text display formats.
[1586] "Methods for analyzing conversation content to infer a user's health status and mood" refers to technologies that collect and analyze conversation content with users as data, and then use that information to infer the user's health status and mood.
[1587] "Means for accumulating user health data and generating periodic reports" refers to technology that accumulates user health data obtained through conversations and other data collection methods, and generates periodic reports based on this data.
[1588] "Means of notifying family members of generated reports" refers to technology that notifies family members of generated health status and mood reports via email or a dedicated app.
[1589] "Means for detecting user emergencies and notifying designated contacts" refers to technologies that use voice recognition technology, sensors, etc., to detect user health emergencies and quickly notify designated contacts (family or medical institutions) of that information.
[1590] Modes for carrying out the invention
[1591] This invention is a dialogue system that recognizes the user's voice and generates appropriate responses, with the aim of alleviating feelings of loneliness among the elderly, monitoring their health status, and providing information to their families. This system is composed of speech recognition technology, natural language processing technology, and generative AI models.
[1592] User registration and initial setup
[1593] The user first starts the device (hereinafter referred to as the terminal) and enters their name, address, emergency contact information, and health information on the initial setup screen. This information is sent from the terminal to the server, which stores the information in a database. The server accumulates user-specific information and uses this information to perform subsequent processing.
[1594] Everyday conversation and data collection
[1595] Users can engage in natural voice interactions with the device. For example, if a user asks, "What's the weather like today?", the device recognizes the speech and converts it into text data. This process utilizes the Google Speech-to-Text API or IBM Watson Speech to Text. The device then analyzes the converted text data and generates an appropriate response using generative AI models such as OpenAI GPT-3 or Microsoft Azure Cognitive Services. The user is then provided with a response such as, "It's sunny today." The conversation is also sent to a server, which analyzes the content to infer the user's health and mood, and stores this information in a database.
[1596] Counseling and health management
[1597] When a user asks a question about their daily health, for example, "I haven't had much of an appetite lately, what should I do?", the device recognizes the speech, converts it into text data, and uses a generative AI model to generate an appropriate answer. The user might receive an answer such as, "It might be because of the heat. Try drinking more water." These questions and answers are also sent to the server and stored in a database.
[1598] Providing information to family members
[1599] The server automatically generates periodic reports on the user's health and mood, and notifies family members. Notifications are sent via email or a dedicated app, using APIs from SendGrid and Twilio. This allows family members to understand the user's status in real time and gain peace of mind.
[1600] Response when a problem occurs
[1601] If a user experiences a sudden change in their health or an emergency, for example, by telling the device "My chest hurts," the device will use voice recognition to convert this information into text. If the analysis determines that it is an emergency, the device will send this information to a server. The server will immediately notify family members and registered emergency contacts, allowing family members to respond quickly.
[1602] As a concrete example, if an elderly user says, "I didn't sleep well last night," that information is sent to the server, and the analysis results are notified to their family. The family can then review the report and take the necessary actions. An example of a prompt message would be, "I haven't been sleeping well lately. What should I do?"
[1603] In this way, the present invention realizes a system that alleviates the user's feelings of loneliness, effectively manages their health, and provides a sense of security to their family.
[1604] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1605] Step 1:
[1606] The user starts up the device. The device automatically displays the initial setup screen. The initial setup screen displays input forms for entering name, address, emergency contact information, and health information.
[1607] Input: User input on the terminal (name, address, emergency contact information, health information)
[1608] Output: Confirmation message displayed on the device upon completion of initial setup.
[1609] Specific operation: The user enters the required information into the input form displayed on the screen and confirms it.
[1610] Step 2:
[1611] The device collects the information entered by the user and sends it to the server using the HTTPS protocol. The server receives the information and stores it in a database.
[1612] Input: Personal information entered by the user (name, address, emergency contact information, health information)
[1613] Output: User information stored in the database
[1614] Specific operation: The terminal converts the input data into JSON format and sends it to the server. The server parses the received data, generates an SQL statement to save it to the database, and executes it.
[1615] Step 3:
[1616] The user asks the device, "What's the weather like today?" The device captures the user's voice using its microphone.
[1617] Input: Voice input from the user (question content)
[1618] Output: Audio file
[1619] Specific operation: The device's microphone captures the user's voice as digital data and temporarily stores it.
[1620] Step 4:
[1621] The device uses the Google Speech-to-Text API to convert the captured audio into text data. The converted text data is stored on the device.
[1622] Input: Audio file
[1623] Output: Text data
[1624] Specific operation: The device sends the audio file to the cloud API, receives the returned text data, and saves it.
[1625] Step 5:
[1626] The terminal analyzes the converted text data and generates an answer using a generative AI model such as OpenAI GPT-3. After an appropriate answer is generated, it is saved in text format.
[1627] Input: Text data (converted question content)
[1628] Output: Text data (generated response)
[1629] Specific operation: The device sends text data to the NLP engine and receives and saves the returned response text.
[1630] Step 6:
[1631] The generated text response is converted to speech using the Google Text-to-Speech API and then sent back to the user.
[1632] Input: Text data (generated response)
[1633] Output: Audio data
[1634] Specific operation: The device sends text data to a cloud API, and the returned audio data is sent to the playback device.
[1635] Step 7:
[1636] The terminal sends the conversation content with the user to the server via the HTTPS protocol. The server analyzes the received data and stores it in a database.
[1637] Input: User and device conversation history
[1638] Output: Conversation history stored in the database
[1639] Specific operation: The terminal converts the conversation data into JSON format and sends it to the server. The server parses the received data, generates SQL statements, and saves them to the database.
[1640] Step 8:
[1641] The user asks a health-related question to the device: "I haven't had much of an appetite lately, what should I do?" The device captures the audio and converts it into text data.
[1642] Input: Voice input from the user (health-related questions)
[1643] Output: Text data
[1644] Specific operation: The user's voice is captured using the device's microphone and converted into text data using the Google Speech-to-Text API.
[1645] Step 9:
[1646] The device analyzes the converted text data and uses a generative AI model to generate an appropriate response. For example, it might generate a response like, "It might be due to the heat. Try drinking more water."
[1647] Input: Text data (converted question content)
[1648] Output: Text data (generated response)
[1649] Specific operation: The device sends text data to the NLP engine and saves the returned response text.
[1650] Step 10:
[1651] The generated response is converted into speech and provided to the user. The user receives a response such as, "It might be because of the heat. Try drinking more water."
[1652] Input: Text data (generated response)
[1653] Output: Audio data
[1654] Specific operation: The device sends text data to the Google Text-to-Speech API and plays the returned audio data.
[1655] Step 11:
[1656] The terminal sends the user's question and answer exchange to the server, which then stores this information in a database.
[1657] Input: User and device question and answer history
[1658] Output: History of questions and answers stored in the database
[1659] Specific operation: The terminal converts the question and answer data into JSON format and sends it to the server. The server parses the data, generates SQL statements, and saves them to the database.
[1660] Step 12:
[1661] The server periodically analyzes information in the database and generates reports on the user's health and mood. These reports utilize libraries such as Python's Pandas library.
[1662] Input: Conversation history and health data stored in the database
[1663] Output: Report on health status and mood
[1664] Specific operation: The server runs a Python script to analyze the stored data and generate periodic reports.
[1665] Step 13:
[1666] The server notifies family members of the generated report. Notifications are sent via email or SMS using APIs such as SendGrid or Twilio.
[1667] Input: Generated report
[1668] Output: Notification to family members (email or SMS)
[1669] Specific operation: The server calls the SendGrid or Twilio APIs to send the generated report to the family.
[1670] Step 14:
[1671] The user reports an emergency, such as "I have chest pain," to the device. The device captures the audio and converts it into text data.
[1672] Input: Voice input from the user (reporting an emergency)
[1673] Output: Text data
[1674] Specific operation: The device captures the audio as digital data and converts it into text data using the Google Speech-to-Text API.
[1675] Step 15:
[1676] The terminal sends the converted text data to the server, which then performs analysis to determine if an emergency has occurred.
[1677] Input: Text data (report of an emergency)
[1678] Output: Emergency Determination Result
[1679] Specific operation: The terminal converts the text data into JSON format and sends it to the server. The server runs an analysis algorithm to determine if it is an emergency.
[1680] Step 16:
[1681] If the server determines an emergency is occurring, it will notify registered emergency contacts. It will also use the Twilio API to notify family members and medical institutions via SMS.
[1682] Input: Emergency situation determination result
[1683] Output: Notification to emergency contacts (SMS)
[1684] Specific operation: The server calls the Twilio API to send a notification containing details of the emergency to the family or healthcare provider.
[1685] The above outlines the specific processing steps of the system. This allows users to alleviate feelings of loneliness, manage their health effectively, and provide reassurance to their families.
[1686] (Application Example 1)
[1687] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1688] There is a need for an effective system to alleviate feelings of loneliness among the elderly, monitor their health, and provide information to their families. In particular, there is a lack of means for the elderly to consult about their health on a daily basis or to receive prompt assistance in emergencies.
[1689] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1690] In this invention, the server includes means for recognizing the user's voice and converting it into text data; means for analyzing the converted text data and generating appropriate responses; means for providing the generated responses to the user; means for analyzing the conversation content and inferring the user's health condition and mood; means for accumulating the user's health data and generating reports periodically; means for notifying family members of the generated reports; means for detecting user emergencies and notifying designated contacts; means for answering health-related questions based on voice input; and means for analyzing the results of the question-answering and using a generative AI model to support the user's health management. This enables support for the daily lives of the elderly, as well as rapid management of their health condition and provision of information to their families.
[1691] "A means of recognizing speech and converting it into text data" refers to a technology that captures the speech spoken by a user as a digital signal, performs analysis processing, and converts that speech into a corresponding string of characters.
[1692] "Means for analyzing converted text data and generating appropriate responses" refers to technology that receives string data converted from speech, understands and analyzes its content, and generates appropriate responses to user questions and statements.
[1693] "Means of providing generated answers to users" refers to means of providing answers obtained through analysis and generation to users in the form of audio or text.
[1694] "Methods for analyzing conversation content to infer a user's health status and mood" refers to methods for evaluating and inferring a user's current health status and emotions by analyzing the user's statements, facial expressions, voice tone, etc., based on the user's dialogue history.
[1695] "Methods for accumulating user health data and generating reports periodically" refers to methods for storing health-related information obtained from users in a database and periodically generating reports summarizing their health status and its trends.
[1696] "Means for notifying family members of generated reports" refers to means of sending and notifying the user's family of the regularly generated user health status reports via email or a dedicated application.
[1697] "Means for detecting user emergencies and notifying designated contacts" refers to a means of quickly notifying pre-registered emergency contacts (e.g., family members or medical institutions) when it is determined that a user is in an emergency situation.
[1698] "A means of providing health-related question-and-answer services based on voice input" refers to a technology that generates responses to specific health-related questions based on voice input from the user.
[1699] "Methods using generative AI models to analyze the results of question-answering and provide support for user health management" refers to methods using generative AI models to analyze user health information collected through voice dialogue and use that data to support user health management.
[1700] This invention relates to a dialogue system for supporting the daily lives of elderly people and effectively managing their health. A specific embodiment includes the following configuration.
[1701] Basic system configuration
[1702] The system of the present invention includes a terminal for user voice input, a server for processing the input voice, and a generative AI model. This configuration enables monitoring of the user's health status and provision of appropriate information.
[1703] Hardware and software usage
[1704] Hardware: Mobile devices such as smartphones that allow the user to input voice commands. This includes a microphone and internet connectivity.
[1705] Software: Python, speech recognition libraries (e.g., SpeechRecognition), natural language processing libraries (e.g., the transformers library).
[1706] Speech recognition and text conversion
[1707] The device has a microphone to receive voice input from the user. When the user speaks, the voice is collected by the device's microphone and converted into text data using a speech recognition library. This conversion can be performed using, for example, the Google Speech Recognition API.
[1708] Text data analysis and response generation
[1709] The converted text data is sent to the server. On the server, a natural language processing library is used to analyze the text data and generate appropriate answers to the user's questions and comments. The generated answers are then provided to the user again in either audio or text format.
[1710] Prediction of health status and data accumulation
[1711] The server has an algorithm built in that analyzes conversation content to infer the user's health status and mood. For example, if a user says, "I haven't had much of an appetite lately, what should I do?", the system will collect information about their health and generate periodic reports.
[1712] Providing information to families and emergency response
[1713] The generated reports are regularly notified to the user's family, for example, via email or a dedicated application. Furthermore, if the user faces an emergency, such as saying "my chest hurts," that information is immediately sent to the server and notified to pre-registered emergency contacts. This notification allows for a swift response.
[1714] Applications of Generative AI Tables
[1715] By using generative AI models, health management support can be further enhanced. When the user's voice input is a specific health question, the generative AI model can provide a more accurate answer. For example, if the prompt is "I haven't had much appetite lately, what should I do?", the generative AI model will generate specific advice such as "It might be because of the heat. Try drinking more water."
[1716] This system will alleviate feelings of loneliness that elderly people experience in their daily lives and allow for effective management of their health. Furthermore, it will enable the rapid provision of information to family members and medical institutions, realizing comprehensive support.
[1717] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1718] Step 1:
[1719] The user speaks into the device. At this time, the device's microphone captures the user's voice. The input is the user's voice data, and the output is prepared voice data for speech recognition. Specifically, the microphone collects the voice data and converts the voice signal into digital data in real time.
[1720] Step 2:
[1721] The device converts the collected audio data into text data using a speech recognition library (e.g., SpeechRecognition). The input is pre-prepared audio data, and the output is string data. Specifically, this data is sent to a cloud-based API, where the audio is analyzed and converted back into text.
[1722] Step 3:
[1723] The converted text data is sent to the server. The server receives the text data and performs analysis using a natural language processing library (e.g., transformers). The input is text data, and the output is the analysis result. Specifically, this analysis prepares the system to generate appropriate answers to the user's questions.
[1724] Step 4:
[1725] The server generates appropriate answers using a generative AI model based on the analysis results. The input is text data containing the user's question, and the output is the answer text. Specifically, it uses a generative AI model (e.g., BERT model) to understand the intent of the question and generate an appropriate answer.
[1726] Step 5:
[1727] The server sends the generated response to the terminal. The input is the generated response text, and the output is the text data sent back to the terminal. Specifically, this response text is sent to the terminal and is ready to be sent back to the user.
[1728] Step 6:
[1729] The terminal provides the user with the received response text. The input is text data received from the server, and the output is the response audio or text provided to the user. Specifically, the terminal speaks the text to the user using a speech output device (e.g., a speaker) or provides it as text information on a display device.
[1730] Step 7:
[1731] The server infers the user's health status and mood based on the conversation content. The input is analyzed text data, and the output is health status or mood evaluation data. Specifically, it uses machine learning algorithms to analyze the text content and infer the user's health status and emotions.
[1732] Step 8:
[1733] The server stores user health data and generates reports periodically. The input is health information obtained from conversations, and the output is a health report generated periodically. Specifically, it stores information in a database and automatically generates reports at regular intervals.
[1734] Step 9:
[1735] The generated report is notified to the family. The input is the generated health report, and the output is notification data sent via email or a dedicated application. Specifically, this involves sending an email with the report attached or notifying the family through a dedicated application.
[1736] Step 10:
[1737] If a user experiences an emergency, the server will notify the designated contacts. The input is text data indicating the emergency, and the output is emergency notification data. Specifically, the server activates a system that immediately sends an alert to the emergency contacts.
[1738] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1739] This invention is a system that recognizes a user's voice, converts it into text data, and then analyzes that data to generate appropriate responses. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it can more accurately predict the user's health condition and mood, and adjust the advice and information provided according to those emotions.
[1740] User registration and initial setup
[1741] The elderly person (hereinafter referred to as the user) turns on the device (hereinafter referred to as the terminal) obtained from the gacha machine and displays the initial setup screen. The user enters their name, address, emergency contact information, and health information, and the terminal sends this information to the server. The server saves the received information in its database and notifies the user via the terminal that the initial setup is complete.
[1742] Everyday conversation and data collection
[1743] Users enjoy natural conversations with the device. For example, they might ask, "What's the weather like today?" The device recognizes the user's voice and converts it into text data. Next, it analyzes the converted text data and generates an appropriate response. During the response generation process, an emotion engine analyzes the user's emotions from their voice and text and generates a response that matches those emotions. For example, if the user asks the question in a cheerful voice, it will provide a positive response such as, "It's sunny today. Have a great day!" At the same time, the conversation content is sent to a server and used to estimate the user's health and mood.
[1744] Counseling and health management
[1745] When a user asks a health-related question, for example, "I haven't had much of an appetite lately, what should I do?", the device recognizes this voice, converts it into text data, and analyzes it. The AI then generates an appropriate response, such as, "It might be because of the heat. Try drinking more water." The emotion engine analyzes the user's emotions and, if the user is feeling down, provides a gentle response such as, "Don't worry, try eating a little at a time." This consultation and response are sent to the server and stored in the health data.
[1746] Providing information to family members
[1747] The server automatically generates periodic reports on the user's health and mood, and notifies family members. Notifications are sent via email or a dedicated app. This allows family members to understand the user's condition in real time and gain peace of mind.
[1748] Response when a problem occurs
[1749] If a user experiences a sudden change in their health or an emergency, they can tell the device, for example, "My chest hurts." The device recognizes this voice and converts it into text data. The emotion engine analyzes the user's urgency, and if it determines it is an emergency, it sends the information to the server. The server immediately notifies family members and designated emergency contacts (e.g., medical institutions). Family members receive the notification and can respond quickly.
[1750] As a concrete example, consider the case of Mr. Tanaka, an elderly person, using this system. When Mr. Tanaka speaks to the device and says, "I didn't sleep well last night," that information is sent to the server, and the analysis results are notified to his family. Furthermore, the emotion engine reads feelings of anxiety and fatigue from the tone of Mr. Tanaka's voice and provides a gentle message such as, "That's worrying. Please rest a little today." Mr. Tanaka's son can then review the report and take any necessary action.
[1751] In this way, the present invention realizes a system that can alleviate the user's feelings of loneliness, effectively manage their health, and provide a sense of security to their family. By combining emotional engines, even more sophisticated responses become possible, enabling the provision of support tailored to the user's emotions.
[1752] The following describes the processing flow.
[1753] Step 1:
[1754] The user turns on the device (terminal) they obtained from the gacha machine.
[1755] Step 2:
[1756] The device displays the initial setup screen. The user enters their name, address, emergency contact information, and health information.
[1757] Step 3:
[1758] The terminal sends the entered information to the server.
[1759] Step 4:
[1760] The server saves the received information to the database.
[1761] Step 5:
[1762] The device notifies the user that the initial setup is complete.
[1763] Step 6:
[1764] The user begins a casual conversation with the device. For example, they might ask, "What's the weather like today?"
[1765] Step 7:
[1766] The device recognizes the user's voice and converts it into text data.
[1767] Step 8:
[1768] The terminal analyzes the converted text data and generates an appropriate response. It generates the response, "It's sunny today."
[1769] Step 9:
[1770] The terminal provides the user with the generated response.
[1771] Step 10:
[1772] The device sends the conversation content to the server.
[1773] Step 11:
[1774] The server analyzes conversation data to infer the user's health status and mood.
[1775] Step 12:
[1776] The server stores the inferred data in the database.
[1777] Step 13:
[1778] The user asks health-related questions to the device. For example, they might ask, "I haven't had much of an appetite lately, what should I do?"
[1779] Step 14:
[1780] The device recognizes the user's voice and converts it into text data.
[1781] Step 15:
[1782] The device analyzes the converted text data, and the AI generates an appropriate response. For example, it might generate a response like, "It might be due to the heat. Try drinking more water."
[1783] Step 16:
[1784] The terminal provides the user with the generated response.
[1785] Step 17:
[1786] The device sends the question and answer data to the server.
[1787] Step 18:
[1788] The server collects health data and incorporates it into reports.
[1789] Step 19:
[1790] The server periodically generates reports on the user's health status and mood.
[1791] Step 20:
[1792] The server notifies family members of the generated report. Notifications are sent via email or a dedicated app.
[1793] Step 21:
[1794] The user can use the device to report a sudden change in their health or an emergency. For example, they might say, "My chest hurts."
[1795] Step 22:
[1796] The device recognizes the user's voice and converts it into text data. It determines that an emergency situation has occurred.
[1797] Step 23:
[1798] The device sends emergency information to the server.
[1799] Step 24:
[1800] The server receives the emergency call and notifies the family and designated emergency contacts.
[1801] Step 25:
[1802] Family members receive the notification and begin taking action.
[1803] Step 26:
[1804] The emotion engine analyzes the user's voice tone and text data to identify their emotions.
[1805] Step 27:
[1806] The device adjusts the tone and content of its responses based on the analysis results of the emotion engine.
[1807] Step 28:
[1808] The device provides users with responses that are tailored to their emotions. For example, if a user is feeling down, it might respond in a gentle tone, "Don't worry, try eating a little at a time."
[1809] Step 29:
[1810] The device sends emotional data to the server.
[1811] Step 30:
[1812] The server analyzes sentiment data and reflects particularly noteworthy changes and trends in a report.
[1813] (Example 2)
[1814] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1815] In modern society, managing the health and daily communication of the elderly are challenges, and for elderly people living alone, loneliness and dealing with changes in their health are particularly difficult. Furthermore, it is difficult for family members and caregivers to understand the elderly person's condition in real time, creating a need for systems that can respond quickly to emergencies. Additionally, a function that analyzes the elderly person's emotions and provides appropriate support based on those emotions is also crucial.
[1816] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1817] In this invention, the server includes means for recognizing the user's voice and converting it into text data; means for analyzing the converted text data and generating an appropriate response; means for providing the generated response to the user; means for analyzing the conversation content and inferring the user's health status and mood; means for accumulating the user's health data and generating reports periodically; means for notifying family members of the generated reports; means for detecting user emergencies and notifying designated contacts; means for analyzing the user's emotions and generating responses appropriate to those emotions; means for analyzing the user's voice and text data to infer emotions and adjusting responses based on those emotions; and means for providing responses in a gentle tone when the user is feeling down. This not only alleviates the user's feelings of loneliness and effectively manages their health, but also provides reassurance to family members and enables support tailored to the user's emotions.
[1818] "Means for recognizing speech and converting it into text data" refers to devices or software that analyze the speech spoken by a user and convert it into a digital text format.
[1819] "Means for analyzing text data and generating appropriate responses" refers to devices or software that understand the content of text data obtained through speech recognition and generate the most appropriate response to a user's question or request.
[1820] "Means of providing generated answers to users" refers to devices or software that have the function of displaying or reading aloud the system-generated answers to the user.
[1821] "Means for analyzing conversation content to infer a user's health status and mood" refers to devices or software that analyze the content and tone of conversations with a user to infer the user's current health status and mental mood.
[1822] "Means for accumulating user health data and generating periodic reports" refers to devices or software that have the function of storing collected user health data and creating periodic reports based on that data.
[1823] "Means of notifying family members of generated reports" refers to devices or software that have the function of sending the created user health report to family members via email or a dedicated app.
[1824] "Means for detecting user emergencies and notifying designated contacts" refers to devices or software that have the function of detecting when a user is in an emergency and automatically sending notifications to pre-designated emergency contacts.
[1825] "Means of analyzing user emotions and generating responses that correspond to those emotions" refers to devices or software that analyze a user's emotions from their tone of voice and word choice, and generate responses that are appropriate to those emotions.
[1826] "Means of analyzing user voice and text data to infer emotions and adjust responses based on those emotions" refers to devices or software that have the function of reading emotions based on the voice and text data spoken by the user and adjusting the system's response content accordingly.
[1827] "Means of providing a gentle tone of response when a user is feeling down" refers to devices or software that have a function to respond in a gentle and friendly tone when the system determines that a user is feeling down.
[1828] This invention relates to a speech recognition system for the elderly that converts the user's voice into text data, analyzes that data, and then uses an emotion engine to recognize emotions and provide appropriate responses and advice. The system uses the following hardware and software.
[1829] User registration and initial setup
[1830] The user first powers on the device and enters their name, address, emergency contact information, and health information using the initial setup screen that appears. The device sends this information to the server. The server stores the received information in a database and notifies the user via the device when the initial setup is complete. A tablet or smart speaker can be used as the device during this process. The transmitted information uses a secure communication protocol such as HTTPS.
[1831] Everyday conversation and data collection
[1832] When a user asks, "What's the weather like today?", the device uses a speech recognition engine (e.g., Google Cloud Speech-to-Text) to convert the speech into text data. This text data is then sent to a server, where the server's analysis engine analyzes its content. A generative AI model (e.g., OpenAI GPT-3) is used for the analysis. During this process, an emotion engine (e.g., IBM Watson Tone Analyzer) analyzes the user's emotions, and for positive voices, it generates a positive response such as, "It's sunny today. Have a nice day!" The response is then sent from the server to the device and communicated to the user via voice or screen display.
[1833] Counseling and health management
[1834] When a user asks, "I haven't had much of an appetite lately, what should I do?", the device uses speech recognition to convert it into text data and send it to the server. The server analyzes the text data, and an emotion engine analyzes the user's emotions. If the user is feeling down, the generative AI model generates a gentle response such as, "Don't worry, try eating a little at a time." The response is provided to the user through the device, and the content is stored on the server as health data.
[1835] Providing information to family members
[1836] The server automatically generates periodic reports on the user's health and mood, and notifies family members. These notifications are sent via email or a dedicated app. Family members receive this information in real time, allowing them to understand the user's condition. This feature helps alleviate feelings of isolation for the user and provides reassurance to family members.
[1837] Response when a problem occurs
[1838] When a user tells the device they are experiencing chest pain, the device recognizes the voice and converts it into text. The server's emotion engine determines the urgency, and if it is determined to be a high-priority emergency, the server immediately notifies family members and designated emergency contacts. The notification is sent via SMS, email, or a dedicated app, allowing family members to respond quickly.
[1839] Examples and prompts for generative AI models
[1840] Specific example
[1841] When an elderly user says, "I didn't sleep well last night," this information is converted into text data by the device and sent to the server. The server's analysis engine and emotion engine analyze this information and notify the family. In addition, the user is provided with a kind message such as, "That's worrying. Please get some rest today," and the family is informed of the user's condition.
[1842] Examples of prompts for generative AI models
[1843] Please generate responses for when a user says, "I've lost my appetite lately, what should I do?" Include a gentle, supportive response for when the user is feeling down.
[1844] answer:
[1845] 1. "It might be due to the heat. Try drinking more water."
[1846] 2. "Don't worry, try eating a little at a time."
[1847] Thus, the present invention is a system that supports the user's health management and provides a sense of security to their family through user voice recognition and emotion analysis.
[1848] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1849] Detailed step-by-step explanation of the program processing
[1850] User registration and initial setup
[1851] Step 1:
[1852] The user powers on the device. The input is pressing the device's power button. The output is the display of the initial setup screen. The device starts up and displays the initial setup screen.
[1853] Step 2:
[1854] The user enters their name, address, emergency contact information, and health information on the initial setup screen. The input is the user's information (name, address, emergency contact information, health information). The output is the display of this information in the device's input fields.
[1855] Step 3:
[1856] The terminal sends the input information to the server. The input is the information entered by the user. The output is the data stream of the information sent to the server. Communication is secure using the HTTPS protocol.
[1857] Step 4:
[1858] The server saves the received information to the database. The input is user information sent from the device. The output is user information saved in the database. Upon successful saving, an initial setup completion message is generated.
[1859] Step 5:
[1860] The server notifies the terminal that the initial setup is complete. The input is the initial setup completion message. The output is the notification to the terminal. The terminal receives the message and displays completion to the user.
[1861] Everyday conversation and data collection
[1862] Step 1:
[1863] The user speaks to the device, saying, "What's the weather like today?" The input is the user's voice. The output is the device receiving the voice input.
[1864] Step 2:
[1865] The device uses a speech recognition engine (e.g., Google Cloud Speech-to-Text) to convert speech into text data. The input is the user's speech data. The output is text data.
[1866] Step 3:
[1867] The terminal sends the converted text data to the server. The input is the text data on the terminal. The output is the text data sent to the server.
[1868] Step 4:
[1869] The server's analysis engine analyzes text data and generates appropriate responses. The input is the text data sent to the server. The output is the generated response data. A generative AI model (e.g., OpenAI GPT-3) is used for analysis.
[1870] Step 5:
[1871] The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions. The input is text data. The output is the emotion analysis result. If the user asks a question in a cheerful voice, it is judged to be a positive emotion.
[1872] Step 6:
[1873] The server adjusts the response based on the analysis results and generates a response. The input is the analysis results and sentiment information. The output is the adjusted response.
[1874] Step 7:
[1875] The server sends the generated response to the terminal. The input is the adjusted response data. The output is the response sent to the terminal.
[1876] Step 8:
[1877] The device displays the received response on the screen and outputs it as audio using a text-to-speech engine (e.g., Amazon Polly). The input is the response data received from the server. The output is displayed to the user and read aloud.
[1878] Counseling and health management
[1879] Step 1:
[1880] The user speaks into the device saying, "I haven't had much of an appetite lately, what should I do?" The input is the user's voice. The output is the device receiving the voice input.
[1881] Step 2:
[1882] The device recognizes the user's voice and converts it into text data. The input is the user's voice data. The output is text data.
[1883] Step 3:
[1884] The terminal sends text data to the server. The input is the text data on the terminal. The output is the text data sent to the server.
[1885] Step 4:
[1886] The server processes text data to analyze health information. The input is the transmitted text data. The output is the health information analysis result.
[1887] Step 5:
[1888] The server's emotion engine analyzes the user's emotions. The input is text data. The output is the emotion analysis result.
[1889] Step 6:
[1890] The server generates appropriate responses based on health information analysis results and emotional information. Using a generative AI model, it generates responses in a gentle tone, such as, "Don't worry, try eating a little at a time." The input is the analysis results and emotional information. The output is the generated response data.
[1891] Step 7:
[1892] The server sends the generated response to the terminal. The input is the generated response data. The output is the response sent to the terminal.
[1893] Step 8:
[1894] The device communicates the received responses to the user and stores the content as health data on the server. Input consists of the response data received from the server and the resulting audio or display. Output is the stored health data.
[1895] Providing information to family members
[1896] Step 1:
[1897] The server automatically generates periodic reports on the user's health status and mood. The input is accumulated health data. The output is automatically generated report data.
[1898] Step 2:
[1899] The server notifies the family of the generated report. Notifications are sent via email or a dedicated app. The input is the generated report data. The output is the notification to the family.
[1900] Response when a problem occurs
[1901] Step 1:
[1902] The user says "My chest hurts" to the device. The input is the user's voice. The output is the device receiving the voice input.
[1903] Step 2:
[1904] The device recognizes speech and converts it into text data. The input is the user's voice data. The output is text data.
[1905] Step 3:
[1906] The terminal sends text data to the server. The input is the text data on the terminal. The output is the transmission of text data to the server.
[1907] Step 4:
[1908] The server's emotion engine determines the urgency. The input is text data. The output is the urgency assessment result.
[1909] Step 5:
[1910] If the server determines the situation is urgent, it will notify family members and designated emergency contacts. Notifications will be sent via SMS or a dedicated app. Input is an emergency notification message. Output is a notification to family members and emergency contacts.
[1911] In this way, the system of the present invention recognizes the user's voice, converts it into text data, generates an appropriate response based on the analysis results, and provides an emotion-appropriate response through an emotion engine. This enables health management for the elderly, information provision to family members, and rapid response to emergencies.
[1912] (Application Example 2)
[1913] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1914] There is a problem in that elderly users and those requiring health management have difficulty choosing appropriate meals that suit their emotions and health condition. Furthermore, there is a lack of means for family members and caregivers to monitor health conditions in real time and respond quickly. Therefore, there is a need for a system that provides accurate meal suggestions based on the user's emotions and health condition, and that can instantly share necessary information with family members and caregivers.
[1915] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[1916] In this invention, the server includes means for recognizing the user's voice and converting it into text data; means for analyzing the converted text data and generating an appropriate response; means for providing the generated response to the user; means for analyzing the conversation content and inferring the user's health condition and mood; means for accumulating the user's health data and generating reports periodically; means for notifying family members of the generated reports; means for detecting the user's emergency and notifying designated contacts; means for suggesting meal menus based on the user's emotional data; and means for automatically placing a food delivery order if the user approves the suggested meal menu. This enables accurate meal suggestions tailored to the user's emotions and health condition, as well as rapid information sharing.
[1917] "Speech recognition" is a technology that converts speech into text data.
[1918] "Text data" refers to string data generated by speech recognition.
[1919] "Emotional data" refers to data obtained as a result of analyzing and evaluating users' emotions.
[1920] "Health status" refers to information indicating the current state of the user's physical and mental health.
[1921] "Mood" refers to information that represents the user's emotional state at that particular time.
[1922] A "meal menu" refers to a selection of appropriate meals suggested based on the user's health condition and mood.
[1923] "Food delivery" is a service that takes orders and delivers meals to customers' homes.
[1924] A "report" is a document that summarizes information about a user's health status and mood.
[1925] An "emergency situation" refers to a serious situation that affects the user's health.
[1926] "Family" refers to the user's relatives and close family members.
[1927] "Notification" is the act of informing other devices or people of specific information.
[1928] "Approval" refers to the act of a user accepting a proposed meal menu.
[1929] A system for realizing this invention includes the following means:
[1930] Speech recognition and analysis
[1931] The user uses a smartphone and speaks into the device, expressing feelings such as "I'm tired today." The user's voice is captured via the microphone and converted into text data using speech recognition technology. For example, the "speech_recognition" library is used for this purpose.
[1932] Emotion analysis
[1933] The converted text data is analyzed using an emotion analysis engine. This engine extracts emotional data from the user's words and evaluates the user's current mood and health status. The "emotion_recognition" library can be used for emotion analysis.
[1934] Predicting health status and mood
[1935] Based on emotional data, the user's health status and mood are inferred. This result is sent to a server and stored in a user-specific database. This allows for the accumulation of user health data, and reports are generated periodically.
[1936] Suggestions for meal menus
[1937] Based on user emotional data and assessments of their health status, meal menus are identified. For example, if a user rates themselves as "tired," a nutritious, energy-boosting meal is suggested. This is selected from a pre-configured database of recommended recipes.
[1938] Automatic ordering and notifications
[1939] Once the user approves the suggested meal menu, the order is automatically sent to the food delivery service. The order details and the user's health information are stored on the server and notified to family members or caregivers via a dedicated app. In the event of an emergency, an immediate notification is sent to the designated contact person.
[1940] Hardware and software used
[1941] Hardware: Smartphones (e.g., iPhone, Android)
[1942] Software: "speech_recognition" library, "emotion_recognition" library
[1943] Examples
[1944] When Ms. Tanaka says "I'm tired today" into her smartphone, the voice is converted into text data and sent to the emotion engine. The engine generates emotion data for "tired" and displays a suggestion: "Here's a recommended energy boost menu: Banana smoothie and chicken sandwich." When Ms. Tanaka approves this, the order is automatically sent and her family is notified.
[1945] Example of a prompt
[1946] "I want to create a food delivery service that analyzes emotions from voice and suggests meals. First, imagine a scenario where a user says, 'I'm tired today.' Based on that, please write a program that suggests the most suitable meal menu."
[1947] This invention aims to realize a system that enables accurate meal suggestions tailored to the user's emotions and health condition, as well as rapid information sharing.
[1948] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1949] Step 1:
[1950] The user speaks into the device.
[1951] Input: User's voice
[1952] Output: Audio data (acoustic signal) from the microphone
[1953] Specific action: The user speaks "I'm tired today" into the smartphone's microphone. The microphone captures this audio.
[1954] Step 2:
[1955] The device performs speech recognition and converts the speech data into text data.
[1956] Input: Audio data (acoustic signal)
[1957] Output: Recognized text data
[1958] Specific operation: The captured audio data is converted into text data by speech recognition software (e.g., the "speech_recognition" library).
[1959] Step 3:
[1960] The device analyzes text data and generates sentiment data.
[1961] Input: Text data
[1962] Output: Emotional data (tired, happy, etc.)
[1963] Specific operation: The generated text data is sent to an emotion analysis engine (e.g., the "emotion_recognition" library), which analyzes the text to determine if the emotion is "tired".
[1964] Step 4:
[1965] The server infers the user's health status and mood from emotional data.
[1966] Input: Sentiment data
[1967] Output: User's health status and mood
[1968] Specific operation: By analyzing emotional data, the user's health status and mood are inferred, and this information is stored in the server's database. For example, emotional data such as "tired" is used to infer a "fatigued state."
[1969] Step 5:
[1970] The server suggests an appropriate meal plan based on the server's estimated health condition and mood.
[1971] Input: User's health status and mood
[1972] Output: Suggested meal menus
[1973] Specific operation: The server selects and suggests a suitable meal menu for the user (e.g., banana smoothie and chicken sandwich) from the recipe information stored in the database.
[1974] Step 6:
[1975] The user approves the suggested meal menu.
[1976] Input: Suggested meal menu
[1977] Output: User approval (confirmation information)
[1978] Specific action: The user reviews the meal menu displayed on their smartphone screen and presses the approve button.
[1979] Step 7:
[1980] The server automatically sends the order to the food delivery service.
[1981] Input: User authorization information
[1982] Output: Food delivery order information
[1983] Specific operation: Based on the user's authorization information, the server automatically sends an order to the food delivery service. The order includes the selected meal menu and the user's delivery address information.
[1984] Step 8:
[1985] The server notifies the family of the generated report.
[1986] Input: User's health status, mood, and dietary information
[1987] Output: Notification report to family
[1988] Specific operation: The server generates reports containing the user's health status, mood, and suggested and approved meal information, and notifies family members via a dedicated app or email.
[1989] The above outlines the processing steps of this system. In each step, data processing or calculations are performed based on the appropriate input data to obtain the final output.
[1990] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1991] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1992] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[1993] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1994] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. In the upper and lower directions of the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. Also, the upper side of the concentric circles is where "pleasant" emotions are located, and the lower side is where "unpleasant" emotions are located. In this way, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[1995] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[1996] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[1997] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[1998] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[1999] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[2000] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[2001] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[2002] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portab...
Claims
1. A means of recognizing a user's voice and converting it into text data, A means of analyzing converted text data and generating an appropriate response, A means of providing the generated answer to the user, A method for analyzing conversation content to infer the user's health status and mood, A means of accumulating user health data and generating reports periodically, A means of notifying family members of the generated report, A means of detecting a user emergency and notifying designated contacts, A system that includes this.
2. The system according to claim 1, which detects changes in the user's health status and mood by analyzing data collected from the user's conversation.
3. The system according to claim 1, wherein the generated response includes advice regarding the user's health.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A