system
A system with speech recognition and biometric monitoring capabilities addresses the communication and emergency response challenges faced by elderly individuals, providing a safe and secure environment.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-01
- Publication Date
- 2026-04-13
AI Technical Summary
Elderly people living alone face challenges such as a lack of communication and delayed responses to emergencies, leading to feelings of loneliness and increased risk in health abnormalities.
A system equipped with speech recognition, biometric data acquisition, and emergency notification capabilities that facilitates communication and rapid response through voice interaction and real-time monitoring of biometric data.
Enables natural communication and rapid emergency response for elderly individuals, ensuring their safety and well-being.
Smart Images

Figure 2026063727000001_ABST
Abstract
Description
Technical Field
[0005]
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In recent years, due to the progress of an aging society, the increase in the number of elderly people living alone has become a problem. Elderly people living alone have problems such as a lack of communication in daily life and difficulty in early response to sudden health abnormalities. If such a situation is left unattended, problems such as an increase in loneliness and a delay in response in an emergency will occur. Therefore, there is a need for an effective system to watch over elderly people living alone, promote communication, and respond promptly in an emergency.
Means for Solving the Problems
[0005] To solve the above problems, the present invention provides the following means. First, the terminal includes a speech recognition means that acquires speech in real time and converts the speech into text. The terminal also includes a transmission means that sends the converted text data to a server, and the server includes an analysis means that analyzes the received text data and generates an appropriate response. Furthermore, the system includes a response transmission means that sends the generated response from the server to the terminal, and a speech synthesis means that converts the response received by the terminal into speech and outputs it to the user. This method enables smooth communication between the elderly and the system.
[0006] In addition, the terminal is equipped with biometric data acquisition means to continuously acquire the user's biometric data (heart rate, body temperature, movement, etc.), and has an anomaly detection means to monitor this data in real time and detect abnormalities. If an anomaly is detected, it has an emergency notification means to send an emergency notification to the server, and the server includes a communication means to receive the emergency notification and send an emergency message to a pre-configured emergency contact. This system enables a rapid response in the event of a health abnormality and ensures the safety of elderly people living alone.
[0007] These measures facilitate communication in daily life and enable a rapid response in the event of sudden health problems, thereby providing an environment in which elderly people living alone can live with peace of mind.
[0008] A "terminal" is a device that interacts directly with the user and performs functions such as voice input, text conversion, data transmission, voice output, and biometric data acquisition.
[0009] A "server" is a central processing unit that receives data transmitted from terminals, analyzes it, generates necessary responses and notifications, and sends them back to the terminals.
[0010] "User" refers to an individual who uses this system, and is particularly targeted at elderly people living alone.
[0011] "Voice recognition means" refers to technology that converts voice signals into text data, and is installed in the terminal.
[0012] "Transmission method" refers to the communication technology used by a terminal to send converted text data and emergency notifications to a server.
[0013] "Analysis means" refers to the technology used by a server to analyze text data received and generate an appropriate response.
[0014] "Response transmission means" refers to communication technology used to send responses generated by a server to a terminal.
[0015] "Speech synthesis means" refers to technology that converts text data into speech and outputs it to the user.
[0016] "Methods for acquiring biometric data" refers to technologies that enable a device to continuously acquire biometric data such as the user's heart rate, body temperature, and movement.
[0017] "Anomaly detection means" refers to technology that monitors acquired biometric data in real time and detects anomalies.
[0018] "Emergency notification means" refers to communication technology used to send emergency notifications to a server when an anomaly is detected.
[0019] "Communication method" refers to the technology that allows a server to receive an emergency notification and send an emergency message to pre-configured emergency contacts.
[0020] A "User ID" refers to unique identification information used to identify a user.
[0021] "Audio data" refers to the digital signal obtained by converting the user's voice into a digital signal.
[0022] "Text data" refers to character information converted by speech recognition technology.
[0023] An "emergency message" refers to the notification content sent from the server to emergency contacts when an anomaly is detected.
[0024] "Emergency contact information" refers to the contact information of family members, medical institutions, etc. that need to be contacted when an abnormality is detected.
Brief Explanation of Drawings
[0025] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14]This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]
[0026] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0027] First, let's explain the terminology used in the following explanation.
[0028] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), and APU (Accelerated Processing Unit).
[0029] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0030] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0031] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0032] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0033] [First Embodiment]
[0034] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0035] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0036] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0037] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0038] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0039] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0040] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0041] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0042] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0043] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0044] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0045] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0046] This invention relates to a system for monitoring elderly people living alone, facilitating communication with them using voice and biometric data, and enabling a rapid response in emergencies. Embodiments of this invention are described in detail below.
[0047] Speech recognition and input processing
[0048] terminal
[0049] When a user speaks into the device, the device acquires the voice through its built-in microphone. The acquired voice data is converted into text data by a speech recognition system. This text data is then sent to the server through the device's transmission system.
[0050] server
[0051] The server receives text data sent from the terminal. The received data is analyzed by an analysis device, and an appropriate response is generated based on the user's intent. The generated response is then sent back to the terminal via the server's response transmission device.
[0052] terminal
[0053] The terminal receives responses from the server, converts them into speech data using speech synthesis technology, and outputs them to the user through a speaker. This allows elderly people to interact with the system in a natural way.
[0054] Examples of conversation partner functions
[0055] User: "What kind of day is it today?"
[0056] Terminal: Acquires audio and converts it into text data, such as "What kind of day is it today?", using speech recognition technology.
[0057] Terminal: Sends text data to the server.
[0058] Server: Analyzes the received text data and generates a response such as "Today is not a special day."
[0059] Server: Sends the generated response to the terminal.
[0060] Terminal: Converts the response data into speech and outputs "Today is not a special day" to the user.
[0061] Emergency contact function
[0062] terminal
[0063] The device has built-in biosensors that continuously acquire the user's heart rate, body temperature, and movement. This biometric data is constantly monitored, and if an abnormality is detected, an emergency notification is generated.
[0064] server
[0065] Emergency notifications are sent from the terminal to the server. The server analyzes the emergency notification and sends an emergency message to pre-registered emergency contacts. The emergency message includes the user ID, terminal ID, and abnormal data.
[0066] Examples of emergency contact functions
[0067] Device: If the user falls and becomes motionless, it detects abnormal heart rate and unresponsive movement.
[0068] Terminal: Immediately send an emergency notification to the server.
[0069] Server: Receives emergency notification and sends SMS to pre-configured family contacts. "User has collapsed. Please respond immediately."
[0070] User (recipient): The family receives the emergency message and immediately begins taking action.
[0071] In this way, the system of the present invention can ensure the safety and communication of elderly people living alone.
[0072] The following describes the processing flow.
[0073] Speech recognition and input processing
[0074] Step 1:
[0075] Terminal: The terminal acquires the user's voice in real time through its built-in microphone.
[0076] Step 2:
[0077] Terminal: Transmits the acquired audio data to a speech recognition device (e.g., a speech recognition engine) and converts the audio data into text data.
[0078] Step 3:
[0079] Terminal: Sends the converted text data to the server using a transmission method. HTTP requests are often used for transmission.
[0080] Conversation partner function
[0081] Step 4:
[0082] Server: Analyzes the received text data using parsing tools (such as a natural language processing engine).
[0083] Step 5:
[0084] Server: Generates an appropriate response based on the analysis results. The response may include fixed statements or information obtained from external APIs.
[0085] Step 6:
[0086] Server: Sends the generated response to the terminal via a response transmission method. JSON format data is often used for transmission.
[0087] Step 7:
[0088] Terminal: Converts the received response text data into speech data using a text-to-speech engine.
[0089] Step 8:
[0090] Terminal: Outputs audio data to the user through the speaker.
[0091] Emergency contact function
[0092] Step 1:
[0093] Device: The device continuously acquires biometric data such as heart rate, body temperature, and movement through built-in biosensors.
[0094] Step 2:
[0095] Terminal: Uses an anomaly detection method to monitor acquired biometric data in real time and detect abnormal values.
[0096] Step 3:
[0097] Terminal: If an anomaly is detected, an emergency notification will be generated. This notification will include the type of anomaly and the user's identification information.
[0098] Step 4:
[0099] Terminal: Sends the generated emergency notification to the server using the emergency notification method.
[0100] Step 5:
[0101] Server: Receives emergency notifications and retrieves a pre-configured list of emergency contacts.
[0102] Step 6:
[0103] Server: Sends an emergency message to emergency contacts based on the information included in the emergency notification. Email, SMS, and phone APIs are often used for sending messages.
[0104] Step 7:
[0105] User (recipient): The family or medical institution will review the received emergency message and begin taking immediate action.
[0106] The above is the program's processing flow, including the specific actions performed at each processing step. Providing such a detailed explanation allows for a clearer understanding of how the entire system works.
[0107] (Example 1)
[0108] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0109] In modern society, elderly people living alone face challenges such as a lack of daily communication and the inability to respond quickly in emergencies. This has led to feelings of loneliness among the elderly and delays in emergency response. The present invention aims to provide a system that enables elderly people to engage in daily conversations in a natural manner and further enables a rapid response in emergencies using biometric data.
[0110] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0111] In this invention, the server includes: speech recognition means for acquiring voice from a terminal in real time and converting the voice into text; transmission means for the terminal to send the converted text data to the server; analysis means for the server to analyze the received text data and generate an appropriate response using natural language processing means; response transmission means for the server to send the generated response to the terminal; speech synthesis means for the terminal to convert the received response into voice and output it to the user; anomaly detection means for monitoring biometric data in real time and detecting anomalies; emergency notification means for sending an emergency notification to the server when an anomaly is detected; and communication means for the server to receive the emergency notification and automatically send an emergency message to a pre-configured emergency contact. This enables elderly people to communicate naturally using voice on a daily basis, and also enables a rapid response in emergencies using biometric data.
[0112] A "terminal" is a device that functions as an interface with the user, collecting, converting, transmitting, and outputting responses for voice data.
[0113] "Speech recognition means" refers to a technology that converts speech into text data, is built into a terminal, and has the function of recognizing the user's voice in real time.
[0114] "Transmission means" refers to the communication means used by a terminal to send converted text data to a server, and includes Wi-Fi modules and internet connection means.
[0115] A "server" is a central control unit that receives data sent from terminals and performs processing, analysis, and response generation.
[0116] "Natural language processing means" refers to technology that analyzes received text data, understands the user's intent, and generates an appropriate response.
[0117] "Analysis means" refers to the technology used by a server to process received text data and generate an appropriate response using natural language processing means.
[0118] "Response transmission means" refers to communication technology used to send responses generated by a server to a terminal.
[0119] "Speech synthesis means" refers to a technology that converts text data received from a server into speech data and outputs a speech response to the user.
[0120] "Methods for acquiring biometric data" refers to sensor technology that continuously acquires data on the user's heart rate, body temperature, and movement.
[0121] "Anomaly detection means" refers to technology that monitors biological data in real time and detects anomalies.
[0122] An "emergency notification system" refers to a technology that immediately sends an emergency notification to a server when an anomaly is detected.
[0123] "Communication method" refers to the technology in which a server receives an emergency notification and automatically sends an emergency message to pre-configured emergency contacts.
[0124] A "generative AI model" refers to artificial intelligence technology used to generate natural dialogue and appropriate responses in abnormal situations.
[0125] A "prompt statement" refers to an input statement used to give instructions to a generative AI model and is used to generate a specific response.
[0126] This invention relates to a system for monitoring elderly people living alone. It facilitates communication with the elderly using voice and biometric data, and enables a rapid response in emergencies. The embodiments of this invention are described in detail below.
[0127] Speech recognition and input processing
[0128] terminal
[0129] When a user speaks into the device, the device acquires the voice through its built-in microphone. The acquired voice data is converted into text data using speech recognition technology (e.g., Google® Cloud Speech-to-Text API). This text data is then sent to the server via the device's transmission module (Wi-Fi module).
[0130] server
[0131] The server uses a Python framework (e.g., Flask or Django) to receive text data sent from the terminal. The received data is parsed by a natural language processing tool (e.g., spaCy) to generate an appropriate response based on the user's intent. The generated response is then sent back to the terminal via the server's response sending mechanism.
[0132] terminal
[0133] The device uses the Google Cloud Text-to-Speech API to convert responses received from the server into audio data, which is then output to the user through the speaker. This allows elderly users to interact with the system in a natural way.
[0134] Specific example
[0135] User: "What kind of day is it today?"
[0136] Terminal: Acquires audio and converts it into text data, such as "What kind of day is it today?", using speech recognition technology.
[0137] Terminal: Sends text data to the server.
[0138] Server: Analyzes the received text data and generates a response such as "Today is not a special day."
[0139] Server: Sends the generated response to the terminal.
[0140] Terminal: Converts the response data into speech and outputs "Today is not a special day" to the user.
[0141] Emergency contact function
[0142] terminal
[0143] The device has built-in biosensors (heart rate sensor, body temperature sensor, motion sensor, etc.) that continuously acquire the user's heart rate, body temperature, and movement. This biometric data is constantly monitored by a microcontroller (e.g., ESP32), and if an anomaly is detected, an emergency notification is generated.
[0144] server
[0145] Emergency notifications are sent from the device to the server. The server analyzes the emergency notification and sends an emergency message to pre-registered emergency contacts using the Twilio API. The emergency message includes the user ID, device ID, and anomaly data.
[0146] Specific example
[0147] If a user collapses and becomes motionless, the system detects abnormal heart rate and unresponsive movement.
[0148] Terminal: Immediately send an emergency notification to the server.
[0149] Server: Receives emergency notification and sends an SMS to pre-configured family contacts using the Twilio API. "User has collapsed. Please respond quickly."
[0150] Examples of prompt statements
[0151] The following are specific examples of prompt statements used in generative AI models.
[0152] In the case of an elderly monitoring system:
[0153] Create text to generate natural conversations with users.
[0154] For example, generate a response to the user's question, "What kind of day is it today?"
[0155] By using these prompts, it becomes possible to provide a system that enables natural interaction with users and allows for rapid response in emergencies.
[0156] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0157] Program processing for speech recognition and input processing
[0158] Step 1:
[0159] terminal
[0160] The user speaks into the device, and the voice is captured through the built-in microphone. This voice data is the input. The captured voice data is sent in real time to a speech recognition system (e.g., Google Cloud Speech-to-Text API) and converted into text data. The data processing here is the conversion from voice to text, and the output is text data.
[0161] Specific actions
[0162] User: "What kind of day is it today?"
[0163] Device: The built-in microphone captures audio and sends it to the speech recognition API.
[0164] Step 2:
[0165] terminal
[0166] The system acquires text data converted by speech recognition and sends it to a server. The input here is the converted text data, and the output is the text data sent to the server. Transmission takes place via the internet through a Wi-Fi module.
[0167] Specific actions
[0168] Terminal: The device sends the text data "What kind of day is it today?" obtained from the voice recognition device to the server via the Wi-Fi module.
[0169] Step 3:
[0170] server
[0171] The server receives text data sent from the terminal using a Python framework (e.g., Flask or Django). The input here is text data, which is parsed using natural language processing tools (e.g., spaCy). The output is the appropriate response resulting from the parsing.
[0172] Specific actions
[0173] Server: Analyzes the text data "What kind of day is it today?" and uses a natural language processing library to understand the user's intent.
[0174] Step 4:
[0175] server
[0176] Based on the analysis results, an appropriate response is generated using a generative AI model (e.g., GPT-3®). The input is the analysis results, and the output is the generated response text.
[0177] Specific actions
[0178] Server: Generates responses such as "Today is not a special day" using an AI model based on the analysis results from a natural language processing library.
[0179] Step 5:
[0180] server
[0181] The generated response is sent to the terminal via the response transmission means. The input is the generated response text, and the output is the text data sent to the terminal.
[0182] Specific actions
[0183] Server: Sends the response text "Today is not a special day" to the terminal.
[0184] Step 6:
[0185] terminal
[0186] The system converts the response received from the server into speech data using a text-to-speech synthesis tool (e.g., Google Cloud Text-to-Speech API). The input is the response text, and the output is speech data. This is output to the user through the built-in speaker.
[0187] Specific actions
[0188] Terminal: Receives the response text "Today is not a special day," converts it into speech data using a speech synthesis method, and outputs it through the speaker.
[0189] Processing of the emergency contact function program
[0190] Step 1:
[0191] terminal
[0192] The device's built-in biosensors (heart rate sensor, body temperature sensor, motion sensor, etc.) continuously acquire the user's biometric data. The input is data from the biosensors, and the output is the continuously acquired biometric data. This data is sent to a microcontroller (e.g., ESP32) for monitoring.
[0193] Specific actions
[0194] Device: The user's heart rate sensor measures a heart rate of 60 BPM.
[0195] Step 2:
[0196] terminal
[0197] The system monitors acquired biometric data and detects anomalies if abnormal values exceed a threshold. The input is continuously acquired biometric data, which is analyzed by an anomaly detection mechanism. The output indicates whether or not an anomaly exists.
[0198] Specific actions
[0199] Terminal: If a heart rate of 30 BPM (below the abnormal capture threshold) is detected, it will be flagged as abnormal.
[0200] Step 3:
[0201] terminal
[0202] When an anomaly is detected, an emergency notification is immediately generated. The input is the result of the anomaly detection, and the output is the emergency notification data. The generated emergency notification includes the user ID and anomaly data.
[0203] Specific actions
[0204] The device detects an abnormal heart rate and generates data that reads "Emergency notification: User ID 123, heart rate 30 BPM".
[0205] Step 4:
[0206] terminal
[0207] The generated emergency notification is sent to the server. The input is the emergency notification data, and the output is the emergency notification sent to the server. Transmission is performed via the Wi-Fi module.
[0208] Specific actions
[0209] Device: Emergency notifications are sent to the server via Wi-Fi.
[0210] Step 5:
[0211] server
[0212] This system receives emergency notifications and automatically sends emergency messages to pre-registered emergency contacts via the Twilio API. The input is the received emergency notification data, and the output is the sent emergency message. The emergency message includes user ID, device ID, and anomaly data.
[0213] Specific actions
[0214] Server: Upon receiving an emergency notification, it sends an SMS message using the Twilio API stating, "A user has collapsed. Please respond quickly."
[0215] As described above, the system of the present invention enables natural communication with the elderly using voice and biometric data, as well as a rapid response in emergencies.
[0216] (Application Example 1)
[0217] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0218] The problem that this invention aims to solve is to provide an environment in which elderly people living alone can live safely and securely. In particular, it aims to enable rapid response in emergencies by linking voice recognition and biometric data, and to monitor changes in the elderly person's physical condition and abnormal situations in real time. Another challenge is to improve communication with the elderly person through a user-friendly interface.
[0219] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0220] In this invention, the server includes an analysis means for analyzing text data and generating an appropriate response, a communication means for receiving emergency notifications and sending emergency messages to pre-configured emergency contacts, and a means for generating responses using an AI model. This enables elderly people living alone to communicate naturally by voice and to respond quickly in emergencies.
[0221] A "terminal" is a device used by a user that has the function of acquiring voice data and sending it to a server, as well as the function of acquiring biometric data.
[0222] "Voice recognition means" refers to technology that converts speech acquired by a terminal into text data.
[0223] "Transmission method" refers to the function by which the terminal sends the converted text data to the server.
[0224] "Analysis means" refers to a technology that analyzes text data received by a server and generates an appropriate response based on the user's intent.
[0225] "Response transmission means" refers to the function by which the server sends the generated response to the terminal.
[0226] "Speech synthesis means" refers to a technology that converts responses received by a terminal from a server into speech data and outputs it to the user as speech.
[0227] "Means of acquiring biometric data" refers to a function in which a device continuously acquires biometric data such as the user's heart rate, body temperature, and movement.
[0228] An "anomaly detection method" is a technology that monitors acquired biometric data in real time and generates an emergency notification when an anomaly is detected.
[0229] An "emergency notification mechanism" is a function that allows a terminal to send an emergency notification to a server when it detects an anomaly.
[0230] "Communication method" refers to a technology where a server receives an emergency notification and sends an emergency message to pre-configured emergency contacts.
[0231] An "AI model" is an artificial intelligence algorithm used by a server to analyze user text data and generate appropriate responses.
[0232] A "prompt message" is input data that is dynamically generated based on user inquiries used by the generative AI model when performing analysis.
[0233] This invention is a system for monitoring elderly people living alone, utilizing voice recognition and biometric data to facilitate communication with them and enable rapid response in emergencies. Details of the invention are shown below.
[0234] System Configuration
[0235] The system primarily consists of "terminals" and "servers." Their respective functions and specific examples are explained below.
[0236] (Device functions)
[0237] 1. Speech recognition means
[0238] The device uses its built-in microphone to capture audio in real time and converts the captured audio into text using speech recognition software (e.g., the speech_recognition library). For example, if a user says "What kind of day is it today?", the audio is converted into the text data "What kind of day is it today?".
[0239] 2. Transmission method
[0240] The terminal sends the converted text data to the server. Therefore, a wireless communication module (e.g., a Wi-Fi module) is required.
[0241] 3. Speech synthesis means
[0242] The terminal converts the response received from the server into audio data using speech synthesis software (e.g., the playsound library) and outputs it to the user through the built-in speaker. For example, if the terminal receives the response "Today is not a special day" from the server, it will convey that to the user as audio.
[0243] 4. Means of acquiring biometric data
[0244] The device uses built-in biosensors (e.g., heart rate sensor, temperature sensor, motion sensor) to continuously acquire data on the user's heart rate, body temperature, and movement.
[0245] 5. Anomaly detection means
[0246] The device monitors acquired biometric data in real time and detects abnormalities. For example, a sudden increase in heart rate or a lack of movement may be considered an abnormality.
[0247] 6. Emergency notification means
[0248] If an anomaly is detected, the terminal immediately sends an emergency notification to the server. A wireless communication module is required for this.
[0249] (Server functions)
[0250] 1. Analysis method
[0251] The server analyzes the received text data and generates an appropriate response. A generative AI model (e.g., the GPT model) is used in this process. For example, in response to the text data "What kind of day is it today?", it generates the response "Today is not a special day."
[0252] 2. Means of contact
[0253] When the server receives an emergency notification, it sends an emergency message to pre-configured emergency contacts. It also includes a mechanism to output the generated emergency message as audio and notify the user.
[0254] 3. Response generation means using an AI model
[0255] The server's analysis method uses a generative AI model to generate appropriate responses based on the user's text data. For example, by inputting a prompt into the generative AI model, a response is dynamically generated.
[0256] Specific example
[0257] The user is wearing smart glasses at home.
[0258] User: "What kind of day is it today?"
[0259] Terminal: Acquires audio, converts it into text data such as "What kind of day is it today?" using speech recognition, and sends it to the server.
[0260] Server: Analyzes the received text data, generates a response saying "Today is not a special day," and sends it to the terminal.
[0261] Terminal: The response from the server is converted into speech data using a speech synthesis method, and the message "Today is not a special day" is output to the user.
[0262] If a user suddenly collapses, the device's biosensors detect a sudden increase in heart rate and cessation of movement, and send an emergency notification to the server.
[0263] Server: Upon receiving an emergency notification, it sends an emergency message to a pre-configured emergency contact (e.g., family member) stating, "The user has collapsed. Please respond immediately."
[0264] Examples of prompts for generative AI models
[0265] User: "What kind of day is it today?"
[0266] The application responded, "Today is not a special day."
[0267] If the user's heart rate suddenly increases afterward, how will an emergency notification be sent?
[0268] Through these components and processes, the present invention provides a system that guarantees the safety and security of the elderly.
[0269] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0270] Step 1:
[0271] The user speaks into the device. The device acquires the voice through its built-in microphone. The acquired voice data is sent as input to a speech recognition system, where the voice is converted into text data.
[0272] Step 2:
[0273] The terminal sends the converted text data to the server via a transmission means. Here, as part of the data processing, the text data is transmitted to the server in an appropriate format.
[0274] Step 3:
[0275] The server analyzes the received text data by means of an analysis means. The analysis means uses a generation AI model to generate an appropriate response based on the input text data. As a specific data operation, a prompt sentence is generated, input into the AI model, and as a result, a text response is generated.
[0276] Step 4:
[0277] The server transmits the generated response back to the terminal through the response transmission means. The data to be transmitted is a text-form response.
[0278] Step 5:
[0279] The terminal converts the response received from the server into voice data by means of voice synthesis means and outputs it to the user through the built-in speaker. The output voice is based on the response content generated by the server.
[0280] Step 6:
[0281] The terminal continuously acquires data on the user's heartbeat, body temperature, and movement by means of a built-in biosensor. These biometric data are sent as inputs to the abnormality detection means. Here, as data processing, the biometric data are analyzed in real time.
[0282] Step 7:
[0283] The abnormality detection means monitors the acquired biometric data in real time and detects whether there is an abnormality. For example, when the heartbeat suddenly rises or there is no movement, it is detected as an abnormality. This detection is the result of data analysis.
[0284] Step 8:
[0285] If an abnormality is detected, the terminal immediately sends an emergency notification to the server using the emergency notification means. This notification includes the user ID, terminal ID, and the detected abnormal data.
[0286] Step 9:
[0287] The server receives an emergency notification and sends an emergency message to a pre-set emergency contact through the communication means. The emergency message contains emergency information such as the user has fallen, which prompts a prompt response.
[0288] Step 10:
[0289] The server outputs the sent emergency message in voice to notify the user. This notification may be made through a phone or a dedicated application.
[0290] Through this series of processes, the system ensures the safety of the user and enables a prompt response in case of an emergency.
[0291] Furthermore, an emotion engine for estimating the user's emotion may be combined. That is, the specific processing unit 290 may estimate the user's emotion using the emotion recognition model 59 and perform specific processing using the user's emotion.
[0292] The present invention relates to a system for monitoring elderly people living alone, promoting communication, and enabling a prompt response in case of an emergency. In particular, the present invention realizes more natural and effective communication by combining an emotion engine for recognizing the user's emotion.
[0293] Speech Recognition and Input Processing
[0294] Terminal
[0295] When the user speaks, the terminal acquires voice through the built-in microphone. The acquired voice data is converted into text data by the voice recognition means. This text data is transmitted to the server by the transmission means of the terminal.
[0296] Server
[0297] The server receives the text data sent from the terminal, and then analyzes the received text data using the analysis means. Based on the analysis, the server generates an appropriate response and sends it back to the terminal via the sending means.
[0298] Terminal
[0299] The terminal converts the received response into voice data using the voice synthesis means and outputs it to the user through the speaker. As a result, the elderly can interact with the system in a natural form.
[0300] Emotion recognition function
[0301] Terminal
[0302] The terminal is equipped with an emotion engine. When the user speaks, the emotion engine recognizes the emotion from the voice and expression. The emotion data is sent to the server together with the text data.
[0303] Server
[0304] The server receives the emotion data and analyzes it together with the text data. Based on the analysis result, a response considering the user's emotion is generated. For example, if the user is judged to be sad, the response will be encouraging.
[0305] Specific example of conversation partner function
[0306] User: "I'm feeling sad today."
[0307] Terminal: Acquire the voice, and convert it into text data "I'm feeling sad today" using the voice recognition means. The emotion engine recognizes the sad emotion from the user's voice.
[0308] Terminal: Send the text data and emotion data to the server.
[0309] Server: Analyze the received data and generate an encouraging response such as "What happened? Please tell me."
[0310] Server: Sends the generated response to the terminal.
[0311] Terminal: Converts response data into speech and outputs to the user, "What's the matter? Please tell me what happened."
[0312] Emergency contact function
[0313] terminal
[0314] The device has built-in biosensors that continuously acquire the user's heart rate, body temperature, and movement. The acquired biometric data is monitored in real time. If an abnormality is detected, an emergency notification is generated.
[0315] server
[0316] Emergency notifications are sent to the server along with emotion data. The server receives the emergency notification and sends an emergency message to pre-registered emergency contacts.
[0317] Examples of emergency contact functions
[0318] Device: If the user falls and becomes motionless, it detects abnormal heart rate and unresponsive movement.
[0319] Terminal: Immediately sends emergency notifications and emotional data to the server.
[0320] Server: Receives emergency notification and sends a message to family contacts saying, "User has collapsed. Please respond quickly."
[0321] User (recipient): The family receives the emergency message and immediately begins taking action.
[0322] Thus, the system of the present invention provides elderly people living alone with an environment in which they can live with peace of mind by simultaneously promoting communication and enabling a rapid response in emergencies. In particular, the emotion engine enables natural dialogue that takes the user's emotions into consideration, thereby improving the quality of life.
[0323] The following describes the processing flow.
[0324] Speech recognition and input processing
[0325] Step 1:
[0326] Terminal: The terminal acquires the user's voice in real time through its built-in microphone.
[0327] Step 2:
[0328] Terminal: Transmits the acquired audio data to a speech recognition device (e.g., a speech recognition engine) and converts the audio data into text data.
[0329] Step 3:
[0330] Terminal: Sends the converted text data to the server using a transmission method. HTTP requests are often used for transmission.
[0331] Emotion recognition function
[0332] Step 4:
[0333] Terminal: Simultaneously, the emotion engine recognizes the user's emotions from their voice and facial expressions. The recognized emotion data is sent to the server along with text data.
[0334] Step 5:
[0335] Server: Analyzes received text data and sentiment data using analytical tools (such as a natural language processing engine).
[0336] Step 6:
[0337] Server: Based on the analysis results, it generates a response that takes the user's emotions into account. For example, if it is determined that the user is sad, the response will be encouraging.
[0338] Step 7:
[0339] Server: Sends the generated response to the terminal via a response transmission method. JSON format data is often used for transmission.
[0340] Step 8:
[0341] Terminal: Receives responses from the server and converts them into speech data using a text-to-speech engine.
[0342] Step 9:
[0343] Terminal: Outputs audio data to the user through the speaker.
[0344] Examples of conversation partner functions
[0345] User: "I'm feeling sad today."
[0346] Step 1: The device acquires the user's voice and converts it into text data, "I feel sad today," using speech recognition technology.
[0347] Step 2: The emotion engine recognizes sad emotions from the user's voice.
[0348] Step 3: Send text data and sentiment data to the server.
[0349] Step 4: The server analyzes the received data and generates an encouraging response such as, "What's wrong? Tell me what happened."
[0350] Step 5: The server sends the generated response to the terminal.
[0351] Step 6: The device converts the response data into speech and outputs to the user, "What's the matter? Please tell me what's wrong."
[0352] Emergency contact function
[0353] Step 1:
[0354] Device: The device continuously acquires biometric data such as heart rate, body temperature, and movement through built-in biosensors.
[0355] Step 2:
[0356] Terminal: Uses an anomaly detection method to monitor acquired biometric data in real time and detect abnormal values.
[0357] Step 3:
[0358] Terminal: If an anomaly is detected, an emergency notification will be generated. This notification will include the type of anomaly and the user's identification information.
[0359] Step 4:
[0360] Terminal: Sends the generated emergency notification to the server using the emergency notification method.
[0361] Step 5:
[0362] Server: Receives emergency notifications and sentiment data, and retrieves a pre-configured list of emergency contacts.
[0363] Step 6:
[0364] Server: Sends an emergency message to emergency contacts based on the information included in the emergency notification. Email, SMS, and phone APIs are often used for sending messages.
[0365] Step 7:
[0366] User (recipient): The family or medical institution will review the received emergency message and begin taking immediate action.
[0367] Examples of emergency contact functions
[0368] Step 1: The device detects the user's abnormal heart rate and cessation of movement.
[0369] Step 2: Send anomaly and emotion data to the server.
[0370] Step 3: The server receives an emergency notification and sends an emergency message to the emergency contact stating, "The user has collapsed. Please respond immediately."
[0371] Step 4: The family receives the emergency message and immediately begins taking action.
[0372] The above outlines the program's processing flow, including the specific actions performed at each processing step. This detailed explanation will allow for a clearer understanding of the overall system's operation.
[0373] (Example 2)
[0374] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0375] When elderly people live alone, there are risks associated with a lack of communication and delayed responses in emergencies. Furthermore, systems that cannot engage in natural conversations that take emotions into consideration can increase the user's mental burden and potentially lower their quality of life. Additionally, if biometric data monitoring and emergency notifications are not effectively implemented, prompt responses become difficult.
[0376] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: speech recognition means for acquiring voice from a terminal in real time and converting the voice into text; emotion recognition means for analyzing emotions from the terminal and generating emotion data; transmission means for the terminal to send the converted text data and emotion data to the server; analysis means for the server to analyze the received text data and emotion data and generate an appropriate response; response transmission means for the server to send the generated response to the terminal; and speech synthesis means for converting the response received by the terminal into voice and outputting it to the user. This enables natural and effective communication with the elderly, allowing for a quick response in emergencies and improving the quality of life for users.
[0377] A "terminal" is a computer device used by users for voice input and output, as well as the collection of biometric data.
[0378] "Voice recognition means" refers to technology or devices that acquire a user's voice in real time on a terminal and convert that voice into text data.
[0379] "Emotion recognition means" refers to technologies and devices that analyze a user's emotions from their voice, facial expressions, etc., and generate emotion data.
[0380] "Transmission means" refers to communication technologies and devices used to send text data and sentiment data from a terminal to a server.
[0381] A "server" is a centralized management system that receives data sent from terminals, analyzes it, and generates appropriate responses.
[0382] "Analysis means" refers to technologies and devices that analyze text data and sentiment data on a server and generate appropriate responses for the user.
[0383] "Response transmission means" refers to communication technology or equipment used to send a response generated by a server to a terminal.
[0384] "Speech synthesis means" refers to a technology or device that converts response data received from a server into speech data at a terminal and outputs it to the user as speech.
[0385] "Methods for acquiring biometric data" refer to technologies and devices that continuously acquire biometric data such as a user's heart rate, body temperature, and movement.
[0386] An "anomaly detection method" refers to a technology or device that monitors acquired biological data in real time and detects anomalies.
[0387] An "emergency notification system" is a technology or device that sends an emergency notification to a server when an anomaly is detected.
[0388] "Communication methods" refer to technologies and devices that, after a server receives an emergency notification, send an emergency message to pre-configured emergency contacts.
[0389] This invention relates to a system that monitors elderly people living alone, facilitates communication, and enables rapid response in emergencies. In particular, this invention achieves more natural and effective communication by combining it with an emotion engine that recognizes the user's emotions.
[0390] Hardware and software to be used
[0391] terminal
[0392] The device includes the following hardware and software:
[0393] High-sensitivity microphone (voice input)
[0394] Speaker (audio output)
[0395] Biosensors (acquisition of heart rate, body temperature, and movement data)
[0396] Speech recognition engine (e.g., Google Speech-to-Text API)
[0397] Emotion recognition engine (e.g., Microsoft® Azure® Emotion API)
[0398] Speech synthesis engine (e.g., Amazon Polly)
[0399] server
[0400] The server includes the following software and features:
[0401] Analysis engine (Python library using NLP technology)
[0402] Response generation engine
[0403] Emergency notification system
[0404] Data receiving and transmitting module
[0405] System operation
[0406] terminal
[0407] When a user speaks, the device acquires the voice through its built-in microphone. The acquired voice data is converted into text data using a speech recognition engine. Simultaneously, an emotion recognition engine analyzes emotions from voice and facial expressions, generating emotion data. This data is transmitted to the server via the device's transmission mechanism.
[0408] server
[0409] The server receives text data and sentiment data sent from the terminal. The received data is analyzed using an analysis engine. Based on the analysis results, the server generates an appropriate response. The generated response is sent to the terminal via a response transmission means.
[0410] terminal
[0411] The terminal converts responses received from the server into speech data using a speech synthesis engine and outputs it to the user through a speaker. This allows elderly people to interact with the system in a natural way.
[0412] Specific example
[0413] Conversation partner function
[0414] User: "I'm feeling sad today."
[0415] Device: It uses the built-in microphone to acquire voice data and converts it into text data, such as "I feel sad today," using a speech recognition engine. The emotion engine then analyzes the user's voice to determine that they are feeling "sad."
[0416] Terminal: Sends the converted text data and sentiment data to the server.
[0417] Server: Analyzes the received data and generates encouraging responses such as, "What's wrong? Please tell me what happened."
[0418] Server: Sends the generated response data to the terminal.
[0419] Terminal: Converts response data into speech and outputs to the user, "What's the matter? Please tell me what happened."
[0420] Emergency contact function
[0421] Device: If the user falls, the built-in biosensors detect abnormal heart rate or unresponsive movement.
[0422] Terminal: Immediately sends emergency notifications and emotional data to the server.
[0423] Server: Upon receiving an emergency notification, it sends the message "User has collapsed. Please respond immediately" to the pre-configured emergency contact.
[0424] Recipient user: The family receives the emergency message and immediately begins taking action.
[0425] Example of a prompt
[0426] Example of prompt text when inputting a "monitoring system for elderly people living alone" into the AI model:
[0427] "The device acquires the user's voice and converts it into text data using a speech recognition engine. Additionally, an emotion engine analyzes the user's emotions and sends this data to a server. The server analyzes the received data, generates an appropriate response, and sends it to the device, which then responds verbally. Furthermore, in emergencies, a biosensor detects an anomaly and sends an emergency notification through the server. Please explain this system."
[0428] Thus, the system of the present invention promotes natural dialogue with the elderly and enables a rapid response in emergencies, thereby improving the safety and quality of life of the elderly.
[0429] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0430] Step 1:
[0431] The device acquires the user's voice. Using the built-in high-sensitivity microphone, it collects voice data when the user says, "I'm feeling sad today."
[0432] Input: User's voice
[0433] Output: Audio data
[0434] Specific operation: The device's operating system converts the audio signal from the microphone into digital data and temporarily stores it in memory.
[0435] Step 2:
[0436] The device uses a speech recognition engine to convert audio data into text data. For example, the speech recognition engine uses the Google Speech-to-Text API.
[0437] Input: Audio data
[0438] Output: Text data
[0439] Specific operation: The speech recognition engine analyzes the audio data, identifies phonemes and morphemes, and generates the text "I feel sad today." The generated text is stored in a separate buffer in memory.
[0440] Step 3:
[0441] The device uses an emotion recognition engine to generate emotion data from voice and text data. For example, it can utilize the Microsoft Azure Emotion API.
[0442] Input: Audio data, text data
[0443] Output: Sentiment data
[0444] Specific operation: The emotion recognition engine analyzes the emotion of "sadness" from the tone of voice and associated text, and generates this emotion data. The generated emotion data is stored in memory along with the text data.
[0445] Step 4:
[0446] The terminal sends the converted text data and sentiment data to the server. This is done using a network transmission module.
[0447] Input: Text data, sentiment data
[0448] Output: Transmitted data
[0449] Specific operation: The transmission module packets text and sentiment data and sends them to the server via the Wi-Fi module over the internet. The data is encrypted to ensure security.
[0450] Step 5:
[0451] The server receives text data and sentiment data sent from the terminal.
[0452] Input: Data to send
[0453] Output: Received data
[0454] Specific operation: The server's receiving module receives packets, decrypts the encrypted data, and obtains text data and sentiment data. This data is stored in a database for analysis.
[0455] Step 6:
[0456] The server analyzes text and sentiment data and generates an appropriate response using an analysis engine.
[0457] Input: Text data, sentiment data
[0458] Output: Response data
[0459] Specific operation: The analysis engine uses NLP technology to analyze text data and sentiment data, and generates the most appropriate response for the user's state, such as "What's wrong? Please tell me what happened."
[0460] Step 7:
[0461] The server sends the generated response data to the terminal. A response transmission method is used.
[0462] Input: Response data
[0463] Output: Transmitted data
[0464] Specific operation: The response data is packetized and sent to the terminal via the network transmission module over the internet. The data is transmitted encrypted.
[0465] Step 8:
[0466] The terminal converts the received response data into audio data and outputs it to the user through the speaker. A speech synthesis engine is used.
[0467] Input: Response data
[0468] Output: Audio data
[0469] Specific operation: The speech synthesis engine analyzes the response data and generates the voice message, "What's the matter? Please tell me what's wrong." The generated voice message is output to the user through the speaker.
[0470] Step 9:
[0471] The device monitors biometric data, using sensors such as heart rate sensors, body temperature sensors, and motion sensors.
[0472] Input: Biometric data
[0473] Output: Acquired data
[0474] Specific operation: The sensor continuously measures the user's heart rate, body temperature, and movement, and records this data in memory. The data is updated in real time, waiting for anomaly detection.
[0475] Step 10:
[0476] If the device detects an anomaly, it sends an emergency notification to the server. For example, it might detect a sudden increase in heart rate or a cessation of movement.
[0477] Input: Acquired data
[0478] Output: Emergency notification data
[0479] Specific operation: When the anomaly detection algorithm detects a sudden increase in heart rate, abnormal fluctuations in body temperature, or cessation of movement, it generates emergency notification data related to these events and sends it to the server.
[0480] Step 11:
[0481] The server receives an emergency notification and sends an emergency message to pre-configured emergency contacts.
[0482] Input: Emergency notification data
[0483] Output: Emergency message
[0484] Specific operation: The server analyzes the received emergency notification data and sends a message stating "User has collapsed. Please respond immediately" to a pre-configured list of emergency contacts (e.g., family or medical institutions). The message is sent via SMS or a dedicated application.
[0485] (Application Example 2)
[0486] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0487] In recent years, the number of elderly people living alone has increased, making their safety and security a major concern. There is also a growing need for systems that alleviate their feelings of loneliness and provide rapid responses in emergencies. However, current systems lack sufficient emotion recognition and emergency notification capabilities, and cannot fully guarantee the psychological and physical safety of the elderly. Therefore, there is a need for natural communication methods that include emotion recognition, as well as rapid emergency response through real-time monitoring of biometric data.
[0488] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a voice recognition means, a transmission means, an analysis means, a response transmission means, a voice synthesis means, and a means for recognizing emotions from the user's voice and facial expressions, equipped with an emotion recognition engine. This enables natural dialogue that takes into account the emotions of the elderly, and also enables a rapid response in emergencies. It also has an emergency notification function and can provide a system that continuously monitors heart rate, body temperature, and movement data, and immediately makes an emergency contact when an abnormality is detected.
[0489] A "terminal" is a device used by a user that has functions such as voice input, speech synthesis, emotion recognition, and acquisition of biometric data.
[0490] "Voice recognition means" refers to a function that converts the voice acquired by the terminal into text data.
[0491] "Transmission method" refers to the function by which the terminal sends the converted text data to the server.
[0492] "Analysis means" refers to a function that analyzes text data and sentiment data received by the server and generates an appropriate response.
[0493] "Response transmission means" refers to the function that sends the response generated by the server to the terminal.
[0494] A "speech synthesis means" is a function that converts the response received by the terminal into speech and outputs it to the user.
[0495] An "emotion recognition engine" is a function that recognizes emotions from the user's voice and facial expressions.
[0496] "Means of acquiring biometric data" refers to a function in which the terminal continuously acquires the user's heart rate, body temperature, and movement.
[0497] An "anomaly detection method" is a function that monitors acquired biometric data in real time and detects sudden increases in heart rate, abnormal fluctuations in body temperature, and cessation of movement.
[0498] An "emergency notification mechanism" is a function that sends an emergency notification to the server when an anomaly is detected.
[0499] "Communication method" refers to the function where the server receives an emergency notification and sends an emergency message to a pre-configured emergency contact.
[0500] This invention is a system that monitors elderly people living alone, facilitates communication, and enables rapid response in emergencies. The system mainly consists of terminals and a server.
[0501] Device functions
[0502] The terminal is equipped with a speech recognition system, a speech synthesis system, an emotion recognition engine, and a biometric data acquisition system. When a user speaks, the voice is acquired through the terminal's microphone and converted into text data by the speech recognition system. This text data is then sent to the server via a transmission system.
[0503] The device also features an emotion recognition engine that recognizes emotions from the user's voice and facial expressions. This emotion data is also sent to the server along with the text data.
[0504] Furthermore, the device is equipped with biometric data acquisition capabilities, continuously monitoring the user's heart rate, body temperature, and movement. If an abnormality is detected, a notification is sent to the server via an emergency notification system.
[0505] Server Functions
[0506] The server has an analysis mechanism to analyze the received text data and sentiment data, and generates an appropriate response. The generated response is sent back to the terminal via a response transmission mechanism. The terminal converts this response into speech using a speech synthesis mechanism and outputs it to the user through a speaker.
[0507] The server has a communication system that, upon receiving an emergency notification, sends an emergency message to pre-configured emergency contacts.
[0508] Hardware and software details
[0509] The system uses the "Speech Recognition API" for speech recognition, the "Network Communication Module" for text data transmission, and the "Emotion Recognition API" for emotion recognition. Additionally, it uses the "Speech Synthesis API" for speech synthesis and the "Biometric Sensor Module" for remote monitoring. The "Notification API" is used for emergency notifications and communication.
[0510] Specific example
[0511] When a user says, "I'm not feeling well today," the device converts the audio data into text data and analyzes it using an emotion recognition engine. The resulting data is sent to a server, which considers the user's emotions and generates an encouraging response such as, "What's wrong? Please tell me what's the matter." The generated response is sent to the device and output to the user as audio.
[0512] Furthermore, if a user collapses, the biosensors built into the device detect abnormal heart rate or unresponsive movement and immediately send an emergency notification to the server. The server receives the emergency notification and sends a message to emergency contacts stating, "A user has collapsed. Please respond quickly."
[0513] Example of a prompt
[0514] "Provide an example where, if a user says 'I'm not feeling well today,' the audio is converted to text, analyzed by an emotion engine, and an appropriate comforting response is generated."
[0515] As described above, this system supports the safe and secure lives of the elderly and can provide prompt assistance when needed. Furthermore, its emotion recognition function enables more natural and effective communication.
[0516] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0517] Step 1:
[0518] The user speaks to the device.
[0519] Input: User's voice data.
[0520] Specific action: The device's microphone picks up the user's voice.
[0521] Step 2:
[0522] The device converts the audio into text data.
[0523] Input: Acquired audio data.
[0524] Specific operation: Use a speech recognition API to convert speech data into text data.
[0525] Output: Text data.
[0526] Step 3:
[0527] The device performs emotion recognition.
[0528] Input: User's voice data and text data.
[0529] Specific operation: Use an emotion recognition API to analyze the user's emotions from their voice and facial expressions.
[0530] Output: Sentiment data.
[0531] Step 4:
[0532] The device sends text data and sentiment data to the server.
[0533] Input: Text data and sentiment data.
[0534] Specific operation: Use the network communication module to send data to the server.
[0535] Output: Confirmation of data transmission to the server.
[0536] Step 5:
[0537] The server analyzes the text data and sentiment data it receives.
[0538] Input: Text data and sentiment data.
[0539] Specific operation: Use analytical tools to perform analysis in order to generate an appropriate response.
[0540] Output: Appropriate response data.
[0541] Step 6:
[0542] The server sends the appropriate response data to the terminal.
[0543] Input: Generated response data.
[0544] Specific operation: Use the response transmission means to send response data to the terminal.
[0545] Output: Confirmation of data transmission to the terminal.
[0546] Step 7:
[0547] The terminal converts the received response data into speech.
[0548] Input: Appropriate response data.
[0549] Specific operation: Use a speech synthesis API to convert text data into speech data.
[0550] Output: Audio data.
[0551] Step 8:
[0552] The device outputs audio data to the user.
[0553] Input: Audio data.
[0554] Specific action: Play audio through the device's speaker.
[0555] Output: Audio output to the user.
[0556] Step 9:
[0557] The device acquires biometric data.
[0558] Input: Biometric data such as the user's heart rate, body temperature, and movement.
[0559] Specific operation: Use a biosensor module to continuously collect data.
[0560] Output: Biometric data.
[0561] Step 10:
[0562] The device monitors biometric data acquired in real time.
[0563] Input: Biometric data.
[0564] Specific operation: Use anomaly detection methods to monitor data for abnormalities.
[0565] Output: Abnormal status data.
[0566] Step 11:
[0567] If an anomaly is detected, the device will send an emergency notification to the server.
[0568] Input: Abnormal situation data.
[0569] Specific action: Use the emergency notification method to send an anomaly notification to the server.
[0570] Output: Sending an emergency notification to the server.
[0571] Step 12:
[0572] The server receives an emergency notification and sends a message to the emergency contact.
[0573] Input: Anomaly notification data.
[0574] Specific action: Use the communication method to send a message to a pre-configured emergency contact.
[0575] Output: Confirmation of emergency message transmission.
[0576] The above describes the processing steps of the system based on this invention.
[0577] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0578] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0579] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0580] [Second Embodiment]
[0581] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0582] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0583] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0584] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0585] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0586] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0587] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0588] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0589] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0590] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0591] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0592] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0593] This invention relates to a system for monitoring elderly people living alone, facilitating communication with them using voice and biometric data, and enabling a rapid response in emergencies. Embodiments of this invention are described in detail below.
[0594] Speech recognition and input processing
[0595] terminal
[0596] When a user speaks into the device, the device acquires the voice through its built-in microphone. The acquired voice data is converted into text data by a speech recognition system. This text data is then sent to the server through the device's transmission system.
[0597] server
[0598] The server receives text data sent from the terminal. The received data is analyzed by an analysis device, and an appropriate response is generated based on the user's intent. The generated response is then sent back to the terminal via the server's response transmission device.
[0599] terminal
[0600] The terminal receives responses from the server, converts them into speech data using speech synthesis technology, and outputs them to the user through a speaker. This allows elderly people to interact with the system in a natural way.
[0601] Examples of conversation partner functions
[0602] User: "What kind of day is it today?"
[0603] Terminal: Acquires audio and converts it into text data, such as "What kind of day is it today?", using speech recognition technology.
[0604] Terminal: Sends text data to the server.
[0605] Server: Analyzes the received text data and generates a response such as "Today is not a special day."
[0606] Server: Sends the generated response to the terminal.
[0607] Terminal: Converts the response data into speech and outputs "Today is not a special day" to the user.
[0608] Emergency contact function
[0609] terminal
[0610] The device has built-in biosensors that continuously acquire the user's heart rate, body temperature, and movement. This biometric data is constantly monitored, and if an abnormality is detected, an emergency notification is generated.
[0611] server
[0612] Emergency notifications are sent from the terminal to the server. The server analyzes the emergency notification and sends an emergency message to pre-registered emergency contacts. The emergency message includes the user ID, terminal ID, and abnormal data.
[0613] Examples of emergency contact functions
[0614] Device: If the user falls and becomes motionless, it detects abnormal heart rate and unresponsive movement.
[0615] Terminal: Immediately send an emergency notification to the server.
[0616] Server: Receives emergency notification and sends SMS to pre-configured family contacts. "User has collapsed. Please respond immediately."
[0617] User (recipient): The family receives the emergency message and immediately begins taking action.
[0618] In this way, the system of the present invention can ensure the safety and communication of elderly people living alone.
[0619] The following describes the processing flow.
[0620] Speech recognition and input processing
[0621] Step 1:
[0622] Terminal: The terminal acquires the user's voice in real time through its built-in microphone.
[0623] Step 2:
[0624] Terminal: Transmits the acquired audio data to a speech recognition device (e.g., a speech recognition engine) and converts the audio data into text data.
[0625] Step 3:
[0626] Terminal: Sends the converted text data to the server using a transmission method. HTTP requests are often used for transmission.
[0627] Conversation partner function
[0628] Step 4:
[0629] Server: Analyzes the received text data using parsing tools (such as a natural language processing engine).
[0630] Step 5:
[0631] Server: Generates an appropriate response based on the analysis results. The response may include fixed statements or information obtained from external APIs.
[0632] Step 6:
[0633] Server: Sends the generated response to the terminal via a response transmission method. JSON format data is often used for transmission.
[0634] Step 7:
[0635] Terminal: Converts the received response text data into speech data using a text-to-speech engine.
[0636] Step 8:
[0637] Terminal: Outputs audio data to the user through the speaker.
[0638] Emergency contact function
[0639] Step 1:
[0640] Device: The device continuously acquires biometric data such as heart rate, body temperature, and movement through built-in biosensors.
[0641] Step 2:
[0642] Terminal: Uses an anomaly detection method to monitor acquired biometric data in real time and detect abnormal values.
[0643] Step 3:
[0644] Terminal: If an anomaly is detected, an emergency notification will be generated. This notification will include the type of anomaly and the user's identification information.
[0645] Step 4:
[0646] Terminal: Sends the generated emergency notification to the server using the emergency notification method.
[0647] Step 5:
[0648] Server: Receives emergency notifications and retrieves a pre-configured list of emergency contacts.
[0649] Step 6:
[0650] Server: Sends an emergency message to emergency contacts based on the information included in the emergency notification. Email, SMS, and phone APIs are often used for sending messages.
[0651] Step 7:
[0652] User (recipient): The family or medical institution will review the received emergency message and begin taking immediate action.
[0653] The above is the program's processing flow, including the specific actions performed at each processing step. Providing such a detailed explanation allows for a clearer understanding of how the entire system works.
[0654] (Example 1)
[0655] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0656] In modern society, elderly people living alone face challenges such as a lack of daily communication and the inability to respond quickly in emergencies. This has led to feelings of loneliness among the elderly and delays in emergency response. The present invention aims to provide a system that enables elderly people to engage in daily conversations in a natural manner and further enables a rapid response in emergencies using biometric data.
[0657] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0658] In this invention, the server includes: speech recognition means for acquiring voice from a terminal in real time and converting the voice into text; transmission means for the terminal to send the converted text data to the server; analysis means for the server to analyze the received text data and generate an appropriate response using natural language processing means; response transmission means for the server to send the generated response to the terminal; speech synthesis means for the terminal to convert the received response into voice and output it to the user; anomaly detection means for monitoring biometric data in real time and detecting anomalies; emergency notification means for sending an emergency notification to the server when an anomaly is detected; and communication means for the server to receive the emergency notification and automatically send an emergency message to a pre-configured emergency contact. This enables elderly people to communicate naturally using voice on a daily basis, and also enables a rapid response in emergencies using biometric data.
[0659] A "terminal" is a device that functions as an interface with the user, collecting, converting, transmitting, and outputting responses for voice data.
[0660] "Speech recognition means" refers to a technology that converts speech into text data, is built into a terminal, and has the function of recognizing the user's voice in real time.
[0661] "Transmission means" refers to the communication means used by a terminal to send converted text data to a server, and includes Wi-Fi modules and internet connection means.
[0662] A "server" is a central control unit that receives data sent from terminals and performs processing, analysis, and response generation.
[0663] "Natural language processing means" refers to technology that analyzes received text data, understands the user's intent, and generates an appropriate response.
[0664] "Analysis means" refers to the technology used by a server to process received text data and generate an appropriate response using natural language processing means.
[0665] "Response transmission means" refers to communication technology used to send responses generated by a server to a terminal.
[0666] "Speech synthesis means" refers to a technology that converts text data received from a server into speech data and outputs a speech response to the user.
[0667] "Methods for acquiring biometric data" refers to sensor technology that continuously acquires data on the user's heart rate, body temperature, and movement.
[0668] "Anomaly detection means" refers to technology that monitors biological data in real time and detects anomalies.
[0669] An "emergency notification system" refers to a technology that immediately sends an emergency notification to a server when an anomaly is detected.
[0670] "Communication method" refers to the technology in which a server receives an emergency notification and automatically sends an emergency message to pre-configured emergency contacts.
[0671] A "generative AI model" refers to artificial intelligence technology used to generate natural dialogue and appropriate responses in abnormal situations.
[0672] A "prompt statement" refers to an input statement used to give instructions to a generative AI model and is used to generate a specific response.
[0673] This invention relates to a system for monitoring elderly people living alone. It facilitates communication with the elderly using voice and biometric data, and enables a rapid response in emergencies. The embodiments of this invention are described in detail below.
[0674] Speech recognition and input processing
[0675] terminal
[0676] When a user speaks into the device, the device acquires the voice through its built-in microphone. The acquired voice data is converted into text data using speech recognition tools (e.g., Google Cloud Speech-to-Text API). This text data is then sent to the server via the device's transmission tools (Wi-Fi module).
[0677] server
[0678] The server uses a Python framework (e.g., Flask or Django) to receive text data sent from the terminal. The received data is parsed by a natural language processing tool (e.g., spaCy) to generate an appropriate response based on the user's intent. The generated response is then sent back to the terminal via the server's response sending mechanism.
[0679] terminal
[0680] The device uses the Google Cloud Text-to-Speech API to convert responses received from the server into audio data, which is then output to the user through the speaker. This allows elderly users to interact with the system in a natural way.
[0681] Specific example
[0682] User: "What kind of day is it today?"
[0683] Terminal: Acquires audio and converts it into text data, such as "What kind of day is it today?", using speech recognition technology.
[0684] Terminal: Sends text data to the server.
[0685] Server: Analyzes the received text data and generates a response such as "Today is not a special day."
[0686] Server: Sends the generated response to the terminal.
[0687] Terminal: Converts the response data into speech and outputs "Today is not a special day" to the user.
[0688] Emergency contact function
[0689] terminal
[0690] The device has built-in biosensors (heart rate sensor, body temperature sensor, motion sensor, etc.) that continuously acquire the user's heart rate, body temperature, and movement. This biometric data is constantly monitored by a microcontroller (e.g., ESP32), and if an anomaly is detected, an emergency notification is generated.
[0691] server
[0692] Emergency notifications are sent from the device to the server. The server analyzes the emergency notification and sends an emergency message to pre-registered emergency contacts using the Twilio API. The emergency message includes the user ID, device ID, and anomaly data.
[0693] Specific example
[0694] If a user collapses and becomes motionless, the system detects abnormal heart rate and unresponsive movement.
[0695] Terminal: Immediately send an emergency notification to the server.
[0696] Server: Receives emergency notification and sends an SMS to pre-configured family contacts using the Twilio API. "User has collapsed. Please respond quickly."
[0697] Examples of prompt statements
[0698] The following are specific examples of prompt statements used in generative AI models.
[0699] In the case of an elderly monitoring system:
[0700] Create text to generate natural conversations with users.
[0701] For example, generate a response to the user's question, "What kind of day is it today?"
[0702] By using these prompts, it becomes possible to provide a system that enables natural interaction with users and allows for rapid response in emergencies.
[0703] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0704] Program processing for speech recognition and input processing
[0705] Step 1:
[0706] terminal
[0707] The user speaks into the device, and the voice is captured through the built-in microphone. This voice data is the input. The captured voice data is sent in real time to a speech recognition system (e.g., Google Cloud Speech-to-Text API) and converted into text data. The data processing here is the conversion from voice to text, and the output is text data.
[0708] Specific actions
[0709] User: "What kind of day is it today?"
[0710] Device: The built-in microphone captures audio and sends it to the speech recognition API.
[0711] Step 2:
[0712] terminal
[0713] The system acquires text data converted by speech recognition and sends it to a server. The input here is the converted text data, and the output is the text data sent to the server. Transmission takes place via the internet through a Wi-Fi module.
[0714] Specific actions
[0715] Terminal: The device sends the text data "What kind of day is it today?" obtained from the voice recognition device to the server via the Wi-Fi module.
[0716] Step 3:
[0717] server
[0718] The server receives text data sent from the terminal using a Python framework (e.g., Flask or Django). The input here is text data, which is parsed using natural language processing tools (e.g., spaCy). The output is the appropriate response resulting from the parsing.
[0719] Specific actions
[0720] Server: Analyzes the text data "What kind of day is it today?" and uses a natural language processing library to understand the user's intent.
[0721] Step 4:
[0722] server
[0723] Based on the analysis results, an appropriate response is generated using a generative AI model (e.g., GPT-3). The input is the analysis results, and the output is the generated response text.
[0724] Specific actions
[0725] Server: Generates responses such as "Today is not a special day" using an AI model based on the analysis results from a natural language processing library.
[0726] Step 5:
[0727] server
[0728] The generated response is sent to the terminal via the response transmission means. The input is the generated response text, and the output is the text data sent to the terminal.
[0729] Specific actions
[0730] Server: Sends the response text "Today is not a special day" to the terminal.
[0731] Step 6:
[0732] terminal
[0733] The system converts the response received from the server into speech data using a text-to-speech synthesis tool (e.g., Google Cloud Text-to-Speech API). The input is the response text, and the output is speech data. This is output to the user through the built-in speaker.
[0734] Specific actions
[0735] Terminal: Receives the response text "Today is not a special day," converts it into speech data using a speech synthesis method, and outputs it through the speaker.
[0736] Processing of the emergency contact function program
[0737] Step 1:
[0738] terminal
[0739] The device's built-in biosensors (heart rate sensor, body temperature sensor, motion sensor, etc.) continuously acquire the user's biometric data. The input is data from the biosensors, and the output is the continuously acquired biometric data. This data is sent to a microcontroller (e.g., ESP32) for monitoring.
[0740] Specific actions
[0741] Device: The user's heart rate sensor measures a heart rate of 60 BPM.
[0742] Step 2:
[0743] terminal
[0744] The system monitors acquired biometric data and detects anomalies if abnormal values exceed a threshold. The input is continuously acquired biometric data, which is analyzed by an anomaly detection mechanism. The output indicates whether or not an anomaly exists.
[0745] Specific actions
[0746] Terminal: If a heart rate of 30 BPM (below the abnormal capture threshold) is detected, it will be flagged as abnormal.
[0747] Step 3:
[0748] terminal
[0749] When an anomaly is detected, an emergency notification is immediately generated. The input is the result of the anomaly detection, and the output is the emergency notification data. The generated emergency notification includes the user ID and anomaly data.
[0750] Specific actions
[0751] The device detects an abnormal heart rate and generates data that reads "Emergency notification: User ID 123, heart rate 30 BPM".
[0752] Step 4:
[0753] terminal
[0754] The generated emergency notification is sent to the server. The input is the emergency notification data, and the output is the emergency notification sent to the server. Transmission is performed via the Wi-Fi module.
[0755] Specific actions
[0756] Device: Emergency notifications are sent to the server via Wi-Fi.
[0757] Step 5:
[0758] server
[0759] This system receives emergency notifications and automatically sends emergency messages to pre-registered emergency contacts via the Twilio API. The input is the received emergency notification data, and the output is the sent emergency message. The emergency message includes user ID, device ID, and anomaly data.
[0760] Specific actions
[0761] Server: Upon receiving an emergency notification, it sends an SMS message using the Twilio API stating, "A user has collapsed. Please respond quickly."
[0762] As described above, the system of the present invention enables natural communication with the elderly using voice and biometric data, as well as a rapid response in emergencies.
[0763] (Application Example 1)
[0764] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0765] The problem that this invention aims to solve is to provide an environment in which elderly people living alone can live safely and securely. In particular, it aims to enable rapid response in emergencies by linking voice recognition and biometric data, and to monitor changes in the elderly person's physical condition and abnormal situations in real time. Another challenge is to improve communication with the elderly person through a user-friendly interface.
[0766] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0767] In this invention, the server includes an analysis means for analyzing text data and generating an appropriate response, a communication means for receiving emergency notifications and sending emergency messages to pre-configured emergency contacts, and a means for generating responses using an AI model. This enables elderly people living alone to communicate naturally by voice and to respond quickly in emergencies.
[0768] A "terminal" is a device used by a user that has the function of acquiring voice data and sending it to a server, as well as the function of acquiring biometric data.
[0769] "Voice recognition means" refers to technology that converts speech acquired by a terminal into text data.
[0770] "Transmission method" refers to the function by which the terminal sends the converted text data to the server.
[0771] "Analysis means" refers to a technology that analyzes text data received by a server and generates an appropriate response based on the user's intent.
[0772] "Response transmission means" refers to the function by which the server sends the generated response to the terminal.
[0773] "Speech synthesis means" refers to a technology that converts responses received by a terminal from a server into speech data and outputs it to the user as speech.
[0774] "Means of acquiring biometric data" refers to a function in which a device continuously acquires biometric data such as the user's heart rate, body temperature, and movement.
[0775] An "anomaly detection method" is a technology that monitors acquired biometric data in real time and generates an emergency notification when an anomaly is detected.
[0776] An "emergency notification mechanism" is a function that allows a terminal to send an emergency notification to a server when it detects an anomaly.
[0777] "Communication method" refers to a technology where a server receives an emergency notification and sends an emergency message to pre-configured emergency contacts.
[0778] An "AI model" is an artificial intelligence algorithm used by a server to analyze user text data and generate appropriate responses.
[0779] A "prompt message" is input data that is dynamically generated based on user inquiries used by the generative AI model when performing analysis.
[0780] This invention is a system for monitoring elderly people living alone, utilizing voice recognition and biometric data to facilitate communication with them and enable rapid response in emergencies. Details of the invention are shown below.
[0781] System Configuration
[0782] The system primarily consists of "terminals" and "servers." Their respective functions and specific examples are explained below.
[0783] (Device functions)
[0784] 1. Speech recognition means
[0785] The device uses its built-in microphone to capture audio in real time and converts the captured audio into text using speech recognition software (e.g., the speech_recognition library). For example, if a user says "What kind of day is it today?", the audio is converted into the text data "What kind of day is it today?".
[0786] 2. Transmission method
[0787] The terminal sends the converted text data to the server. Therefore, a wireless communication module (e.g., a Wi-Fi module) is required.
[0788] 3. Speech synthesis means
[0789] The terminal converts the response received from the server into audio data using speech synthesis software (e.g., the playsound library) and outputs it to the user through the built-in speaker. For example, if the terminal receives the response "Today is not a special day" from the server, it will convey that to the user as audio.
[0790] 4. Means of acquiring biometric data
[0791] The device uses built-in biosensors (e.g., heart rate sensor, temperature sensor, motion sensor) to continuously acquire data on the user's heart rate, body temperature, and movement.
[0792] 5. Anomaly detection means
[0793] The device monitors acquired biometric data in real time and detects abnormalities. For example, a sudden increase in heart rate or a lack of movement may be considered an abnormality.
[0794] 6. Emergency notification means
[0795] If an anomaly is detected, the terminal immediately sends an emergency notification to the server. A wireless communication module is required for this.
[0796] (Server functions)
[0797] 1. Analysis method
[0798] The server analyzes the received text data and generates an appropriate response. A generative AI model (e.g., the GPT model) is used in this process. For example, in response to the text data "What kind of day is it today?", it generates the response "Today is not a special day."
[0799] 2. Means of contact
[0800] When the server receives an emergency notification, it sends an emergency message to pre-configured emergency contacts. It also includes a mechanism to output the generated emergency message as audio and notify the user.
[0801] 3. Response generation means using an AI model
[0802] The server's analysis method uses a generative AI model to generate appropriate responses based on the user's text data. For example, by inputting a prompt into the generative AI model, a response is dynamically generated.
[0803] Specific example
[0804] The user is wearing smart glasses at home.
[0805] User: "What kind of day is it today?"
[0806] Terminal: Acquires audio, converts it into text data such as "What kind of day is it today?" using speech recognition, and sends it to the server.
[0807] Server: Analyzes the received text data, generates a response saying "Today is not a special day," and sends it to the terminal.
[0808] Terminal: The response from the server is converted into speech data using a speech synthesis method, and the message "Today is not a special day" is output to the user.
[0809] If a user suddenly collapses, the device's biosensors detect a sudden increase in heart rate and cessation of movement, and send an emergency notification to the server.
[0810] Server: Upon receiving an emergency notification, it sends an emergency message to a pre-configured emergency contact (e.g., family member) stating, "The user has collapsed. Please respond immediately."
[0811] Examples of prompts for generative AI models
[0812] User: "What kind of day is it today?"
[0813] The application responded, "Today is not a special day."
[0814] If the user's heart rate suddenly increases afterward, how will an emergency notification be sent?
[0815] Through these components and processes, the present invention provides a system that guarantees the safety and security of the elderly.
[0816] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0817] Step 1:
[0818] The user speaks into the device. The device acquires the voice through its built-in microphone. The acquired voice data is sent as input to a speech recognition system, where the voice is converted into text data.
[0819] Step 2:
[0820] The terminal sends the converted text data to the server via a transmission means. Here, as part of the data processing, the text data is transmitted to the server in an appropriate format.
[0821] Step 3:
[0822] The server analyzes the received text data using an analysis tool. The analysis tool uses a generative AI model to generate an appropriate response based on the input text data. Specifically, the data calculation involves generating a prompt sentence, inputting it into the AI model, and generating a text response as a result.
[0823] Step 4:
[0824] The server sends the generated response back to the terminal via the response transmission means. The data transmitted is a text-based response.
[0825] Step 5:
[0826] The terminal converts the response received from the server into audio data using a speech synthesis system and outputs it to the user through its built-in speaker. The output audio is based on the response content generated by the server.
[0827] Step 6:
[0828] The device continuously acquires data on the user's heart rate, body temperature, and movement using built-in biosensors. This biometric data is sent as input to an anomaly detection system. Here, the biometric data is analyzed in real time as part of the data processing.
[0829] Step 7:
[0830] The anomaly detection system monitors acquired biometric data in real time and detects whether there are any abnormalities. For example, a sudden increase in heart rate or a lack of movement will be detected as an anomaly. This detection is the result of data analysis.
[0831] Step 8:
[0832] If an anomaly is detected, the terminal immediately sends an emergency notification to the server using the emergency notification mechanism. This notification includes the user ID, terminal ID, and detected anomaly data.
[0833] Step 9:
[0834] The server receives an emergency notification and sends an emergency message to pre-configured emergency contacts via the designated communication channels. The emergency message includes urgent information, such as that the user has collapsed, prompting a swift response.
[0835] Step 10:
[0836] The server will output the transmitted emergency message as an audio message to notify the user. This notification may be made via telephone or a dedicated application.
[0837] This series of processes ensures user safety and enables rapid response in emergencies.
[0838] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0839] This invention relates to a system that monitors elderly people living alone, facilitates communication, and enables rapid response in emergencies. In particular, this invention achieves more natural and effective communication by combining it with an emotion engine that recognizes the user's emotions.
[0840] Speech recognition and input processing
[0841] terminal
[0842] When a user speaks, the device acquires the voice through its built-in microphone. The acquired voice data is converted into text data by a speech recognition device. This text data is then sent to the server by the device's transmission device.
[0843] server
[0844] The server receives text data sent from the terminal and then analyzes the received text data using an analysis device. Based on the analysis, the server generates an appropriate response and sends that response back to the terminal via the transmission device.
[0845] terminal
[0846] The terminal converts received responses into speech data using speech synthesis technology and outputs it to the user through a speaker. This allows elderly people to interact with the system in a natural way.
[0847] Emotion recognition function
[0848] terminal
[0849] The device is equipped with an emotion engine. When the user speaks, the emotion engine recognizes emotions from their voice and facial expressions. The emotion data is sent to the server along with the text data.
[0850] server
[0851] The server receives emotion data and analyzes it along with text data. Based on the analysis results, it generates a response that takes the user's emotions into account. For example, if the user is determined to be sad, the response will be encouraging.
[0852] Examples of conversation partner functions
[0853] User: "I'm feeling sad today."
[0854] Terminal: Acquires audio and converts it into text data, such as "I feel sad today," using speech recognition. The emotion engine recognizes the sad emotion from the user's voice.
[0855] Terminal: Sends text data and sentiment data to the server.
[0856] Server: Analyzes the received data and generates encouraging responses such as, "What's wrong? Please tell me what happened."
[0857] Server: Sends the generated response to the terminal.
[0858] Terminal: Converts response data into speech and outputs to the user, "What's the matter? Please tell me what happened."
[0859] Emergency contact function
[0860] terminal
[0861] The device has built-in biosensors that continuously acquire the user's heart rate, body temperature, and movement. The acquired biometric data is monitored in real time. If an abnormality is detected, an emergency notification is generated.
[0862] server
[0863] Emergency notifications are sent to the server along with emotion data. The server receives the emergency notification and sends an emergency message to pre-registered emergency contacts.
[0864] Examples of emergency contact functions
[0865] Device: If the user falls and becomes motionless, it detects abnormal heart rate and unresponsive movement.
[0866] Terminal: Immediately sends emergency notifications and emotional data to the server.
[0867] Server: Receives emergency notification and sends a message to family contacts saying, "User has collapsed. Please respond quickly."
[0868] User (recipient): The family receives the emergency message and immediately begins taking action.
[0869] Thus, the system of the present invention provides elderly people living alone with an environment in which they can live with peace of mind by simultaneously promoting communication and enabling a rapid response in emergencies. In particular, the emotion engine enables natural dialogue that takes the user's emotions into consideration, thereby improving the quality of life.
[0870] The following describes the processing flow.
[0871] Speech recognition and input processing
[0872] Step 1:
[0873] Terminal: The terminal acquires the user's voice in real time through its built-in microphone.
[0874] Step 2:
[0875] Terminal: Transmits the acquired audio data to a speech recognition device (e.g., a speech recognition engine) and converts the audio data into text data.
[0876] Step 3:
[0877] Terminal: Sends the converted text data to the server using a transmission method. HTTP requests are often used for transmission.
[0878] Emotion recognition function
[0879] Step 4:
[0880] Terminal: Simultaneously, the emotion engine recognizes the user's emotions from their voice and facial expressions. The recognized emotion data is sent to the server along with text data.
[0881] Step 5:
[0882] Server: Analyzes received text data and sentiment data using analytical tools (such as a natural language processing engine).
[0883] Step 6:
[0884] Server: Based on the analysis results, it generates a response that takes the user's emotions into account. For example, if it is determined that the user is sad, the response will be encouraging.
[0885] Step 7:
[0886] Server: Sends the generated response to the terminal via a response transmission method. JSON format data is often used for transmission.
[0887] Step 8:
[0888] Terminal: Receives responses from the server and converts them into speech data using a text-to-speech engine.
[0889] Step 9:
[0890] Terminal: Outputs audio data to the user through the speaker.
[0891] Examples of conversation partner functions
[0892] User: "I'm feeling sad today."
[0893] Step 1: The device acquires the user's voice and converts it into text data, "I feel sad today," using speech recognition technology.
[0894] Step 2: The emotion engine recognizes sad emotions from the user's voice.
[0895] Step 3: Send text data and sentiment data to the server.
[0896] Step 4: The server analyzes the received data and generates an encouraging response such as, "What's wrong? Tell me what happened."
[0897] Step 5: The server sends the generated response to the terminal.
[0898] Step 6: The device converts the response data into speech and outputs to the user, "What's the matter? Please tell me what's wrong."
[0899] Emergency contact function
[0900] Step 1:
[0901] Device: The device continuously acquires biometric data such as heart rate, body temperature, and movement through built-in biosensors.
[0902] Step 2:
[0903] Terminal: Uses an anomaly detection method to monitor acquired biometric data in real time and detect abnormal values.
[0904] Step 3:
[0905] Terminal: If an anomaly is detected, an emergency notification will be generated. This notification will include the type of anomaly and the user's identification information.
[0906] Step 4:
[0907] Terminal: Sends the generated emergency notification to the server using the emergency notification method.
[0908] Step 5:
[0909] Server: Receives emergency notifications and sentiment data, and retrieves a pre-configured list of emergency contacts.
[0910] Step 6:
[0911] Server: Sends an emergency message to emergency contacts based on the information included in the emergency notification. Email, SMS, and phone APIs are often used for sending messages.
[0912] Step 7:
[0913] User (recipient): The family or medical institution will review the received emergency message and begin taking immediate action.
[0914] Examples of emergency contact functions
[0915] Step 1: The device detects the user's abnormal heart rate and cessation of movement.
[0916] Step 2: Send anomaly and emotion data to the server.
[0917] Step 3: The server receives an emergency notification and sends an emergency message to the emergency contact stating, "The user has collapsed. Please respond immediately."
[0918] Step 4: The family receives the emergency message and immediately begins taking action.
[0919] The above outlines the program's processing flow, including the specific actions performed at each processing step. This detailed explanation will allow for a clearer understanding of the overall system's operation.
[0920] (Example 2)
[0921] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0922] When elderly people live alone, there are risks associated with a lack of communication and delayed responses in emergencies. Furthermore, systems that cannot engage in natural conversations that take emotions into consideration can increase the user's mental burden and potentially lower their quality of life. Additionally, if biometric data monitoring and emergency notifications are not effectively implemented, prompt responses become difficult.
[0923] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: speech recognition means for acquiring voice from a terminal in real time and converting the voice into text; emotion recognition means for analyzing emotions from the terminal and generating emotion data; transmission means for the terminal to send the converted text data and emotion data to the server; analysis means for the server to analyze the received text data and emotion data and generate an appropriate response; response transmission means for the server to send the generated response to the terminal; and speech synthesis means for converting the response received by the terminal into voice and outputting it to the user. This enables natural and effective communication with the elderly, allowing for a quick response in emergencies and improving the quality of life for users.
[0924] A "terminal" is a computer device used by users for voice input and output, as well as the collection of biometric data.
[0925] "Voice recognition means" refers to technology or devices that acquire a user's voice in real time on a terminal and convert that voice into text data.
[0926] "Emotion recognition means" refers to technologies and devices that analyze a user's emotions from their voice, facial expressions, etc., and generate emotion data.
[0927] "Transmission means" refers to communication technologies and devices used to send text data and sentiment data from a terminal to a server.
[0928] A "server" is a centralized management system that receives data sent from terminals, analyzes it, and generates appropriate responses.
[0929] "Analysis means" refers to technologies and devices that analyze text data and sentiment data on a server and generate appropriate responses for the user.
[0930] "Response transmission means" refers to communication technology or equipment used to send a response generated by a server to a terminal.
[0931] "Speech synthesis means" refers to a technology or device that converts response data received from a server into speech data at a terminal and outputs it to the user as speech.
[0932] "Methods for acquiring biometric data" refer to technologies and devices that continuously acquire biometric data such as a user's heart rate, body temperature, and movement.
[0933] An "anomaly detection method" refers to a technology or device that monitors acquired biological data in real time and detects anomalies.
[0934] An "emergency notification system" is a technology or device that sends an emergency notification to a server when an anomaly is detected.
[0935] "Communication methods" refer to technologies and devices that, after a server receives an emergency notification, send an emergency message to pre-configured emergency contacts.
[0936] This invention relates to a system that monitors elderly people living alone, facilitates communication, and enables rapid response in emergencies. In particular, this invention achieves more natural and effective communication by combining it with an emotion engine that recognizes the user's emotions.
[0937] Hardware and software to be used
[0938] terminal
[0939] The device includes the following hardware and software:
[0940] High-sensitivity microphone (voice input)
[0941] Speaker (audio output)
[0942] Biosensors (acquisition of heart rate, body temperature, and movement data)
[0943] Speech recognition engine (e.g., Google Speech-to-Text API)
[0944] Emotion recognition engine (e.g., Microsoft Azure Emotion API)
[0945] Speech synthesis engine (e.g., Amazon Polly)
[0946] server
[0947] The server includes the following software and features:
[0948] Analysis engine (Python library using NLP technology)
[0949] Response generation engine
[0950] Emergency notification system
[0951] Data receiving and transmitting module
[0952] System operation
[0953] terminal
[0954] When a user speaks, the device acquires the voice through its built-in microphone. The acquired voice data is converted into text data using a speech recognition engine. Simultaneously, an emotion recognition engine analyzes emotions from voice and facial expressions, generating emotion data. This data is transmitted to the server via the device's transmission mechanism.
[0955] server
[0956] The server receives text data and sentiment data sent from the terminal. The received data is analyzed using an analysis engine. Based on the analysis results, the server generates an appropriate response. The generated response is sent to the terminal via a response transmission means.
[0957] terminal
[0958] The terminal converts responses received from the server into speech data using a speech synthesis engine and outputs it to the user through a speaker. This allows elderly people to interact with the system in a natural way.
[0959] Specific example
[0960] Conversation partner function
[0961] User: "I'm feeling sad today."
[0962] Device: It uses the built-in microphone to acquire voice data and converts it into text data, such as "I feel sad today," using a speech recognition engine. The emotion engine then analyzes the user's voice to determine that they are feeling "sad."
[0963] Terminal: Sends the converted text data and sentiment data to the server.
[0964] Server: Analyzes the received data and generates encouraging responses such as, "What's wrong? Please tell me what happened."
[0965] Server: Sends the generated response data to the terminal.
[0966] Terminal: Converts response data into speech and outputs to the user, "What's the matter? Please tell me what happened."
[0967] Emergency contact function
[0968] Device: If the user falls, the built-in biosensors detect abnormal heart rate or unresponsive movement.
[0969] Terminal: Immediately sends emergency notifications and emotional data to the server.
[0970] Server: Upon receiving an emergency notification, it sends the message "User has collapsed. Please respond immediately" to the pre-configured emergency contact.
[0971] Recipient user: The family receives the emergency message and immediately begins taking action.
[0972] Example of a prompt
[0973] Example of prompt text when inputting a "monitoring system for elderly people living alone" into the AI model:
[0974] "The device acquires the user's voice and converts it into text data using a speech recognition engine. Additionally, an emotion engine analyzes the user's emotions and sends this data to a server. The server analyzes the received data, generates an appropriate response, and sends it to the device, which then responds verbally. Furthermore, in emergencies, a biosensor detects an anomaly and sends an emergency notification through the server. Please explain this system."
[0975] Thus, the system of the present invention promotes natural dialogue with the elderly and enables a rapid response in emergencies, thereby improving the safety and quality of life of the elderly.
[0976] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0977] Step 1:
[0978] The device acquires the user's voice. Using the built-in high-sensitivity microphone, it collects voice data when the user says, "I'm feeling sad today."
[0979] Input: User's voice
[0980] Output: Audio data
[0981] Specific operation: The device's operating system converts the audio signal from the microphone into digital data and temporarily stores it in memory.
[0982] Step 2:
[0983] The device uses a speech recognition engine to convert audio data into text data. For example, the speech recognition engine uses the Google Speech-to-Text API.
[0984] Input: Audio data
[0985] Output: Text data
[0986] Specific operation: The speech recognition engine analyzes the audio data, identifies phonemes and morphemes, and generates the text "I feel sad today." The generated text is stored in a separate buffer in memory.
[0987] Step 3:
[0988] The device uses an emotion recognition engine to generate emotion data from voice and text data. For example, it can utilize the Microsoft Azure Emotion API.
[0989] Input: Audio data, text data
[0990] Output: Sentiment data
[0991] Specific operation: The emotion recognition engine analyzes the emotion of "sadness" from the tone of voice and associated text, and generates this emotion data. The generated emotion data is stored in memory along with the text data.
[0992] Step 4:
[0993] The terminal sends the converted text data and sentiment data to the server. This is done using a network transmission module.
[0994] Input: Text data, sentiment data
[0995] Output: Transmitted data
[0996] Specific operation: The transmission module packets text and sentiment data and sends them to the server via the Wi-Fi module over the internet. The data is encrypted to ensure security.
[0997] Step 5:
[0998] The server receives text data and sentiment data sent from the terminal.
[0999] Input: Data to send
[1000] Output: Received data
[1001] Specific operation: The server's receiving module receives packets, decrypts the encrypted data, and obtains text data and sentiment data. This data is stored in a database for analysis.
[1002] Step 6:
[1003] The server analyzes text and sentiment data and generates an appropriate response using an analysis engine.
[1004] Input: Text data, sentiment data
[1005] Output: Response data
[1006] Specific operation: The analysis engine uses NLP technology to analyze text data and sentiment data, and generates the most appropriate response for the user's state, such as "What's wrong? Please tell me what happened."
[1007] Step 7:
[1008] The server sends the generated response data to the terminal. A response transmission method is used.
[1009] Input: Response data
[1010] Output: Transmitted data
[1011] Specific operation: The response data is packetized and sent to the terminal via the network transmission module over the internet. The data is transmitted encrypted.
[1012] Step 8:
[1013] The terminal converts the received response data into audio data and outputs it to the user through the speaker. A speech synthesis engine is used.
[1014] Input: Response data
[1015] Output: Audio data
[1016] Specific operation: The speech synthesis engine analyzes the response data and generates the voice message, "What's the matter? Please tell me what's wrong." The generated voice message is output to the user through the speaker.
[1017] Step 9:
[1018] The device monitors biometric data, using sensors such as heart rate sensors, body temperature sensors, and motion sensors.
[1019] Input: Biometric data
[1020] Output: Acquired data
[1021] Specific operation: The sensor continuously measures the user's heart rate, body temperature, and movement, and records this data in memory. The data is updated in real time, waiting for anomaly detection.
[1022] Step 10:
[1023] If the device detects an anomaly, it sends an emergency notification to the server. For example, it might detect a sudden increase in heart rate or a cessation of movement.
[1024] Input: Acquired data
[1025] Output: Emergency notification data
[1026] Specific operation: When the anomaly detection algorithm detects a sudden increase in heart rate, abnormal fluctuations in body temperature, or cessation of movement, it generates emergency notification data related to these events and sends it to the server.
[1027] Step 11:
[1028] The server receives an emergency notification and sends an emergency message to pre-configured emergency contacts.
[1029] Input: Emergency notification data
[1030] Output: Emergency message
[1031] Specific operation: The server analyzes the received emergency notification data and sends a message stating "User has collapsed. Please respond immediately" to a pre-configured list of emergency contacts (e.g., family or medical institutions). The message is sent via SMS or a dedicated application.
[1032] (Application Example 2)
[1033] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[1034] In recent years, the number of elderly people living alone has increased, making their safety and security a major concern. There is also a growing need for systems that alleviate their feelings of loneliness and provide rapid responses in emergencies. However, current systems lack sufficient emotion recognition and emergency notification capabilities, and cannot fully guarantee the psychological and physical safety of the elderly. Therefore, there is a need for natural communication methods that include emotion recognition, as well as rapid emergency response through real-time monitoring of biometric data.
[1035] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a voice recognition means, a transmission means, an analysis means, a response transmission means, a voice synthesis means, and a means for recognizing emotions from the user's voice and facial expressions, equipped with an emotion recognition engine. This enables natural dialogue that takes into account the emotions of the elderly, and also enables a rapid response in emergencies. It also has an emergency notification function and can provide a system that continuously monitors heart rate, body temperature, and movement data, and immediately makes an emergency contact when an abnormality is detected.
[1036] A "terminal" is a device used by a user that has functions such as voice input, speech synthesis, emotion recognition, and acquisition of biometric data.
[1037] "Voice recognition means" refers to a function that converts the voice acquired by the terminal into text data.
[1038] "Transmission method" refers to the function by which the terminal sends the converted text data to the server.
[1039] "Analysis means" refers to a function that analyzes text data and sentiment data received by the server and generates an appropriate response.
[1040] "Response transmission means" refers to the function that sends the response generated by the server to the terminal.
[1041] A "speech synthesis means" is a function that converts the response received by the terminal into speech and outputs it to the user.
[1042] An "emotion recognition engine" is a function that recognizes emotions from the user's voice and facial expressions.
[1043] "Means of acquiring biometric data" refers to a function in which the terminal continuously acquires the user's heart rate, body temperature, and movement.
[1044] An "anomaly detection method" is a function that monitors acquired biometric data in real time and detects sudden increases in heart rate, abnormal fluctuations in body temperature, and cessation of movement.
[1045] An "emergency notification mechanism" is a function that sends an emergency notification to the server when an anomaly is detected.
[1046] "Communication method" refers to the function where the server receives an emergency notification and sends an emergency message to a pre-configured emergency contact.
[1047] This invention is a system that monitors elderly people living alone, facilitates communication, and enables rapid response in emergencies. The system mainly consists of terminals and a server.
[1048] Device functions
[1049] The terminal is equipped with a speech recognition system, a speech synthesis system, an emotion recognition engine, and a biometric data acquisition system. When a user speaks, the voice is acquired through the terminal's microphone and converted into text data by the speech recognition system. This text data is then sent to the server via a transmission system.
[1050] The device also features an emotion recognition engine that recognizes emotions from the user's voice and facial expressions. This emotion data is also sent to the server along with the text data.
[1051] Furthermore, the device is equipped with biometric data acquisition capabilities, continuously monitoring the user's heart rate, body temperature, and movement. If an abnormality is detected, a notification is sent to the server via an emergency notification system.
[1052] Server Functions
[1053] The server has an analysis mechanism to analyze the received text data and sentiment data, and generates an appropriate response. The generated response is sent back to the terminal via a response transmission mechanism. The terminal converts this response into speech using a speech synthesis mechanism and outputs it to the user through a speaker.
[1054] The server has a communication system that, upon receiving an emergency notification, sends an emergency message to pre-configured emergency contacts.
[1055] Hardware and software details
[1056] The system uses the "Speech Recognition API" for speech recognition, the "Network Communication Module" for text data transmission, and the "Emotion Recognition API" for emotion recognition. Additionally, it uses the "Speech Synthesis API" for speech synthesis and the "Biometric Sensor Module" for remote monitoring. The "Notification API" is used for emergency notifications and communication.
[1057] Specific example
[1058] When a user says, "I'm not feeling well today," the device converts the audio data into text data and analyzes it using an emotion recognition engine. The resulting data is sent to a server, which considers the user's emotions and generates an encouraging response such as, "What's wrong? Please tell me what's the matter." The generated response is sent to the device and output to the user as audio.
[1059] Furthermore, if a user collapses, the biosensors built into the device detect abnormal heart rate or unresponsive movement and immediately send an emergency notification to the server. The server receives the emergency notification and sends a message to emergency contacts stating, "A user has collapsed. Please respond quickly."
[1060] Example of a prompt
[1061] "Provide an example where, if a user says 'I'm not feeling well today,' the audio is converted to text, analyzed by an emotion engine, and an appropriate comforting response is generated."
[1062] As described above, this system supports the safe and secure lives of the elderly and can provide prompt assistance when needed. Furthermore, its emotion recognition function enables more natural and effective communication.
[1063] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1064] Step 1:
[1065] The user speaks to the device.
[1066] Input: User's voice data.
[1067] Specific action: The device's microphone picks up the user's voice.
[1068] Step 2:
[1069] The device converts the audio into text data.
[1070] Input: Acquired audio data.
[1071] Specific operation: Use a speech recognition API to convert speech data into text data.
[1072] Output: Text data.
[1073] Step 3:
[1074] The device performs emotion recognition.
[1075] Input: User's voice data and text data.
[1076] Specific operation: Use an emotion recognition API to analyze the user's emotions from their voice and facial expressions.
[1077] Output: Sentiment data.
[1078] Step 4:
[1079] The device sends text data and sentiment data to the server.
[1080] Input: Text data and sentiment data.
[1081] Specific operation: Use the network communication module to send data to the server.
[1082] Output: Confirmation of data transmission to the server.
[1083] Step 5:
[1084] The server analyzes the text data and sentiment data it receives.
[1085] Input: Text data and sentiment data.
[1086] Specific operation: Use analytical tools to perform analysis in order to generate an appropriate response.
[1087] Output: Appropriate response data.
[1088] Step 6:
[1089] The server sends the appropriate response data to the terminal.
[1090] Input: Generated response data.
[1091] Specific operation: Use the response transmission means to send response data to the terminal.
[1092] Output: Confirmation of data transmission to the terminal.
[1093] Step 7:
[1094] The terminal converts the received response data into speech.
[1095] Input: Appropriate response data.
[1096] Specific operation: Use a speech synthesis API to convert text data into speech data.
[1097] Output: Audio data.
[1098] Step 8:
[1099] The device outputs audio data to the user.
[1100] Input: Audio data.
[1101] Specific action: Play audio through the device's speaker.
[1102] Output: Audio output to the user.
[1103] Step 9:
[1104] The device acquires biometric data.
[1105] Input: Biometric data such as the user's heart rate, body temperature, and movement.
[1106] Specific operation: Use a biosensor module to continuously collect data.
[1107] Output: Biometric data.
[1108] Step 10:
[1109] The device monitors biometric data acquired in real time.
[1110] Input: Biometric data.
[1111] Specific operation: Use anomaly detection methods to monitor data for abnormalities.
[1112] Output: Abnormal status data.
[1113] Step 11:
[1114] If an anomaly is detected, the device will send an emergency notification to the server.
[1115] Input: Abnormal situation data.
[1116] Specific action: Use the emergency notification method to send an anomaly notification to the server.
[1117] Output: Sending an emergency notification to the server.
[1118] Step 12:
[1119] The server receives an emergency notification and sends a message to the emergency contact.
[1120] Input: Anomaly notification data.
[1121] Specific action: Use the communication method to send a message to a pre-configured emergency contact.
[1122] Output: Confirmation of emergency message transmission.
[1123] The above describes the processing steps of the system based on this invention.
[1124] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1125] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1126] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[1127] [Third Embodiment]
[1128] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[1129] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[1130] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1131] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[1132] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1133] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1134] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1135] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1136] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1137] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1138] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1139] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[1140] This invention relates to a system for monitoring elderly people living alone, facilitating communication with them using voice and biometric data, and enabling a rapid response in emergencies. Embodiments of this invention are described in detail below.
[1141] Speech recognition and input processing
[1142] terminal
[1143] When a user speaks into the device, the device acquires the voice through its built-in microphone. The acquired voice data is converted into text data by a speech recognition system. This text data is then sent to the server through the device's transmission system.
[1144] server
[1145] The server receives text data sent from the terminal. The received data is analyzed by an analysis device, and an appropriate response is generated based on the user's intent. The generated response is then sent back to the terminal via the server's response transmission device.
[1146] terminal
[1147] The terminal receives responses from the server, converts them into speech data using speech synthesis technology, and outputs them to the user through a speaker. This allows elderly people to interact with the system in a natural way.
[1148] Examples of conversation partner functions
[1149] User: "What kind of day is it today?"
[1150] Terminal: Acquires audio and converts it into text data, such as "What kind of day is it today?", using speech recognition technology.
[1151] Terminal: Sends text data to the server.
[1152] Server: Analyzes the received text data and generates a response such as "Today is not a special day."
[1153] Server: Sends the generated response to the terminal.
[1154] Terminal: Converts the response data into speech and outputs "Today is not a special day" to the user.
[1155] Emergency contact function
[1156] terminal
[1157] The device has built-in biosensors that continuously acquire the user's heart rate, body temperature, and movement. This biometric data is constantly monitored, and if an abnormality is detected, an emergency notification is generated.
[1158] server
[1159] Emergency notifications are sent from the terminal to the server. The server analyzes the emergency notification and sends an emergency message to pre-registered emergency contacts. The emergency message includes the user ID, terminal ID, and abnormal data.
[1160] Examples of emergency contact functions
[1161] Device: If the user falls and becomes motionless, it detects abnormal heart rate and unresponsive movement.
[1162] Terminal: Immediately send an emergency notification to the server.
[1163] Server: Receives emergency notification and sends SMS to pre-configured family contacts. "User has collapsed. Please respond immediately."
[1164] User (recipient): The family receives the emergency message and immediately begins taking action.
[1165] In this way, the system of the present invention can ensure the safety and communication of elderly people living alone.
[1166] The following describes the processing flow.
[1167] Speech recognition and input processing
[1168] Step 1:
[1169] Terminal: The terminal acquires the user's voice in real time through its built-in microphone.
[1170] Step 2:
[1171] Terminal: Transmits the acquired audio data to a speech recognition device (e.g., a speech recognition engine) and converts the audio data into text data.
[1172] Step 3:
[1173] Terminal: Sends the converted text data to the server using a transmission method. HTTP requests are often used for transmission.
[1174] Conversation partner function
[1175] Step 4:
[1176] Server: Analyzes the received text data using parsing tools (such as a natural language processing engine).
[1177] Step 5:
[1178] Server: Generates an appropriate response based on the analysis results. The response may include fixed statements or information obtained from external APIs.
[1179] Step 6:
[1180] Server: Sends the generated response to the terminal via a response transmission method. JSON format data is often used for transmission.
[1181] Step 7:
[1182] Terminal: Converts the received response text data into speech data using a text-to-speech engine.
[1183] Step 8:
[1184] Terminal: Outputs audio data to the user through the speaker.
[1185] Emergency contact function
[1186] Step 1:
[1187] Device: The device continuously acquires biometric data such as heart rate, body temperature, and movement through built-in biosensors.
[1188] Step 2:
[1189] Terminal: Uses an anomaly detection method to monitor acquired biometric data in real time and detect abnormal values.
[1190] Step 3:
[1191] Terminal: If an anomaly is detected, an emergency notification will be generated. This notification will include the type of anomaly and the user's identification information.
[1192] Step 4:
[1193] Terminal: Sends the generated emergency notification to the server using the emergency notification method.
[1194] Step 5:
[1195] Server: Receives emergency notifications and retrieves a pre-configured list of emergency contacts.
[1196] Step 6:
[1197] Server: Sends an emergency message to emergency contacts based on the information included in the emergency notification. Email, SMS, and phone APIs are often used for sending messages.
[1198] Step 7:
[1199] User (recipient): The family or medical institution will review the received emergency message and begin taking immediate action.
[1200] The above is the program's processing flow, including the specific actions performed at each processing step. Providing such a detailed explanation allows for a clearer understanding of how the entire system works.
[1201] (Example 1)
[1202] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1203] In modern society, elderly people living alone face challenges such as a lack of daily communication and the inability to respond quickly in emergencies. This has led to feelings of loneliness among the elderly and delays in emergency response. The present invention aims to provide a system that enables elderly people to engage in daily conversations in a natural manner and further enables a rapid response in emergencies using biometric data.
[1204] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1205] In this invention, the server includes: speech recognition means for acquiring voice from a terminal in real time and converting the voice into text; transmission means for the terminal to send the converted text data to the server; analysis means for the server to analyze the received text data and generate an appropriate response using natural language processing means; response transmission means for the server to send the generated response to the terminal; speech synthesis means for the terminal to convert the received response into voice and output it to the user; anomaly detection means for monitoring biometric data in real time and detecting anomalies; emergency notification means for sending an emergency notification to the server when an anomaly is detected; and communication means for the server to receive the emergency notification and automatically send an emergency message to a pre-configured emergency contact. This enables elderly people to communicate naturally using voice on a daily basis, and also enables a rapid response in emergencies using biometric data.
[1206] A "terminal" is a device that functions as an interface with the user, collecting, converting, transmitting, and outputting responses for voice data.
[1207] "Speech recognition means" refers to a technology that converts speech into text data, is built into a terminal, and has the function of recognizing the user's voice in real time.
[1208] "Transmission means" refers to the communication means used by a terminal to send converted text data to a server, and includes Wi-Fi modules and internet connection means.
[1209] A "server" is a central control unit that receives data sent from terminals and performs processing, analysis, and response generation.
[1210] "Natural language processing means" refers to technology that analyzes received text data, understands the user's intent, and generates an appropriate response.
[1211] "Analysis means" refers to the technology used by a server to process received text data and generate an appropriate response using natural language processing means.
[1212] "Response transmission means" refers to communication technology used to send responses generated by a server to a terminal.
[1213] "Speech synthesis means" refers to a technology that converts text data received from a server into speech data and outputs a speech response to the user.
[1214] "Methods for acquiring biometric data" refers to sensor technology that continuously acquires data on the user's heart rate, body temperature, and movement.
[1215] "Anomaly detection means" refers to technology that monitors biological data in real time and detects anomalies.
[1216] An "emergency notification system" refers to a technology that immediately sends an emergency notification to a server when an anomaly is detected.
[1217] "Communication method" refers to the technology in which a server receives an emergency notification and automatically sends an emergency message to pre-configured emergency contacts.
[1218] A "generative AI model" refers to artificial intelligence technology used to generate natural dialogue and appropriate responses in abnormal situations.
[1219] A "prompt statement" refers to an input statement used to give instructions to a generative AI model and is used to generate a specific response.
[1220] This invention relates to a system for monitoring elderly people living alone. It facilitates communication with the elderly using voice and biometric data, and enables a rapid response in emergencies. The embodiments of this invention are described in detail below.
[1221] Speech recognition and input processing
[1222] terminal
[1223] When a user speaks into the device, the device acquires the voice through its built-in microphone. The acquired voice data is converted into text data using speech recognition tools (e.g., Google Cloud Speech-to-Text API). This text data is then sent to the server via the device's transmission tools (Wi-Fi module).
[1224] server
[1225] The server uses a Python framework (e.g., Flask or Django) to receive text data sent from the terminal. The received data is parsed by a natural language processing tool (e.g., spaCy) to generate an appropriate response based on the user's intent. The generated response is then sent back to the terminal via the server's response sending mechanism.
[1226] terminal
[1227] The device uses the Google Cloud Text-to-Speech API to convert responses received from the server into audio data, which is then output to the user through the speaker. This allows elderly users to interact with the system in a natural way.
[1228] Specific example
[1229] User: "What kind of day is it today?"
[1230] Terminal: Acquires audio and converts it into text data, such as "What kind of day is it today?", using speech recognition technology.
[1231] Terminal: Sends text data to the server.
[1232] Server: Analyzes the received text data and generates a response such as "Today is not a special day."
[1233] Server: Sends the generated response to the terminal.
[1234] Terminal: Converts the response data into speech and outputs "Today is not a special day" to the user.
[1235] Emergency contact function
[1236] terminal
[1237] The device has built-in biosensors (heart rate sensor, body temperature sensor, motion sensor, etc.) that continuously acquire the user's heart rate, body temperature, and movement. This biometric data is constantly monitored by a microcontroller (e.g., ESP32), and if an anomaly is detected, an emergency notification is generated.
[1238] server
[1239] Emergency notifications are sent from the device to the server. The server analyzes the emergency notification and sends an emergency message to pre-registered emergency contacts using the Twilio API. The emergency message includes the user ID, device ID, and anomaly data.
[1240] Specific example
[1241] If a user collapses and becomes motionless, the system detects abnormal heart rate and unresponsive movement.
[1242] Terminal: Immediately send an emergency notification to the server.
[1243] Server: Receives emergency notification and sends an SMS to pre-configured family contacts using the Twilio API. "User has collapsed. Please respond quickly."
[1244] Examples of prompt statements
[1245] The following are specific examples of prompt statements used in generative AI models.
[1246] In the case of an elderly monitoring system:
[1247] Create text to generate natural conversations with users.
[1248] For example, generate a response to the user's question, "What kind of day is it today?"
[1249] By using these prompts, it becomes possible to provide a system that enables natural interaction with users and allows for rapid response in emergencies.
[1250] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1251] Program processing for speech recognition and input processing
[1252] Step 1:
[1253] terminal
[1254] The user speaks into the device, and the voice is captured through the built-in microphone. This voice data is the input. The captured voice data is sent in real time to a speech recognition system (e.g., Google Cloud Speech-to-Text API) and converted into text data. The data processing here is the conversion from voice to text, and the output is text data.
[1255] Specific actions
[1256] User: "What kind of day is it today?"
[1257] Device: The built-in microphone captures audio and sends it to the speech recognition API.
[1258] Step 2:
[1259] terminal
[1260] The system acquires text data converted by speech recognition and sends it to a server. The input here is the converted text data, and the output is the text data sent to the server. Transmission takes place via the internet through a Wi-Fi module.
[1261] Specific actions
[1262] Terminal: The device sends the text data "What kind of day is it today?" obtained from the voice recognition device to the server via the Wi-Fi module.
[1263] Step 3:
[1264] server
[1265] The server receives text data sent from the terminal using a Python framework (e.g., Flask or Django). The input here is text data, which is parsed using natural language processing tools (e.g., spaCy). The output is the appropriate response resulting from the parsing.
[1266] Specific actions
[1267] Server: Analyzes the text data "What kind of day is it today?" and uses a natural language processing library to understand the user's intent.
[1268] Step 4:
[1269] server
[1270] Based on the analysis results, an appropriate response is generated using a generative AI model (e.g., GPT-3). The input is the analysis results, and the output is the generated response text.
[1271] Specific actions
[1272] Server: Generates responses such as "Today is not a special day" using an AI model based on the analysis results from a natural language processing library.
[1273] Step 5:
[1274] server
[1275] The generated response is sent to the terminal via the response transmission means. The input is the generated response text, and the output is the text data sent to the terminal.
[1276] Specific actions
[1277] Server: Sends the response text "Today is not a special day" to the terminal.
[1278] Step 6:
[1279] terminal
[1280] The system converts the response received from the server into speech data using a text-to-speech synthesis tool (e.g., Google Cloud Text-to-Speech API). The input is the response text, and the output is speech data. This is output to the user through the built-in speaker.
[1281] Specific actions
[1282] Terminal: Receives the response text "Today is not a special day," converts it into speech data using a speech synthesis method, and outputs it through the speaker.
[1283] Processing of the emergency contact function program
[1284] Step 1:
[1285] terminal
[1286] The device's built-in biosensors (heart rate sensor, body temperature sensor, motion sensor, etc.) continuously acquire the user's biometric data. The input is data from the biosensors, and the output is the continuously acquired biometric data. This data is sent to a microcontroller (e.g., ESP32) for monitoring.
[1287] Specific actions
[1288] Device: The user's heart rate sensor measures a heart rate of 60 BPM.
[1289] Step 2:
[1290] terminal
[1291] The system monitors acquired biometric data and detects anomalies if abnormal values exceed a threshold. The input is continuously acquired biometric data, which is analyzed by an anomaly detection mechanism. The output indicates whether or not an anomaly exists.
[1292] Specific actions
[1293] Terminal: If a heart rate of 30 BPM (below the abnormal capture threshold) is detected, it will be flagged as abnormal.
[1294] Step 3:
[1295] terminal
[1296] When an anomaly is detected, an emergency notification is immediately generated. The input is the result of the anomaly detection, and the output is the emergency notification data. The generated emergency notification includes the user ID and anomaly data.
[1297] Specific actions
[1298] The device detects an abnormal heart rate and generates data that reads "Emergency notification: User ID 123, heart rate 30 BPM".
[1299] Step 4:
[1300] terminal
[1301] The generated emergency notification is sent to the server. The input is the emergency notification data, and the output is the emergency notification sent to the server. Transmission is performed via the Wi-Fi module.
[1302] Specific actions
[1303] Device: Emergency notifications are sent to the server via Wi-Fi.
[1304] Step 5:
[1305] server
[1306] This system receives emergency notifications and automatically sends emergency messages to pre-registered emergency contacts via the Twilio API. The input is the received emergency notification data, and the output is the sent emergency message. The emergency message includes user ID, device ID, and anomaly data.
[1307] Specific actions
[1308] Server: Upon receiving an emergency notification, it sends an SMS message using the Twilio API stating, "A user has collapsed. Please respond quickly."
[1309] As described above, the system of the present invention enables natural communication with the elderly using voice and biometric data, as well as a rapid response in emergencies.
[1310] (Application Example 1)
[1311] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1312] The problem that this invention aims to solve is to provide an environment in which elderly people living alone can live safely and securely. In particular, it aims to enable rapid response in emergencies by linking voice recognition and biometric data, and to monitor changes in the elderly person's physical condition and abnormal situations in real time. Another challenge is to improve communication with the elderly person through a user-friendly interface.
[1313] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1314] In this invention, the server includes an analysis means for analyzing text data and generating an appropriate response, a communication means for receiving emergency notifications and sending emergency messages to pre-configured emergency contacts, and a means for generating responses using an AI model. This enables elderly people living alone to communicate naturally by voice and to respond quickly in emergencies.
[1315] A "terminal" is a device used by a user that has the function of acquiring voice data and sending it to a server, as well as the function of acquiring biometric data.
[1316] "Voice recognition means" refers to technology that converts speech acquired by a terminal into text data.
[1317] "Transmission method" refers to the function by which the terminal sends the converted text data to the server.
[1318] "Analysis means" refers to a technology that analyzes text data received by a server and generates an appropriate response based on the user's intent.
[1319] "Response transmission means" refers to the function by which the server sends the generated response to the terminal.
[1320] "Speech synthesis means" refers to a technology that converts responses received by a terminal from a server into speech data and outputs it to the user as speech.
[1321] "Means of acquiring biometric data" refers to a function in which a device continuously acquires biometric data such as the user's heart rate, body temperature, and movement.
[1322] An "anomaly detection method" is a technology that monitors acquired biometric data in real time and generates an emergency notification when an anomaly is detected.
[1323] An "emergency notification mechanism" is a function that allows a terminal to send an emergency notification to a server when it detects an anomaly.
[1324] "Communication method" refers to a technology where a server receives an emergency notification and sends an emergency message to pre-configured emergency contacts.
[1325] An "AI model" is an artificial intelligence algorithm used by a server to analyze user text data and generate appropriate responses.
[1326] A "prompt message" is input data that is dynamically generated based on user inquiries used by the generative AI model when performing analysis.
[1327] This invention is a system for monitoring elderly people living alone, utilizing voice recognition and biometric data to facilitate communication with them and enable rapid response in emergencies. Details of the invention are shown below.
[1328] System Configuration
[1329] The system primarily consists of "terminals" and "servers." Their respective functions and specific examples are explained below.
[1330] (Device functions)
[1331] 1. Speech recognition means
[1332] The device uses its built-in microphone to capture audio in real time and converts the captured audio into text using speech recognition software (e.g., the speech_recognition library). For example, if a user says "What kind of day is it today?", the audio is converted into the text data "What kind of day is it today?".
[1333] 2. Transmission method
[1334] The terminal sends the converted text data to the server. Therefore, a wireless communication module (e.g., a Wi-Fi module) is required.
[1335] 3. Speech synthesis means
[1336] The terminal converts the response received from the server into audio data using speech synthesis software (e.g., the playsound library) and outputs it to the user through the built-in speaker. For example, if the terminal receives the response "Today is not a special day" from the server, it will convey that to the user as audio.
[1337] 4. Means of acquiring biometric data
[1338] The device uses built-in biosensors (e.g., heart rate sensor, temperature sensor, motion sensor) to continuously acquire data on the user's heart rate, body temperature, and movement.
[1339] 5. Anomaly detection means
[1340] The device monitors acquired biometric data in real time and detects abnormalities. For example, a sudden increase in heart rate or a lack of movement may be considered an abnormality.
[1341] 6. Emergency notification means
[1342] If an anomaly is detected, the terminal immediately sends an emergency notification to the server. A wireless communication module is required for this.
[1343] (Server functions)
[1344] 1. Analysis method
[1345] The server analyzes the received text data and generates an appropriate response. A generative AI model (e.g., the GPT model) is used in this process. For example, in response to the text data "What kind of day is it today?", it generates the response "Today is not a special day."
[1346] 2. Means of contact
[1347] When the server receives an emergency notification, it sends an emergency message to pre-configured emergency contacts. It also includes a mechanism to output the generated emergency message as audio and notify the user.
[1348] 3. Response generation means using an AI model
[1349] The server's analysis method uses a generative AI model to generate appropriate responses based on the user's text data. For example, by inputting a prompt into the generative AI model, a response is dynamically generated.
[1350] Specific example
[1351] The user is wearing smart glasses at home.
[1352] User: "What kind of day is it today?"
[1353] Terminal: Acquires audio, converts it into text data such as "What kind of day is it today?" using speech recognition, and sends it to the server.
[1354] Server: Analyzes the received text data, generates a response saying "Today is not a special day," and sends it to the terminal.
[1355] Terminal: The response from the server is converted into speech data using a speech synthesis method, and the message "Today is not a special day" is output to the user.
[1356] If a user suddenly collapses, the device's biosensors detect a sudden increase in heart rate and cessation of movement, and send an emergency notification to the server.
[1357] Server: Upon receiving an emergency notification, it sends an emergency message to a pre-configured emergency contact (e.g., family member) stating, "The user has collapsed. Please respond immediately."
[1358] Examples of prompts for generative AI models
[1359] User: "What kind of day is it today?"
[1360] The application responded, "Today is not a special day."
[1361] If the user's heart rate suddenly increases afterward, how will an emergency notification be sent?
[1362] Through these components and processes, the present invention provides a system that guarantees the safety and security of the elderly.
[1363] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1364] Step 1:
[1365] The user speaks into the device. The device acquires the voice through its built-in microphone. The acquired voice data is sent as input to a speech recognition system, where the voice is converted into text data.
[1366] Step 2:
[1367] The terminal sends the converted text data to the server via a transmission means. Here, as part of the data processing, the text data is transmitted to the server in an appropriate format.
[1368] Step 3:
[1369] The server analyzes the received text data using an analysis tool. The analysis tool uses a generative AI model to generate an appropriate response based on the input text data. Specifically, the data calculation involves generating a prompt sentence, inputting it into the AI model, and generating a text response as a result.
[1370] Step 4:
[1371] The server sends the generated response back to the terminal via the response transmission means. The data transmitted is a text-based response.
[1372] Step 5:
[1373] The terminal converts the response received from the server into audio data using a speech synthesis system and outputs it to the user through its built-in speaker. The output audio is based on the response content generated by the server.
[1374] Step 6:
[1375] The device continuously acquires data on the user's heart rate, body temperature, and movement using built-in biosensors. This biometric data is sent as input to an anomaly detection system. Here, the biometric data is analyzed in real time as part of the data processing.
[1376] Step 7:
[1377] The anomaly detection system monitors acquired biometric data in real time and detects whether there are any abnormalities. For example, a sudden increase in heart rate or a lack of movement will be detected as an anomaly. This detection is the result of data analysis.
[1378] Step 8:
[1379] If an anomaly is detected, the terminal immediately sends an emergency notification to the server using the emergency notification mechanism. This notification includes the user ID, terminal ID, and detected anomaly data.
[1380] Step 9:
[1381] The server receives an emergency notification and sends an emergency message to pre-configured emergency contacts via the designated communication channels. The emergency message includes urgent information, such as that the user has collapsed, prompting a swift response.
[1382] Step 10:
[1383] The server will output the transmitted emergency message as an audio message to notify the user. This notification may be made via telephone or a dedicated application.
[1384] This series of processes ensures user safety and enables rapid response in emergencies.
[1385] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1386] This invention relates to a system that monitors elderly people living alone, facilitates communication, and enables rapid response in emergencies. In particular, this invention achieves more natural and effective communication by combining it with an emotion engine that recognizes the user's emotions.
[1387] Speech recognition and input processing
[1388] terminal
[1389] When a user speaks, the device acquires the voice through its built-in microphone. The acquired voice data is converted into text data by a speech recognition device. This text data is then sent to the server by the device's transmission device.
[1390] server
[1391] The server receives text data sent from the terminal and then analyzes the received text data using an analysis device. Based on the analysis, the server generates an appropriate response and sends that response back to the terminal via the transmission device.
[1392] terminal
[1393] The terminal converts received responses into speech data using speech synthesis technology and outputs it to the user through a speaker. This allows elderly people to interact with the system in a natural way.
[1394] Emotion recognition function
[1395] terminal
[1396] The device is equipped with an emotion engine. When the user speaks, the emotion engine recognizes emotions from their voice and facial expressions. The emotion data is sent to the server along with the text data.
[1397] server
[1398] The server receives emotion data and analyzes it along with text data. Based on the analysis results, it generates a response that takes the user's emotions into account. For example, if the user is determined to be sad, the response will be encouraging.
[1399] Examples of conversation partner functions
[1400] User: "I'm feeling sad today."
[1401] Terminal: Acquires audio and converts it into text data, such as "I feel sad today," using speech recognition. The emotion engine recognizes the sad emotion from the user's voice.
[1402] Terminal: Sends text data and sentiment data to the server.
[1403] Server: Analyzes the received data and generates encouraging responses such as, "What's wrong? Please tell me what happened."
[1404] Server: Sends the generated response to the terminal.
[1405] Terminal: Converts response data into speech and outputs to the user, "What's the matter? Please tell me what happened."
[1406] Emergency contact function
[1407] terminal
[1408] The device has built-in biosensors that continuously acquire the user's heart rate, body temperature, and movement. The acquired biometric data is monitored in real time. If an abnormality is detected, an emergency notification is generated.
[1409] server
[1410] Emergency notifications are sent to the server along with emotion data. The server receives the emergency notification and sends an emergency message to pre-registered emergency contacts.
[1411] Examples of emergency contact functions
[1412] Device: If the user falls and becomes motionless, it detects abnormal heart rate and unresponsive movement.
[1413] Terminal: Immediately sends emergency notifications and emotional data to the server.
[1414] Server: Receives emergency notification and sends a message to family contacts saying, "User has collapsed. Please respond quickly."
[1415] User (recipient): The family receives the emergency message and immediately begins taking action.
[1416] Thus, the system of the present invention provides elderly people living alone with an environment in which they can live with peace of mind by simultaneously promoting communication and enabling a rapid response in emergencies. In particular, the emotion engine enables natural dialogue that takes the user's emotions into consideration, thereby improving the quality of life.
[1417] The following describes the processing flow.
[1418] Speech recognition and input processing
[1419] Step 1:
[1420] Terminal: The terminal acquires the user's voice in real time through its built-in microphone.
[1421] Step 2:
[1422] Terminal: Transmits the acquired audio data to a speech recognition device (e.g., a speech recognition engine) and converts the audio data into text data.
[1423] Step 3:
[1424] Terminal: Sends the converted text data to the server using a transmission method. HTTP requests are often used for transmission.
[1425] Emotion recognition function
[1426] Step 4:
[1427] Terminal: Simultaneously, the emotion engine recognizes the user's emotions from their voice and facial expressions. The recognized emotion data is sent to the server along with text data.
[1428] Step 5:
[1429] Server: Analyzes received text data and sentiment data using analytical tools (such as a natural language processing engine).
[1430] Step 6:
[1431] Server: Based on the analysis results, it generates a response that takes the user's emotions into account. For example, if it is determined that the user is sad, the response will be encouraging.
[1432] Step 7:
[1433] Server: Sends the generated response to the terminal via a response transmission method. JSON format data is often used for transmission.
[1434] Step 8:
[1435] Terminal: Receives responses from the server and converts them into speech data using a text-to-speech engine.
[1436] Step 9:
[1437] Terminal: Outputs audio data to the user through the speaker.
[1438] Examples of conversation partner functions
[1439] User: "I'm feeling sad today."
[1440] Step 1: The device acquires the user's voice and converts it into text data, "I feel sad today," using speech recognition technology.
[1441] Step 2: The emotion engine recognizes sad emotions from the user's voice.
[1442] Step 3: Send text data and sentiment data to the server.
[1443] Step 4: The server analyzes the received data and generates an encouraging response such as, "What's wrong? Tell me what happened."
[1444] Step 5: The server sends the generated response to the terminal.
[1445] Step 6: The device converts the response data into speech and outputs to the user, "What's the matter? Please tell me what's wrong."
[1446] Emergency contact function
[1447] Step 1:
[1448] Device: The device continuously acquires biometric data such as heart rate, body temperature, and movement through built-in biosensors.
[1449] Step 2:
[1450] Terminal: Uses an anomaly detection method to monitor acquired biometric data in real time and detect abnormal values.
[1451] Step 3:
[1452] Terminal: If an anomaly is detected, an emergency notification will be generated. This notification will include the type of anomaly and the user's identification information.
[1453] Step 4:
[1454] Terminal: Sends the generated emergency notification to the server using the emergency notification method.
[1455] Step 5:
[1456] Server: Receives emergency notifications and sentiment data, and retrieves a pre-configured list of emergency contacts.
[1457] Step 6:
[1458] Server: Sends an emergency message to emergency contacts based on the information included in the emergency notification. Email, SMS, and phone APIs are often used for sending messages.
[1459] Step 7:
[1460] User (recipient): The family or medical institution will review the received emergency message and begin taking immediate action.
[1461] Examples of emergency contact functions
[1462] Step 1: The device detects the user's abnormal heart rate and cessation of movement.
[1463] Step 2: Send anomaly and emotion data to the server.
[1464] Step 3: The server receives an emergency notification and sends an emergency message to the emergency contact stating, "The user has collapsed. Please respond immediately."
[1465] Step 4: The family receives the emergency message and immediately begins taking action.
[1466] The above outlines the program's processing flow, including the specific actions performed at each processing step. This detailed explanation will allow for a clearer understanding of the overall system's operation.
[1467] (Example 2)
[1468] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1469] When elderly people live alone, there are risks associated with a lack of communication and delayed responses in emergencies. Furthermore, systems that cannot engage in natural conversations that take emotions into consideration can increase the user's mental burden and potentially lower their quality of life. Additionally, if biometric data monitoring and emergency notifications are not effectively implemented, prompt responses become difficult.
[1470] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: speech recognition means for acquiring voice from a terminal in real time and converting the voice into text; emotion recognition means for analyzing emotions from the terminal and generating emotion data; transmission means for the terminal to send the converted text data and emotion data to the server; analysis means for the server to analyze the received text data and emotion data and generate an appropriate response; response transmission means for the server to send the generated response to the terminal; and speech synthesis means for converting the response received by the terminal into voice and outputting it to the user. This enables natural and effective communication with the elderly, allowing for a quick response in emergencies and improving the quality of life for users.
[1471] A "terminal" is a computer device used by users for voice input and output, as well as the collection of biometric data.
[1472] "Voice recognition means" refers to technology or devices that acquire a user's voice in real time on a terminal and convert that voice into text data.
[1473] "Emotion recognition means" refers to technologies and devices that analyze a user's emotions from their voice, facial expressions, etc., and generate emotion data.
[1474] "Transmission means" refers to communication technologies and devices used to send text data and sentiment data from a terminal to a server.
[1475] A "server" is a centralized management system that receives data sent from terminals, analyzes it, and generates appropriate responses.
[1476] "Analysis means" refers to technologies and devices that analyze text data and sentiment data on a server and generate appropriate responses for the user.
[1477] "Response transmission means" refers to communication technology or equipment used to send a response generated by a server to a terminal.
[1478] "Speech synthesis means" refers to a technology or device that converts response data received from a server into speech data at a terminal and outputs it to the user as speech.
[1479] "Methods for acquiring biometric data" refer to technologies and devices that continuously acquire biometric data such as a user's heart rate, body temperature, and movement.
[1480] An "anomaly detection method" refers to a technology or device that monitors acquired biological data in real time and detects anomalies.
[1481] An "emergency notification system" is a technology or device that sends an emergency notification to a server when an anomaly is detected.
[1482] "Communication methods" refer to technologies and devices that, after a server receives an emergency notification, send an emergency message to pre-configured emergency contacts.
[1483] This invention relates to a system that monitors elderly people living alone, facilitates communication, and enables rapid response in emergencies. In particular, this invention achieves more natural and effective communication by combining it with an emotion engine that recognizes the user's emotions.
[1484] Hardware and software to be used
[1485] terminal
[1486] The device includes the following hardware and software:
[1487] High-sensitivity microphone (voice input)
[1488] Speaker (audio output)
[1489] Biosensors (acquisition of heart rate, body temperature, and movement data)
[1490] Speech recognition engine (e.g., Google Speech-to-Text API)
[1491] Emotion recognition engine (e.g., Microsoft Azure Emotion API)
[1492] Speech synthesis engine (e.g., Amazon Polly)
[1493] server
[1494] The server includes the following software and features:
[1495] Analysis engine (Python library using NLP technology)
[1496] Response generation engine
[1497] Emergency notification system
[1498] Data receiving and transmitting module
[1499] System operation
[1500] terminal
[1501] When a user speaks, the device acquires the voice through its built-in microphone. The acquired voice data is converted into text data using a speech recognition engine. Simultaneously, an emotion recognition engine analyzes emotions from voice and facial expressions, generating emotion data. This data is transmitted to the server via the device's transmission mechanism.
[1502] server
[1503] The server receives text data and sentiment data sent from the terminal. The received data is analyzed using an analysis engine. Based on the analysis results, the server generates an appropriate response. The generated response is sent to the terminal via a response transmission means.
[1504] terminal
[1505] The terminal converts responses received from the server into speech data using a speech synthesis engine and outputs it to the user through a speaker. This allows elderly people to interact with the system in a natural way.
[1506] Specific example
[1507] Conversation partner function
[1508] User: "I'm feeling sad today."
[1509] Device: It uses the built-in microphone to acquire voice data and converts it into text data, such as "I feel sad today," using a speech recognition engine. The emotion engine then analyzes the user's voice to determine that they are feeling "sad."
[1510] Terminal: Sends the converted text data and sentiment data to the server.
[1511] Server: Analyzes the received data and generates encouraging responses such as, "What's wrong? Please tell me what happened."
[1512] Server: Sends the generated response data to the terminal.
[1513] Terminal: Converts response data into speech and outputs to the user, "What's the matter? Please tell me what happened."
[1514] Emergency contact function
[1515] Device: If the user falls, the built-in biosensors detect abnormal heart rate or unresponsive movement.
[1516] Terminal: Immediately sends emergency notifications and emotional data to the server.
[1517] Server: Upon receiving an emergency notification, it sends the message "User has collapsed. Please respond immediately" to the pre-configured emergency contact.
[1518] Recipient user: The family receives the emergency message and immediately begins taking action.
[1519] Example of a prompt
[1520] Example of prompt text when inputting a "monitoring system for elderly people living alone" into the AI model:
[1521] "The device acquires the user's voice and converts it into text data using a speech recognition engine. Additionally, an emotion engine analyzes the user's emotions and sends this data to a server. The server analyzes the received data, generates an appropriate response, and sends it to the device, which then responds verbally. Furthermore, in emergencies, a biosensor detects an anomaly and sends an emergency notification through the server. Please explain this system."
[1522] Thus, the system of the present invention promotes natural dialogue with the elderly and enables a rapid response in emergencies, thereby improving the safety and quality of life of the elderly.
[1523] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1524] Step 1:
[1525] The device acquires the user's voice. Using the built-in high-sensitivity microphone, it collects voice data when the user says, "I'm feeling sad today."
[1526] Input: User's voice
[1527] Output: Audio data
[1528] Specific operation: The device's operating system converts the audio signal from the microphone into digital data and temporarily stores it in memory.
[1529] Step 2:
[1530] The device uses a speech recognition engine to convert audio data into text data. For example, the speech recognition engine uses the Google Speech-to-Text API.
[1531] Input: Audio data
[1532] Output: Text data
[1533] Specific operation: The speech recognition engine analyzes the audio data, identifies phonemes and morphemes, and generates the text "I feel sad today." The generated text is stored in a separate buffer in memory.
[1534] Step 3:
[1535] The device uses an emotion recognition engine to generate emotion data from voice and text data. For example, it can utilize the Microsoft Azure Emotion API.
[1536] Input: Audio data, text data
[1537] Output: Sentiment data
[1538] Specific operation: The emotion recognition engine analyzes the emotion of "sadness" from the tone of voice and associated text, and generates this emotion data. The generated emotion data is stored in memory along with the text data.
[1539] Step 4:
[1540] The terminal sends the converted text data and sentiment data to the server. This is done using a network transmission module.
[1541] Input: Text data, sentiment data
[1542] Output: Transmitted data
[1543] Specific operation: The transmission module packets text and sentiment data and sends them to the server via the Wi-Fi module over the internet. The data is encrypted to ensure security.
[1544] Step 5:
[1545] The server receives text data and sentiment data sent from the terminal.
[1546] Input: Data to send
[1547] Output: Received data
[1548] Specific operation: The server's receiving module receives packets, decrypts the encrypted data, and obtains text data and sentiment data. This data is stored in a database for analysis.
[1549] Step 6:
[1550] The server analyzes text and sentiment data and generates an appropriate response using an analysis engine.
[1551] Input: Text data, sentiment data
[1552] Output: Response data
[1553] Specific operation: The analysis engine uses NLP technology to analyze text data and sentiment data, and generates the most appropriate response for the user's state, such as "What's wrong? Please tell me what happened."
[1554] Step 7:
[1555] The server sends the generated response data to the terminal. A response transmission method is used.
[1556] Input: Response data
[1557] Output: Transmitted data
[1558] Specific operation: The response data is packetized and sent to the terminal via the network transmission module over the internet. The data is transmitted encrypted.
[1559] Step 8:
[1560] The terminal converts the received response data into audio data and outputs it to the user through the speaker. A speech synthesis engine is used.
[1561] Input: Response data
[1562] Output: Audio data
[1563] Specific operation: The speech synthesis engine analyzes the response data and generates the voice message, "What's the matter? Please tell me what's wrong." The generated voice message is output to the user through the speaker.
[1564] Step 9:
[1565] The device monitors biometric data, using sensors such as heart rate sensors, body temperature sensors, and motion sensors.
[1566] Input: Biometric data
[1567] Output: Acquired data
[1568] Specific operation: The sensor continuously measures the user's heart rate, body temperature, and movement, and records this data in memory. The data is updated in real time, waiting for anomaly detection.
[1569] Step 10:
[1570] If the device detects an anomaly, it sends an emergency notification to the server. For example, it might detect a sudden increase in heart rate or a cessation of movement.
[1571] Input: Acquired data
[1572] Output: Emergency notification data
[1573] Specific operation: When the anomaly detection algorithm detects a sudden increase in heart rate, abnormal fluctuations in body temperature, or cessation of movement, it generates emergency notification data related to these events and sends it to the server.
[1574] Step 11:
[1575] The server receives an emergency notification and sends an emergency message to pre-configured emergency contacts.
[1576] Input: Emergency notification data
[1577] Output: Emergency message
[1578] Specific operation: The server analyzes the received emergency notification data and sends a message stating "User has collapsed. Please respond immediately" to a pre-configured list of emergency contacts (e.g., family or medical institutions). The message is sent via SMS or a dedicated application.
[1579] (Application Example 2)
[1580] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1581] In recent years, the number of elderly people living alone has increased, making their safety and security a major concern. There is also a growing need for systems that alleviate their feelings of loneliness and provide rapid responses in emergencies. However, current systems lack sufficient emotion recognition and emergency notification capabilities, and cannot fully guarantee the psychological and physical safety of the elderly. Therefore, there is a need for natural communication methods that include emotion recognition, as well as rapid emergency response through real-time monitoring of biometric data.
[1582] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a voice recognition means, a transmission means, an analysis means, a response transmission means, a voice synthesis means, and a means for recognizing emotions from the user's voice and facial expressions, equipped with an emotion recognition engine. This enables natural dialogue that takes into account the emotions of the elderly, and also enables a rapid response in emergencies. It also has an emergency notification function and can provide a system that continuously monitors heart rate, body temperature, and movement data, and immediately makes an emergency contact when an abnormality is detected.
[1583] A "terminal" is a device used by a user that has functions such as voice input, speech synthesis, emotion recognition, and acquisition of biometric data.
[1584] "Voice recognition means" refers to a function that converts the voice acquired by the terminal into text data.
[1585] "Transmission method" refers to the function by which the terminal sends the converted text data to the server.
[1586] "Analysis means" refers to a function that analyzes text data and sentiment data received by the server and generates an appropriate response.
[1587] "Response transmission means" refers to the function that sends the response generated by the server to the terminal.
[1588] A "speech synthesis means" is a function that converts the response received by the terminal into speech and outputs it to the user.
[1589] An "emotion recognition engine" is a function that recognizes emotions from the user's voice and facial expressions.
[1590] "Means of acquiring biometric data" refers to a function in which the terminal continuously acquires the user's heart rate, body temperature, and movement.
[1591] An "anomaly detection method" is a function that monitors acquired biometric data in real time and detects sudden increases in heart rate, abnormal fluctuations in body temperature, and cessation of movement.
[1592] An "emergency notification mechanism" is a function that sends an emergency notification to the server when an anomaly is detected.
[1593] "Communication method" refers to the function where the server receives an emergency notification and sends an emergency message to a pre-configured emergency contact.
[1594] This invention is a system that monitors elderly people living alone, facilitates communication, and enables rapid response in emergencies. The system mainly consists of terminals and a server.
[1595] Device functions
[1596] The terminal is equipped with a speech recognition system, a speech synthesis system, an emotion recognition engine, and a biometric data acquisition system. When a user speaks, the voice is acquired through the terminal's microphone and converted into text data by the speech recognition system. This text data is then sent to the server via a transmission system.
[1597] The device also features an emotion recognition engine that recognizes emotions from the user's voice and facial expressions. This emotion data is also sent to the server along with the text data.
[1598] Furthermore, the device is equipped with biometric data acquisition capabilities, continuously monitoring the user's heart rate, body temperature, and movement. If an abnormality is detected, a notification is sent to the server via an emergency notification system.
[1599] Server Functions
[1600] The server has an analysis mechanism to analyze the received text data and sentiment data, and generates an appropriate response. The generated response is sent back to the terminal via a response transmission mechanism. The terminal converts this response into speech using a speech synthesis mechanism and outputs it to the user through a speaker.
[1601] The server has a communication system that, upon receiving an emergency notification, sends an emergency message to pre-configured emergency contacts.
[1602] Hardware and software details
[1603] The system uses the "Speech Recognition API" for speech recognition, the "Network Communication Module" for text data transmission, and the "Emotion Recognition API" for emotion recognition. Additionally, it uses the "Speech Synthesis API" for speech synthesis and the "Biometric Sensor Module" for remote monitoring. The "Notification API" is used for emergency notifications and communication.
[1604] Specific example
[1605] When a user says, "I'm not feeling well today," the device converts the audio data into text data and analyzes it using an emotion recognition engine. The resulting data is sent to a server, which considers the user's emotions and generates an encouraging response such as, "What's wrong? Please tell me what's the matter." The generated response is sent to the device and output to the user as audio.
[1606] Furthermore, if a user collapses, the biosensors built into the device detect abnormal heart rate or unresponsive movement and immediately send an emergency notification to the server. The server receives the emergency notification and sends a message to emergency contacts stating, "A user has collapsed. Please respond quickly."
[1607] Example of a prompt
[1608] "Provide an example where, if a user says 'I'm not feeling well today,' the audio is converted to text, analyzed by an emotion engine, and an appropriate comforting response is generated."
[1609] As described above, this system supports the safe and secure lives of the elderly and can provide prompt assistance when needed. Furthermore, its emotion recognition function enables more natural and effective communication.
[1610] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1611] Step 1:
[1612] The user speaks to the device.
[1613] Input: User's voice data.
[1614] Specific action: The device's microphone picks up the user's voice.
[1615] Step 2:
[1616] The device converts the audio into text data.
[1617] Input: Acquired audio data.
[1618] Specific operation: Use a speech recognition API to convert speech data into text data.
[1619] Output: Text data.
[1620] Step 3:
[1621] The device performs emotion recognition.
[1622] Input: User's voice data and text data.
[1623] Specific operation: Use an emotion recognition API to analyze the user's emotions from their voice and facial expressions.
[1624] Output: Sentiment data.
[1625] Step 4:
[1626] The device sends text data and sentiment data to the server.
[1627] Input: Text data and sentiment data.
[1628] Specific operation: Use the network communication module to send data to the server.
[1629] Output: Confirmation of data transmission to the server.
[1630] Step 5:
[1631] The server analyzes the text data and sentiment data it receives.
[1632] Input: Text data and sentiment data.
[1633] Specific operation: Use analytical tools to perform analysis in order to generate an appropriate response.
[1634] Output: Appropriate response data.
[1635] Step 6:
[1636] The server sends the appropriate response data to the terminal.
[1637] Input: Generated response data.
[1638] Specific operation: Use the response transmission means to send response data to the terminal.
[1639] Output: Confirmation of data transmission to the terminal.
[1640] Step 7:
[1641] The terminal converts the received response data into speech.
[1642] Input: Appropriate response data.
[1643] Specific operation: Use a speech synthesis API to convert text data into speech data.
[1644] Output: Audio data.
[1645] Step 8:
[1646] The device outputs audio data to the user.
[1647] Input: Audio data.
[1648] Specific action: Play audio through the device's speaker.
[1649] Output: Audio output to the user.
[1650] Step 9:
[1651] The device acquires biometric data.
[1652] Input: Biometric data such as the user's heart rate, body temperature, and movement.
[1653] Specific operation: Use a biosensor module to continuously collect data.
[1654] Output: Biometric data.
[1655] Step 10:
[1656] The device monitors biometric data acquired in real time.
[1657] Input: Biometric data.
[1658] Specific operation: Use anomaly detection methods to monitor data for abnormalities.
[1659] Output: Abnormal status data.
[1660] Step 11:
[1661] If an anomaly is detected, the device will send an emergency notification to the server.
[1662] Input: Abnormal situation data.
[1663] Specific action: Use the emergency notification method to send an anomaly notification to the server.
[1664] Output: Sending an emergency notification to the server.
[1665] Step 12:
[1666] The server receives an emergency notification and sends a message to the emergency contact.
[1667] Input: Anomaly notification data.
[1668] Specific action: Use the communication method to send a message to a pre-configured emergency contact.
[1669] Output: Confirmation of emergency message transmission.
[1670] The above describes the processing steps of the system based on this invention.
[1671] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1672] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1673] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[1674] [Fourth Embodiment]
[1675] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[1676] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1677] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1678] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[1679] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1680] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1681] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1682] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[1683] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1684] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1685] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1686] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1687] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1688] This invention relates to a system for monitoring elderly people living alone, facilitating communication with them using voice and biometric data, and enabling a rapid response in emergencies. Embodiments of this invention are described in detail below.
[1689] Speech recognition and input processing
[1690] terminal
[1691] When a user speaks into the device, the device acquires the voice through its built-in microphone. The acquired voice data is converted into text data by a speech recognition system. This text data is then sent to the server through the device's transmission system.
[1692] server
[1693] The server receives text data sent from the terminal. The received data is analyzed by an analysis device, and an appropriate response is generated based on the user's intent. The generated response is then sent back to the terminal via the server's response transmission device.
[1694] terminal
[1695] The terminal receives responses from the server, converts them into speech data using speech synthesis technology, and outputs them to the user through a speaker. This allows elderly people to interact with the system in a natural way.
[1696] Examples of conversation partner functions
[1697] User: "What kind of day is it today?"
[1698] Terminal: Acquires audio and converts it into text data, such as "What kind of day is it today?", using speech recognition technology.
[1699] Terminal: Sends text data to the server.
[1700] Server: Analyzes the received text data and generates a response such as "Today is not a special day."
[1701] Server: Sends the generated response to the terminal.
[1702] Terminal: Converts the response data into speech and outputs "Today is not a special day" to the user.
[1703] Emergency contact function
[1704] terminal
[1705] The device has built-in biosensors that continuously acquire the user's heart rate, body temperature, and movement. This biometric data is constantly monitored, and if an abnormality is detected, an emergency notification is generated.
[1706] server
[1707] Emergency notifications are sent from the terminal to the server. The server analyzes the emergency notification and sends an emergency message to pre-registered emergency contacts. The emergency message includes the user ID, terminal ID, and abnormal data.
[1708] Examples of emergency contact functions
[1709] Device: If the user falls and becomes motionless, it detects abnormal heart rate and unresponsive movement.
[1710] Terminal: Immediately send an emergency notification to the server.
[1711] Server: Receives emergency notification and sends SMS to pre-configured family contacts. "User has collapsed. Please respond immediately."
[1712] User (recipient): The family receives the emergency message and immediately begins taking action.
[1713] In this way, the system of the present invention can ensure the safety and communication of elderly people living alone.
[1714] The following describes the processing flow.
[1715] Speech recognition and input processing
[1716] Step 1:
[1717] Terminal: The terminal acquires the user's voice in real time through its built-in microphone.
[1718] Step 2:
[1719] Terminal: Transmits the acquired audio data to a speech recognition device (e.g., a speech recognition engine) and converts the audio data into text data.
[1720] Step 3:
[1721] Terminal: Sends the converted text data to the server using a transmission method. HTTP requests are often used for transmission.
[1722] Conversation partner function
[1723] Step 4:
[1724] Server: Analyzes the received text data using parsing tools (such as a natural language processing engine).
[1725] Step 5:
[1726] Server: Generates an appropriate response based on the analysis results. The response may include fixed statements or information obtained from external APIs.
[1727] Step 6:
[1728] Server: Sends the generated response to the terminal via a response transmission method. JSON format data is often used for transmission.
[1729] Step 7:
[1730] Terminal: Converts the received response text data into speech data using a text-to-speech engine.
[1731] Step 8:
[1732] Terminal: Outputs audio data to the user through the speaker.
[1733] Emergency contact function
[1734] Step 1:
[1735] Device: The device continuously acquires biometric data such as heart rate, body temperature, and movement through built-in biosensors.
[1736] Step 2:
[1737] Terminal: Uses an anomaly detection method to monitor acquired biometric data in real time and detect abnormal values.
[1738] Step 3:
[1739] Terminal: If an anomaly is detected, an emergency notification will be generated. This notification will include the type of anomaly and the user's identification information.
[1740] Step 4:
[1741] Terminal: Sends the generated emergency notification to the server using the emergency notification method.
[1742] Step 5:
[1743] Server: Receives emergency notifications and retrieves a pre-configured list of emergency contacts.
[1744] Step 6:
[1745] Server: Sends an emergency message to emergency contacts based on the information included in the emergency notification. Email, SMS, and phone APIs are often used for sending messages.
[1746] Step 7:
[1747] User (recipient): The family or medical institution will review the received emergency message and begin taking immediate action.
[1748] The above is the program's processing flow, including the specific actions performed at each processing step. Providing such a detailed explanation allows for a clearer understanding of how the entire system works.
[1749] (Example 1)
[1750] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1751] In modern society, elderly people living alone face challenges such as a lack of daily communication and the inability to respond quickly in emergencies. This has led to feelings of loneliness among the elderly and delays in emergency response. The present invention aims to provide a system that enables elderly people to engage in daily conversations in a natural manner and further enables a rapid response in emergencies using biometric data.
[1752] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1753] In this invention, the server includes: speech recognition means for acquiring voice from a terminal in real time and converting the voice into text; transmission means for the terminal to send the converted text data to the server; analysis means for the server to analyze the received text data and generate an appropriate response using natural language processing means; response transmission means for the server to send the generated response to the terminal; speech synthesis means for the terminal to convert the received response into voice and output it to the user; anomaly detection means for monitoring biometric data in real time and detecting anomalies; emergency notification means for sending an emergency notification to the server when an anomaly is detected; and communication means for the server to receive the emergency notification and automatically send an emergency message to a pre-configured emergency contact. This enables elderly people to communicate naturally using voice on a daily basis, and also enables a rapid response in emergencies using biometric data.
[1754] A "terminal" is a device that functions as an interface with the user, collecting, converting, transmitting, and outputting responses for voice data.
[1755] "Speech recognition means" refers to a technology that converts speech into text data, is built into a terminal, and has the function of recognizing the user's voice in real time.
[1756] "Transmission means" refers to the communication means used by a terminal to send converted text data to a server, and includes Wi-Fi modules and internet connection means.
[1757] A "server" is a central control unit that receives data sent from terminals and performs processing, analysis, and response generation.
[1758] "Natural language processing means" refers to technology that analyzes received text data, understands the user's intent, and generates an appropriate response.
[1759] "Analysis means" refers to the technology used by a server to process received text data and generate an appropriate response using natural language processing means.
[1760] "Response transmission means" refers to communication technology used to send responses generated by a server to a terminal.
[1761] "Speech synthesis means" refers to a technology that converts text data received from a server into speech data and outputs a speech response to the user.
[1762] "Methods for acquiring biometric data" refers to sensor technology that continuously acquires data on the user's heart rate, body temperature, and movement.
[1763] "Anomaly detection means" refers to technology that monitors biological data in real time and detects anomalies.
[1764] An "emergency notification system" refers to a technology that immediately sends an emergency notification to a server when an anomaly is detected.
[1765] "Communication method" refers to the technology in which a server receives an emergency notification and automatically sends an emergency message to pre-configured emergency contacts.
[1766] A "generative AI model" refers to artificial intelligence technology used to generate natural dialogue and appropriate responses in abnormal situations.
[1767] A "prompt statement" refers to an input statement used to give instructions to a generative AI model and is used to generate a specific response.
[1768] This invention relates to a system for monitoring elderly people living alone. It facilitates communication with the elderly using voice and biometric data, and enables a rapid response in emergencies. The embodiments of this invention are described in detail below.
[1769] Speech recognition and input processing
[1770] terminal
[1771] When a user speaks into the device, the device acquires the voice through its built-in microphone. The acquired voice data is converted into text data using speech recognition tools (e.g., Google Cloud Speech-to-Text API). This text data is then sent to the server via the device's transmission tools (Wi-Fi module).
[1772] server
[1773] The server uses a Python framework (e.g., Flask or Django) to receive text data sent from the terminal. The received data is parsed by a natural language processing tool (e.g., spaCy) to generate an appropriate response based on the user's intent. The generated response is then sent back to the terminal via the server's response sending mechanism.
[1774] terminal
[1775] The device uses the Google Cloud Text-to-Speech API to convert responses received from the server into audio data, which is then output to the user through the speaker. This allows elderly users to interact with the system in a natural way.
[1776] Specific example
[1777] User: "What kind of day is it today?"
[1778] Terminal: Acquires audio and converts it into text data, such as "What kind of day is it today?", using speech recognition technology.
[1779] Terminal: Sends text data to the server.
[1780] Server: Analyzes the received text data and generates a response such as "Today is not a special day."
[1781] Server: Sends the generated response to the terminal.
[1782] Terminal: Converts the response data into speech and outputs "Today is not a special day" to the user.
[1783] Emergency contact function
[1784] terminal
[1785] The device has built-in biosensors (heart rate sensor, body temperature sensor, motion sensor, etc.) that continuously acquire the user's heart rate, body temperature, and movement. This biometric data is constantly monitored by a microcontroller (e.g., ESP32), and if an anomaly is detected, an emergency notification is generated.
[1786] server
[1787] Emergency notifications are sent from the device to the server. The server analyzes the emergency notification and sends an emergency message to pre-registered emergency contacts using the Twilio API. The emergency message includes the user ID, device ID, and anomaly data.
[1788] Specific example
[1789] If a user collapses and becomes motionless, the system detects abnormal heart rate and unresponsive movement.
[1790] Terminal: Immediately send an emergency notification to the server.
[1791] Server: Receives emergency notification and sends an SMS to pre-configured family contacts using the Twilio API. "User has collapsed. Please respond quickly."
[1792] Examples of prompt statements
[1793] The following are specific examples of prompt statements used in generative AI models.
[1794] In the case of an elderly monitoring system:
[1795] Create text to generate natural conversations with users.
[1796] For example, generate a response to the user's question, "What kind of day is it today?"
[1797] By using these prompts, it becomes possible to provide a system that enables natural interaction with users and allows for rapid response in emergencies.
[1798] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1799] Program processing for speech recognition and input processing
[1800] Step 1:
[1801] terminal
[1802] The user speaks into the device, and the voice is captured through the built-in microphone. This voice data is the input. The captured voice data is sent in real time to a speech recognition system (e.g., Google Cloud Speech-to-Text API) and converted into text data. The data processing here is the conversion from voice to text, and the output is text data.
[1803] Specific actions
[1804] User: "What kind of day is it today?"
[1805] Device: The built-in microphone captures audio and sends it to the speech recognition API.
[1806] Step 2:
[1807] terminal
[1808] The system acquires text data converted by speech recognition and sends it to a server. The input here is the converted text data, and the output is the text data sent to the server. Transmission takes place via the internet through a Wi-Fi module.
[1809] Specific actions
[1810] Terminal: The device sends the text data "What kind of day is it today?" obtained from the voice recognition device to the server via the Wi-Fi module.
[1811] Step 3:
[1812] server
[1813] The server receives text data sent from the terminal using a Python framework (e.g., Flask or Django). The input here is text data, which is parsed using natural language processing tools (e.g., spaCy). The output is the appropriate response resulting from the parsing.
[1814] Specific actions
[1815] Server: Analyzes the text data "What kind of day is it today?" and uses a natural language processing library to understand the user's intent.
[1816] Step 4:
[1817] server
[1818] Based on the analysis results, an appropriate response is generated using a generative AI model (e.g., GPT-3). The input is the analysis results, and the output is the generated response text.
[1819] Specific actions
[1820] Server: Generates responses such as "Today is not a special day" using an AI model based on the analysis results from a natural language processing library.
[1821] Step 5:
[1822] server
[1823] The generated response is sent to the terminal via the response transmission means. The input is the generated response text, and the output is the text data sent to the terminal.
[1824] Specific actions
[1825] Server: Sends the response text "Today is not a special day" to the terminal.
[1826] Step 6:
[1827] terminal
[1828] The system converts the response received from the server into speech data using a text-to-speech synthesis tool (e.g., Google Cloud Text-to-Speech API). The input is the response text, and the output is speech data. This is output to the user through the built-in speaker.
[1829] Specific actions
[1830] Terminal: Receives the response text "Today is not a special day," converts it into speech data using a speech synthesis method, and outputs it through the speaker.
[1831] Processing of the emergency contact function program
[1832] Step 1:
[1833] terminal
[1834] The device's built-in biosensors (heart rate sensor, body temperature sensor, motion sensor, etc.) continuously acquire the user's biometric data. The input is data from the biosensors, and the output is the continuously acquired biometric data. This data is sent to a microcontroller (e.g., ESP32) for monitoring.
[1835] Specific actions
[1836] Device: The user's heart rate sensor measures a heart rate of 60 BPM.
[1837] Step 2:
[1838] terminal
[1839] The system monitors acquired biometric data and detects anomalies if abnormal values exceed a threshold. The input is continuously acquired biometric data, which is analyzed by an anomaly detection mechanism. The output indicates whether or not an anomaly exists.
[1840] Specific actions
[1841] Terminal: If a heart rate of 30 BPM (below the abnormal capture threshold) is detected, it will be flagged as abnormal.
[1842] Step 3:
[1843] terminal
[1844] When an anomaly is detected, an emergency notification is immediately generated. The input is the result of the anomaly detection, and the output is the emergency notification data. The generated emergency notification includes the user ID and anomaly data.
[1845] Specific actions
[1846] The device detects an abnormal heart rate and generates data that reads "Emergency notification: User ID 123, heart rate 30 BPM".
[1847] Step 4:
[1848] terminal
[1849] The generated emergency notification is sent to the server. The input is the emergency notification data, and the output is the emergency notification sent to the server. Transmission is performed via the Wi-Fi module.
[1850] Specific actions
[1851] Device: Emergency notifications are sent to the server via Wi-Fi.
[1852] Step 5:
[1853] server
[1854] This system receives emergency notifications and automatically sends emergency messages to pre-registered emergency contacts via the Twilio API. The input is the received emergency notification data, and the output is the sent emergency message. The emergency message includes user ID, device ID, and anomaly data.
[1855] Specific actions
[1856] Server: Upon receiving an emergency notification, it sends an SMS message using the Twilio API stating, "A user has collapsed. Please respond quickly."
[1857] As described above, the system of the present invention enables natural communication with the elderly using voice and biometric data, as well as a rapid response in emergencies.
[1858] (Application Example 1)
[1859] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1860] The problem that this invention aims to solve is to provide an environment in which elderly people living alone can live safely and securely. In particular, it aims to enable rapid response in emergencies by linking voice recognition and biometric data, and to monitor changes in the elderly person's physical condition and abnormal situations in real time. Another challenge is to improve communication with the elderly person through a user-friendly interface.
[1861] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1862] In this invention, the server includes an analysis means for analyzing text data and generating an appropriate response, a communication means for receiving emergency notifications and sending emergency messages to pre-configured emergency contacts, and a means for generating responses using an AI model. This enables elderly people living alone to communicate naturally by voice and to respond quickly in emergencies.
[1863] A "terminal" is a device used by a user that has the function of acquiring voice data and sending it to a server, as well as the function of acquiring biometric data.
[1864] "Voice recognition means" refers to technology that converts speech acquired by a terminal into text data.
[1865] "Transmission method" refers to the function by which the terminal sends the converted text data to the server.
[1866] "Analysis means" refers to a technology that analyzes text data received by a server and generates an appropriate response based on the user's intent.
[1867] "Response transmission means" refers to the function by which the server sends the generated response to the terminal.
[1868] "Speech synthesis means" refers to a technology that converts responses received by a terminal from a server into speech data and outputs it to the user as speech.
[1869] "Means of acquiring biometric data" refers to a function in which a device continuously acquires biometric data such as the user's heart rate, body temperature, and movement.
[1870] An "anomaly detection method" is a technology that monitors acquired biometric data in real time and generates an emergency notification when an anomaly is detected.
[1871] An "emergency notification mechanism" is a function that allows a terminal to send an emergency notification to a server when it detects an anomaly.
[1872] "Communication method" refers to a technology where a server receives an emergency notification and sends an emergency message to pre-configured emergency contacts.
[1873] An "AI model" is an artificial intelligence algorithm used by a server to analyze user text data and generate appropriate responses.
[1874] A "prompt message" is input data that is dynamically generated based on user inquiries used by the generative AI model when performing analysis.
[1875] This invention is a system for monitoring elderly people living alone, utilizing voice recognition and biometric data to facilitate communication with them and enable rapid response in emergencies. Details of the invention are shown below.
[1876] System Configuration
[1877] The system primarily consists of "terminals" and "servers." Their respective functions and specific examples are explained below.
[1878] (Device functions)
[1879] 1. Speech recognition means
[1880] The device uses its built-in microphone to capture audio in real time and converts the captured audio into text using speech recognition software (e.g., the speech_recognition library). For example, if a user says "What kind of day is it today?", the audio is converted into the text data "What kind of day is it today?".
[1881] 2. Transmission method
[1882] The terminal sends the converted text data to the server. Therefore, a wireless communication module (e.g., a Wi-Fi module) is required.
[1883] 3. Speech synthesis means
[1884] The terminal converts the response received from the server into audio data using speech synthesis software (e.g., the playsound library) and outputs it to the user through the built-in speaker. For example, if the terminal receives the response "Today is not a special day" from the server, it will convey that to the user as audio.
[1885] 4. Means of acquiring biometric data
[1886] The device uses built-in biosensors (e.g., heart rate sensor, temperature sensor, motion sensor) to continuously acquire data on the user's heart rate, body temperature, and movement.
[1887] 5. Anomaly detection means
[1888] The device monitors acquired biometric data in real time and detects abnormalities. For example, a sudden increase in heart rate or a lack of movement may be considered an abnormality.
[1889] 6. Emergency notification means
[1890] If an anomaly is detected, the terminal immediately sends an emergency notification to the server. A wireless communication module is required for this.
[1891] (Server functions)
[1892] 1. Analysis method
[1893] The server analyzes the received text data and generates an appropriate response. A generative AI model (e.g., the GPT model) is used in this process. For example, in response to the text data "What kind of day is it today?", it generates the response "Today is not a special day."
[1894] 2. Means of contact
[1895] When the server receives an emergency notification, it sends an emergency message to pre-configured emergency contacts. It also includes a mechanism to output the generated emergency message as audio and notify the user.
[1896] 3. Response generation means using an AI model
[1897] The server's analysis method uses a generative AI model to generate appropriate responses based on the user's text data. For example, by inputting a prompt into the generative AI model, a response is dynamically generated.
[1898] Specific example
[1899] The user is wearing smart glasses at home.
[1900] User: "What kind of day is it today?"
[1901] Terminal: Acquires audio, converts it into text data such as "What kind of day is it today?" using speech recognition, and sends it to the server.
[1902] Server: Analyzes the received text data, generates a response saying "Today is not a special day," and sends it to the terminal.
[1903] Terminal: The response from the server is converted into speech data using a speech synthesis method, and the message "Today is not a special day" is output to the user.
[1904] If a user suddenly collapses, the device's biosensors detect a sudden increase in heart rate and cessation of movement, and send an emergency notification to the server.
[1905] Server: Upon receiving an emergency notification, it sends an emergency message to a pre-configured emergency contact (e.g., family member) stating, "The user has collapsed. Please respond immediately."
[1906] Examples of prompts for generative AI models
[1907] User: "What kind of day is it today?"
[1908] The application responded, "Today is not a special day."
[1909] If the user's heart rate suddenly increases afterward, how will an emergency notification be sent?
[1910] Through these components and processes, the present invention provides a system that guarantees the safety and security of the elderly.
[1911] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1912] Step 1:
[1913] The user speaks into the device. The device acquires the voice through its built-in microphone. The acquired voice data is sent as input to a speech recognition system, where the voice is converted into text data.
[1914] Step 2:
[1915] The terminal sends the converted text data to the server via a transmission means. Here, as part of the data processing, the text data is transmitted to the server in an appropriate format.
[1916] Step 3:
[1917] The server analyzes the received text data using an analysis tool. The analysis tool uses a generative AI model to generate an appropriate response based on the input text data. Specifically, the data calculation involves generating a prompt sentence, inputting it into the AI model, and generating a text response as a result.
[1918] Step 4:
[1919] The server sends the generated response back to the terminal via the response transmission means. The data transmitted is a text-based response.
[1920] Step 5:
[1921] The terminal converts the response received from the server into audio data using a speech synthesis system and outputs it to the user through its built-in speaker. The output audio is based on the response content generated by the server.
[1922] Step 6:
[1923] The device continuously acquires data on the user's heart rate, body temperature, and movement using built-in biosensors. This biometric data is sent as input to an anomaly detection system. Here, the biometric data is analyzed in real time as part of the data processing.
[1924] Step 7:
[1925] The anomaly detection system monitors acquired biometric data in real time and detects whether there are any abnormalities. For example, a sudden increase in heart rate or a lack of movement will be detected as an anomaly. This detection is the result of data analysis.
[1926] Step 8:
[1927] If an anomaly is detected, the terminal immediately sends an emergency notification to the server using the emergency notification mechanism. This notification includes the user ID, terminal ID, and detected anomaly data.
[1928] Step 9:
[1929] The server receives an emergency notification and sends an emergency message to pre-configured emergency contacts via the designated communication channels. The emergency message includes urgent information, such as that the user has collapsed, prompting a swift response.
[1930] Step 10:
[1931] The server will output the transmitted emergency message as an audio message to notify the user. This notification may be made via telephone or a dedicated application.
[1932] This series of processes ensures user safety and enables rapid response in emergencies.
[1933] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1934] This invention relates to a system that monitors elderly people living alone, facilitates communication, and enables rapid response in emergencies. In particular, this invention achieves more natural and effective communication by combining it with an emotion engine that recognizes the user's emotions.
[1935] Speech recognition and input processing
[1936] terminal
[1937] When a user speaks, the device acquires the voice through its built-in microphone. The acquired voice data is converted into text data by a speech recognition device. This text data is then sent to the server by the device's transmission device.
[1938] server
[1939] The server receives text data sent from the terminal and then analyzes the received text data using an analysis device. Based on the analysis, the server generates an appropriate response and sends that response back to the terminal via the transmission device.
[1940] terminal
[1941] The terminal converts received responses into speech data using speech synthesis technology and outputs it to the user through a speaker. This allows elderly people to interact with the system in a natural way.
[1942] Emotion recognition function
[1943] terminal
[1944] The device is equipped with an emotion engine. When the user speaks, the emotion engine recognizes emotions from their voice and facial expressions. The emotion data is sent to the server along with the text data.
[1945] server
[1946] The server receives emotion data and analyzes it along with text data. Based on the analysis results, it generates a response that takes the user's emotions into account. For example, if the user is determined to be sad, the response will be encouraging.
[1947] Examples of conversation partner functions
[1948] User: "I'm feeling sad today."
[1949] Terminal: Acquires audio and converts it into text data, such as "I feel sad today," using speech recognition. The emotion engine recognizes the sad emotion from the user's voice.
[1950] Terminal: Sends text data and sentiment data to the server.
[1951] Server: Analyzes the received data and generates encouraging responses such as, "What's wrong? Please tell me what happened."
[1952] Server: Sends the generated response to the terminal.
[1953] Terminal: Converts response data into speech and outputs to the user, "What's the matter? Please tell me what happened."
[1954] Emergency contact function
[1955] terminal
[1956] The device has built-in biosensors that continuously acquire the user's heart rate, body temperature, and movement. The acquired biometric data is monitored in real time. If an abnormality is detected, an emergency notification is generated.
[1957] server
[1958] Emergency notifications are sent to the server along with emotion data. The server receives the emergency notification and sends an emergency message to pre-registered emergency contacts.
[1959] Examples of emergency contact functions
[1960] Device: If the user falls and becomes motionless, it detects abnormal heart rate and unresponsive movement.
[1961] Terminal: Immediately sends emergency notifications and emotional data to the server.
[1962] Server: Receives emergency notification and sends a message to family contacts saying, "User has collapsed. Please respond quickly."
[1963] User (recipient): The family receives the emergency message and immediately begins taking action.
[1964] Thus, the system of the present invention provides elderly people living alone with an environment in which they can live with peace of mind by simultaneously promoting communication and enabling a rapid response in emergencies. In particular, the emotion engine enables natural dialogue that takes the user's emotions into consideration, thereby improving the quality of life.
[1965] The following describes the processing flow.
[1966] Speech recognition and input processing
[1967] Step 1:
[1968] Terminal: The terminal acquires the user's voice in real time through its built-in microphone.
[1969] Step 2:
[1970] Terminal: Transmits the acquired audio data to a speech recognition device (e.g., a speech recognition engine) and converts the audio data into text data.
[1971] Step 3:
[1972] Terminal: Sends the converted text data to the server using a transmission method. HTTP requests are often used for transmission.
[1973] Emotion recognition function
[1974] Step 4:
[1975] Terminal: Simultaneously, the emotion engine recognizes the user's emotions from their voice and facial expressions. The recognized emotion data is sent to the server along with text data.
[1976] Step 5:
[1977] Server: Analyzes received text data and sentiment data using analytical tools (such as a natural language processing engine).
[1978] Step 6:
[1979] Server: Based on the analysis results, it generates a response that takes the user's emotions into account. For example, if it is determined that the user is sad, the response will be encouraging.
[1980] Step 7:
[1981] Server: Sends the generated response to the terminal via a response transmission method. JSON format data is often used for transmission.
[1982] Step 8:
[1983] Terminal: Receives responses from the server and converts them into speech data using a text-to-speech engine.
[1984] Step 9:
[1985] Terminal: Outputs audio data to the user through the speaker.
[1986] Examples of conversation partner functions
[1987] User: "I'm feeling sad today."
[1988] Step 1: The device acquires the user's voice and converts it into text data, "I feel sad today," using speech recognition technology.
[1989] Step 2: The emotion engine recognizes sad emotions from the user's voice.
[1990] Step 3: Send text data and sentiment data to the server.
[1991] Step 4: The server analyzes the received data and generates an encouraging response such as, "What's wrong? Tell me what happened."
[1992] Step 5: The server sends the generated response to the terminal.
[1993] Step 6: The device converts the response data into speech and outputs to the user, "What's the matter? Please tell me what's wrong."
[1994] Emergency contact function
[1995] Step 1:
[1996] Device: The device continuously acquires biometric data such as heart rate, body temperature, and movement through built-in biosensors.
[1997] Step 2:
[1998] Terminal: Uses an anomaly detection method to monitor acquired biometric data in real time and detect abnormal values.
[1999] Step 3:
[2000] Terminal: If an anomaly is detected, an emergency notification will be generated. This notification will include the type of anomaly and the user's identification information.
[2001] Step 4:
[2002] Terminal: Sends the generated emergency notification to the server using the emergency notification method.
[2003] Step 5:
[2004] Server: Receives emergency notifications and sentiment data, and retrieves a pre-configured list of emergency contacts.
[2005] Step 6:
[2006] Server: Sends an emergency message to emergency contacts based on the information included in the emergency notification. Email, SMS, and phone APIs are often used for sending messages.
[2007] Step 7:
[2008] User (recipient): The family or medical institution will review the received emergency message and begin taking immediate action.
[2009] Examples of emergency contact functions
[2010] Step 1: The device detects the user's abnormal heart rate and cessation of movement.
[2011] Step 2: Send anomaly and emotion data to the server.
[2012] Step 3: The server receives an emergency notification and sends an emergency message to the emergency contact stating, "The user has collapsed. Please respond immediately."
[2013] Step 4: The family receives the emergency message and immediately begins taking action.
[2014] The above outlines the program's processing flow, including the specific actions performed at each processing step. This detailed explanation will allow for a clearer understanding of the overall system's operation.
[2015] (Example 2)
[2016] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[2017] When elderly people live alone, there are risks associated with a lack of communication and delayed responses in emergencies. Furthermore, systems that cannot engage in natural conversations that take emotions into consideration can increase the user's mental burden and potentially lower their quality of life. Additionally, if biometric data monitoring and emergency notifications are not effectively implemented, prompt responses become difficult.
[2018] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: speech recognition means for acquiring voice from a terminal in real time and converting the voice into text; emotion recognition means for analyzing emotions from the terminal and generating emotion data; transmission means for the terminal to send the converted text data and emotion data to the server; analysis means for the server to analyze the received text data and emotion data and generate an appropriate response; response transmission means for the server to send the generated response to the terminal; and speech synthesis means for converting the response received by the terminal into voice and outputting it to the user. This enables natural and effective communication with the elderly, allowing for a quick response in emergencies and improving the quality of life for users.
[2019] A "terminal" is a computer device used by users for voice input and output, as well as the collection of biometric data.
[2020] "Voice recognition means" refers to technology or devices that acquire a user's voice in real time on a terminal and convert that voice into text data.
[2021] "Emotion recognition means" refers to technologies and devices that analyze a user's emotions from their voice, facial expressions, etc., and generate emotion data.
[2022] "Transmission means" refers to communication technologies and devices used to send text data and sentiment data from a terminal to a server.
[2023] A "server" is a centralized management system that receives data sent from terminals, analyzes it, and generates appropriate responses.
[2024] "Analysis means" refers to technologies and devices that analyze text data and sentiment data on a server and generate appropriate responses for the user.
[2025] "Response transmission means" refers to communication technology or equipment used to send a response generated by a server to a terminal.
[2026] "Speech synthesis means" refers to a technology or device that converts response data received from a server into speech data at a terminal and outputs it to the user as speech.
[2027] "Methods for acquiring biometric data" refer to technologies and devices that continuously acquire biometric data such as a user's heart rate, body temperature, and movement.
[2028] An "anomaly detection method" refers to a technology or device that monitors acquired biological data in real time and detects anomalies.
[2029] An "emergency notification system" is a technology or device that sends an emergency notification to a server when an anomaly is detected.
[2030] "Communication methods" refer to technologies and devices that, after a server receives an emergency notification, send an emergency message to pre-configured emergency contacts.
[2031] This invention relates to a system that monitors elderly people living alone, facilitates communication, and enables rapid response in emergencies. In particular, this invention achieves more natural and effective communication by combining it with an emotion engine that recognizes the user's emotions.
[2032] Hardware and software to be used
[2033] terminal
[2034] The device includes the following hardware and software:
[2035] High-sensitivity microphone (voice input)
[2036] Speaker (audio output)
[2037] Biosensors (acquisition of heart rate, body temperature, and movement data)
[2038] Speech recognition engine (e.g., Google Speech-to-Text API)
[2039] Emotion recognition engine (e.g., Microsoft Azure Emotion API)
[2040] Speech synthesis engine (e.g., Amazon Polly)
[2041] server
[2042] The server includes the following software and features:
[2043] Analysis engine (Python library using NLP technology)
[2044] Response generation engine
[2045] Emergency notification system
[2046] Data receiving and transmitting module
[2047] System operation
[2048] terminal
[2049] When a user speaks, the device acquires the voice through its built-in microphone. The acquired voice data is converted into text data using a speech recognition engine. Simultaneously, an emotion recognition engine analyzes emotions from voice and facial expressions, generating emotion data. This data is transmitted to the server via the device's transmission mechanism.
[2050] server
[2051] The server receives text data and sentiment data sent from the terminal. The received data is analyzed using an analysis engine. Based on the analysis results, the server generates an appropriate response. The generated response is sent to the terminal via a response transmission means.
[2052] terminal
[2053] The terminal converts responses received from the server into speech data using a speech synthesis engine and outputs it to the user through a speaker. This allows elderly people to interact with the system in a natural way.
[2054] Specific example
[2055] Conversation partner function
[2056] User: "I'm feeling sad today."
[2057] Device: It uses the built-in microphone to acquire voice data and converts it into text data, such as "I feel sad today," using a speech recognition engine. The emotion engine then analyzes the user's voice to determine that they are feeling "sad."
[2058] Terminal: Sends the converted text data and sentiment data to the server.
[2059] Server: Analyzes the received data and generates encouraging responses such as, "What's wrong? Please tell me what happened."
[2060] Server: Sends the generated response data to the terminal.
[2061] Terminal: Converts response data into speech and outputs to the user, "What's the matter? Please tell me what happened."
[2062] Emergency contact function
[2063] Device: If the user falls, the built-in biosensors detect abnormal heart rate or unresponsive movement.
[2064] Terminal: Immediately sends emergency notifications and emotional data to the server.
[2065] Server: Upon receiving an emergency notification, it sends the message "User has collapsed. Please respond immediately" to the pre-configured emergency contact.
[2066] Recipient user: The family receives the emergency message and immediately begins taking action.
[2067] Example of a prompt
[2068] Example of prompt text when inputting a "monitoring system for elderly people living alone" into the AI model:
[2069] "The device acquires the user's voice and converts it into text data using a speech recognition engine. Additionally, an emotion engine analyzes the user's emotions and sends this data to a server. The server analyzes the received data, generates an appropriate response, and sends it to the device, which then responds verbally. Furthermore, in emergencies, a biosensor detects an anomaly and sends an emergency notification through the server. Please explain this system."
[2070] Thus, the system of the present invention promotes natural dialogue with the elderly and enables a rapid response in emergencies, thereby improving the safety and quality of life of the elderly.
[2071] The flow of the specific processing in Example 2 will be explained using Figure 13.
[2072] Step 1:
[2073] The device acquires the user's voice. Using the built-in high-sensitivity microphone, it collects voice data when the user says, "I'm feeling sad today."
[2074] Input: User's voice
[2075] Output: Audio data
[2076] Specific operation: The device's operating system converts the audio signal from the microphone into digital data and temporarily stores it in memory.
[2077] Step 2:
[2078] The device uses a speech recognition engine to convert audio data into text data. For example, the speech recognition engine uses the Google Speech-to-Text API.
[2079] Input: Audio data
[2080] Output: Text data
[2081] Specific operation: The speech recognition engine analyzes the audio data, identifies phonemes and morphemes, and generates the text "I feel sad today." The generated text is stored in a separate buffer in memory.
[2082] Step 3:
[2083] The device uses an emotion recognition engine to generate emotion data from voice and text data. For example, it can utilize the Microsoft Azure Emotion API.
[2084] Input: Audio data, text data
[2085] Output: Sentiment data
[2086] Specific operation: The emotion recognition engine analyzes the emotion of "sadness" from the tone of voice and associated text, and generates this emotion data. The generated emotion data is stored in memory along with the text data.
[2087] Step 4:
[2088] The terminal sends the converted text data and sentiment data to the server. This is done using a network transmission module.
[2089] Input: Text data, sentiment data
[2090] Output: Transmitted data
[2091] Specific operation: The transmission module packets text and sentiment data and sends them to the server via the Wi-Fi module over the internet. The data is encrypted to ensure security.
[2092] Step 5:
[2093] The server receives text data and sentiment data sent from the terminal.
[2094] Input: Data to send
[2095] Output: Received data
[2096] Specific operation: The server's receiving module receives packets, decrypts the encrypted data, and obtains text data and sentiment data. This data is stored in a database for analysis.
[2097] Step 6:
[2098] The server analyzes text and sentiment data and generates an appropriate response using an analysis engine.
[2099] Input: Text data, sentiment data
[2100] Output: Response data
[2101] Specific operation: The analysis engine uses NLP technology to analyze text data and sentiment data, and generates the most appropriate response for the user's state, such as "What's wrong? Please tell me what happened."
[2102] Step 7:
[2103] The server sends the generated response data to the terminal. A response transmission method is used.
[2104] Input: Response data
[2105] Output: Transmitted data
[2106] Specific operation: The response data is packetized and sent to the terminal via the network transmission module over the internet. The data is transmitted encrypted.
[2107] Step 8:
[2108] The terminal converts the received response data into audio data and outputs it to the user through the speaker. A speech synthesis engine is used.
[2109] Input: Response data
[2110] Output: Audio data
[2111] Specific operation: The speech synthesis engine analyzes the response data and generates the voice message, "What's the matter? Please tell me what's wrong." The generated voice message is output to the user through the speaker.
[2112] Step 9:
[2113] The device monitors biometric data, using sensors such as heart rate sensors, body temperature sensors, and motion sensors.
[2114] Input: Biometric data
[2115] Output: Acquired data
[2116] Specific operation: The sensor continuously measures the user's heart rate, body temperature, and movement, and records this data in memory. The data is updated in real time, waiting for anomaly detection.
[2117] Step 10:
[2118] If the device detects an anomaly, it sends an emergency notification to the server. For example, it might detect a sudden increase in heart rate or a cessation of movement.
[2119] Input: Acquired data
[2120] Output: Emergency notification data
[2121] Specific operation: When the anomaly detection algorithm detects a sudden increase in heart rate, abnormal fluctuations in body temperature, or cessation of movement, it generates emergency notification data related to these events and sends it to the server.
[2122] Step 11:
[2123] The server receives an emergency notification and sends an emergency message to pre-configured emergency contacts.
[2124] Input: Emergency notification data
[2125] Output: Emergency message
[2126] Specific operation: The server analyzes the received emergency notification data and sends a message stating "User has collapsed. Please respond immediately" to a pre-configured list of emergency contacts (e.g., family or medical institutions). The message is sent via SMS or a dedicated application.
[2127] (Application Example 2)
[2128] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[2129] In recent years, the number of elderly people living alone has increased, making their safety and security a major concern. There is also a growing need for systems that alleviate their feelings of loneliness and provide rapid responses in emergencies. However, current systems lack sufficient emotion recognition and emergency notification capabilities, and cannot fully guarantee the psychological and physical safety of the elderly. Therefore, there is a need for natural communication methods that include emotion recognition, as well as rapid emergency response through real-time monitoring of biometric data.
[2130] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a voice recognition means, a transmission means, an analysis means, a response transmission means, a voice synthesis means, and a means for recognizing emotions from the user's voice and facial expressions, equipped with an emotion recognition engine. This enables natural dialogue that takes into account the emotions of the elderly, and also enables a rapid response in emergencies. It also has an emergency notification function and can provide a system that continuously monitors heart rate, body temperature, and movement data, and immediately makes an emergency contact when an abnormality is detected.
[2131] A "terminal" is a device used by a user that has functions such as voice input, speech synthesis, emotion recognition, and acquisition of biometric data.
[2132] "Voice recognition means" refers to a function that converts the voice acquired by the terminal into text data.
[2133] "Transmission method" refers to the function by which the terminal sends the converted text data to the server.
[2134] "Analysis means" refers to a function that analyzes text data and sentiment data received by the server and generates an appropriate response.
[2135] "Response transmission means" refers to the function that sends the response generated by the server to the terminal.
[2136] A "speech synthesis means" is a function that converts the response received by the terminal into speech and outputs it to the user.
[2137] An "emotion recognition engine" is a function that recognizes emotions from the user's voice and facial expressions.
[2138] "Means of acquiring biometric data" refers to a function in which the terminal continuously acquires the user's heart rate, body temperature, and movement.
[2139] An "anomaly detection method" is a function that monitors acquired biometric data in real time and detects sudden increases in heart rate, abnormal fluctuations in body temperature, and cessation of movement.
[2140] An "emergency notification mechanism" is a function that sends an emergency notification to the server when an anomaly is detected.
[2141] "Communication method" refers to the function where the server receives an emergency notification and sends an emergency message to a pre-configured emergency contact.
[2142] This invention is a system that monitors elderly people living alone, facilitates communication, and enables rapid response in emergencies. The system mainly consists of terminals and a server.
[2143] Device functions
[2144] The terminal is equipped with a speech recognition system, a speech synthesis system, an emotion recognition engine, and a biometric data acquisition system. When a user speaks, the voice is acquired through the terminal's microphone and converted into text data by the speech recognition system. This text data is then sent to the server via a transmission system.
[2145] The device also features an emotion recognition engine that recognizes emotions from the user's voice and facial expressions. This emotion data is also sent to the server along with the text data.
[2146] Furthermore, the device is equipped with biometric data acquisition capabilities, continuously monitoring the user's heart rate, body temperature, and movement. If an abnormality is detected, a notification is sent to the server via an emergency notification system.
[2147] Server Functions
[2148] The server has an analysis mechanism to analyze the received text data and sentiment data, and generates an appropriate response. The generated response is sent back to the terminal via a response transmission mechanism. The terminal converts this response into speech using a speech synthesis mechanism and outputs it to the user through a speaker.
[2149] The server has a communication system that, upon receiving an emergency notification, sends an emergency message to pre-configured emergency contacts.
[2150] Hardware and software details
[2151] The system uses the "Speech Recognition API" for speech recognition, the "Network Communication Module" for text data transmission, and the "Emotion Recognition API" for emotion recognition. Additionally, it uses the "Speech Synthesis API" for speech synthesis and the "Biometric Sensor Module" for remote monitoring. The "Notification API" is used for emergency notifications and communication.
[2152] Specific example
[2153] When a user says, "I'm not feeling well today," the device converts the audio data into text data and analyzes it using an emotion recognition engine. The resulting data is sent to a server, which considers the user's emotions and generates an encouraging response such as, "What's wrong? Please tell me what's the matter." The generated response is sent to the device and output to the user as audio.
[2154] Furthermore, if a user collapses, the biosensors built into the device detect abnormal heart rate or unresponsive movement and immediately send an emergency notification to the server. The server receives the emergency notification and sends a message to emergency contacts stating, "A user has collapsed. Please respond quickly."
[2155] Example of a prompt
[2156] "Provide an example where, if a user says 'I'm not feeling well today,' the audio is converted to text, analyzed by an emotion engine, and an appropriate comforting response is generated."
[2157] As described above, this system supports the safe and secure lives of the elderly and can provide prompt assistance when needed. Furthermore, its emotion recognition function enables more natural and effective communication.
[2158] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[2159] Step 1:
[2160] The user speaks to the device.
[2161] Input: User's voice data.
[2162] Specific action: The device's microphone picks up the user's voice.
[2163] Step 2:
[2164] The device converts the audio into text data.
[2165] Input: Acquired audio data.
[2166] Specific operation: Use a speech recognition API to convert speech data into text data.
[2167] Output: Text data.
[2168] Step 3:
[2169] The device performs emotion recognition.
[2170] Input: User's voice data and text data.
[2171] Specific operation: Use an emotion recognition API to analyze the user's emotions from their voice and facial expressions.
[2172] Output: Sentiment data.
[2173] Step 4:
[2174] The device sends text data and sentiment data to the server.
[2175] Input: Text data and sentiment data.
[2176] Specific operation: Use the network communication module to send data to the server.
[2177] Output: Confirmation of data transmission to the server.
[2178] Step 5:
[2179] The server analyzes the text data and sentiment data it receives.
[2180] Input: Text data and sentiment data.
[2181] Specific operation: Use analytical tools to perform analysis in order to generate an appropriate response.
[2182] Output: Appropriate response data.
[2183] Step 6:
[2184] The server sends the appropriate response data to the terminal.
[2185] Input: Generated response data.
[2186] Specific operation: Use the response transmission means to send response data to the terminal.
[2187] Output: Confirmation of data transmission to the terminal.
[2188] Step 7:
[2189] The terminal converts the received response data into speech.
[2190] Input: Appropriate response data.
[2191] Specific operation: Use a speech synthesis API to convert text data into speech data.
[2192] Output: Audio data.
[2193] Step 8:
[2194] The device outputs audio data to the user.
[2195] Input: Audio data.
[2196] Specific action: Play audio through the device's speaker.
[2197] Output: Audio output to the user.
[2198] Step 9:
[2199] The device acquires biometric data.
[2200] Input: Biometric data such as the user's heart rate, body temperature, and movement.
[2201] Specific operation: Use a biosensor module to continuously collect data.
[2202] Output: Biometric data.
[2203] Step 10:
[2204] The device monitors biometric data acquired in real time.
[2205] Input: Biometric data.
[2206] Specific operation: Use anomaly detection methods to monitor data for abnormalities.
[2207] Output: Abnormal status data.
[2208] Step 11:
[2209] If an anomaly is detected, the device will send an emergency notification to the server.
[2210] Input: Abnormal situation data.
[2211] Specific action: Use the emergency notification method to send an anomaly notification to the server.
[2212] Output: Sending an emergency notification to the server.
[2213] Step 12:
[2214] The server receives an emergency notification and sends a message to the emergency contact.
[2215] Input: Anomaly notification data.
[2216] Specific action: Use the communication method to send a message to a pre-configured emergency contact.
[2217] Output: Confirmation of emergency message transmission.
[2218] The above describes the processing steps of the system based on this invention.
[2219] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[2220] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2221] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[2222] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[2223] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[2224] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[2225] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[2226] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[2227] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[2228] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[2229] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[2230] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[2231] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[2232] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[2233] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[2234] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[2235] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[2236] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[2237] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[2238] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[2239] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[2240] The following is further disclosed regarding the embodiments described above.
[2241] (Claim 1)
[2242] A speech recognition means that acquires audio in real time and converts the audio into text,
[2243] The terminal includes a transmission means for sending the converted text data to a server,
[2244] The server includes an analysis means that analyzes the text data it receives and generates an appropriate response,
[2245] A response transmission means that transmits the response generated by the server to the terminal,
[2246] A speech synthesis means that converts the response received by the terminal into speech and outputs it to the user,
[2247] A system that includes this.
[2248] (Claim 2)
[2249] A means for acquiring biometric data in which a terminal continuously acquires the user's biometric data,
[2250] An anomaly detection means for monitoring the biological data in real time and detecting abnormalities,
[2251] An emergency notification means that sends an emergency notification to the server when an anomaly is detected,
[2252] The server receives an emergency notification and sends an emergency message to a pre-configured emergency contact;
[2253] The system according to claim 1, including the following:
[2254] (Claim 3)
[2255] The means for acquiring biometric data from the device is a means for acquiring heart rate, body temperature, and movement data.
[2256] The abnormality detection means is a means for detecting a sudden increase in heart rate, abnormal fluctuations in body temperature, and cessation of movement.
[2257] The system according to claim 2.
[2258] "Example 1"
[2259] (Claim 1)
[2260] A speech recognition means that acquires audio in real time and converts the audio into text,
[2261] The terminal includes a transmission means for sending the converted text data to a server,
[2262] The server analyzes the text data it receives and generates an appropriate response using natural language processing means.
[2263] A response transmission means that transmits the response generated by the server to the terminal,
[2264] A speech synthesis means that converts the response received by the terminal into speech and outputs it to the user,
[2265] A system that includes this.
[2266] (Claim 2)
[2267] A means for acquiring biometric data in which a terminal continuously acquires the user's biometric data,
[2268] An anomaly detection means for monitoring the biological data in real time and detecting abnormaliti...
Claims
1. A speech recognition means that acquires audio in real time and converts the audio into text, The terminal includes a transmission means for sending the converted text data to a server, The server includes an analysis means that analyzes the text data it receives and generates an appropriate response, A response transmission means that transmits the response generated by the server to the terminal, A speech synthesis means that converts the response received by the terminal into speech and outputs it to the user, A system that includes this.
2. A means for acquiring biometric data in which a terminal continuously acquires the user's biometric data, An anomaly detection means for monitoring the biological data in real time and detecting abnormalities, An emergency notification means that sends an emergency notification to the server when an anomaly is detected, The server receives an emergency notification and sends an emergency message to a pre-configured emergency contact; The system according to claim 1, including the following:
3. The means for acquiring biometric data from the device is a means for acquiring heart rate, body temperature, and movement data. The abnormality detection means is a means for detecting a sudden increase in heart rate, abnormal fluctuations in body temperature, and cessation of movement. The system according to claim 2.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A